Transaction Details
SuccessPayment
Details
Type
Payment
Result
tesSUCCESS
Timestamp
5 months ago (2/21/2026, 3:25:11 PM)
Fee
0.00001 XRP
Position
#8 in ledger
Amount
0.000001 XRP
Memo
t is feasible to initially assign random parameters to new nodes then use Reinforcement Learning to update the parameters of the proliferated neural network. Using MCTS to search on the current state, adopting a self-play mode where actions are chosen based on the Transformer evaluator, and under the premise of limited search depth, continuously backpropagate and optimize the parameters of the Transformer and embedding matrix using the improved policy distribution and return signals obtained from the MCTS search; when the reward (e.g., significant compression) exceeds a threshold, the important reasoning chain or structure used is treated as a new object.Like AlphaZero's algorithm, here the Transformer also simultaneously acts as the evaluator of the current reasoning state. The difference lies in that we might also use the Transformer to generate word vectors for candidate new words. The system calculates the distance between the outpu
