The application provides a cache-restricted LT code
degree distribution optimization method based on
reinforcement learning, comprising the following steps: S1, generating a source symbol sequence with a length of k; S2, constructing a
reinforcement learning agent for determining a
degree distribution of a transmitting end, the
reinforcement learning agent is composed of a 2-node input layer, a 2k-node
hidden layer and a k-node output layer; S3, updating a current degree value through the reinforcement
learning agent, and performing LT coding on the degree value and the source symbol by an
encoder to obtain a coded symbol and send the coded symbol to a receiving end decoder; S4, judging the degree value of the received coded symbol according to the currently recovered source symbol by the decoder, and performing a cache or decoding according to the degree value; S5, training the reinforcement
learning agent by using the reinforcement learning environment in step S3, and finally obtaining a
degree distribution that makes the reward tend to be maximum through iterative training. The application can obtain a degree distribution with smaller redundancy overhead required for
full recovery under the condition of cache restriction.