Local repair code construction method and device based on reinforcement learning and storage medium
By automatically optimizing the structure of local repair codes based on reinforcement learning, the complexity and computational cost of traditional LRC construction methods are solved, and more efficient and flexible code construction is achieved, which improves repair performance and storage efficiency.
Patent Information
- Application Number
- CN202510285512.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In practical applications, traditional local repair code (LRC) construction methods face problems such as complex construction process, large calculations, and difficulty in balancing storage efficiency and repair capabilities.
The local repair code is constructed based on reinforcement learning, and the reinforcement learning model is automatically optimized and the code design is dynamically adjusted to adapt to different application environments.
The repair performance, storage efficiency and bandwidth utilization of local repair codes are improved, the computational complexity and manual intervention in traditional methods are reduced, and a more efficient and flexible construction solution is achieved.
Smart Images

Figure CN120223099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of error - correcting codes and artificial intelligence, and particularly relates to a method, device, and storage medium for constructing locally repairable codes based on reinforcement learning. Background Art
[0002] Error - correcting codes are crucial in modern communication and information storage systems, mainly ensuring the recovery of information under noise interference through redundant information. Traditional error - correcting codes, such as Huffman coding, Reed - Solomon codes, etc., are widely used in fields such as wireless communication and optical fiber communication, and can effectively improve the reliability of data transmission. In recent years, locally repairable codes (LRCs), as a new type of error - correcting code, have received extensive attention because they can repair only the damaged data part, improving the repair efficiency and storage utilization. Although LRCs have good theoretical properties, their traditional construction methods (such as algebraic construction methods, graph - theoretic construction methods, combinatorial design methods) still face challenges in practical applications, mainly manifested in complex construction processes, large computational amounts, and difficulties in balancing storage efficiency and repair capabilities.
[0003] To solve these problems, artificial intelligence technology, especially reinforcement learning, has begun to be applied to the construction of error - correcting codes. Different from traditional methods that rely on manual derivation, reinforcement learning can automatically adjust the code structure according to different application requirements and environmental conditions through adaptive optimization, thereby improving the construction efficiency, reducing the computational complexity, and flexibly coping with dynamic environments. This emerging technology has shown significant advantages in the construction of locally repairable codes and can achieve more efficient and flexible construction schemes. Summary of the Invention
[0004] A method, device, and storage medium for constructing locally repairable codes based on reinforcement learning proposed by the present invention aim to solve at least the problems of insufficient flexibility and intelligence in the construction of LRCs in related technologies.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for constructing locally repairable codes based on reinforcement learning, the specific steps include: Step S01: Input the code length n and dimension k to determine the size of the locally repairable code to be constructed. The generator matrix of the locally repairable code is obtained by splicing a k×k - dimensional identity matrix and a k×(n - k) - dimensional first matrix, where k is an integer greater than 0, and n is an integer greater than k; Among them, in the finite field the linear code C is a subspace in the linear space where q is the order of the finite field; if the dimension of this linear subspace is k, then the linear code C is called an [n, k] code, where n is the code length and k is the dimension of the code; Step S02: Determine the number of neural network nodes in the input layer and the number of neural network nodes in the output layer of the reinforcement learning model according to the code length n and the dimension k, to obtain a target reinforcement learning model, where the number of neural network nodes in the input layer is k×n, and the number of neural network nodes in the output layer is k×(n - k); Step S03: Concatenate the identity matrix and the initialized first matrix to obtain an initialized state matrix; and input the initialized state matrix into the target reinforcement learning model; Step S04: According to the target first matrix output by the target reinforcement learning model, concatenate the identity matrix and the target first matrix to obtain the current generating matrix; Step S05: Pass the generating matrix obtained in Step S04 to the evaluation module; Step S06: Steps S04 and S05 will be repeatedly executed until the set maximum number of iterations is reached; in each round of iteration, the reinforcement learning algorithm will continuously optimize the construction process of the local repair code, so that the generated local repair code is gradually improved and has better local repair performance and storage efficiency.
[0006] Further, in Step S02, it also includes initializing the prediction network and the policy network, setting the learning rate of the reinforcement learning model and binding the optimizer. The learning rate of the policy network is 1e - 4, the learning rate of the prediction network is 1e - 3, and the optimizers of both the policy network and the prediction network are Adam optimizers.
[0007] On the other hand, the present invention also discloses a computer - readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the above - mentioned method.
[0008] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which when executed by the processor, causes the processor to execute the steps of the above - mentioned method.
[0009] As can be seen from the above technical solutions, the present invention proposes a construction method for locally repairable codes (LRCs) based on reinforcement learning, aiming to solve problems such as low efficiency, poor adaptability, and unstable repair performance in existing LRC construction methods. Existing methods usually rely on manual derivation and adjustment, lacking the ability of automated and adaptive optimization, and unable to meet the increasingly complex requirements in modern communication and storage systems. By introducing reinforcement learning, the present invention can automatically optimize the structure of LRCs, dynamically adjust the code design to adapt to different application environments, significantly improving the repair performance, storage efficiency, and bandwidth utilization. At the same time, the reinforcement learning algorithm reduces the cumbersome calculation and tuning processes in traditional methods by continuously optimizing the construction strategy, improving the construction efficiency and reducing the computational complexity. In summary, the present invention provides an efficient, flexible, and adaptive LRC construction method with stronger fault tolerance and robustness, which can be widely applied to various modern communication and storage systems.
[0010] Generally speaking, the present invention is an innovative solution proposed precisely to solve several problems in traditional LRC construction methods. First of all, traditional methods often rely on complex mathematical derivations and are difficult to quickly adapt to changes in requirements, while the present invention can automatically adjust the code structure, reduce manual intervention, and enhance flexibility. Secondly, in the face of large-scale data or complex networks, the computational complexity of traditional methods is relatively high, while the present invention reduces the computational complexity through an intelligent optimization mechanism and improves the repair performance of the system at the same time. Finally, the method based on the present invention can flexibly respond in different communication and storage environments, maximizing the optimization of storage efficiency, repair ability, and bandwidth utilization, and has stronger application potential.
[0011] Therefore, the present invention can not only solve the inherent problems in traditional methods but also provide a more efficient, flexible, and adaptive construction solution. With the continuous development of artificial intelligence technology, this innovative method is expected to become an important direction for LRC design and play a more important role in various communication and storage systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of the LRC construction method based on reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0014] The method for constructing locally repairable codes based on reinforcement learning described in this embodiment has a flowchart as Figure 1 shown, and includes the steps: Input parameters: relevant parameters such as code length, dimension, maximum number of iterations, and learning rate of the reinforcement learning algorithm. According to the input parameters, control the output of the algorithm, the size of the model, the number of runs, and the learning performance of the model to ensure the effectiveness and efficiency of the construction process.
[0015] According to the two parameters of code length and dimension, generate a construction module, which is mainly composed of a reinforcement learning algorithm and a neural network model; at the same time, generate an evaluation module, which is used to evaluate the performance of the constructed code and feedback relevant information.
[0016] In the continuous construction-evaluation process, the reinforcement learning algorithm continuously adjusts its parameters through the neural network model. The evaluation module optimizes the reinforcement learning strategy according to the feedback information of the model, so that the performance of the constructed local repair code is continuously improved until the predetermined convergence condition is reached.
[0017] Take the initialized state matrix as the state matrix to be updated, and input the state matrix to be updated and the reward to be updated into the agent module; Update the parameters in the agent according to the state matrix to be updated and the reward to be updated; Determine the first matrix to be updated according to the action matrix to be updated output by the agent module after updating the parameters and the order of the finite field to which the elements in the local repair code belong; After iterating a preset number of times, determine the first matrix to be updated as the target first matrix.
[0018] Among them, the first input of the agent module is only the initialized state matrix, without a reward, or the reward is 0. In each subsequent iteration, the state matrix and reward will be updated, and then the updated state matrix and reward will be used as the input of the agent module, and the agent will update its own parameters according to this input for iteration until after iterating a preset number of times, stop updating, and then determine the first matrix to be updated as the target first matrix.
[0019] Specifically, it includes the following steps Step S01: Input the code length n and dimension k to determine the size of the local repair code to be constructed. The generator matrix of the local repair code is composed of a k×k identity matrix and a k×(n - k) first matrix Matrices are spliced.
[0020] Among them, k is an integer greater than 0, and n is an integer greater than k. In the finite field The linear code (or linear block code) C on is a linear space a subspace, where q is the order of the finite field. In particular, when q = 2, it is called a binary linear code. If the dimension of this linear subspace is k, then C is called an [n, k] code, where n is the code length and k is the dimension of the code.
[0021] The generator matrix of the locally repairable code (LRC code) can be divided into two parts, namely, a k×k dimensional identity matrix and a k×(n - k) dimensional first matrix. For example, the identity matrix is denoted as I, the first matrix is denoted as P, and the generator matrix is denoted as G. Taking q = 2, n = 6, and k = 4 as an example: If , P = Then
[0022] It should be noted that the way of concatenating the above identity matrix I and the first matrix P left and right is just an example, and there can be other concatenation methods. For example , P = Then
[0023] Step S02: Determine the number of neural network nodes in the input layer and the number of neural network nodes in the output layer of the reinforcement learning model according to the code length n and the dimension k, to obtain the target reinforcement learning model. The number of neural network nodes in the input layer is k×n, and the number of neural network nodes in the output layer is k×(n - k).
[0024] Furthermore, the number of neural network nodes in the hidden layer of the reinforcement learning model can be determined according to experience, neither too large nor too small. For example, it can be 2×n×k, 3×n×k, 4×n×k, etc. At the same time, the method also includes initializing the prediction network and the policy network, setting the learning rate of the reinforcement learning model and binding the optimizer. The learning rate of the policy network is 1e - 4, the learning rate of the prediction network is 1e - 3, and the optimizers of both the policy network and the prediction network are Adam optimizers.
[0025] Step S03: Concatenate the identity matrix and the initialized first matrix to obtain the initialized state matrix; and input the initialized state matrix into the target reinforcement learning model.
[0026] After initialization, the elements of the first matrix can be set to preset values. For example, all elements are 0. Since the input of the reinforcement learning model is the state, the generator matrix of the LRC code can be used as the state vector and input into the reinforcement learning model.
[0027] Step S04: According to the target first matrix output by the target reinforcement learning model, concatenate the identity matrix and the target first matrix to obtain the current generator matrix; Concatenate the identity matrix and the target first matrix to obtain the generator matrix of the LRC code, and then generate the LRC code according to the linear combination or linear space of the generator matrix.
[0028] Compared with the existing coding theory, the LRC code constructed through the reinforcement learning model does not construct a new code based on the generator matrix of the existing code. Therefore, parameters such as the performance, code length, and dimension of the existing code will not affect the constructed LRC code. Secondly, when using the reinforcement learning model to generate the LRC code, the input code length and dimension can be arbitrary without mandatory specific conditions. Therefore, it is more flexible when constructing the LRC code. Finally, the model can be connected to the decoder to accurately evaluate the performance of the code under a certain channel condition and can achieve automatic computer operation.
[0029] If the order of the finite field to which the elements in the local repair code belong is 2, then for the elements in the to-be-updated action matrix output by the agent module, those with values greater than or equal to 0.5 are taken as 1, and those less than 0.5 are taken as 0, to obtain the to-be-updated first matrix. The judgment threshold of 0.5 is just an example and can also be other values, such as 0.49, 0.48, 0.51, 0.52, which can be determined according to the actual situation.
[0030] For example, when the order is 2, the code length n = 6, the dimension k = 4, and the action matrix is a k*(n - k) matrix. The to-be-updated action matrix is as follows:
[0031] Take those with values greater than or equal to 0.5 as 1 and those less than 0.5 as 0, to obtain the to-be-updated first matrix as follows:
[0032] Concatenate the to-be-updated first matrix and the identity matrix to obtain the to-be-updated state matrix.
[0033] Step S05: Transmit the matrix with updated state to the evaluation module. First, calculate the key performance indicators of the LRC code according to the generator matrix, including locality, availability, and minimum distance. These three indicators are the core criteria for evaluating the quality of the LRC code. The evaluation module will check whether the current state meets the above three criteria through corresponding calculations on the matrix. The specific evaluation process includes checking matrix operations and quantifying the repair ability and system redundancy by calculating indicators such as the minimum distance and locality.
[0034] Based on the evaluation results, the module calculates a reward, which is obtained according to the comprehensive performance of the local repair degree, availability, and minimum distance. If the current matrix meets the design requirements of the LRC code, that is, it has a high local repair degree, strong availability, and appropriate minimum distance, the reward is positive; conversely, if the performance of the matrix does not meet these criteria, the reward is zero or negative. The setting of the reward can be determined based on experience and should not be too complex to facilitate the learning of the agent. For example, when judging the updated first matrix, if the number of 1s in each row is equal to the corresponding local repair degree, the reward is +1, and if the number of 1s in each column is equal to the corresponding reliability, the reward is +1. The initial reward is equal to the minimum distance. The setting of the reward varies slightly according to the input code length and dimension, that is, it is determined according to the actual situation.
[0035] The goal of the reinforcement learning model is to maximize the reward. If the goal is not the LRC code, the reward is relatively low. This model can perform positive incentives and rapid iterations to find the target LRC code, improving the calculation speed.
[0036] Update the parameters in the prediction network according to the loss function. The loss function is equal to the sum of the squares of the differences between the corresponding elements in f(state) and f’(state), where the output of the target network is f(state) and the output of the prediction network is f’(state). Further, the number of input nodes of the target network and the prediction network is k×(2n - k), and the number of output nodes can be customized, for example, it is 1. The parameters of the target network are frozen, so there is no need to update them.
[0037] Step S06, step S04, and step S05 will be repeatedly executed until the set maximum number of iterations is reached; in each round of iteration, the reinforcement learning algorithm will continuously optimize the construction process of the LRC code, making the generated LRC code gradually improved and having better local repair performance and storage efficiency. Finally, when the maximum number of iterations reaches the predetermined value, the algorithm terminates, the model converges, outputs the optimal LRC code construction scheme, and gives the final local repair code generation matrix.
[0038] The partial generation matrix constructed is as follows:
[0039] Taking the above generation matrix as an example, the first information symbol [1, 0, 0, 0, 0] Τ (column vector) is missing due to noise or other factor interferences. According to the generation matrix, it has a reliability of t = 4 repair sets {2, 6}, {4, 10}, {3, 13}, {5, 15}. We note that there are r = 2 other code symbols with locality in each repair set, that is, for [1, 0, 0, 0, 0] ΤFor this missing information symbol, we can repair it with the other two symbols. Taking {2,6} as an example, the second and sixth symbols are [0,1,0,0,0] Τ and [1,1,0,0,0] Τ In the binary field (only 0 and 1, 1 + 1 = 0), [0,1,0,0,0] Τ + [1,1,0,0,0] Τ = [1,0,0,0,0] Τ In this way, we have restored the missing information [1,0,0,0,0] using the repair set {2,6} Τ . The same applies to other repair sets and other information symbols.
[0040]
[0041] The LRC parameters generated by this method are shown in Table 1: Table 1 Generated LRC experimental data
[0042] As described in the above generating matrix, the code length n represents the number of columns of the generating matrix, the dimension k represents the number of rows of the generating matrix, the reliability Availability(t) is the number of repair sets for a single information symbol, the locality Locality(r) is the number of other symbols required to repair a single information symbol. Refer to the example of the generating matrix above. The minimum distance d is the minimum code weight in the space generated by the current generating matrix. Taking [1,0,0] as an example, the weight is 1 (the number of 1s in the vector). The above table shows the parameters of the generating matrix obtained through the reinforcement learning model, which can refer to the partial generating matrix constructed above.
[0043] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0044] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of the above method.
[0045] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer causes the computer to execute any of the above-described reinforcement learning-based local repair code construction methods.
[0046] It should be understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention. For the explanations, examples, and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods.
[0047] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0048] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0049] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding description in the method embodiment.
[0050] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A local repair code construction method based on reinforcement learning, characterized in that: The following steps are included: Step S01, input code length n and dimension k to determine the size of the local repair code to be constructed, and the generator matrix of the local repair code is obtained by concatenating a k×k-dimensional unit matrix and a k×(nk)-dimensional first matrix, where k is an integer greater than 0 and n is an integer greater than k; Among them, in the finite field The linear code C on the linear space The subspace in , where q is the order of the finite field; if the dimension of this linear subspace is k, then the linear code C is called an [n, k] code, where n is the code length and k is the dimension of the code; Step S02, determining the number of input layer neural network nodes and the number of output layer neural network nodes of the reinforcement learning model according to the code length n and the dimension k, and obtaining a target reinforcement learning model, wherein the number of input layer neural network nodes is k×n, and the number of output layer neural network nodes is k×(nk); Step S03, concatenating the unit matrix and the initialized first matrix to obtain an initialized state matrix; and inputting the initialized state matrix into the target reinforcement learning model; Step S04: according to the target first matrix output by the target reinforcement learning model, concatenate the unit matrix and the target first matrix to obtain the current generation matrix; Step S05, passing the generation matrix obtained in step S04 to the evaluation module; Step S06, step S04 and step S05 will be repeatedly executed until the set maximum number of iterations is reached; in each round of iteration, the reinforcement learning algorithm will continuously optimize the construction process of the local repair code, so that the generated local repair code is gradually improved and has better local repair performance and storage efficiency.
2. The local repair code construction method based on reinforcement learning according to claim 1 is characterized in that: Step S02 also includes initializing the prediction network and the policy network, setting the learning rate and binding optimizer of the reinforcement learning model, the learning rate of the policy network is 1e-4, the learning rate of the prediction network is 1e-3, and the optimizers of the policy network and the prediction network are both Adam optimizers.
3. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to claim 1 or 2.
4. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to claim 1 or 2.