Software architecture evolution path planning system and method based on reinforcement learning
Through the path planning method based on reinforcement learning, the problem of inconvenience in moving in complex environments in traditional software architecture evolution path planning systems is solved, and efficient and adaptive path planning in complex environments is achieved.
Patent Information
- Application Number
- CN202510432023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional software architecture evolution path planning systems are difficult to guide software movement safely and effectively in complex dynamic environments, and there are limitations in large-scale scenarios.
The path planning method based on reinforcement learning is adopted, and the path planning is carried out by simulated annealing evolution path map. The reinforcement learning module interacts with the environment to learn the optimal strategy to generate actual paths and analyze and compare modules to output the optimal path.
Improves the robustness and accuracy of path planning, and adaptively adjusts strategies in complex environments without the need to build an environment map in advance, ensuring that the software continues to learn and improve during the task execution.
Smart Images

Figure CN120355082A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software planning, and particularly to a software architecture evolution path planning system and method based on reinforcement learning. Background Art
[0002] In today's rapidly developing information technology era, software systems have become the infrastructure for the operation of modern society. However, with the expansion of software scale and the increase in complexity, how to effectively plan the evolution of software has become an important challenge.
[0003] As the skeleton of a software system, the software architecture supports the entire software system and provides important guarantees for it to have many good characteristics. Through the software architecture, the overall complexity and variability of the software system can be effectively controlled. The cost of software detection and modification based on the software architecture is relatively low. The evolution of the software architecture can better ensure the consistency and correctness of software evolution, and significantly reduce the cost of software evolution, making the evolution of the software system more convenient. The software architecture is the basis of software design and can be regarded as the blueprint of the entire system, which helps to guide the development process and achieve goals such as system reliability and maintainability. The software architecture provides a common high-level abstraction, enabling most relevant personnel of the system to communicate and collaborate based on this abstraction, forming a unified understanding and mutual understanding, so as to rationally manage large and complex systems. The software architecture clarifies the constraints on system implementation, determines the organizational structure of the development and maintenance organizations, and restricts the quality attributes of the system. By studying the software architecture, the quality of software can be predicted, and reasoning and controlling changes can be made simpler.
[0004] However, the traditional software architecture evolution path planning system has the following disadvantages:
[0005] The traditional software architecture evolution path planning system safely and effectively guides the movement of software in a complex environment. Traditional path planning methods usually rely on specific rules or algorithms, but these methods may have limitations when dealing with complex dynamic environments or large-scale scenarios;
[0006] The present invention uses reinforcement learning to perform path planning on the simulated annealing evolution path graph. Reinforcement learning can effectively handle environments with obstacles and dynamic targets because it does not need to pre-construct an environmental map, but directly obtains the optimal strategy through learning and adaptively adjusts the strategy according to changes in the environment, thereby improving the robustness of path planning. The software continuously learns and improves the strategy during the process of executing tasks without the need for a large amount of offline training in advance, and has high accuracy and security. Summary of the Invention
[0007] The purpose of the present invention is to provide a software architecture evolution path planning system and method based on reinforcement learning, so as to solve the problem that the traditional software architecture evolution path planning system guides the software to move safely and effectively in a complex environment. The traditional path planning methods are usually based on specific rules or algorithms, but these methods may have limitations when dealing with complex dynamic environments or large-scale scenarios.
[0008] To achieve the above purpose, the present invention provides the following technical solutions: A software architecture evolution path planning system based on reinforcement learning, including a planning system, the planning system includes a software evolution module, a path graph module, a path planning module, a reinforcement learning module, and an analysis and comparison module;
[0009] The software evolution module evolves the software to adapt to the new requirements of users, changes in the business environment and operating environment. The path graph module constructs a software architecture evolution graph and a software architecture evolution path graph respectively. The path planning module maps the information in the simulated annealing evolution path graph to various factors in reinforcement learning to complete software path planning. The reinforcement learning module model learns the optimal strategy by interacting with the environment, and then the model uses the learned strategy to generate the actual path. The analysis and comparison module analyzes the software evolution path and outputs the optimal software evolution path.
[0010] As a preferred technical solution of the present invention, the software evolution module includes an object evolution sub-module, a message evolution sub-module, a composite fragment evolution sub-module, a constraint evolution sub-module, a static evolution sub-module, and a dynamic evolution sub-module;
[0011] The object evolution sub-module evolves numerous attributes, such as interfaces, types, semantics. The message evolution sub-module evolves the name, source object, target object, and time sequence. The composite fragment evolution sub-module evolves together with the message and belongs to the category of connectors. The constraint evolution sub-module adds and deletes constraint information. The static evolution sub-module consults documents, analyzes the architecture, identifies the system composition and the mutual relationship between elements, extracts the abstract representation form of the system, tests the evolved system, and finds the errors and deficiencies therein. The dynamic evolution sub-module during the operation process, changes in user requirements or adjustments to the system's own service quality will trigger changes in software behavior.
[0012] As a preferred technical solution of the present invention, the path graph module includes a software evolution graph sub-module and a software architecture graph sub-module;
[0013] The software evolution graph sub-module constructs a software evolution graph through Github to know the relationship between the previous and current versions of the software system. The software architecture graph sub-module obtains the software architecture by reverse and converts the software evolution graph into a software architecture evolution graph.
[0014] As a preferred technical solution of the present invention, the path planning module includes a dimension attribute sub-module, an action sequence group sub-module, and an attribute calculation sub-module;
[0015] The dimension attribute sub-module uses attributes of five dimensions to measure the software architecture of each version, namely the number of atomic components, the size of atomic components, the size of the software architecture, the number of effective code lines, and the number of Java files. The action sequence sub-module uses the change log to record the software architecture change action sequence between adjacent versions. The attribute calculation sub-module uses the five-dimensional attributes of the software architecture to calculate the change rate and the initial reward value between adjacent versions.
[0016] As a preferred technical solution of the present invention, the reinforcement learning module includes a Q-leaning sub-module and a Monte Carlo tree search sub-module;
[0017] The Q-leaning sub-module understands the environment through long-term interactive learning and finds out how to maximize the reward function in various situations to form a planning strategy. The Monte Carlo tree search sub-module generates an actual path according to the learned strategy to ensure finding the optimal path efficiently in a complex environment.
[0018] As a preferred technical solution of the present invention, the analysis and comparison module includes a path analysis module, a path comparison module, and a result output module;
[0019] The path analysis module analyzes the software evolution path, the path comparison module compares other software evolution paths, and the result output module outputs the optimal software evolution path.
[0020] The usage method of the software architecture evolution path planning system based on reinforcement learning of the present invention includes the following steps:
[0021] Step 1. Software evolution: Construct a software evolution graph through Github to know the relationship between the previous and current versions of the software system. Secondly, obtain the software architecture by reverse engineering and convert the software evolution graph into a software architecture evolution graph;
[0022] Step 2. Path planning: Convert the software architecture evolution graph into a software architecture evolution path and use reinforcement learning to perform path planning on the simulated annealing evolution path graph;
[0023] Step 3. Result output: Perform comparative analysis according to the planning content and output the optimal software evolution path.
[0024] As a preferred technical solution of the present invention, the reinforcement learning in step two is specifically the Q-value update rule, Monte Carlo tree search, SARSA, and DQN. At each time step, Q-learning updates the Q-value according to the following update rule, and the calculation formula is as follows:
[0025] Q(s,a)←Q(s,a)+α·[r+γ·max α' Q(s',a')-Q(s,a)],
[0026] where Q(s,a) is the Q-value of taking action a in state s; α is the learning rate, which controls the step size of Q-value update; r is the immediate reward obtained after taking action a in state s; γ is the discount factor, which is used to weigh the importance of the current reward and future rewards; s' is the next state transferred after taking action a from state s; a' is the optimal action taken in the next state s'. The Monte Carlo tree search helps the algorithm decide which child node to explore next in the selection stage, and the calculation formula is as follows:
[0027]
[0028] where, represents the node, and S i is the average value size; c is a constant, usually taking the value of 2; N is the total number of explorations; n i is the number of explorations of node S i The SARSA calculation formula is as follows:
[0029] Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A)),
[0030] where R is the reward of the current state, which represents the immediate reward obtained after executing action A in state S; γ is the discount factor, which is used to weigh the importance of the current reward and future rewards; Q(S′,A′): the action value estimate of the next state; α is the learning rate, which determines the amplitude of each update. The DQN calculation formula is as follows:
[0031] Q(s,a)=r+γ*max_a'Q_target(s',a'),
[0032] where, Q _target (s', a') is the Q-value calculated by the target network, while Q(s,a) is the Q-value calculated by the online network.
[0033] As a preferred technical solution of the present invention, the specific process of the reinforcement learning in step two is as follows: in path planning, the state represents the position coordinates where the software is located, and the action represents the software moving in the up, down, left, and right directions. The Q-value is initialized to a small random value or zero:
[0034] S1. Initialize the state s to the starting position;
[0035] S2. Select an action a using the ε-greedy policy, i.e., with probability ε, select a random action, and with probability 1 - ε, select the action with the highest Q-value;
[0036] S3. Execute the action a and observe the immediate reward r and the next state s';
[0037] S4. Update the Q-value according to the Q-value update rule;
[0038] S5. Update the state to s';
[0039] S6. Repeat steps S2 - S5 until the target state is reached.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] Using reinforcement learning for path planning of the simulated annealing evolution path graph, reinforcement learning can effectively handle complex environments with obstacles and dynamic targets. Because it does not need to pre-construct an environmental map, but directly obtains the optimal policy through learning and adaptively adjusts the policy according to the changes in the environment, thereby improving the robustness of path planning. The software continuously learns and improves the policy during the process of executing tasks, without the need for a large amount of offline training in advance, and has high accuracy and security. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the architecture of the planning system of the present invention;
[0043] Figure 2 It is a schematic diagram of the architecture of the software evolution module of the present invention;
[0044] Figure 3 It is a schematic diagram of the architecture of the path graph module of the present invention;
[0045] Figure 4 It is a schematic diagram of the architecture of the path planning module of the present invention;
[0046] Figure 5 It is a schematic diagram of the architecture of the reinforcement learning module of the present invention;
[0047] Figure 6 It is a schematic diagram of the architecture of the analysis and comparison module of the present invention;
[0048] Figure 7 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] Please refer to Figure 1-7 , the present invention provides a software architecture evolution path planning system based on reinforcement learning, including a planning system. The planning system includes a software evolution module, a path graph module, a path planning module, a reinforcement learning module, and an analysis and comparison module;
[0051] The software evolution module evolves the software to adapt to the new requirements of users, the changes in the business environment and the operating environment. The path graph module constructs a software architecture evolution graph and a software architecture evolution path graph respectively. The path planning module maps the information in the simulated annealing evolution path graph to each factor in reinforcement learning to complete software path planning. The reinforcement learning module model learns the optimal strategy by interacting with the environment. Then, the model uses the learned strategy to generate the actual path. The analysis and comparison module analyzes the software evolution path and outputs the optimal software evolution path.
[0052] The software evolution module includes an object evolution sub-module, a message evolution sub-module, a composite fragment evolution sub-module, a constraint evolution sub-module, a static evolution sub-module, and a dynamic evolution sub-module;
[0053] The object evolution sub-module evolves numerous attributes, such as interfaces, types, semantics. The message evolution sub-module evolves names, source objects, target objects, and time sequences. The composite fragment evolution sub-module, like messages, belongs to the category of connectors and evolves. The constraint evolution sub-module adds and deletes constraint information. The static evolution sub-module consults documents, analyzes the architecture, identifies the system composition and the mutual relationships between elements, extracts the abstract representation form of the system, tests the evolved system, and finds the errors and deficiencies therein. The dynamic evolution sub-module during operation, changes in user requirements or adjustments to the system's own service quality will trigger changes in software behavior.
[0054] The path graph module includes a software evolution graph sub-module and a software architecture graph sub-module;
[0055] The software evolution graph sub-module constructs a software evolution graph through Github to know the relationship between the previous and current versions of the software system. The software architecture graph sub-module obtains the software architecture by reverse and converts the software evolution graph into a software architecture evolution graph.
[0056] The path planning module includes a dimension attribute sub-module, an action sequence group sub-module, and an attribute calculation sub-module;
[0057] The dimension attribute sub-module uses the attributes of five dimensions to measure the software architecture of each version, including the number of atomic components, the size of atomic components, the size of the software architecture, the number of effective code lines, and the number of Java files. The action sequence sub-module uses the change log to record the software architecture change action sequence between adjacent versions. The attribute calculation sub-module uses the five-dimensional attributes of the software architecture to calculate the change rate and the initial reward value between adjacent versions.
[0058] The reinforcement learning module includes a Q-learning sub-module and a Monte Carlo tree search sub-module;
[0059] The Q-learning sub-module understands the environment through long-term interactive learning and finds out how to maximize the reward function in various situations to form a planning strategy. The Monte Carlo tree search sub-module generates an actual path according to the learned strategy to ensure finding the optimal path efficiently in a complex environment.
[0060] The analysis and comparison module includes a path analysis module, a path comparison module, and a result output module;
[0061] The path analysis module analyzes the software evolution path, the path comparison module compares other software evolution paths, and the result output module outputs the optimal software evolution path.
[0062] The usage method of the software architecture evolution path planning system based on reinforcement learning of the present invention includes the following steps:
[0063] Step 1, software evolution: construct a software evolution graph through Github to know the relationship between the previous and current versions of the software system. Secondly, obtain the software architecture by reverse engineering and convert the software evolution graph into a software architecture evolution graph;
[0064] Step 2, path planning: convert the software architecture evolution graph into a software architecture evolution path and use reinforcement learning to perform path planning on the simulated annealing evolution path graph;
[0065] Step 3, result output: perform comparative analysis according to the planning content and output the optimal software evolution path.
[0066] The reinforcement learning in Step 2 is specifically the Q-value update rule, Monte Carlo tree search, SARSA, and DQN. At each time step, Q-learning updates the Q-value according to the following update rule, and the calculation formula is as follows:
[0067] Q(s,a)←Q(s,a)+α·[r+γ·max α' Q(s',a')-Q(s,a)],
[0068] Among them, Q(s,a) is the Q-value of taking action a in state s; α is the learning rate, which controls the step size of Q-value update; r is the immediate reward obtained after taking action a in state s; γ is the discount factor, which is used to weigh the importance of current rewards and future rewards; s' is the next state transferred from state s after taking action a; a' is the optimal action taken in the next state s'. Monte Carlo tree search helps the algorithm decide which child node to explore next during the selection phase, and the calculation formula is as follows:
[0069]
[0070] Among them, represents the node, S i the average value size of; c is a constant, usually taking the value of 2; N is the total number of explorations; n i is the number of explorations of node S i The exploration times of, the SARSA calculation formula is as follows:
[0071] Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A)),
[0072] Among them, R is the reward of the current state, representing the immediate reward obtained after executing action A in state S; γ is the discount factor, which is used to weigh the importance of current rewards and future rewards; Q(S′,A′): the action value estimate of the next state; α is the learning rate, which determines the amplitude of each update. The DQN calculation formula is as follows:
[0073] Q(s,a)=r+γ*max_a'Q_target(s',a'),
[0074] Among them, Q _target (s', a') is the Q-value calculated through the target network, while Q(s,a) is the Q-value calculated through the online network.
[0075] In step two, the specific process of reinforcement learning is as follows. In path planning, the state represents the position coordinates where the software is located, and the action represents the software moving in the up, down, left, and right directions. Initialize the Q-value to a small random value or zero:
[0076] S1. Initialize the state s to the starting position;
[0077] S2. Select action a, using the ε-greedy strategy, that is, select a random action with a probability of ε, and select the action with the highest Q-value with a probability of 1-ε;
[0078] S3. Execute action a, observe the immediate reward r and the next state s';
[0079] S4. Update the Q-value according to the Q-value update rule;
[0080] S5. Update the state to s'.
[0081] S6. Repeat steps S2 - S5 until the target state is reached.
[0082] In the present invention, the object evolution sub - module evolves numerous attributes such as interfaces, types, and semantics. The message evolution sub - module evolves names, source objects, target objects, and time sequences. The composite fragment evolution sub - module, which belongs to the same category as messages in connectors, evolves. The constraint evolution sub - module adds and deletes constraint information. The static evolution sub - module consults documents, analyzes the architecture, identifies the system composition and the relationships between elements, extracts the abstract representation form of the system, tests the evolved system, and finds the errors and deficiencies therein. The dynamic evolution sub - module, during the running process, changes in user requirements or the adjustment of the system's own service quality will trigger changes in software behavior. The software evolution graph sub - module constructs a software evolution graph through Github to know the relationship between the previous and current versions of the software system. The software architecture graph sub - module obtains the software architecture reversely and converts the software evolution graph into a software architecture evolution graph. The dimensional attribute sub - module uses five - dimensional attributes to measure the software architecture of each version, including the number of atomic components, the size of atomic components, the size of the software architecture, the number of effective code lines, and the number of Java files. The action sequence sub - module uses the change log to record the action sequence of the software architecture changes between adjacent versions. The attribute calculation sub - module calculates the change rate and the initial reward value between adjacent versions using the five - dimensional attributes of the software architecture. The Q - learning sub - module understands the environment through long - term interactive learning and finds out how to maximize the reward function in various situations to form a planning strategy. The Monte Carlo tree search sub - module generates the actual path according to the learned strategy to ensure finding the optimal path efficiently in a complex environment. At each time step, Q - learning updates the Q - value according to the following update rule, and the calculation formula is as follows:
[0083] Q(s,a)←Q(s,a)+α·[r+γ·max α' Q(s',a') - Q(s,a)],
[0084] where Q(s,a) is the Q - value of taking action a in state s; α is the learning rate, which controls the step size of Q - value update; r is the immediate reward obtained after taking action a in state s; γ is the discount factor, which is used to balance the importance of the current reward and future rewards; s' is the next state transferred from state s after taking action a; a' is the optimal action taken in the next state s'. CTS helps the algorithm decide which child node to explore next in the selection stage, and the calculation formula is as follows:
[0085] UCB1(S_i)=\overline{V_i}+c\sqrt{\frac{\log N}{n_i}}, \quad c = 2,
[0086] where \overline{V_i} represents the average value of node S_i; c is a constant, usually taking the value of 2; N is the total number of explorations; n_i is the number of explorations of node S_i, and the SARSA calculation formula is as follows:
[0087] Q(S,A) \leftarrow Q(S,A)+\alpha(R+\gamma Q(S',A') - Q(S,A)),
[0088] where R is the reward in the current state, representing the immediate reward obtained after performing action A in state S; \gamma is the discount factor, used to weigh the importance of the current reward and future rewards; Q(S',A'): the action value estimate of the next state; \alpha is the learning rate, which determines the magnitude of each update. The DQN calculation formula is as follows:
[0089] Q(s,a)=r+\gamma*\max_{a'}Q_{target}(s',a'),
[0090] where Q _target (s',a') is the Q value calculated by the target network, while Q(s,a) is the Q value calculated by the online network.
[0091] Example: The first step in architecture evolution is to classify changes in user requirements. At this stage, developers first need to comprehensively understand the changes in requirements and classify them into existing system components. If there is no existing component corresponding to a change in a requirement, then a mark needs to be made to clarify that this part of the requirement lacks appropriate support and may require the development or adjustment of new components in the future. This process of requirement classification helps to systematically understand the nature of the changes and also provides a basis for subsequent architecture adjustments. The key to this step is to ensure that each change can be matched with a part of the system to avoid omission or duplication of work. After clarifying the requirement changes, the next step is to formulate a detailed evolution plan. This plan needs to cover every aspect of the evolution process, clarify the goals, resource requirements, schedule, and possible risks at each stage. The evolution plan is not just a time management tool; it also helps the development team anticipate potential problems and provides guidance for future implementation. At this stage, the team should cooperate closely with stakeholders such as product managers, developers, and testing teams to ensure the comprehensiveness and feasibility of the plan. Under the guidance of the evolution plan, developers decide whether to modify existing components, add new components, or delete components that are no longer needed based on the classification of requirement changes. The core of this step is how to flexibly adjust the components according to the overall design of the system to meet new requirements. When modifying existing components, special attention needs to be paid not to damage the stability and compatibility of the existing system; when adding new components, it is necessary to ensure that they can be well integrated with the existing system; when deleting components, it is necessary to ensure that the removed part does not affect the function or performance of the system. As components are added, modified, or deleted, the interactions between components in the system, such as data flow and control flow, must be updated accordingly. The interactions between components are one of the cores of architecture design, which ensures the coordinated operation of the system. At this stage, developers need to reconstruct relationships such as control flow and data flow to ensure that the new system can still operate efficiently and stably. This is not just a simple connection between components but also involves the overall optimization and balance of the system to ensure that each component can play its due role in the new architecture.
[0092] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A software architecture evolution path planning system based on reinforcement learning, including a planning system, characterized in that: The planning system includes a software evolution module, a path graph module, a path planning module, a reinforcement learning module, and an analysis and comparison module; The software evolution module evolves the software to adapt to new user requirements, changes in the business environment, and the operating environment. The path graph module constructs a software architecture evolution graph and a software architecture evolution path graph respectively. The path planning module maps the information in the simulated annealing evolution path graph to various factors in reinforcement learning to complete software path planning. The reinforcement learning module model learns the optimal policy by interacting with the environment. Then, the model uses the learned policy to generate the actual path. The analysis and comparison module analyzes the software evolution path and outputs the optimal software evolution path.
2. The software architecture evolution path planning system based on reinforcement learning according to claim 1, wherein: The software evolution module includes an object evolution sub-module, a message evolution sub-module, a composite fragment evolution sub-module, a constraint evolution sub-module, a static evolution sub-module, and a dynamic evolution sub-module; The object evolution sub-module evolves numerous attributes, such as interfaces, types, and semantics. The message evolution sub-module evolves names, source objects, target objects, and time sequences. The composite fragment evolution sub-module, which belongs to the same category as messages in terms of connectors, evolves. The constraint evolution sub-module adds and deletes constraint information. The static evolution sub-module consults documents, analyzes the architecture, identifies the system composition and the interrelationships between elements, extracts the abstract representation form of the system, tests the evolved system, and finds the errors and deficiencies therein. The dynamic evolution sub-module during operation, changes in user requirements or adjustments to the system's own service quality will trigger changes in software behavior.
3. The software architecture evolution path planning system based on reinforcement learning according to claim 1, characterized in that: The path graph module includes a software evolution graph sub-module and a software system graph sub-module; The software evolution graph sub-module constructs a software evolution graph through Github to know the relationship between the previous and current versions of the software system. The software system graph sub-module obtains the software architecture by reverse engineering and converts the software evolution graph into a software architecture evolution graph.
4. The software architecture evolution path planning system based on reinforcement learning according to claim 1, characterized in that: The path planning module includes a dimension attribute sub-module, an action sequence group sub-module, and an attribute calculation sub-module; The dimension attribute sub-module uses attributes in five dimensions to measure the software architecture of each version, namely the number of atomic components, the size of atomic components, the size of the software architecture, the number of effective code lines, and the number of Java files. The action sequence sub-module uses the change log to record the software architecture change action sequence between adjacent versions. The attribute calculation sub-module uses the five-dimensional attributes of the software architecture to calculate the change rate and the initial reward value between adjacent versions.
5. The software architecture evolution path planning system based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning module includes a Q-learning sub-module and a Monte Carlo tree search sub-module; The Q-learning sub-module understands the environment through long-term interactive learning and finds out how to maximize the reward function in various situations to form a planning strategy. The Monte Carlo tree search sub-module generates the actual path according to the learned strategy to ensure finding the optimal path efficiently in a complex environment.
6. The software architecture evolution path planning system based on reinforcement learning according to claim 1, characterized in that: The analysis and comparison module includes a path analysis module, a path comparison module, and a result output module; The path analysis module analyzes the software evolution path, the path comparison module compares other software evolution paths, and the result output module outputs the optimal software evolution path.
7. The usage method of the software architecture evolution path planning system based on reinforcement learning according to any one of claims 1-6, characterized in that, It includes the following steps: Step 1, software evolution: After the software basic information is collected and input, the software evolution graph is constructed through Github to know the relationship between the previous and current versions of the software system. Secondly, the software architecture is obtained through reverse engineering and the software evolution graph is converted into a software architecture evolution graph; Step 2, path planning: The software architecture evolution graph is converted into a software architecture evolution path and reinforcement learning is used to perform path planning on the simulated annealing evolution path graph; Step 3, result output: Compare and analyze according to the planned content and output the optimal software evolution path.
8. The method for using the software architecture evolution path planning system based on reinforcement learning according to claim 7, characterized in that: The reinforcement learning in Step 2 is specifically the Q-value update rule, Monte Carlo tree search, SARSA, and DQN. At each time step, Q-learning updates the Q-value according to the following update rule, and the calculation formula is as follows: Q(s,a) ← Q(s,a) + α·[r + γ·max α' Q(s',a') - Q(s,a)], Where Q(s,a) is the Q-value of taking action a in state s; α is the learning rate, which controls the step size of Q-value update; r is the immediate reward obtained after taking action a in state s; γ is the discount factor, which is used to weigh the importance of the current reward and future rewards; s' is the next state transferred from state s after taking action a; a' is the optimal action taken in the next state s'. The Monte Carlo tree search helps the algorithm decide which child node to explore next in the selection stage, and the calculation formula is as follows: Among them, represents the node, S i is the average value size; c is a constant, usually taking the value of 2; N is the total number of explorations; n i is the number of explorations of the node S i The SARSA calculation formula is as follows: Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A)), Where, R is the reward of the current state, indicating the immediate reward obtained after executing action A in state S; γ is the discount factor, which is used to weigh the importance of the current reward and future rewards; Q(S′,A′): the action value estimate of the next state; α is the learning rate, which determines the amplitude of each update. The calculation formula of DQN is as follows: Q(s,a)=r+γ*max_a'Q_target(s',a'), Among them, Q _target (s', a') is the Q value calculated by the target network, while Q(s, a) is the Q value calculated by the online network.
9. The method of using the software architecture evolution path planning system based on reinforcement learning according to claim 7, characterized in that: The specific process of the reinforcement learning in Step 2 is that in path planning, the state represents the position coordinates of the software, and the action represents the software moving in the up, down, left, and right directions. The Q-value is initialized to a small random value or zero: S1. Initialize the state s to the starting position; S2. Select action a, using the ε-greedy strategy, that is, select a random action with a probability of ε and select the action with the highest Q-value with a probability of 1-ε; S3. Execute action a and observe the immediate reward r and the next state s'; S4. Update the Q-value according to the Q-value update rule; S5. Update the state to s’; S6. Repeat steps S2-S5 until the target state is reached.