Software change influence modeling and simulation analysis method based on large model
By using a large-scale model-based software change impact modeling method, a software model is constructed and quantitative evaluation indicators are generated, which solves the problem of difficulty in quantitatively evaluating the impact of changes in large-scale software development, realizes automated analysis and rapid testing, and saves resources.
Patent Information
- Application Number
- CN202511454703.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-06
AI Technical Summary
Existing technologies make it difficult to quantitatively assess the impact of software changes in large-scale software development, leading to wasted time and resources in full regression testing and a lack of methods to quickly identify affected functions.
A software change impact modeling method based on large models is adopted. By constructing a software model, analyzing the differences between the models before and after the change, and using depth-first search and expression calculation, quantitative evaluation indicators are generated, including call frequency deviation, impact intensity, impact range, and network diameter ratio.
It enables automated quantitative analysis of the impact of software changes, reduces test items, saves time and costs, improves testing efficiency, and quickly identifies affected functions.
Smart Images

Figure CN121277809A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software change impact modeling technology, and in particular to a method for software change impact modeling and simulation analysis based on a large model. Background Technology
[0002] In the development of large-scale software, after software changes, it is necessary to test the affected functions and verify the correctness of the changes before submission. However, as the system's functions, scale, and complexity continue to expand, the impact of software changes becomes increasingly difficult to analyze. Therefore, in actual engineering, a method of testing all test items is often adopted, which is full regression testing. While this method is effective, it results in a significant waste of time and human resources. Therefore, studying the impact of software changes on functionality can help maintenance personnel quickly identify affected functions, eliminating the need to repeatedly test unaffected parts during testing. This effectively reduces test items and time, saves testing costs and resources, and has significant research value and practical significance.
[0003] Large Language Models (LLMs) are deep learning models trained on large amounts of text data, enabling them to generate natural language text or understand the meaning of language text. With the rapid development of LLMs, they have been widely applied in the field of software engineering. LLMs have demonstrated superior performance in various tasks related to code understanding and generation, and have numerous applications in software development, software quality assurance, and software maintenance, providing new opportunities for traditional software change impact analysis techniques. Traditional software change impact analysis techniques use various testing tools to analyze the internal dependencies of software or use machine learning methods to predict the potential impact of changes. However, these analytical methods can only analyze programs written in specific languages and can only provide the names of potentially affected elements, lacking indicators to assess the degree of impact. This makes traditional analytical methods primarily qualitative, lacking specific quantitative analytical capabilities.
[0004] Therefore, there is an urgent need for a method for modeling and simulating the impact of software changes based on large models. Summary of the Invention
[0005] This invention provides a method for modeling and simulating the impact of software changes based on a large model, in order to solve the above-mentioned problems existing in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A method for modeling and simulating the impact of software changes based on a large model, comprising:
[0008] S1: Obtain the software source code, construct a structured prompt word containing software model element definitions, mapping rules, model structure templates and analysis examples, input the prompt word and source code into the large model, and output a software model containing a set of nodes and a set of edges;
[0009] S2: Analyze the software model to construct the software network, compare the edge attributes of the software model before and after the change to determine the change point, traverse the software network through depth-first search to perform variable passing and expression calculation, record the usage of change marker variables, and generate statistical data on the number of calls to nodes and edges and the number of calls affected.
[0010] S3: Substitute the statistical data into the calculation formulas for call frequency deviation, impact intensity, impact range, and network diameter ratio to calculate the values of four quantitative evaluation indicators.
[0011] Step S1 includes:
[0012] S11: Define the names, input variables, global variables, and type attributes of nodes in the software model; define the starting node, ending node, output variables, constraints, and type attributes of edges.
[0013] S12: Define rules for directly extracting visible elements from source code, and define rules for deriving output expressions through variable substitution and deriving constraints through logical operations;
[0014] S13: Construct the software model definition, extraction rules, JSON format template, and sample code into prompt words, and input them into the large model along with the source code to be analyzed.
[0015] Step S2 includes:
[0016] S21: Analyze the node-edge relationships of the software model to construct the network topology, store the initial values of global variables, compare the differences between the output expressions and constraint expressions of the edges before and after the change, and mark the changed variables;
[0017] S22: Starting from the initial node, assign random input values, visit nodes in depth-first order, verify edge constraints, and perform expression calculation and variable passing on edges that meet the conditions;
[0018] S23: Detect the participation of change marker variables during expression evaluation, assign the evaluation result of the change variables to the change marker, and accumulate the number of times each element is called and the number of times the change affects the expression.
[0019] Step S3 includes:
[0020] S31: Substitute the number of calls before and after the change into the frequency deviation formula to calculate the call change rate of each element, and substitute the number of affected calls and the total number of calls into the influence strength formula to calculate the influence ratio;
[0021] S32: Calculate the percentage of affected nodes by counting the number of nodes affected by the call, and calculate the ratio of the maximum shortest path length of the sub-network of the affected nodes to the length of the entire network.
[0022] S33: Output the specific values of call frequency deviation, impact intensity, impact range, and network diameter ratio as the results of the change impact assessment.
[0023] The rule-making process in step S12 includes:
[0024] Lexical analysis is used to extract the function name and parameter list from the function signature, and syntax analysis is used to extract global variable references and function call statements from the function body.
[0025] For output variable expressions, trace back from the calling parameters to their assignment statements, recursively replacing local variables in the rvalues until only input parameters and global variables are included;
[0026] For constraints, collect the branch conditions on the call path, combine them through conjunction, and perform a recursive substitution process on the variables in the conditions.
[0027] In step S22, different processing is performed depending on the edge type:
[0028] The edge call, which returns no value, directly calculates the output expression and passes it to the target node;
[0029] The request returns a value that the calling edge passes the parameters to the target node and pushes that node onto the call stack.
[0030] The return value is passed by popping the source node from the call stack and passing the return value.
[0031] The loop calls identify the set of values for the loop control variable and create an independent call instance for each value;
[0032] The global variable modification edge writes the new value to the corresponding position in the global variable storage area.
[0033] The call frequency deviation is the ratio of the difference in the number of calls to nodes and edges in the software model before and after the change to the number of calls before the modification. The calculation formula is as follows:
[0034]
[0035] in Indicates the frequency deviation of calls. The number of calls after the change. This represents the number of calls made before the change.
[0036] The impact strength is the percentage of node or edge calls affected by the software change, calculated as follows:
[0037]
[0038] in The number of calls affected by the change. Total number of calls;
[0039] The diameter is the maximum value of the shortest path length between any two nodes, and the diameter ratio is the ratio of the diameter of the software subnetwork formed by the affected functions to the diameter of the complete software network. The calculation method is as follows:
[0040]
[0041] in, For the diameter of the complete software network, The diameter of the software subnetwork formed by the affected functions.
[0042] The process of constructing prompt words in step S13:
[0043] The JSON format template defines the hierarchical structure of the software model, including the top-level node array and edge array, as well as the attribute fields of each element;
[0044] The example code covers typical scenarios for five node types and five edge types, and each example is accompanied by the corresponding model extraction results;
[0045] The source code to be analyzed is organized in the form of a dictionary with file paths as keys and code text as values, and supports cross-file call relationship analysis.
[0046] The change marking process in step S21 includes:
[0047] Traverse all edges in the model before and after the change, and filter edge pairs where the starting and ending nodes are the same;
[0048] Compare the output variable expressions of each pair of edges one by one, and add the variables with different expression texts to the change point set;
[0049] Compare the constraint expressions of the edge pairs; if there are differences, add all output variables of that edge to the change point set.
[0050] Write the set of change points as attributes into the corresponding edge object for use in simulation traversal.
[0051] Compared with the prior art, the present invention has the following advantages:
[0052] A method for modeling and simulating the impact of software changes based on a large model is proposed. This method enables automated software modeling, change location identification, and dynamic simulation during software maintenance, ultimately achieving quantitative analysis of the software change impact. Compared to traditional full regression testing in software maintenance, this analysis method helps maintenance personnel quickly identify the functions affected by changes, effectively reducing test items, shortening maintenance time, and saving maintenance costs. Compared to traditional qualitative analysis of software change impacts, this method uses indicators to assess the impact of changes, providing a basis for maintenance personnel to allocate test resources and improving testing efficiency.
[0053] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention.
[0054] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is a schematic diagram of the structure of the prompt words in an embodiment of the present invention;
[0057] Figure 2 This is a schematic diagram of the simulation preparation process in an embodiment of the present invention;
[0058] Figure 3 This is a flowchart of the software dynamic simulation algorithm in an embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the edge type parsing strategy in an embodiment of the present invention. Figure 1 ;
[0060] Figure 5 This is a schematic diagram of the edge type parsing strategy in an embodiment of the present invention. Figure 2 ;
[0061] Figure 6 This is a schematic diagram of the edge type parsing strategy in an embodiment of the present invention. Figure 3 ;
[0062] Figure 7 This is a schematic diagram of the edge type parsing strategy in an embodiment of the present invention. Figure 4 . Detailed Implementation
[0063] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0064] This invention provides a method for modeling and simulating the impact of software changes based on a large model, including:
[0065] S1: Obtain the software source code, construct a structured prompt word containing software model element definitions, mapping rules, model structure templates and analysis examples, input the prompt word and source code into the large model, and output a software model containing a set of nodes and a set of edges;
[0066] S2: Analyze the software model to construct the software network, compare the edge attributes of the software model before and after the change to determine the change points, traverse the software network through depth-first search to perform variable passing and expression calculation, record the usage of change marker variables, and generate statistical data on the number of calls to nodes and edges and the number of calls affected; perform differentiated parsing according to edge type, including edges that call without return value, edges that request return value, edges that pass return value, edges that call in a loop (the range of values for the loop variable needs to be identified and an independent call instance needs to be created for each value), and edges that modify global variables (pointing to virtual global variable nodes and updating the global storage area).
[0067] S3: Substitute the statistical data into the calculation formulas for call frequency deviation, impact intensity, impact range, and network diameter ratio to calculate the values of four quantitative evaluation indicators. The impact range calculation requires statistical analysis of the proportion of affected nodes, and the network diameter ratio calculation requires extracting the sub-networks of affected nodes and calculating their diameters using the shortest path algorithm.
[0068] The working principle and beneficial effects of the above technical solution are as follows: Step 1: Software model generation method based on large model: This part proposes the definition of each component in the software model from the perspective of software change impact analysis, and designs a mapping method from software source code to software model. Combined with hint engineering guidance, it achieves automated software model generation based on the large model. The specific research steps are as follows:
[0069] Step 1.1: Software Model Composition. The software model is a collection of elements extracted from the software source code that are relevant to the change impact analysis. The software model is defined as follows:
[0070] in, It is a software model, a network based on function call relationships. These are the nodes of the model, i.e., the functions of the software. These are the edges of the model, representing the function call behavior of the software. Their specific properties are as follows:
[0071] for It consists of four attributes:
[0072] (1) , is the name of the node, i.e., the name of the function;
[0073] (2) , which are the input variables of the function;
[0074] (3) This is a global variable used by the function;
[0075] (4) , where is the type of the node, i.e., the type of the function. Here, is used... This represents the software's initial function. This represents the software termination function. This indicates a function that has no return value. This indicates a function that returns a value. This indicates a function that neither calls other functions nor is called by other functions.
[0076] for It consists of five attributes:
[0077] (1) , is the starting node of the edge, i.e., the function that initiates the call;
[0078] (2) , is the terminating node of the edge, i.e., the function being called;
[0079] (3) This is for passing variables during the call;
[0080] (4) , which are the constraints that the call must satisfy;
[0081] (5) , where is the type of the edge, i.e., the type of the function call. Here, uses This indicates a call that has no return value. This indicates a call that requests a return value from a terminating function. This indicates that the return value will be passed back to the call to the initial function. This indicates that the starting function repeatedly calls the ending function through loop logic. Specifically, it uses... This indicates modifications to global variables within the starting function; in this case, its terminating node is a virtual node.
[0082] Step 1.2: Mapping Method from Software Source Code to Software Model. The mapping method is the way to abstract and extract various elements from the software source code to form the software model. In the software model, , , , , and This can be obtained directly by observing the software's source code, and and The software's operational logic needs to be extracted from its source code. The detailed steps are as follows:
[0083] Step 1.2.1: The acquisition of parameters. In the software source code, the call that passes parameters specifies the name of the parameter. For a parameter, it must appear as the result of some calculation in the preceding text. Therefore, this calculation expression is used to represent the parameter. At the same time, we look for the variables in the preceding text that are calculated from, and perform the same substitution until all variables involved in the calculation are function input variables or global variables. At this point, for each passed parameter, we can obtain an expression that is calculated only by the function's input variables and global variables. The set of such expressions is the function's input variables. The value of .
[0084] Step 1.2.2: The acquisition of constraints is achieved through logical expressions used in the software source code. For each function call, the intersection of the logical expressions representing the constraints is taken, and then... Similarly, by replacing all variables in the expression with expressions that only involve the function's input variables and global variables in the calculation, the final logical expression is obtained as follows: The value of .
[0085] Step 1.3: Software Model Generation Method Based on Hint Engineering. This embodiment uses hint engineering to guide the large model to complete the mapping behavior in Step 1.2, dividing the hint words into two parts: description and input, such as... Figure 1 As shown, the structure of the prompt words is illustrated:
[0086] (1) Explanation: This section contains four elements. The task objective describes the definition of each element in the software model using natural language; the mapping rules describe how to abstract and extract the elements defined in the task objective from the software, and specify the renaming rules to prevent variable naming conflicts. This rule uses five elements to form a new name, namely the relative path of the file, the name of the function, the data type of the variable, the source of the variable (input variable, return value, or global variable), and the order in which the variable appears in the function; the model structure provides a JSON format to guide the large model in organizing the software model; the analysis examples provide a combination of several software snippets and human analysis results. These examples contain all possible software model element types to help the large model learn how to extract the software model.
[0087] (2) Input. This section uses a key-value pair format to input the software to be analyzed, where the key is the relative path of the file and the value is the code text in the file. This input enables large models to autonomously analyze function call relationships across files, realizing the analysis of large software.
[0088] Step 2: Simulation Method for Software Change Impact Analysis
[0089] Based on the software model obtained in step 1, a simulation algorithm is designed to achieve model-driven dynamic software simulation, while simultaneously recording preliminary statistical data on the impact of software changes. The specific research steps are as follows:
[0090] Step 2.1: Preparations for software dynamic simulation. Preparations include... Figure 2 As shown, the process consists of three steps: network construction, global variable recording, and change point injection, as detailed below:
[0091] Step 2.1.1: Network Construction. Based on the software model from Step 1, design simulation software to read the software model file, analyze the relationships between the elements of the software model, and construct the software network.
[0092] Step 2.1.2: Recording global variables. If global variables exist in the software model, record their initial values when the software starts running, for use in initializing the simulation environment in subsequent simulations.
[0093] Step 2.1.3: Change Point Injection. First, compare the software model before and after the change to find edges that are inconsistent. For edges with the same start and end nodes, compare their... and Attributes. If If an expression with an output variable changes, then that output variable is considered a change point. If If the logical expression changes, all output variables are considered as change points. Finally, the information of the change points is written into the corresponding edge elements of the software network.
[0094] Step 2.2: Software dynamic simulation algorithm. The simulation algorithm is as follows: Figure 3 As shown, a depth-first search rule is used to traverse the software network and reconstruct the software's execution logic. The specific steps are as follows:
[0095] Step 2.2.1: Random Input Values. At the start of the simulation, a starting node of the software is randomly selected, and a random number is used to generate the input value for that node. This random number has the same values and distribution patterns as the input variables that may be input during the actual operation of the software.
[0096] Step 2.2.2: Proceed to the next node. Pass the variable to the node as its input variable, which will be used for the calculation of various expressions in the future.
[0097] Step 2.2.3: Check constraints. Find all edges that start from this node, and check each edge to see if it satisfies the constraints. The logical judgment result is true. If true, the corresponding edge is parsed.
[0098] Step 2.2.4: Edge parsing. In the simulation, the static paths recorded by the software model need to be expanded into dynamic paths by combining them with the input variables. For the five edge types proposed in Step 1.1, this embodiment designs the following parsing strategy:
[0099] (1) Parsing of type edges. For example... Figure 4 As shown, under the condition that the constraints are satisfied, the output variable is calculated directly from the calculation expression using the input variables and global variables.
[0100] (2) and Parsing of type edges. Because... type of edge and Edges of type appear in pairs, therefore, as Figure 5 As shown, when designing the algorithm, the output variable of the starting node is first passed through... Edges of type [type] are passed to the terminal node, which then determines which edge to use. The edge type is returned, and finally the output variables of the terminating node are returned to the starting node, and these variables are added to the input variables of the starting node.
[0101] (3) Parsing of type edges. For example... Figure 6 As shown, first determine the loop variable in the calculation expression, and then generate several lines based on all possible values of the loop variable. For each edge of type , calculate the output variable for each edge, and finally pass the variables in sequence.
[0102] (4) Parsing of type edges. For example... Figure 7 As shown, to enable changes to global variables, a virtual global variable node is introduced and integrated into the software network. When a global variable changes, the starting node initiates an update to the existing global variable.
[0103] Step 2.2.5: Output Variable Calculation. After parsing the edges in Step 2.2.4, perform calculations on all expressions to obtain the output variables of the starting node.
[0104] Step 2.2.6: Impact Assessment of Changes. In the calculations of Step 2.2.5, check whether variables affected by the change were used. If so, the calculated output variable is also considered an affected variable, and this check will be included in the next iteration.
[0105] Step 2.2.7: Exit the node. The starting node passes the output variable to the ending node through the edge. At this point, it exits the starting node and prepares to visit the ending node, using it as the next starting node. Steps 2.2.2 to 2.2.7 are repeated. This continues until no edge satisfying the constraints is found in step 2.2.3. Then, it returns to the previous node and resolves the next edge that meets the requirements according to step 2.2.4.
[0106] Step 2.2.8: Simulation End. A simulation is considered complete when all edges satisfying the constraints in all nodes have been resolved. At this point, the number of times each edge was called during the simulation, and the number of calls affected by changes, can be counted. For all nodes, their total number of calls and the number of calls affected by changes are the sum of the total number of calls to all incoming edges and the number of calls affected by changes. This data will serve as the raw data for the calculations in Step 3.
[0107] Step 3: Evaluation Methods for Software Change Impact Analysis
[0108] Based on step 2, which recorded the number of calls to each node and edge in the software model, including the number of calls affected by changes, four indicators are proposed to quantitatively analyze the impact of changes. The meaning and calculation methods of these indicators are given. The specific indicators are described below:
[0109] Step 3.1: Frequency deviation Definition and calculation. The call frequency deviation is the ratio of the difference in the number of calls to nodes and edges in the software model before and after the change to the number of calls before the modification, and the calculation method is as follows:
[0110]
[0111] in, The number of calls before the change. This represents the number of calls after the change. A positive value indicates that the node or edge has been called more frequently after modification. A negative value indicates that the frequency of calls to the node or edge has decreased after modification, and The larger the absolute value, the greater the change in call frequency caused by the influence of the node or edge.
[0112] Step 3.2: Influence Intensity The definition and calculation of the impact intensity. The impact intensity is the proportion of calls to nodes or edges affected by the software change, calculated as follows:
[0113]
[0114] in, The number of calls affected by the change. This represents the total number of calls.
[0115] The larger the value, the greater the impact of the change on the node or edge. Specifically, when... When this occurs, it indicates that the node or edge will definitely be affected by the change.
[0116] Step 3.3: Scope of Impact The definition and calculation of the impact range. The impact range is the ratio of the affected functions to all functions in the software, calculated as follows:
[0117]
[0118] in, It is a set of software functions, with the numerator representing the software functions affected by the changes. The larger the value, the more functions are affected by the change.
[0119] Step 3.4: Software network influences diameter ratio Definition and calculation of . The software network is the network structure of the software model, the diameter is the maximum value of the shortest path length between any two nodes, and the diameter ratio is the ratio of the diameter of the software sub-network formed by the affected functions to the diameter of the complete software network, calculated as follows:
[0120]
[0121] in, For the diameter of the complete software network, The diameter of the software subnetwork formed by the affected functions. The larger the value, the greater the diffusion of the impact of software changes within the software network.
[0122] In another embodiment, step S1 includes:
[0123] S11: Define the names, input variables, global variables, and type attributes of nodes in the software model; define the starting node, ending node, output variable, constraint conditions, and type attributes of edges; among which the edge type includes the global type, and its ending node is a virtual global variable node;
[0124] S12: Define rules for directly extracting visible elements from source code, and define rules for deriving output expressions through variable substitution and deriving constraints through logical operations;
[0125] S13: Construct the software model definition, extraction rules, JSON format template, and sample code into prompt words, and input them into the large model along with the source code to be analyzed.
[0126] The working principle and beneficial effects of the above technical solution are as follows: In step S11, the system defines the basic element structure of the software model. Node elements contain four attributes: name (function name), input (function input parameters), global_value (global variable used by the function), and type (function type, including start function, end function, normal function with no return value, return function with return value, and idle function). Edge elements contain five attributes: the starting node n_source (function initiating the call), the ending node n_target (function being called), output (variable passing during the call), condition (call constraints), and type (call type, including normal call with no return value, call requesting a return value, callback passing a return value, loop call, and global variable modification).
[0127] In step S12, the system establishes mapping rules from source code to model elements. For directly visible elements such as function names, parameter lists, and global variable references, these are extracted directly through lexical and syntactic analysis. For output variable expressions, the system traces back from the call parameters to their assignment statements, recursively replacing local variables in the expression until only input parameters and global variables are included. For constraints, the system collects branch conditions along the call path, combines them logically, and recursively replaces variables. To prevent naming conflicts, a new name is constructed using five elements: file path, function name, data type, variable source, and order of appearance.
[0128] In step S13, the system integrates the model structure defined in S11 and the mapping rules established in S12 into structured prompts. The prompts consist of four parts: a description of the task objective, an explanation of the mapping rules, a JSON template, and sample code. The examples cover all possible node and edge types, and each example is accompanied by corresponding model extraction results to assist the large model's learning and extraction process. The source code to be analyzed is organized in dictionary form with file paths as keys and code text as values, supporting cross-file relationship analysis.
[0129] The process of constructing prompt words: The JSON format template defines the hierarchical structure of the software model, including the top-level node array and edge array, as well as the attribute fields of each element; the sample code covers typical scenarios of five node types (start, end, normal, return, idle) and five edge types (normal, call, callback, loop, global), and each example is accompanied by the corresponding model extraction results; the source code to be analyzed is organized in the form of a dictionary with file path as the key and code text as the value, supporting cross-file call relationship analysis.
[0130] In another embodiment, step S2 includes:
[0131] S21: Analyze the node-edge relationships of the software model to construct the network topology, store the initial values of global variables, compare the differences between the output expressions and constraint expressions of the edges before and after the change, and mark the changed variables;
[0132] S22: Starting from the initial node, assign random input values, visit nodes in depth-first order, verify edge constraints, and perform expression calculation and variable passing on edges that meet the conditions;
[0133] S23: Detect the participation of change marker variables during expression evaluation, assign the evaluation result of the change variables to the change marker, and accumulate the number of times each element is called and the number of times the change affects the expression.
[0134] The working principle and beneficial effects of the above technical solution are as follows: In step S21, the system first parses the node-edge relationships in the software model and constructs a directed graph-like network topology. It reads and stores the initial values of all global variables to provide environment initialization data for subsequent simulations. By comparing edges in the model before and after the change where the starting and ending nodes are the same, it checks the differences between their output variable expressions and constraint expressions. If the output expression changes, the corresponding output variable is added to the change point set; if the constraint expressions differ, all output variables of that edge are added to the change point set. The change point information is written as an attribute to the corresponding edge object. The change marking process includes: traversing all edges in the model before and after the change, filtering edge pairs where the starting and ending nodes are the same; comparing the output variable expressions of each edge pair, adding variables with different expression texts to the change point set; comparing the constraint expressions of the edge pairs, adding all output variables of that edge to the change point set if differences exist; and writing the change point set as an attribute to the corresponding edge object for querying during simulation traversal.
[0135] In step S22, the system performs a simulation traversal starting from the starting node. A random number generator is used to assign random values to the input parameters of the starting node, conforming to the actual operating rules. Nodes in the network are visited in a depth-first search order. For each node, all edges that serve as the starting node are searched, and the constraint condition logical expressions of each edge are verified to be true. For edges that satisfy the conditions, the corresponding expression calculation and variable passing operations are performed according to their type. When a node has no outgoing edges that satisfy the constraints, the traversal backtracks to the previous node.
[0136] In step S23, the system monitors the participation of change-marked variables in real time during expression calculation. When a calculated expression uses a variable with a change mark, the resulting variable is also marked with a change mark, thus tracking the propagation of the change's impact. Throughout the simulation, the system continuously accumulates the call count for each node and edge, while also separately counting the number of calls affected by the change. These statistics will serve as the basis for subsequent quantitative analysis.
[0137] In another embodiment, step S3 includes:
[0138] S31: Substitute the number of calls before and after the change into the frequency deviation formula to calculate the call change rate of each element, and substitute the number of affected calls and the total number of calls into the influence strength formula to calculate the influence ratio;
[0139] S32: Calculate the percentage of affected nodes by counting the number of nodes affected by the call, and calculate the ratio of the maximum shortest path length of the sub-network of the affected nodes to the length of the entire network.
[0140] S33: Output the specific values of call frequency deviation, impact intensity, impact range, and network diameter ratio as the results of the change impact assessment.
[0141] The working principle and beneficial effects of the above technical solution are as follows: In step S31, the system first calculates two indicators: call frequency deviation and impact intensity. For call frequency deviation, the number of calls after the change (n_after) is subtracted from the number of calls before the change (n_before), then divided by the number of calls before the change (n_before) and multiplied by 100% to obtain the frequency change rate in percentage form. A positive value indicates an increase in call frequency, and a negative value indicates a decrease; the absolute value reflects the degree of change. For impact intensity, the number of affected calls (n_affected) is divided by the total number of calls (n) and multiplied by 100% to obtain the impact percentage; the larger the value, the more severely the element is affected by the change.
[0142] In step S32, the system calculates two network-level metrics: the impact range and the network diameter ratio. The impact range is calculated by dividing the number of nodes with affected calls (i.e., n_affected > 0) by the total number of nodes to obtain the proportion of affected nodes. The network diameter ratio requires first extracting the sub-networks formed by all affected nodes, using the shortest path algorithm to calculate the maximum shortest path length between any two nodes in the sub-network as the sub-network diameter, then calculating the diameter of the complete network, and dividing the two to obtain the diameter ratio. The impact range proportion is calculated by counting the number of nodes with affected calls, and the ratio of the maximum shortest path length (using the shortest path algorithm) of the affected node sub-network to the length of the complete network is calculated.
[0143] In step S33, the system integrates the calculation results of the four indicators to form a complete change impact assessment report. The call frequency deviation reflects the changing trend of the usage frequency of each element; the impact intensity quantifies the severity of the impact on each element; the impact scope assesses the ripple effect of the change in the software; and the network diameter ratio measures the depth of the impact propagation in the call relationship network. These four indicators comprehensively characterize the impact features of the software change from different dimensions.
[0144] In another embodiment, the rule-making process in step S12 includes:
[0145] Lexical analysis is used to extract the function name and parameter list from the function signature, and syntax analysis is used to extract global variable references and function call statements from the function body.
[0146] For output variable expressions, trace back from the calling parameters to their assignment statements, recursively replacing local variables in the rvalues until only input parameters and global variables are included;
[0147] For constraints, collect the branch conditions on the call path, combine them through conjunction, and perform a recursive substitution process on the variables in the conditions;
[0148] The rules for renaming variables use a combination of five elements: file path, function name, data type, variable source, and order of appearance.
[0149] The working principle and beneficial effects of the above technical solution are as follows: During the rule-making process, the system scans the source code using a lexical analyzer to identify function definition statements and extract the function name and parameter list from the function signature. An abstract syntax tree is constructed using a syntax analyzer, traversing the statement nodes within the function body to identify global variable references and function call statements.
[0150] For deriving the expression for the output variable, the system starts from the actual argument of the function call statement and traces backwards to the most recent assignment statement for that variable. After obtaining the expression on the right-hand side of the assignment statement, it checks whether the expression contains local variables. If a local variable exists, it continues to trace backwards to the source of the local variable's assignment and replaces the variable in the current expression with its assignment expression. This recursive replacement process is repeated until the expression contains only function input parameters and global variables, resulting in the final output expression.
[0151] To derive the constraints, the system analyzes all execution paths from the function entry point to the calling statement, collecting conditional statements (such as if, while, etc.) along the path. These conditions are combined using logical AND operations to form complete constraints that the call must satisfy. Recursive substitution is also performed on the variables in the constraint expressions, replacing local variables with their calculated expressions, ultimately yielding constraints containing only input parameters and global variables.
[0152] In another embodiment, step S22 performs different processing based on the edge type:
[0153] The edge call, which returns no value, directly calculates the output expression and passes it to the target node;
[0154] The request returns a value that the calling edge passes the parameters to the target node and pushes that node onto the call stack.
[0155] The return value is passed by popping the source node from the call stack and passing the return value.
[0156] The loop calls identify the set of values for the loop control variable and create an independent call instance for each value;
[0157] The global variable modification edge writes the new value to the corresponding position in the global variable storage area (virtual global variable node).
[0158] The working principle and beneficial effects of the above technical solution are as follows: The system executes differentiated processing strategies based on the type attributes of the edges. For edges with no return value (normal type), when the constraints are met, the input variables and global variables of the current node are directly substituted into the output expression for calculation, and the calculation result is passed as a parameter to the target node, completing the one-way variable transfer.
[0159] Request return value call edges (of type call) and return value passing edges (of type callback) need to be processed in coordination. When a call type edge is encountered, the system passes the parameters to the target node and pushes the current node onto the call stack to save the context. After the target node completes execution, it pops the source node from the call stack via the callback type edge, passes the return value back to the source node, and adds it to its input variable set.
[0160] Loop-type edges require identifying the loop control variable and its value range. The system analyzes the loop condition expression, determines all possible values of the loop variable, and creates an independent call instance for each value. The loop-type edge is expanded into multiple normal-type edges, and the output variable of each instance is calculated and passed sequentially.
[0161] Global variable modification edges (global type) are implemented by introducing virtual global variable nodes. When a global variable assignment operation is detected, a global type edge pointing to the virtual node is created, the new value is written to the corresponding position in the global variable storage area, and the global state is updated for subsequent nodes to access.
[0162] In another embodiment, the call frequency deviation is the ratio of the difference in the number of calls to nodes and edges in the software model before and after the change to the number of calls before the modification, and the calculation formula is:
[0163]
[0164] in Indicates the frequency deviation of calls. The number of calls after the change. This represents the number of calls made before the change.
[0165] The impact strength is the percentage of node or edge calls affected by the software change, calculated as follows:
[0166] in The number of calls affected by the change. Total number of calls;
[0167] The scope of impact is the ratio of affected functions to all functions in the software, calculated as follows:
[0168]
[0169] in, It is a set of software functions, with the numerator representing the software functions affected by the changes. The larger the value, the more functions are affected by the change.
[0170] The diameter is the maximum value of the shortest path length between any two nodes, and the diameter ratio is the ratio of the diameter of the software subnetwork formed by the affected functions to the diameter of the complete software network. The calculation method is as follows:
[0171]
[0172] in, For the diameter of the complete software network, The diameter of the software subnetwork formed by the affected functions.
[0173] The working principle and beneficial effects of the above technical solution are as follows: The Invocation Frequency Deviation (ICD) quantifies the change in the usage frequency of a software element before and after a modification. During calculation, the number of calls after the modification (n_after) is subtracted from the number of calls before the modification (n_before) to obtain the change in the number of calls. This change is then normalized by dividing by the number of calls before the modification (n_before), and finally multiplied by 100% to convert it to a percentage. A positive ICD value indicates that the element is called more frequently after the modification, while a negative value indicates a decrease in calls. The absolute value reflects the degree of drastic change.
[0174] Impact Intensity II measures the degree of impact of a change on a software element. It is calculated by counting the number of calls to the element affected by the change (n_affected) out of all calls to that element, dividing this number by the total number of calls (n), and then multiplying by 100%. The closer the II value is to 100%, the more susceptible the element is to the change. When II equals 100%, it means that every call to the element is affected by the change.
[0175] The network diameter ratio (IDR) reflects the extent to which the impact of a change spreads within the software call network. The diameter is defined as the maximum value of the shortest path length between any two nodes in the network, representing the maximum propagation distance. The system first calculates the diameter (diameter(G)) of the complete software network G, then extracts the sub-network G_sub composed of all affected functions and calculates its diameter (diameter(G_sub)). Dividing the two and multiplying by 100% yields the diameter ratio. A larger IDR value indicates a longer propagation path for the change's impact and a wider affected network area.
[0176] In another embodiment, the process of constructing the prompt word in step S13:
[0177] The JSON format template defines the hierarchical structure of the software model, including the top-level node array and edge array, as well as the attribute fields of each element;
[0178] The example code covers typical scenarios for five node types and five edge types, and each example is accompanied by the corresponding model extraction results;
[0179] The source code to be analyzed is organized in the form of a dictionary with file paths as keys and code text as values, and supports cross-file call relationship analysis.
[0180] The working principle and beneficial effects of the above technical solution are as follows: During the construction of prompt words, the JSON format template defines a standardized hierarchical structure for the software model. The top level contains two arrays, `nodes` and `edges`, which store the collections of nodes and edges, respectively. Node objects contain four attribute fields: `name`, `input`, `global_value`, and `type`, while edge objects contain five attribute fields: `source`, `target`, `output`, `condition`, and `type`. This structured format ensures the consistency and parsability of the large model output.
[0181] The example code is carefully designed to cover typical programming scenarios for five node types (start, end, normal, return, idle) and five edge types (normal, call, callback, loop, global). Each example includes a specific source code snippet and the corresponding correct model extraction result. Through these examples, the large model can learn the mapping relationship between different code patterns and model elements, improving extraction accuracy.
[0182] The source code to be analyzed is organized in dictionary form, with relative file paths as keys and complete code text as values. This organization allows large models to identify function call relationships across files, supporting the analysis of large, multi-file software projects. Path information is also used to generate unique variable identifiers, avoiding conflicts between variables with the same name in different files.
[0183] In another embodiment, the change mark process in step S21 includes:
[0184] Traverse all edges in the model before and after the change, and filter edge pairs where the starting and ending nodes are the same;
[0185] Compare the output variable expressions of each pair of edges one by one, and add the variables with different expression texts to the change point set;
[0186] Compare the constraint expressions of the edge pairs; if there are differences, add all output variables of that edge to the change point set.
[0187] Write the set of change points as attributes into the corresponding edge object for use in simulation traversal.
[0188] The working principle and beneficial effects of the above technical solution are as follows: The change marking process first traverses all edges in the two software models before and after the change, establishing the correspondence between the edges. By comparing the start and end node attributes of each edge, edges that exist in both models and connect the same pair of nodes are selected. These edge pairs may contain changes.
[0189] For each selected edge pair, the system compares its output variable expression text one by one. A string comparison algorithm is used to detect whether the expression has changed. If the calculated expression of an output variable is inconsistent before and after the change, the variable identifier is added to the change point set. This comparison can accurately locate which variables' calculation logic has been modified.
[0190] Constraint comparison uses a similar method, comparing the text of the `condition` attribute of edge pairs. If any difference exists in the constraint expressions, it means the execution conditions of the call have changed, potentially leading to different execution paths. In this case, all output variables of that edge are added to the change point set, because the constraint change may affect all data passed through that edge.
[0191] Finally, the system writes the constructed set of change points as new attributes into the corresponding edge objects. During subsequent simulation traversal, it can quickly query whether each edge contains change points and which specific variables are affected, achieving precise tracking of the impact of changes.
[0192] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from the spirit and scope of this invention.
Claims
1. A large model-based software change impact modeling and simulation analysis method, characterized in that, Comprise: S1: obtain software source code, build structured prompts containing software model element definition, mapping rule, model structure template and analysis example, input prompts and source code into large model, output software model containing node set and edge set; S2: parse software model to build software network, compare edge attributes of software model before and after change to determine change point, perform variable passing and expression calculation by traversing software network through depth-first search, record use of change marker variable, generate statistical data of node and edge call frequency and affected call frequency; S3: substitute statistical data into calculation formula of call frequency deviation, influence strength, influence range and network diameter ratio to calculate four quantitative evaluation index values.
2. The large model based software change impact modeling and simulation analysis method of claim 1, wherein, S1 step includes: S11: define name, input variable, global variable and type attribute of node in software model, define start node, end node, output variable, constraint condition and type attribute of edge; S12: develop rules for directly extracting visible elements from source code, develop rules for deriving output expression through variable substitution and deriving constraint condition through logical operation; S13: build software model definition, extraction rule, JSON format template and example code into prompts, input prompts into large model together with source code to be analyzed. 3.The large model based software change impact modeling and simulation analysis method of claim 1, wherein, S2 step includes: S21: parse node and edge relationship of software model to build network topology, store initial value of global variable, compare differences between output expression and constraint expression of edge before and after change to mark change variable; S22: assign input random value to start node, access node in depth-first order, verify edge constraint condition, perform expression calculation and variable passing on edge meeting condition; S23: detect participation of change marker variable during expression calculation, assign calculation result using change variable to change marker, accumulate call frequency and change influence frequency of each element.
4. The large model based software change impact modeling and simulation analysis method of claim 1, wherein, S3 step includes: S31: substitute call frequency before and after change into frequency deviation formula to calculate call change rate of each element, substitute affected call frequency and total frequency into influence strength formula to calculate influence proportion; S32: count number of nodes with affected call to calculate influence range proportion, calculate ratio of maximum and minimum path length of affected node subnetwork to complete network; S33: output specific values of call frequency deviation, influence strength, influence range and network diameter ratio as change influence evaluation results.
5. The large model based software change impact modeling and simulation analysis method of claim 2, wherein, Rule development process of S12 step includes: Extract function name and parameter list in function signature through lexical analysis, extract global variable reference and function call statement in function body through syntax analysis; For output variable expression, trace back its assignment statement from call parameter, recursively replace local variable in right value until only input parameter and global variable are left; For constraint condition, collect branch conditions on call path, combine through conjunction operation, and perform recursive replacement process on variables in condition.
6. The large model based software change impact modeling and simulation analysis method of claim 3, wherein, S22 step performs different processing according to edge type: No return value call edge directly calculates output expression and passes to target node; Request return value call edge passes parameter to target node and pushes this node into call stack; Return value passing edges pop the source node from the call stack and pass the return value; Loop calling edges identify the value set of the loop control variable and create independent calling instances for each value; Global variable modifying edges write the new value to the corresponding location in the global variable storage.
7. The large model based software change impact modeling and simulation analysis method of claim 1, wherein, The calling frequency deviation is the ratio of the difference in the number of calls before and after the change to the number of calls before the change for a node or edge in the software model, and the calculation formula is: wherein represents the call frequency deviation, is the call frequency after the change, is the call frequency before the change; The impact strength is the proportion of the number of calls affected by the change in the node or edge call after the software change, and the calculation method is as follows: wherein is the number of calls affected by the change, is the total number of calls; The diameter is the maximum value of the shortest path length between any two nodes, and the diameter ratio is the ratio of the diameter of the software sub-network affected by the change to the diameter of the complete software network, and the calculation method is as follows: wherein, is the diameter of the complete software network, is the diameter of the software subnetwork of affected functions.
8. The large model based software change impact modeling and simulation analysis method of claim 2, wherein, The construction process of the prompt word in S13 step: The JSON format template defines the hierarchical structure of the software model, including the node array and edge array of the top layer, and the attribute fields of each element; The example code covers typical scenarios of five node types and five edge types, and each example is equipped with the corresponding model extraction result. The source code to be analyzed is organized in the form of a dictionary with file path as key and code text as value, supporting cross-file call relationship analysis.
9. The large model based software change impact modeling and simulation analysis method of claim 3, wherein, The change marking process in S21 step includes: Traverse all edges in the model before and after the change, and filter the edge pairs with the same start and end nodes; Compare the output variable expressions of the edge pairs one by one, and add the variables with different expression texts to the change point set; Compare the constraint condition expressions of the edge pairs, and if there is a difference, add all the output variables of the edge to the change point set; Write the change point set as an attribute to the corresponding edge object for query use during simulation traversal.