Association relation determination method and device, computer equipment, storage medium and computer program product
By constructing a unified knowledge graph and combining it with a large language model, the problem of inefficient determination of the relationship between software requirements and source code was solved, achieving efficient and accurate determination of the relationship and improving the traceability and impact analysis capabilities of software projects.
Patent Information
- Application Number
- CN202511158588.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies are inefficient in determining the relationship between software requirements and source code, lack a unified data structure and a global relational view, struggle to handle semantically complex or historically evolving scenarios, lack integration with large language models, and fail to meet the diverse and interactive tracing needs of developers.
By constructing a unified knowledge graph, including generating tracing links between requirement knowledge graphs based on the source code and requirement documents of the projects to be processed, generating the graph, and combining it with a large language model, the target result information corresponding to the query request is output.
It improves the efficiency of determining the relationship between software requirements and source code, enabling rapid identification of which specific functional point will affect which specific business requirement. It achieves fine-grained linking from requirements to the method level, enhancing the traceability and impact analysis capabilities of software projects.
Smart Images

Figure CN121029969A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of software engineering, and in particular to a correlation relationship determination method and device, computer equipment, a storage medium and a computer program product. BACKGROUND
[0002] With the continuous expansion of the scale and the improvement of the complexity of software systems, the correlation relationship between the requirement documents and the source code in the project is increasingly complex. In the traditional software development process, the tracking of requirements and code mainly relies on manual maintenance of requirement tracking matrices or manual annotation, which has problems such as low efficiency, easy omission, and difficulty in dynamic updating.
[0003] At present, the automatic requirement-code tracking and intelligent analysis based on structured knowledge attempt to use static code analysis, natural language processing, data mining and other methods to assist requirement tracking, but there are generally the following deficiencies: requirement and code knowledge modeling is fragmented, lacking a unified data structure and global correlation view; the tracking link is established depending on rules or shallow text matching, which is difficult to cope with semantic complexity or historical evolution scenarios; there is a lack of intelligent question and answer and automatic explanation capabilities combined with large language models, making it difficult to meet the diverse and interactive tracking needs of developers. Therefore, there is a problem of low efficiency in the process of determining the correlation relationship between software requirements and source code. SUMMARY
[0004] Therefore, it is necessary to provide a correlation relationship determination method, device, computer equipment, computer readable storage medium and computer program product to solve the technical problem of low efficiency in the process of determining the correlation relationship between software requirements and source code.
[0005] In a first aspect, the present application provides a correlation relationship determination method, comprising:
[0006] constructing a code knowledge graph based on the source code of a to-be-processed project, constructing a requirement knowledge graph based on the requirement document of the to-be-processed project, and mining historical data of the to-be-processed project to establish a tracking link between the code knowledge graph and the requirement knowledge graph, thereby generating a unified knowledge graph;
[0007] receiving a query request of a user, searching in the unified knowledge graph based on the query request, and obtaining structured context information;
[0008] based on the structured context information and a preset prompt template, obtaining enhanced prompt information, sending the enhanced prompt information to a large language model, and outputting target result information corresponding to the query request based on the large language model; the target result information includes the correlation relationship between the corresponding target requirement and the target source code.
[0009] In one embodiment, receiving a user's query request and retrieving structured context information from the unified knowledge graph based on the query request includes: obtaining anchor nodes in the unified knowledge graph based on the user's query request; the anchor nodes include at least one starting node in the unified knowledge graph; starting from the anchor node, traversing the first node in the unified knowledge graph based on tracing links to obtain a first traversal result, completing the retrieval in the unified knowledge graph, and obtaining the structured context information.
[0010] In one embodiment, obtaining enhanced prompt information based on the structured context information and a preset prompt template, sending the enhanced prompt information to a large language model, and outputting the target result information corresponding to the query request based on the large language model includes: obtaining a first node located in the code knowledge graph and a second node located in the requirement knowledge graph based on the structured context information and the preset prompt template; traversing the code knowledge graph and the requirement knowledge graph based on the query request, respectively, using the first node and the second node as starting points, to obtain a second traversal result and a third traversal result; obtaining the enhanced prompt information based on the first traversal result, the second traversal result, and the third traversal result, and sending the enhanced prompt information to the large language model, and outputting the target result information corresponding to the query request based on the large language model.
[0011] In one embodiment, before traversing the first node in the unified knowledge graph based on the tracking link, starting from the anchor node, to obtain the first traversal result, the method includes: extracting a requirement identifier from the historical data of the project to be processed; obtaining the requirement task information corresponding to the requirement identifier, and matching the requirement task information with a second node in the requirement knowledge graph to obtain a first matching result; obtaining the modified code entity in the historical data, and matching it with a third node in the code knowledge graph to obtain a second matching result; and establishing a tracking link between the code knowledge graph and the requirement knowledge graph based on the first matching result and the second matching result.
[0012] In one embodiment, the step of constructing a code knowledge graph based on the source code of the project to be processed includes: parsing the source code of the project to be processed to obtain code entities of a preset type, attribute information corresponding to the code entities, and dependency relationships between multiple code entities; mapping each code entity to multiple first nodes, and mapping the dependency relationships to multiple edges of the multiple first nodes to generate a code graph; and performing node pruning and pattern merging operations on the code graph to obtain the code knowledge graph.
[0013] In one embodiment, constructing a requirement knowledge graph based on the requirement document of the project to be processed includes: parsing the requirement document of the project to be processed to obtain structured semantic elements; mapping the structured semantic elements to a predefined standardized six-tuple structure to obtain a structured requirement set and a set of relationships corresponding to the structured requirement set; mapping the structured requirement set to multiple second nodes and the set of relationships to multiple edges of the multiple second nodes; and constructing the knowledge graph based on the multiple second nodes and the multiple edges of the multiple second nodes.
[0014] Secondly, this application also provides an apparatus for determining association relationships, comprising:
[0015] The knowledge graph construction module is used to construct a code knowledge graph based on the source code of the project to be processed, and a requirement knowledge graph based on the requirement document of the project to be processed; and to mine the historical data of the project to be processed to establish a tracing link between the code knowledge graph and the requirement knowledge graph, and generate a unified knowledge graph.
[0016] The information acquisition module is used to receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information;
[0017] The information determination module is used to obtain enhanced prompt information based on the structured context information and the preset prompt template, send the enhanced prompt information to the large language model, and output the target result information corresponding to the query request based on the large language model; the target result information includes the correlation between the corresponding target requirements and the target source code.
[0018] Thirdly, this application also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0019] A code knowledge graph is constructed based on the source code of the project to be processed, and a requirement knowledge graph is constructed based on the requirement document of the project to be processed; and historical data of the project to be processed is mined to establish a tracing link between the code knowledge graph and the requirement knowledge graph, thereby generating a unified knowledge graph;
[0020] Receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information;
[0021] Based on the structured context information and the preset prompt template, enhanced prompt information is obtained and sent to the large language model. Based on the large language model, the target result information corresponding to the query request is output; the target result information includes the relationship between the corresponding target requirement and the target source code.
[0022] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0023] A code knowledge graph is constructed based on the source code of the project to be processed, and a requirement knowledge graph is constructed based on the requirement document of the project to be processed; and historical data of the project to be processed is mined to establish a tracing link between the code knowledge graph and the requirement knowledge graph, thereby generating a unified knowledge graph;
[0024] Receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information;
[0025] Based on the structured context information and the preset prompt template, enhanced prompt information is obtained and sent to the large language model. Based on the large language model, the target result information corresponding to the query request is output; the target result information includes the relationship between the corresponding target requirement and the target source code.
[0026] Fifthly, this application also provides a computer program product, which includes a computer program that, when executed by a processor, performs the following steps:
[0027] A code knowledge graph is constructed based on the source code of the project to be processed, and a requirement knowledge graph is constructed based on the requirement document of the project to be processed; and historical data of the project to be processed is mined to establish a tracing link between the code knowledge graph and the requirement knowledge graph, thereby generating a unified knowledge graph;
[0028] Receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information;
[0029] Based on the structured context information and the preset prompt template, enhanced prompt information is obtained and sent to the large language model. Based on the large language model, the target result information corresponding to the query request is output; the target result information includes the relationship between the corresponding target requirement and the target source code.
[0030] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for determining association relationships first construct a code knowledge graph based on the source code of the project to be processed, and then construct a requirement knowledge graph based on the requirement documents of the project to be processed. Historical data of the project to be processed is then mined to establish tracing links between the code knowledge graph and the requirement knowledge graph, generating a unified knowledge graph. Next, user query requests are received, and a search is performed in the unified knowledge graph based on the query requests to obtain structured context information. Finally, based on the structured context information and a preset prompt template, enhanced prompt information is obtained and sent to a large language model. The large language model then outputs the target result information corresponding to the query request. The target result information contains the association relationship between the corresponding target requirement and the target source code. This improves the efficiency in determining the association relationship between software requirements and source code. In the above process, the unified knowledge graph can integrate isolated requirements, code and historical data, and establish fine-grained links from requirements to the method level. This makes the accuracy and depth of impact analysis far exceed existing technologies. It can not only quickly determine which "file" or "module" will be affected by modifying a certain piece of code, but also accurately locate which specific functional point of which specific business requirement will be affected, thus improving the efficiency of determining the relationship between software requirements and source code. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a diagram illustrating the application environment of the association determination method in one embodiment;
[0033] Figure 2 This is a flowchart illustrating a method for determining association relationships in one embodiment;
[0034] Figure 3 This is a flowchart illustrating the steps for determining the association relationship in one embodiment;
[0035] Figure 4 This is a flowchart illustrating the method for determining association relationships in another embodiment;
[0036] Figure 5 This is a flowchart illustrating the method for determining association relationships in yet another embodiment;
[0037] Figure 6 This is a structural block diagram of an association determination device in one embodiment;
[0038] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0040] As software systems grow in scale and complexity, the relationships between requirements documents and source code become increasingly complex. In traditional software development processes, requirement and code tracing relies heavily on manual maintenance of requirement tracing matrices or manual annotation, resulting in low efficiency, susceptibility to omissions, and difficulty in dynamic updates. In recent years, with the development of technologies such as knowledge graphs, natural language processing, and large language models, automated requirement-code tracing and intelligent analysis based on structured knowledge has become a research hotspot. While some studies have attempted to utilize static code analysis, natural language processing, and data mining to assist requirement tracing, they generally suffer from the following shortcomings: fragmented requirement and code knowledge modeling lacks a unified data structure and a global relational view; tracing link establishment relies on rules or shallow text matching, making it difficult to handle semantically complex or historically evolving scenarios; and a lack of intelligent question-answering and automated explanation capabilities combined with large language models fails to meet the diverse and interactive tracing needs of developers. Therefore, there is an urgent need for a technical solution that can automatically and structurally extract, model, and correlate requirement and code knowledge, and support intelligent retrieval and question-answering to improve the traceability, impact analysis, and knowledge reuse capabilities of software projects.
[0041] Existing technologies mainly include knowledge graph-based software requirement tracing systems and information retrieval (IR)-based requirement-code link recommendation methods. Typical existing technologies include: knowledge graph construction methods based on static code analysis and requirement document processing; these technologies typically use Abstract Syntax Trees (ASTs) or other static analysis tools to parse source code, extract code entities and their dependencies, and construct a code knowledge graph; simultaneously, they use natural language processing to extract elements from requirement documents, forming a requirement knowledge graph. Some studies attempt to establish a preliminary tracing relationship between the two through text similarity and rule matching, achieving semi-automatic retrieval and recommendation from requirements to code; and information retrieval (IR)-based methods; such as using vector space models, TF-IDF (Term Frequency-Inverse Document Frequency), word embedding, etc., to calculate the similarity between requirement descriptions and code comments, identifiers, etc., and automatically recommend relevant code fragments. However, all of these existing technologies suffer from inefficiency in determining the relationship between requirements and target source code.
[0042] To address the aforementioned technical problems, embodiments of this application provide a method for determining association relationships, which can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with the large language model 104 via a network. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc.
[0043] This application provides a method for determining association relationships, which can be applied to, for example... Figure 1 In the application environment shown, specifically, terminal 102 constructs a code knowledge graph based on the source code of the project to be processed and a requirement knowledge graph based on the requirement document of the project to be processed; it also mines historical data of the project to be processed to establish a tracing link between the code knowledge graph and the requirement knowledge graph, generating a unified knowledge graph; then terminal 102 receives the user's query request, searches the unified knowledge graph based on the query request, and obtains structured context information; finally, terminal 102 obtains enhanced prompt information based on the structured context information and a preset prompt template, sends the enhanced prompt information to the large language model 104, and outputs the target result information corresponding to the query request based on the large language model 104; the target result information contains the correlation between the corresponding target requirement and the target source code.
[0044] In one exemplary embodiment, such as Figure 2 As shown, a method for determining association relationships is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps 202 to 206. Wherein:
[0045] Step 202: Construct a code knowledge graph based on the source code of the project to be processed, and construct a requirement knowledge graph based on the requirement document of the project to be processed; and mine the historical data of the project to be processed to establish a tracing link between the code knowledge graph and the requirement knowledge graph to generate a unified knowledge graph.
[0046] Source code can be the text-based implementation code of a software system, including source files, resource files, scripts, build configurations, test code, and documentation comments in various languages; a code knowledge graph is a graph structure that organizes code-related entities and their relationships, with "code" as the main body, used to represent semantic information such as code structure, dependencies, changes, and evolution; a requirements document is a collection of documents that describe the functions, constraints, acceptance criteria, etc., that the system should have in natural language or structured form; a requirements knowledge graph is a knowledge graph built around requirements documents and related business concepts, used to express semantic relationships and tracking information between requirements and between requirements and implementation; historical data is historical information related to the project to be processed, usually from test systems or deployment records, etc.
[0047] Step 204: Receive the user's query request, and search the unified knowledge graph based on the query request to obtain structured context information.
[0048] Among them, the query request is a retrieval instruction initiated by the user, describing the specific content and scope of the relevant information that the user wishes to obtain from the unified knowledge graph; the structured context information is a set of structured information retrieved from the unified knowledge graph to support subsequent reasoning and generate target result information.
[0049] Specifically, query requests can cover a variety of tracing scenarios, such as: tracing from requirements to code ("Which code implements the 'encrypted storage of user passwords' requirement?"); tracing from code to requirements ("What business requirements are implemented in the LoginController.java file?"); impact analysis ("Which functional requirements will be affected by modifying the OrderService class?"), etc.
[0050] Step 206: Based on the structured context information and the preset prompt template, obtain the enhanced prompt information, send the enhanced prompt information to the large language model, and output the target result information corresponding to the query request based on the large language model; the target result information contains the relationship between the corresponding target requirements and the target source code.
[0051] The preset prompt templates can be a pre-designed set of templates used to transform structured contextual information into input suitable for understanding and reasoning by a large language model. The templates can include background information, target tasks, constraints, evidence chains, output format requirements, etc. The large language model is a language model based on large-scale pre-trained text and fine-tuned data, which has the ability to understand natural language, reason, and generate structured text. The target result information is the set of results output by the large language model based on structured contextual information and preset prompt templates. It can include the relationship between the target requirement and the target source code. The relationship can be used to describe the semantic and implementation mapping between the target requirement and the target source code, as well as their coherence and causal relationship in terms of evidence chains, change records, test results, etc.
[0052] The aforementioned method for determining relationships first constructs a code knowledge graph based on the source code of the project to be processed, and a requirement knowledge graph based on the requirement documents of the project to be processed. Then, historical data of the project to be processed is mined to establish tracing links between the code knowledge graph and the requirement knowledge graph, generating a unified knowledge graph. Next, user query requests are received, and a search is performed in the unified knowledge graph based on the query requests to obtain structured context information. Finally, based on the structured context information and a preset prompt template, enhanced prompt information is obtained and sent to a large language model. The large language model then outputs the target result information corresponding to the query request. The target result information contains the relationship between the corresponding target requirement and the target source code. This improves the efficiency in determining the relationship between software requirements and source code. In the above process, the unified knowledge graph can integrate isolated requirements, code and historical data, and establish fine-grained links from requirements to the method level. This makes the accuracy and depth of impact analysis far exceed existing technologies. It can not only quickly determine which "file" or "module" will be affected by modifying a certain piece of code, but also accurately locate which specific functional point of which specific business requirement will be affected, thus improving the efficiency of determining the relationship between software requirements and source code.
[0053] In one exemplary embodiment, a user's query request is received, and a search is performed in the unified knowledge graph based on the query request to obtain structured context information, including:
[0054] Based on the user's query request, anchor nodes in the unified knowledge graph are obtained; anchor nodes include at least one starting node in the unified knowledge graph; starting from the anchor node, the first node in the unified knowledge graph is traversed based on the tracing links to obtain the first traversal result, thus completing the retrieval in the unified knowledge graph and obtaining structured context information.
[0055] Anchor nodes are the starting reference points in the knowledge graph retrieval process. They are the core nodes corresponding to the user's query request in the knowledge graph and must contain at least one starting node. The starting node is the most basic and core node among the anchor nodes and is the initial starting point for traversing the unified knowledge graph. Tracing links refer to navigating from one node to another along the relationships between nodes in the knowledge graph, using predefined entity relationships in the knowledge graph. The first traversal result refers to the set formed after traversing from the anchor node to all first nodes by tracing links.
[0056] In this embodiment, a complete logical chain for knowledge graph query and information extraction is formed, from determining the anchor node, traversing the first node by following links, and then generating structured context information. Using the anchor node as the starting point, following links as the means, and structured context information as the final output, each step is closely linked to achieve efficient and accurate information retrieval within the unified knowledge graph.
[0057] Furthermore, in one embodiment, enhanced prompt information is obtained based on structured context information and a preset prompt template. This enhanced prompt information is then sent to a large language model, which outputs the target result information corresponding to the query request, including:
[0058] Step S302: Based on structured context information and a preset prompt template, obtain the first located node in the code knowledge graph and the second located node in the requirement knowledge graph; Step S304: Starting from the first located node and the second located node, respectively, traverse the code knowledge graph and the requirement knowledge graph based on the query request to obtain the second traversal result and the third traversal result; Step S306: Based on the first traversal result, the second traversal result, and the third traversal result, obtain the enhanced prompt information, send the enhanced prompt information to the large language model, and output the target result information corresponding to the query request based on the large language model.
[0059] The second traversal result refers to the set of associated information obtained by traversing the code knowledge graph based on the query request, starting from the first node already located in the code knowledge graph; the third traversal result refers to the set of associated information obtained by traversing the demand knowledge graph based on the query request, starting from the second node already located in the demand knowledge graph; the preset prompt template is a predefined structured text framework used to standardize information extraction or generation, which can contain fixed formats and variables to be filled; the enhanced prompt information is a comprehensive prompt content formed by integrating the first, second, and third traversal results and used to input the large language model; the target result information is the final result output by the large language model based on the enhanced prompt information, which directly corresponds to the query request.
[0060] In this embodiment, nodes in the code and the requirement domain are located by using a preset prompt template. Then, the code knowledge graph and the requirement knowledge graph are traversed to obtain the second and third traversal results. Finally, multi-source information is integrated to form enhanced prompt information, which is input into the large language model to obtain the target result. This process realizes the combination of the structured knowledge of the knowledge graph and the generation capability of the large language model, improving the accuracy and richness of the final result.
[0061] In an exemplary embodiment, before obtaining the first traversal result by traversing the first node in the unified knowledge graph based on the tracking links, starting from the anchor node, the process includes:
[0062] Step S402: Extract requirement identifiers from the historical data of the project to be processed; Step S404: Obtain the requirement task information corresponding to the requirement identifier, and match the requirement task information with the second node in the requirement knowledge graph to obtain the first matching result; Step S406: Obtain the modified code entities in the historical data, and match them with the third node in the code knowledge graph to obtain the second matching result; Step S408: Based on the first matching result and the second matching result, establish a tracking link between the code knowledge graph and the requirement knowledge graph.
[0063] The requirement identifier is a symbol, number, or string used to uniquely identify a requirement task; the requirement task information is the specific content corresponding to the requirement identifier, which may include detailed information such as the requirement description, priority, responsible person, creation time, associated scenarios, and acceptance criteria.
[0064] More specifically, historical data refers to various types of data generated and accumulated during the past development, operation, and maintenance of the project to be processed, including but not limited to requirement documents and modification logs; the first matching result is the matching result obtained by comparing and associating the requirement task information with the second node in the requirement knowledge graph, used to confirm the correspondence between the two; the second matching result is the matching result obtained by comparing and associating the "modified code entity" in the historical data with the third node in the code knowledge graph, used to confirm the correspondence between the two; the modified code entity refers to the code recorded in the historical data that has been adjusted, updated, or refactored during the project development or maintenance process.
[0065] In this embodiment, the requirement identifier and the modified code entity are retrieved from the historical data of the project to be processed and matched with the second node of the requirement knowledge graph and the third node of the code knowledge graph, respectively. Finally, a tracking link between the two is established based on the matching results. The above process realizes the accurate association between requirements and code and improves the efficiency of obtaining association results.
[0066] More specifically, in one embodiment, constructing a code knowledge graph based on the source code of the project to be processed includes: parsing the source code of the project to be processed to obtain code entities of a preset type, attribute information corresponding to the code entities, and dependencies between multiple code entities; mapping each code entity to multiple first nodes and mapping dependencies to multiple edges of the multiple first nodes to generate a code graph; and performing node pruning and pattern merging operations on the code graph to obtain the code knowledge graph.
[0067] Among them, the preset type of code entity refers to the type of code element to be extracted in advance before constructing the code knowledge graph; attribute information is detailed data describing the characteristics of the code entity; dependency relationship refers to the association relationship such as calling, referencing, inheritance, and inclusion between multiple code entities; the code graph is an intermediate product that represents code entities and their dependencies in a graphical structure, and is the foundation for constructing the code knowledge graph. Nodes correspond to preset type code entities, and edges correspond to the dependency relationships between code entities; node pruning operation refers to filtering and removing redundant, irrelevant, or secondary nodes in the code graph; and pattern merging operation refers to integrating and standardizing nodes or relationship patterns with similar structures and semantics in the code graph.
[0068] As an example, common types of code entities with default types include functions, classes, variables, and modules; the attributes of functions include parameter list, return type, module, and function description; the attributes of classes include parent class, member variables, and member methods.
[0069] In this embodiment, a preliminary code graph is formed by parsing the source code to extract preset code entities, attributes, and dependencies. Then, redundancy is removed by node pruning and the structure is unified by pattern merging, finally resulting in a standardized code knowledge graph. The above process realizes the transformation of code from unstructured text to structured knowledge, laying the foundation for the subsequent confirmation of relationships.
[0070] In one embodiment, a requirements knowledge graph is constructed based on the requirements document of the project to be processed, including:
[0071] Parse the requirements document of the project to be processed to obtain structured semantic elements; map the structured semantic elements to a predefined standardized six-tuple structure to obtain a structured set of requirements and a set of relationships corresponding to the structured set of requirements; map the structured set of requirements to multiple second nodes and the set of relationships to multiple edges of multiple second nodes; construct a knowledge graph based on the multiple second nodes and multiple edges of multiple second nodes.
[0072] Among them, the requirements document is a document that records information such as user requirements, functional goals, business logic, and constraints in the project to be processed; the structured semantic elements are core information units with clear meaning and logical structure parsed from the requirements document, which can be used to reflect the key content and attributes of the requirements; the standardized six-tuple structure is a predefined structured framework used to standardize the representation of requirements information, which consists of six dimensions of elements.
[0073] More specifically, the set of relationships refers to the logical connections between different requirement items in a structured set of requirements, such as dependency, inclusion, causality, etc., which are used to reflect the inherent connections between requirements.
[0074] In this embodiment, structured semantic elements are extracted by parsing the requirement document and mapped to standardized six-tuples to form a structured requirement set. At the same time, the set of relationships between requirements is sorted out. Then, the requirement set is transformed into a second node, and the relationships are transformed into edges between nodes. Finally, a requirement knowledge graph is constructed, realizing the transformation of requirement information from natural language description to structured knowledge.
[0075] This application provides a method for determining association relationships. To better understand the process of the above-described method for determining association relationships, in conjunction with... Figure 5 As shown below, the specific process of a method for determining the association relationship in this application is described in detail, including the following steps:
[0076] Step S502: Construct a code knowledge graph.
[0077] This step aims to parse the source code of the project to be processed into a structured code knowledge graph, which can be accomplished through a code knowledge extraction module, which includes a parsing and extraction unit, a code graph construction unit, and a code graph optimization unit.
[0078] First, the parsing and extraction unit receives a set of source code files as input. The unit performs deep syntax analysis on the source code files using a parser such as a Tree-Sitter parser or an Abstract Syntax Tree (AST) parser. During this process, code entities of preset types are extracted. These code entities can specifically include: projects, packages, classes, methods, variables, and external dependencies. Simultaneously, attribute information corresponding to each code entity is obtained, including at least: the entity's type, scope, or function signature. Furthermore, the dependencies between code entities are extracted, including at least: function call relationships, class inheritance relationships, interface implementation relationships, or variable reference relationships.
[0079] Then, the code graph construction unit receives the output information from the parsing and extraction unit. Each code entity is abstractly modeled as a node in the graph, and the dependencies between entities are abstractly modeled as directed edges connecting the corresponding nodes. By integrating the analysis results of all source code files, an initial code graph that reflects the entire code repository structure is constructed.
[0080] Finally, the initial code graph is optimized by the code graph optimization unit to improve graph quality and query efficiency. Specific optimization operations include, but are not limited to: performing node pruning to remove redundant nodes with no actual dependencies in the graph (e.g., uncalled private methods or unreferenced variables); and performing pattern merging to identify and merge specific code patterns that occur repeatedly in the graph, thereby simplifying the graph structure and reducing complexity.
[0081] Step S504: Construct a requirement knowledge graph.
[0082] This step aims to transform unstructured natural language requirement documents into structured requirement knowledge graphs, which can be accomplished through a requirement knowledge extraction module. This module includes a requirement text structure parsing unit, a requirement relationship extraction unit, and a requirement knowledge graph construction unit.
[0083] Specifically, the requirement text structure parsing unit receives a set of requirement documents or text fragments as input. First, a series of Natural Language Processing (NLP) analyses are performed on the requirement text, including word segmentation, part-of-speech tagging, named entity recognition, sentence component analysis, and dependency relation analysis. Then, key structured semantic elements are extracted from the analysis results using semantic parsing techniques. Further, these semantic elements are mapped to a predefined standardized six-tuple structure, represented as: Req=<agent,operation,input,output,constraint,event> In this unit, agent represents the executing entity, operation represents the core operation, input represents the input conditions, output represents the output result, constraint represents the constraint condition, and event represents the triggering event. The output of this unit is a structured set of requirements.
[0084] Then, the structured set of requirements is received as input by the requirement relationship extraction unit. By analyzing the semantic elements of single or multiple requirement six-tuples, and selectively combining requirement metadata (such as creation time, responsible person, etc.), logical relationships between requirements are identified and labeled. The types of relationships include: hierarchical relationships, for example, when a requirement Req_i is a detailed description of another requirement Req_a, a "refinement (isRefinedBy)" relationship is established from Req_i to Req_a; temporal dependency relationships, for example, when the triggering event of a requirement Req_a is the core operation of another requirement Req_j, a "precondition (hasPrecondition)" relationship is established from Req_j to Req_a; other logical relationships, including but not limited to blocking relationships, repetition relationships, verification relationships, or association relationships.
[0085] Finally, the structured requirement set and the extracted set of relationships are received through the requirement knowledge graph construction unit. Each structured requirement is modeled as a node in the knowledge graph, and the node's attribute information includes its corresponding six-tuple semantic elements. The relationships between requirements are modeled as directed edges connecting corresponding nodes, and the edge type is used to represent the specific relationship type. Finally, the generated requirement knowledge graph is optimized, including deduplication, merging redundant relationships, and detecting circular dependencies.
[0086] Step S506: Cross-graph joint modeling to generate a unified knowledge graph.
[0087] This step aims to integrate the Requirement Knowledge Graph (RKG) and Code Knowledge Graph (CKG) generated in the previous steps, and to establish a tracing link between the two by mining historical data, ultimately forming a unified knowledge graph.
[0088] Specifically, the first step is to acquire and parse historical development data: gain access to the project's version control system (such as Git) history and task management system (such as Jira, Trello). Traverse the project's Git commit history and use regular expressions or pattern matching techniques to parse each commit message to extract any requirement identifiers (such as REQ-123, TICKET-456) that may be contained within.
[0089] Then, the requirements are associated with code entities. If the submission information contains a requirement identifier, the corresponding task management system is queried through the application programming interface (API) to obtain the requirement task details corresponding to the identifier and match them with the requirement node in RKG. At the same time, the code diffs changed in this submission are parsed to identify the specific code entities that have been modified (down to the method or class level) and matched with the code nodes in CKG.
[0090] Next, establish tracing links: For each commit that is successfully associated with a requirement node, establish a new directed edge between the modified code node and the corresponding requirement node. This edge represents an "implements" relationship, which we call a tracing link. All tracing links form a set E_trace.
[0091] Finally, the UKG (Unified Knowledge Graph) is constructed and stored: the node set V_req of the RKG and the node set V_code of the CKG are merged to obtain the final node set V_ukg = V_req∪V_code of the UKG. The edge set E_req of the RKG, the edge set E_code of the CKG, and the tracing link set E_trace generated in the above steps are merged to obtain the final edge set E_ukg of the UKG = E_req∪E_code∪E_trace. Finally, the constructed UKG is stored in a graph database for subsequent queries.
[0092] Step S508: Search within the unified knowledge graph, construct enhanced hints, and generate answers.
[0093] When a user initiates a tracking query, the system executes an online interactive tracking process. This process is based on the Retrieval-Augmented Generation (RAG) framework.
[0094] Specifically, the system first receives and parses user queries. The system receives query requests input by the user through a natural language interface. These queries can cover various tracing scenarios, such as: tracing from requirements to code ("Which code implements the 'encrypted user password storage' requirement?"), tracing from code to requirements ("What business requirements are implemented in the LoginController.java file?"), or impact analysis ("Which functionalities will be affected by modifying the OrderService class?").
[0095] Then, a retrieval is performed within the unified knowledge graph (UKG) to extract the most relevant contextual information to the user's query, forming a subgraph. This includes locating anchor nodes: based on the user's query intent and entities, one or more starting nodes, i.e., anchor nodes, are located in the UKG. For example, for the query "'User password encrypted storage' requirement...", the system will locate the node representing "user password encrypted storage" in the requirement node section of the UKG; cross-graph traversal: starting from the anchor node, traversal is performed along pre-established tracing links (E_trace, i.e., implements edges). For example, starting from the "password encryption" requirement node, traversal can reach code nodes such as UserService.java::hashPassword that implement this requirement; context expansion: to provide a more comprehensive answer, the retrieval process will further expand the traversal within each knowledge graph. In the CKG part, starting from the located code nodes, traversal is performed along dependency edges such as calls or called by to obtain its upstream and downstream code context. In the RKG section, starting from the located requirement node, traverse along logical relationship edges such as refinement (isRefinedBy) or premise (hasPrecondition) to obtain its associated requirement context; Information summary: package all relevant nodes (including requirement and code entities) and their relationship edges collected in the above traversal process into a structured context information object.
[0096] Then, augmentation takes the structured context information retrieved in the above steps, organizes it, and injects it into a preset prompt template to build an information-rich and context-clear augmented prompt, which will be sent to a large language model (LLM).
[0097] Finally, in the generation phase, the large language model receives and processes the enhanced prompts built in the previous steps. Based on the precise and structured contextual information provided in the prompts, the model generates a logically clear, detailed, and human-language-compliant answer that directly and accurately responds to the user's original query (query request) from the previous steps. For example, the model will explicitly indicate what the implementation code for each requirement is, its specific call relationships, and its logical connections with other requirements.
[0098] Through the above embodiments, requirements, code, and historical development data can be effectively integrated to construct a comprehensive knowledge graph. Utilizing RAG technology, accurate, efficient, and intelligent bidirectional requirement-code tracing and impact analysis are achieved, significantly improving the efficiency and quality of software development and maintenance. Furthermore, the unified knowledge graph integrates previously isolated requirements, code, and historical data, establishing fine-grained links from requirements down to the method / class level. This makes the accuracy and depth of impact analysis far exceed existing technologies. Users can not only know which "file" or "module" a modification to a piece of code will affect, but also precisely pinpoint which specific functional point of which specific business requirement will be affected. Conversely, when a business requirement is changed, the system can clearly list all underlying code units and their dependencies that need to be modified or verified. This bidirectional, multi-hop, and fine-grained analytical capability provides unprecedented decision support for software change risk assessment, test case design, and code refactoring.
[0099] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0100] Based on the same inventive concept, this application also provides an association relationship determination apparatus for implementing the association relationship determination method described above. The solution provided by this apparatus is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the association relationship determination apparatus provided below can be found in the limitations of the association relationship determination method described above, and will not be repeated here.
[0101] In one exemplary embodiment, such as Figure 6 As shown, a device for determining association relationships is provided, including a map construction module 601, an information acquisition module 602, and an information determination module 603, wherein:
[0102] The knowledge graph construction module 601 is used to build a code knowledge graph based on the source code of the project to be processed, and a requirement knowledge graph based on the requirement document of the project to be processed; and to mine the historical data of the project to be processed to establish a tracing link between the code knowledge graph and the requirement knowledge graph, and generate a unified knowledge graph.
[0103] The information acquisition module 602 is used to receive user query requests, and to search in the unified knowledge graph based on the query requests to obtain structured context information.
[0104] The information determination module 603 is used to obtain enhanced prompt information based on structured context information and preset prompt templates, send the enhanced prompt information to the large language model, and output the target result information corresponding to the query request based on the large language model; the target result information contains the relationship between the corresponding target requirements and the target source code.
[0105] Furthermore, in one embodiment, the information acquisition module 602 is also used to acquire anchor nodes in the unified knowledge graph based on the user's query request; the anchor nodes include at least one starting node in the unified knowledge graph; starting from the anchor nodes, the first node in the unified knowledge graph is traversed based on the tracing links to obtain the first traversal result, thereby completing the retrieval in the unified knowledge graph and obtaining structured context information.
[0106] Furthermore, in one embodiment, the information acquisition module 602 is also used to acquire the first node located in the code knowledge graph and the second node located in the requirement knowledge graph based on the structured context information and the preset prompt template; taking the first node and the second node located as starting points respectively, traversing the code knowledge graph and the requirement knowledge graph based on the query request to obtain the second traversal result and the third traversal result; obtaining enhanced prompt information based on the first traversal result, the second traversal result and the third traversal result, and sending the enhanced prompt information to the large language model, and outputting the target result information corresponding to the query request based on the large language model.
[0107] Furthermore, in one embodiment, the information acquisition module 602 is also used to extract requirement identifiers from the historical data of the project to be processed; obtain requirement task information corresponding to the requirement identifiers, and match the requirement task information with the second node in the requirement knowledge graph to obtain a first matching result; obtain modified code entities in the historical data, and match them with the third node in the code knowledge graph to obtain a second matching result; and establish a tracking link between the code knowledge graph and the requirement knowledge graph based on the first matching result and the second matching result.
[0108] Furthermore, in one embodiment, the graph construction module 601 is also used to parse the source code of the project to be processed, obtain code entities of a preset type, attribute information corresponding to the code entities, and dependency relationships between multiple code entities; map each code entity to multiple first nodes, and map the dependency relationships to multiple edges of the multiple first nodes to generate a code graph; perform node pruning and pattern merging operations on the code graph to obtain a code knowledge graph.
[0109] Furthermore, in one embodiment, the graph construction module 601 is also used to parse the requirement document of the project to be processed to obtain structured semantic elements; map the structured semantic elements to a predefined standardized six-tuple structure to obtain a structured requirement set and a set of relationships corresponding to the structured requirement set; map the structured requirement set to multiple second nodes and the set of relationships to multiple edges of the multiple second nodes; and construct a knowledge graph based on the multiple second nodes and the multiple edges of the multiple second nodes.
[0110] The modules in the aforementioned relationship determination device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0111] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data for determining association relationships. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a method for determining association relationships.
[0112] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0113] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0114] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0115] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0117] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0118] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0119] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for determining association relationships, characterized in that, The method includes: A code knowledge graph is constructed based on the source code of the project to be processed, and a requirement knowledge graph is constructed based on the requirement document of the project to be processed; and historical data of the project to be processed is mined to establish a tracing link between the code knowledge graph and the requirement knowledge graph, thereby generating a unified knowledge graph; Receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information; Based on the structured context information and the preset prompt template, enhanced prompt information is obtained and sent to the large language model. Based on the large language model, the target result information corresponding to the query request is output; the target result information includes the relationship between the corresponding target requirement and the target source code.
2. The method according to claim 1, characterized in that, The process of receiving a user's query request and then searching the unified knowledge graph based on the query request to obtain structured context information includes: Based on the user's query request, anchor nodes in the unified knowledge graph are obtained; the anchor nodes include at least one starting node in the unified knowledge graph. Starting from the anchor node, the first node in the unified knowledge graph is traversed based on the tracing links to obtain the first traversal result, thus completing the retrieval in the unified knowledge graph and obtaining the structured context information.
3. The method according to claim 2, characterized in that, The enhanced prompt information is obtained based on the structured context information and the preset prompt template. This enhanced prompt information is then sent to the large language model, and the target result information corresponding to the query request is output based on the large language model, including: Based on the structured context information and the preset prompt template, the first node located in the code knowledge graph and the second node located in the requirement knowledge graph are obtained; Starting from the first and second located nodes respectively, the code knowledge graph and the demand knowledge graph are traversed based on the query request to obtain the second and third traversal results. Based on the first traversal result, the second traversal result, and the third traversal result, the enhanced prompt information is obtained and sent to the large language model. Based on the large language model, the target result information corresponding to the query request is output.
4. The method according to claim 2, characterized in that, Before obtaining the first traversal result by traversing the first node in the unified knowledge graph based on the tracking links, starting from the anchor node, the process includes: Extract the demand identifier from the historical data of the project to be processed; Obtain the requirement task information corresponding to the requirement identifier, and match the requirement task information with the second node in the requirement knowledge graph to obtain the first matching result; Obtain the modified code entities from the historical data and match them with the third node in the code knowledge graph to obtain a second matching result; Based on the first matching result and the second matching result, a tracing link is established between the code knowledge graph and the requirement knowledge graph.
5. The method according to claim 1, characterized in that, The construction of the code knowledge graph based on the source code of the project to be processed includes: Parse the source code of the project to be processed to obtain code entities of a preset type, attribute information corresponding to the code entities, and dependency relationships between multiple code entities; Each of the code entities is mapped to multiple first nodes, and the dependency relationship is mapped to multiple edges of the multiple first nodes to generate a code graph; Perform node pruning and pattern merging operations on the code graph to obtain the code knowledge graph.
6. The method according to claim 1, characterized in that, The construction of a requirements knowledge graph based on the requirements document of the project to be processed includes: Parse the requirements document of the project to be processed to obtain structured semantic elements; The structured semantic elements are mapped to a predefined standardized six-tuple structure to obtain a structured set of requirements and a set of association relationships corresponding to the structured set of requirements; The structured set of requirements is mapped to multiple second nodes, and the set of relationships is mapped to multiple edges of the multiple second nodes; The knowledge graph is constructed based on multiple second nodes and multiple edges of the second nodes.
7. A device for determining association relationships, characterized in that, The device includes: The knowledge graph construction module is used to construct a code knowledge graph based on the source code of the project to be processed, and a requirement knowledge graph based on the requirement document of the project to be processed; and to mine the historical data of the project to be processed to establish a tracing link between the code knowledge graph and the requirement knowledge graph, and generate a unified knowledge graph. The information acquisition module is used to receive user query requests, and search the unified knowledge graph based on the query requests to obtain structured context information; The information determination module is used to obtain enhanced prompt information based on the structured context information and the preset prompt template, send the enhanced prompt information to the large language model, and output the target result information corresponding to the query request based on the large language model; the target result information includes the correlation between the corresponding target requirements and the target source code.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Cited By
Code retrieval method and device and related equipment
CN121277885A
Software tracing method and system based on multi-agent collaborative decision
CN121680790A
Software traceability method and system based on multi-agent collaborative decision-making
CN121680790B
Code data management method based on artificial intelligence and related device
CN122018956A
Cross-service system intelligent question answering and process automation method and system
CN122153017A