Code Management System and Method Based on Software Requirement Design

By building a code knowledge graph warehouse and deep learning model, the complexity and efficiency problems of code management in the existing technology are solved, and efficient reuse and recommendation of code fragments in different environments are achieved, and development efficiency is improved.

CN120162034BActive Publication Date: 2025-07-29湖南砺石科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510641163.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-29
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing code management methods rely on the experience of developers, and are difficult to meet the needs of complex software systems, cannot effectively reduce repetitive work, and lack systematic code reuse, migration and stability evaluation, which affects development efficiency and software maintenance.

Method used

By extracting the transferable index of code snippets, a code knowledge graph warehouse is built, and the deep learning model of the Transformer architecture monitors the key token characteristics of code snippets in real time, predicting and recommending the next code snippet to be written.

Benefits of technology

It realizes structured management of code snippets, improves the accuracy and applicability of code recommendations, ensures compatibility and stability in different programming languages, software frameworks and execution environments, reduces repetitive work for developers, and improves coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162034B_ABST
    Figure CN120162034B_ABST
Patent Text Reader

Abstract

The present invention discloses a code management method based on software requirement design, which relates to the field of computer technology and includes: extracting a number of target code segments from software code documents according to the migration index, and performing ontology library design of a knowledge graph for the target code segments; constructing a code knowledge graph warehouse for managing the target code based on the obtained ontology library after design; the background monitors the key Token features of the code segments currently written by the developer in real time and inputs them into a deep learning model of the Transformer architecture to predict the functional Token of the next code segment to be written; inputting the functional Token of the next code segment to be written into the code knowledge graph warehouse to feedback and output the recommended code segments of the next code segment to be written; the present invention is beneficial to reducing the repetitive work of developers and improving the coding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computers, and specifically relates to a code management system and method based on software requirement design. Background Art

[0002] With the continuous growth of the complexity and scale of software systems, the reuse and management of code have become key factors affecting software development efficiency and quality; in the actual development process, developers need to share code snippets between different modules, or expand and optimize based on existing code to improve development efficiency; however, the existing code management methods mainly rely on the experience of developers, and reuse code through manual search, copy and paste, etc., which are difficult to meet the increasingly complex requirements of software systems and cannot effectively reduce the repetitive work of developers; at the same time, in large software projects, there is a lack of systematic evaluation criteria for code reuse rate, portability and stability, resulting in an increase in the complexity of code management, affecting development efficiency and software maintainability; while the traditional method of retrieving and recommending code snippets based on static matching of code texts fails to effectively measure the portability and reuse value of code snippets in different programming languages, software frameworks and execution environments, affecting the applicability and practical feasibility of code in different scenarios. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a code management method based on software requirement design.

[0004] To achieve the above object, the present invention provides a code management method based on software requirement design, including:

[0005] Step 1: Extract a number of target code snippets from the software code document according to the migration index, and design the ontology library of the knowledge graph for the target code snippets;

[0006] Step 2: Construct a code knowledge graph warehouse for managing the target code based on the ontology library obtained after design;

[0007] Step 3: During the software development stage, the background monitors the key Token features of the code snippets currently written by the developer in real time, and inputs them into the deep learning model of the Transformer architecture to predict the functional Token of the next code snippet to be written;

[0008] Step 4: Input the functional Token of the next code snippet to be written into the code knowledge graph warehouse to feedback and output the recommended code snippet of the next code snippet to be written.

[0009] Preferably, before extracting a number of target code snippets from the software code document according to the migration index, it includes:

[0010] In each software code document, the start and end positions of multiple groups of code are parsed according to the abstract syntax tree;

[0011] The code between the start and end positions of each group of code is used as a code snippet, and multiple code snippets are obtained;

[0012] Calculate the recurrence rate of each code snippet in all software code documents;

[0013] Mark the code snippets with a recurrence rate greater than or equal to the preset recurrence rate threshold as candidate code snippets, and mark the code snippets with a recurrence rate less than the preset recurrence rate threshold as non-candidate code snippets.

[0014] Preferably, the calculating the recurrence rate of each code snippet in all software code documents includes:

[0015] In each software code document, each code snippet is transformed into a code hash fingerprint by using a hash function, and a number of code hash fingerprints are obtained;

[0016] Select any one of the code hash fingerprints among a number of code hash fingerprints as the first hash fingerprint, and use the remaining code hash fingerprints among the number of code hash fingerprints as the second hash fingerprints, and multiple second hash fingerprints are obtained;

[0017] Calculate the fingerprint similarity coefficient between the first hash fingerprint and each second hash fingerprint, and obtain multiple fingerprint similarity coefficients for the first hash fingerprint;

[0018] Wherein, the calculation method of the fingerprint similarity coefficient is as follows:

[0019] ;

[0020] In the formula: represents the fingerprint similarity coefficient, represents the i-th hash value in the first hash fingerprint, represents the i-th hash value in the second hash fingerprint, represents the dimension of the hash fingerprint, represents the maximum value function, represents a constant;

[0021] Compare each fingerprint similarity coefficient with the fingerprint similarity coefficient threshold, and count the number of fingerprint similarity coefficients greater than or equal to the fingerprint similarity coefficient threshold to obtain the recurrence number of the first hash fingerprint;

[0022] Calculate the ratio of the recurrence number of the first hash fingerprint to the number of all code hash fingerprints to obtain the recurrence rate of the code snippet corresponding to the first hash fingerprint.

[0023] Preferably, the method for obtaining the transferability index is as follows:

[0024] In the testing phase, capture the running success rate of each candidate code snippet in different test environments;

[0025] Calculate the average value of the running success rates in each test environment to obtain the average running success rate of each candidate code snippet in different test environments;

[0026] Take the average running success rate as the transferability index of the candidate code snippet to obtain the transferability index of each candidate code snippet.

[0027] Preferably, extracting several target code snippets from the software code document according to the transferability index includes:

[0028] Compare the transferability index of each candidate code snippet with a preset transferability index threshold;

[0029] If the transferability index is greater than or equal to the transferability index threshold, mark the corresponding candidate code snippet as a target code snippet;

[0030] If the transferability index is less than the transferability index threshold, mark the corresponding candidate code snippet as a non-target code snippet;

[0031] Repeat the above steps until the comparison of all candidate code snippets is completed to obtain several target code snippets.

[0032] Preferably, the ontology library design of the knowledge graph for the target code snippets includes:

[0033] Obtain all the target code snippets required for constructing the knowledge graph, and take each target code snippet as a first target entity to obtain several first target entities;

[0034] Extract the header files of each target code snippet through a static analysis tool, and map the function Tokens of the target code snippet according to a predefined library function table, and take the target code snippet as a second target entity;

[0035] Combine any first target entity and second target entity into an entity pair;

[0036] Traverse each entity pair into a preset relationship library, and match a predefined inter-entity relationship for each entity pair to obtain multiple inter-entity relationships;

[0037] Take all the first target entities, second target entities, and inter-entity relationships as elements in the ontology library, and design the ontology library.

[0038] Preferably, the method for constructing the code knowledge graph repository is as follows:

[0039] Take the first target entity and the second target entity in the ontology library as nodes in the knowledge graph, and take the predefined relationships between entities as the edges connecting the nodes in the associated knowledge graph to form the basic structure of the knowledge graph;

[0040] Import the basic structure into the Neo4j database for storage to obtain a code knowledge graph repository for managing target code.

[0041] Preferably, the key Token features include multiple structural Tokens and non-structural Tokens. The structural Tokens include function names, class names, loop types, and conditional statement types; the non-structural Tokens include API call types, file operation types, and database operation types;

[0042] Among them, the generation method of the deep learning model with the Transformer architecture is as follows:

[0043] Obtain historical code function training data, and divide the historical code function training data into a code function training set and a code function test set. The historical code function training data includes the key Token features of code snippets and their corresponding function Tokens;

[0044] Construct a regression network with the Transformer architecture. Take the key Token features in the code function training set as the input of the regression network, and take the function Tokens as the output of the regression network, and train the regression network to obtain the original regression network;

[0045] Use the code function test set to verify the original regression network model, and output the original regression network with an error greater than or equal to the preset test error threshold as the deep learning model with the Transformer architecture.

[0046] A code management system based on software requirement design is implemented based on the above-mentioned code management method based on software requirement design, and includes:

[0047] An ontology design module for extracting a number of target code snippets from software code documents according to the transferability index, and performing ontology library design of the knowledge graph for the target code snippets;

[0048] A graph construction module for constructing a code knowledge graph repository for managing target code based on the ontology library obtained after design;

[0049] A model prediction module, which is used to monitor the key Token features of the code snippets currently written by developers in real time in the background during the software development stage, and input them into a deep learning model with a Transformer architecture to predict the functional Tokens of the next code snippet to be written.

[0050] An extended recommendation module, which is used to input the functional Tokens of the next code snippet to be written into the code knowledge graph repository to feedback and output the recommended code snippets of the next code snippet to be written.

[0051] An electronic device includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the code management method based on software requirements designed above.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] The present invention structurally manages code snippets based on a code knowledge graph, enabling the storage, retrieval, and recommendation of code snippets to no longer rely solely on keyword matching, but rather combining the structural information, functional descriptions, and their context information of the code to achieve an accurate mapping between code snippets and functional requirements, improving the accuracy and applicability of code recommendations; secondly, by calculating the code transferability index, the compatibility and stability of code snippets are evaluated under different programming languages, software frameworks, and execution environments to ensure the applicability of the code in cross-language, cross-framework, and cross-platform environments, improving the reuse value of code snippets; furthermore, static analysis is used to extract the key Token information of the code, combined with a deep learning model based on the Transformer structure, to analyze the structural and functional characteristics of the code snippets, predict the expansion requirements of the code, and provide a code completion solution that conforms to the development scenario based on the code knowledge graph; enabling the system to automatically identify the context of the current code when developers write code, intelligently predict the next code content, reduce the repetitive work of developers, and improve the coding efficiency. Description of the Drawings

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic flowchart of the method of the present invention;

[0056] Figure 2 It is a schematic structural diagram of the system of the present invention;

[0057] Figure 3 It is a schematic structural diagram of the electronic device of the present invention. Specific embodiments

[0058] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] Please refer to Figure 1 , an embodiment of the first aspect of the present invention provides a code management method based on software requirement design, including:

[0060] Step 1: Extract a number of target code fragments from the software code document according to the migration index, and design the ontology library of the knowledge graph for the target code fragments;

[0061] It should be noted that: there are several software code documents (such as API documents, technical descriptions, source code files, and open-source projects, etc.) collected in advance by technicians stored in the system database, and each software code document contains the complete source code of a software program;

[0062] In implementation, before extracting a number of target code fragments from the software code document according to the migration index, it includes:

[0063] In each software code document, parse the start and end positions of multiple groups of code according to the Abstract Syntax Tree (AST);

[0064] It should be understood that: the Abstract Syntax Tree (AST) is a structured representation of the source code, which converts the code into a tree structure, where each node represents the syntax components of the code, such as function definitions, variable declarations, operators, control statements, etc.; AST is widely used in compilers, interpreters, code analysis tools, and can be used for tasks such as code parsing, optimization, transformation, and static analysis tools; in the task of parsing the start and end positions of the code, each code fragment (such as functions, loops, conditional judgments, etc.) will be parsed into a node in the AST tree, and its start line number, column number, and end line number will be recorded;

[0065] Take the code between the start and end positions of each group of code as a code fragment to obtain multiple code fragments;

[0066] Calculate the recurrence rate of each code fragment in all software code documents;

[0067] Among them, calculating the recurrence rate of each code snippet in all software code documents includes:

[0068] In each software code document, use a hash function to convert each code snippet into a code hash fingerprint, obtaining a number of code hash fingerprints;

[0069] It should be understood that: A hash function is a mathematical algorithm that can map an input of any length to an output of a fixed length (hash value). The hash function includes but is not limited to one of MD5, SHA-256, or SHA-1, etc.;

[0070] Select any one of the number of code hash fingerprints as the first hash fingerprint, and use the remaining code hash fingerprints among the number of code hash fingerprints as the second hash fingerprints, obtaining a number of second hash fingerprints;

[0071] Calculate the fingerprint similarity coefficient between the first hash fingerprint and each second hash fingerprint, obtaining a number of fingerprint similarity coefficients for the first hash fingerprint;

[0072] Specifically, the calculation method of the fingerprint similarity coefficient is as follows:

[0073] ;

[0074] In the formula: represents the fingerprint similarity coefficient, represents the i-th hash value in the first hash fingerprint, represents the i-th hash value in the second hash fingerprint, represents the dimension of the hash fingerprint, represents the maximum value function, represents a constant;

[0075] Compare each fingerprint similarity coefficient with the fingerprint similarity coefficient threshold, and count the number of fingerprint similarity coefficients greater than or equal to the fingerprint similarity coefficient threshold, obtaining the recurrence count of the first hash fingerprint;

[0076] Calculate the ratio of the recurrence count of the first hash fingerprint to the number of all code hash fingerprints, obtaining the recurrence rate of the code snippet corresponding to the first hash fingerprint;

[0077] Mark the code snippets with a recurrence rate greater than or equal to the preset recurrence rate threshold as candidate code snippets, and mark the code snippets with a recurrence rate less than the preset recurrence rate threshold as non-candidate code snippets;

[0078] In implementation, the method for obtaining the transferability index is as follows:

[0079] In the test phase, capture the running success rate of each candidate code snippet in different test environments;

[0080] It should be noted that the different test environments include, but are not limited to, cross - language compatibility testing, cross - framework adaptability testing, cross - execution - environment adaptability testing, resource adaptability testing, and dependency compatibility testing, etc. (see Table 1 below).

[0081] Table 1: Example data table of test environments

[0082]

[0083] Calculate the mean of the running success rates in each test environment to obtain the average running success rate of each candidate code snippet in different test environments.

[0084] Among them, the specific calculation formula is as follows: ; In the formula: represents the transferability index, is the total number of test environments, is the running success rate of the i - th test environment;

[0085] Take the average running success rate as the transferability index of the candidate code snippet to obtain the transferability index of each candidate code snippet.

[0086] It should be understood that the transferability index is used to measure the generality and stability of the code snippet to determine whether a certain code snippet is suitable as a reusable high - quality code snippet. In other words, it is to use the transferability index to judge the usability of the code snippet in actual applications, that is, whether the code snippet can be reused in different software development scenarios, different technical frameworks, and different business requirements, and whether it can run stably in the actual environment, rather than being limited to a single project or a specific environment.

[0087] Specifically, extracting several target code snippets from the software code document according to the transferability index includes:

[0088] Compare the transferability index of each candidate code snippet with a preset transferability - index threshold.

[0089] If the transferability index is greater than or equal to the transferability - index threshold, mark the corresponding candidate code snippet as a target code snippet;

[0090] If the transferability index is less than the transferability - index threshold, mark the corresponding candidate code snippet as a non - target code snippet;

[0091] Repeat the above steps until the comparison of all candidate code snippets is completed to obtain several target code snippets;

[0092] Exemplarily, if the migration index threshold of a candidate code snippet is 0.5, and the migration index < 0.5, then this candidate code snippet is not recommended as the target code snippet, which means that this candidate code snippet is not suitable for generalization (reuse). If it is applied in actual situations, it may instead affect the efficiency of software writing;

[0093] In implementation, the ontology library design for the target code snippet includes:

[0094] Obtain all target code snippets required for constructing the knowledge graph, and take each target code snippet as the first target entity to obtain a number of first target entities;

[0095] Extract the header file of each target code snippet through a static analysis tool, and map the function Token of this target code snippet according to a predefined library function table, and take the target code snippet as the second target entity;

[0096] It should be understood that: without running the code, the static analysis tool can parse the structure of the source code and analyze the header file of the code snippet. The static analysis tool includes but is not limited to Clang Static Analyzer, libclang, and GCC-MM, etc.; the library function table is a mapping table used to correspond the library (Library) or header file (HeaderFile) in the programming language with its main function (function Token); to help quickly infer the function Token of the code snippet through the import statement or header file (#include) without having to parse the entire code logic;

[0097] Combine any first target entity and second target entity into an entity pair;

[0098] Traverse each entity pair in a preset relationship library, match a predefined relationship between entities for each entity pair, and obtain multiple relationships between entities;

[0099] It should be noted that: in the relationship library, technicians have predefined the mapping relationships between several relationships between entities (such as call relationship, data flow relationship, and dependency relationship, etc.) and entity pairs. In each mapping relationship, there is only one entity pair and its corresponding relationship between entities. The entity pair and its corresponding relationship between entities are artificially associated and bound in advance by technicians; it is worth noting that when the traversal result input of the relationship library is empty, there is no relationship between entities for this entity pair;

[0100] Take all the first target entities, second target entities, and relationships between entities as elements in the ontology library, and design the ontology library.

[0101] Step 2: Based on the ontology library obtained after design, construct a code knowledge graph repository for managing target code;

[0102] In implementation, the construction method of the code knowledge graph repository is as follows:

[0103] Take the first target entity and the second target entity in the ontology library as nodes in the knowledge graph, and take the predefined relationships between entities as the edges connecting the nodes in the knowledge graph to form the basic structure of the knowledge graph;

[0104] Import the basic structure into the Neo4j database for storage to obtain a code knowledge graph repository for managing target code.

[0105] Step 3: In the software development stage, the background monitors the key Token features of the code snippets currently written by the developer in real time and inputs them into the deep learning model of the Transformer architecture to predict the functional Token of the next code snippet to be written;

[0106] Among them, the key Token features include multiple structural Tokens and non-structural Tokens. The structural Tokens include but are not limited to function names, class names, loop types, and conditional statement types, etc.; the non-structural Tokens include but are not limited to API call types (such as "HTTP request", "network connection"), file operation types (such as "file reading", "file copying"), and database operation types (such as "transaction rollback", "transaction commit");

[0107] It should be understood that: the structural Tokens and non-structural Tokens are collected based on existing parsing tools, including but not limited to Tree-sitter, libclang, and Clang Static Analyzer, etc.;

[0108] In implementation, the generation method of the deep learning model of the Transformer architecture is as follows:

[0109] Obtain historical code function training data, and divide the historical code function training data into a code function training set and a code function test set. The historical code function training data includes the key Token features of the code snippets and their corresponding functional Tokens;

[0110] It should be understood that: the key Token features and their corresponding functional Tokens in the historical code function training data are actually collected and recorded by technical personnel according to historical experimental data;

[0111] Construct a regression network with a Transformer architecture, using the key Token features in the code function training set as the input of the regression network, and the function Token as the output of the regression network, and train the regression network to obtain the original regression network;

[0112] Use the code function test set to verify the model of the original regression network, and output the original regression network with an error greater than or equal to the preset test error threshold as the deep learning model with a Transformer architecture;

[0113] It can be understood that the key Tokens (structural Tokens + non-structural Tokens) of the previous code snippet can reflect the function of the previous code snippet, and because the code is in accordance with the logic, the Tokens of the previous code snippet can reflect the state of the current code, and the next code is usually a natural continuation of the previous code; this enables the prediction of the function Tokens of the next code snippet based on the structural and non-structural Tokens of the previous code snippet.

[0114] Step 4: Input the function Tokens of the next code snippet to be written into the code knowledge graph repository to feedback and output the recommended code snippets for the next code snippet to be written;

[0115] It should be noted that there may be multiple recommended code snippets output by the code knowledge graph repository. Therefore, the next code snippet to be written is determined manually by the software writer from multiple recommended code snippets.

[0116] Please refer to Figure 2 , based on the same inventive concept, the second aspect embodiment of the present invention provides a code management system based on software requirement design. For the details not described in this embodiment, please refer to the relevant parts in Embodiment 1. The system includes:

[0117] An ontology design module, configured to extract a number of target code snippets from the software code document according to the transferable index, and perform ontology library design of the knowledge graph on the target code snippets;

[0118] A graph construction module, configured to construct a code knowledge graph repository for managing target codes according to the ontology library obtained after design;

[0119] A model prediction module, configured to, during the software development stage, monitor in real time the key Token features of the code snippet currently written by the developer in the background, and input them into the deep learning model with a Transformer architecture to predict the function Tokens of the next code snippet to be written;

[0120] An extended recommendation module is used to input the function tokens of the next code snippet to be written into the code knowledge graph repository, so as to feedback and output the recommended code snippet of the next code snippet to be written.

[0121] Please refer to Figure 3 , in the third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the code management method based on software requirements provided by any one of the above methods.

[0122] Since the electronic device introduced in the content of this embodiment is the electronic device used to implement a code management method based on software requirements in the embodiments of the present application, based on a code management method based on software requirements introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device used in a code management method based on software requirements in the embodiments of the present application, it falls within the scope of protection of the present application.

[0123] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on the computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired network or a wireless network. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0124] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only one way, and in actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0127] Some of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to get a formula closest to the actual situation; the preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0128] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A code management method based on software requirement design, characterized in that Including: Step 1: Extract several target code snippets from the software code document according to the transferability index, and design the ontology library of the knowledge graph for the target code snippets; Among them, before extracting several target code snippets from the software code document according to the transferability index, it includes: In each software code document, parse the start and end positions of multiple groups of code according to the abstract syntax tree; Take the code between the start and end positions of each group of code as a code snippet, and obtain multiple code snippets; Calculate the recurrence rate of each code snippet in all software code documents; Mark the code snippets with a recurrence rate greater than or equal to the preset recurrence rate threshold as candidate code snippets, and mark the code snippets with a recurrence rate less than the preset recurrence rate threshold as non-candidate code snippets; Among them, the method for obtaining the transferability index is as follows: In the test phase, capture the running success rate of each candidate code snippet in different test environments; the different test environments include cross-language compatibility test, cross-framework adaptability test, cross-execution environment adaptability test, resource adaptability test, and dependency compatibility test; Calculate the average value of the running success rate in each test environment to obtain the average running success rate of each candidate code snippet in different test environments; Take the average running success rate as the transferability index of the candidate code snippet to obtain the transferability index of each candidate code snippet; Among them, the design of the ontology library of the knowledge graph for the target code snippets includes: Obtain all target code snippets required to build the knowledge graph, and take each target code snippet as the first target entity to obtain several first target entities; Extract the header file of each target code snippet through a static analysis tool, and map the function Token of the target code snippet according to the predefined library function table, and take the target code snippet as the second target entity; Combine any first target entity and second target entity into an entity pair; Traverse each entity pair in the preset relationship library, and match the predefined relationship between entities for each entity pair to obtain multiple relationships between entities; Take all the first target entities, second target entities, and relationships between entities as elements in the ontology library, and design the ontology library; Step 2: Build a code knowledge graph warehouse for managing target code based on the designed ontology library; Step 3: In the software development phase, the background monitors the key Token features of the code snippet currently written by the developer in real time, and inputs them into the deep learning model of the Transformer architecture to predict the function Token of the next code snippet to be written; Step 4: Input the function Token of the next code snippet to be written into the code knowledge graph warehouse to feedback and output the recommended code snippet of the next code snippet to be written.

2. The code management method based on software requirement design according to claim 1, characterized in that, The calculation of the recurrence rate of each code snippet in all software code documents includes: In each software code document, use the hash function to convert each code snippet into a code hash fingerprint to obtain several code hash fingerprints; Select any one of several code hash fingerprints as the first hash fingerprint, and use the remaining code hash fingerprints among the several code hash fingerprints as the second hash fingerprints to obtain multiple second hash fingerprints; Calculate the fingerprint similarity coefficients between the first hash fingerprint and each second hash fingerprint to obtain multiple fingerprint similarity coefficients for the first hash fingerprint; Among them, the calculation method of the fingerprint similarity coefficient is as follows: ; Wherein: represents the fingerprint similarity coefficient, represents the i-th hash value in the first hash fingerprint, represents the i-th hash value in the second hash fingerprint, represents the dimension of the hash fingerprint, represents the maximum value function, represents a constant; Compare each fingerprint similarity coefficient with the fingerprint similarity coefficient threshold, and count the number of fingerprint similarity coefficients greater than or equal to the fingerprint similarity coefficient threshold to obtain the reproduction number of the first hash fingerprint; Calculate the ratio of the reproduction number of the first hash fingerprint to the number of all code hash fingerprints to obtain the reproduction rate of the code segment corresponding to the first hash fingerprint.

3. The code management method based on software requirement design according to claim 2, wherein The extraction of several target code segments from the software code document according to the transferability index includes: Compare the transferability index of each candidate code segment with a preset transferability index threshold; If the transferability index is greater than or equal to the transferability index threshold, mark the corresponding candidate code segment as a target code segment; If the transferability index is less than the transferability index threshold, mark the corresponding candidate code segment as a non-target code segment; Repeat the above steps until the comparison of all candidate code segments is completed to obtain several target code segments.

4. The code management method based on software requirements design according to claim 3, wherein The construction method of the code knowledge graph repository is as follows: Use the first target entity and the second target entity in the ontology library as the nodes in the knowledge graph, and use the predefined relationships between entities as the edges connecting the nodes in the knowledge graph to form the basic structure of the knowledge graph; Import the basic structure into the Neo4j database for storage to obtain a code knowledge graph repository for managing target codes.

5. The code management method based on software requirements design according to claim 4, characterized in that, The key Token features include multiple structural Tokens and non-structural Tokens. The structural Tokens include function names, class names, loop types, and conditional statement types; the non-structural Tokens include API call types, file operation types, and database operation types; Among them, the generation method of the deep learning model with the Transformer architecture is as follows: Obtain historical code function training data, and divide the historical code function training data into a code function training set and a code function test set. The historical code function training data includes the key Token features of code segments and their corresponding functional Tokens; Construct a regression network with the Transformer architecture, use the key Token features in the code function training set as the input of the regression network, and use the functional Tokens as the output of the regression network to train the regression network to obtain the original regression network; Use the code function test set to verify the model of the original regression network, and output the original regression network with a test error greater than or equal to the preset test error threshold as the deep learning model with the Transformer architecture.

6. A code management system based on software requirement design, implemented based on the code management method based on software requirement design described in any one of claims 1-5, characterized in that, Include: An ontology design module for extracting several target code segments from the software code document according to the transferability index and designing the ontology library of the knowledge graph for the target code segments; A graph construction module for constructing a code knowledge graph repository for managing target code based on the ontology library obtained after design; A model prediction module for, during the software development stage, real-time monitoring in the background of the key Token features of the code snippet currently being written by the developer and inputting them into a deep learning model with a Transformer architecture to predict the functional Token of the next code snippet to be written; An extension recommendation module for inputting the functional Token of the next code snippet to be written into the code knowledge graph repository to feedback and output a recommended code snippet for the next code snippet to be written.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the code management method based on software requirements design according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent dynamic prompting method and system for online code writing

    CN119149001A

  • Intelligent auxiliary method and system for software development

    CN119149045A