A specific project software code knowledge management platform and construction method thereof

By building a software code knowledge management platform for specific projects and establishing a mapping model between code and knowledge, the knowledge management problems of specific projects and fields in software code are solved, and the synchronous update and sharing of code and knowledge are realized, improving code quality and maintenance efficiency.

CN113986340BActive Publication Date: 2025-05-09FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111180973.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-05-09
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and update specific projects and domain knowledge in software code, resulting in difficulty in code understanding, problem positioning and quality management.

Method used

Build a project-specific software code knowledge management platform, and combine code knowledge extraction with code specification inspection to achieve synchronous update and sharing of code and knowledge by establishing a mapping model between code and knowledge.

Benefits of technology

It realizes the organic combination of code and knowledge, supports document generation on demand, code understanding and problem positioning, saves the cost of acquiring knowledge, and improves code quality and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113986340B_ABST
    Figure CN113986340B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of software engineering and intelligent software development and maintenance, and is specifically a specific project software code knowledge management platform and a construction method thereof. The platform includes a code and knowledge mapping model, a seed knowledge and traceability relationship module, a code automatic knowledge extraction module, and a code quality inspection feedback module; the code and knowledge mapping model is used to establish a mapping mode and specification between code and knowledge, and to clarify the types of code elements and knowledge types that need to be included in the platform management; the seed knowledge and traceability relationship module is used to obtain some seed knowledge and establish the initial traceability relationship between these seed knowledge and the code; the platform constructed by the present invention iteratively updates and improves the quality of software code and knowledge through the code automatic knowledge extraction and code quality inspection feedback modules. By connecting the constructed knowledge management platform to the project code base, the entire platform can continue to evolve and form a closed loop of positive promotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software engineering and intelligent software development and maintenance, and specifically relates to a specific project software code knowledge management platform and a construction method thereof. Background Art

[0002] With the rapid development of the software industry and the Internet industry, many companies have encountered a series of problems caused by existing codes while their businesses are developing rapidly. Due to the very long development and maintenance cycle and the alternation of new and old developers, many software have not accumulated relevant specific project and domain knowledge well during the continuous development and maintenance process. This will cause development and maintenance activities such as code understanding, adding new features, locating problems, and managing code quality to become very difficult while software code continues to accumulate. These problems have become a huge pain point for many companies. On the other hand, although software documentation is a good form of knowledge accumulation, developers generally ignore document writing and software documentation does not bring substantial short-term benefits. In practice, few software projects can truly form a complete and continuously updated set of documents.

[0003] In fact, the acquisition and update of software code knowledge has attracted the attention of academia and industry. The academic community proposed On-Demand Documentation Generation to study how to dynamically generate documents required by developers on demand, and the industry proposed the idea of ​​Live Documentation to guide developers to write documents that can evolve with the code. However, these two ideas have not made real breakthroughs in actual engineering practice, and still cannot solve the problem that code knowledge requires a lot of manpower to maintain and is difficult to update. In addition, although documents are a form that can carry specific project and domain knowledge contained in the code, they are often decoupled from the code, and the mapping relationship between code and knowledge cannot be traced through documents. This means that although some code-related knowledge can be understood by reading documents, developers and maintainers still need to spend a lot of time mapping this knowledge to the corresponding code. This results in documents not being of much help in many actual development and maintenance tasks (such as problem location), which is also one of the important reasons why developers and maintainers are not very motivated to write and update documents. Therefore, it is very important to build a solution that can efficiently manage code knowledge. Summary of the invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a specific project software code knowledge management platform and a construction method thereof. The present invention incorporates software code and specific project (including domain) knowledge into platform management by constructing a set of software code knowledge management platforms, thereby supporting the synchronous evolution and update of software code and knowledge, and supporting upper-level applications such as on-demand document generation, software code understanding, and software problem location and analysis. The core of the present invention is to construct a mapping model between software code and knowledge, and combine code knowledge extraction with code specification inspection. On the one hand, knowledge is extracted from the code and incorporated into the platform for management through a knowledge extraction algorithm based on code analysis, and knowledge with more obvious patterns is summarized and abstracted into a code specification template; on the other hand, the specification template is used to check the existing and newly added code, detect the existing specification problems in the code and improve the code quality, which in turn promotes the extraction of knowledge. Through the mutual promotion and continuous iteration of the above two aspects, the software code and knowledge are organically combined and managed.

[0005] The technical solution of the present invention is specifically described as follows.

[0006] The present invention provides a specific project software code knowledge management platform, which includes a code and knowledge mapping model, a seed knowledge and traceability relationship module, a code automatic knowledge extraction module and a code quality inspection feedback module; wherein:

[0007] The code and knowledge mapping model is used to establish the mapping mode and specification between code and knowledge, and to clarify the types of code elements and knowledge that need to be included in the platform management; it is divided into three parts according to different knowledge types:

[0008] The domain glossary includes a number of specific project and domain terminology knowledge, and is the basic knowledge base of the platform, which is mapped to the naming of various identifiers in the code;

[0009] The business knowledge base includes a number of business knowledge, which is mapped to the code execution logic including method calls, condition judgments, and attribute status modifications in the code, and is used to support code understanding and problem location applications;

[0010] The specification knowledge base includes several specification templates based on code specification knowledge, which are mapped to similar identifier naming, repeated branch statements and template class pattern implementations with similar functions in the code, and are used for code knowledge extraction in the code automatic extraction module and code quality inspection in the code quality inspection feedback module;

[0011] The seed knowledge and traceability module is used to obtain some seed knowledge and establish the initial traceability relationship between the seed knowledge and the code; the seed knowledge includes an initial specific project and domain term list, some business knowledge, and some standard templates based on code standard knowledge;

[0012] The automatic code extraction module combines code analysis, template matching and machine learning technology, adopts the idea of ​​bootstrapping to automatically extract the required knowledge from the code through seed knowledge, updates the corresponding knowledge base and establishes new traceability relationships;

[0013] The code quality check feedback module uses the updated specification template of the code automatic extraction module to check the quality issues in the code and feedback to the developer for modification.

[0014] In the present invention, in the seed knowledge and traceability relationship module, some seed knowledge is first obtained through expert knowledge extraction and analysis of existing documents, and then the initial traceability relationship between the seed knowledge and the code is established by combining code static analysis with manual confirmation.

[0015] In the present invention, in the automatic code extraction module, the code static analysis technology is used to parse out the identifiers, functions or method elements in the code, the standard template is read from the standard knowledge base and the standard template is matched with the code to obtain the terminology knowledge that can be directly recognized, and then the vocabulary mining is performed by machine learning to further enrich the domain terminology table, and the mined new terms are combined with the existing domain terminology knowledge and manual verification; the concept association, operation logic and state transition in the software are extracted by combining code static analysis and dynamic analysis, and the extraction results are processed and abstracted in combination with rules to form business knowledge, and the extraction results are selectively confirmed by manual intervention; then, based on the existing code and knowledge, the template is extracted by code clone detection and diff analysis to form code standard knowledge after manual confirmation; finally, the traceability relationship between knowledge and code is established by analyzing the source of the corresponding knowledge.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] The code knowledge management model proposed by the present invention is more valuable than the living document concept currently proposed by the industry. The documents proposed by the living document still face the problem of requiring a lot of manual maintenance and difficulty in knowledge sharing. The code and knowledge managed in the code knowledge platform constructed by the present invention are organically combined and directly rely on the code library, so the code and knowledge can be updated and shared synchronously, saving the cost of knowledge acquisition for project personnel at different links.

[0018] The platform constructed by the present invention can manage code elements such as identifiers, functions, files and modules in the software as well as the domain terminology, business knowledge base and specification knowledge base involved, and establish a traceability relationship between code and knowledge.

[0019] Based on the domain knowledge provided by domain experts, the present invention constructs a software code knowledge management platform for specific projects in a way of bidirectional mapping and synchronous evolution of code and knowledge. It achieves the dual purposes of software development knowledge extraction and code specification management through automatic extraction of code knowledge and code specification compliance analysis based on code specification management, and supports software development knowledge applications such as on-demand generation of software project documents, auxiliary code understanding, and code problem location.

[0020] The platform constructed by the present invention iteratively updates and improves the quality of software code and knowledge through automatic code knowledge extraction and code quality inspection feedback modules. By connecting the constructed knowledge management platform to the project code base, the continuous evolution of the entire platform can be achieved and a closed loop of positive promotion can be formed. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A high-level architecture diagram of the software code knowledge management platform constructed for the present invention. DETAILED DESCRIPTION

[0022] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and embodiments.

[0023] The present invention builds a code knowledge management platform to integrate software code and specific project and domain knowledge into platform management, thereby supporting the synchronous evolution and update of code and knowledge, and supporting upper-level applications such as document generation, code understanding, and code problem location. The architecture diagram of the software code knowledge management platform constructed by the present invention is as follows: Figure 1 shown.

[0024] The content managed by the entire platform is divided into three parts: software code, knowledge, and traceability. The software code level mainly includes identifiers, functions (or methods), files, and modules in the code; the knowledge level mainly includes specific project and domain terminology knowledge, business knowledge, and code specification knowledge; the traceability relationship is mainly used to maintain the mapping relationship between code and knowledge.

[0025] The main functional modules of the platform are divided into two parts: the automatic code extraction module, which is mainly responsible for automatically extracting the required knowledge from the code and updating the corresponding knowledge base; the code quality inspection feedback module, which mainly uses the standard templates in the standard knowledge base to check the quality problems in the code and feedback to the developer for modification.

[0026] The present invention implements the construction of a specific project software code knowledge management platform and the synchronous evolution of knowledge and code through the following five steps.

[0027] 1) Design of the code and knowledge mapping model. The mapping model is mainly used to establish the mapping mode and specification between code and knowledge, so as to clarify which code elements and what types of knowledge need to be included in the management of the platform. This design is common to different software. The mapping model we designed is mainly divided into three parts according to different knowledge types. The domain glossary is the basic knowledge base of the entire platform, which is mainly mapped to the naming of various identifiers in the code, so that all software developers have the same domain cognition and consistent domain knowledge; the business knowledge base mainly contains the concepts, relationships and related business logic involved in some important functions in the software, which is mainly mapped to the code execution logic such as method calls, conditional judgments, and attribute status modifications in the code, and can be used to support code understanding and problem location applications; the specification knowledge base mainly contains some specification templates related to software code implementation, which is mapped to similar identifier naming, repeated branch statements, and template classes with similar functions in the code, and can be used for code knowledge extraction and code quality inspection.

[0028] 2) Construction of seed knowledge and traceability relationship module. For a software project, some high-quality seed knowledge is obtained through expert knowledge extraction and analysis of existing documents. These seed knowledge include an initial domain term list, some high-quality business knowledge and some high-quality specification templates. Then, the relationship between these knowledge and the code is established by combining code static analysis with manual confirmation. In a preferred embodiment, for expert knowledge, we use modeling tools such as UML to express and refine domain knowledge. For existing software documents, we use AutoPhrase, a vocabulary mining tool commonly used in the field of self-language processing, to mine terms and relationship extraction tools to extract conceptual relationships.

[0029] 3) Construction of automatic code knowledge extraction module. Code static analysis technology is used to parse out identifiers, functions (methods) and other elements in the code. By reading templates from the standard knowledge base and matching the templates with the code, knowledge such as terms that can be directly recognized is obtained. Then, vocabulary mining is performed using machine learning and other methods to further enrich the glossary. The mined new terms will be verified in combination with existing knowledge and manual verification. For business knowledge, a combination of code static analysis and dynamic analysis will be used to extract concept associations, operating logic, state transitions, etc. in the software, and the extraction results will be processed and abstracted in combination with rules to form business knowledge. The extraction results can be confirmed by manual intervention at an appropriate time. Subsequently, based on the existing code and knowledge, templates are extracted through code clone detection and diff analysis, and then standardized knowledge is formed after manual confirmation. Finally, the traceability relationship between knowledge and code is established by analyzing the source of the corresponding knowledge.

[0030] In the preferred embodiment, regarding code analysis, taking Java code as an example, in the static analysis part, we mainly use JavaParser to extract the code abstract syntax tree and identifiers. We use soot to generate the intermediate representation of the code and extract the static call graph, data flow and control flow of the code. For dynamic analysis, we use BTrace to implement the Java bytecode instrumentation and track dynamic running data.

[0031] In a preferred embodiment, regarding automatic knowledge extraction from code, for terminology knowledge, a tool is used to split camel case names in identifiers into words and extract all possible N-grams. Existing terms are then used as seeds to automatically annotate training data using remote supervision and train classifiers to determine whether N-grams are terms. For business knowledge, we analyze the concepts involved in related functions and the associations between concepts from the call graph, data flow, and control flow of the code, and formulate corresponding extraction rules based on different software. At the same time, we construct state transitions and state machines in the software based on the running data tracked in dynamic analysis. For normative knowledge, we improve Sourcercc to detect clones of different granularities, and use gumtree to perform diff analysis on cloned code to form a normative template.

[0032] 4) Code quality inspection feedback module construction. Based on the existing specification knowledge base, the specification template is used to match the existing code and the new code to identify the code that does not meet the specifications. On this basis, the classifier is trained to determine whether the code that does not meet the specifications is the code with quality problems, and the problem code is reported to the relevant developers and maintenance personnel to promote the improvement of code quality.

[0033] For codes that do not match the specification template, we use a deep learning model to extract features from the corresponding code and specification template, and then train a classification model based on the extracted features. Specifically, the deep learning model used is a graph-based neural network, and the model input is the program dependency graph corresponding to the code and the corresponding specification template.

[0034] 5) Code library access and platform iterative evolution. In order to enable the entire platform to continuously and automatically evolve, the two core modules of the platform are connected to the appropriate location of the software code library and trigger rules are set. When the code is updated or other trigger conditions are met, the corresponding module will perform corresponding processing and update the knowledge managed in the platform.

[0035] In the present invention, the evolution process of the entire specific project software code knowledge management platform is as follows:

[0036] In the initial stage of the platform, the management content of the platform is mainly the stock code, and the knowledge and traceability relationship are missing. At this time, a small amount of knowledge will be manually injected as the initial seed, and this knowledge will be associated with the relevant code elements. Subsequently, combined with code analysis, template matching and machine learning technology, the idea of ​​bootstrapping is adopted to extract new knowledge and establish new traceability relationships through these seeds. In the early stage of the platform, certain manual confirmation methods are used to ensure that the extracted knowledge and the established traceability relationships are reliable. In this process, when new specification templates are summarized and added to the knowledge base, these templates will be used to check the code and report the code with quality problems. As the platform evolves, the knowledge base will gradually be enriched, the code quality will gradually improve, and manual intervention will gradually decrease. Through the iterative execution and mutual positive feedback of the two core functional modules, the entire platform will form a closed loop.

Claims

1. A specific project software code knowledge management platform, characterized by: It includes code and knowledge mapping model, seed knowledge and traceability relationship module, code automatic knowledge extraction module and code quality inspection feedback module; among which: The code and knowledge mapping model is used to establish the mapping mode and specification between code and knowledge, and to clarify the types of code elements and knowledge that need to be included in the platform management; it is divided into three parts according to different knowledge types: The domain glossary includes a number of specific project and domain terminology knowledge, and is the basic knowledge base of the platform, which is mapped to the naming of various identifiers in the code; The business knowledge base includes a number of business knowledge, which is mapped to the code execution logic including method calls, condition judgments, and attribute status modifications in the code, and is used to support code understanding and problem location applications; The specification knowledge base includes several specification templates based on code specification knowledge, which are mapped to similar identifier naming, repeated branch statements and template class pattern implementations with similar functions in the code, and are used for code knowledge extraction in the code automatic extraction module and code quality inspection in the code quality inspection feedback module; The seed knowledge and traceability module is used to obtain some seed knowledge and establish the initial traceability relationship between the seed knowledge and the code; the seed knowledge includes an initial specific project and domain term list, some business knowledge, and some standard templates based on code standard knowledge; The automatic code extraction module combines code analysis, template matching and machine learning technology, adopts the idea of ​​bootstrapping to automatically extract the required knowledge from the code through seed knowledge, updates the corresponding knowledge base, and establishes a new traceability relationship; The code quality check feedback module uses the updated standard template of the code automatic extraction module to check the quality problems in the code and feedback to the developer for modification; among them: In the automatic code extraction module, the code static analysis technology is used to parse out the identifiers, functions or method elements in the code, the standard template is read from the standard knowledge base and the standard template is matched with the code to obtain the terminology knowledge that can be directly recognized, and then the vocabulary mining method is used by machine learning to further enrich the domain terminology list, and the mined new terms will be combined with the existing domain terminology knowledge and manual verification; the concept association, operation logic and state transition in the software are extracted by combining code static analysis and dynamic analysis, and the extraction results are processed and abstracted in combination with rules to form business knowledge, and the extraction results are selectively confirmed by manual intervention; then, based on the existing code and knowledge, the template is extracted through code clone detection and diff analysis to form code standard knowledge after manual confirmation; finally, the traceability relationship between knowledge and code is established by analyzing the source of the corresponding knowledge.

2. The code knowledge management platform according to claim 1, characterized in that: In the seed knowledge and traceability relationship module, some seed knowledge is first obtained through expert knowledge extraction and analysis of existing documents, and then the initial traceability relationship between this seed knowledge and the code is established by combining code static analysis with manual confirmation.

3. The code knowledge management platform according to claim 1, characterized in that: In the code quality inspection feedback module, based on the existing specification knowledge base, specification templates are used to match existing codes and new codes to identify non-compliant codes. On this basis, classifiers are trained to determine whether non-compliant codes are codes with quality issues, and problematic codes are reported to relevant developers and maintenance personnel to promote the improvement of code quality.

4. A method for constructing a code knowledge management platform according to claim 1, characterized in that: The following five steps are used to build a software code knowledge management platform for a specific project and to synchronize the evolution of knowledge and code. The specific steps are as follows: 1) Design code and knowledge mapping model; 2) Construct seed knowledge and traceability relationship module; 3) Build a code automatic knowledge extraction module; 4) Build a code quality check feedback module; 5) Code library access and platform iterative evolution In order to enable the entire platform to continue to evolve automatically, the code automatic knowledge extraction module and the code quality inspection feedback module are connected to the appropriate location of the software code library and trigger rules are set. When the code is updated or other trigger conditions are met, the corresponding module will perform corresponding processing and update the knowledge managed in the platform.

Citation Information

Patent Citations

  • Software project and third-party library knowledge graph construction method for software system

    CN111241307A

  • Method and system for establishing open source project knowledge graph

    CN111949800A