Code digital twinning construction method based on multi-source heterogeneous knowledge extraction
By building a code digital twin model, using large models to extract high-level knowledge of software systems from multi-source heterogeneous data, it solves the problem that developers find it difficult to understand the knowledge of complex software systems, and achieves efficient development and maintenance support.
Patent Information
- Application Number
- CN202510459580.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-01
AI Technical Summary
In complex software systems, it is difficult for developers to efficiently understand and maintain domain knowledge. Existing tools have limited capabilities in parsing the high-level semantics and domain knowledge behind the code, resulting in high development and maintenance costs and inefficiency.
The large model combined with program analysis technology is used to automatically extract concepts, functional characteristics and design decision-making knowledge from multi-source heterogeneous data, build a code digital twin model, realize the full-link semantic mapping of high-level knowledge and underlying code units, and support development and maintenance tasks.
It improves developers' understanding of complex software systems, reduces the complexity of development and maintenance, provides intelligent development and maintenance support, significantly improves efficiency and reduces costs.
Smart Images

Figure CN120407011A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software engineering technology, and specifically relates to a method for constructing a code digital twin. Background Art
[0002] With the rapid development of information technology, the scale and complexity of modern software systems continue to rise. Take the Linux operating system, for example. Its source code volume has reached tens of millions of lines, encompassing hundreds of subsystems and modules. Within such a vast and highly complex codebase, developers face unprecedented cognitive challenges. Accurately understanding the domain concepts, functional features, and design decisions underlying a system has become a core task in software development and maintenance. Domain knowledge not only defines the key terminology, high-level semantics, and functional objectives of a system but also plays a vital role in development and maintenance tasks such as feature identification, problem troubleshooting, analysis of design decisions, and code implementation and bug fixes. For example, in the Linux kernel, core concepts such as "virtual memory," "network stack," and "file system" are not only fundamental to understanding the functionality of key modules and subsystems, but also carry a wealth of functional features, design considerations, and code implementation details.
[0003] However, this domain knowledge is often not explicitly reflected in direct software artifacts such as code and documentation. Instead, it is scattered across heterogeneous data sources such as feature descriptions, mailing lists, and change histories, and may not even be fully documented. Developers need to mine these data sources and combine them with expert experience to establish connections between code and domain knowledge. However, relying on manual labor to complete this process is not only time-consuming and labor-intensive, but also prone to missing key information, increasing development difficulty and creating maintenance barriers. Furthermore, as software systems continue to evolve, domain knowledge needs to be continuously accumulated and dynamically updated. Manual maintenance alone is difficult to keep up with, further exacerbating the challenges of understanding and management.
[0004] While existing static analysis and code navigation tools can reveal the structural relationships of code (such as dependencies and call graphs), they are limited in their ability to parse the high-level semantics and domain knowledge underlying the code. In recent years, with the rapid development of big model technology, automated knowledge extraction techniques based on natural language processing and deep semantic modeling have provided a new approach to solving this problem. Big models can uniformly analyze heterogeneous data sources, automatically extract high-level domain knowledge, and deeply associate it with the code. For example, when analyzing the Linux kernel mailing list, the big model, with its powerful semantic parsing capabilities, can capture the important information implicit in design decisions and accurately map it to the relevant code, providing more intuitive understanding support.
[0005] The present invention gives full play to the advantages of large models in semantic understanding, and combines program analysis techniques to automatically extract concepts, functional features, and design decision knowledge from multi-source heterogeneous data, and constructs a full-link semantic mapping between high-level knowledge and low-level code units. Finally, by intelligently constructing a digital twin of a complex software system, it provides efficient support for development and maintenance tasks, significantly improves development efficiency, reduces maintenance costs, and enhances developers' overall understanding of the system. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for constructing a code digital twin based on multi-source heterogeneous knowledge extraction with high development efficiency and low maintenance cost.
[0007] The method for constructing a code digital twin based on multi-source heterogeneous knowledge extraction provided by the present invention comprehensively uses domain concept extraction, software functional feature knowledge extraction, high-level design decision analysis, and code units and their structural relationships to construct a code digital twin model for the conceptual representation of software domain knowledge that is co-constructed, shared, and synchronously evolved with the code, helping developers understand the domain background knowledge, functional features, and design knowledge behind complex software code, so as to better understand the software and effectively complete development and maintenance tasks. This digital twin structures domain knowledge into two major categories: skeleton knowledge for software artifacts, and explanatory knowledge with design principles as the core. This digital twin can accurately express the high-level structure and design decisions of a software system; improve developers' understanding of complex software systems, provide intelligent development and maintenance support, and reduce the complexity of software development and maintenance.
[0008] The present invention first defines the code digital twin and related terms.
[0009] Software artifacts and related assets: They are the "physical" resources of a software system that developers can directly access, including source code, feature description lists, commit history, issue tracking records, mailing lists, change logs, etc.
[0010] Key knowledge elements: They are high-level conceptual assets that support software, including domain concepts, functions, and design principles (design decisions at different levels). Among them:
[0011] (1) Domain concept: It refers to the background concept knowledge used to understand the requirements and design solutions of a software system (for example, the basic operating system concepts involved in the Linux kernel).
[0012] (2) Function: It refers to the core functions implemented in code and documents, usually spanning multiple functions, files, or software packages.
[0013] (3) Design principle: The explanations of design decisions, such as API selection or module division, are usually hidden in unstructured resources such as commit messages, issue discussions, and mailing lists.
[0014] Code digital twin: Integrate key knowledge elements and their relationships through a unified representation method and closely associate with software artifacts. It includes the following components:
[0015] (1) Artifact-oriented skeletal knowledge: A high-level representation constructed using concepts and functions to provide a shared understanding of the system.
[0016] (2) Explanatory knowledge centered on design principles: Enhance the in-depth understanding of the system and decision support through detailed explanatory knowledge.
[0017] (3) Artifact-knowledge mapping and co-evolution: Ensure the consistency of skeletal knowledge and explanatory knowledge with the evolution of software artifacts.
[0018] The method for constructing a code digital twin based on multi-source heterogeneous knowledge extraction provided by the present invention specifically comprises the following steps:
[0019] (1) Knowledge representation design: Combine structured and unstructured knowledge representation methods to design different knowledge representation models for conceptual knowledge, functional characteristic knowledge, and design principle knowledge;
[0020] (2) Artifact-oriented skeletal knowledge extraction: Use large models to extract concepts and functional characteristics from documents and code from top to bottom, and summarize functional logic from code from bottom to top;
[0021] (3) Explanatory knowledge extraction centered on design principles: Extract design decisions and their bases from unstructured corpora through large models;
[0022] (4) Artifact-knowledge mapping construction: Map the extracted concepts and functional characteristics to software code elements;
[0023] (5) Support for co-evolution and incremental update: Connect the construction process with the evolution of software artifacts to ensure the synchronous evolution of the knowledge base and software artifacts.
[0024] Furthermore:
[0025] In the knowledge representation design described in step (1), the code digital twin uses multiple knowledge representation methods to support different categories of knowledge elements; specifically including:
[0026] Structured representation: That is, the form of knowledge representation under the specification of a specific schema, such as Knowledge Graph, Knowledge Card, and Knowledge Frame, etc. By modeling the relationships and hierarchies between knowledge units, efficient organization and convenient query of knowledge are realized. For example, a knowledge graph schema defining domain concepts and their relationships can be combined with domain expert knowledge (for example, classifying domain concepts and relationships, defining top-level concepts). In addition, a knowledge frame (including necessary attributes) can be defined to store functional features and different types of design decisions. The interaction mode between the knowledge graph and the knowledge frame can also be integrated by defining unified identifiers for knowledge units (such as concepts and functional features).
[0027] Unstructured representation: That is, text description, including natural language explanations and unstructured documents, providing a flexible and context-rich representation form. Unstructured representation is not restricted by predefined schemas and can capture knowledge in a more detailed and descriptive manner. For example, functional descriptions can be extracted from a list of features, and explanatory sentences of domain concepts can be stored.
[0028] Collaborative representation: That is, combining structured and unstructured knowledge representations to create a comprehensive view, integrating the advantages of both, and enhancing the overall understanding and practicality of knowledge.
[0029] The artifact-oriented skeletal knowledge extraction described in step (2);
[0030] This stage includes two parts: With the help of human experts, concepts and functional features are extracted from documents in a top-down manner, and functional logic is extracted from source code in a bottom-up manner.
[0031] Top-down extraction: Based on the concept knowledge schema defined in the structured representation, large language models (LLMs) are used to extract concepts and their predefined relationships from the original document text, and a schema-independent large language model method is adopted to capture additional concept knowledge, such as open relationships between concepts and descriptive texts, etc. To support the above extraction process, a set of concept knowledge extraction and verification frameworks will be constructed, covering key modules such as concept extraction, relationship extraction, explanation extraction, concept linking, concept fusion, and quality assurance, to ensure the accuracy and comprehensiveness of the extraction results. Subsequently, for functional feature knowledge, based on large language models, the original functional descriptions (such as text descriptions in the feature list) are structured into a predefined functional knowledge frame format, specifically including trigger requirement extraction, core concept role identification, implementation mechanism, sub-step decomposition, etc. A functional knowledge extraction and verification framework relying on large language models is constructed to efficiently extract relevant knowledge from the feature description list and establish a deep association between requirement feature knowledge and concept knowledge.
[0032] Bottom-up extraction: First, use static analysis tools to extract the structural information of the code, and then use large models to hierarchically summarize functions, function call chains, modules, and other cross-cutting code elements to extract functional logic from them.
[0033] The design decision-centered explanatory knowledge extraction described in step (3);
[0034] In this stage, large language models (LLMs) are used to extract design principle knowledge based on predefined knowledge framework patterns. First, use large models to identify relevant information from a large amount of unstructured corpora (e.g., a large number of email discussions), and then structure this information into a predefined knowledge framework format through large models. Specifically, the extraction of design principle knowledge includes identifying decision candidates and their pros and cons comparisons, the final design selection, decision-making bases, and other core contents, summarizing the logic by which design decisions are triggered by specific requirements or characteristics (triggers), and extracting the concepts involved and their specific roles (Roles). For this purpose, build a large model-based design knowledge extraction and verification framework to automatically extract design knowledge from code submission information and mailing lists, and accurately establish deep associations between requirement design knowledge, feature knowledge, and concept knowledge.
[0035] The artifact-knowledge mapping construction described in step (4);
[0036] In this stage, large language models (LLMs) are used to associate concepts between artifacts and knowledge elements, and map the functional characteristics between software artifacts and knowledge elements; specifically including:
[0037] Concept semantic matching: Utilize the capabilities of large models in text and code understanding, combine the context of concept texts with the context of code elements, and accurately map high-level domain concepts to code elements and related descriptions to ensure the semantic association between concepts and code units.
[0038] Function mapping: Map the top-down extracted functional characteristic knowledge to the functional logic summarized from the code bottom-up. Among them, code commit serves as a key bridge. Code commits are often cited in feature descriptions and mailing lists, and code commit records are naturally associated with modified files, functions, and code blocks. By combining the semantic understanding and matching capabilities of large models, establish a direct connection between high-level feature knowledge and code functional logic. In addition, further analyze the semantic of code functions through large models, construct an accurate association between the breakdown of feature steps and code implementation, and form a semantic link between feature decomposition and code implementation.
[0039] The co-evolution and incremental update described in step (5);
[0040] The construction process must support incremental updates to keep alignment with the real-time status of software artifacts and related assets. Specifically, by seamlessly integrating the version management system of the code repository, the Issue tracking system, and the email system into the platform, real-time monitoring of code and related data is achieved, and an incremental update mechanism is adopted to ensure that the knowledge base can be updated synchronously with code evolution. In addition, through continuous feedback from domain experts or users, the knowledge system is optimized and verified to improve the accuracy and reliability of knowledge. At the same time, associated knowledge is automatically extracted based on user interactions and incorporated into the knowledge base to achieve dynamic accumulation and continuous optimization of knowledge, ensuring that the knowledge base always maintains the characteristics of high efficiency, accuracy, and evolvability during software development and maintenance. Brief Description of the Drawings
[0041] Figure 1 It is a high-level overview diagram of the method for constructing the code digital twin of the present invention. Detailed Implementation Manner
[0042] (1) High-level knowledge schema construction
[0043] To extract high-level domain knowledge from text data, it is first necessary to design and construct an efficient schema. With the participation of domain experts, based on the knowledge architecture of the target domain, a schema including domain concepts, requirement characteristics, and design decisions is constructed. The schema structure should cover various concepts, relationships, and explanations, specifically including the definitions of domain terms, the relationships between requirements and characteristics, the candidate options of design decisions and their advantages and disadvantages, etc. To ensure that these concept categories can accurately reflect the actual needs of the target domain, analysis methods such as Theme Analysis or Open-coding can be introduced to manually define the concept categories.
[0044] In this process, through the theme analysis of the literature, requirement documents, and design documents of the target domain, domain experts can identify the recurring themes or patterns in the documents and use them as candidate concept categories. Through careful analysis of these candidates, experts can define representative and accurate concept categories and formulate specific descriptions, relationships, and explanations for them. In addition, using the Open-coding method, experts can gradually mark and classify requirements, designs, and domain concepts to extract potential knowledge structures from the underlying text. This method helps experts systematically organize and code the concepts in the text and discover potential knowledge that is not explicitly expressed in the domain.
[0045] During the implementation process, experts participate in the design of the schema through professional tools and interfaces to ensure that each knowledge category can accurately reflect the actual needs of the target domain. The feedback and participation of experts will be carried out through regular review and discussion sessions to ensure that each domain document and design document can be accurately mapped to the predefined schema. Through this method of injecting expert knowledge, the extraction results of knowledge not only have high practicality and domain matching, but also help the system better understand the domain background and professional terms, thus laying a foundation for subsequent knowledge extraction and analysis.
[0046] (2) Code Static Analysis and Semantic Annotation
[0047] During the implementation process, first, the source code is parsed by static analysis tools (such as Clang, LLVM) to extract the basic structural units in the code, such as modules, classes, functions, variables, etc., and at the same time, a relationship graph between these units is constructed, including function calls, module dependencies, etc. For each code unit, a large model (such as GPT-4 or other semantic understanding models) is used to perform semantic parsing on it to understand its function, behavior, and implementation logic. In the specific implementation process, appropriate model prompting engineering techniques (such as Few-Shot In-context Learning, Chain-of-Thoughts, etc.) are used to perform semantic annotation on identifiers (such as the meanings of variables, functions, classes, etc.), describe the functions of code blocks and functions, and parse the logic of branch paths. This process is completed through automated scripts or pipelines to ensure that the structural and semantic information of the code can be completely extracted and converted into multi-dimensional semantic tags as the basis for subsequent mapping and association.
[0048] (3) Construction of a Knowledge Extraction Pipeline Based on a Large Model
[0049] It is to combine a large model with a rule-driven extraction framework to construct an automated knowledge extraction pipeline. This pipeline first grabs relevant data from multiple data sources (such as requirement documents, design documents, mailing lists, code submission information, etc.). The grabbed data will be preprocessed and standardized, including operations such as text cleaning and low-quality text filtering, to ensure data quality.
[0050] Next, the pipeline performs knowledge extraction on text data based on large language models (such as GPT-4), with a focus on extracting domain concepts, requirement characteristics, and design decisions. Through a pre-designed schema, the extracted domain knowledge is structured and classified, and an attempt is made to link the extracted conceptual knowledge with existing knowledge bases (such as Wikipedia or Wikidata) to ensure the accuracy and depth of the concepts. During this process, the pipeline not only focuses on the extraction of schema-defined knowledge but also identifies and extracts schema-free knowledge, such as open relationships between concepts and explanations of open concepts, to ensure the comprehensiveness and flexibility of the knowledge. One of the cores of the above process lies in designing a reasonable large language model knowledge extraction method to fully leverage the advantages of different prompt engineering methods.
[0051] Quality assurance is required for each extraction step. For example, in the concept extraction stage, the pipeline combines methods such as reference corpora and data comparison to ensure that the extracted concepts are accurate and representative. For complex knowledge types (such as requirement characteristics and design decisions), based on the initial extraction, the pipeline relies on the feedback of domain experts for manual correction and verification. This step ensures, through expert review and feedback, that the finally extracted knowledge highly matches the actual content of the target domain.
[0052] For different types of knowledge, corresponding knowledge persistence methods are designed. Specifically, high-level knowledge such as domain knowledge, requirement characteristics, and design decisions needs to be persisted in a database and shared with other systems or tools (such as large language models, code analysis platforms). At this time, a combination of a graph database (such as Neo4j) and a relational database (such as MySQL) can be used to store knowledge for efficient querying, updating, and associating different types of knowledge. In addition, for large-scale text data and extracted knowledge, a distributed storage system (such as HDFS or a database cluster) can be used to support the efficient processing and management of data. In this way, knowledge can not only be stored for a long time but also be flexibly updated and queried according to development needs, supporting knowledge management and iteration throughout the software life cycle.
[0053] (4) Artifact-Knowledge Mapping Construction
[0054] The goal of knowledge mapping and traceability relationship construction is to ensure an accurate mapping path from code to high-level domain knowledge. During the implementation process, first, the code semantics are deeply parsed by a large model, and the extracted code units are mapped to predefined domain concepts. The specific method is to generate semantic descriptions of code units through the large model, and combine context information (such as function calls, module dependencies, etc.), and based on large model prompting engineering, semantically match these code elements with relevant concepts in the domain (such as functional modules, system components, etc.). At the same time, a two-way mapping framework is designed. Through code commit records, based on large model prompting engineering, the semantic descriptions of code elements are associated with requirement features and design decisions. Experts will provide corrections for complex or ambiguous mapping relationships during the mapping process to ensure the accuracy of the mapping. Finally, a traceability link from code to high-level knowledge is established to ensure that each code snippet can be traced back to its underlying requirements, features, or design decisions.
[0055] Implementing the above method on Linux Kernel can extract approximately 80,000 domain concepts from about 40,000 feature descriptions. Each concept contains multiple aliases and is respectively linked to general knowledge bases such as code repositories, documents, and Wikipedia. In addition, by extracting commit messages and mailing lists, approximately 5,000 to 10,000 pieces of design knowledge at different levels can also be obtained. Based on the above knowledge, combined with generative artificial intelligence, knowledge views from multiple information sources and with multiple perspectives can be provided to users on demand, thus effectively improving the automation and intelligence levels in the maintenance process of complex software such as Linux Kernel.
Claims
1. A method for constructing a code digital twin based on multi-source heterogeneous knowledge extraction, characterized in that, Comprehensively apply domain concept extraction, software function feature knowledge extraction, high-level design decision analysis, and code units and their structural relationships to construct a code digital twin model for the conceptual representation of software domain knowledge that co-constructs, shares, and evolves synchronously with the code, helping developers understand the domain background knowledge, function features, and design knowledge behind complex software code, so as to better understand the software and effectively complete development and maintenance tasks; This digital twin structures domain knowledge into two major categories: skeletal knowledge oriented to software artifacts and explanatory knowledge centered on design principles; this digital twin can accurately express the high-level structure and design decisions of a software system, and the specific steps are as follows: (1) Knowledge representation design: Combine structured and unstructured knowledge representation methods to design different knowledge representation models for conceptual knowledge, function feature knowledge, and design principle knowledge; (2) Artifact-oriented skeletal knowledge extraction: Use large models to extract concepts and function features from documents and code from top to bottom, and summarize function logic from code from bottom to top; (3) Design principle-centered explanatory knowledge extraction: Extract design decisions and their bases from unstructured corpora through large models; (4) Artifact-knowledge mapping construction: Map the extracted concepts and function features to software code elements; (5) Support for co-evolution and incremental updates: Connect the construction process with the evolution of software artifacts to ensure the synchronous evolution of the knowledge base and software artifacts.
2. The method for constructing a code digital twin according to claim 1, wherein For the knowledge representation design described in step (1), the code digital twin uses multiple knowledge representations to support different categories of knowledge elements; specifically including: Structured representation: That is, a knowledge representation form under specific schema specifications, which realizes the efficient organization and convenient query of knowledge by modeling the relationships and hierarchies between knowledge units; Unstructured representation: That is, text description, including natural language explanations and unstructured documents, providing a flexible and context-rich representation form; unstructured representation is not restricted by predefined schemas and can capture knowledge more meticulously and descriptively; Collaborative representation: That is, combine structured and unstructured knowledge representations to create a comprehensive view, combine the advantages of both, and enhance the overall understanding and practicality of knowledge.
3. The method for constructing a code digital twin according to claim 2, wherein For the artifact-oriented skeletal knowledge extraction described in step (2), it specifically includes two parts: With the help of artificial experts, extract concepts and function features from documents in a top-down manner, and extract function logic from source code in a bottom-up manner; Top-down extraction, that is, based on the concept knowledge schema defined in the structured representation, use large language models (LLMs) to extract concepts and their predefined relationships from the original document text, and adopt a schema-independent large model method to capture additional concept knowledge; To support the above extraction process, a set of concept knowledge extraction and verification frameworks will be constructed, covering key modules such as concept extraction, relationship extraction, explanation extraction, concept linking, concept fusion, and quality assurance, to ensure the accuracy and comprehensiveness of the extraction results; Subsequently, for functional feature knowledge, the original functional descriptions are structured into a predefined functional knowledge framework format based on large models, specifically including trigger requirement extraction, core concept role identification, implementation mechanism, and sub-step decomposition; a functional knowledge extraction and verification framework relying on large models is constructed to extract relevant knowledge from the feature description list and establish a deep association between requirement feature knowledge and concept knowledge; For bottom-up extraction, first, a static analysis tool is used to extract the structural information of the code, and then a large model is used to hierarchically summarize functions, function call chains, modules, and other cross-cutting code elements to extract functional logic from them.
4. The code digital twin construction method according to claim 3, wherein For the design decision-centered explanatory knowledge extraction described in step (3), a large model (LLMs) is specifically used to extract design principle knowledge based on a predefined knowledge framework pattern; first, the large model is used to identify relevant information from a large amount of unstructured corpora, and then this information is structured into a predefined knowledge framework format by the large model; specifically, the extraction of design principle knowledge includes identifying decision candidates and their pros and cons comparison, the final design selection, the basis for decision-making, summarizing the logic by which design decisions are triggered by specific requirements or features, and extracting the concepts involved and their specific roles; for this purpose, a set of design knowledge extraction and verification frameworks based on large models is constructed to automatically extract design knowledge from code commit information and mailing lists and accurately establish a deep association between requirement design knowledge, feature knowledge, and concept knowledge.
5. The method for constructing a code digital twin according to claim 4, wherein For the construction of the artifact-knowledge mapping described in step (4), a large model (LLMs) is specifically used to associate concepts between artifacts and knowledge elements and map the functional features between software artifacts and knowledge elements; specifically including: Concept semantic matching: Utilize the capabilities of large models in text and code understanding, combine the context of concept texts with the context of code elements, and accurately map high-level domain concepts to code elements and related descriptions to ensure the semantic association between concepts and code units; Functional mapping: Map the top-down extracted functional feature knowledge to the functional logic summarized from the code bottom-up; among them, code commits serve as the key bridge; feature descriptions and mailing lists reference code commits, and code commit records are naturally associated with the modified files, functions, and code blocks. By combining the semantic understanding and matching capabilities of large models, a direct connection between high-level feature knowledge and code functional logic is established; in addition, through further analysis of the code functional semantics by the large model, an accurate association between feature step decomposition and code implementation is constructed to form a semantic link between feature decomposition and code implementation.
6. The method for constructing a code digital twin according to claim 5, wherein The co-evolution and incremental update described in step (5) include building a process to support incremental update to maintain alignment with the real-time status of software artifacts and related assets. Specifically, by seamlessly integrating the version management system of the code repository, the Issue tracking system, and the email system into the platform, real-time monitoring of code and related data is achieved, and an incremental update mechanism is adopted to ensure that the knowledge base can be updated synchronously with code evolution. In addition, through continuous feedback from domain experts or users, the knowledge system is optimized and verified to improve the accuracy and reliability of knowledge. At the same time, associated knowledge is automatically extracted based on user interactions and incorporated into the knowledge base to achieve dynamic accumulation and continuous optimization of knowledge, ensuring that the knowledge base always maintains the characteristics of high efficiency, accuracy, and evolvability during software development and maintenance.
7. The method for constructing a code digital twin according to claim 6, wherein The normalized Schema has a structure covering various concepts, relationships, and explanations, specifically including the definitions of domain terms, the relationships between requirements and characteristics, the candidate options for design decisions and their advantages and disadvantages. To ensure that these concept categories can accurately reflect the actual needs of the target domain, topic analysis or Open-coding analysis methods are introduced to manually define the concept categories. In this process, through topic analysis of the literature, requirement documents, and design documents in the target domain, domain experts identify the recurring themes or patterns in the documents and use them as candidate concept categories. Through careful analysis of these candidates, experts define representative and accurate concept categories and formulate specific descriptions, relationships, and explanations for them. In addition, using the Open-coding method, experts gradually label and classify requirements, designs, and domain concepts to extract potential knowledge structures from the underlying text.
8. The method for constructing a code digital twin according to claim 7, wherein It also involves building a knowledge extraction pipeline based on a large model, and this pipeline includes: First, relevant data is crawled from multiple data sources; the crawled data is preprocessed and standardized, including text cleaning and low-quality text filtering operations to ensure data quality. Next, knowledge extraction is performed on the text data based on the large model, with a focus on extracting domain concepts, requirement characteristics, and design decisions; through a pre-designed schema, the extracted domain knowledge is structured and classified, and the extracted concept knowledge is linked to the existing knowledge base to ensure the accuracy and depth of the concepts. In this process, the pipeline not only focuses on the extraction of schema-defined knowledge but also identifies and extracts schema-free knowledge to ensure the comprehensiveness and flexibility of knowledge. Quality assurance is required for each extraction link; including in the concept extraction stage, the pipeline combines a reference corpus and data comparison methods to ensure that the extracted concepts are accurate and representative; for complex knowledge types, based on the preliminary extraction, relying on the feedback of domain experts, manual correction and verification are carried out. For different types of knowledge, corresponding knowledge persistence methods are designed; specifically, high-level knowledge such as domain knowledge, requirement characteristics, and design decisions needs to be persisted in the database and shared with other systems or tools; at this time, a combination of graph database and relational database is used to store knowledge for efficient querying, updating, and associating different types of knowledge; in addition, for large-scale text data and extracted knowledge, a distributed storage system is used to support the efficient processing and management of data; in this way, knowledge can not only be stored for a long time, but also be flexibly updated and queried according to development needs, supporting knowledge management and iteration throughout the software life cycle.