Electronic test instrument knowledge base construction method based on large language model

Through the large language model combined with the field ontology construction method, structured prompt word templates are designed, which solves the problems of complexity and high maintenance costs of knowledge graph construction in the existing technology, and realizes efficient information extraction and dynamic updates, improving the testing guarantee efficiency of electronic testing instruments.

CN120541237APending Publication Date: 2025-08-26THE 41ST INST OF CHINA ELECTRONICS TECH GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510624399.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing knowledge graph construction methods are complex, there are many data types, inconsistent knowledge representation, and complex data source processing methods, resulting in low information extraction efficiency and high dynamic update and maintenance costs, making it difficult to meet the diversified and complex testing needs of electronic testing instruments.

Method used

A large language model is used to combine the field ontology construction method to design structured prompt word templates, extract core concepts through TF-IDF and TextRank algorithms, use OpenIE to perform unsupervised relationship extraction, and form a structured knowledge base in combination with graph databases and traditional databases to achieve efficient information extraction and dynamic updates.

Benefits of technology

It improves the efficiency of information extraction, reduces maintenance costs, forms a high-quality electronic testing instrument knowledge base, and improves the testing guarantee efficiency and data application value of testing instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541237A_ABST
    Figure CN120541237A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic test instrument knowledge base construction method based on a large language model, and belongs to the technical field of electronic test instruments. According to the method, the construction of the electronic test instrument domain ontology is completed by adopting a technical path of fusing a term frequency and inverse document frequency (TF-IDF) algorithm, a text ranking algorithm (TextRank), an open information extraction method (OpenIE) and domain expert auditing, so that the problem of non-uniform representation of industry knowledge is solved; the field information extraction cue word design is adopted, the constructed field ontology is combined, automatic extraction of the technical field information of the electronic test instrument is achieved on the basis of a large language model, and the problems of low information extraction efficiency and the like are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic testing, and in particular relates to a method for constructing an electronic testing instrument knowledge base based on a large language model. Background Art

[0002] Electronic testing technology and instrumentation are methods and means for qualitative or quantitative measurement of electrical and non-electrical quantities using modern electronic technology, addressing the full life cycle testing needs of electronic equipment development, production, and assurance. With the increasing sophistication, integration, and digitalization of electronic equipment, the testing tasks affecting the entire life cycle of electronic equipment are becoming increasingly diverse and complex. Traditional testing technology is primarily based on "blind" testing of instrument panel parameters. The testing process relies heavily on the professional quality and experience of the testers, and the test instruments are complex to operate. This makes it difficult for the traditional low-efficiency, high-cost test assurance model to meet current testing needs.

[0003] A knowledge graph is a large-scale knowledge network that describes concepts, entities, and their relationships in the objective world in a structured form. This expresses information in a form closer to human cognition, providing a way to better organize, manage, and understand massive amounts of information. Based on test technology terminology, expert knowledge, test instrument manuals, and various business data, a knowledge graph for electronic test instruments is constructed to form a high-quality, structured test knowledge base. This standardizes the expression of various test elements in the test process and the relationships between test elements, completing the construction of an information-based testing technology knowledge infrastructure. This strengthens the digital information foundation for intelligent testing, enhances the application value of test data, further empowers test instruments, and strives to improve the test assurance efficiency of test instruments and systems. However, due to the highly domain-specific nature of the test instrument knowledge, test technology methods, and test business data covered in the electronic testing field, the entity types, relationship types, and test semantic association application scenarios are unique, and the forms and methods of knowledge extraction tasks such as entity recognition and relationship extraction vary significantly. Therefore, the general knowledge graph construction process and methods are not directly applicable to the electronic testing field. Large language model technology has outstanding advantages in interactive task understanding and content generation. Information extraction based on large language models provides new ideas for the unified output of multi-type knowledge extraction tasks, and provides a flexible means for the high-precision and high-coverage construction of electronic test instrument knowledge graphs.

[0004] The existing knowledge graph construction method is complex and involves many modules, such as Figure 1As shown in the figure, the three different types of data are processed separately. Entity extraction, relationship extraction, and attribute extraction are performed on semi-structured and unstructured data. Structured data and third-party database data are integrated. Then, knowledge reasoning technology is used to entity link the data obtained above. The quality of the entity linking results is further evaluated, and knowledge is updated. Finally, data is stored and the knowledge graph is constructed.

[0005] The shortcomings of the existing technology are mainly reflected in: first, there are many data types (for example, non-structured data such as word documents / pdf documents, semi-structured data such as html / xml, and structured data such as MySQL relational database files), the knowledge representation is not unified, and the knowledge is difficult to use later; second, the data source processing method is complex, and information extraction methods need to be designed separately. There is a dependency between the entity and relationship extraction modules, resulting in low information extraction efficiency and high costs for dynamic knowledge updates and maintenance. Summary of the Invention

[0006] In response to the above technical problems existing in the prior art, the present invention proposes a method for constructing an electronic test instrument knowledge base based on a large language model. The method has a reasonable design, overcomes the shortcomings of the prior art, and has good effects.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for constructing an electronic test instrument knowledge base based on a large language model comprises the following steps:

[0009] Step 1: domain ontology construction;

[0010] Step 2: Design the prompt word template;

[0011] Based on the domain ontology, a structured prompt word template is constructed, including four parts: overview, process, dependency, and control. The domain ontology structure and information extraction examples are embedded in the template to guide the large language model to perform information extraction tasks.

[0012] Step 3: Large language model information extraction;

[0013] Using prompt word templates and combining large language models, we extract entities, attributes, and relationships from multi-source data in the field of electronic test instrument technology and generate a set of fact triples.

[0014] Step 4: Knowledge base construction and storage;

[0015] The extracted triples are stored in a graph database and integrated with a traditional relational database to form a structured electronic test instrument knowledge base, realizing unified knowledge management and dynamic updating.

[0016] Preferably, in step 1, domain ontology construction includes the following steps:

[0017] Step 1.1: Use the TF-IDF algorithm and the TextRank algorithm to summarize and extract the core concepts in the field of electronic test instruments;

[0018] Step 1.2: Perform unsupervised relationship extraction on unstructured or semi-structured data using the Open Information Extraction method (OpenIE) to screen meaningful relationships.

[0019] Step 1.3: Combine the general knowledge graph ontology and domain expert review to optimize entity types, attribute associations, and relationship semantic expressions, and complete ontology modeling using the Protégé tool;

[0020] Step 1.4: Form an electronic test instrument technology field ontology covering entity definition, attribute definition, relationship definition and concept hierarchy division.

[0021] Preferably, in step 2, the structured design of the prompt word template includes the following steps:

[0022] Step 2.1: Overview: Define the task context, goals, and the role of the large language model;

[0023] Step 2.2: Process: Clarify the information extraction steps, rules, and execution order;

[0024] Step 2.3: Dependencies: Specify the required tools, knowledge bases, and data materials;

[0025] Step 2.4: Control: Limit the sampling scope, quality requirements and error avoidance rules.

[0026] Preferably, in step 3, extracting large language model information includes the following steps:

[0027] Step 3.1: Label a small number of sample sets needed for the information extraction process;

[0028] Step 3.2: Focusing on the task of constructing a knowledge graph in the field of electronic testing instrument technology, further design and construct positive and negative sample sets for the information extraction process;

[0029] Step 3.3: Train the model based on the labeled positive and negative sample sets to improve the extraction coverage and accuracy of entities, attributes, and relationships.

[0030] Preferably, in step 4, the knowledge storage uses a graph database to store triple relationships and is associated with a relational database to form a multimodal structured knowledge base.

[0031] Preferably, the large language model is deepseek-r1:7b, which realizes information extraction and semantic understanding in the field of electronic test instruments through domain ontology constraints and prompt word guidance.

[0032] The beneficial technical effects brought about by the present invention are:

[0033] (1) The electronic test instrument domain ontology is constructed by integrating the term frequency and inverse document frequency algorithm (TF-IDF), the text ranking algorithm (TextRank), the open information extraction method (OpenIE), and the review of domain experts. This solves the problem of inconsistent representation of domain industry knowledge and avoids the problems of semantic inconsistency and lack of professional concepts caused by general ontology.

[0034] (2) For data sources with different data types / data formats (such as non-structured data such as word documents / pdf documents, semi-structured data such as html / xml, and structured data such as MySQL relational database files), information extraction prompts are designed that meet the data characteristics in the field of electronic testing instrument technology. Information extraction is completed based on the deepseek-r1:7b large language model to form a high-quality fact triple set, avoiding the separate design and optimization of entity / relationship / attribute extraction methods, the dependency between entity and relationship extraction modules, etc., which leads to low information extraction efficiency, dynamic knowledge update, and high maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Build a block diagram for a knowledge graph;

[0036] Figure 2 Construct a block diagram for the electronic test instrument knowledge base based on a large language model;

[0037] Figure 3 Design entity attribute relationship diagram for the ontology in the field of electronic test instrument technology;

[0038] Figure 4 Design a conceptual hierarchy diagram for the ontology in the field of electronic test instrument technology;

[0039] Figure 5 Construct a prompt design schematic for the electronic test instrument knowledge graph. DETAILED DESCRIPTION

[0040] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0041] The technical solution process of the present invention is as follows Figure 2 As shown, the specific implementation is as follows.

[0042] 1. Domain-local construction

[0043] When constructing ontologies in the field of electronic test instrument technology, in order to model scenario elements as accurately as possible, it's necessary to understand business details, such as how to split complex entities and distinguish between attributes and entity relationships, while also maintaining a high-level understanding of business abstraction. This often requires domain experts to evaluate the conceptual system and relationships before a relatively complete ontology model can be formed. However, due to the large number of ontologies in the testing field, their complex relationships, and the subjective nature of entity type design, along with the lack of unified standards, manually constructed ontology models often suffer from flaws. Therefore, the key to constructing a measurement knowledge graph is how to efficiently and automatically acquire factual knowledge while maintaining acceptable accuracy and designing a domain ontology.

[0044] This paper uses the TF-IDF algorithm and the TextRank algorithm to summarize and extract the core concepts in the test field. It uses the OpenIE method to extract unsupervised open relationships from the test instrument manuals, and then finds meaningful relationships from them. At the same time, it draws on the ontology knowledge of the general knowledge graph and iteratively optimizes the expression accuracy and knowledge coverage of the domain ontology under the guidance of experts. Finally, it uses the Protégé tool to complete the ontology construction, and finally completes the domain entity definition, attribute definition, relationship definition, concept hierarchy division, etc. Figure 3 、 Figure 4 This is a schematic diagram of an example of body design in the field of electronic test instrument technology.

[0045] The domain ontology was constructed and successfully stored in the ontology library. The domain knowledge contained in this ontology library will provide key clues for the Large Language Model (LLM), enabling it to more accurately extract graph triples. This paradigm design enables the LLM to better understand the knowledge background and contextual information in the field of electronic test instrument technology during the extraction of entities, relationships, and attributes, thus laying a solid foundation for building and updating the knowledge graph.

[0046] 2. Prompt Engineering Template Design and Information Extraction Based on Large Language Model

[0047] Prompt engineering is a new field that focuses on creating and optimizing prompts so that large language models can adapt more effectively to a variety of different applications and research fields. Prompts are natural language input sequences for large language models and need to be created based on specific tasks built on domain knowledge graphs. Prompts can be composed of multiple elements, including instructions, background information, and input text. Instructions are a phrase used to guide the model to perform a specific task. Background information provides information related to the input text, which helps the model understand and process the task. Input text is the main text data that the model needs to process. By designing and optimizing prompts to guide large language models to more accurately understand user intent, higher quality information can be generated.

[0048] To improve the large language model's ability to understand user intent, prompts must be structured. This paper employs the following design paradigm to structure prompts, which is divided into four parts: overview, process, dependency, and control. This ensures that the prompt design has a clear logic and organizational structure, enabling accurate and efficient extraction of test element triples.

[0049] (1) Overview: Describe the roles that users or artificial intelligence (AI) need to play and the tasks to be completed in a specific context to achieve a specific goal. This section provides the context of the problem and specifies the identities and tasks of the participants.

[0050] (2) Process: This describes the specific steps and processes required to complete the task. It describes how participants will use their intelligence skills, which rules they will follow, and in what order they will perform the task. This part includes the specific execution process of the task, as well as the required intelligent behaviors and rules.

[0051] (3) Dependencies: List the tools, knowledge, and materials required to complete the task. This section describes the relevant resources and information involved in the execution of the task, including the tools and techniques used, the required knowledge areas, and the data or materials processed.

[0052] (4) Control: This section specifies the requirements and restrictions for the task execution process, including both positive and negative requirements. This section specifies the criteria and standards for task execution, as well as the risks and restrictions that may be involved. The control section may also include supervision and management measures for the task execution process.

[0053] Using the defined electronic test instrument technology field ontology structure, based on the test instrument manual and other corpora, we annotate a small number of sample sets needed for the information extraction process. Aiming at the task of building a knowledge graph in the electronic test instrument technology field, we further design and construct positive and negative sample sets for the information extraction process. Then, we embed the previously defined electronic test instrument technology field ontology structure and information extraction examples in the prompt, and provide relevant instructions for building the electronic test instrument knowledge graph, such as Figure 5 A large language model (deepseek-r1:7b is used in this paper) is combined with defined prompt words and text corpus to achieve efficient and accurate information triple extraction, complete the construction of the electronic test instrument knowledge graph, store it in the Neo4j graph database, and jointly form a structured, high-quality electronic test instrument knowledge base with traditional relational databases in the field.

[0054] The key points and protection points of the present invention are:

[0055] 1. The present invention proposes a knowledge graph ontology construction method for the field of electronic test instrument technology. For the subject ontology of the field of electronic test instrument technology, coverage and accuracy are very important evaluation indicators. In the case that the current Chinese ontology automatic construction technology is still immature, combined with the business data characteristics of the test field, the term frequency and inverse document frequency algorithm (Term Frequency-Inverse Document Frequency, TF-IDF) and TextRank algorithm are used to achieve the induction and extraction of core concepts in the test field. The open information extraction method (Open Information Extraction, OpenIE) is used to perform unsupervised open relationship extraction of test instrument knowledge, encyclopedic knowledge, domain expert knowledge, etc., and then find meaningful relationships from them. At the same time, drawing on the ontology knowledge of the general knowledge graph, the expression accuracy and knowledge coverage of the domain ontology are iteratively optimized under the guidance of experts.

[0056] 2. This paper proposes a domain information extraction method based on ontology prompts and large language model reasoning in the field of electronic test instruments. Aiming at information extraction tasks such as domain entity recognition and relationship / attribute extraction, this method labels a small amount of data based on a custom domain ontology and designs domain information extraction prompts. It then utilizes the large language model DeepSeek-r1:7b to efficiently extract information from the electronic test instrument domain, generating a high-quality set of fact triples.

[0057] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.

Claims

1. A method for constructing an electronic test instrument knowledge base based on a large language model, characterized by: The following steps are involved: Step 1: domain ontology construction; Step 2: Design the prompt word template; Based on the domain ontology, a structured prompt word template is constructed, including four parts: overview, process, dependency, and control. The domain ontology structure and information extraction examples are embedded in the template to guide the large language model to perform information extraction tasks. Step 3: Large language model information extraction; Using prompt word templates and combining large language models, we extract entities, attributes, and relationships from multi-source data in the field of electronic test instrument technology and generate a set of fact triples. Step 4: Knowledge base construction and storage; The extracted triples are stored in a graph database and integrated with a traditional relational database to form a structured electronic test instrument knowledge base, realizing unified knowledge management and dynamic updating.

2. The method for constructing an electronic test instrument knowledge base based on a large language model according to claim 1, characterized in that: In step 1, domain ontology construction includes the following steps: Step 1.1: Use the TF-IDF algorithm and the TextRank algorithm to summarize and extract the core concepts in the field of electronic test instruments; Step 1.2: Perform unsupervised relationship extraction on unstructured or semi-structured data using the Open Information Extraction method (OpenIE) to screen meaningful relationships. Step 1.3: Combine the general knowledge graph ontology and domain expert review to optimize entity types, attribute associations, and relationship semantic expressions, and complete ontology modeling using the Protégé tool; Step 1.4: Form an electronic test instrument technology field ontology covering entity definition, attribute definition, relationship definition and concept hierarchy division.

3. The method for constructing an electronic test instrument knowledge base based on a large language model according to claim 1, wherein: In step 2, the structured design of the prompt word template includes the following steps: Step 2.1: Overview: Define the task context, goals, and the role of the large language model; Step 2.2: Process: Clarify the information extraction steps, rules, and execution order; Step 2.3: Dependencies: Specify the required tools, knowledge bases, and data materials; Step 2.4: Control: Limit the sampling scope, quality requirements and error avoidance rules.

4. The method for constructing an electronic test instrument knowledge base based on a large language model according to claim 1, wherein: In step 3, the large language model information extraction includes the following steps: Step 3.1: Label a small number of sample sets needed for the information extraction process; Step 3.2: Focusing on the task of constructing a knowledge graph in the field of electronic testing instrument technology, further design and construct positive and negative sample sets for the information extraction process; Step 3.3: Train the model based on the labeled positive and negative sample sets to improve the extraction coverage and accuracy of entities, attributes, and relationships.

5. The method for constructing an electronic test instrument knowledge base based on a large language model according to claim 1, characterized in that: In step 4, knowledge storage uses a graph database to store triple relationships and associates it with a relational database to form a multimodal structured knowledge base.

6. The method for constructing an electronic test instrument knowledge base based on a large language model according to claim 1, characterized in that: The large language model is deepseek-r1:7b, which realizes semantic understanding and information extraction in the field of electronic test instrument technology through domain ontology constraints and prompt word guidance.