Information extraction task-oriented cue word design and optimization method and system
By designing and optimizing prompt words based on domain ontology, combined with cognitive linguistics and large language models, the standardization problem of the information extraction system was solved, the accuracy and efficiency of information extraction were improved, and rapid adaptation across tasks and domains was achieved.
Patent Information
- Application Number
- CN202510739421.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing information extraction systems lack a standardized prompt word structure, resulting in inconsistent results, high complexity, poor interoperability, and difficulty in adapting to different tasks and fields. Language ambiguity and ambiguity affect accuracy and efficiency, the model lacks portability, the cost of manual design is high, and it is difficult to quickly adapt to new fields or tasks.
The prompt word design and optimization method based on domain ontology designs prompt words through the principles of cognitive linguistics, combines large language models and ontology databases, performs multi-round dialogue optimization, forms a standardized prompt word structure, and improves semantic understanding and transferability.
It improves the accuracy, reusability and response speed of information extraction tasks, enhances the stability and user-friendliness of the system, reduces development and maintenance costs, and achieves rapid adaptation across tasks and fields.
Smart Images

Figure CN120671823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cognitive intelligence technology, and in particular to a prompt word design and optimization method and system for information extraction tasks. Background Art
[0002] Hint engineering, also known as contextual hinting, is a technology that guides large language models to produce efficient and accurate output by designing, experimenting with, and optimizing input hints for pre-trained language models without updating model weights or parameters. In recent years, research has explored techniques such as zero-shot hinting, small-sample hinting, thought chain hinting, and cognitive hinting. However, current practical applications of hint engineering still rely primarily on manual design and empirical adjustments, lacking a systematic approach to generating and optimizing hints, particularly in specific domains.
[0003] In the field of natural language processing, information extraction is a key technology for converting unstructured text into structured data. Information extraction technology can identify and extract information such as entities, relationships, and events from text, and store it in a structured database for further analysis and application. With the development of artificial intelligence technology, information extraction technology has been widely used in many fields such as finance, medicine, and law. Although existing information extraction systems have made progress in certain areas, they still face many challenges. For example, how to process domain-specific knowledge to improve extraction accuracy, how to adapt to new domains and languages, and how to effectively integrate domain ontologies into the information extraction process. In addition, the performance and efficiency of existing systems still need to be improved when processing large-scale data and real-time information extraction tasks.
[0004] Prompt engineering, as an interactive method relying on natural language, focuses on designing appropriate prompts to guide language models in generating desired outputs. However, existing information extraction systems face the challenge of a lack of standardized prompt word structures. This lack of standardization leads to inconsistent results and makes it difficult to adapt prompt words to different tasks and domains. Without a unified standard, each system or developer may adopt a different prompt word design approach, which not only increases system complexity but also limits interoperability and scalability.
[0005] Furthermore, the lack of a standardized cue word structure makes cross-system migration and integration difficult, as each system requires redesigning and adapting the cue words to suit its specific architecture and needs. This fragmented approach hinders the rapid development and widespread adoption of the technology, requiring extensive duplication of effort and customized development rather than building upon it for innovation. Therefore, developing and promoting a standardized cue word structure is crucial for improving the efficiency, adaptability, and user-friendliness of information extraction systems.
[0006] The ambiguity and polysemy of language are inherent challenges in natural language processing, and they pose a particularly significant challenge in engineering. These characteristics make misunderstandings prone to occur during communication, especially in technical communications and document interpretation, where divergent interpretations can lead to project delays or failure. Furthermore, these characteristics of language severely impact human-computer interoperability. During information extraction and human-computer interaction, machines often struggle to accurately capture the nuances and implicit meanings of language, resulting in the system's inability to correctly understand and respond to user commands or queries. This limitation not only restricts the application scope of artificial intelligence technology but also increases development and maintenance costs, as additional work is required to address and resolve issues arising from linguistic ambiguity. Therefore, developing technologies that can understand and handle linguistic ambiguity and polysemy is crucial for improving the accuracy of information extraction and enhancing the naturalness and efficiency of human-computer interaction.
[0007] For example, the word "load," commonly used in engineering, can mean either "load" or "load," depending on the context. Without further clarification in instructions or documentation, this can lead to misunderstandings among different engineers. Another example is the instruction "start system," which, without clear context, could mean different actions: starting a software system, starting a hardware device, or even starting a network service. Ambiguous instructions can easily lead to misoperation.
[0008] People with different backgrounds and expertise may interpret the same command or technical description differently, leading to misunderstandings and inefficiencies in project communication. Language ambiguity hinders collaboration between different teams on a project, potentially leading to duplication of work, wasted time, or incorrect decisions.
[0009] Another challenge facing existing information extraction technologies is the poor transferability between different information extraction tasks. In practical applications, information extraction tasks often involve diverse scenarios and requirements, such as entity recognition, relationship extraction, and event detection. However, existing information extraction models are often designed for specific tasks, lacking flexibility and generalization capabilities, and are difficult to transfer from one task to another. This limitation means that every time a new task is faced, data collection, model design, and training must be started from scratch, resulting in wasted resources and inefficiency.
[0010] Another consequence of poor transferability is that models struggle to leverage knowledge gained in one domain or task to improve performance in another. This ability to transfer knowledge is crucial for building intelligent systems that can quickly adapt to new environments and tasks. The lack of effective transfer learning strategies limits the application of information extraction techniques in a wider range of scenarios, especially when data resources are limited or tasks change frequently.
[0011] In the field of natural language processing, especially in information extraction tasks, the design of prompt words has a direct impact on model performance. Existing systems often rely on domain experts to manually design prompt words, a time-consuming, labor-intensive, and costly approach. Due to the lack of automated mechanisms, these systems struggle to quickly adapt to new domains or tasks. Each domain change or task update requires redesigning and testing the prompt words, resulting in slow response and an inability to meet rapidly changing market demands. Furthermore, the subjectivity of manually designed prompt words can lead to inconsistent and unstable model performance. Summary of the Invention
[0012] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0013] The present invention proposes a domain ontology-based prompt word design and optimization method for information extraction tasks. This method covers the design of standardized prompt word technology and related system design from a large amount of natural language text at the pragmatic, semantic, and grammatical levels, so as to achieve the output of the large language model more in line with the expected purpose.
[0014] Another object of the present invention is to propose a prompt word design and optimization system for information extraction tasks based on domain ontology.
[0015] To achieve the above objectives, the present invention proposes a method for designing and optimizing prompt words for information extraction tasks based on domain ontology, comprising:
[0016] Preprocess relevant texts in the field to obtain preprocessed text data, and build an environment system between the front-end interface, large language model and ontology database;
[0017] Based on the principles of cognitive linguistics, prompt words are designed. When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the domain ontology according to the prompt words, and then matched with the ontology database to obtain preliminary information matching results;
[0018] Feedback information is obtained through multiple rounds of dialogue results of a large language model and manual proofreading results based on the preliminary information matching results, and the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system are optimized based on the feedback information.
[0019] The domain ontology-based prompt word design and optimization method for information extraction tasks according to an embodiment of the present invention may also have the following additional technical features:
[0020] In one embodiment of the present invention, when a query request is input into the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the domain ontology according to the prompt word, and the information is matched with the ontology database to obtain preliminary information matching results, including:
[0021] Utilize a large language model to receive data requests submitted through a front-end interface; wherein the data requests include product descriptions and query conditions;
[0022] After the large language model receives the data request transmitted from the front end, it performs text processing and information extraction on the input pre-processed text data using pre-designed prompt words to extract preliminary entity information;
[0023] By calling the ontology database, the extracted preliminary entity information is matched and verified with the knowledge in the called ontology database to obtain a preliminary information matching result.
[0024] In one embodiment of the present invention, the method further includes:
[0025] Determine the role of the agent in the conversation or task execution;
[0026] Determine the macro-task context in which the agent performs the task;
[0027] Based on the role and the macro task background, the instruction that best matches the current task is selected from a preset instruction set.
[0028] In one embodiment of the present invention, the method further includes:
[0029] Use a large language model to analyze the received text and use named entity recognition technology to identify key entities in the text;
[0030] Match key entities with concepts in the domain ontology, and determine the attributes corresponding to each entity based on the definitions and relationships in the ontology, so as to determine the type of each entity through the hierarchical structure and classification system in the ontology.
[0031] In one embodiment of the present invention, the method further includes:
[0032] An output template for the output result of the large language model is defined. The output template specifies the structure and format of the output result. The template contains grammatical components, including the subject, predicate, and object for a text field, or the date format for a date field.
[0033] In one embodiment of the present invention, feedback information is obtained from the results of multiple rounds of dialogue using a large language model and the manual proofreading results based on the preliminary information matching results, and the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system are optimized based on the feedback information, including:
[0034] The first round of dialogue is conducted with the large language model. The large language model receives the request information entered by the user and extracts key information to obtain the extraction results.
[0035] The accuracy of the model output is checked by manually reviewing the first round of extraction results to obtain manual proofreading results;
[0036] Improve the prompt words based on the manual proofreading results and input them into the large language model again to obtain the model output results;
[0037] Manually proofread and provide feedback on the model output results;
[0038] The prompt word design is optimized based on the proofread model output results and feedback information, and the semantic network in the environmental system is adjusted in combination with domain ontology knowledge.
[0039] To achieve the above objectives, the present invention further proposes a prompt word design and optimization system for information extraction tasks based on domain ontology, comprising:
[0040] The data preprocessing and environment building module is used to preprocess relevant texts in the field to obtain preprocessed text data, and build an environment system between the front-end interface, large language model and ontology database;
[0041] The information extraction and matching module is used to design prompt words based on the principles of cognitive linguistics. When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the prompt words based on the domain ontology, and then match it with the ontology database to obtain preliminary information matching results;
[0042] The proofreading and optimization module is used to obtain feedback information through the results of multiple rounds of dialogue with the large language model and the manual proofreading results based on the preliminary information matching results, and optimize the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system according to the feedback information.
[0043] The domain ontology-based prompt word design and optimization method and system for information extraction tasks, according to embodiments of the present invention, improve the accuracy, reusability, and responsiveness of information extraction tasks by standardizing prompt words. By applying cognitive linguistic principles, a formalized prompt word structure is formed to enhance the stability of information extraction task output. By leveraging the domain ontology model, a large language model is assisted in accurately understanding the semantics of specialized vocabulary, improving the accuracy of information extraction tasks. Furthermore, a dynamic prompt word generation and adaptation mechanism is established, enhancing the system's reusability across various information extraction tasks.
[0044] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0046] Figure 1 is a flowchart of a method for designing and optimizing prompt words for information extraction tasks based on domain ontology according to an embodiment of the present invention;
[0047] Figure 2 This is a flowchart of prompt word design based on e-commerce information extraction and cognitive linguistics principles according to an embodiment of the present invention;
[0048] Figure 3 4 is a structural diagram of a prompt word design and optimization system for information extraction tasks based on domain ontology according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0050] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0051] The following describes a method and system for designing and optimizing prompt words for information extraction tasks based on domain ontology according to an embodiment of the present invention with reference to the accompanying drawings.
[0052] Figure 1FIG. 1 is a flow chart of a method for designing and optimizing prompt words for information extraction tasks based on domain ontology according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0053] S1, preprocess the relevant texts in the field to obtain preprocessed text data, and build an environment system between the front-end interface, large language model and ontology database;
[0054] S2, based on the principles of cognitive linguistics, designs prompt words. When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the domain ontology according to the prompt words, and then matches it with the ontology database to obtain preliminary information matching results;
[0055] S3, obtaining feedback information through multiple rounds of dialogue results of the large language model and manual proofreading results based on the preliminary information matching results, and optimizing the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system based on the feedback information.
[0056] like Figure 2 The figure below illustrates the workflow for designing formalized prompts, using the e-commerce information extraction task as an example. The solution primarily involves: leveraging cognitive linguistics principles to construct a basic framework for formalized prompts, forming a fundamental structure spanning pragmatics, semantics, and grammar; automating the construction of the pragmatic and semantic layers within the prompts by invoking instruction sets and domain ontologies; and providing rapid feedback through interaction with a large language model and evaluation of extraction results, continuously improving the prompt design and the construction of the semantic network within the system.
[0057] It is understood that this method improves the model's ability to understand and execute specific tasks by designing and adjusting input prompts. This approach combines different levels of linguistics to ensure that the prompt words are not only correct in terms of linguistic structure, but also more consistent with the task requirements at the semantic and pragmatic levels, thereby guiding the large model to produce more accurate and useful results. The specific steps of this method are as follows:
[0058] In one embodiment of the present invention, relevant texts are preprocessed and an environment required for performing the information extraction task is built.
[0059] During the text preprocessing phase, it's important to ensure that the selected text belongs to a specific field and comes from a reliable source. For example, information extraction in the e-commerce field typically targets only specific types of documents, such as product detail pages and user reviews. The scope of these documents should be limited based on the task requirements.
[0060] Considering the performance limitations of large language models, users also need to strictly control the number of input tokens (usually around 8,000 tokens) to ensure that it is within the model's processing capabilities. If it exceeds the limitations of the large language model, the long initial text content should be appropriately segmented.
[0061] After that, the terms should be normalized. For example, "iPhone 13" and "Apple 13" both refer to the same product, and should be treated uniformly to eliminate the interference caused by the diversity of terms.
[0062] To further improve data quality, textual noise should be removed so that the model can focus more on key data features. For example, by pre-defining a set of relevant keywords (such as product, price, brand, and specifications), relevant content can be identified. Other content (such as promotional advertisements and substitute products) is treated as noise and removed. This series of preprocessing steps can provide clean, accurate, and high-quality data for subsequent analysis and model training.
[0063] The purpose of setting up the environment is to ensure the connection between the front-end interface, the large language model, and the ontology database. Based on this environment, the large language model receives data requests submitted through the front-end interface. After receiving the data request from the front-end, the large language model performs text processing and information extraction on the input text using pre-designed prompt words to extract preliminary entity information. By calling the ontology database, the extracted preliminary entity information is matched and verified with the knowledge in the called ontology database to obtain preliminary information matching results.
[0064] Specifically, users submit requests, such as product descriptions and query conditions, through the front-end interface, and the front-end sends the data to the back-end large language model. After receiving the request, the back-end large language model performs text processing and information extraction using the given prompt words. Subsequently, the large language model matches and verifies the extracted entities (such as brand, model, price, etc.) with the knowledge in the ontology database by calling the ontology database to ensure data consistency and accuracy. At the same time, the ontology database can also provide domain knowledge reasoning to help the model better understand and extract information. For example, query the product category by brand name, or filter products by price range. Finally, the extraction results are returned to the front-end interface for display to the user. This process ensures the efficiency, accuracy and domain consistency of the information extraction task.
[0065] In one embodiment of the present invention, the prompt word structure is designed based on the principles of cognitive linguistics to form prompts at different levels of abstraction for information extraction tasks.
[0066] Exemplarily, the role of the agent in the dialogue or task execution is determined; the macro-task context when the agent performs the task is determined; and the instruction that best matches the current task is selected from a preset instruction set based on the role and the macro-task context.
[0067] Specifically, at the pragmatic level, the agent's role must first be determined; then a broad task context must be provided; and finally, the specific task to be performed is extracted from the instruction set. In this step, the innovative approach of assigning agent roles based on the instruction set is proposed. This approach helps the agent use appropriate language style and intonation, as well as provide relevant domain expertise, thereby improving the naturalness and efficiency of the interaction.
[0068] Exemplarily, a large language model is used to analyze the received text, and named entity recognition technology is used to identify key entities in the text; the key entities are matched with concepts in the domain ontology, and the attributes corresponding to each entity are determined based on the definitions and relationships in the ontology, so as to judge the type of each entity through the hierarchical structure and classification system in the ontology.
[0069] Specifically, at the semantic level, domain ontologies are used to help refine relevant semantics (e.g., definitions of specialized terms and axioms between specialized terms), generating more complete text to enhance the model's understanding of the textual content. In this step, key entities in the text must first be identified. By matching these entities with concepts in the ontology, the model can determine which category each entity belongs to and what its corresponding attributes are, which in turn supplements and improves the original textual content. The main innovation lies in the integration of semantically enhanced understanding technology based on domain ontologies.
[0070] Exemplarily, an output template of the output result of the large language model is defined, where the output template specifies the structure and format of the output result. The template contains grammatical components including the subject, predicate and object for a text field, or the date format for a date field.
[0071] Specifically, at the grammatical level, clear output templates need to be defined. These templates specify the structure and format of the output results. Templates should include necessary grammatical components, such as the subject, predicate, and object for text fields, or the date format (e.g., YYYY-MM-DD, MM / DD / YYYY, etc.) for date fields to maintain consistency while allowing for flexible adaptation to different content.
[0072] The prompt word structure design of the embodiment of the present invention realizes flexible format output that cannot be achieved by traditional technologies, and improves the portability of large language models.
[0073] In one embodiment of the present invention, a feedback mechanism is formed through multiple rounds of dialogue with a large language model and the results of manual proofreading to continuously optimize the prompt word design and improve the semantic network built based on the domain ontology.
[0074] Specifically, during the first round of conversation with the model, the model will extract information based on the initial input requirements. For example, if you enter a product description, the system will try to extract basic information such as product name, price, brand, and category.
[0075] Manual proofreading is done by manually reviewing the first round of extraction results to check the accuracy of the model output. For example, some product information may not be accurately extracted, or some field formats may not meet requirements. Proofreaders will modify the results, point out errors or incompleteness in the model, and provide improvement requirements.
[0076] Based on the results of the previous round of manual proofreading, the prompt wording is improved and fed back into the model to obtain the model output. For example, the prompt wording can be changed to "Ensure that the output contains the product brand, price, color, etc., and is returned in a specific format."
[0077] After obtaining the model output from the previous step, manual proofreaders review the model output again and conduct final verification to ensure that the model-generated results meet domain requirements. If there are new field requirements or additions, proofreaders will further refine the prompt words and semantic network.
[0078] Based on the output and feedback from the previous round, further optimize the prompt word design and adjust the semantic network in the environment system by combining domain ontology knowledge. Repeat the above steps until the final extraction results reach the ideal state in terms of accuracy, format consistency, and business relevance.
[0079] According to the prompt word design and optimization method for information extraction tasks based on domain ontology according to an embodiment of the present invention, the role of the intelligent agent is set based on the instruction set, that is, the role played by the intelligent agent is determined by the specific application field and the instructions in the instruction set. Semantic understanding is enhanced based on the domain ontology, and the original text is completed or corrected by automatically identifying relevant information in the domain ontology to obtain good semantics. Output templates are defined to achieve flexible output and realize the portability of large language models. At the same time, the principles of cognitive linguistics are used to form a formal prompt word structure to improve the stability of the output of information extraction tasks; with the help of the domain ontology model, the large language model is assisted in the accurate semantic understanding of professional vocabulary, thereby improving the accuracy of information extraction tasks; a dynamic generation and adaptation mechanism of prompt words is formed to improve the reusability of the system between various information extraction tasks.
[0080] In order to implement the above embodiment, Figure 3As shown, this embodiment also provides a prompt word design and optimization system 10 for information extraction tasks based on domain ontology, including:
[0081] The data preprocessing and environment building module 100 is used to preprocess relevant texts in the field to obtain preprocessed text data, and to build an environment system between the front-end interface, the large language model and the ontology database;
[0082] The information extraction and matching module 200 is used to design prompt words based on the principles of cognitive linguistics. When a query request is input into the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the prompt words based on the domain ontology, and then match it with the ontology database to obtain preliminary information matching results;
[0083] The proofreading and optimization module 300 is used to obtain feedback information through the results of multiple rounds of dialogue of the large language model and the manual proofreading results based on the preliminary information matching results, and optimize the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system according to the feedback information.
[0084] Furthermore, the information extraction and matching module 200 is further configured to:
[0085] Utilize a large language model to receive data requests submitted through a front-end interface; wherein the data requests include product descriptions and query conditions;
[0086] After the large language model receives the data request transmitted from the front end, it performs text processing and information extraction on the input pre-processed text data using pre-designed prompt words to extract preliminary entity information;
[0087] By calling the ontology database, the extracted preliminary entity information is matched and verified with the knowledge in the called ontology database to obtain a preliminary information matching result.
[0088] Furthermore, it is also used for:
[0089] Determine the role of the agent in the conversation or task execution;
[0090] Determine the macro-task context in which the agent performs the task;
[0091] Based on the role and the macro task background, the instruction that best matches the current task is selected from a preset instruction set.
[0092] Furthermore, it is also used for:
[0093] Use a large language model to analyze the received text and use named entity recognition technology to identify key entities in the text;
[0094] Match key entities with concepts in the domain ontology, and determine the attributes corresponding to each entity based on the definitions and relationships in the ontology, so as to determine the type of each entity through the hierarchical structure and classification system in the ontology.
[0095] According to the prompt word design and optimization system for information extraction tasks based on domain ontology according to an embodiment of the present invention, the role of the intelligent agent is set based on the instruction set, that is, the role played by the intelligent agent is determined by the specific application field and the instructions in the instruction set. Semantic understanding is enhanced based on the domain ontology, and the original text is completed or corrected by automatically identifying relevant information in the domain ontology to obtain good semantics. Output templates are defined to achieve flexible output and realize the portability of large language models. At the same time, the principles of cognitive linguistics are used to form a formal prompt word structure to improve the stability of the output of information extraction tasks; with the help of the domain ontology model, the large language model is assisted in the accurate semantic understanding of professional vocabulary, thereby improving the accuracy of information extraction tasks; a dynamic generation and adaptation mechanism of prompt words is formed to improve the reusability of the system between various information extraction tasks.
[0096] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0097] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A method for designing and optimizing prompt words for information extraction tasks based on domain ontology, characterized in that: include: Preprocess relevant texts in the field to obtain preprocessed text data, and build an environment system between the front-end interface, large language model and ontology database; Based on the principles of cognitive linguistics, prompt words are designed. When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the domain ontology according to the prompt words, and then matched with the ontology database to obtain preliminary information matching results; Feedback information is obtained through multiple rounds of dialogue results of a large language model and manual proofreading results based on the preliminary information matching results, and the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system are optimized based on the feedback information.
2. The method according to claim 1, characterized in that When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the domain ontology according to the prompt word, and then matched with the ontology database to obtain preliminary information matching results, including: Utilize a large language model to receive data requests submitted through a front-end interface; wherein the data requests include product descriptions and query conditions; After the large language model receives the data request transmitted from the front end, it performs text processing and information extraction on the input pre-processed text data using pre-designed prompt words to extract preliminary entity information; By calling the ontology database, the extracted preliminary entity information is matched and verified with the knowledge in the called ontology database to obtain a preliminary information matching result.
3. The method according to claim 1, characterized in that The method further comprises: Determine the role of the agent in the conversation or task execution; Determine the macro-task context in which the agent performs the task; Based on the role and the macro task background, the instruction that best matches the current task is selected from a preset instruction set.
4. The method according to claim 1, wherein The method further comprises: Use a large language model to analyze the received text and use named entity recognition technology to identify key entities in the text; Match key entities with concepts in the domain ontology, and determine the attributes corresponding to each entity based on the definitions and relationships in the ontology, so as to determine the type of each entity through the hierarchical structure and classification system in the ontology.
5. The method according to claim 1, wherein The method further comprises: An output template for the output result of the large language model is defined. The output template specifies the structure and format of the output result. The template contains grammatical components, including the subject, predicate, and object for a text field, or the date format for a date field.
6. The method according to claim 1, characterized in that Feedback information is obtained through the results of multiple rounds of dialogue using a large language model and the manual proofreading results based on the preliminary information matching results. The design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system are optimized based on the feedback information, including: The first round of dialogue is conducted with the large language model. The large language model receives the request information entered by the user and extracts key information to obtain the extraction results. The accuracy of the model output is checked by manually reviewing the first round of extraction results to obtain manual proofreading results; Improve the prompt words based on the manual proofreading results and input them into the large language model again to obtain the model output results; Manually proofread and provide feedback on the model output results; The prompt word design is optimized based on the proofread model output results and feedback information, and the semantic network in the environmental system is adjusted in combination with domain ontology knowledge.
7. A prompt word design and optimization system for information extraction tasks based on domain ontology, characterized by: include: The data preprocessing and environment building module is used to preprocess relevant texts in the field to obtain preprocessed text data, and build an environment system between the front-end interface, large language model and ontology database; The information extraction and matching module is used to design prompt words based on the principles of cognitive linguistics. When a query request is entered on the front-end interface, the pre-processed text data is sent to the large language model to extract information based on the prompt words based on the domain ontology, and then match it with the ontology database to obtain preliminary information matching results; The proofreading and optimization module is used to obtain feedback information through the results of multiple rounds of dialogue with the large language model and the manual proofreading results based on the preliminary information matching results, and optimize the design structure of the prompt words and the semantic network constructed based on the domain ontology in the optimization environment system according to the feedback information.
8. The system according to claim 7, characterized in that The information extraction and matching module is also used for: Utilize a large language model to receive data requests submitted through a front-end interface; wherein the data requests include product descriptions and query conditions; After the large language model receives the data request transmitted from the front end, it performs text processing and information extraction on the input pre-processed text data using pre-designed prompt words to extract preliminary entity information; By calling the ontology database, the extracted preliminary entity information is matched and verified with the knowledge in the called ontology database to obtain a preliminary information matching result.
9. The system according to claim 7, wherein: Also used for: Determine the role of the agent in the conversation or task execution; Determine the macro-task context in which the agent performs the task; Based on the role and the macro task background, the instruction that best matches the current task is selected from a preset instruction set.
10. The system according to claim 7, wherein: Also used for: Use a large language model to analyze the received text and use named entity recognition technology to identify key entities in the text; Match key entities with concepts in the domain ontology, and determine the attributes corresponding to each entity based on the definitions and relationships in the ontology, so as to determine the type of each entity through the hierarchical structure and classification system in the ontology.
Citation Information
Patent Citations
Prompt word optimization method and device, storage medium and program product
CN119088936A
Question and answer method and device based on large language model, electronic equipment and computer program product
CN119886332A
Extracting information from reports using large language models
US20240411994A1
Task-type dialogue response method and apparatus
WO2025107850A1
Cited By
Bidding document information extraction method
CN121031593A
A method for extracting information from bidding documents
CN121031593B
Literature processing method and device, equipment, storage medium and program product
CN121636704A
General field scientific data extraction method and system for science and technology papers
CN121960512A
A general extraction method and system for scientific data in the field of scientific papers
CN121960512B