A method and device for constructing an interactive task knowledge base based on a large language model
Through the interactive task knowledge base construction method based on large language model, the problem of low efficiency and insufficient accuracy of task knowledge base construction in the existing technology is solved, and efficient and accurate automated construction is achieved to adapt to the needs of different application scenarios.
Patent Information
- Application Number
- CN202510075297.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the prior art, task knowledge base construction is inefficient and insufficiently accurate, making it difficult to adapt to rapidly changing needs, especially when switching scenarios in semantic communication systems, the knowledge base needs to be reconstructed.
An interactive task knowledge base construction method based on a large language model is adopted, a task corpus request instruction is generated through the first keyword prompt template, a task corpus is obtained using the large language model, and a text comparison analysis is used to determine whether the corpus meets the preset requirements. If not satisfied, generate fine-tuning instructions to adjust until the requirements are met.
It realizes efficient and precise automation of the task knowledge base, improves adaptability to different application scenarios, reduces the dependence on the manual participation of experts in the field, and improves the repeatability and scalability of the construction process.
Smart Images

Figure CN119513228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semantic communication technology, and in particular to a method and device for constructing an interactive task knowledge base based on a large language model. Background Art
[0002] As an emerging communication technology, semantic communication, with its unique advantages, demonstrates tremendous potential and is a key development direction in the future communications field. Task knowledge bases, as a key component, can significantly enhance the performance of semantic communication systems, enabling them to better understand and respond to communication needs. These bases are crucial for achieving efficient and accurate information transfer in semantic communication. Therefore, solving the problem of building knowledge bases is crucial for the current development and application of semantic communication technology.
[0003] Traditional knowledge base construction methods include manual, automatic, and semi-automatic construction. While manual construction ensures knowledge accuracy and authority, it relies heavily on the expertise and subjective judgment of domain experts. This approach suffers from low construction efficiency and is difficult to adapt to the large-scale, fast-paced demands of knowledge updating. Furthermore, due to the limitations of experts' knowledge and subjective biases, the knowledge base may be limited in coverage and lack objectivity. Automated construction methods use information extraction techniques to automatically extract knowledge from massive amounts of data, alleviating the manual burden. However, when faced with text data with diverse structures and complex semantics, automatic construction methods struggle to accurately identify entities, attributes, and relationships, and are unable to handle complex situations such as implicit relationships and ambiguous expressions. Semi-automatic construction methods combine the advantages of manual and automatic methods, but still suffer from inefficiency and error-proneness when processing large amounts of complex data. However, task knowledge bases are a key component of task-oriented semantic communication systems, and they often need to be rebuilt when switching scenarios. Therefore, optimizing the efficiency and accuracy of task knowledge base construction is a key component of the development and application of semantic communication technology. Summary of the Invention
[0004] The present invention provides a method and device for constructing an interactive task knowledge base based on a large language model, which is used to solve the problems of low construction efficiency, insufficient accuracy and difficulty in adapting to rapidly changing needs in the existing technology, realize efficient and accurate automatic construction of the task knowledge base, and improve its adaptability to different application scenarios.
[0005] In a first aspect, the present invention provides a method for constructing an interactive task knowledge base based on a large language model, comprising: generating a task corpus request instruction for a target task based on a first keyword prompt template, inputting the task corpus request instruction into the large language model, and obtaining the task corpus for the target task; matching a target standard task corpus from a preset standard task corpus according to the task corpus request instruction, performing text comparison analysis on the task corpus using the target standard task corpus, and judging whether the task corpus meets the preset task corpus requirements; when it is determined that the task corpus meets the preset task corpus requirements, generating an entity extraction instruction based on a second keyword prompt template, and inputting the entity extraction instruction into the large language model to obtain key semantic information of the task corpus; the key semantic information comprises at least one entity information, and each entity information includes an entity target focused on by the task and a task relevance score of the entity target; based on the key semantic information of the task corpus, generating a task knowledge base for the task target using the large language model.
[0006] According to a method for constructing an interactive task knowledge base based on a large language model provided by the present invention, a target standard task corpus is matched from a preset standard task corpus according to a task corpus request instruction, and a text comparison analysis is performed on the task corpus using the target standard task corpus to determine whether the task corpus meets the preset task corpus requirements, including: parsing the task corpus request instruction and extracting matching information; the matching information includes at least one of the following information: the task subject of the target task and keywords of the task content; matching the target standard task corpus from the standard task corpus based on the matching information; extracting feature vectors from the task corpus and the target standard task corpus, and calculating the similarity of the feature vectors of the two; when the similarity of the feature vectors is greater than or equal to a preset threshold, determining that the task corpus meets the preset task corpus requirements; when the similarity of the feature vectors is less than the preset threshold, determining that the task corpus does not meet the preset task corpus requirements.
[0007] According to a method for constructing an interactive task knowledge base based on a large language model provided by the present invention, task corpus for a target task is obtained, including: determining the source of the task corpus; determining the initial publisher based on the source of the task corpus, and determining whether the initial publisher is within a preset publisher range; eliminating task corpus whose initial publisher is not within the preset publisher range, and retaining task corpus whose initial publisher is within the preset publisher range.
[0008] According to a method for constructing an interactive task knowledge base based on a large language model provided by the present invention, when it is determined that the task corpus does not meet the preset task corpus requirements, a third keyword prompt template is used to generate fine-tuning instructions, and the fine-tuning instructions are input into the large language model to fine-tune the large language model until the task corpus output by the large language model meets the preset task corpus requirements; wherein the third keyword prompt template is used to describe the fine-tuning criteria for the large language model.
[0009] According to a method for constructing an interactive task knowledge base based on a large language model provided by the present invention, after obtaining the key semantic information of the task corpus, the method further includes: using a pre-built classification model to classify the entity information in the key semantic information; the classification types include entity omission, entity misrecognition, entity misrecognition, task relevance score misrecognition and correct classification; based on the classification results, a fourth keyword prompt template is generated, and the large language model is fine-tuned so that the proportion of entity information in the key semantic information output by the large language model that belongs to the correct classification is greater than a preset proportion.
[0010] According to the present invention, a method for constructing an interactive task knowledge base based on a large language model is provided. Based on the key semantic information of the task corpus, a task knowledge base for the task target is generated using the large language model, including: generating program writing instructions based on a fifth keyword prompt template, inputting the program writing instructions into the large language model to obtain a target program; executing the target program; and using the target program to generate a visual task knowledge base based on the key semantic information.
[0011] According to the present invention, a method for constructing an interactive task knowledge base based on a large language model also includes: testing the visualized task knowledge base; if there is no problem with the test result, completing the construction of the task knowledge base; if there is a problem with the test result, fine-tuning the large language model.
[0012] In a second aspect, the present invention further provides an interactive task knowledge base construction device based on a large language model, comprising:
[0013] A first processing module is configured to generate a task corpus request instruction for a target task based on the first keyword prompt template, input the task corpus request instruction into the large language model, and obtain the task corpus for the target task;
[0014] The second processing module is used to match the target standard task corpus from the preset standard task corpus according to the task corpus request instruction, perform text comparison analysis on the task corpus using the target standard task corpus, and determine whether the task corpus meets the preset task corpus requirements;
[0015] The third processing module is used to generate an entity extraction instruction based on the second keyword prompt template when it is determined that the task corpus meets the preset task corpus requirements, and input the entity extraction instruction into the large language model to obtain key semantic information of the task corpus; the key semantic information is at least one entity information,
[0016] Each entity information includes the entity target of the task and the task relevance score of the entity target;
[0017] The fourth processing module is used to generate a task knowledge base of the task target based on the key semantic information of the task corpus using a large language model.
[0018] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for constructing an interactive task knowledge base based on a large language model as described above are implemented.
[0019] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for constructing an interactive task knowledge base based on a large language model.
[0020] The method and device for constructing an interactive task knowledge base based on a large language model provided by the present invention have the following advantages compared with the prior art:
[0021] The method and device for constructing an interactive task knowledge base based on a large language model provided by the present invention have the following significant beneficial effects compared to the existing technology:
[0022] (1) Through automation and intelligent means, with the help of large language models, the need for human intervention is greatly reduced, thereby accelerating the construction process of the task knowledge base. Compared with traditional manual or semi-automatic methods, this method can process large-scale data more quickly and adapt to the fast-paced knowledge update needs.
[0023] (2) By leveraging the powerful natural language processing capabilities of the large language model and the prompt word engineering, we can more accurately parse text content and extract task-related entities, attributes, and relationships, ensuring the high quality of the generated task knowledge base information. In addition, by performing standard comparison analysis on the generated task corpus, its accuracy and consistency are further guaranteed.
[0024] (3) Through automation and semi-automation, the technology of this application reduces the dependence on manual participation of domain experts, reduces construction costs, and improves the repeatability and scalability of the construction process.
[0025] (4) Introducing a user feedback loop allows end users to directly participate in the optimization process of the knowledge base. This not only improves the practicality of the system, but also enhances user satisfaction and participation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 It is a flowchart of the method for constructing an interactive task knowledge base based on a large language model provided by the present invention;
[0028] Figure 2 This is a schematic diagram of a framework for establishing a disaster awareness task knowledge base in a drone aerial photography scenario provided by the present invention;
[0029] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0031] It should be noted that, in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include a ..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0032] The terms "first," "second," and the like in this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, so that embodiments of this application can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and the like generally distinguish objects of a class and do not limit the number of objects; for example, the first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.
[0033] The following combination Figure 1-Figure 3 The method and device for constructing an interactive task knowledge base based on a large language model provided by an embodiment of the present invention are described.
[0034] Figure 1 This is a flow chart of the method for constructing an interactive task knowledge base based on a large language model provided by the present invention. Figure 1 As shown, including but not limited to the following steps:
[0035] Step 101: Generate a task corpus request instruction for a target task based on a first keyword prompt template, input the task corpus request instruction into a large language model, and obtain the task corpus for the target task;
[0036] The first keyword prompt template is used to describe the requirements of the target task. For example, please collect news reports and official announcements related to earthquake drone post-disaster rescue.
[0037] To further refine the process of acquiring task corpora for the target task, this paper introduces a verification step for the source and initial publisher of the task corpora. This process ensures that the acquired task corpora are not only relevant to the target task, but also come from reliable, pre-defined publishers. Specifically, the following steps are included:
[0038] (1) Determine the source of the task corpus:
[0039] Identify and determine all potential data sources that may contain relevant information for the target task. These sources may include but are not limited to public online resources, internal databases, professional literature repositories, social media platforms, etc.
[0040] These data sources are classified according to different criteria (such as reliability, authority, timeliness, etc.) for subsequent screening.
[0041] (2) Determine the initial publisher based on the source of the task corpus and judge whether it is within the preset publisher range:
[0042] For each task corpus in the data source, trace it back to its original publisher. This can be achieved by parsing metadata, looking up author information, or using other identifiers.
[0043] The identified initial publisher is compared with a pre-defined publisher range. The pre-defined publisher range can be a trusted list or a blacklist based on factors such as organizational structure, individual qualifications, and historical contributions.
[0044] (3) Eliminate the task corpus whose initial publisher is not within the pre-defined publisher range, and retain the task corpus whose initial publisher is within the pre-defined publisher range:
[0045] The present invention can implement strict filtering rules to retain only task corpora from a preset range of publishers.
[0046] Through this screening method, it is ensured that the task corpus ultimately used to construct the task knowledge base has a high degree of credibility and authority, thereby improving the quality and reliability of the entire knowledge base.
[0047] Step 102: matching target standard task corpus from a preset standard task corpus according to the task corpus request instruction, performing text comparison analysis on the task corpus using the target standard task corpus, and determining whether the task corpus meets the preset task corpus requirements;
[0048] As an optional embodiment, step 102 can be further refined into the following specific steps to improve the accuracy and reliability of the acquired task corpus:
[0049] (1) Parse the task corpus request instructions and extract matching information:
[0050] The matching information includes at least one of the following information: a task theme of the target task and keywords of the task content.
[0051] The present invention can analyze the natural language in the request instruction and extract key phrases or sentences describing the core content of the task as the task subject.
[0052] The present invention can utilize text processing technology (such as TF-IDF algorithm) to automatically extract keywords of task content from task corpus request instructions.
[0053] (2) Match the target standard task corpus from the standard task corpus based on the matching information:
[0054] The present invention can use the extracted task themes and keywords as query conditions to quickly search a pre-set standard task corpus. The corpus should have been indexed and optimized so that the most relevant target standard task corpus can be found efficiently.
[0055] Optionally, if there are too many preliminary search results, the candidate set can be narrowed down by adding more filtering conditions (such as time range, source type, etc.) to ensure that the target standard task corpus finally selected is as close to the requirements of the target task as possible.
[0056] (3) Extract feature vectors from the task corpus and the target standard task corpus, and calculate the similarity of their feature vectors:
[0057] For the task corpus and the corresponding target standard task corpus, natural language processing technology and machine learning models (such as Word2Vec and BERT) are applied to convert the text into numerical feature vectors. These feature vectors can capture the semantic information of the text, including text structure and contextual relationships.
[0058] Use an appropriate similarity measurement method (such as cosine similarity) to calculate the feature vector similarity between the task corpus and its matched target standard task corpus.
[0059] (4) Determine whether the preset task corpus requirements are met:
[0060] Based on a pre-set similarity threshold, determine whether the calculated feature vector similarity reaches or exceeds this threshold. If the similarity is greater than or equal to the preset threshold (similarity threshold), the task corpus is considered to meet the preset task corpus requirements; conversely, if the similarity is lower than the preset threshold, it is considered not to meet the requirements.
[0061] Through the above steps, the present invention not only enhances the accuracy of task corpus matching but also introduces automated and intelligent means to ensure the efficiency of the matching process. This method is particularly suitable for processing large-scale text data and can maintain high accuracy and reliability in complex application scenarios. In addition, by dynamically adjusting the similarity threshold and other parameters, the system can continuously optimize its performance to meet the specific needs of different tasks.
[0062] Step 103: When it is determined that the task corpus meets the preset task corpus requirements, an entity extraction instruction is generated based on the second keyword prompt template, and the entity extraction instruction is input into the large language model to obtain key semantic information of the task corpus; the key semantic information is at least one entity information, and each entity information includes an entity target that the task focuses on and a task relevance score of the entity target.
[0063] The second keyword prompt template is used to describe the extraction method and output form of key semantic information, for example:
[0064] Please refer to the case and use [] to mark the entities in the corpus, and summarize the typical tasks in the corpus in the form of triples of task execution subject-entity target-score, task focus entity target, and task relevance score of the target. The case is as follows: At 10:01 on September 6, [Model A UAV] flew to the disaster area again to scout the B River tributary [Lake C], [house collapse], [tunnel collapse], [water pipe rupture], [landslide], etc., and completed the communication guarantee and real-time transmission of audio and video information in areas such as D Township and E Village, and realized the real-time transmission of disaster survey information back to the command center of the Ministry of Emergency Management, providing strong support for the scientific and timely deployment of earthquake relief work. The key semantic information in the form of triples is as follows:
[0065] Model A UAV-C Lake-8
[0066] Model A UAV-collapsed house-10
[0067] Model A UAV-collapsed tunnel-6
[0068] Model A UAV-Broken Water Pipe-4
[0069] Model A UAV-Landslide Mountain-5
[0070] As an optional embodiment, when it is determined that the task corpus does not meet the preset task corpus requirements, a third keyword prompt template is used to generate fine-tuning instructions, and the fine-tuning instructions are input into the large language model to fine-tune the large language model until the task corpus output by the large language model meets the preset task corpus requirements; wherein the third keyword prompt template is used to describe the fine-tuning criteria for the large language model.
[0071] Specifically, the present invention designs and uses a third keyword prompt template based on the differences between the characteristics of the unverified task corpus and the preset task corpus requirements. This template should detail aspects that need adjustment, such as domain-specific terminology, depth of contextual understanding, and entity recognition accuracy.
[0072] Based on the third keyword hint template, a set of detailed fine-tuning instructions are automatically generated. These instructions are designed to guide the large language model on how to improve its output to better match the requirements of the target task.
[0073] The generated fine-tuning instructions are fed into the large language model as input, triggering its internal parameter adjustment mechanism. This establishes an iterative feedback loop, allowing the system to reassess whether the task corpus meets the preset requirements based on the output after each fine-tuning. If it still does not meet the requirements, new fine-tuning instructions are generated and the above process is repeated.
[0074] By introducing this dynamic fine-tuning mechanism, the present invention can not only effectively solve the problem of the initial generation of task corpus not meeting the requirements, but also provide an opportunity for continuous self-optimization for the large language model.
[0075] For example, if the task corpus "More than a decade ago, we could only manually take local photos to understand the disaster situation. Now, we not only have a panoramic view of the disaster, but also it only takes a dozen minutes at the fastest" lacks key data and is determined not to meet the preset task corpus requirements, then a third keyword prompt template is designed:
[0076] The above data lacks key information. Please collect data related to the work completed by drones in post-disaster rescue, such as "At 10:01 on September 6, Model A drone flew to the disaster area again, reconnoitering Lake C, a tributary of River B, collapsed buildings, collapsed tunnels, broken water pipes, and landslides. It also provided communication support and real-time transmission of audio and video information in areas such as Township D and Village E. This enabled real-time transmission of disaster survey results back to the Ministry of Emergency Management's command center, providing strong support for the scientific and timely deployment of earthquake relief efforts."
[0077] Step 104: Based on the key semantic information of the task corpus, a task knowledge base of the task target is generated using the large language model.
[0078] Optionally, in step 104, the process of generating a task knowledge base of the task target using the large language model based on the key semantic information of the task corpus can be further refined into the following specific steps:
[0079] (1) Generate programming instructions based on the fifth keyword prompt template
[0080] Design of the fifth keyword prompt template: Based on the task requirements and the characteristics of key semantic information, design and use the fifth keyword prompt template. This template should describe in detail the program logic, functional modules, and visualization requirements to be generated.
[0081] (2) Input the programming instructions into the large language model to obtain the target program
[0082] The generated program writing instructions are provided as input to a pre-trained large language model; based on the received instructions, the large language model automatically generates one or more target programs that meet the task requirements.
[0083] (3) Execute the target program
[0084] Set up an appropriate runtime environment, including but not limited to installing necessary dependency libraries and configuring the server, to ensure that the target program can run properly in the specified environment. For example, set up the runtime environment for a Python program.
[0085] Start the target program and let it generate a task knowledge base based on key semantic information. During this process, the program automatically completes data loading, processing, and analysis, and constructs a structured knowledge representation.
[0086] (4) Generate a visual task knowledge base
[0087] After the target program is executed, one or more visual interfaces will be generated. These interfaces can be in the form of charts, tables, graphical networks, etc., which are used to display the information in the task knowledge base.
[0088] Through the above steps, the present invention not only realizes the automatic generation of task knowledge base from key semantic information, but also introduces visualization elements, which greatly enhances the readability and practicality of the knowledge base.
[0089] Based on the content of the above embodiment, as an optional embodiment, the method for constructing an interactive task knowledge base based on a large language model provided by the present invention further includes:
[0090] (1) Testing the visualized task knowledge base;
[0091] The present invention can perform knowledge question and answer on the task knowledge base based on the large language model to realize the testing of the task knowledge base.
[0092] (2) If the test results are normal, complete the construction of the task knowledge base; if there are problems with the test results, fine-tune the large language model.
[0093] If the output result is correct, it is determined that there is no problem; otherwise, there is a problem with the test and the large language model needs to be fine-tuned based on the existing problem.
[0094] Based on the above embodiments, as an optional embodiment, the present invention provides a method for constructing an interactive task knowledge base based on a large language model. After obtaining the key semantic information of the task corpus, it further introduces entity information classification and targeted fine-tuning mechanisms. The following is a description of the specific steps of this embodiment:
[0095] (1) Use the pre-built classification model to classify the entity information in the key semantic information:
[0096] Use a pre-trained classification model to analyze entity information from key semantic information obtained from a large language model. The classification model should be able to identify multiple classification types, including but not limited to:
[0097] Entity omission: Failure to identify task-related entities.
[0098] Entity misidentification: An item that is incorrectly identified as an entity but is not.
[0099] Entity misidentification: An entity that is incorrectly identified, such as misclassifying one entity as another.
[0100] Task relevance score misidentification: The task relevance score of the entity is inaccurate.
[0101] Correctly classified: An entity that is accurately identified and classified.
[0102] Classification result output: The classification model outputs the classification label of each entity information and records these labels for subsequent processing.
[0103] The classification model is pre-trained based on entity information samples and corresponding sample labels.
[0104] (2) Generate the fourth keyword prompt template based on the classification results:
[0105] Based on the classification results, analyze the main issues with entity recognition in current large language models. For example, if a large number of entities are missed, the model's understanding of domain-specific terminology needs to be strengthened; if there are many entity misidentifications, the context parsing capability may need to be improved.
[0106] Based on the above analysis, design and generate a fourth keyword suggestion template. This template will guide how to adjust the large language model to reduce or eliminate the identified issues in future outputs. The template should be specific and include detailed adjustment directions and the desired results.
[0107] (3) Fine-tune the large language model:
[0108] The fourth keyword suggestion template is converted into specific fine-tuning instructions and fed into the large language model to initiate the fine-tuning process. A feedback loop is established, allowing the system to re-evaluate the key semantic information output after each fine-tuning, particularly the classification of the entities within it. Through continuous iteration, the proportion of correctly classified entities is gradually increased until it exceeds the preset percentage.
[0109] Through this process of continuous learning and self-optimization, the present invention not only solves the inaccuracy problem that may occur when the task corpus and its key semantic information are initially generated, but also enhances the understanding ability and adaptability of the large language model for specific task scenarios.
[0110] Figure 2 This is a schematic diagram of the framework for establishing a disaster awareness task knowledge base in the drone aerial photography scenario provided by the present invention, such as Figure 2 As shown, the technical solution of the present invention is further described below in conjunction with specific application scenarios.
[0111] Step 1: Propose initial requirements to the large language model based on the key word template.
[0112] Corresponding keyword template: Please collect news reports and official announcements related to earthquake drone post-disaster rescue.
[0113] Step 2: Perform a sample review of the corpus output by the large language model and check whether it meets the preset task corpus requirements in combination with the target standard task corpus. If so, proceed to Step 4; otherwise, continue to Step 3.
[0114] Step 3: It was discovered that some of the corpus lacked key data (over a decade ago, we could only manually capture partial photos to understand the disaster situation. Now, we have panoramic images of the disaster, and they can be captured in just over ten minutes). The operator then proposed fine-tuning criteria for the large language model based on the key word template. Then, we proceeded to Step 2.
[0115] Corresponding key word template:
[0116] The above data lacks key information. Please collect data related to the work completed by drones in post-disaster rescue, such as "At 10:01 on September 6, Model A drone flew to the disaster area again, reconnoitering Lake C, a tributary of River B, collapsed buildings, collapsed tunnels, broken water pipes, and landslides. It also provided communication support and real-time transmission of audio and video information in areas such as Township D and Village E. This enabled real-time transmission of disaster survey results back to the Ministry of Emergency Management's command center, providing strong support for the scientific and timely deployment of earthquake relief efforts."
[0117] Step 4: Propose the initial requirement of entity extraction to the large language model based on the key word template.
[0118] Corresponding key word template:
[0119] Please refer to the example and use [] to mark the entities in the corpus, and summarize the task execution subject, task focus entity target, and task relevance score of the target in the corpus in the form of a triple of task execution subject-entity target-score. The example is as follows: At 10:01 on September 6, [Model A UAV] flew to the disaster area again to scout the B River tributary [Lake C], [house collapse], [tunnel collapse], [water pipe rupture], [landslide], etc., and completed the communication guarantee and real-time transmission of audio and video information in areas such as D Township and E Village, and realized the real-time transmission of disaster survey information back to the command center of the Ministry of Emergency Management, providing strong support for the scientific and timely deployment of earthquake relief work. The key semantic information in the form of triples is as follows:
[0120] Model A UAV-C Lake-8
[0121] Model A UAV-collapsed house-10
[0122] Model A UAV-collapsed tunnel-6
[0123] Model A UAV-Broken Water Pipe-4
[0124] Model A UAV-Landslide Mountain-5
[0125] Step 5: Use the classification model to sample and review the triples output by the large language model to check for problems. If there are no problems, proceed to Step 7; otherwise, continue to Step 6.
[0126] Step 6: Review reveals that some triples (drone-real-time imagery-10, drone-communication service-8) do not directly represent the operator of the drone in the aerial photography. Based on the key word template, fine-tune the large language model. Then proceed to Step 5.
[0127] Corresponding key word template:
[0128] The above triples contain entities that cannot be seen in drone aerial imagery. Please adjust the triples to ensure that the entities are all targets that can be directly seen in drone aerial imagery, such as Lake C, collapsed houses, damaged roads, and broken water pipes.
[0129] Step 7: Propose programming requirements to the large language model based on the key word template.
[0130] Keyword template: Please write a program according to Python programming specifications, with the input being the above triples and the output being a visual knowledge base.
[0131] Step 8: Import the program output by the large language model into Python and run it to generate a visual knowledge base.
[0132] Step 9: Conduct knowledge questions based on the large language model and test the task knowledge base.
[0133] Question: Output all entities of interest in flood disaster relief and their relevance scores;
[0134] Answer: Trapped people - 9; rescue vessels - 8; temporary shelters - 8...
[0135] Step 10: Analyze the test results. If there are no problems, complete the knowledge base construction; if there are problems, go to step 11.
[0136] Step 11: The operator submits modification requests to the large language model based on the key prompt word template, and then proceeds to step 8.
[0137] Keyword template: The above knowledge base has XXX issues, please modify it according to XXX standards.
[0138] On the other hand, the present invention also provides an interactive task knowledge base construction device based on a large language model, the device comprising:
[0139] A first processing module is configured to generate a task corpus request instruction for a target task based on the first keyword prompt template, input the task corpus request instruction into the large language model, and obtain the task corpus for the target task;
[0140] The second processing module is used to match the target standard task corpus from the preset standard task corpus according to the task corpus request instruction, perform text comparison analysis on the task corpus using the target standard task corpus, and determine whether the task corpus meets the preset task corpus requirements;
[0141] A third processing module is configured to, upon determining that the task corpus meets the preset task corpus requirements, generate an entity extraction instruction based on the second keyword prompt template, and input the entity extraction instruction into the large language model to obtain key semantic information of the task corpus; the key semantic information comprises at least one entity information, each entity information including an entity target focused on by the task and a task relevance score of the entity target;
[0142] The fourth processing module is used to generate a task knowledge base of the task target based on the key semantic information of the task corpus using a large language model.
[0143] It should be noted that the interactive task knowledge base construction device based on a large language model provided in an embodiment of the present invention can execute the interactive task knowledge base construction method based on a large language model described in any of the above embodiments during specific operation, which will not be described in detail in this embodiment.
[0144] Figure 3 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the interactive task knowledge base construction method based on the large language model.
[0145] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the interactive task knowledge base construction method based on a large language model provided in the above embodiments.
[0146] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the interactive task knowledge base construction method based on a large language model provided in the above embodiments.
[0147] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for constructing an interactive task knowledge base based on a large language model, characterized in that: include: Generate a task corpus request instruction for the target task based on the first keyword prompt template, input the task corpus request instruction into the large language model, and obtain the task corpus for the target task; According to the task corpus request instruction, a target standard task corpus is matched from a preset standard task corpus, and a text comparison analysis is performed on the task corpus using the target standard task corpus to determine whether the task corpus meets the preset task corpus requirements; including: parsing the task corpus request instruction and extracting matching information; the matching information includes at least one of the following information: a task subject of the target task and keywords of the task content; matching the target standard task corpus from the standard task corpus based on the matching information; extracting feature vectors from the task corpus and the target standard task corpus, and calculating the similarity of the feature vectors of the two; When it is determined that the task corpus meets the preset task corpus requirements, an entity extraction instruction is generated based on the second keyword prompt template, and the entity extraction instruction is input into the large language model to obtain key semantic information of the task corpus; the key semantic information includes at least one entity information, each entity information includes an entity target focused on by the task and a task relevance score of the entity target; Based on the key semantic information of the task corpus, a task knowledge base of the task target is generated using a large language model; After obtaining the key semantic information of the task corpus, it also includes: Use the pre-built classification model to classify entity information in key semantic information; classification types include entity omission, entity misrecognition, entity misrecognition, task relevance score misrecognition, and correct classification; Based on the classification result, a fourth keyword prompt template is generated, and the large language model is fine-tuned so that the proportion of entity information in the key semantic information output by the large language model that belongs to the correct classification is greater than the preset proportion.
2. The method for constructing an interactive task knowledge base based on a large language model according to claim 1, characterized in that: Also includes: When the similarity of the feature vectors is greater than or equal to a preset threshold, it is determined that the task corpus meets the preset task corpus requirements; When the similarity of the feature vectors is less than a preset threshold, it is determined that the task corpus does not meet the preset task corpus requirements.
3. The method for constructing an interactive task knowledge base based on a large language model according to claim 1, characterized in that: Obtain the task corpus of the target task, including: Determine the source of the task corpus; Determine the initial publisher based on the source of the task corpus and determine whether the initial publisher is within the preset publisher range; The task corpus whose initial publisher is not within the pre-defined publisher range is eliminated, and the task corpus whose initial publisher is within the pre-defined publisher range is retained.
4. The method for constructing an interactive task knowledge base based on a large language model according to claim 1, characterized in that: When it is determined that the task corpus does not meet the preset task corpus requirements, a third keyword prompt template is used to generate a fine-tuning instruction, and the fine-tuning instruction is input into the large language model to fine-tune the large language model until the task corpus output by the large language model meets the preset task corpus requirements; The third keyword prompt template is used to describe the fine-tuning standard for the large language model.
5. The method for constructing an interactive task knowledge base based on a large language model according to claim 1, characterized in that: Based on the key semantic information of the task corpus, a large language model is used to generate a task knowledge base of the task target, including: generating a program writing instruction based on the fifth keyword prompt template, inputting the program writing instruction into the large language model, and obtaining a target program; Execute the target program; the target program is used to generate a visualized task knowledge base based on key semantic information.
6. The method for constructing an interactive task knowledge base based on a large language model according to claim 5, characterized in that: Also includes: Testing the visualized task knowledge base; If there are no problems with the test results, complete the construction of the task knowledge base; Fine-tune the large language model if there are issues with the test results.
7. An interactive task knowledge base construction device based on a large language model, characterized in that: include: A first processing module is used to generate a task corpus request instruction for a target task based on a first keyword prompt template, input the task corpus request instruction into a large language model, and obtain the task corpus for the target task; The second processing module is used to match the target standard task corpus from the preset standard task corpus according to the task corpus request instruction, and use the target standard task corpus to perform text comparison analysis on the task corpus to determine whether the task corpus meets the preset task corpus requirements; including: parsing the task corpus request instruction and extracting matching information; the matching information includes at least one of the following information: the task subject of the target task and the keywords of the task content; matching the target standard task corpus from the standard task corpus based on the matching information; extracting feature vectors from the task corpus and the target standard task corpus, and calculating the similarity of the feature vectors of the two; A third processing module is used to generate an entity extraction instruction based on the second keyword prompt template when it is determined that the task corpus meets the preset task corpus requirements, and input the entity extraction instruction into the large language model to obtain key semantic information of the task corpus; the key semantic information includes at least one entity information, each entity information includes an entity target focused on by the task and a task relevance score of the entity target; The fourth processing module is used to generate a task knowledge base of the task target using a large language model based on the key semantic information of the task corpus; After obtaining the key semantic information of the task corpus, it also includes: Use the pre-built classification model to classify entity information in key semantic information; classification types include entity omission, entity misrecognition, entity misrecognition, task relevance score misrecognition, and correct classification; Based on the classification result, a fourth keyword prompt template is generated, and the large language model is fine-tuned so that the proportion of entity information in the key semantic information output by the large language model that belongs to the correct classification is greater than the preset proportion.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for constructing an interactive task knowledge base based on a large language model as described in any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for constructing an interactive task knowledge base based on a large language model as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model-based code generation method and device, storage medium and electronic equipment
CN117873453A
Automatic construction method and device of knowledge base, electronic equipment and storage medium
CN118312577A
Meteorological environment information crawling and analyzing method based on large language model
CN119248986A