Emergency question and answer large model construction method, system, equipment and medium
By building an emergency corpus database and specific training of the generative large language model, the problem of low emergency response efficiency in the existing technology is solved, and rapid and accurate emergency knowledge generation and emergency response are achieved, improving the adaptability and performance in the emergency field.
Patent Information
- Application Number
- CN202510093801.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Existing generative large models are difficult to quickly and accurately generate practical emergency knowledge and methods to deal with emergencies, resulting in poor response efficiency for emergency events.
By collecting relevant corpus data in the emergency field, building an emergency corpus database, and using a generative large language model to perform short sentence segmentation, quality improvement processing and zero-order optimization training on the corpus data, we obtain the emergency vertical large language model.
It realizes the rapid and accurate generation of practical emergency knowledge and methods to deal with emergencies, improves the response efficiency of emergency events, and improves the adaptability and performance in the emergency field.
Smart Images

Figure CN119938862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for constructing an emergency question and answer large model. Background Art
[0002] In today's rapidly developing information age, emergency science popularization has become an important part of social management. Emergency science popularization involves various types of emergency situations such as natural disasters, public health events, and sudden accidents. It faces problems such as scattered data, fragmented knowledge, and insufficient response timeliness. How to quickly and accurately obtain information and respond has become a major challenge facing society.
[0003] With the rapid progress of big data and artificial intelligence technology, emergency knowledge is highly professional, the classification standards at all levels are different, the tasks of managers are complicated, and there are more customized needs. The traditional emergency response mechanism has gradually exposed its lack of flexibility and poor adaptability, and new technical means are urgently needed to improve it. In recent years, the development of generative models has brought new possibilities to the emergency field, especially in the application of natural language processing (NLP), which has shown great potential.
[0004] Generative big models can generate content similar to real data by simulating its distribution, thus having higher flexibility in information acquisition and generation. This model can not only understand the input text information, but also generate relevant responses based on the context. However, the current generative big models are difficult to target emergency scenarios, and it is difficult to quickly and accurately generate practical emergency knowledge and methods to deal with emergencies, which leads to poor response efficiency in emergency events. Summary of the invention
[0005] In view of this, the present invention provides a method, system, device and medium for constructing an emergency question and answer big model, which solves the technical problem that the current generative big model is difficult to target emergency scenarios, and is difficult to quickly and accurately generate practical emergency knowledge and methods for dealing with emergencies, which leads to poor response efficiency to emergency events.
[0006] The first aspect of the present invention provides a method for constructing an emergency question-answering large model, comprising:
[0007] Collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and build an emergency corpus database;
[0008] Using a generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data;
[0009] Using the generative large language model to improve the quality of the plurality of seed corpus data to generate high-quality fine-tuning corpus data;
[0010] The generative large language model is subjected to zero-order optimization training according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
[0011] Preferably, the collecting of a large amount of relevant corpus data in the emergency field, classifying the relevant corpus data, and constructing an emergency corpus database includes:
[0012] Collect multiple pieces of relevant corpus data in the emergency field by crawling public data;
[0013] For each piece of the relevant corpus data, extract keywords from the relevant corpus data using a keyword library under a preset emergency category to obtain a plurality of keywords of the relevant corpus data under the preset emergency category;
[0014] Determine the number of keyword categories and the number of keywords of the relevant corpus data under the preset emergency category according to the multiple keywords of the relevant corpus data under the preset emergency category;
[0015] The preset emergency category to which the relevant corpus data belongs is determined according to the number of keyword categories and the number of keywords of the relevant corpus data under the preset emergency category, and an emergency corpus database is constructed according to the relevant corpus data and the preset emergency category to which it belongs.
[0016] Preferably, the generative large language model is used to segment the emergency corpus database into short sentences to obtain a plurality of seed corpus data, including:
[0017] The prompt word engineering technology is used to guide the generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data.
[0018] Preferably, the using the generative large language model to perform quality improvement processing on the plurality of seed corpus data to generate high-quality fine-tuning corpus data includes:
[0019] Performing context learning on the seed corpus data to generate multiple example corpus data;
[0020] Based on prompt word engineering technology, a plurality of the example corpus data are input into the generative large language model, and the generative large language model is guided to generate emergency instruction fine-tuning data;
[0021] By using self-instruction generation technology and context learning technology, the generative large language model is guided to expand the emergency instruction fine-tuning data to obtain emergency instruction fine-tuning extension data, and the emergency instruction fine-tuning extension data is used as the high-quality fine-tuning corpus data.
[0022] Preferably, the using the generative large language model to perform quality improvement processing on the plurality of seed corpus data to generate high-quality fine-tuning corpus data includes:
[0023] Two generative large language models are used as a user-generated large language model and an emergency expert-generated large language model respectively;
[0024] Based on the prompt word engineering technology, the seed corpus data is used to enable the user-generated large language model and the emergency expert-generated large language model to conduct multiple rounds of dialogue to generate emergency multi-round dialogue data, and the emergency multi-round dialogue data is used as the high-quality fine-tuning corpus data.
[0025] Preferably, the performing zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain the emergency vertical domain large language model includes:
[0026] Generating mixed corpus data by mixing the high-quality fine-tuning corpus data with a preset general corpus;
[0027] Based on a preset minimum loss function, the generative large language model is trained iteratively using the mixed corpus data in zero-order optimization, and after the iteration stops, the emergency vertical domain large language model is generated.
[0028] Preferably, the method further comprises: performing zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model, and then:
[0029] Vectorizing the query text pre-input into the emergency vertical domain large language model to obtain a query text vector;
[0030] The similarity between the query text vector and a plurality of relevant corpus vectors in the emergency corpus database is calculated by using cosine similarity; wherein the relevant corpus vector is obtained by vectorizing the relevant corpus data in the emergency corpus database;
[0031] According to the calculation result of the similarity, a relevant corpus vector having the highest similarity to the query text vector is selected, and the relevant corpus data corresponding to the selected relevant corpus vector is concatenated with the query text to obtain a concatenated query text;
[0032] The concatenated query text is retrieved through the emergency vertical domain large language model to generate a response text corresponding to the query text.
[0033] In a second aspect, the present invention further provides a system for constructing an emergency question-answering large model, comprising:
[0034] A corpus construction module is used to collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and construct an emergency corpus database;
[0035] A corpus segmentation module is used to segment the emergency corpus database into short sentences using a generative large language model to obtain a plurality of seed corpus data;
[0036] A corpus quality improvement module, used to use the generative large language model to improve the quality of the plurality of seed corpus data to generate high-quality fine-tuning corpus data;
[0037] The emergency model training module is used to perform zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
[0038] In a third aspect, the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing an emergency question and answer large model as described in the first aspect.
[0039] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the steps of the method for constructing an emergency question and answer large model as described in the first aspect.
[0040] It can be seen from the above technical scheme that the present invention collects a large amount of relevant corpus data in the emergency field and classifies them to construct an emergency corpus database, uses a generative large language model to segment the emergency corpus database into short sentences, uses the generative large language model to improve the quality of multiple seed corpus data obtained by short sentence segmentation, and performs zero-order optimization training on the generative large language model based on the high-quality fine-tuning corpus data obtained by the quality improvement process to obtain an emergency vertical domain large language model, so as to quickly and accurately generate practical emergency knowledge and methods for dealing with emergencies, improve the response efficiency of emergency events, and thus enhance the adaptability and performance of the emergency field. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0042] Figure 1 An application environment for a method for constructing a large emergency question-and-answer model provided by an embodiment of the present invention;
[0043] Figure 2 A flowchart of a method for constructing an emergency question and answer large model provided by an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of the structure of a system for building an emergency question and answer model according to an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] The method for constructing an emergency question-answering model provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the generative large language model communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or it can be placed on the cloud or other network servers. Server 102 collects a large amount of relevant corpus data in the emergency field, classifies the relevant corpus data, and constructs an emergency corpus database; uses the generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data; uses the generative large language model to improve the quality of multiple seed corpus data to generate high-quality fine-tuning corpus data; performs zero-order optimization training on the generative large language model based on the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model. Server 102 can be an independent physical server, or it can be a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud computing services.
[0048] like Figure 2 As shown, the embodiment of the present application provides a method for constructing an emergency question and answer model, and the method is applied to Figure 1 The server 102 in the example is used as an example to illustrate the method, which includes the following steps S1 to S4. Among them:
[0049] Step S1: collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and build an emergency corpus database.
[0050] Among them, by crawling relevant emergency field data as source knowledge, and widely collecting emergency-related corpus data from open source large-scale corpora for classification, an emergency corpus database is constructed.
[0051] Specifically, the step S1 collects a large amount of relevant corpus data in the emergency field, classifies the relevant corpus data, and constructs an emergency corpus database, including:
[0052] Step S101: Collect multiple pieces of relevant corpus data in the emergency field by crawling public data.
[0053] Among them, the source of public data can come from the government's public emergency guidance materials and high-quality data sets of online communities, and crawl public data through crawlers and other technologies to collect multiple relevant corpus data in the emergency field.
[0054] Step S102: for each piece of relevant corpus data, extract keywords from the relevant corpus data using a keyword library under a preset emergency category to obtain a plurality of keywords of the relevant corpus data under the preset emergency category.
[0055] Among them, tasks in the emergency field can be defined and classified. For example, emergency dialogue tasks can be divided into six categories: ["pre-prevention", "incident response", "incident handling", "post-disaster recovery process", and "rules and regulations query"]. For each preset emergency category, a keyword library is established based on the emergency professional knowledge under the preset emergency category.
[0056] Since the public data covers a wide range of fields, the embodiment of the present application utilizes a "multi-keyword matching" strategy and uses a keyword library under a preset emergency category to extract keywords from relevant corpus data to obtain multiple keywords matched by the relevant corpus data under the preset emergency category.
[0057] Step S103: Determine the number of keyword categories and the number of keywords of the relevant corpus data under the preset emergency category according to the multiple keywords of the relevant corpus data under the preset emergency category.
[0058] Step S104: determine the preset emergency category to which the relevant corpus data belongs according to the number of keyword categories and the number of keywords under the preset emergency category of the relevant corpus data, and construct an emergency corpus database according to the relevant corpus data and the preset emergency category to which it belongs.
[0059] Among them, when the keywords matched by the emergency keywords under a preset emergency category of the relevant corpus data exceed a set threshold, the relevant corpus data will be classified into a preset emergency category, and an emergency corpus database will be constructed through the relevant corpus data and the preset emergency category to which it belongs.
[0060] Exemplarily, taking the data screening of the "pre-prevention" task as an example, the keywords selected by the present invention are {"fire", "collapse", "prevention", "inspection", "maintenance", ...}, the category threshold of the keywords is set to 2, and the keyword number threshold is set to 3. Based on the above keywords, the public corpus data set is matched with multiple keywords, and the relevant corpus text is "Before using electrical tools, it is necessary to check whether the electrical circuit meets the equipment requirements to prevent electrical fires. Regularly inspect and maintain the electrical circuit to ensure the safe and reliable use of the tools. Operators must be certified to work and strictly follow the operating procedures to avoid electrical accidents." Among them, the three categories of keywords "fire", "inspection", and "maintenance" appear in the relevant corpus text, and 4 keywords ("fire", two "inspections", and "maintenance") appear in statistics. By comparison, the keyword category threshold and keyword number threshold requirements are met. Therefore, the relevant corpus text is incorporated into the "pre-prevention" task category.
[0061] Step S2: Use a generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data.
[0062] Among them, since the length of the relevant corpus data in the emergency corpus database in step S1 is different, there may be longer text corpuses, and they often contain knowledge in multiple emergency fields. In order to improve the accuracy and efficiency of subsequent training, the embodiment of the present application divides the length of the relevant corpus data into short sentences.
[0063] Specifically, the embodiment of the present application utilizes prompt word engineering technology to guide the generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data.
[0064] Among them, prompt engineering is to guide the generative large language model to produce specific types of outputs through designed input text or "prompt words" to meet actual needs. Among them, specific types include but are not limited to length restrictions, topic coherence, structural integrity and context relevance.
[0065] For example, the relevant corpus data is (I) Classification of explosion hazardous places: Explosion hazardous places are divided into two categories according to the physical state of explosive substances: gas explosion hazardous places and dust explosion hazardous places. (II) Classification of explosion hazardous places: The classification principle of explosion hazardous places is to divide them into areas of different danger levels according to the frequency, duration and degree of danger of the occurrence of explosive substances. 1. Regional level of gas explosion hazardous places Places where explosive gas, flammable vapor and air are mixed to form explosive gas mixtures are divided into three regional levels according to the degree of their danger. (1) Level 0 area (referred to as Zone 0, the same below) Under normal circumstances, places where explosive gas mixtures appear continuously, frequently in a short period of time or exist for a long time. (2) Level 1 area (referred to as Zone 1, the same below) Under normal circumstances, places where explosive gas mixtures are likely to appear . (3) Level 3 area (abbreviated as Zone 3, the same below) is a place where explosive gas mixture cannot appear under normal circumstances and only appears occasionally for a short time under abnormal circumstances. Note: Normal conditions refer to the normal start-up, stop, normal operation and maintenance of equipment. Abnormal conditions refer to the possibility of equipment failure or misoperation. 2. Regional levels of dust explosion hazardous places Places where explosive dust and combustible fibers mix with air to form explosive mixtures are divided into two regional levels according to the degree of danger. (1) Level 10 area: a place where, under normal circumstances, a mixture of explosive dust or combustible fibers and air may appear continuously, frequently in a short period of time, or exist for a long time. (2) Level 11 area: a place where, under normal circumstances, a mixture of explosive dust or combustible fibers and air cannot appear and only appears occasionally for a short time under abnormal circumstances.
[0066] The prompt word is set to "The following is a section of relevant corpus data, which contains rich emergency knowledge. It can be reasonably decomposed into 5-10 sub-knowledge based on relevance." The set prompt word can be used to reasonably decompose the above-mentioned relevant corpus data into 5-10 sub-knowledge.
[0067] Step S3: Use a generative large language model to improve the quality of multiple seed corpus data to generate high-quality fine-tuning corpus data.
[0068] It is understandable that due to the large number of seed corpus data and the fact that most of the corpus data is not accurate, high-quality, and has a standard format of instruction data, in order to make the training more accurate and reliable, the seed corpus data needs to be improved to obtain high-quality fine-tuning corpus data that meets the standard format and is of high quality.
[0069] In some embodiments, the crawled relevant knowledge is used to generate instruction-response pairs and their task types through context learning methods. The number of each task type is counted, and in order to make the model training more balanced, the instruction generalization method is used to generalize the instructions for tasks with fewer samples to obtain batch instruction data. The language model is fine-tuned based on the instruction data to improve the model's instruction recognition and response capabilities.
[0070] Specifically, in step S3, the generative large language model is used to improve the quality of multiple seed corpus data to generate high-quality fine-tuning corpus data, including:
[0071] Step S301: Perform context learning on the seed corpus data to generate multiple example corpus data.
[0072] Among them, based on the original seed corpus data, the context learning method is used to generate diverse and qualified example corpus data.
[0073] Step S302: Based on the prompt word engineering technology, multiple example corpus data are input into the generative large language model to guide the generative large language model to generate emergency instruction fine-tuning data.
[0074] in,
[0075] Through the prompt word engineering technology, multiple sample corpus data are input into the generative large language model, thereby guiding the generative large language model to generate accurate, high-quality and standardized instruction data. At the same time, in order to statistically analyze the data distribution of the generated instruction fine-tuning data set, the present invention also requires the model to output the task classification corresponding to the generated instruction data.
[0076] For example, the prompt words are set as: I am cleaning the data of a large language model that can be used for emergency management. The large language model needs to act as an expert and provide reasonable feedback for the emergency situation consultation input by the user. I will provide you with a text and its related keywords and task fields. You need to generate three instruction pairs for a given task scenario based on this text and keywords. These instruction pairs can be well applied to the training of the large language model instruction following. Each instruction pair is separated by two line breaks. Note that you should keep as many original details in the text as possible and do not modify or delete them by yourself. #Instruction pair output format:
[0077] (instruction content)", "input":{"instruction":"(user input, temporarily blank)" "output"(the answer to the instruction)", ""type":"(the specific task type, should be one of [(all task types under this category))"}#Example:
[0078] #Question:
[0079] Text: (txt text);
[0080] Keywords: (txt title);
[0081] Task area: (major task category);
[0082] Please generate three corresponding instruction pairs from the perspective of professionals in related fields. The answers should be brief and clear. The instructions should be general requirements or questions based on keywords and text. The output should be the answer to the materials and instructions.
[0083] Step S303: Use the self-instruction generation technology and context learning technology to guide the generative large language model to expand the emergency instruction fine-tuning data to obtain the emergency instruction fine-tuning extension data, and use the emergency instruction fine-tuning extension data as high-quality fine-tuning corpus data.
[0084] Among them, since the collected relevant corpus data itself has the phenomenon of unbalanced data distribution, the instruction fine-tuning data set generated in step S22 will also have the problem of unbalanced data distribution, which can easily lead to overfitting in model training. Therefore, this application uses self-instruction generation technology and context learning technology as complementary technologies, and guides the generative large language model to learn features and rules from existing data for task data with less data distribution, and generates similar emergency instruction fine-tuning data, thereby expanding the emergency instruction fine-tuning data, alleviating the phenomenon of unbalanced data distribution, and using the emergency instruction fine-tuning extended data as high-quality fine-tuning corpus data.
[0085] For example, the specific prompt words used in the self-command generation technology are set as: You are an expert in constructing command pairs and an expert in the field of [fire emergency-information delivery notification]. Please provide 10 different command pairs that meet the following requirements:
[0086] - The content of the instruction pair should be related to [Fire Emergency-Information Delivery Notice];
[0087] - The instruction pair has the same format as the provided instruction pair sample;
[0088] - The content of the instruction pair cannot be plagiarized from the content in the sample;
[0089] - The instructions should meet the format of {\n"instruction":"(instruction content)",\n"input":"(user input, temporarily blank)",\n"output":"(the answer to the instruction, which must be detailed, comprehensive and not less than 200 words)",\n"type":"information transmission notification"\n}';
[0090] The following are some sample instruction pairs:
[0091] {sample}
[0092] Please generate {x} instruction pairs, each separated by two newline characters.
[0093] In some embodiments, the step S3 of using a generative large language model to improve the quality of multiple seed corpus data to generate high-quality fine-tuning corpus data also includes:
[0094] Step S311: using two generative large language models as a user-generated large language model and an emergency expert-generated large language model respectively.
[0095] Step S312: Based on the prompt word engineering technology, the seed corpus data is used to enable the user-generated large language model and the emergency expert-generated large language model to conduct multiple rounds of dialogue, generate emergency multi-round dialogue data, and use the emergency multi-round dialogue data as high-quality fine-tuning corpus data.
[0096] Among them, based on the prompt word engineering technology, two generative large language models are guided to play the roles of users and emergency experts respectively, to serve as a user-generated large language model and an emergency expert-generated large language model, and through the two generative large language models, conversations are carried out around the seed corpus data, thereby constructing emergency multi-round dialogue data as high-quality fine-tuning corpus data.
[0097] Exemplarily, the prompt words used in the multi-round dialogue data generation are set as:
[0098] Role:
[0099] Emergency experts: have rich experience in emergency management, are responsible for answering users' questions and providing scientific emergency advice.
[0100] User: Ask questions about emergency management and first aid knowledge and seek help.
[0101] Please conduct the following multiple rounds of conversations based on the emergency knowledge broken down above.
[0102] The following is a long and in-depth multi-round conversation (at least 10 rounds) between the user and the emergency expert in this scenario, including greetings and ending. Please try to make the current user's question related to the emergency expert's previous reply or your previous question, such as: adding pronoun references to the current conversation, the user correcting the emergency expert's misunderstanding, and asking further questions based on the expert's reply. The following is an example (JSON format data) {{Selected example}}.
[0103] Step S4: Perform zero-order optimization training on the generative large language model based on high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
[0104] It should be noted that since retraining a new generative large language model is relatively complex and inefficient, the embodiment of the present application performs zero-order optimization training on the generative large language model through high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model that focuses on emergency scenarios, thereby utilizing the emergency vertical domain large language model to query and generate practical emergency knowledge and methods for dealing with emergencies, thereby helping to improve the public's safety awareness and self-rescue and mutual rescue capabilities, and can achieve consultation, decision support, and emotional support question-and-answer tasks in multiple scenarios such as pre-prevention, in-process handling, and post-disaster recovery, reflecting high standards in terms of safety, practicality, and standardization in emergency science popularization.
[0105] Specifically, in step S4, the generative large language model is trained with zero-order optimization according to the high-quality fine-tuning corpus data to obtain the emergency vertical domain large language model, including:
[0106] Step S401: generating mixed corpus data by mixing high-quality fine-tuning corpus data with preset general corpus.
[0107] In order to balance the domain specificity and versatility of the model, the embodiment of the present application adopts a mixed training method of high-quality fine-tuning corpus data and general corpus, and the mixing ratio is determined according to the following formula:
[0108]
[0109] in, represents the amount of emergency corpus data, Indicates the amount of general corpus data.
[0110] Exemplarily, general corpus (that is, corpus covering various fields) and high-quality fine-tuning corpus data are mixed in a ratio of 1:4 to avoid a small deviation in the trained large language model.
[0111] Step S402: Based on a preset minimum loss function, the generative large language model is iteratively trained using the mixed corpus data through zero-order optimization. After the iteration stops, an emergency vertical domain large language model is generated.
[0112] Among them, the next token prediction (NTP) task is used to pre-train the Qwen2 model to enhance the model's professional language generation ability and context understanding ability in the emergency field. The preset minimum loss function is:
[0113]
[0114] In the formula, represents the loss function for the next word prediction, and the goal is to minimize this loss function. Represents the first words, Indicates the model's response to the next word. The predicted probability, Represents the model Dimension parameters.
[0115] It should be noted that the generative large language model in the embodiment of the present application is a general large language model, such as Qwen2, gpt, qwen, llama and other large language models.
[0116] In this embodiment, the zero-order optimization algorithm (ZOA) is used to efficiently fine-tune the parameters of the generative large language model to improve its adaptability and generation ability in the emergency field. Specifically, the high-dimensional model overall weight matrix Decomposed into two low-rank matrices and In this way, the originally large model parameter matrix Can be In the form of and The dimension of is much smaller than that of the original weight matrix, which significantly reduces the number of parameters that need to be updated and greatly reduces the demand for computing resources. In the actual training process, a small batch data set obtained by randomly screening the data is used. right and Zero-order fine-tuning is performed to make parameter updates to minimize the loss function value.
[0117] The core of the zero-order optimization algorithm is to use the function value difference to update the parameters, without explicitly calculating the gradient, which greatly reduces the computational complexity. The gradient estimation formula for parameter update is as follows:
[0118]
[0119] In the formula, is a Gaussian distributed random variable, Characterizes the degree of disturbance, Represents the model Dimension parameters, Indicated by the distribution Mini-batch dataset The obtained local loss function.
[0120] It should be noted that the embodiment of the present application collects and classifies a large amount of relevant corpus data in the emergency field, constructs an emergency corpus database, uses a generative large language model to segment the emergency corpus database into short sentences, uses the generative large language model to improve the quality of multiple seed corpus data obtained by short sentence segmentation, and performs zero-order optimization training on the generative large language model based on the high-quality fine-tuning corpus data obtained by the quality improvement processing to obtain an emergency vertical domain large language model, so as to quickly and accurately generate practical emergency knowledge and methods for dealing with emergencies, improve the response efficiency of emergency events, and thus enhance the adaptability and performance of the emergency field.
[0121] The embodiment of this application builds a refined emergency scenario knowledge question-and-answer system with traceability. The present invention significantly improves the performance of the emergency science popularization system under the complexity of scenarios and knowledge traceability requirements, provides reliable technical support for efficient response and knowledge dissemination of emergency events, and has certain reference value for other vertical domain knowledge management and generation tasks.
[0122] In some embodiments, the method for constructing an emergency question and answer model provided in the embodiments of the present application further includes:
[0123] Step S501: vectorize the query text pre-input into the emergency vertical domain large language model to obtain a query text vector.
[0124] Among them, the query text refers to the question text input by the user into the emergency vertical domain large language model according to needs.
[0125] Step S502: Calculate the similarity between the query text vector and multiple related corpus vectors in the emergency corpus database using cosine similarity; wherein the related corpus vector is obtained by vectorizing the related corpus data in the emergency corpus database.
[0126] Among them, the semantic vectorization model acge_text_embedding can be used to vectorize the query text and the relevant corpus data in the emergency corpus database. Specifically, the relevant corpus data text in the emergency corpus database Converted into high-dimensional related corpus vector , its vectorized calculation formula is:
[0127]
[0128] in, It is a semantic vectorization model to ensure that the vector can accurately capture the semantic information in the corpus.
[0129] Query text entered by the user The same vectorization is performed to obtain the vector Then, the present invention calculates the relevant corpus vector by cosine similarity With query text vector Similarity:
[0130]
[0131] In the formula, is the relevant corpus vector With query text vector The similarity.
[0132] Among them, when the corpus text is vectorized, the title of the corpus text is retained and not vectorized, so that in the subsequent answering process, the title (such as the network source of the corpus text, etc.) is directly output, and the response generated by the emergency vertical domain large language model and the title of the corpus data used are fed back to the user.
[0133] Step S503: select the relevant corpus vector with the highest similarity to the query text vector according to the similarity calculation result, and concatenate the relevant corpus data corresponding to the selected relevant corpus vector with the query text to obtain a concatenated query text.
[0134] Among them, the relevant corpus data corresponding to the screened relevant corpus vectors can be concatenated with the query text to form a complete sequence input, and input into the emergency vertical domain large language model, thereby improving the use of text vectorization technology to retrieve corpus data with high vector similarity to user input as evidence support, thereby improving the refinement and accuracy of the retrieval.
[0135] Step S504: search the concatenated query text through the emergency vertical domain large language model to generate a response text corresponding to the query text.
[0136] Among them, the spliced query text is searched and queried through the emergency vertical domain large language model, and the response text corresponding to the query text is generated to achieve traceable and refined emergency scenario knowledge questions and answers, which significantly improves the performance of the emergency science popularization system under the complexity of the scene and the need for knowledge traceability, and provides reliable technical support for efficient response and knowledge dissemination of emergency events. Based on the above data, the generative large language model is trained to enable the emergency vertical domain large language model to have professional knowledge in the emergency field. During the user query process, the model will improve the reliability of the response through the method of retrieval enhancement generation.
[0137] For example, during the query process, if you input how to escape safely in case of fire, the response text is: In case of fire, you should remain calm, call 119 immediately, observe whether you can escape from the fire, and cover your mouth and nose with a wet towel.
[0138] —The above answers come from: "XX Province Fire Protection Guide Manual".
[0139] Based on the same inventive concept, an embodiment of the present application also provides an emergency question and answer big model construction system for implementing the above-mentioned emergency question and answer big model construction method.
[0140] The implementation solution for solving the problem provided by the system is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more emergency question and answer big model construction system embodiments provided below can be referred to the limitations on the emergency question and answer big model construction method above, and will not be repeated here.
[0141] like Figure 3 As shown, the embodiment of the present application provides a system for building an emergency question and answer model, including:
[0142] The corpus construction module 100 is used to collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and construct an emergency corpus database;
[0143] The corpus segmentation module 200 is used to segment the emergency corpus database into short sentences using a generative large language model to obtain a plurality of seed corpus data;
[0144] A corpus quality improvement module 300 is used to improve the quality of multiple seed corpus data using a generative large language model to generate high-quality fine-tuning corpus data;
[0145] The emergency model training module 400 is used to perform zero-order optimization training on the generative large language model based on high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
[0146] In some embodiments, the corpus construction module 100 is used to collect multiple pieces of relevant corpus data in the emergency field by crawling public data;
[0147] For each piece of relevant corpus data, extract keywords from the relevant corpus data using a keyword library under a preset emergency category to obtain multiple keywords of the relevant corpus data under the preset emergency category;
[0148] According to a plurality of keywords in the relevant corpus data under the preset emergency category, determining the number of keyword categories and the number of keywords in the relevant corpus data under the preset emergency category;
[0149] The preset emergency category to which the relevant corpus data belongs is determined according to the number of keyword categories and the number of keywords in the preset emergency category, and an emergency corpus database is constructed according to the relevant corpus data and the preset emergency category to which it belongs.
[0150] In some embodiments, the corpus segmentation module 200 is used to utilize prompt word engineering technology to guide the generative large language model to perform short sentence segmentation on the emergency corpus database to obtain a plurality of seed corpus data.
[0151] In some embodiments, the corpus quality improvement module 300 is used to perform context learning on the seed corpus data to generate multiple example corpus data;
[0152] Based on the prompt word engineering technology, multiple sample corpus data are input into the generative large language model to guide the generative large language model to generate emergency command fine-tuning data;
[0153] By using self-instruction generation technology and context learning technology, the generative large language model is guided to expand the emergency instruction fine-tuning data to obtain the emergency instruction fine-tuning extended data, and the emergency instruction fine-tuning extended data is used as high-quality fine-tuning corpus data.
[0154] In some embodiments, the corpus quality improvement module 300 is used to use two generative large language models as a user-generated large language model and an emergency expert-generated large language model respectively;
[0155] Based on the prompt word engineering technology, the seed corpus data is used to enable the user-generated large language model and the emergency expert-generated large language model to conduct multi-round conversations, generate emergency multi-round conversation data, and use the emergency multi-round conversation data as high-quality fine-tuning corpus data.
[0156] In some embodiments, the emergency model training module 400 is used to generate mixed corpus data by mixing high-quality fine-tuning corpus data with preset general corpus;
[0157] Based on the preset minimum loss function, the generative large language model is trained iteratively using mixed corpus data for zero-order optimization. After the iteration stops, an emergency vertical domain large language model is generated.
[0158] In some embodiments, the system includes:
[0159] The text vector module is used to vectorize the query text pre-input into the emergency vertical domain large language model to obtain the query text vector;
[0160] A similarity calculation module is used to calculate the similarity between the query text vector and multiple related corpus vectors in the emergency corpus database using cosine similarity; wherein the related corpus vector is obtained by vectorizing the related corpus data in the emergency corpus database;
[0161] A text concatenation module is used to select the relevant corpus vector with the highest similarity to the query text vector according to the similarity calculation result, and concatenate the relevant corpus data corresponding to the selected relevant corpus vector with the query text to obtain a concatenated query text;
[0162] The text retrieval module is used to retrieve the concatenated query text through the emergency vertical domain large language model and generate the response text corresponding to the query text.
[0163] like Figure 4 As shown, an embodiment of the present application provides an electronic device, the electronic device 10 includes a memory 20 and a processor 30, the memory 20 stores a computer program, and when the computer program is executed by the processor 30, the processor 30 executes the steps of the method for constructing an emergency question and answer large model as in any of the above embodiments.
[0164] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed, the steps of the method for constructing an emergency question and answer large model as in any of the above embodiments are implemented.
[0165] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, electronic device and computer storage medium can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0166] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or apparatus.
[0167] In several embodiments provided by the present invention, it is understood that each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and a part of a module, a program segment or a code includes one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
[0168] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0169] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0170] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for executing all or part of the steps of the method described in each embodiment of the present invention through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program codes.
[0172] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a large emergency question-answering model, characterized in that: include: Collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and build an emergency corpus database; Using a generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data; Using the generative large language model to improve the quality of the plurality of seed corpus data to generate high-quality fine-tuning corpus data; The generative large language model is subjected to zero-order optimization training according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
2. The method for constructing an emergency question-answering model according to claim 1, characterized in that: The collecting of a large amount of relevant corpus data in the emergency field, classifying the relevant corpus data, and constructing an emergency corpus database includes: Collect multiple pieces of relevant corpus data in the emergency field by crawling public data; For each piece of the relevant corpus data, extract keywords from the relevant corpus data using a keyword library under a preset emergency category to obtain a plurality of keywords of the relevant corpus data under the preset emergency category; Determine the number of keyword categories and the number of keywords of the relevant corpus data under the preset emergency category according to the multiple keywords of the relevant corpus data under the preset emergency category; The preset emergency category to which the relevant corpus data belongs is determined according to the number of keyword categories and the number of keywords of the relevant corpus data under the preset emergency category, and an emergency corpus database is constructed according to the relevant corpus data and the preset emergency category to which it belongs.
3. The method for constructing an emergency question-answering model according to claim 1, characterized in that: The generative large language model is used to segment the emergency corpus database into short sentences to obtain a plurality of seed corpus data, including: The prompt word engineering technology is used to guide the generative large language model to segment the emergency corpus database into short sentences to obtain multiple seed corpus data.
4. The method for constructing an emergency question-answering model according to claim 1, characterized in that: The using the generative large language model to perform quality improvement processing on the plurality of seed corpus data to generate high-quality fine-tuning corpus data includes: Performing context learning on the seed corpus data to generate multiple example corpus data; Based on prompt word engineering technology, a plurality of the example corpus data are input into the generative large language model, and the generative large language model is guided to generate emergency instruction fine-tuning data; By using self-instruction generation technology and context learning technology, the generative large language model is guided to expand the emergency instruction fine-tuning data to obtain emergency instruction fine-tuning extension data, and the emergency instruction fine-tuning extension data is used as the high-quality fine-tuning corpus data.
5. The method for constructing an emergency question-answering large model according to claim 1 or 4, characterized in that: The using the generative large language model to perform quality improvement processing on the plurality of seed corpus data to generate high-quality fine-tuning corpus data includes: Two generative large language models are used as user-generated large language model and emergency expert-generated large language model respectively; Based on the prompt word engineering technology, the seed corpus data is used to enable the user-generated large language model and the emergency expert-generated large language model to conduct multiple rounds of dialogue to generate emergency multi-round dialogue data, and the emergency multi-round dialogue data is used as the high-quality fine-tuning corpus data.
6. The method for constructing an emergency question-answering model according to claim 1, characterized in that: The step of performing zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model includes: Generating mixed corpus data by mixing the high-quality fine-tuning corpus data with a preset general corpus; Based on a preset minimum loss function, the generative large language model is trained iteratively using the mixed corpus data in zero-order optimization, and after the iteration stops, the emergency vertical domain large language model is generated.
7. The method for constructing an emergency question-answering model according to claim 1, characterized in that: The step of performing zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model further includes: Vectorizing the query text pre-input into the emergency vertical domain large language model to obtain a query text vector; The similarity between the query text vector and a plurality of relevant corpus vectors in the emergency corpus database is calculated by using cosine similarity; wherein the relevant corpus vector is obtained by vectorizing the relevant corpus data in the emergency corpus database; According to the calculation result of the similarity, a relevant corpus vector having the highest similarity to the query text vector is selected, and the relevant corpus data corresponding to the selected relevant corpus vector is concatenated with the query text to obtain a concatenated query text; The concatenated query text is retrieved through the emergency vertical domain large language model to generate a response text corresponding to the query text.
8. A large model construction system for emergency question and answer, characterized in that: include: A corpus construction module is used to collect a large amount of relevant corpus data in the emergency field, classify the relevant corpus data, and construct an emergency corpus database; A corpus segmentation module is used to segment the emergency corpus database into short sentences using a generative large language model to obtain a plurality of seed corpus data; A corpus quality improvement module, used to use the generative large language model to improve the quality of the plurality of seed corpus data to generate high-quality fine-tuning corpus data; The emergency model training module is used to perform zero-order optimization training on the generative large language model according to the high-quality fine-tuning corpus data to obtain an emergency vertical domain large language model.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the method for constructing an emergency question and answer large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method for constructing an emergency question and answer large model are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Model training method, data processing method, equipment and storage medium
CN117520842A
Emergency plan generation system and method based on large language model
CN117668155A
Dam emergency response rule question and answer recommendation system construction method based on large language model
CN118332076A
Cited By
Construction method and system of vertical domain large language model, electronic equipment and storage medium
CN121543719A