Ocean early warning intelligent question answering method and related device based on large language model

Through the intelligent question-and-answer method of marine early warning based on large language models, the vectorized corpus in the task chain and vector database is matched, and the answer is generated using incremental pre-training and instruction fine-tuning of marine early warning large language model, which solves the problem of inefficient knowledge retrieval in the field of marine early warning, and realizes the intelligent question-and-answer and timeliness of marine early warning.

CN119179756BActive Publication Date: 2025-05-02NAT MARINE ENVIRONMENTAL FORECASTING CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411230423.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2025-05-02
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

The prior art knowledge retrieval, sharing, updating and application in the field of marine early warning is inefficient, and the model is relatively single in terms of the expression of the answer content, which cannot meet the diversified needs and timeliness requirements of marine early warning services.

Method used

The marine early warning intelligent question-and-answer method based on the large language model is adopted. By obtaining user input questions, matching vectorized corpus in the task chain and vector database, the marine early warning large language model is generated by incremental pre-training and instruction fine-tuning.

Benefits of technology

It realizes intelligent Q&A for marine early warning, improves the ability to adapt to the diversified needs of marine early warning services, and meets the accuracy and timeliness requirements of marine early warning services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119179756B_ABST
    Figure CN119179756B_ABST
Patent Text Reader

Abstract

The present application discloses a method and related devices for intelligent question and answer of ocean early warning based on a large language model, which relates to the technical field of intelligent question and answer of ocean early warning, matches the questions input by users with each type of task chain and each vectorized corpus in a vector database, and fills the questions and target vectorized corpus into the prompt word template of the target task chain, and uses the ocean early warning large language model to output the answers to the questions, which can be applied to the work of ocean early warning business to realize intelligent question and answer of ocean early warning, and vectorizes the ocean early warning corpus to build a vector database, retrieves and reorders the vector database according to the questions input by users to obtain the target vectorized corpus, and can constrain and trace the answers of the ocean early warning large language model, and when the ocean early warning corpus is updated, the vector database is updated to meet the accuracy and timeliness requirements of the ocean early warning business.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of ocean early warning intelligent question answering, and in particular to an ocean early warning intelligent question answering method and related devices based on a large language model. Background Art

[0002] As the most extensive component of the earth's ecosystem, the ocean has a profound impact on the origin and development of human civilization. However, the vastness of the ocean also contains tremendous power. Marine disasters such as storm surges, tsunamis, waves, and red tides cause huge losses of life and property to human activities at sea and near the coast, and seriously affect the development of marine economic industries such as fisheries, shipping, tourism, and marine engineering. Therefore, in order to achieve marine disaster prevention and mitigation, it is very important to develop marine early warning technology.

[0003] Marine early warning expertise is the scientific basis for marine disaster prevention and mitigation. It not only records the systematic observation of marine phenomena, but also provides data and theoretical support for marine early warning technology. Marine early warning expertise promotes exchanges and cooperation among marine early warning workers and researchers around the world, and promotes the dissemination and innovation of marine early warning technology. However, in the field of marine disaster prevention and mitigation, it faces challenges brought by the breadth and complexity of the knowledge system. Traditional manual knowledge retrieval methods require a lot of time for manual search, screening and sorting, which is inefficient and highly dependent on the user's knowledge background and experience. The retrieved knowledge needs to go through tedious processing, analysis and processing before it can be transformed into effective applications in actual marine early warning business, which undoubtedly increases the time cost and operational complexity. Therefore, there is an urgent need for a new way to retrieve, share, update and apply marine early warning knowledge.

[0004] In recent years, the general large language model (LLM) represented by ChatGPT has demonstrated its powerful capabilities in the field of natural language processing, and some models for specific professional application fields have also emerged, which provides a new idea for solving the above problems. However, the existing models in specific professional application fields generally have the following problems and cannot be directly applied to marine early warning business: (1) Limitations of content expression: The existing models are relatively simple in the expression of answer content, which limits their ability to adapt to the diversified needs of marine early warning business; (2) Model hallucination problem and knowledge lag: The existing models have obvious lags in knowledge update mechanism, and the trained models can no longer input new knowledge. At the same time, there is a lack of constraints and traceability on the model's answers, and the model is prone to hallucination problems when answering, which makes it difficult to meet the accuracy and timeliness requirements of the marine early warning business. Summary of the invention

[0005] The purpose of this application is to provide an ocean early warning intelligent question and answer method and related devices based on a large language model, which can be applied to ocean early warning business work, realize intelligent question and answer of ocean early warning, and meet the accuracy and timeliness requirements of ocean early warning business.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] In a first aspect, the present application provides an ocean early warning intelligent question-answering method based on a large language model, and the ocean early warning intelligent question-answering method based on a large language model includes:

[0008] Get user input questions;

[0009] Matching the question with the basic information of each type of task chain to obtain a target task chain matching the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge quizzes, task chains for true-false tasks in professional knowledge quizzes, task chains for short-answer tasks in professional knowledge quizzes, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain;

[0010] The question is matched with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database;

[0011] The target vectorized corpus is filled into the background knowledge input of the prompt word template of the target task chain, and the question is filled into the question input of the prompt word template of the target task chain to obtain a filled template, and the filled template is used as input to answer the question using the ocean early warning large language model output to obtain an answer; the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model.

[0012] In a second aspect, the present application provides an ocean early warning intelligent question-answering device based on a large language model, and the ocean early warning intelligent question-answering device based on a large language model includes:

[0013] Question input module, used to obtain questions input by users;

[0014] A task chain matching module is used to match the question with the basic information of each type of task chain to obtain a target task chain that matches the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge quizzes, task chains for true-or-false tasks in professional knowledge quizzes, task chains for short-answer tasks in professional knowledge quizzes, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain;

[0015] A vectorized corpus matching module is used to match the question with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database;

[0016] An answer generation module is used to fill the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, fill the question into the question input of the prompt word template of the target task chain, obtain a filled template, and use the filled template as input to answer the question using the ocean early warning large language model output to obtain an answer; the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model.

[0017] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned ocean early warning intelligent question-answering method based on a large language model.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned ocean early warning intelligent question-answering method based on a large language model.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned ocean early warning intelligent question and answer method based on a large language model.

[0020] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0021] The present application provides an ocean early warning intelligent question-answering method and related devices based on a large language model, which matches the questions input by the user with the basic information of each type of task chain to obtain a target task chain matching the question, matches the questions input by the user with each vectorized corpus in a vector database to obtain a target vectorized corpus matching the question, fills the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, fills the question into the question input of the prompt word template of the target task chain to obtain a filled-in template, and uses the filled-in template as input to output the answer to the question using the ocean early warning large language model. Since the types of task chains include task chains of multiple-choice questions in professional knowledge question and answer, task chains of true-false questions in professional knowledge question and answer, task chains of short-answer questions in professional knowledge question and answer, task chains of copywriting generation tasks and task chains of summary tasks, and the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model, it can complete professional knowledge question and answer tasks (multiple-choice questions, true-false questions, short-answer questions), copywriting generation tasks and summary tasks, realize intelligent question and answer of ocean early warning, solve the problem that the existing model has a relatively single expression form in the answer content, and improve the ability to adapt to the diversified needs of ocean early warning business. By setting up a vector database, before inputting the question into the ocean early warning large language model, the target vectorized corpus matching the question can be determined based on the nearest neighbor search and re-ranking method, and the background knowledge of the question can be determined, so that the ocean early warning large language model can determine the answer based on the background knowledge and the question, solve the problem that the existing model lacks constraints and traceability on the model's answer, and the model is prone to hallucinations when answering, and meet the accuracy requirements of the ocean early warning business. By setting up vectorization processing for each updated ocean warning corpus to update the vector database, the vector database can always contain the latest ocean warning corpus, solving the problem of obvious lag in the knowledge updating mechanism of the existing model and meeting the timeliness requirements of the ocean warning business. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 An application environment diagram of an ocean early warning intelligent question-answering method based on a large language model provided in Example 1 of the present application.

[0024] Figure 2A flow chart of an intelligent question-answering method for ocean early warning based on a large language model provided in Example 1 of the present application.

[0025] Figure 3 A flowchart of matching questions and task chains provided in Example 1 of the present application.

[0026] Figure 4 A schematic diagram of the process of matching questions with vectorized corpus provided in Example 1 of the present application.

[0027] Figure 5 A schematic diagram of the process of constructing the large language model for ocean early warning provided in Example 1 of the present application.

[0028] Figure 6 A schematic diagram of the process of constructing the vector database provided in Example 1 of the present application.

[0029] Figure 7 A schematic diagram of the functional modules of an ocean early warning intelligent question-answering device based on a large language model provided in Example 2 of the present application.

[0030] Figure 8 A schematic diagram of the functional modules of another marine early warning intelligent question-answering device based on a large language model provided in Example 2 of the present application.

[0031] Fig. 9 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. Example

[0033] The ocean early warning intelligent question-answering method based on a large language model provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the question of the user input to be processed to the server. After the server receives the question of the user input to be processed, for the question of the user input to be processed, the server matches the question with the basic information of each type of task chain, obtains the target task chain matching the question, matches the question with each vectorized corpus in the vector database based on the nearest neighbor search and reordering method, obtains the target vectorized corpus matching the question, fills the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, fills the question into the question input of the prompt word template of the target task chain, obtains the filled template, and uses the filled template as input, and uses the ocean early warning large language model to output the answer to the question. The server can feedback the answer to the question of the user input to be processed to the terminal.

[0034] In addition, in some embodiments, the ocean early warning intelligent question and answer method based on the large language model can also be implemented independently by a server or a terminal. For example, the terminal can directly process the questions input by the user to be processed, or the server can obtain the questions input by the user to be processed from the data storage system and process the questions input by the user to be processed.

[0035] The terminal may be, but is not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0036] like Figure 2 As shown, a marine early warning intelligent question-answering method based on a large language model is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 Taking the server in as an example, the ocean early warning intelligent question answering method based on the large language model includes the following steps:

[0037] Step S1, obtaining a question input by a user.

[0038] Step S2, matching the question with the basic information of each type of task chain to obtain a target task chain matching the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge questions and answers, task chains for true-or-false tasks in professional knowledge questions and answers, task chains for short-answer tasks in professional knowledge questions and answers, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain.

[0039] Step S3, based on the nearest neighbor search and re-ranking method, the question is matched with each vectorized corpus in the vector database to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database.

[0040] Step S4, fill the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, fill the question into the question input of the prompt word template of the target task chain, obtain the filled template, and use the filled template as input to answer the question using the ocean early warning big language model output to obtain the answer; the ocean early warning big language model is obtained after incremental pre-training and instruction fine-tuning of the big language model.

[0041] By implementing the above steps S1 to S4, this embodiment can be applied to the marine early warning business work, and multiple types of task chains are pre-set, and the types of task chains include task chains of multiple-choice questions in professional knowledge questions and answers, task chains of true-or-false questions in professional knowledge questions and answers, task chains of short-answer questions in professional knowledge questions and answers, task chains of copywriting tasks, and task chains of summary tasks. The questions input by the user are matched with the basic information of each type of task chain to obtain a target task chain matching the question, and the questions input by the user are matched with each vectorized corpus in the vector database to obtain a target vectorized corpus matching the question. Subsequently, the question and the target vectorized corpus are input into the prompt word template of the target task chain, and the answer obtained by answering the question is output by the marine early warning large language model. By setting multiple types of task chains, and the marine early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model, professional knowledge question and answer tasks (multiple-choice questions, true-or-false questions, short-answer questions), copywriting tasks, and summary tasks can be completed, and intelligent question and answer of marine early warning can be realized, solving the problem that the existing model has a relatively single expression form in the answer content, and improving the ability to adapt to the diversified needs of marine early warning business. By setting up a vector database, before inputting the question into the ocean warning large language model, the target vectorized corpus matching the question can be determined based on the nearest neighbor search and re-ranking method, and the background knowledge of the question can be determined, so that the ocean warning large language model can determine the answer based on the background knowledge and the question, solving the problem that the existing model lacks constraints and traceability on the model's answers, and the model is prone to hallucinations when answering, meeting the accuracy requirements of the ocean warning business. By setting the updated ocean warning corpus to be vectorized whenever the ocean warning corpus is updated to update the vector database, the vector database can always contain the latest ocean warning corpus, solving the problem that the existing model has obvious lags in the knowledge update mechanism, and meeting the timeliness requirements of the ocean warning business.

[0042] In traditional question-answering systems, different models are usually needed to perform different tasks, such as classification tasks and prediction tasks. This embodiment relies on the multi-application capabilities of the large language model and only needs to build different prompt word templates for different tasks. Different ocean early warning task logics can be realized through link combination, that is, preset prompt word templates for multiple tasks, and combine them according to the ocean early warning tasks to build a task chain. This embodiment designs two different task chain forms: (1) user input question + prompt word template 1 → large model output; (2) user input question + prompt word template 1 → large model output 1 + prompt word template 2 → large model output 2.

[0043] In this embodiment, the types of task chains include task chains of multiple-choice tasks in professional knowledge quizzes, task chains of true-or-false tasks in professional knowledge quizzes, task chains of short-answer tasks in professional knowledge quizzes, task chains of text generation tasks, and task chains of summary tasks. Among them, the task chains of multiple-choice tasks in professional knowledge quizzes, the task chains of true-or-false tasks in professional knowledge quizzes, and the task chains of short-answer tasks in professional knowledge quizzes adopt the first task chain form, the task chain of text generation tasks adopts the second task chain form, and the task chain of summary tasks adopts the first task chain form. Multiple-choice questions refer to single-choice questions.

[0044] The prompt word template of the task chain includes background knowledge input, task description and question input. Background knowledge input is used to input background knowledge (i.e., target vectorized corpus), task description is used to determine the specific requirements of the task, and question input is used to input questions.

[0045] For example, for the multiple-choice task in the professional knowledge question and answer, the first task chain form is used, and its prompt word template is as follows:

[0046] "[INST] < <sys>> \n"

[0047] "You are an oceanographer.\n"

[0048] "<< / sys> >\n\n"

[0049] "The following is background knowledge:\n"

[0050] "{context}"

[0051] "\n"

[0052] "This is a multiple-choice question. Based on the above background knowledge, please select the correct answer to be filled in the brackets among the four options A, B, C, and D, and output the option letter of the answer. The following is a multiple-choice question: {question}."

[0053] "[ / INST]"

[0054] Among them, [INST], < <sys> >、<< / sys> > and [ / INST] are special tokens that mark the various parts of the prompt word template. [INST] represents the beginning of the user prompt word, < <sys> > represents the beginning of the system prompt word, << / sys> > represents the end of the system prompt, [ / INST] represents the end of the user prompt; \n represents a line break; {context} represents background knowledge input; "This is a multiple-choice question. Based on the above background knowledge, please select the correct answer to be filled in the brackets from the four options A, B, C, and D, and output the option letter of the answer" represents the task description; {question} represents the question input.

[0055] For the judgment question task and the short-answer question task in the professional knowledge quiz, the first task chain form is used. The prompt word template is similar to the prompt word template for the multiple-choice question task in the professional knowledge quiz. Only the task description needs to be changed. For the judgment question task in the professional knowledge quiz, the task description of the prompt word template can be "This is a judgment question. Please determine whether the following question is correct based on the above background knowledge, and output the correct or wrong answer"; for the short-answer question task in the professional knowledge quiz, the task description of the prompt word template can be "This is a short-answer question. Please answer the following question based on the above background knowledge and output the answer."

[0056] For the copywriting generation task, the second task chain form is used, which consists of two prompt word templates. The user inputs the question and prompt word template 1 and inputs it into the ocean early warning language model to obtain model answer 1. The model answer 1 and prompt word template 2 are combined and input into the ocean early warning language model to obtain model answer 2. For example, for the ocean early warning disaster process summary generation task, its prompt word template 1 is as follows:

[0057] "[INST] < <sys>>\n"

[0058] "You are a marine forecaster.\n"

[0059] "<< / sys> >\n\n"

[0060] "The following is background knowledge:\n"

[0061] "{context}"

[0062] "\n"

[0063] "Please describe the actual situation of the waves based on the above background knowledge and weather situation field. The following is the weather situation field: {question}."

[0064] "[ / INST]"

[0065] Among them, “Please describe the actual situation of the waves based on the above background knowledge and weather conditions” represents the task description.

[0066] The prompt word template 2 is as follows:

[0067] "[INST] < <sys>>\n"

[0068] "You are a marine forecaster.\n"

[0069] "<< / sys> >\n\n"

[0070] "The following is background knowledge:\n"

[0071] "{context}"

[0072] "\n"

[0073] "Please generate a summary of the sea wave warning service based on the above background knowledge and the live sea wave scene. The following is the live sea wave scene: {question}."

[0074] "[ / INST]"

[0075] Among them, "Please generate a summary of the wave warning service based on the above background knowledge and the actual wave situation" represents the task description.

[0076] For the summary task, the first task chain form is used, and its prompt word template is as follows:

[0077] "[INST] < <sys>>\n"

[0078] "You are a marine forecaster.\n"

[0079] "<< / sys> >\n\n"

[0080] "The following is background knowledge:\n"

[0081] "{context}"

[0082] "\n"

[0083] "Please summarize the disaster process based on the above background knowledge, including the impact time, impact range, disaster category and characteristics of the disaster process. The following is the disaster information and summary requirements: {question}."

[0084] Among them, "Please summarize the disaster process based on the above background knowledge, including the impact time, impact range, disaster type and characteristics of the disaster process" represents the task description.

[0085] In the process of implementing ocean early warning question and answer, intent recognition plays a vital role. It ensures that the user's specific needs can be accurately matched with the corresponding task chain, that is, the user's questions are matched to the corresponding task chain through intent recognition. To this end, this embodiment sets basic information for each task chain according to the ocean early warning service. The basic information includes task name and task description, for example:

[0086] {

[0087] "name": "Generate sea wave prediction text for the next four weeks"

[0088] "description": "Based on the weather situation field and historical wave disaster statistics, the effective wave height of China's coastal waters in the next four weeks is predicted and a forecast report is generated."

[0089] }

[0090] Among them, "name" is the task name; "description" is the task description.

[0091] like Figure 3As shown, on the basis of setting the basic information of each type of task chain, this embodiment adopts a dynamic routing strategy based on text similarity. This dynamic routing strategy measures the semantic similarity between the question input by the user and the basic information of each type of task chain by calculating the normalized cosine similarity (its value range is limited to [0, 1]) between the text vector of the question input by the user and the basic information of each type of task chain, sorts these similarity values ​​from high to low, and selects the task chain with the highest similarity, thereby realizing user intent recognition and completing the dynamic matching of the question input by the user and the task chain. In addition, this embodiment also sets a default task chain, which is used to execute this default task chain when the question input by the user matches poorly with each existing type of task chain (that is, the normalized cosine similarity is lower than the preset threshold, which can be 0.5). The default task chain uses the first task chain form, and its prompt word template is:

[0092] "[INST] < <sys>>\n"

[0093] "You are a marine forecaster.\n"

[0094] "<< / sys> >\n\n"

[0095] "The following is background knowledge:\n"

[0096] "{context}"

[0097] "\n"

[0098] "Please answer the following question based on the above background knowledge: {question}."

[0099] "[ / INST]"

[0100] Then in S2, the question is matched with the basic information of each type of task chain to obtain a target task chain matching the question, specifically including: vectorizing the basic information of the question and each type of task chain respectively, specifically using a word embedding model to vectorize the basic information of the question and each type of task chain to obtain a vectorized question and the vectorized basic information of each type of task chain; for each type of task chain, calculating the similarity between the vectorized question and the vectorized basic information of the task chain to obtain the similarity corresponding to the task chain; judging whether the maximum value of all similarities is greater than a preset threshold; if so, taking the task chain corresponding to the maximum value as the target task chain; if not, selecting a default task chain as the target task chain, and the prompt word template of the default task chain includes background knowledge input and question input.

[0101] This embodiment utilizes retrieval enhancement generation technology to extract background knowledge from a vector database. Specifically, the question is matched with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus that matches the question. The target vectorized corpus is the background knowledge.

[0102] like Figure 4As shown, based on the nearest neighbor search and re-ranking method, the question is matched with each vectorized corpus in the vector database to obtain the target vectorized corpus matching the question, which specifically includes:

[0103] (1) Vectorize the problem to obtain a vectorized problem.

[0104] Specifically, the word embedding model is used to vectorize the problem to obtain a vectorized problem.

[0105] (2) Taking the vectorization problem as input, a graph-based approximate nearest neighbor search algorithm is used to search in the vector database to obtain multiple preliminary vectorized corpora that match the vectorization problem.

[0106] By using the Query interface of the vector database, the most relevant top-K1 records are roughly recalled in the vector database through vector similarity matching, that is, K1 preliminary vectorized corpora are obtained. This embodiment specifically uses a graph-based approximate nearest neighbor search algorithm, Hierarchical Navigable Small World (HNSW), to achieve efficient similarity search in a large-scale vector database. The HNSW algorithm constructs a novel dynamic graph by mapping the vectorized corpus to a multi-level pyramid structure according to an exponential decay probability function. The HNSW algorithm starts from the top layer (sparse layer) of the dynamic graph and gradually traverses downward to the bottom layer (dense layer) until the closest vectorized corpus is identified, and K1 preliminary vectorized corpora are obtained. This dynamic graph construction mechanism gives the HNSW algorithm adaptability when processing dynamic data sets that need to be updated regularly, such as marine early warning business information. In addition, the hierarchical characteristics of the HNSW algorithm allow it to quickly narrow the search scope, and significantly improve the retrieval efficiency by excluding a large number of irrelevant data points at a high level. It not only optimizes the search performance, but also ensures the response speed and accuracy when processing large-scale data sets.

[0107] (3) Reordering the multiple preliminary vectorized corpora, and screening the multiple preliminary vectorized corpora based on the reordering results to obtain the target vectorized corpora that matches the vectorization problem.

[0108] In view of the extensiveness and diversity of the professional knowledge corpus in the field of marine early warning, and the high requirements of intelligent question answering for real-time interactive response, the vector database search technology using the HNSW algorithm, while pursuing retrieval efficiency, may sacrifice retrieval accuracy to a certain extent due to its random characteristics, which leads to errors in the relevance ranking of the top-K1 records of rough recall; at the same time, when generating the vector database, the word embedding technology slices the marine early warning corpus and then maps it to a vector, which will cause a certain degree of semantic information loss; the above two factors may have an adverse effect on the accuracy and reliability of the large language model when generating answers. Therefore, this embodiment further adopts a re-ranking mechanism (Re-ranking) to further refine the sorting of the top-K1 records of the rough recall, obtain the top-K2 records, and obtain K2 target vectorized corpora. This mechanism significantly improves the accuracy of the large language model in capturing the business scenarios of the input questions by giving priority to displaying information that is more closely related to the input questions, thereby ensuring that the generated answers are more accurate and appropriate.

[0109] This embodiment uses the bce-reranker-base model to perform the Re-ranking mechanism. The core principle of this model is to use the cross-entropy encoder to perform in-depth feature extraction on the user input questions and the top-K1 records initially recalled, and calculate the text similarity between them, and select K2 records with greater similarity as the final top-K2 records. It is worth noting that since the reranking process only involves the top-K1 records, its data scale is much smaller than the scale of the entire vector database. Therefore, compared with the preliminary recall stage, the impact of reranking on system performance is minimal. In addition, by discarding a certain number of records with the lowest relevance in the reranking stage, the efficiency of large language model reasoning is further optimized.

[0110] This embodiment inserts the question and the target vectorized corpus into the prompt word template of the target task chain, and inputs the ocean early warning large language model to generate an answer. Specifically, the target vectorized corpus is filled into the background knowledge input of the prompt word template of the target task chain, and the question is filled into the question input of the prompt word template of the target task chain to obtain the filled template. The filled template is used as input, and the answer obtained by answering the question using the ocean early warning large language model output.

[0111] This embodiment uses the top-K2 records selected as background knowledge, and inserts them into the prompt word template in the target task chain together with the original question asked by the user, so as to enrich the context of the question and reorganize its logic, and finally complete the prompt word project. This process not only enriches the context of the question, but also enhances its professionalism and coherence. Subsequently, this optimized new question (i.e., the filled template) is used as input and provided to the ocean early warning language model for processing. The ocean early warning language model will comprehensively utilize this information to generate accurate, comprehensive and professional answers. Finally, these answers will be packaged and returned to the user. This method not only improves the quality and relevance of the answers, but also ensures that users can obtain richer and more in-depth information.

[0112] This embodiment collects Chinese natural language data of marine early warning, constructs a marine early warning professional knowledge corpus, further constructs a marine early warning multi-task instruction set based on manual annotation + instruction expansion technology, performs incremental pre-training on a large language model based on the marine early warning professional knowledge corpus, further fine-tunes the incremental pre-trained large language model based on the marine early warning multi-task instruction set, and constructs a marine early warning large language model, such as Figure 5 As shown in FIG. 1 , the method for constructing a large language model for ocean early warning includes:

[0113] (1) Preprocess the Chinese natural language data of ocean warning to obtain the ocean warning corpus, and then combine all the ocean warning corpora into an ocean warning professional knowledge corpus.

[0114] The Chinese natural language data of ocean warning include eight categories of Chinese natural language data related to ocean warning business, including monographs, academic papers, textbooks, marine disaster bulletins, technical manuals, business reports, industry standards, and Internet information. The natural language processing technologies such as text cleaning, deduplication, denoising and purification are used to pre-process the Chinese natural language data of ocean warning to obtain the ocean warning corpus. The ocean warning corpus is in the form of unstructured and unlabeled natural language text data. All ocean warning corpora are stored to form an ocean warning professional knowledge corpus.

[0115] The existing model has the problem of insufficient language adaptability. This is because the existing model is built based on an English public corpus dataset and is not suitable for the Chinese ocean early warning business scenario. This embodiment builds an ocean early warning professional knowledge corpus based on the Chinese natural language data of ocean early warnings, which can solve the above problem well.

[0116] (2) Annotate the marine early warning professional knowledge corpus and generate a seed instruction set. The seed instruction set includes multiple seed instructions, and the seed instructions include questions and answers.

[0117] In this embodiment, 2350 question-answer instructions were manually annotated by experts in the marine early warning industry as a seed instruction set. The seed instruction set is in JSON format, and the structure of the seed instruction is as follows:

[0118] {

[0119] "id": "seek_task_983"

[0120] "name": "Tsunami Alert Product Description"

[0121] "instruction": "Details of the contents that should be included in the tsunami warning product as specified in the Tsunami Warning Product Production Specifications."

[0122] "instances": [

[0123] {

[0124] "input": " "

[0125] "output": "Tsunami warning products should include the following: product identification, including issuing unit, warning level, product title, release time, product number, issuer and consultant information; text content, including warning conclusion, tsunami source information assessment results, tsunami risk assessment results, tsunami confirmation and defense guidelines; tsunami forecast and monitoring forms, including tsunami monitoring forms and fixed-point forecast forms; tsunami forecast maps, including deep sea forecast maps, shore segment forecast maps and inundation forecast maps."

[0126] } ]

[0128] "is_classification": false

[0129] }

[0130] Among them, "id" is the unique identifier of the seed instruction, representing the seed instruction number; "name" is the subject of the seed instruction; "instruction" is the description or explanation of the seed instruction, which details the purpose and usage scenarios of the seed instruction, corresponding to the content of the question asked by the user to the intelligent question-answering system; "input" is the additional input data in the user's question; "output" is the accurate and professional answer expected to be output by the large language model, that is, the answer; "is_classification" is a Boolean value, indicating whether this seed instruction is used for classification tasks. If it is true, it means that the seed instruction is used for classification tasks, such as yes-or-no question answering, multiple-choice question answering, information extraction, etc.; if it is false, it means that the seed instruction is used for other types of tasks, such as content creation, text summarization, reading comprehension, text proofreading, grammar checking, etc., which is used to calculate the ratio of classification tasks to non-classification tasks when evaluating the sample balance of the instruction set.

[0131] (3) Expand the seed instruction set to obtain an extended instruction set, and review and revise the extended instruction set to obtain an ocean early warning multi-task instruction set.

[0132] The seed instruction set is highly professional, but its scale is not enough to meet the generalization requirements of the large language model for the multi-task of ocean early warning. The seed instruction set needs to be expanded to achieve data enhancement. The seed instruction set is expanded to obtain the expanded instruction set, which specifically includes: using the second largest language model deployed locally to expand the seed instruction set to obtain the expanded instruction set. The second largest language model can be chatglm3-6B.

[0133] To ensure data security and compliance, this embodiment deploys the open source Chinese language model chatglm3-6B locally as the second largest language model for seed instruction set enhancement. The seed instruction set is expanded through chatglm3-6B to obtain an extended instruction set, which replaces the use of large-scale human annotated data training to save labor costs. When the seed instruction set is expanded through chatglm3-6B, the input is the seed instruction set and the output is the extended instruction set. For a seed instruction input, chatglm3-6B will output multiple instructions related to the content of the seed instruction and with different forms.

[0134] The existing model has data security and compliance issues. This is because the existing model relies on foreign non-open source model APIs to build instruction sets, which requires uploading data to foreign servers to complete the construction. This poses information security risks and cannot achieve autonomous control. This embodiment uses the second largest language model deployed locally to expand the seed instruction set to obtain an extended instruction set, which can solve the above problems well.

[0135] For the expanded instruction set, this embodiment further performs manual sampling review and revision, removes low-quality instructions caused by the chatglm3-6B model hallucination, and finally obtains 33,035 instructions, converts the instruction format to the Stanford Alpaca standard format, and achieves sample balance between classified tasks / non-classified tasks and marine professional disciplines, and finally obtains a rich and high-quality marine early warning multi-task instruction set. Among them, marine professional disciplines include marine environment, marine chemistry and marine physics.

[0136] For unstructured and unlabeled ocean early warning corpus, this embodiment uses manual annotation + instruction expansion + sampling manual review and revision technology to generate a labeled ocean early warning multi-task instruction set for the ocean early warning business.

[0137] In the ocean early warning multi-task command set, the command samples of the classification task are as follows:

[0138] {

[0139] "input": "Which wave numerical prediction model is suitable for the refined calculation of waves in shallow coastal waters?\nA. WAM mode\nB. WaveWatch III mode\nC. SWAN mode\nD. NOWAMO mode"

[0140] "instruction": "Answer the following multiple-choice questions. Choose A or B or C or D as the answer"

[0141] "output": "C, SWAN model, is a numerical prediction model with high wave simulation accuracy in nearshore shallow waters. It is based on the research results of the third-generation wave model and uses the spectral balance equation to simulate the generation of waves. It can consider a variety of factors that affect waves, such as the refraction effect caused by terrain and the nonlinear interaction between waves. In addition, the SWAN model also considers the dissipation of breaking energy caused by depth limitation, so that it can better adapt to shallow water environments."

[0142] }

[0143] (4) The first language model is incrementally pre-trained using the ocean early warning professional knowledge corpus to obtain an ocean early warning pre-training model.

[0144] Incremental pre-training is the pre-training of a large language model on a large-scale unlabeled marine early warning professional knowledge corpus in an unsupervised learning manner. The first large language model used in this embodiment is the open source Chinese large language model chinese-llama-aplaca2-13B. Of course, other types of large language models can also be used. The first large language model already has the basic understanding ability of common human natural language such as semantics and grammar, but it lacks performance in downstream tasks in professional fields. By inheriting the language ability advantage of the first large language model and injecting marine early warning professional knowledge, the training cost can be greatly reduced. The marine early warning pre-training model after incremental pre-training has the language comprehension ability and long context continuation ability related to the marine early warning professional field, but does not have the ability of question-answering dialogue and the ability to complete specific marine early warning tasks.

[0145] (5) The ocean early warning multi-task instruction set is used to fine-tune the instructions of the ocean early warning pre-training model to obtain the ocean early warning large language model.

[0146] The ocean early warning pre-training model is not designed and optimized for a specific task. In order to further enable the ocean early warning pre-training model to learn the specific language habits, terms and concepts of the ocean early warning and the answer paradigm of the task, and guide the model to achieve accurate and professional answers, this embodiment will fine-tune the instructions in a supervised learning manner based on the ocean early warning multi-task instruction set for the ocean early warning pre-training model, and specifically adopt the Parameter-EfficientFine-Tuning (PEFT) technology. The core strategy of this technology is to freeze most of the model parameters and only fine-tune the instructions for a small part of the model parameters. The instruction fine-tuning updates the parameters of the self-attention module in the large language model to prevent the model from losing the knowledge learned in the incremental pre-training stage, thereby causing catastrophic forgetting. When PEFT is used for instruction fine-tuning, the input is the question of the instruction in the ocean early warning multi-task instruction set, including "input" and "instruction", and the label is the answer to the instruction in the ocean early warning multi-task instruction set, including "output". The use of PEFT technology for fine-tuning instructions not only allows the model to maintain its original human natural language understanding ability and memory of ocean early warning professional knowledge, but also further expands the model's capabilities, enabling it to solve a variety of professional tasks for ocean early warning. At the same time, fine-tuning instructions for only a small number of parameters of the ocean early warning pre-training model can greatly reduce computing costs and achieve efficient use of resources. The model after instruction fine-tuning already has the skills to deeply understand, carefully analyze and comprehensively utilize various text materials, perform logical reasoning, and integrate knowledge to solve complex ocean early warning problems or tasks.

[0147] In this embodiment, the vectorized corpus is obtained by vectorizing the marine early warning corpus. The marine early warning professional knowledge corpus including the marine early warning corpus is stored by vectorization technology to obtain a vector database, such as Figure 6 As shown, the method for constructing a vector database includes:

[0148] (1) Slice the ocean warning corpus to obtain multiple text data.

[0149] The ocean warning corpus is sliced ​​and recursively decomposed into smaller units to obtain multiple text data, which helps to capture the semantic information in the ocean warning corpus.

[0150] (2) Use the word embedding model to process text data to obtain vectorized corpus.

[0151] This embodiment uses word embedding technology to convert text data into a vector form of fixed dimension to accurately capture the semantic connection between words. Word embedding technology is a core method in the field of natural language processing. It expresses semantic similarity by mapping words into continuous vectors in a high-dimensional space, so that words with similar semantics are closer in the vector space. Compared with the traditional bag-of-words model or the unique hot encoding method, word embedding technology provides a dense vector representation, in which each word is represented by a continuous vector rather than a sparse vector, which significantly reduces the computational complexity. This embodiment uses an offline Chinese word embedding model text2vec-base-chinese, which is suitable for Chinese marine early warning corpus. The model can map text data to a 768-dimensional dense vector space, and use cosine similarity to quantify the similarity between text vectors. By minimizing the cosine distance of positive sample pairs (semantically similar sentence pairs) during training, and maximizing the cosine distance of negative sample pairs (semantically dissimilar sentence pairs), the semantic distinction ability of the model is optimized. Through the application of the above technologies, efficient vectorization processing of marine warning corpus can be achieved, laying a solid foundation for subsequent text analysis and information retrieval tasks.

[0152] (3) Store the vectorized corpus in a vector database.

[0153] The Chroma vector database is used to store vectorized corpus, including creating database objects and data sets. The Chroma vector database can quickly output contexts with high weights of similarity to user questions through similarity fuzzy search in high-dimensional vector space. Compared with unstructured corpus storage methods or traditional relational databases that require row-column-based search, vector databases measure the similarity of data item vectors and use distance or angle calculations to quickly locate the data items closest to the query vector. The significant advantage of this method is that it goes beyond traditional keyword matching and realizes deep semantic retrieval, thereby significantly improving the accuracy and relevance of retrieval results. It has obvious advantages in retrieval speed, especially when processing large-scale data, which greatly improves the user experience of the question-answering system. It should be emphasized that the vector database uses a persistent method to store data on the hard disk, rather than the usual use as an in-memory database, which ensures the continuous expansion and updating of the knowledge of the question-answering system.

[0154] The entity field design of each vectorized corpus in the vector database is as follows:

[0155] Collection name: Ocean Early Warning Dataset

[0156] Primary key field: document_id (INT64, primary key)

[0157] Vector field: embedding_vector (FLOAT_VECTOR, storing text content)

[0158] Metadata fields:

[0159] title(VARCHAR, title)

[0160] author (VARCHAR, author)

[0161] year (INT64, year of publication / writing / release)

[0162] abstract(TEXT, text abstract)

[0163] Among them, the primary key field is the serial number of the vectorized corpus, which serves as the unique identifier of the vectorized corpus; the vector field is used to store the text content of the vectorized corpus; the title, author, year and abstract in the metadata field can be empty.

[0164] Through the above steps, this embodiment uses text embedding technology to embed the unstructured marine warning corpus of the marine warning professional knowledge corpus into words, converts it into structured data in the form of low-dimensional dense vectors, and constructs a vector database. Whenever the marine warning corpus is updated, the updated marine warning corpus is vectorized to update the vector database.

[0165] The business work of the marine early warning industry has high requirements for timeliness. In the application of intelligent question and answer, users may ask questions that require real-time data support, such as the task of generating a summary copy of the typhoon process. This type of content creation has extremely high requirements for the timeliness of information. Although the trained marine early warning language model has a deep professional knowledge reserve and can process and understand historical data and general information, the existing marine early warning language model may not be able to provide sufficiently accurate answers to the newly generated content after the model is deployed, such as regular updates of marine forecasts, detailed analysis of specific weather processes, the latest developments in industry dynamics, or cutting-edge advances in early warning technology. In order to meet this challenge, this embodiment adopts a knowledge dynamic update mechanism based on a vector database, which stores the existing knowledge and the relevant knowledge regularly updated by the early warning business into the vector database through text vectorization technology, so as to realize the continuous updating and expansion of the model knowledge. Even if the marine early warning language model has not learned the relevant knowledge, the corresponding content can be directly retrieved from the vector database as the background knowledge input for the model answer. This method does not require retraining, is lower in cost and faster than the common large model fine-tuning, and guarantees the real-time update of the knowledge of the large model by updating the database. Among them, the updated relevant knowledge includes the real-time data stream of marine environmental forecasts, the latest forecasts released by the meteorological department, updates of industry reports, and academic papers related to early warning technology published in scientific research journals, that is, the updated content of Chinese natural language data of marine early warnings.

[0166] This embodiment first collects Chinese natural language data of ocean early warning, constructs an ocean early warning professional knowledge corpus, constructs an ocean early warning multi-task instruction set based on manual annotation + instruction expansion technology, performs incremental pre-training on the large language model based on the ocean early warning professional knowledge corpus, and further fine-tunes the instructions of the incremental pre-trained model based on the ocean early warning multi-task instruction set to construct an ocean early warning large language model. The unstructured ocean early warning corpus of the ocean early warning professional knowledge corpus is embedded in words through text embedded technology, converted into structured data in the form of low-dimensional dense vectors, and a vector database is constructed. Preset multi-task prompt word templates, and combine them according to ocean early warning tasks to construct a task chain. When applied, the questions asked by the user are matched to the corresponding task chain through intent recognition, and the retrieval enhancement generation technology is used to extract background knowledge from the vector database, insert the background knowledge into the task chain, input the ocean early warning large language model, and generate answers.

[0167] The intelligent question-answering method for marine early warning alarms in this embodiment has the following main capabilities:

[0168] (1) Ocean Early Warning Professional Knowledge Q&A: It can handle professional consultations in the field of ocean early warning and provide detailed intelligent analysis and answers based on scientific principles. Users can ask a variety of questions including short answer questions, multiple choice questions and true or false questions.

[0169] (2) Marine early warning document generation: It can automatically generate professional documents for specific marine early warning business scenarios according to the specific needs of the business. For example, it can generate a "typhoon process wave disaster response summary report" based on user requests to support decision-making and post-event evaluation.

[0170] (3) Summary of marine disaster processes: Ability to summarize disaster situations based on specific marine disaster processes and business needs.

[0171] In order to evaluate the professional knowledge question-answering ability of this embodiment, an objective evaluation experiment was conducted. The experimental environment was executed on a high-performance GPU server cluster, and the operating system environment was Red Hat Enterprise Linux Server release4.8.5. During the reasoning process, a single NVIDIA A800-SXM4-80GB graphics processing unit was configured. The experiment selected undergraduate examination questions of marine science majors in domestic higher education institutions to construct a test question bank to conduct a horizontal comparative evaluation between the method of this embodiment and the existing domestic open source large-scale language models. The test question bank consists of 20 multiple-choice questions and 40 judgment questions, a total of 60 questions. The content of the questions covers basic theory to applied technology, including marine science foundations, physical oceanography, marine geology, marine meteorology, marine environmental science, and marine ecology and other disciplines in the field of marine science. The evaluation results are shown in Table 1, which confirms that the method of this embodiment has better accuracy.

[0172] name Large language model for ocean warning Chinese-llama-alpaca2-13B Chatglm3-6B Accuracy 71.7% 53.3% 48.3%

[0173] As shown below, it is a sample answer to the short answer question, which proves that the method of this embodiment has good performance.

[0174] Input: What are the wave observation data from my country's ocean stations?

[0175] Response: The ocean stations that observe waves in my country's coastal waters include: Dalian, Qinhuangdao, Tanggu, Longkou, Chengshantou, Beihuangcheng, Xiaomaidao, Lianyungang, Lvsi, Shengshan, Dachen, Nanji, Beishuang, Xiamen, Zhelang, Naozhou, Boao, Dongfang and Fangchenggang. Observation methods include manual observation and instrument automatic observation. The elements of manual observation of waves at the ocean station include: 1 / 10 big wave height and period, maximum wave height and period, wave direction, etc. The observation time range is 08:00 to 17:00, and the observation time interval is 3 hours. The elements of automatic observation of waves by ocean station instruments include: sampling interval, average wave height and period, 1 / 10 big wave height and period, maximum wave height and period, effective wave height and period, wave direction, etc. The observation time range is 24 hours, and the observation time interval is 1 hour.

[0176] This embodiment provides an intelligent question-answering method for marine early warning based on a large language model. Through a fluent human-computer dialogue, it realizes business needs such as forecast service assistant, historical disaster analysis, and professional knowledge question-answering, and improves the ability of early warning knowledge retrieval, management, and application.

[0177] An embodiment of the present application also provides an application scenario, which applies the above-mentioned ocean early warning intelligent question and answer method based on a large language model. Specifically, the ocean early warning intelligent question and answer method based on a large language model provided in this embodiment can be applied in the ocean early warning intelligent question and answer scenario. The ocean early warning intelligent question and answer scenario includes a question acquisition link, a question processing link and a result display link. The question acquisition link is used to obtain the questions input by the user, the question processing link is used to process the questions input by the user and obtain answers, and the result display link is used to display the answers to the user. The ocean early warning intelligent question and answer method based on a large language model provided in this embodiment belongs to the question processing link. Example

[0178] Based on the same inventive concept, the embodiment of the present application also provides a large language model-based ocean early warning intelligent question-answering device for implementing the large language model-based ocean early warning intelligent question-answering method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiment of the ocean early warning intelligent question-answering device based on a large language model provided below can refer to the limitations of the ocean early warning intelligent question-answering method based on a large language model above, and will not be repeated here.

[0179] like Figure 7 As shown, an intelligent question-answering device for ocean early warning based on a large language model is provided, and the intelligent question-answering device for ocean early warning based on a large language model includes:

[0180] The question input module is used to obtain questions input by users.

[0181] A task chain matching module is used to match the question with the basic information of each type of task chain to obtain a target task chain that matches the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge questions and answers, task chains for true-or-false tasks in professional knowledge questions and answers, task chains for short-answer tasks in professional knowledge questions and answers, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain.

[0182] A vectorized corpus matching module is used to match the question with each vectorized corpus in the vector database based on the nearest neighbor search and reordering method to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database.

[0183] An answer generation module is used to fill the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, fill the question into the question input of the prompt word template of the target task chain, obtain a filled template, and use the filled template as input to answer the question using the ocean early warning large language model output to obtain an answer; the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model.

[0184] like Figure 8 As shown, another ocean early warning intelligent question-answering device based on a large language model is provided, and the ocean early warning intelligent question-answering device based on a large language model includes:

[0185] The corpus collection module is used to collect Chinese natural language data on marine early warnings. After preprocessing steps such as text cleaning, deduplication, denoising and purification, a high-quality and highly available marine early warning professional knowledge corpus is formed.

[0186] The instruction set building module is used to build an ocean early warning multi-task instruction set based on the ocean early warning professional knowledge corpus and using manual annotation + multi-task instruction expansion technology.

[0187] The model training module is used to perform incremental pre-training on the large language model based on the ocean early warning professional knowledge corpus, and further fine-tune the instructions of the incremental pre-trained model based on the ocean early warning multi-task instruction set to build an ocean early warning large language model.

[0188] The data storage module is used to embed the unstructured marine warning corpus of the marine warning professional knowledge corpus into words through text embedding technology, convert it into structured data in the form of low-dimensional dense vectors, and build a vector database.

[0189] The model reasoning module is used to preset multi-task prompt word templates, combine them according to the ocean early warning tasks, build task chains, and match the questions asked by users to the corresponding task chains through intent recognition; use retrieval enhancement generation technology to extract background knowledge from the vector database, insert the background knowledge into the task chain, input the ocean early warning large language model, and generate answers. Example

[0190] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig. 9As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an intelligent question-answering method for marine early warning based on a large language model is implemented.

[0191] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0192] In an exemplary embodiment, a computer device is also provided, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the ocean early warning intelligent question-answering method based on a large language model as described in Example 1. Example

[0193] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for intelligent question-answering of ocean early warning based on a large language model described in Example 1 is implemented. Example

[0194] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the ocean early warning intelligent question-answering method based on a large language model described in Example 1.

[0195] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0196] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. An intelligent question-answering method for ocean early warning based on a large language model, characterized in that: The ocean early warning intelligent question answering method based on the large language model includes: Get user input questions; Matching the question with the basic information of each type of task chain to obtain a target task chain matching the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge quizzes, task chains for true-false tasks in professional knowledge quizzes, task chains for short-answer tasks in professional knowledge quizzes, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain; The question is matched with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database; The target vectorized corpus is filled into the background knowledge input of the prompt word template of the target task chain, and the question is filled into the question input of the prompt word template of the target task chain to obtain a filled template, and the filled template is used as input to answer the question using the ocean early warning large language model output to obtain an answer; the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model; Match the question with the basic information of each type of task chain to obtain a target task chain matching the question, specifically including: Performing vectorization processing on the problem and basic information of each type of task chain respectively to obtain vectorized problem and vectorized basic information of each type of task chain; For each type of task chain, calculate the similarity between the vectorized problem and the vectorized basic information of the task chain to obtain the similarity corresponding to the task chain; Determine whether the maximum value of all the similarities is greater than a preset threshold; If yes, then the task chain corresponding to the maximum value is used as the target task chain; If not, a default task chain is selected as the target task chain; the prompt word template of the default task chain includes background knowledge input and question input; Matching the question with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question, specifically including: Vectorizing the problem to obtain a vectorized problem; Taking the vectorization problem as input, searching in a vector database using a graph-based approximate nearest neighbor search algorithm to obtain a plurality of preliminary vectorization corpora matching the vectorization problem; The multiple preliminary vectorized corpora are reordered, and the multiple preliminary vectorized corpora are screened based on the reordering result to obtain a target vectorized corpus matching the vectorization problem.

2. The ocean early warning intelligent question answering method based on a large language model according to claim 1 is characterized in that: The vector database construction method includes: Slice the marine warning corpus to obtain multiple text data; Processing the text data using a word embedding model to obtain vectorized corpus; The vectorized corpus is stored in a vector database.

3. The ocean early warning intelligent question answering method based on a large language model according to claim 1 is characterized in that: The method for constructing the large language model for ocean early warning includes: Preprocessing the marine warning Chinese natural language data to obtain marine warning corpus, and forming a marine warning professional knowledge corpus with all the marine warning corpus; Annotating the marine early warning professional knowledge corpus to generate a seed instruction set; the seed instruction set includes a plurality of seed instructions, and the seed instructions include questions and answers; Expanding the seed instruction set to obtain an extended instruction set, and reviewing and revising the extended instruction set to obtain an ocean early warning multi-task instruction set; Using the ocean early warning professional knowledge corpus to incrementally pre-train the first language model to obtain an ocean early warning pre-training model; The ocean early warning multi-task instruction set is used to fine-tune the instructions of the ocean early warning pre-training model to obtain an ocean early warning large language model.

4. The ocean early warning intelligent question answering method based on a large language model according to claim 3 is characterized in that: Expanding the seed instruction set to obtain an extended instruction set specifically includes: expanding the seed instruction set using a second largest language model deployed locally to obtain an extended instruction set.

5. An intelligent question-answering device for marine early warning based on a large language model, characterized in that: The ocean early warning intelligent question-answering device based on the large language model includes: Question input module, used to obtain questions input by users; A task chain matching module is used to match the question with the basic information of each type of task chain to obtain a target task chain that matches the question; the types of task chains include task chains for multiple-choice tasks in professional knowledge quizzes, task chains for true-or-false tasks in professional knowledge quizzes, task chains for short-answer tasks in professional knowledge quizzes, task chains for text generation tasks, and task chains for summary tasks; the prompt word template of the task chain includes background knowledge input, task description, and question input; the basic information includes the task name and task description of the task chain; A vectorized corpus matching module is used to match the question with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question; the vectorized corpus is obtained by vectorizing the ocean warning corpus; whenever the ocean warning corpus is updated, the updated ocean warning corpus is vectorized to update the vector database; an answer generation module, for filling the target vectorized corpus into the background knowledge input of the prompt word template of the target task chain, filling the question into the question input of the prompt word template of the target task chain, obtaining a filled template, and using the filled template as input to answer the question using the ocean early warning large language model output to obtain an answer; the ocean early warning large language model is obtained after incremental pre-training and instruction fine-tuning of the large language model; Match the question with the basic information of each type of task chain to obtain a target task chain matching the question, specifically including: Performing vectorization processing on the problem and basic information of each type of task chain respectively to obtain vectorized problem and vectorized basic information of each type of task chain; For each type of task chain, calculate the similarity between the vectorized problem and the vectorized basic information of the task chain to obtain the similarity corresponding to the task chain; Determine whether the maximum value of all the similarities is greater than a preset threshold; If yes, then the task chain corresponding to the maximum value is used as the target task chain; If not, a default task chain is selected as the target task chain; the prompt word template of the default task chain includes background knowledge input and question input; Matching the question with each vectorized corpus in the vector database based on the nearest neighbor search and re-ranking method to obtain a target vectorized corpus matching the question, specifically including: Vectorizing the problem to obtain a vectorized problem; Taking the vectorization problem as input, searching in a vector database using a graph-based approximate nearest neighbor search algorithm to obtain a plurality of preliminary vectorization corpora matching the vectorization problem; The multiple preliminary vectorized corpora are reordered, and the multiple preliminary vectorized corpora are screened based on the reordering result to obtain a target vectorized corpus matching the vectorization problem.

6. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the ocean early warning intelligent question-answering method based on a large language model as described in any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the ocean early warning intelligent question-answering method based on a large language model described in any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the ocean early warning intelligent question-answering method based on a large language model described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Policy question and answer method and system based on large language model and knowledge graph technology

    CN117725170A

  • Ocean emergency event processing method and device

    CN118014805A