Question-answering model training method, question-answering method and apparatus, and terminal device

By introducing answer samples of specific questions and their cited documents into a large-scale pre-trained language model for augmented training, the problem of lack of reference documents for answers is solved, achieving high accuracy and authority of the answers and improving the clarity of legal liability attribution.

WO2026036531A1PCT designated stage Publication Date: 2026-02-19SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129964
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-14
Filing Date
2024-11-05
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Current large-scale pre-trained language models lack reference documentation when outputting answers, resulting in low credibility and authority, especially in situations where the attribution of legal responsibility is unclear, as they cannot provide clear evidence.

Method used

By acquiring specific questions in a specific domain and their corresponding answers with cited documents, the basic question-answering model is enhanced and trained to generate the target question-answering model, which can more accurately understand user questions and generate answers with cited documents.

Benefits of technology

It improves the accuracy and authority of answers to specific questions, enabling the target question-answering model to generate answers that meet specific requirements and include cited documents, thereby enhancing the credibility and persuasiveness of decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129964_19022026_PF_FP_ABST
    Figure CN2024129964_19022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applicable to the technical field of computers, and provides a question-answering model training method, a question-answering method and apparatus, and a terminal device. The question-answering model training method comprises: acquiring a first question-answering data set, the first question-answering data set comprising two or more police affairs enhancement samples, and the police affairs enhancement samples each comprising: a police affairs question and an answer carrying a reference document corresponding to the police affairs question; using a police affairs question and a corresponding answer carrying a reference document in a police affairs enhancement sample as a model input sample and a model output sample, respectively, performing enhancement training on a basic question-answering model, and obtaining a target question-answering model, the basic question-answering model being obtained by training on the basis of a second question-answering data set, and the second question-answering data set comprising a plurality of questions and an answer corresponding to each question. By means of the described method, the target question-answering model has the capability of generating answer information carrying a reference document, and can provide more accurate and authoritative service support in subsequent police affairs work.
Need to check novelty before this filing date? Find Prior Art

Description

[Rule 26 correction 09.04.2025] Question and answer model training method, question and answer method, device and terminal equipment

[0001] [Rule 26 correction 09.04.2025] This application claims priority to the Chinese patent application No. 202411117628.X, filed on August 14, 2024, and entitled "Question and answer model training method, question and answer method, device and terminal equipment", the entire content of which is incorporated herein by reference. [Rule 26 correction 09.04.2025] TECHNICAL FIELD

[0002] [Rule 26 correction 09.04.2025] The present application belongs to the field of computer technology, and particularly relates to a question and answer model training method, a question and answer method, a device and a terminal equipment. [Rule 26 correction 09.04.2025] BACKGROUND

[0003] [Rule 26 correction 09.04.2025] With the rapid development of artificial intelligence technology, especially the excellent performance of large-scale pre-training language models in diversified tasks, the application of large-scale pre-training language models in specific work is increasingly widespread. In specific work, the use of large-scale pre-training language models can realize knowledge question and answer, that is, through intelligent analysis of the consultation sentence or actual scene text input in the large-scale pre-training language model, the large-scale pre-training language model understands and analyzes the input content, and generates the corresponding answer or suggestion.

[0004] [Rule 26 correction 09.04.2025] However, the current large-scale pre-training language model does not provide the reference document corresponding to the answer when outputting the answer, and due to the lack of reference document support, the credibility and authority of the output answer are reduced; in the context involving legal liability, it may also lead to unclear attribution of legal liability, and cannot provide clear evidence support for legal liability. [Rule 26 correction 09.04.2025] SUMMARY

[0005] [Rule 26 correction 09.04.2025] The embodiments of the present application provide a question and answer model training method, a question and answer method, a device and a terminal equipment, which can solve the problems of low credibility and authority caused by the current large-scale pre-training language model not providing the reference document corresponding to the answer when outputting the answer, and unclear attribution of legal liability.

[0006] [Rule 26 correction 09.04.2025] In a first aspect, the embodiments of the present application provide a question and answer model training method, comprising:

[0007] [According to Rule 26, correct on 09.04.2025] obtaining a first question and answer dataset; the first question and answer dataset comprises two or more specific augmented samples, the specific augmented sample comprising: a specific question and an answer corresponding to the specific question with a reference document;

[0008] [According to Rule 26, correct on 09.04.2025] performing augmented training on the basis question and answer model to obtain a target question and answer model, by taking the specific question and the corresponding answer with the reference document in the specific augmented sample as a model input sample and a model output sample respectively; the basis question and answer model is trained based on a second question and answer dataset, the second question and answer dataset comprising a plurality of questions and an answer corresponding to each question.

[0009] [According to Rule 26, correct on 09.04.2025] In a possible implementation manner of the first aspect, before obtaining the first question and answer dataset, further comprising:

[0010] [According to Rule 26, correct on 09.04.2025] searching at least one document related to the specific question from a preset knowledge base for the input specific question;

[0011] [According to Rule 26, correct on 09.04.2025] inputting the specific question, the searched document related to the specific question, and a preset guide text into a preset language model to obtain an answer corresponding to the specific question with a reference document, the guide text being used to guide the result output mode of the preset language model;

[0012] [According to Rule 26, correct on 09.04.2025] generating the specific augmented sample based on the specific question and the corresponding answer with the reference document.

[0013] [According to Rule 26, correct on 09.04.2025] In a possible implementation manner of the first aspect, after obtaining the target question and answer model, further comprising:

[0014] [According to Rule 26, correct on 09.04.2025] obtaining a preference dataset; the preference dataset comprises positive sample data and negative sample data; the positive sample data comprises a specific question sample and a corresponding reference answer, the reference answer with a reference document; the negative sample data comprises the specific question sample and an answer output by the target question and answer model based on the specific question sample;

[0015] [According to Rule 26, correct on 09.04.2025] performing preference alignment optimization training on the target question and answer model based on the preference dataset to obtain an optimized target question and answer model.

[0016] [Rule 26 Corrected on 09.04.2025] In a possible implementation manner of the first aspect, the preference alignment optimization training of the target question and answer model based on the preference data set comprises:

[0017] [Rule 26 Corrected on 09.04.2025] inputting a specific question sample in the positive sample data into the current target question and answer model to obtain a prediction result corresponding to the positive sample data, and calculating a first loss between the prediction result corresponding to the positive sample data and the reference answer;

[0018] [Rule 26 Corrected on 09.04.2025] inputting a specific question sample in the negative sample data into the current target question and answer model to obtain a prediction result corresponding to the negative sample data, and calculating a second loss between the prediction result corresponding to the negative sample data and the corresponding answer in the negative sample data;

[0019] [Rule 26 Corrected on 09.04.2025] performing weighted summation on the first loss and the second loss to obtain a target loss, wherein the weight of the second loss is less than the weight of the first loss;

[0020] [Rule 26 Corrected on 09.04.2025] iteratively updating model parameters of the current target question and answer model according to the target loss until the target question and answer model meets a preset first training stop condition.

[0021] [Rule 26 Corrected on 09.04.2025] In a possible implementation manner of the first aspect, the obtaining of the first question and answer data set comprises:

[0022] [Rule 26 Corrected on 09.04.2025] obtaining a corresponding number of specific enhanced samples in a preset proportion based on the number of questions contained in the second question and answer data set to obtain the first question and answer data set, wherein the number of specific enhanced samples contained in the first question and answer data set is less than the number of questions contained in the second question and answer data set.

[0023] [Rule 26 Corrected on 09.04.2025] In a possible implementation manner of the first aspect, the enhancement training of the basic question and answer model by taking the specific question in the specific enhanced sample and the corresponding answer with the reference document as model input sample and model output sample respectively comprises:

[0024] [Rule 26 Corrected on 09.04.2025] inputting the specific question in the specific enhanced sample into the basic question and answer model to obtain an initial prediction answer output by the basic question and answer model;

[0025] [According to Rule 26, correct on 09.04.2025] calculate the loss between the initial predicted answer and the corresponding model output sample;

[0026] [According to Rule 26, correct on 09.04.2025] iteratively update the model parameters of the basic question and answer model according to the loss between the initial predicted answer and the corresponding model output sample until the basic question and answer model meets the preset second training stop condition.

[0027] [According to Rule 26, correct on 09.04.2025] In a second aspect, the embodiments of the present application provide a question and answer method, comprising:

[0028] [According to Rule 26, correct on 09.04.2025] obtaining a target question;

[0029] [According to Rule 26, correct on 09.04.2025] input the target question into a target question and answer model to obtain an answer with a reference document corresponding to the target question, wherein the target question and answer model is trained by the question and answer model training method of any one of the above first aspect.

[0030] [According to Rule 26, correct on 09.04.2025] In a third aspect, the embodiments of the present application provide a question and answer model training device, comprising:

[0031] [According to Rule 26, correct on 09.04.2025] a first obtaining module for obtaining a first question and answer data set; the first question and answer data set includes two or more specific enhanced samples, the specific enhanced sample includes: a specific question and an answer with a reference document corresponding to the specific question;

[0032] [According to Rule 26, correct on 09.04.2025] an enhanced training module for taking the specific question in the specific enhanced sample and the corresponding answer with the reference document as model input sample and model output sample respectively, and performing enhanced training on a basic question and answer model to obtain a target question and answer model; the basic question and answer model is trained based on a second question and answer data set, and the second question and answer data set includes a plurality of questions and answers corresponding to each question.

[0033] [According to Rule 26, correct on 09.04.2025] In a fourth aspect, the embodiments of the present application provide a question and answer device, comprising:

[0034] [According to Rule 26, correct on 09.04.2025] a second obtaining module for obtaining a target question;

[0035] [Rule 26 amended on 09.04.2025] The generation module is configured to input the target question into a target question answering model to obtain an answer corresponding to the target question with a reference document, wherein the target question answering model is trained by using the question answering model training method according to any one of the first aspect.

[0036] [Rule 26 amended on 09.04.2025] In a fifth aspect, the embodiments of the present application provide a terminal device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the question answering model training method according to any one of the first aspect when executing the computer program; or the processor implements the question answering method according to the second aspect when executing the computer program.

[0037] [Rule 26 amended on 09.04.2025] In a sixth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the question answering model training method according to any one of the first aspect; or the computer program is executed by the processor to implement the question answering method according to the second aspect.

[0038] [Rule 26 amended on 09.04.2025] In a seventh aspect, the embodiments of the present application provide a computer program product, which, when running on a terminal device, causes the terminal device to execute the question answering model training method according to any one of the first aspect; or, when running on a terminal device, causes the terminal device to execute the question answering method according to the second aspect.

[0039] [Rule 26 amended on 09.04.2025] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0040] [Rule 26 is corrected on 09.04.2025] The first aspect of the embodiment of the present application provides a question and answer model training method. First, by obtaining specific enhanced samples containing specific questions and answers with reference documents corresponding to the specific questions in a specific field, reliable source information can be attached to the specific questions, so that the answers corresponding to the specific questions have high accuracy and authority. Then, by introducing the specific questions in the specific enhanced samples as model input samples and taking the answers with reference documents corresponding to the specific questions as model output samples, the basic question and answer model is enhanced and trained, so that the enhanced and trained basic question and answer model (i.e. the target question and answer model) can more accurately understand the user's question and quickly generate answers that meet specific requirements and have authority, that is, the target question and answer model has the ability to generate answers with reference documents. The second aspect of the embodiment of the present application provides a question and answer method. When a specific person is processing a target question, the target question and answer model can be used to generate answers with reference documents corresponding to the target question, and the answers with reference documents generated by the target question and answer model can be used to make more wise and accurate decisions to improve the credibility and persuasiveness of the decisions made by the specific person when processing the target question.

[0041] [Rule 26 is corrected on 09.04.2025] It can be understood that the beneficial effects of the above-mentioned third aspect, fifth aspect to seventh aspect can be referred to the related description in the above-mentioned first aspect, which will not be repeated here. The beneficial effects of the above-mentioned fourth aspect, fifth aspect to seventh aspect can be referred to the related description in the above-mentioned second aspect, which will not be repeated here. [Rule 26 is corrected on 09.04.2025] BRIEF DESCRIPTION OF DRAWINGS

[0042] [Rule 26 is corrected on 09.04.2025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0043] [Rule 26 is corrected on 09.04.2025] Fig. 1 is a flowchart of a question and answer model training method according to an embodiment of the present application;

[0044] [Rule 26 is corrected on 09.04.2025] Fig. 2 is a flowchart of another question and answer model training method according to an embodiment of the present application;

[0045] [Rule 26 is corrected on 09.04.2025] Fig. 3 is a flowchart of generating a specific enhanced sample according to an embodiment of the present application.

[0046] [According to Rule 26 Correction 09.04.2025] FIG. 4 is a flowchart of a sub-step of step S105 according to an embodiment of the present application;

[0047] [According to Rule 26 Correction 09.04.2025] FIG. 5 is a flowchart of another method for training a question and answer model according to an embodiment of the present application;

[0048] [According to Rule 26 Correction 09.04.2025] FIG. 6 is a flowchart of a sub-step of step S107 according to an embodiment of the present application;

[0049] [According to Rule 26 Correction 09.04.2025] FIG. 7 is a schematic diagram of an overall method for ORPO optimization according to an embodiment of the present application;

[0050] [According to Rule 26 Correction 09.04.2025] FIG. 8 is a flowchart of a method for question and answer according to an embodiment of the present application;

[0051] [According to Rule 26 Correction 09.04.2025] FIG. 9 is a schematic diagram of a device for training a question and answer model according to an embodiment of the present application;

[0052] [According to Rule 26 Correction 09.04.2025] FIG. 10 is a schematic diagram of a specific device for question and answer according to an embodiment of the present application;

[0053] [According to Rule 26 Correction 09.04.2025] FIG. 11 is a schematic diagram of a terminal device according to an embodiment of the present application. [According to Rule 26 Correction 09.04.2025] DETAILED DESCRIPTION

[0054] [According to Rule 26 Correction 09.04.2025] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as a particular system architecture, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0055] [According to Rule 26 Correction 09.04.2025] It is to be understood that the terminology "including", when used in the present specification and in the accompanying claims, does not exclude the presence of other features, integers, steps, operations, elements, and / or groups thereof.

[0056] [Corresponding to Rule 26 Correction 09.04.2025] It should also be understood that the term "and / or" as used in the specification and the appended claims, means any one or more of the associated listed items, as well as all possible combinations of the items, and includes these combinations.

[0057] [Corresponding to Rule 26 Correction 09.04.2025] As used in this application specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0058] [Corresponding to Rule 26 Correction 09.04.2025] In addition, in the description of the application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0059] [Corresponding to Rule 26 Correction 09.04.2025] In the present application specification, the reference "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in additional some embodiments" and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized.

[0060] [Corresponding to Rule 26 Correction 09.04.2025] With the rapid development of artificial intelligence technology, especially the excellent performance of large-scale pre-training language models in diversified tasks, its application in specific work is increasingly widespread. In specific work, the use of large-scale pre-training language models can realize knowledge question and answer, that is, through intelligent analysis of the consultation sentence or actual scene text input in the large-scale pre-training language model, the large-scale pre-training language model understands and analyzes the input content, and generates the corresponding answer or suggestion.

[0061] [According to Rule 26, the amendment is made on 09.04.2025] However, the current large-scale pre-training language model does not provide the reference document corresponding to the answer when outputting the answer, and due to the lack of reference document support, the credibility and authority of the output answer are reduced; in the context of legal liability, it may also cause the attribution of legal liability to be unclear, and cannot provide clear evidence support for legal liability.

[0062] [According to Rule 26, the amendment is made on 09.04.2025] In order to solve the above technical problems, the embodiment of the application provides a question and answer model training method, first, by acquiring specific enhanced samples containing specific questions and answers with reference documents corresponding to the specific questions in a specific field, reliable source information can be attached to the specific questions, so that the answers corresponding to the specific questions have high accuracy and authority; then, by introducing the specific question in the specific enhanced sample as the model input sample, and taking the answer with the reference document corresponding to the specific question as the model output sample, the basic question and answer model is enhanced and trained, so that the basic question and answer model after enhancement training (namely the target question and answer model) can more accurately understand the user's question and quickly generate answers that meet the specific requirements and have authority, that is, the target question and answer model has the ability to generate answers with reference documents.

[0063] [According to Rule 26, the amendment is made on 09.04.2025] In order to illustrate the technical solutions described in the application, specific embodiments are used to illustrate the technical solutions described in the application.

[0064] [According to Rule 26, the amendment is made on 09.04.2025] Figure 1 shows a flowchart of the question and answer model training method provided by an embodiment of the application. Referring to Figure 1, the question and answer model training method includes steps S104 and S105.

[0065] [According to Rule 26, the amendment is made on 09.04.2025] Step S104: acquiring a first question and answer data set; the first question and answer data set includes two or more specific enhanced samples, and the specific enhanced sample includes: a specific question and an answer with a reference document corresponding to the specific question.

[0066] [Amended according to Rule 26 09.04.2025] To quickly obtain specific enhanced samples, first, a large number of specific question and answer data pairs are collected. In specific scenarios such as legal consultation, case investigation, and public safety, specific question and answer data generated in specific work can be collected through various different sources. Among them, various different sources include open source legal data, legal regulation documents, legal question and answer books, etc. After obtaining a large amount of specific question and answer data, by analyzing and processing the specific question and answer data, a large number of specific question and answer data pairs can be obtained. Each specific question and answer data pair usually includes three main parts: a specific question, a document related to the specific question, and an answer corresponding to the specific question. The specific question is a specific question related to the specific work; the document related to the specific question is reference material supporting or explaining the answer to the specific question, which has high authority and accuracy; the answer corresponding to the specific question is a specific answer to the specific question.

[0067] [Amended according to Rule 26 09.04.2025] Specifically, for specific question and answer data in open source legal data, there are many data pairs that already have a direct "question and answer" structure, that is, specific question and answer data already exists in the format of specific question and answer data pairs. Extracting these specific question and answer data pairs can directly obtain the specific question and answer data pairs contained in the open source legal data.

[0068] [Amended according to Rule 26 09.04.2025] For specific question and answer data in legal regulation documents, these legal regulation documents often do not have a direct "question and answer" structure, so specific question and answer data pairs contained in legal regulation documents need to be extracted. The specific way to extract specific question and answer data pairs from legal regulation documents is as follows: First, split according to paragraphs, chapters, or specific keywords to identify specific related parts; then use regular expressions to quickly locate text blocks containing specific keywords or patterns; finally, use other natural language processing techniques (such as named entity recognition, syntax analysis, relationship extraction, etc.) to extract specific information related to the specific to obtain specific question and answer data pairs contained in legal regulation documents.

[0069] [Amended according to Rule 26 09.04.2025] For specific question and answer data in legal question and answer books, it is often stored in epub format files. To obtain specific question and answer data pairs in legal question and answer books, first, use a special epub document parsing tool to extract the text content; then use an HTML and XML parsing library like BeautifulSoup to further parse and structure the extracted text content; finally, analyze the document structure (such as title, paragraph, list, etc.) to identify and extract specific question and answer data pairs in legal question and answer books.

[0070] [According to Rule 26, correct on 09.04.2025] After a large number of specific question and answer data pairs are collected, a large number of specific question and answer data pairs can be used to build a preset knowledge base, i.e. a specific knowledge base, so that specific personnel can query and reference the content related to the specific work at any time. For example, in a specific work, specific personnel can quickly query related specific knowledge, actual cases, laws and regulations, etc. in the specific knowledge base through keywords, question types, etc. to make correct decisions when handling cases or providing legal advice. When specific personnel handle cases or provide legal advice, they need to reference relevant laws and regulations or actual cases to make decisions. Through the reference function of the specific knowledge base, specific personnel can easily find and reference the found laws and regulations, etc. as important reference and reference sources, and thus ensure that the decisions or suggestions made have sufficient basis and authority.

[0071] [According to Rule 26, correct on 09.04.2025] Building a preset knowledge base using a large number of specific question and answer data pairs specifically includes: first, extracting specific question parts from all collected specific question and answer data pairs, then, performing text preprocessing on all extracted specific questions, including removing punctuation, removing stop words, removing irrelevant information such as numbers, and performing processing such as word segmentation and lemmatization, so as to obtain text data suitable for word embedding model processing; then, using BGE model or other language model as text vector model, inputting the preprocessed text data into BGE model to obtain vector representation of each word or phrase in the preprocessed specific question in high-dimensional space. These vectors can capture the semantic relationship between words or phrases, and the vectors of all words or phrases are combined in some form (such as taking average, sum, weighted sum, etc.) to obtain the text vector of the entire specific question, i.e. to obtain the first text vector for representing the semantic features of the specific question.

[0072] [According to Rule 26, correct on 09.04.2025] Next, after obtaining all first text vectors, all first text vectors are stored in the FAISS vector library. Each first text vector corresponds to a related document and a corresponding answer, and in the FAISS vector library, the association of the first text vector with the related document and the answer usually exists in the form of meta information, such as document ID or answer ID, which is used to find the actual document and answer content in the preset knowledge base. The preset knowledge base includes the FAISS vector library and other forms of data storage and retrieval mechanism. For example, the preset knowledge base stores the first text vector and the meta information, as well as the document, answer, law and regulation content corresponding to the first text vector. The content stored in the preset knowledge base is usually stored in the form of database, file system or other form for easy retrieval and display.

[0073] [According to Rule 26, correct 09.04.2025] In this embodiment, by converting the specific question into a first text vector, and storing each first text vector, the document related to the first text vector, and the answer corresponding to the first text vector into the preset knowledge base, the rapidity and accuracy of vector operation can be utilized to quickly locate the similar specific question and answer data pair in the specific knowledge base to the new specific question encountered in the specific work, so as to improve the retrieval efficiency and reduce the waiting time of the specific personnel when querying information. That is, when querying a new specific question generated in a specific work in the preset knowledge base, the specific question related to the new specific question that already exists in the preset knowledge base can be found by calculating the similarity between the text vector corresponding to the new specific question and the first text vector stored in the preset knowledge base. Once the specific question related to the new specific question that already exists in the preset knowledge base is found, the document associated with the new specific question and the corresponding answer can be directly extracted from the preset knowledge base for decision-making on the new specific question.

[0074] [According to Rule 26, correct 09.04.2025] In some embodiments, FIG. 2 shows a flowchart of another question and answer model training method provided by an embodiment of the present application. Referring to FIG. 2, before step S104, the question and answer model training method further includes steps S101 to S103.

[0075] [According to Rule 26, correct 09.04.2025] Step S101: searching at least one document related to the input specific question from the preset knowledge base.

[0076] [According to Rule 26, correct 09.04.2025] Specifically, when an input specific question is obtained, the input specific question can be a new specific question, that is, it can not have corresponding related documents and corresponding answers in the preset knowledge base. Due to the diversity of natural language text and the richness of semantics, the new specific question can have semantic association or similarity with some specific questions in the preset knowledge base. Therefore, the technology of text vectorization and similarity search can be used to search at least one specific question related to the input specific question from the preset knowledge base according to the input specific question, and the related documents and corresponding answers corresponding to the at least one specific question related to the input specific question, so as to provide relevant context information and reference answers for answering the input specific question.

[0077] [According to Rule 26, correct 09.04.2025] Specifically, searching at least one document related to the input specific question from the preset knowledge base according to the input specific question includes: first, text preprocessing is performed on the input specific question, including removing irrelevant information such as punctuation marks, stop words, numbers, and performing processing such as word segmentation and morphological restoration, so as to convert the input specific question into a text form more conducive to the understanding of the preset language model; then, the BGE model is used to vectorize the preprocessed input specific question. The BGE model will convert the input specific question into a high-dimensional vector, i.e. a second text vector, which can represent the semantic features of the input specific question; after obtaining the second text vector, the similarity between the second text vector and all first text vectors in the preset knowledge base is calculated respectively, and the similarity scores corresponding to the second text vector and each first text vector are obtained, so as to find specific questions related to the input specific question from the preset knowledge base.

[0078] [According to Rule 26, correct 09.04.2025] Finally, the specific question related to the input specific question, the document related to the input specific question and the corresponding answer are determined according to the similarity score, i.e. the specific question and answer data pair related to the input specific question. Specifically, once the similarity scores corresponding to the second text vector and each first text vector are calculated, the specific question and answer data pair related to the input specific question can be determined according to these similarity scores. First, the calculated similarity scores are sorted from high to low, so that the specific question most related to the input specific question will be ranked first. Then, according to the actual situation, the top k specific questions with high similarity scores, related documents and corresponding answers are selected from the sorting results, and the selected results are taken as the specific question and answer data pair related to the input specific question.

[0079] [According to Rule 26, correct 09.04.2025] Step S102: input the input specific question, the searched document related to the input specific question and the preset guide text into the preset language model to obtain an answer corresponding to the input specific question with a reference document, and the guide text is used to guide the result output mode of the preset language model.

[0080] [Rule 26 Corrected on 09.04.2025] Specifically, after obtaining the specific question and answer data pair related to the input specific question, the relevant documents and preset guide texts corresponding to each specific question and answer data pair related to the input specific question are extracted, and when there are k specific question and answer data pairs related to the input specific question, there are also k relevant documents and preset guide texts. The value of k can be determined according to actual conditions. For example, k is 8. Then, the input specific question, the first k documents related to the input specific question, and the preset guide texts contained in each relevant document are input into the preset language model. The preset language model understands and analyzes the input data, and summarizes and induces the results of understanding and analysis, and can generate answer information with reference documents corresponding to the input specific question.

[0081] [Rule 26 Corrected on 09.04.2025] Specifically, the preset language model can be a ChatGPT model or other language model. The input specific question, the first k documents related to the input specific question, and the preset guide texts contained in each relevant document are input into the ChatGPT model or other language model. The ChatGPT model or other language model uses its powerful natural language processing capabilities and deep learning algorithms to deeply analyze and understand the input specific question, the first k relevant documents, and the corresponding preset guide texts. In the process of generating answers, ChatGPT will pay special attention to the use of relevant documents. That is, key information, legal provisions or actual cases are extracted from relevant documents and directly integrated into answers to generate a detailed, authoritative and grammatically correct answer information. The answer information not only contains the direct answer to the input specific question, but also contains the background, evidence documents and context-related reference documents related to the input specific question, making the answer more persuasive and credible.

[0082] [Rule 26 Corrected on 09.04.2025] Step S103: generating a specific enhanced sample based on the input specific question and the corresponding answer with reference documents.

[0083] [Rule 26 Corrected on 09.04.2025] Specifically, each specific enhanced sample is composed of an input specific question and an answer with reference documents generated by a preset language model. When a new specific question is encountered in a specific work, a specific enhanced sample containing the new specific question and the corresponding answer with reference documents can be obtained according to the above steps S101 to S103.

[0084] [Rule 26 Corrected on 09.04.2025] In one specific example, FIG. 3 shows a flowchart of generating a specific enhanced sample according to an embodiment of the present application. Referring to FIG. 3, the input specific question (i.e., the query question) is "A production and operation unit has more than 100 employees, but the unit does not set up a safety production management agency or employ full-time safety production management personnel, and does not carry out safety production education and training. In this case, does it violate the safety production law?"

[0085] [Rule 26 Corrected on 09.04.2025] According to the input specific question, the first 8 specific question-answer data pairs related to the input specific question (not shown) are found from the pre-set knowledge base, corresponding related documents and pre-set guide texts in each related document. Among them, the corresponding related documents are document [1], document [2], document [3], document [4], document [5], document [6], document [7] and document [8] respectively.

[0086] [Rule 26 Corrected on 09.04.2025] The pre-set guide text corresponding to document [1] (i.e., the answer corresponding to the specific question of document [1]) is: Article 24 of the Safety Production Law (2021) Mines, metal smelting, construction, transportation units and production, operation, storage, loading units of dangerous goods shall set up safety production management agencies or employ full-time safety production management personnel. Other production and operation units other than those specified in the preceding paragraph shall set up safety production management agencies or employ full-time safety production management personnel if the number of employees exceeds 100; if the number of employees is less than 100, full-time or part-time safety production management personnel shall be employed.

[0087] [Rule 26 Corrected on 09.04.2025] The pre-set guide text corresponding to document [2] is: Article 107 of the Safety Production Law (2021) The employees of a production and operation unit shall not be assigned to their post safety responsibilities, shall not be subject to management, shall not violate safety production rules and regulations or operating procedures, and shall be subject to criticism and education by the production and operation unit, and shall be subject to disciplinary action in accordance with relevant rules and regulations; if a crime is committed, criminal responsibility shall be investigated in accordance with relevant provisions of the Criminal Law.

[0088] [Rule 26 Corrected on 09.04.2025] The pre-set guide text corresponding to document [3] is: Article 25 of the Safety Production Law (2021) The safety production management agencies and safety production management personnel of a production and operation unit shall perform the following duties:

[0089] [Rule 26 Corrected on 09.04.2025] (1) Organize or participate in the formulation of safety production rules and regulations, operating procedures and production safety accident emergency rescue plans of the unit;

[0090] [According to Rule 26, correct 09.04.2025] (ii) Organize or participate in the safety production education and training of the unit, and record the safety production education and training truthfully;

[0091] [According to Rule 26, correct 09.04.2025] (iii) Organize the identification and assessment of dangerous sources, and supervise the implementation of safety management measures for major dangerous sources in the unit;

[0092] [According to Rule 26, correct 09.04.2025] (iv) Organize or participate in the unit's emergency rescue drill;

[0093] [According to Rule 26, correct 09.04.2025] (v) Check the safety production situation of the unit, timely investigate the hidden dangers of production safety accidents, and put forward suggestions for improving safety production management;

[0094] [According to Rule 26, correct 09.04.2025] (vi) Stop and correct the behaviors of violating orders, forcibly ordering dangerous operations, and violating operation procedures;

[0095] [According to Rule 26, correct 09.04.2025] (vii) Supervise the implementation of safety production rectification measures in the unit.

[0096] [According to Rule 26, correct 09.04.2025] Production and operation units may set up full-time safety production responsible persons to assist the main responsible person of the unit in performing safety production management duties.

[0097] [According to Rule 26, correct 09.04.2025] The preset guide text corresponding to Document [4] is: Article 20 of the Safety Production Law (2021) Production and operation units shall have the safety production conditions prescribed in this Law and relevant laws, administrative regulations or industry standards; those without safety production conditions shall not engage in production and operation activities.

[0098] [Corrected according to Rule 26 09.04.2025] The corresponding preset guide text of document [5] is: According to the provisions of Article 83 of the Safety Production Law of the People's Republic of China, if a production and business operation unit has any of the following behaviors and causes serious consequences, it shall be investigated for criminal responsibility in accordance with the relevant provisions of the Criminal Law: (1) The construction project of a mine or the construction project for the production and storage of dangerous goods is not designed with safety facilities or the safety facility design is not submitted to the relevant department for examination and approval as required; (2) The construction unit of a mine construction project or a construction project for the production and storage of dangerous goods is not constructed in accordance with the approved safety facility design; (3) The safety facilities of a mine construction project or a construction project for the production and storage of dangerous goods have not been inspected and passed before the project is completed and put into production or use; (4) No obvious safety warning signs are installed on production and business operation sites and relevant facilities and equipment with relatively large dangerous factors; (5) The installation, use, detection, modification and scrapping of safety equipment do not meet the industry standards; (6) No regular maintenance, maintenance and periodic detection are provided for safety equipment; (7) No industry-standard labor protection articles are provided for employees; (8) Special equipment and containers, transportation tools for dangerous goods have not been detected, inspected and qualified by a professional institution, obtained a safety use certificate or safety mark, and put into use; (9) The use of dangerous process, equipment that is ordered to be eliminated or prohibited for production safety is used. Of course, if there is such behavior but no serious consequences, it does not constitute a crime. The corresponding punishment is to order to correct within a certain period of time, and if the correction is not made within the period of time, to order to stop construction or to stop production and rectification, and can also be imposed with a fine of not more than 50,000 yuan.

[0099] [Corrected according to Rule 26 09.04.2025] The corresponding preset guide text of document [6] is: Article 22 of the Safety Production Law (2021) The whole personnel safety production responsibility system of a production and business operation unit shall clearly define the responsible personnel, responsibility scope and assessment standards of each post.

[0100] [Corrected according to Rule 26 09.04.2025] A production and business operation unit shall establish a corresponding mechanism to strengthen the supervision and examination of the implementation of the whole personnel safety production responsibility system and ensure the implementation of the whole personnel safety production responsibility system.

[0101] [Corrected according to Rule 26 09.04.2025] The corresponding preset guide text of document [7] is: Article 81 of the Safety Production Law (2021) A production and business operation unit shall formulate an emergency rescue plan for production safety accidents and link up with the production safety accident emergency rescue plan formulated by the organizer, and organize drills regularly.

[0102] [Corrected according to Rule 26 09.04.2025] The corresponding preset guidance text of document [8] is: Article 96 of the Safety Production Law (2021) Other responsible persons and safety production management personnel of production and operation units shall not perform the safety production management responsibilities prescribed by this Law. They shall be ordered to correct within a certain period, and fined between 10,000 yuan and 30,000 yuan; if they cause a production safety accident, their qualifications related to safety production shall be suspended or revoked, and they shall be fined between 20% and 50% of their annual income of the previous year; if they commit a crime, they shall be investigated for criminal responsibility in accordance with relevant provisions of the Criminal Law.

[0103] [Corrected according to Rule 26 09.04.2025] After obtaining the 8 relevant documents corresponding to the specific question related to the input specific question and the preset guidance text in each relevant document, the input specific question, 8 relevant documents and the preset guidance text in each relevant document are input into the ChatGPT model, the ChatGPT model understands and analyzes the input specific question, 8 relevant documents and the preset guidance text in each relevant document, and summarizes the results after understanding and analysis, and generates an answer containing cited documents: "Yes, the production and operation unit violates the Safety Production Law. According to Article 24 of the Safety Production Law, production and operation units with more than 100 employees shall set up a safety production management institution or appoint full-time safety production management personnel[1]. Therefore, the unit has violated the law by not setting up a safety production management institution or appointing full-time safety production management personnel.\n\nIn addition, according to Article 25 of the Safety Production Law, the safety production management institution or safety production management personnel of the production and operation unit shall perform the duties of organizing or participating in the safety production education and training of the unit[3]. The unit has also violated this legal provision by not conducting safety production education and training. If the unit causes a production safety accident due to violation of the Safety Production Law, according to Article 96 of the Safety Production Law, other responsible persons and safety production management personnel may be punished[8]. In some serious cases, it may also constitute a crime, and shall be investigated for criminal responsibility in accordance with the Criminal Law[5]."

[0104] [Corrected according to Rule 26 09.04.2025] As can be seen, the ChatGPT model obtains the answer information containing cited documents by understanding and analyzing the input specific question and the preset guidance text in documents [1] to [8], which not only gives a correct answer to the question of whether the input specific question is illegal, but also quotes the contents in documents [1], [3], [5] and [8] as the judgment basis and evidence documents for decision-making, increasing the credibility and authority of the answer to the decision-making.

[0105] [According to Rule 26, correct on 09.04.2025] Specifically, when each input specific question and its corresponding answer information containing reference documents are generated, they are considered as specific enhanced samples. As these specific enhanced samples accumulate, when a certain number of specific enhanced samples are accumulated, a sufficient number of specific enhanced samples can be used as the first question and answer data set.

[0106] [According to Rule 26, correct on 09.04.2025] Step S105: The specific question in the specific enhanced sample and the corresponding answer with the reference document are respectively taken as the model input sample and the model output sample, and the basic question and answer model is enhanced training to obtain the target question and answer model; the basic question and answer model is trained based on the second question and answer data set, and the second question and answer data set includes multiple questions and answers corresponding to each question.

[0107] [According to Rule 26, correct on 09.04.2025] Specifically, the basic question and answer model has universality, which is trained by the second question and answer data set collected from multiple sources, which includes questions and corresponding answers, wherein the questions in the second question and answer data set include but are not limited to daily life, technology, culture, history and specific fields. In one specific example, the questions and corresponding answers in non-specific fields such as daily life, technology, culture, history, etc. are taken as general answer data pairs, and the questions and corresponding answers in specific fields are taken as specific question and answer data pairs. When training the basic question and answer model using the second question and answer data set, the training data contained in the second question and answer data set includes general answer data pairs and specific question and answer data pairs, so the basic question and answer model has general question and answer ability and specific question and answer ability.

[0108] [According to Rule 26, correct on 09.04.2025] After training the basic question and answer model according to the questions and corresponding answers in various fields, since the basic question and answer model only has general question and answer ability and specific question and answer ability, it does not have the ability to generate answers with reference documents corresponding to specific field questions, so after inputting specific questions into the basic question and answer model, only answers corresponding to specific questions can be obtained, and it is not possible to generate answers with reference documents that meet specific requirements and have authority for specific questions. Therefore, after obtaining the first question and answer data set constructed by specific enhanced samples, the specific enhanced samples in the first question and answer data set can be used to further enhance the training of the basic question and answer model, that is, the specific question model input sample in the specific enhanced sample, the answer with reference documents corresponding to the specific question in the specific enhanced sample is taken as the model output sample, and the basic question and answer model is further enhanced training, so that the enhanced training of the basic question and answer model (i.e. the target question and answer model) has the ability to generate answers with reference documents.

[0109] [Amended according to Rule 26 09.04.2025] In some embodiments, obtaining the first question and answer data set comprises:

[0110] [Amended according to Rule 26 09.04.2025] Step S1041: obtaining a corresponding number of specific enhanced samples at a preset ratio based on the number of questions contained in the second question and answer data set, to obtain the first question and answer data set, wherein the number of specific enhanced samples contained in the first question and answer data set is less than the number of questions contained in the second question and answer data set.

[0111] [Amended according to Rule 26 09.04.2025] In one specific example, the question in a specific field and the corresponding answer with reference documents are taken as specific enhanced question and answer data pairs, the above-mentioned general answer data pairs, specific question and answer data pairs and specific enhanced question and answer data pairs are taken as three different kinds of sample data, the sample data corresponding to the number of general answer data pairs and specific question and answer data pairs are taken as the second question and answer data set, and the basic question and answer model is trained using the second question and answer data set, so that the basic question and answer model has general question and answer and specific question and answer capabilities. Then, a corresponding number of specific enhanced question and answer data pairs are obtained at a preset ratio, and the specific enhanced samples corresponding to the number of specific enhanced question and answer data pairs are taken as the first question and answer data set, and the basic question and answer model is further enhanced using the first question and answer data set, so that the enhanced basic question and answer model has the ability to generate more authoritative answer information with reference documents for specific questions.

[0112] [Amended according to Rule 26 09.04.2025] In an exemplary manner, the preset ratio of the number of general answer data pairs, the number of specific question and answer data pairs and the number of specific enhanced question and answer data pairs is set to 2:1:1, that is, the preset ratio of the number of sample data in the second question and answer data set to the number of sample data in the first question and answer data set is set to 3:1, and the corresponding number of specific enhanced samples are obtained at a preset ratio of 3:1, so as to ensure that the model does not pay excessive attention to the characteristics of the specific enhanced samples while being exposed to a sufficient number of specific enhanced samples to enable the model to have the ability to generate more authoritative answer information with reference documents for specific questions, so as to maintain the original general question and answer capabilities and specific question and answer capabilities of the model.

[0113] [Amended according to Rule 26 09.04.2025] In some embodiments, FIG. 4 shows a flowchart of the sub-steps of step S105 provided by an embodiment of the present application. Referring to FIG. 4, the specific question in the specific enhanced sample and the corresponding answer with reference documents are taken as model input samples and model output samples respectively, and the basic question and answer model is enhanced and trained, including steps S1051 to S1053.

[0114] [Amended according to Rule 26 on 09.04.2025] Step S1051: input the specific question in the specific enhanced sample into the basic question and answer model to obtain an initial predicted answer output by the basic question and answer model.

[0115] [Amended according to Rule 26 on 09.04.2025] Specifically, after obtaining a corresponding number of specific enhanced samples, the specific question in each specific enhanced sample is input into the basic question and answer model trained using the second question and answer data set. The basic question and answer model uses the current model parameters obtained by training to encode the input specific question and convert it into a numerical form that the model can understand. Then, the model decodes or reasons the encoded result to generate an initial predicted answer for the input specific question. At this time, since the basic question and answer model may not be adequately trained in the specific field, the initial predicted answer may not be completely accurate or comprehensive.

[0116] [Amended according to Rule 26 on 09.04.2025] Step S1052: calculate the loss between the initial predicted answer and the corresponding model output sample.

[0117] [Amended according to Rule 26 on 09.04.2025] Specifically, in order to evaluate the gap between the initial predicted answer generated by the model and the answer provided in the specific enhanced sample with the reference document (i.e. the true answer), the initial predicted answer is compared with the answer provided in the specific enhanced sample with the reference document, and the loss between the initial predicted answer and the answer provided in the specific enhanced sample with the reference document is calculated, for example, the cross-entropy loss between the initial predicted answer and the answer provided in the specific enhanced sample with the reference document is calculated to obtain the difference between the predicted result of the model and the true answer.

[0118] [Amended according to Rule 26 on 09.04.2025] Step S1053: iteratively update the model parameters of the basic question and answer model according to the loss between the initial predicted answer and the corresponding model output sample until the basic question and answer model meets the preset second training stopping condition.

[0119] [Rule 26 Correction 09.04.2025] Specifically, according to the loss value between the calculated initial predicted answer and the corresponding model output sample, the loss value is propagated back to each layer of the basic question and answer model through the back propagation algorithm, so as to calculate the gradient of each layer parameter. Then, using an optimization algorithm (such as Adam, SGD, etc.) to update the parameters of the basic question and answer model according to these gradients. The optimization algorithm adjusts the values of the parameters to minimize the loss between the initial predicted answer and the corresponding model output sample, and the process of steps S1051 to S1053 is repeated several times until the basic question and answer model meets the preset second training stopping condition. Among them, the second training stopping condition includes: when the change of the loss between the initial predicted answer and the corresponding model output sample is less than a certain threshold, that is, the basic question and answer model has converged, the training stops; or when the basic question and answer model is enhanced training, if the training round reaches the preset maximum training round, the training stops.

[0120] [Rule 26 Correction 09.04.2025] In some embodiments, Figure 5 shows a flowchart of another question and answer model training method provided by an embodiment of the present application. Referring to Figure 5, after step S105, the question and answer model training method further includes steps S106 and S107.

[0121] [Rule 26 Correction 09.04.2025] Step S106: Obtain a preference dataset; the preference dataset includes positive sample data and negative sample data; the positive sample data includes a specific question sample and a corresponding reference answer, and the reference answer has a reference document; the negative sample data includes a specific question sample and an answer output by the target question and answer model based on the specific question sample.

[0122] [Rule 26 Correction 09.04.2025] Specifically, after further enhancing the training of the basic question and answer model with the first question and answer dataset, so that the target question and answer model has the ability to generate more authoritative answer information with reference documents for specific questions, the target question and answer model still has a high probability of generating incorrect answers. Therefore, after obtaining the target question and answer model, a preference dataset is obtained to optimize the target question and answer model using the preference dataset, so that the accuracy of the optimized target question and answer model in generating answer information with reference documents is higher.

[0123] [Rule 26 Correction 09.04.2025] The preference dataset includes positive sample data and negative sample data, the positive sample data contains a specific question sample and a corresponding reference answer, and the reference answer has a reference document; the negative sample data contains a specific question sample and an answer output by the target question and answer model based on the question sample, and the answer is likely to be an incorrect answer.

[0124] [Amended according to Rule 26 on September 4, 2025] Step S107: based on the preference dataset, the target question and answer model is trained for preference alignment optimization to obtain an optimized target question and answer model.

[0125] [Amended according to Rule 26 on September 4, 2025] Wherein, preference alignment optimization is a method widely used in large-scale language models, aiming to make the output of the model more in line with human preferences. Specifically, preference alignment optimization adjusts and optimizes the parameters and behaviors of the model through a series of technical means and strategies to generate more accurate, reasonable, useful and human-expected answers. Methods of preference alignment optimization include human preference optimization algorithm (Reinforcement Learning from Human Feedback, RLHF), direct preference optimization (Direct Preference Optimization, DPO), identity preference optimization (Identity Preference Optimisation, IPO) and odds ratio preference optimization (Odds Ratio Preference Optimization, ORPO) methods. The embodiments of the present application can select appropriate methods of preference alignment optimization according to actual conditions.

[0126] [Amended according to Rule 26 on September 4, 2025] In one specific example, since the ORPO optimization method has the characteristics of lower cost than the RLHF optimization method, and has the characteristics of simpler operation than the DPO optimization method and the IPO optimization method, and does not need to train an additional preference model, the embodiments of the present application can select the ORPO optimization method to perform preference alignment optimization on the target question and answer model. Specifically, the ORPO optimization method combines instruction fine-tuning and preference alignment into a single process and introduces OR loss in the loss function to achieve preference alignment, thereby avoiding training an additional preference model and simplifying the operation process.

[0127] [Amended according to Rule 26 on September 4, 2025] In some embodiments, Figure 6 shows a flowchart of the sub-steps of step S107 provided by an embodiment of the present application. Referring to Figure 6, step S107 further includes steps S1071 to S1074.

[0128] [Amended according to Rule 26 on September 4, 2025] Step S1071: inputting a specific question sample in the positive sample data into the current target question and answer model to obtain a predicted result corresponding to the positive sample data, and calculating a first loss between the predicted result corresponding to the positive sample data and the reference answer.

[0129] [Amended according to Rule 26 on 09.04.2025] In a specific example, when the ORPO optimization method is selected to perform preference alignment optimization on the target question answering model, the difference between the model prediction result and the reference answer is calculated by inputting the positive sample data (including the specific question sample and its corresponding reference answer).

[0130] [Amended according to Rule 26 on 09.04.2025] Specifically, the specific question sample in each positive sample data is input into the current target question answering model trained by the first question answering data set, and the current target question answering model uses the current trained model parameters to encode the input specific question sample in the positive sample data, and converts it into a numerical form that the current target question answering model can understand. Then, the current target question answering model decodes or infers the encoded result to generate a corresponding prediction result for the input specific question sample in the positive sample data. At this time, since the current target question answering model may not completely cover all cases during training, the prediction result generated for the input specific question sample in the positive sample data may not be completely accurate or comprehensive. In order to evaluate the gap between the prediction result generated by the current target question answering model for the input specific question sample in the positive sample data and the corresponding reference answer, the prediction result generated for the input specific question sample in the positive sample data is compared with the corresponding reference answer, and the loss between the prediction result generated for the input specific question sample in the positive sample data and the corresponding reference answer is calculated, such as the cross-entropy loss between the prediction result generated for the input specific question sample in the positive sample data and the corresponding reference answer, to obtain the difference between the prediction result generated by the current target question answering model for the input specific question sample in the positive sample data and the corresponding reference answer, that is, the first loss

[0131] [Amended according to Rule 26 on 09.04.2025] Step S1072: input the specific question sample in the negative sample data into the current target question answering model to obtain the prediction result corresponding to the negative sample data, and calculate the second loss between the prediction result corresponding to the negative sample data and the corresponding answer in the negative sample data.

[0132] [Amended according to Rule 26 on 09.04.2025] In a specific example, when the ORPO optimization method is selected to perform preference alignment optimization on the target question answering model, the OR loss (Odds Ratio Loss, OR) is also calculated by inputting the negative sample data (including the specific question sample and the answer output by the target question answering model based on the specific question sample), and the OR loss is used to weakly punish the unpopular generated content and strongly reward the popular generated content.

[0133] [According to Rule 26, correct on 09.04.2025] Specifically, input each specific question sample in the negative sample data into the current target question answering model that has been trained by the first question answering data set, and the current target question answering model uses the current training model parameters to encode the input specific question sample in the negative sample data, and converts it into a numerical form that the current target question answering model can understand, and then the current target question answering model decodes or reasons the encoded result, and generates a corresponding prediction result for the input specific question sample in the negative sample data. Then, calculate the OR loss between the prediction result generated for the input specific question sample in the negative sample data and the answer provided in the negative sample data, i.e. the second loss

[0134] [According to Rule 26, correct on 09.04.2025] In one specific example, calculating the second loss between the prediction result corresponding to the negative sample data and the corresponding answer in the negative sample data includes steps S10721 to S10724.

[0135] [According to Rule 26, correct on 09.04.2025] Step S10721: For a specific question sequence x in the input negative sample data, calculate the average log probability logP θ (y|x) of the output sequence y (containing m tokens) generated by the current target question answering model for sequence x:

[0136] [According to Rule 26, correct on 09.04.2025] Wherein, θ represents the parameters of the current target question answering model, y t represents the t-th token in the output sequence y, y <t represents all tokens in y located before y t ; logP θ (y|x) is used to measure the probability of the current target question answering model generating a given sequence y.

[0137] [According to Rule 26, correct on 09.04.2025] Step S10722: Calculate the ratio of the conditional probability of the current target question answering model generating a given sequence y to the conditional probability of not generating a given sequence y (i.e. generating any other sequence) oddS θ (y|x):

[0138] [According to Rule 26, correct on 09.04.2025] In this embodiment, oddS θ (y|x) is the ratio between the probability of the current target question answering model generating the target sequence y and the probability of generating all other sequences; oddS θ(y|x) reflects the degree of preference of the current target question answering model for the generated target sequence y.

[0139] [Revised according to Rule 26, 09.04.2025] Step S10723: Calculate the positive sample sequence y generated by the current target question answering model. w and generate negative sample sequence y l The ratio of OR θ (y w y l )for:

[0140] [Revised according to Rule 26, 09.04.2025] Wherein, oddS θ (y w |x) represents the sequence of positive samples y generated by the current target question-answering model given the model parameters θ and the input sequence x. w The ratio of odds; θ (y l |x) represents the negative sample sequence y generated by the current target question answering model given model parameters θ and input sequence x. l The ratio; OR θ (y w ,y l It measures the difference in preference between positive and negative samples in the current target question answering model.

[0141] [Revised according to Rule 26 09.04.2025] Step S10724: Define the second loss for::

[0142] [Revised according to Rule 26 09.04.2025] Where σ represents the sigmoid function.

[0143] [Revised according to Rule 26 09.04.2025] Step S1073: Perform a weighted summation of the first loss and the second loss to obtain the target loss, wherein the weight of the second loss is less than the weight of the first loss.

[0144] [Revised according to Rule 26, 09.04.2025] In a specific example, the target loss L is defined. ORPO for:

[0145] [Revised according to Detailed Rule 26, 09.04.2025] Among them, The first loss, For the second loss, λ is a hyperparameter used to balance the first and second losses. For (x,y) w ,yl ) on all possible data points to calculate L ORPO the average or expectation of L

[0146] [According to Rule 26 Correction 09.04.2025] In this embodiment, since the target loss of the ORPO optimization method contains both the first loss and the second loss, after obtaining the first loss and the second loss , in order to balance the first loss and the second loss , the weighted sum of the first loss and the second loss can obtain the target loss L ORPO In this embodiment, the weight of the second loss is set to be less than the weight of the first loss. For example, the weight of the first loss is set to 1, and the weight of the second loss is set to 0.1.

[0147] [According to Rule 26 Correction 09.04.2025] It should be noted that setting the weight of the first loss to 1 and the weight of the second loss to 0.1 is only a specific example, and the weights of the first loss and the second loss can be adjusted according to actual conditions.

[0148] [According to Rule 26 Correction 09.04.2025] Step S1074: iteratively updating the model parameters of the current target question answering model according to the target loss until the target question answering model meets the preset first training stopping condition.

[0149] [According to Rule 26 Correction 09.04.2025] Specifically, according to the calculated target loss, the target loss is backpropagated to each layer of the current target question answering model through the backpropagation algorithm, so as to calculate the gradient of each layer parameter. Then, using an optimization algorithm (such as Adam, SGD, etc.) to update the parameters of the current target question answering model according to these gradients. The optimization algorithm will adjust the values of the parameters to minimize the target loss, and the process of steps S1071 to S1074 will be repeated several times until the current target question answering model meets the preset first training stopping condition. The first training stopping condition includes: when the change of the target loss is less than a certain threshold, that is, the current target question answering model has converged, the training stops; or when the target question answering model is performing preference alignment optimization training, if the training round reaches the preset maximum training round, the training stops.

[0150] [Rule 26 amended on 09.04.2025] In a specific example, as shown in FIG. 7, in the ORPO optimization method, two samples are used as input data, respectively denoted as positive sample data and negative sample data, wherein the two sample data correspond to the same specific problem, and the answers to the specific problem are different, the answer of the positive sample data corresponds to the expected answer, and the answer of the negative sample data corresponds to the answer not expected. In the ORPO optimization process, the goal is to adjust the target loss of the target question and answer model so that the target question and answer model is more inclined to the answer corresponding to the positive sample data when outputting the predicted answer. Among them, the answer corresponding to the positive sample data can come from the reference answer correctly labeled for the specific problem; or, it comes from the correction of the answer generated by the target question and answer model using a language model such as ChatGPT model, and the corrected answer is used as the answer corresponding to the positive sample data. The answer corresponding to the negative sample data can come from the answer generated by the target question and answer model.

[0151] [Rule 26 amended on 09.04.2025] FIG. 8 shows a flowchart of a question and answer method provided by an embodiment of the application. Referring to FIG. 8, the question and answer method includes steps S201 to S202.

[0152] [Rule 26 amended on 09.04.2025] Step S201: Obtain a target question.

[0153] [Rule 26 amended on 09.04.2025] When a new specific question is generated in a specific work, the generated new specific question is used as the target specific question. Specifically, the new specific question can be obtained from news media, social media, etc. in a timely manner by monitoring news media, social media, etc.; or, the new specific question can be obtained in a timely manner through regular social surveys, public opinion surveys, and public security situation assessments.

[0154] [Rule 26 amended on 09.04.2025] Step S202: input the target question into the target question and answer model to obtain the answer with reference documents corresponding to the target question, and the target question and answer model is trained by the above question and answer model training method.

[0155] [Rule 26 amended on 09.04.2025] The question and answer method provided by the embodiments of the application can generate answers with reference documents corresponding to the target question by using the target question and answer model when a specific person is handling the target question, and can make more wise and accurate decisions based on the answers with reference documents generated by the target question and answer model, to improve the credibility and persuasiveness of the decisions made by the specific person when handling the target question.

[0156] [Rule 26 Correction 09.04.2025] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0157] [Rule 26 Correction 09.04.2025] Corresponding to the specific question and answer model creation method described in the above embodiments, FIG. 9 is a structural block diagram of a question and answer model training device provided by the embodiments of the present application, only the parts related to the embodiments of the present application are shown for the convenience of description.

[0158] [Rule 26 Correction 09.04.2025] Referring to FIG. 9, the question and answer model training device comprises:

[0159] [Rule 26 Correction 09.04.2025] The first acquisition module 91 is configured to acquire a first question and answer data set; the first question and answer data set comprises two or more specific enhanced samples, and each specific enhanced sample comprises a specific question and an answer with a reference document corresponding to the specific question.

[0160] [Rule 26 Correction 09.04.2025] The enhanced training module 92 is configured to take the specific question in the specific enhanced sample and the corresponding answer with the reference document as a model input sample and a model output sample respectively, and perform enhanced training on a basic question and answer model to obtain a target question and answer model; the basic question and answer model is trained based on a second question and answer data set, and the second question and answer data set comprises a plurality of questions and an answer corresponding to each question.

[0161] [Rule 26 Correction 09.04.2025] In some embodiments, the question and answer model training device further comprises:

[0162] [Rule 26 Correction 09.04.2025] The search module 93 is configured to search at least one document related to the specific question from a preset knowledge base for an input specific question.

[0163] [Rule 26 Correction 09.04.2025] The input and output module 94 is configured to input the specific question, the searched document related to the specific question, and a preset guide text into a preset language model to obtain an answer with a reference document corresponding to the specific question, and the guide text is used to guide the result output mode of the preset language model.

[0164] [Rule 26 Correction 09.04.2025] The specific enhanced sample generation module 95 is configured to generate the specific enhanced sample based on the specific question and the corresponding answer with the reference document.

[0165] [Amended according to Rule 26 09.04.2025] In some embodiments, the question and answer model training apparatus further comprises:

[0166] [Amended according to Rule 26 09.04.2025] a preference dataset obtaining module 96, configured to obtain a preference dataset; the preference dataset comprises positive sample data and negative sample data; the positive sample data comprises a specific question sample and a corresponding reference answer, the reference answer being associated with a reference document; the negative sample data comprises the specific question sample and an answer output by the target question and answer model based on the specific question sample.

[0167] [Amended according to Rule 26 09.04.2025] a preference alignment optimization training module 97, configured to perform preference alignment optimization training on the target question and answer model based on the preference dataset, to obtain an optimized target question and answer model.

[0168] [Amended according to Rule 26 09.04.2025] In some embodiments, the preference alignment optimization training module 97 is further configured to:

[0169] [Amended according to Rule 26 09.04.2025] input the specific question sample in the positive sample data into the current target question and answer model to obtain a predicted result corresponding to the positive sample data, and calculate a first loss between the predicted result corresponding to the positive sample data and the reference answer.

[0170] [Amended according to Rule 26 09.04.2025] input the specific question sample in the negative sample data into the current target question and answer model to obtain a predicted result corresponding to the negative sample data, and calculate a second loss between the predicted result corresponding to the negative sample data and the corresponding answer in the negative sample data.

[0171] [Amended according to Rule 26 09.04.2025] perform weighted summation on the first loss and the second loss to obtain a target loss, wherein the weight of the second loss is less than the weight of the first loss.

[0172] [Amended according to Rule 26 09.04.2025] iteratively update the model parameters of the current target question and answer model according to the target loss until the target question and answer model satisfies a preset first training stop condition.

[0173] [Amended according to Rule 26 09.04.2025] In some embodiments, the first obtaining module 91 is further configured to:

[0174] [According to Rule 26, correct on 09.04.2025] Based on the number of questions contained in the second question and answer data set, a corresponding number of specific augmented samples are obtained at a preset ratio to obtain the first question and answer data set, wherein the number of specific augmented samples contained in the first question and answer data set is less than the number of questions contained in the second question and answer data set.

[0175] [According to Rule 26, correct on 09.04.2025] In some embodiments, the augmented training module 92 is further used to:

[0176] [According to Rule 26, correct on 09.04.2025] The specific question in the specific augmented sample is input into the basic question and answer model to obtain an initial predicted answer output by the basic question and answer model.

[0177] [According to Rule 26, correct on 09.04.2025] The loss between the initial predicted answer and the corresponding model output sample is calculated.

[0178] [According to Rule 26, correct on 09.04.2025] The model parameters of the basic question and answer model are iteratively updated according to the loss between the initial predicted answer and the corresponding model output sample until the basic question and answer model satisfies a preset second training stopping condition.

[0179] [According to Rule 26, correct on 09.04.2025] Corresponding to the question and answer method described in the above embodiments, FIG. 10 is a structural block diagram of a question and answer device provided by the embodiments of the present application. For ease of illustration, only parts related to the embodiments of the present application are shown.

[0180] [According to Rule 26, correct on 09.04.2025] Referring to FIG. 10, the question and answer device comprises:

[0181] [According to Rule 26, correct on 09.04.2025] The second acquisition module 101 is configured to acquire a target question.

[0182] [According to Rule 26, correct on 09.04.2025] The generation module 102 is configured to input the target question into a target question and answer model to obtain an answer with a reference document corresponding to the target question, wherein the target question and answer model is trained by using the above question and answer model training method.

[0183] [According to Rule 26, correct on 09.04.2025] It should be noted that the information interaction, execution process, etc. between the above devices / units, since based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought by them can be referred to the method embodiments part, which will not be repeated here.

[0184] [Amended according to Rule 26 on 09.04.2025] In addition, the question and answer model training apparatus shown in FIG. 9 / the question and answer apparatus shown in FIG. 10 can be a software unit, a hardware unit, or a combination of software and hardware built into an existing terminal device, can be integrated into the terminal device as an independent plug-in, or can exist as an independent terminal device.

[0185] [Amended according to Rule 26 on 09.04.2025] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the above division of functional units and modules is taken as an example, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the above-described functions. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual differentiation, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0186] [Amended according to Rule 26 on 09.04.2025] FIG. 11 is a structural schematic diagram of a terminal device provided by an embodiment of the present application. As shown in FIG. 11, the terminal device 11 of this embodiment includes at least one processor 110 (only one processor is shown in FIG. 11), a memory 111, and a computer program 112 stored in the memory 111 and executable on the at least one processor 110, and the processor 110 implements the steps in any of the foregoing question and answer model training methods / question and answer methods.

[0187] [Amended according to Rule 26 on 09.04.2025] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The terminal device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that FIG. 11 is only an example of the terminal device 11, and does not limit the terminal device 11, and can include more or fewer components than shown, or combine certain components, or different components, for example, can also include an input / output device, a network access device, and the like.

[0188] [According to Rule 26, the processor 110 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.

[0189] [According to Rule 26, the memory 111 can be an internal storage unit of the terminal device 11 in some embodiments, such as a hard disk or a memory of the terminal device 11. The memory 111 can also be an external storage device of the terminal device 11 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 111 can include both an internal storage unit and an external storage device of the terminal device 11. The memory 111 is used to store an operating system, application programs, a boot loader, data, and other programs, such as program codes of the computer program, etc. The memory 111 can also be used to temporarily store data that has been output or will be output.

[0190] [According to Rule 26, the computer readable storage medium of the embodiments of the present application stores a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments.

[0191] [According to Rule 26, the computer program product of the embodiments of the present application is used to run on a terminal device, so that the terminal device is enabled to implement the steps in each of the method embodiments when executed.

[0192] [Rule 26 Correction 09.04.2025] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0193] [Rule 26 Correction 09.04.2025] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can refer to the relevant description of other embodiments.

[0194] [Rule 26 Correction 09.04.2025] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0195] [According to Rule 26 Correction 09.04.2025] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other manners. For example, the apparatus / terminal device embodiments described above are merely schematic; for example, the division of the modules or units is merely logical function division; there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0196] [According to Rule 26 Correction 09.04.2025] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0197] [According to Rule 26 Correction 09.04.2025] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. [Amended according to Rule 26 09.04.2025] A method for training a question and answer model, characterized in that, The method comprises the following steps: obtaining a first question and answer data set; the first question and answer data set comprises two or more specific enhanced samples, and each specific enhanced sample comprises a specific question and an answer corresponding to the specific question with a reference document; performing enhanced training on a basic question and answer model based on the specific question and the corresponding answer with the reference document in the specific enhanced sample to obtain a target question and answer model; the basic question and answer model is trained based on a second question and answer data set, and the second question and answer data set comprises a plurality of questions and an answer corresponding to each question.

2. [Amended according to Rule 26 09.04.2025] The question and answer model training method according to claim 1, characterized in that, Before obtaining the first question and answer data set, the method further comprises the following steps: searching, from a preset knowledge base, at least one document related to a specific question input by a user; inputting the specific question, the searched document related to the specific question, and a preset guide text into a preset language model to obtain an answer corresponding to the specific question with a reference document; the guide text is used to guide the output mode of the preset language model; generating the specific enhanced sample based on the specific question and the corresponding answer with the reference document.

3. The method of claim 1 or 2, wherein the method further comprises: After obtaining the target question and answer model, the method further comprises the following steps: obtaining a preference data set; the preference data set comprises positive sample data and negative sample data; the positive sample data comprises a specific question sample and a reference answer corresponding to the specific question sample, and the reference answer has a reference document; the negative sample data comprises the specific question sample and an answer output by the target question and answer model based on the specific question sample; performing preference alignment optimization training on the target question and answer model based on the preference data set to obtain an optimized target question and answer model.

4. The method of claim 3, wherein the method further comprises: The preference alignment optimization training on the target question and answer model based on the preference data set comprises the following steps: inputting the specific question sample in the positive sample data into the current target question and answer model to obtain a predicted result corresponding to the positive sample data, calculating a first loss between the predicted result corresponding to the positive sample data and the reference answer; inputting the specific question sample in the negative sample data into the current target question and answer model to obtain a predicted result corresponding to the negative sample data, calculating a second loss between the predicted result corresponding to the negative sample data and the answer corresponding to the negative sample data in the negative sample data; performing weighted summation on the first loss and the second loss to obtain a target loss, wherein the weight of the second loss is less than the weight of the first loss; iteratively updating the model parameters of the current target question and answer model according to the target loss until the target question and answer model satisfies a preset first training stop condition.

5. [Amended according to Rule 26 09.04.2025] The question and answer model training method according to claim 1 or 2, characterized in that, The method of obtaining the first question and answer data set comprises the following steps: obtaining a corresponding number of specific enhanced samples at a preset ratio based on the number of questions contained in the second question and answer data set to obtain the first question and answer data set, wherein the number of specific enhanced samples contained in the first question and answer data set is less than the number of questions contained in the second question and answer data set.

6. The method of claim 5, wherein the method further comprises: The specific question in the specific enhanced sample and the corresponding answer with the cited document are input as a model input sample and a model output sample, respectively, to perform enhanced training on the basic question and answer model, including: inputting the specific question in the specific enhanced sample into the basic question and answer model to obtain an initial predicted answer output by the basic question and answer model; calculating the loss between the initial predicted answer and the corresponding model output sample; iteratively updating the model parameters of the basic question and answer model according to the loss between the initial predicted answer and the corresponding model output sample until the basic question and answer model meets a preset second training stop condition.

7. [Amended according to Rule 26 09.04.2025] A question and answer method, characterized in that, including: obtaining a target question; inputting the target question into a target question and answer model to obtain an answer with a cited document corresponding to the target question, wherein the target question and answer model is trained by using the question and answer model training method in any one of claims 1-6.

8. [Amended according to Rule 26 09.04.2025] A question and answer model training apparatus, characterized by, including: a first obtaining module configured to obtain a first question and answer data set; the first question and answer data set includes two or more specific enhanced samples, and each specific enhanced sample includes a specific question and an answer with a cited document corresponding to the specific question; an enhanced training module configured to input the specific question in the specific enhanced sample and the corresponding answer with the cited document as a model input sample and a model output sample, respectively, to perform enhanced training on a basic question and answer model, and obtain a target question and answer model; the basic question and answer model is trained based on a second question and answer data set, and the second question and answer data set includes a plurality of questions and an answer corresponding to each question.

9. [Amended according to Rule 26 09.04.2025] A question and answer device characterized in that, including: a second obtaining module configured to obtain a target question; a generating module configured to input the target question into a target question and answer model to obtain an answer with a cited document corresponding to the target question, wherein the target question and answer model is trained by using the question and answer model training method in any one of claims 1-6.

10. [Amended according to Rule 26 09.04.2025] A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the question and answer model training method in any one of claims 1-6; or the processor executes the computer program to implement the question and answer method in claim 7.

Citation Information

Patent Citations

  • Retrieval type intelligent question and answer system and method for coal mine safety regulations

    CN114020862A

  • Generation method and device of intelligent question and answer model, computing equipment and storage medium

    CN114547267A

  • Question and answer model construction method, knowledge base creation method, question and answer search method and electronic equipment

    CN116842151A

  • Intelligent questioning and answering method for government affairs

    CN118153686A

  • Mineral knowledge question answering method and system based on large language model

    CN118193708A