Question and answer training method and system of dialogue model, electronic equipment and storage medium

Through the EZ-Train framework automated transformation and similarity evaluation algorithm, the Q&A pairs are combined with GPT, BERT and LLM model training, the problem of time-consuming and unstable quality in the existing technology is solved, and efficient, flexible, and multi-domain adaptable Q&A dialogue training is achieved.

CN120508611APending Publication Date: 2025-08-19GUANGZHOU HUADI CREATIVE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516666.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the prior art, the Q&A dialogue training method is time-consuming and labor-intensive, and the generation quality is uneven, the rule-based method lacks flexibility, and it is difficult to deal with complex dialogue scenarios and diverse user input.

Method used

The EZ-Train framework is used to automatically convert unstructured text into structured Q&A pairs, combine the similarity evaluation algorithm to process Q&A pairs, and use GPT, BERT and LLM models for training to support text interaction, voice interaction and digital human applications.

Benefits of technology

It improves the efficiency and quality of Q&A training, reduces redundancy, enhances the adaptability and flexibility of the dialogue model, reduces training costs, and is suitable for digital human Q&A dialogue training in multiple fields and topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508611A_ABST
    Figure CN120508611A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer training method and system for a dialogue model, electronic equipment and a storage medium. The method comprises the steps of obtaining target data input by a target object; the target data comprises at least one of a target file and a first question and answer pair; when the target file exists in the target data, converting the target file into a question and answer pair data set; the question-answer pair data set comprises a plurality of second question-answer pairs; when the first question-answer pair still exists in the target data, combining the first question-answer pair and a second question-answer pair in the question-answer pair data set based on similarity evaluation to obtain a question-answer pair training set; and inputting the question and answer pair training set into the target model for training, and generating a dialogue model. The question-answer training quality and efficiency of the dialogue model can be effectively improved, and the method can be widely applied to the technical field of question-answer training of the dialogue model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of question-answering training for dialogue models, and in particular to a method, system, electronic device, and storage medium for question-answering training for dialogue models. Background Art

[0002] In recent years, with the rapid development of artificial intelligence (AI) technology, digital humans, as an important vehicle for virtual interaction, have been widely used in numerous fields. One of their core functions is to enable natural and fluent conversations between digital humans and users, and question-and-answer dialogue training is a key step in achieving this. Traditional question-and-answer dialogue training methods rely primarily on manually composing question-and-answer pairs. This approach is not only time-consuming and labor-intensive, but also produces inconsistent quality, making it difficult to meet the demands of large-scale dialogue training. Furthermore, while some rule-based methods can improve dialogue quality to a certain extent, they lack flexibility and adaptability, making them difficult to handle complex dialogue scenarios and diverse user input. Summary of the Invention

[0003] The main purpose of the embodiments of the present invention is to propose a question-answering training method, system, electronic device and storage medium for a dialogue model, in order to solve at least one problem of the prior art. The present invention can improve the training efficiency of question-answering training of the dialogue model.

[0004] To achieve the above objectives, an embodiment of the present invention provides a question-answering training method for a dialogue model, the method comprising:

[0005] Obtaining target information input by the target subject; the target information includes at least one of a target file and a first question-answer pair;

[0006] When a target file exists in the target data, the target file is converted into a question-answer pair dataset; the question-answer pair dataset includes a plurality of second question-answer pairs;

[0007] If the first question-answer pair still exists in the target data, the first question-answer pair and the second question-answer pair in the question-answer pair dataset are merged based on similarity evaluation to obtain a question-answer pair training set.

[0008] Input the question-answer pair training set into the target model for training to generate a dialogue model;

[0009] Among them, the target models include GPT, BERT and LLM, and the application forms of dialogue models include text interaction, voice interaction and digital humans.

[0010] In some embodiments, obtaining target information input by the target object includes at least one of the following steps:

[0011] In response to an upload instruction of a target object, obtaining a target file in a target format;

[0012] In response to a manual input instruction of the target object, a first question-answer pair is obtained.

[0013] In some embodiments, converting a target file into a question-answer pair dataset comprises the following steps:

[0014] Based on the preset ratio of repeated question-answer pairs, a generative pre-trained transformer is used to extract key information from the file content of the target file, and then the key information is converted into a question-answer format to obtain a second question-answer pair.

[0015] The second question-answer pairs corresponding to all key information are summarized and sorted to obtain a question-answer pair dataset.

[0016] In some embodiments, merging the first question-answer pair and the second question-answer pair in the question-answer pair dataset based on similarity evaluation includes the following steps:

[0017] Evaluate the similarity between the first question-answer pair and each second question-answer pair based on a preset similarity algorithm; the similarity results include different, identical, conflicting, and similar;

[0018] Based on the similarity results, the corresponding first question-answer pair and the second question-answer pair are adaptively merged; the adaptive merging process includes retaining or deleting one question-answer pair in both question-answer pairs.

[0019] In some embodiments, obtaining a similarity result between the first question-answer pair and each second question-answer pair based on a preset similarity algorithm includes the following steps:

[0020] evaluating question similarity between the question in the first question-answer pair and the question in each second question-answer pair based on a similarity algorithm;

[0021] Evaluate the answer similarity between the answer in the first question-answer pair and the answer in each second question-answer pair based on a similarity algorithm; both question similarity and answer similarity include same, different, and similar;

[0022] When the question similarity is different, it is determined that the corresponding first question-answer pair and the second question-answer pair are different;

[0023] When the question similarity and the answer similarity are the same, it is determined that the corresponding first question-answer pair and the second question-answer pair are the same;

[0024] When the question similarities are the same and the answer similarities are different, it is determined that the corresponding first question-answer pair and the second question-answer pair are in conflict;

[0025] When the question similarity is similar, it is determined that the corresponding first question-answer pair and the second question-answer pair are similar.

[0026] In some embodiments, adaptively merging the corresponding first question-answer pair and the second question-answer pair based on the similarity result includes the following steps:

[0027] When the similarity results are the same, the corresponding first question-answer pair or the second question-answer pair is deleted;

[0028] When the similarity results are different, the corresponding first question-answer pair and the second question-answer pair are retained;

[0029] When the similarity result is a conflict, a first prompt message is sent to the target object according to the corresponding first question-answer pair and the second question-answer pair to obtain a first decision instruction of the target object, and a first target processing is performed on the corresponding first question-answer pair and the second question-answer pair in response to the first decision instruction;

[0030] When the similarity result is similar, a second prompt message is sent to the target object according to the corresponding first question-answer pair and the second question-answer pair to obtain a second decision instruction of the target object, and a second target processing is performed on the corresponding first question-answer pair and the second question-answer pair in response to the second decision instruction; the first target processing and the second target processing include retaining or deleting one question-answer pair in both question-answer pairs.

[0031] In some embodiments, the method further comprises the following steps:

[0032] In response to an additional input instruction of the target object, obtaining additional information input by the target object; the additional information includes at least one of an additional file and a first additional question-answer pair;

[0033] When there is an additional file in the additional data, convert the additional file into an additional question-answer pair dataset; the additional question-answer pair dataset includes multiple second additional question-answer pairs;

[0034] When the additional data still contains a first additional question-answer pair, the first additional question-answer pair and the second additional question-answer pair in the additional question-answer pair dataset are merged based on the similarity evaluation to obtain an additional set of question-answer pairs;

[0035] Based on the similarity evaluation, the question-answer pair training set and the question-answer pair additional set are merged to obtain the question-answer pair additional training set;

[0036] The question-answer pair training set is updated through the additional training set of question-answer pairs, and then the digital human question-answer dialogue model is optimized and trained using the question-answer pair training set to update the digital human question-answer dialogue model.

[0037] To achieve the above objectives, another aspect of an embodiment of the present invention provides a question-answering training system for a dialogue model, the system comprising:

[0038] The first module is configured to obtain target information input by a target subject; the target information includes at least one of a target file and a first question-answer pair;

[0039] The second module is configured to convert a target file into a question-answer pair dataset when the target file exists in the target data; the question-answer pair dataset includes a plurality of second question-answer pairs;

[0040] The third module is used to merge the first question-answer pair with the second question-answer pair in the question-answer pair dataset based on similarity evaluation when the first question-answer pair still exists in the target data to obtain a question-answer pair training set;

[0041] The fourth module is used to input the question-answer pair training set into the target model for training and generate a dialogue model.

[0042] In some embodiments, the system further includes a fifth module, specifically configured to perform the following operations:

[0043] In response to an additional input instruction of the target object, obtaining additional information input by the target object; the additional information includes at least one of an additional file and a first additional question-answer pair;

[0044] When there is an additional file in the additional data, convert the additional file into an additional question-answer pair dataset; the additional question-answer pair dataset includes multiple second additional question-answer pairs;

[0045] When the additional data still contains a first additional question-answer pair, the first additional question-answer pair and the second additional question-answer pair in the additional question-answer pair dataset are merged based on the similarity evaluation to obtain an additional set of question-answer pairs;

[0046] Based on the similarity evaluation, the question-answer pair training set and the question-answer pair additional set are merged to obtain the question-answer pair additional training set;

[0047] The question-answer pair training set is updated through the additional training set of question-answer pairs, and then the digital human question-answer dialogue model is optimized and trained using the question-answer pair training set to update the digital human question-answer dialogue model.

[0048] To achieve the above object, another aspect of an embodiment of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above method when executing the computer program.

[0049] To achieve the above object, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0050] Embodiments of the present invention include at least the following advantageous effects: The present invention provides a question-answering training method, system, electronic device, and storage medium for a conversational model. The method comprises obtaining target data input by a target subject; the target data comprises at least one of a target file and a first question-answer pair; when the target file exists in the target data, converting the target file into a question-answer pair dataset; the question-answer pair dataset comprises multiple second question-answer pairs; when the first question-answer pair also exists in the target data, merging the first question-answer pair with the second question-answer pairs in the question-answer pair dataset based on similarity evaluation to obtain a question-answer pair training set; and inputting the question-answer pair training set into a target model for training to generate a conversational model. The target model comprises GPT, BERT, and LLM, and the conversational model is applied in text interaction, voice interaction, and digital human. By automatically converting the target file into a structured question-answer pair dataset, the present invention generates high-quality second question-answer pairs in batches, effectively improving data processing efficiency compared to manual compilation, which is time-consuming, labor-intensive, and inefficient. Furthermore, the present invention eliminates redundant question-answer pairs through similarity calculation, avoiding model overfitting caused by repeated training data. The present invention can effectively improve the question-answering training efficiency of the conversational model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of a question-answering training method for a dialogue model provided by an embodiment of the present invention;

[0052] Figure 2 Schematic diagram of the overall process of the question-answering training method for the dialogue model provided by an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of the structure of the question-answering training system for the dialogue model provided by an embodiment of the present invention;

[0054] Figure 4 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are merely examples of systems and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0056] It will be understood that the terms "first," "second," and the like used in the present invention may be used to describe various concepts in the present invention, but unless otherwise specified, these concepts are not limited by these terms. These terms are merely used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the terms "if" and "if" as used herein may be interpreted as "at the time of," "when," or "in response to a determination."

[0057] The terms "at least one", "plurality", "each", "any", etc. used in the present invention include at least one, two or more, multiple, two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0058] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art to which the present invention pertains. The terms used in the present invention are for the purpose of describing the embodiments of the present invention only and are not intended to limit the present invention.

[0059] To facilitate the description of the technical solutions of the present invention, the following glossary of proprietary technologies that may be applied in the embodiments of the present invention is provided:

[0060] EZ-Train (EZ-Train Training Framework): A GPT-based automated Q&A training framework designed to efficiently generate conversational data for training digital humans. Features include: 1. Automatically converting unstructured text (such as PDFs) into a structured Q&A format; 2. Handling different types of Q&A (distinct, repetitive, and similar); and 3. Providing a user-friendly training method for non-engineers.

[0061] Q&A Pair: A structured conversation unit consisting of a question (Q) and a corresponding answer (A).

[0062] Digital Human: A virtual human based on artificial intelligence technology (such as large language models) that can interact and communicate through natural language.

[0063] Repeated Q&A pair ratio (β): The threshold ratio used to determine whether the generated Q&A sufficiently covers the original content, ensuring that the generated dataset is both sufficient and not excessive.

[0064] The question-and-answer training method for a dialogue model provided in an embodiment of the present invention relates to the technical field of question-and-answer training for dialogue models. The question-and-answer training method for a dialogue model provided in an embodiment of the present invention can be applied to a terminal or a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and an in-vehicle terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or as a server cluster or distributed system consisting of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the question-and-answer training method for a dialogue model, etc., but is not limited to the above forms.

[0065] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0066] Figure 1 This is an optional flowchart of the question-answering training method for the dialogue model provided by an embodiment of the present invention. Figure 1 The method may include but is not limited to steps S100 to S400.

[0067] S100, obtaining target information input by the target object;

[0068] The target data includes at least one of a target file and a first question-answer pair;

[0069] It should be noted that, in some embodiments, step S100 may include at least one of the following steps: obtaining a target file in a target format in response to an upload instruction from the target object; obtaining a first question-answer pair in response to a manual input instruction from the target object.

[0070] For example, in some specific implementations, users can enter data into the EZ-Train framework by manually entering question-answer pairs or uploading pre-formatted files such as PDF files or Word files. In some optional implementations, the user's voice commands can be converted into text question-answer pairs to obtain the user-entered question-answer pairs, further enhancing the convenience of data entry.

[0071] S200: When a target file exists in the target data, convert the target file into a question-answer pair dataset;

[0072] The question-answer pair dataset includes multiple second question-answer pairs;

[0073] It should be noted that, in some embodiments, converting the target file into a question-answer pair dataset may include the following steps: based on a preset ratio of repeated question-answer pairs, using a generative pre-trained transformer to extract key information from the file content of the target file, and then converting the key information into a question-answer format to obtain a second question-answer pair; summarizing and arranging the second question-answer pairs corresponding to all key information to obtain a question-answer pair dataset.

[0074] For example, in some specific implementations, the EZ-Train framework automatically converts uploaded files in a pre-set format into a unified question-and-answer format. This process utilizes Generative Pre-Trained Transformers (GPT) technology, which can extract key information from unstructured file content and automatically generate relevant questions and answers.

[0075] Among them, in some specific application scenarios, in order to ensure that the question-answer pairs can fully cover the file content, the EZ-Train framework introduces a ratio β to control the number of question-answer pairs generated. Through experimental testing, it is determined that when β>50%, the proportion of repeated or similar questions and answers in the newly generated question-answer pairs is high. At this time, the generation of new question-answer pairs can be stopped, thereby ensuring the quality of the question-answer pairs and the training efficiency. Preferably, the embodiment of the present invention can set β to 58% (stop generating when the proportion of repeated question-answer pairs reaches 58%).

[0076] In addition, in some optional implementations, in addition to controlling the number of question-answer pairs generated using the ratio β, other methods can be used to evaluate the coverage and saturation of question-answer pairs. For example, by analyzing the semantic distribution and keyword coverage of the question-answer pairs, it can be determined whether to continue generating new question-answer pairs. In addition, a manual review process can be introduced. After a certain number of question-answer pairs are generated, professionals will evaluate the quality and coverage of the question-answer pairs and decide whether to stop generating based on the evaluation results.

[0077] S300: When the first question-answer pair still exists in the target data, the first question-answer pair is merged with the second question-answer pair in the question-answer pair dataset based on similarity evaluation to obtain a question-answer pair training set.

[0078] It should be noted that, in some embodiments, merging the first question-answer pair and the second question-answer pair in the question-answer pair data set based on similarity evaluation may include the following steps: evaluating the similarity results of the first question-answer pair and each second question-answer pair based on a preset similarity algorithm; the similarity results include different, identical, conflicting and similar; adaptively merging the corresponding first question-answer pair and the second question-answer pair based on the similarity results; the adaptive merging includes retaining or deleting one question-answer pair in both question-answer pairs.

[0079] For example, in some specific implementations, when both manual input and file upload are used for data input, the EZ-Train framework enters the merge phase. This phase primarily addresses four types of question-answer pairs: different, identical, conflicting, and similar. For different question-answer pairs, corresponding merge strategies are applied based on their similarity.

[0080] Among them, in some embodiments, obtaining the similarity results of the first question and answer pair and each second question and answer pair based on a preset similarity algorithm can include the following steps: evaluating the question similarity between the question in the first question and answer pair and the question in each second question and answer pair based on the similarity algorithm; evaluating the answer similarity between the answer in the first question and answer pair and the answer in each second question and answer pair based on the similarity algorithm; both the question similarity and the answer similarity include same, different and similar; when the question similarity is different, determining that the corresponding first question and answer pair and the second question and answer pair are different; when the question similarity and the answer similarity are both the same, determining that the corresponding first question and answer pair and the second question and answer pair are the same; when the question similarity is the same and the answer similarity is different, determining that the corresponding first question and answer pair and the second question and answer pair are in conflict; when the question similarity is similar, determining that the corresponding first question and answer pair and the second question and answer pair are similar. Specifically, the results of question similarity and answer similarity can be determined by mapping the range of similarity values obtained by evaluating the similarity algorithm (for example, greater than 0.9 is the same, less than 0.1 is different, and the range of 0.1 to 0.9 is similar. This is only an example and should not be regarded as a limitation of the present invention).

[0081] For example, in some specific implementations, when processing question-answer pairs, the EZ-Train framework uses algorithms such as cosine similarity and Levenshtein distance, combined with GPT technology, to evaluate the similarity of question-answer pairs. By calculating the similarity value between question-answer pairs, the relationship between question-answer pairs (different, identical, conflicting, and similar) is determined, providing a basis for merging processing and ensuring the quality and training effect of question-answer pairs. In addition, although the cosine similarity and Levenshtein distance algorithms have certain advantages in evaluating the similarity of question-answer pairs, they may face efficiency bottlenecks when processing large-scale data. Therefore, more efficient similarity evaluation algorithms can be explored, such as similarity metric learning methods based on deep learning, which learn the similarity metric between question-answer pairs by training a deep neural network model, thereby improving the efficiency of the algorithm while ensuring the accuracy of the evaluation.

[0082] Among them, in some embodiments, adaptive merging processing is performed on the corresponding first question and answer pair and the second question and answer pair based on the similarity results, which may include the following steps: when the similarity results are the same, the corresponding first question and answer pair or the second question and answer pair is deleted; when the similarity results are different, the corresponding first question and answer pair and the second question and answer pair are retained; when the similarity results are conflicting, a first prompt message is sent to the target object according to the corresponding first question and answer pair and the second question and answer pair to obtain a first decision instruction of the target object, and the first target processing is performed on the corresponding first question and answer pair and the second question and answer pair in response to the first decision instruction; when the similarity results are similar, a second prompt message is sent to the target object according to the corresponding first question and answer pair and the second question and answer pair to obtain a second decision instruction of the target object, and the second target processing is performed on the corresponding first question and answer pair and the second question and answer pair in response to the second decision instruction; the first target processing and the second target processing include retaining or deleting one question and answer pair for both question and answer pairs.

[0083] For example, in some specific implementations, no special processing is performed for different question-answer pairs; for the same question-answer pairs, one question-answer pair is retained and duplicate question-answer pairs are deleted; for similar question-answer pairs, their similarity is evaluated using GPT technology, and whether to merge or retain them is decided based on the similarity value and a user-set threshold (such as 0.5), so as to reduce the redundancy of question-answer pairs and improve training efficiency.

[0084] Furthermore, in some optional implementations, when processing identical and similar question-answer pairs, in addition to using GPT technology to assess similarity, other natural language processing algorithms, such as word embedding-based similarity calculation and semantic role labeling, can be combined to more accurately determine the relationship between question-answer pairs, thereby implementing a more refined merging strategy. Furthermore, a user feedback mechanism can be introduced to allow users to participate in the decision-making process of question-answer pair merging, allowing adjustments and optimizations to be made based on their actual needs and preferences.

[0085] In some specific application scenarios, when manual input and file upload are used for data input, the EZ-Train framework enters the merging phase. This phase mainly processes four types of question-answer pairs (different, identical, conflicting, and similar):

[0086] 1. For different questions, no special treatment is done (i.e., both corresponding question-answer pairs are retained);

[0087] 2. For the same questions and answers, keep one question and answer and delete the duplicate questions and answers;

[0088] 3. For the same question and different answers, display them as conflicts and let the user select the required question and answer;

[0089] 4. For similar questions, the similarity between two questions or two answers is evaluated through GPT or other effective algorithm technology to determine whether the two questions or two answers are similar, and ultimately allow users to select the questions and answers they need.

[0090] S400: Input the question-answer pair training set into the target model for training to generate a dialogue model.

[0091] The target models include GPT, BERT, and LLM, and the application forms of the dialogue models include text interaction, voice interaction, and digital human.

[0092] For example, in some specific implementations, the target models include, but are not limited to, GPT, BERT, and LLM. Specifically, GPT utilizes the Transformer decoder structure to generate coherent text through autoregression. BERT (Bidirectional Encoding Representation Model) captures contextual associations through a bidirectional attention mechanism and outputs semantic encoding vectors for use in downstream tasks. LLM (Large Language Model), with its hundreds of billions of parameters, achieves domain-specific adaptation through instruction fine-tuning.

[0093] The application forms of dialogue models include but are not limited to text interaction, voice interaction and digital humans. Specifically: a text interaction example can be a webpage embedded chat window, which calls the dialogue model to generate a reply after receiving user text input. Specifically, the front-end can call the REST API to interact with the back-end model, and combine with the intent recognition module to achieve multi-round dialogue management. Voice interaction can be achieved by a smart speaker device converting the user's voice into text through speech recognition (ASR), generating reply text after parsing by the dialogue model, and then synthesizing the voice output through TTS. Digital human interaction can be achieved by a 3D virtual image driven by motion capture equipment to synchronize lips, combining the reply text generated by the dialogue model with facial expressions and movements, and further combining computer vision (expression generation algorithm) and speech synthesis technology to achieve a multimodal interactive experience.

[0094] In some alternative implementations, the GPT model leverages its powerful language generation capabilities and pre-trained knowledge to learn and optimize question-answer pairs, enabling the conversational model to generate natural and accurate question-answer dialogues. Furthermore, while the GPT model excels in generating question-answer dialogues, other similar pre-trained language models, such as BERT and RoBERTa, can also be considered, with appropriate fine-tuning and optimization to adapt them to the needs of question-answer dialogue training. Furthermore, hybrid model architectures can be constructed by combining the strengths of multiple models to further improve the quality and performance of question-answer dialogues.

[0095] In some embodiments, the method may further include the following steps: in response to an additional input instruction of the target object, obtaining additional data input by the target object; the additional data includes at least one of an additional file and a first additional question-answer pair; when an additional file exists in the additional data, converting the additional file into an additional question-answer pair dataset; the additional question-answer pair dataset includes multiple second additional question-answer pairs; when a first additional question-answer pair still exists in the additional data, merging the first additional question-answer pair and the second additional question-answer pair in the additional question-answer pair dataset based on similarity evaluation to obtain an additional set of question-answer pairs; merging the question-answer pair training set and the question-answer pairs in the additional set of question-answer pairs based on similarity evaluation to obtain an additional training set of question-answer pairs; updating the question-answer pair training set through the additional training set of question-answer pairs, and then optimizing and training the digital human question-answer dialogue model using the question-answer pair training set to update the digital human question-answer dialogue model. Specifically, the process steps for converting and merging additional data are the same as the processing logic of the aforementioned target data, and will not be repeated here.

[0096] For example, in some specific implementations, after the dialogue model training is completed, the user can continue to input new data into the EZ-Train framework as needed, and through the same input and conversion process, merge the new question-answer pairs into the existing question-answer pairs to further train and optimize the dialogue model, thereby continuously improving the dialogue model's dialogue capabilities.

[0097] In order to explain the principle of the technical solution of the present invention in detail, the overall process of the present invention is described below in combination with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and cannot be regarded as a limitation of the present invention.

[0098] First, it's important to note that pre-training techniques and fine-tuning methods can generate highly coherent and contextually relevant natural language responses. For example, some technologies further improve conversation quality by introducing more data and parameters. Other technologies provide a large-scale corpus (3TB of data) to support pre-training effective language models, such as the 3-billion-parameter Transformer-XL, which significantly improves Chinese natural language processing capabilities. Other technologies propose a pre-trained conversation model that can generate personalized responses from sparse data, enriching conversation content through personal attribute embedding and attention routing. Furthermore, some technologies achieve near-human-level conversation quality by learning from exchanges on Reddit. However, these models and systems still face challenges in terms of the quality and diversity of training data.

[0099] Although existing generative dialogue systems have made significant progress in generating natural language responses, challenges remain in efficiently converting text data into question-answer dialogue pairs for digital human deployment. Related technologies have the following shortcomings:

[0100] 1. These methods are not friendly to non-engineering technicians and lack ease of use.

[0101] 2. The existing generation process requires a lot of work to generate interactive digital humans, which limits the application of digital humans in a wider range of fields.

[0102] 3. Existing generation methods often cannot effectively handle the same, similar, and different questions and their corresponding answers when processing question-answer pairs.

[0103] 4. This results in redundancy and inconsistency in the generated question-answer pairs, affecting training efficiency and conversation quality.

[0104] In view of this, the present invention aims to improve the conversational performance of digital humans, enabling them to interact with users more naturally and efficiently, while reducing training costs and time, and promoting the application of digital humans in more fields. Figure 2 As shown, the technology of the present invention can be implemented as follows:

[0105] This paper proposes a method for training digital human question-answering dialogues based on the EZ-Train framework. This method aims to efficiently convert user-provided files (using PDF format as an example) into structured question-answer pairs, optimize the training process, and improve both training efficiency and dialogue quality. The specific technical solution is as follows:

[0106] S1. Data input and conversion:

[0107] Users can enter data into the EZ-Train framework by manually entering question-answer pairs or uploading PDF files. The framework automatically converts uploaded PDF files into a unified question-answer format. This process utilizes Generative Pre-Train Transformer (GPT) technology, which extracts key information from unstructured PDF content and automatically generates relevant questions and answers.

[0108] S2. Question and answer pair merging process:

[0109] When both manual input and PDF file upload are used for data input, the EZ-Train framework enters the merging phase. This phase primarily addresses four types of question-answer pairs: different, identical, conflicting, and similar. Different question-answer pairs are not treated specially; for identical pairs, one pair is retained and duplicates are deleted; conflicting pairs are merged based on user decision-making; and for similar pairs, GPT technology is used to evaluate their similarity. The decision to merge or retain the pairs is based on the similarity value and a user-defined threshold (e.g., 0.5), reducing redundancy in question-answer pairs and improving training efficiency.

[0110] S3.GPT training:

[0111] The merged question-answer pairs are fed into the GPT model for training, generating a digital human question-answering dialogue model. The GPT model leverages its powerful language generation capabilities and pre-trained knowledge to learn and optimize the question-answer pairs, enabling the digital human to generate natural and accurate question-answering dialogues.

[0112] S4. Additional training:

[0113] After the training of the digital human question-answering dialogue model is completed, users can continue to input new data into the EZ-Train framework as needed. Through the same input and conversion process, the new question-answer pairs can be merged into the existing question-answer pairs, and the digital human question-answering dialogue model can be further trained and optimized to continuously improve the digital human's conversational capabilities.

[0114] S5. Question-Answer Pair Similarity Evaluation:

[0115] When processing question-answer pairs, the EZ-Train framework uses algorithms such as cosine similarity and Levenshtein distance, combined with GPT technology, to assess the similarity of question-answer pairs. By calculating the similarity between question-answer pairs, the relationship between them (different, identical, similar) is determined, providing a basis for merging and ensuring the quality of the question-answer pairs and the effectiveness of training.

[0116] S6. Question-Answer Pair Generation Control:

[0117] To ensure that the Q&A pairs fully cover the content of PDF files, the EZ-Train framework introduces a ratio β to control the number of Q&A pairs generated. Experimental testing has determined that when β exceeds 50%, the proportion of duplicate or similar Q&A pairs in the generated Q&A pairs is high. At this point, the generation of new Q&A pairs can be stopped, thus ensuring the quality of the Q&A pairs and training efficiency.

[0118] Specifically, the workflow of the EZ-Train framework of an embodiment of the present invention includes: (a) manual input and / or PDF file uploading and automatic conversion into question and answer format; (b) merging manual input and PDF conversion into a final question and answer corpus for question and answer training; (c) additional training.

[0119] For example, in some specific application scenarios, the technical effects of the present invention are described in detail in combination with the experimental verification results of the method of the embodiment of the present invention:

[0120] One embodiment of the present invention involves question-and-answer dialogue training for a digital humanoid called "Cantonese Chef." This digital humanoid aims to promote and introduce Cantonese cuisine and dim sum culture, including its history, culinary techniques, and dietary customs. Using the EZ-Train framework, relevant PDF files and manually entered question-and-answer pairs were processed and trained, generating 98 question-and-answer pairs covering various aspects of Cantonese cuisine and dim sum culture. During the training process, the EZ-Train framework effectively addressed differences, similarities, and differences between question-and-answer pairs, ensuring both quality and training efficiency. The trained "Cantonese Chef" digital humanoid is able to engage in natural and fluent conversations with users, accurately answering their questions about Cantonese cuisine and dim sum culture and providing them with a wealth of knowledge and information. Another embodiment involves applying the EZ-Train framework to question-and-answer dialogue training for a digital humanoid called "Incubation Center Ambassador." This digital humanoid is deployed at the entrance of an incubation center to provide visitors with information about the center's projects, collaborations, and achievements. By uploading the PDF file provided by the incubation center, the EZ-Train framework quickly generated question-answer pairs and trained a digital human question-answer dialogue model that can meet the needs of actual applications. In actual deployment, the digital human can provide visitors with accurate and detailed information in real time, improving the visitor experience and understanding of the incubation center. In addition, the present invention also tested and verified the performance of the EZ-Train framework in processing question-answer pairs of different topics (such as Mediterranean diet, Asian artificial intelligence art, China travel, etc.). Under each topic, the EZ-Train framework can effectively generate 180 pairs of question-answer pairs, and successfully handle a variety of different, identical and similar question-answer pairs, demonstrating its wide applicability and efficiency in different fields and topics.

[0121] In summary, the present invention proposes a digital human question-answering dialogue training method based on the EZ-Train framework. Through automated question-answer pair generation, merging processing and GPT training, efficient, convenient and high-quality question-answering dialogue training is achieved, solving the problems of low efficiency of manually writing question-answer pairs, lack of flexibility of rule-based methods, and generation of question-answering dialogues using pre-training models and fine-tuning techniques in the existing technology. The present invention innovatively introduces a ratio β to control the number of question-answer pairs generated. Through experimental tests, it is determined that when β>58%, the proportion of repeated or similar questions in the newly generated question-answer pairs is high, thereby effectively avoiding excessive redundancy of question-answer pairs, improving training efficiency, and ensuring that the question-answer pairs can fully cover the content of the input file. In addition, the present invention designs an efficient question-answer pair similarity evaluation algorithm, which uses algorithms such as cosine similarity and Levenshtein distance combined with GPT technology to accurately evaluate the similarity of question-answer pairs, providing a reliable basis for the merging processing of question-answer pairs, ensuring the quality of question-answer pairs and training effect, and the time complexity of the algorithm is determined to be O(n 2 ), with good efficiency and performance when processing large-scale question-answer pairs. This invention provides a user-friendly question-answering dialogue training method, enabling even non-engineering technicians to easily conduct digital human question-answering dialogue training. This lowers the threshold for digital human technology application and facilitates its widespread application and promotion in various fields.

[0122] Compared with the prior art, the present invention has at least the following beneficial effects:

[0123] 1. High quality:

[0124] This invention uses advanced GPT technology and a reasonable similarity evaluation algorithm to generate natural, accurate, and high-quality question-answer pairs, improving the conversational performance of digital humans, enabling them to better understand and answer users' questions and provide a better interactive experience. Specifically:

[0125] Natural language processing capabilities: GPT technology can generate natural and fluent text, making question-answer pairs closer to human expression.

[0126] Similarity assessment: Using algorithms such as cosine similarity and Levenshtein distance, we accurately assess the similarity of question-answer pairs, ensuring that the generated question-answer pairs are of high quality and consistent.

[0127] Optimization processing: Automatically remove duplicates and merge similar question-answer pairs to reduce redundancy and improve the accuracy and consistency of question-answer pairs.

[0128] 2. Strong adaptability:

[0129] This invention is applicable to digital human question-answering dialogue training in various fields and topics. Whether it is professional knowledge question-answering or information consultation in daily life, the EZ-Train framework can quickly generate targeted question-answer pairs to meet the digital human dialogue needs in different scenarios. Specifically:

[0130] Multi-domain support: Experimental verification shows that the EZ-Train framework can effectively handle question-answer pairs in different fields, such as Mediterranean diet, Asian AI art, and Chinese travel.

[0131] Fast generation: Under each topic, the EZ-Train framework can quickly generate 180 question-answer pairs, significantly improving generation efficiency.

[0132] Flexibility: Allows users to adjust the generated question-answer pairs according to their specific needs, ensuring their adaptability.

[0133] 3. Scalability:

[0134] It allows users to add new data for additional training at any time, continuously updating and optimizing the digital human's question-answering dialogue model, enabling the digital human to continuously learn and adapt to new knowledge and information, and maintain the timeliness and accuracy of the dialogue. Specifically:

[0135] Continuous learning: Users can upload new files or manually enter new question-answer pairs at any time for additional training through the EZ-Train framework.

[0136] Dynamic updates: New data is automatically merged into existing question-answer pairs, ensuring that digital humans can continuously learn new knowledge and information.

[0137] Efficient integration: Through automated data processing and merging mechanisms, new data can be quickly integrated into existing question-answer pairs, improving training efficiency.

[0138] 4. Reduce training costs:

[0139] Compared with traditional question-answering dialogue training methods, this invention reduces the demand for large amounts of computing resources, lowers the hardware and labor costs during the training process, makes digital human question-answering dialogue training more economical and feasible, and is conducive to the widespread application and promotion of digital human technology. Specifically:

[0140] Efficient Algorithms: Through optimized algorithms and strategies, we ensure efficiency and accuracy when processing large-scale question-answer pairs and reduce the consumption of computing resources.

[0141] User-friendly: The EZ-Train framework provides a simple and easy-to-use operation interface, making it easy for non-engineering technicians to get started, reducing labor costs.

[0142] Economical: By reducing the need for large amounts of computing resources, hardware costs are reduced, making digital human question-and-answer dialogue training more economical and feasible.

[0143] 5. Lower the threshold for non-engineer users:

[0144] This invention supports code-free training of digital humans. Through a simple and easy-to-use operating interface and automated processing flow, non-engineering technicians can also easily get started without writing code or performing complex settings. Specifically:

[0145] Code-free training: Users can train by manually entering question-answer pairs or uploading files without writing code.

[0146] Automated processing: The EZ-Train framework automatically handles data conversion, merging, and training, allowing users to complete training with simple operations.

[0147] User-friendly: Detailed operation guides and help documents are provided to ensure that users can quickly master the training process.

[0148] 6. Improve Q&A generation efficiency:

[0149] Compared with manually constructed datasets, the Q&A generation speed of the present invention is increased by at least 3 times, significantly improving training efficiency. Specifically:

[0150] Fast generation: Through GPT technology and automated processing, the EZ-Train framework is able to generate a large number of high-quality question-answer pairs in a short period of time.

[0151] Efficient processing: Automatically remove duplicates and merge similar question-answer pairs to reduce redundancy and improve generation efficiency.

[0152] Batch generation: Supports batch generation of question and answer pairs. Users can quickly generate multiple question and answer pairs by clicking a button, further improving the generation speed.

[0153] 7. Optimize training data quality:

[0154] This invention optimizes the quality of training data and improves the consistency and accuracy of question-answer pairs by automatically removing duplicates and merging similar question-answer pairs. Specifically:

[0155] Automatic deduplication: Through similarity evaluation algorithms, duplicate question and answer pairs are automatically identified and removed to reduce redundancy.

[0156] Merge similar question-answer pairs: Automatically merge similar question-answer pairs to ensure their consistency and accuracy.

[0157] Data optimization: Through optimized algorithms and strategies, we ensure that the generated question-answer pairs are of high quality and good consistency, thereby improving training results.

[0158] 8. Improve the quality of digital human interaction:

[0159] This invention can automatically expand the content of conversations based on domain knowledge, improve the interactive quality of digital humans, enable digital humans to better understand and answer users' questions, and provide a better interactive experience. Specifically:

[0160] Domain knowledge expansion: Through GPT technology, the conversation content is automatically expanded based on domain knowledge, ensuring that digital humans can provide richer information.

[0161] Intelligent answering: Through an optimized question-answer generation mechanism, digital humans can answer users' questions more intelligently and provide more accurate answers.

[0162] Interaction optimization: Through continuous learning and optimization, we ensure that the interaction quality of digital humans continues to improve and provide a better user experience.

[0163] In summary, this invention significantly improves the efficiency and quality of digital human question-answering dialogue training through advanced technology and optimized algorithms, reduces the usage threshold and training costs, and has broad application prospects and important practical significance.

[0164] On the other hand, Figure 3 As shown, an embodiment of the present invention further provides a question-answering training system 900 for a dialogue model, which may include:

[0165] The first module 901 is configured to obtain target information input by a target subject; the target information includes at least one of a target file and a first question-answer pair;

[0166] The second module 902 is configured to convert a target file into a question-answer pair dataset when the target file exists in the target data; the question-answer pair dataset includes a plurality of second question-answer pairs;

[0167] The third module 903 is configured to, when the first question-answer pair still exists in the target data, merge the first question-answer pair with the second question-answer pair in the question-answer pair dataset based on similarity evaluation to obtain a question-answer pair training set;

[0168] The fourth module 904 is used to input the question-answer pair training set into the target model for training to generate a dialogue model.

[0169] In some embodiments, the system may further include a fifth module, specifically configured to perform the following operations:

[0170] In response to an additional input instruction of the target object, obtaining additional information input by the target object; the additional information includes at least one of an additional file and a first additional question-answer pair;

[0171] When there is an additional file in the additional data, convert the additional file into an additional question-answer pair dataset; the additional question-answer pair dataset includes multiple second additional question-answer pairs;

[0172] When the additional data still contains a first additional question-answer pair, the first additional question-answer pair and the second additional question-answer pair in the additional question-answer pair dataset are merged based on the similarity evaluation to obtain an additional set of question-answer pairs;

[0173] Based on the similarity evaluation, the question-answer pair training set and the question-answer pair additional set are merged to obtain the question-answer pair additional training set;

[0174] The question-answer pair training set is updated through the additional training set of question-answer pairs, and then the digital human question-answer dialogue model is optimized and trained using the question-answer pair training set to update the digital human question-answer dialogue model.

[0175] The contents of the method embodiments of the present invention are all applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.

[0176] An embodiment of the present invention further provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned question-and-answer training method for a conversational model. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.

[0177] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0178] See also Figure 4 , Figure 4 The hardware structure of an electronic device 1000 according to another embodiment is shown. The electronic device 1000 includes:

[0179] The processor 1001 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.

[0180] Memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 1002 and is called by processor 1001 to execute the question-answering training method for the dialogue model of the embodiments of the present invention.

[0181] Input / output interface 1003, used to implement information input and output;

[0182] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0183] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );

[0184] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .

[0185] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the question-answering training method of the above-mentioned dialogue model.

[0186] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0187] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0188] The embodiment of the present invention provides a question-answering training method for a dialogue model, a question-answering training system for a dialogue model, an electronic device, and a storage medium. The method obtains target data input by a target object; the target data includes a target file and at least one of a first question-answer pair; when a target file exists in the target data, the target file is converted into a question-answer pair dataset; the question-answer pair dataset includes multiple second question-answer pairs; when a first question-answer pair still exists in the target data, the first question-answer pair and the second question-answer pair in the question-answer pair dataset are merged based on a similarity assessment to obtain a question-answer pair training set; the question-answer pair training set is input into the target model for training to generate a dialogue model. The present invention automatically converts the target file into a structured question-answer pair dataset. Compared with manual writing, which is time-consuming, labor-intensive, and has low productivity, the present invention effectively improves data processing efficiency by batch-generating high-quality second question-answer pairs. In addition, the present invention eliminates redundant question-answer pairs through similarity calculation to avoid model overfitting caused by repeated training data. The present invention can effectively improve the question-answering training quality of the dialogue model.

[0189] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0190] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0191] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0192] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0193] The terms "first," "second," "third," "fourth," and the like (if any) in the description of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or apparatus.

[0194] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0195] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, and can be electrical, mechanical or other forms.

[0196] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0197] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0198] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store programs.

[0199] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the invention is not limited thereby. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the invention should be within the scope of the invention.

Claims

1. A question-answering training method for a dialogue model, characterized in that: The method comprises the following steps: Obtaining target information input by the target subject; the target information includes at least one of a target file and a first question-answer pair; When the target file exists in the target data, converting the target file into a question-answer pair dataset; the question-answer pair dataset includes a plurality of second question-answer pairs; When the first question-answer pair still exists in the target material, merging the first question-answer pair with the second question-answer pair in the question-answer pair dataset based on similarity evaluation to obtain a question-answer pair training set; Inputting the question-answer pair training set into the target model for training to generate a dialogue model; Among them, the target models include GPT, BERT and LLM, and the application forms of the dialogue model include text interaction, voice interaction and digital human.

2. The question-answering training method for a dialogue model according to claim 1, characterized in that: The step of obtaining the target information input by the target object includes at least one of the following steps: In response to an upload instruction of the target object, obtaining the target file in a target format; In response to a manual input instruction of the target object, the first question-answer pair is obtained.

3. The question-answering training method for a dialogue model according to claim 1, characterized in that: Converting the target file into a question-answer pair dataset comprises the following steps: Based on a preset ratio of repeated question-answer pairs, extract key information from the file content of the target file using a generative pre-trained transformer, and then convert the key information into a question-answer format to obtain the second question-answer pair; The second question-answer pairs corresponding to all the key information are aggregated and collated to obtain the question-answer pair dataset.

4. The question-answering training method for a dialogue model according to claim 1, characterized in that: The merging process of the first question-answer pair and the second question-answer pair in the question-answer pair dataset based on similarity evaluation includes the following steps: Evaluate the similarity results between the first question-answer pair and each of the second question-answer pairs based on a preset similarity algorithm; the similarity results include different, identical, conflicting, and similar; Based on the similarity result, adaptive merging processing is performed on the corresponding first question-answer pair and the second question-answer pair; the adaptive merging processing includes retaining or deleting one question-answer pair in both question-answer pairs.

5. The question-answering training method for a dialogue model according to claim 4, characterized in that: The obtaining of a similarity result between the first question-answer pair and each of the second question-answer pairs based on a preset similarity algorithm comprises the following steps: evaluating, based on the similarity algorithm, the question similarity between the question in the first question-answer pair and the question in each of the second question-answer pairs; Evaluate the answer similarity between the answer in the first question-answer pair and the answer in each of the second question-answer pairs based on the similarity algorithm; both the question similarity and the answer similarity include same, different, and similar; When the question similarity is different, determining that the corresponding first question-answer pair and the second question-answer pair are different; When the question similarity and the answer similarity are both the same, determining that the corresponding first question-answer pair and the second question-answer pair are the same; When the question similarities are the same and the answer similarities are different, determining that the corresponding first question-answer pair and the second question-answer pair are in conflict; When the question similarity is similar, it is determined that the corresponding first question-answer pair and the second question-answer pair are similar.

6. The question-answering training method for a dialogue model according to claim 4, characterized in that: The adaptive merging process of the corresponding first question-answer pair and the second question-answer pair based on the similarity result includes the following steps: When the similarity results are the same, deleting the corresponding first question-answer pair or the second question-answer pair; When the similarity result is different, retaining the corresponding first question-answer pair and second question-answer pair; When the similarity result is a conflict, sending a first prompt message to the target object according to the corresponding first question-answer pair and the second question-answer pair to obtain a first decision instruction of the target object, and performing a first target processing on the corresponding first question-answer pair and the second question-answer pair in response to the first decision instruction; When the similarity result is similar, a second prompt message is sent to the target object according to the corresponding first question-answer pair and the second question-answer pair to obtain a second decision instruction of the target object, and a second target processing is performed on the corresponding first question-answer pair and the second question-answer pair in response to the second decision instruction; the first target processing and the second target processing include retaining or deleting one question-answer pair for both question-answer pairs.

7. The question-answering training method for a dialogue model according to claim 1, characterized in that: The method further comprises the following steps: In response to the additional input instruction of the target object, obtaining additional information input by the target object; the additional information includes at least one of an additional file and a first additional question-answer pair; When the additional file exists in the additional material, converting the additional file into an additional question-answer pair dataset; the additional question-answer pair dataset includes a plurality of second additional question-answer pairs; When the first additional question-answer pair still exists in the additional material, merging the first additional question-answer pair with the second additional question-answer pair in the additional question-answer pair dataset based on the similarity evaluation to obtain an additional set of question-answer pairs; Merging the question-answer pair training set and the question-answer pair additional set based on the similarity evaluation to obtain the question-answer pair additional training set; The question-answer pair training set is updated by using the additional training set of the question-answer pairs, and then the digital human question-answer dialogue model is optimized and trained using the question-answer pair training set to update the digital human question-answer dialogue model.

8. A question-answering training system for a dialogue model, characterized in that: The system comprises: The first module is configured to obtain target information input by a target subject; the target information includes at least one of a target file and a first question-answer pair; A second module is configured to convert the target file into a question-answer pair dataset when the target file exists in the target data; the question-answer pair dataset includes a plurality of second question-answer pairs; A third module is configured to, when the first question-answer pair still exists in the target data, merge the first question-answer pair with the second question-answer pair in the question-answer pair dataset based on similarity evaluation to obtain a question-answer pair training set; The fourth module is used to input the question-answer pair training set into the target model for training to generate a dialogue model.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.