Data processing method and apparatus, and device and readable storage medium
Through the method of classifying the data and sorting the learning difficulty of the Q&A, the overfitting problem of large language models in the supervision fine-tuning stage is solved, and the accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- PCT/CN2024/136156
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-12-02
- Publication Date
- 2025-07-10
AI Technical Summary
During the supervised fine-tuning stage, the existing large language models directly use complex samples to adjust model parameters, resulting in the model overfitting complex samples and being unable to effectively learn simple sample features, which reduces the model accuracy and generation ability.
Through Q&A, the category information set is classified into the Q&A data, the training sample set is generated and the learning difficulty is sorted, and the model parameters are gradually adjusted to improve the model accuracy.
提高了模型在复杂和简单样本中的泛化能力,避免了过度拟合,增强了模型在多样化问答场景中的应用能力。
Smart Images

Figure CN2024136156_10072025_PF_FP_ABST
Abstract
Description
Data processing method, device, equipment and readable storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 5, 2024, with application number 202410024726.2 and application name “Data processing method, device, equipment and readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, and readable storage medium. Background Art
[0003] During the supervised fine-tuning phase of existing large-scale language models, a large number of samples are first obtained. These large samples are mixed with simple and complex samples. When the model does not have a certain level of understanding ability, directly adjusting the model parameters through complex samples will cause the model to over-adjust the parameters to fit the complex samples, resulting in the inability to learn the features of simple samples and thus unable to be applied to simple or normal samples. In other words, the model training is suboptimal, which reduces the model's accuracy and affects the model's generation ability. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method, apparatus, device, and readable storage medium, which can improve the model accuracy of a target question-answering model.
[0005] On the one hand, an embodiment of the present application provides a data processing method, including:
[0006] Obtain M question-answer pair data and a question-answer pair category information set, classify the M question-answer pair data into categories based on the question-answer pair category information set, and obtain category labels corresponding to the M question-answer pair data; the question-answer pair category information set includes the category labels; M is a positive integer;
[0007] Based on the M question-answer pairs and the class labels corresponding to the M question-answer pairs, generate N training sample sets and the sample set class labels corresponding to the N training sample sets; N is a positive integer; the input training sample in each training sample set is associated with at least one question-answer pair in the M question-answer pairs;
[0008] Sort the N training sample sets by the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence; the learning difficulty refers to the difficulty of the initial question-answering model training to learn the input training samples with the sample set category labels;
[0009] In the training sample sequence, input training samples are obtained in sequence from the sorted training sample set, and the model parameters of the initial question-answering model are adjusted in sequence through the input training samples obtained in sequence until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions.
[0010] Among them, based on the question-answer pair category information set, the M question-answer pair data are divided into categories, and the category labels corresponding to the M question-answer pair data are obtained, including:
[0011] Input M question-answer pairs and their category information set into the sample partitioning model. In the sample partitioning model, keyword vectors are generated based on the category information set, and feature extraction is performed on the M question-answer pairs to obtain feature word vectors corresponding to each of the M question-answer pairs.
[0012] Based on the feature word vector and keyword vector, determine the category labels corresponding to the M question-answer pairs.
[0013] Among them, feature extraction is performed on M question-answer pairs to obtain the feature word vectors corresponding to the M question-answer pairs, including:
[0014] Filter the M question-answer pair data for business stop words, and perform text segmentation on the filtered M question-answer pair data to obtain P question-answer text segments; P is a positive integer greater than or equal to M;
[0015] Part-of-speech tagging is performed on each of the P question-answer text segments to obtain P question-answer part-of-speech information. Based on the P question-answer text segments and the P question-answer part-of-speech information, feature word vectors corresponding to each of the M question-answer pair data are generated.
[0016] The N training sample sets include the training sample set of the first stage, the training sample set of the second stage, and the training sample set of the third stage;
[0017] Based on the M question-answer pairs and the corresponding class labels of the M question-answer pairs, N training sample sets and the corresponding sample set class labels of the N training sample sets are generated, including:
[0018] Obtain basic question-answer pair data from the M question-answer pair data, divide the basic question-answer pair data into the training sample set of the first stage, and determine the class label corresponding to the basic question-answer pair data as the sample set class label corresponding to the training sample set of the first stage;
[0019] Obtain progressive question-answer pair data from the M question-answer pair data, determine the progressive question-answer pair data as the training sample set for the second stage, generate progressive thinking chain sample category labels, and determine the progressive thinking chain sample category labels as the sample set category labels corresponding to the training sample set for the second stage; the learning difficulty of the progressive thinking chain sample category labels is greater than the learning difficulty of the category labels corresponding to the basic question-answer pair data;
[0020] A first multi-round question-answer pair data is obtained from M question-answer pair data, a second multi-round question-answer pair data is generated based on the basic question-answer pair data, the first multi-round question-answer pair data and the second multi-round question-answer pair data are determined as the training sample set of the third stage, a multi-round category label is generated, and the multi-round category label is determined as the sample set category label corresponding to the training sample set of the third stage; the learning difficulty corresponding to the multi-round category label is greater than the learning difficulty of the progressive thinking chain sample category label.
[0021] The number of basic question-answer pair data is H, where H is a positive integer less than or equal to M; the training sample set of the first stage also includes a short dialogue training sample set; and the method further includes:
[0022] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0023] According to the context information, S interrelated basic question-answer pair data are spliced into short dialogue training samples, the short dialogue training samples are added to the short dialogue training sample set, and the short dialogue label is determined as the sample set category label corresponding to the short dialogue training sample set; S is a positive integer less than or equal to H.
[0024] The second round of question-answer pair data is generated based on the basic question-answer pair data, including:
[0025] The basic question-answer pair data is spliced to obtain the spliced multi-round question-answer pair data;
[0026] Reconstruct the basic question-answer pair data to obtain reconstructed multi-round question-answer pair data;
[0027] The spliced multi-round question-answer pair data and the reconstructed multi-round question-answer pair data are determined as the second multi-round question-answer pair data.
[0028] The number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The basic question-answer pair data are spliced to obtain spliced multi-round question-answer pair data, including:
[0029] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0030] Obtain Q basic question-answer pair data based on context information, and splice the Q basic question-answer pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic relevance between the Q basic question-answer pair data is greater than a first relevance threshold;
[0031] Obtain T basic question-answer pair data based on context information, and splice the T basic question-answer pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic relevance between the T basic question-answer pair data is less than the second relevance threshold, and the first relevance threshold is greater than or equal to the second relevance threshold;
[0032] The following conversation data and random conversation data are determined as spliced multi-round question-answer pair data.
[0033] The number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The basic question-answer pair data is reconstructed to obtain reconstructed multi-round question-answer pair data, including:
[0034] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category labels contained in the question-answer pair category information set; the deep evolutionary data is associated with the evolutionary category labels;
[0035] Obtain context information corresponding to H basic question-answer pair data, and rewrite the mutually related basic question-answer pair data into breadth-evolved data based on the context information; the context information is used to indicate the semantic relevance between the H basic question-answer pair data;
[0036] Deep evolution data and broad evolution data are determined as reconstructed multi-round question-answer pair data.
[0037] Among them, based on the evolutionary constraint information, the basic question-answer pair data is rewritten into deep evolutionary data, including:
[0038] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into transition evolution data;
[0039] If the question-and-answer rounds of the transition evolution data are less than the question-and-answer round threshold, the transition evolution data will continue to be rewritten based on the evolution constraint information until the question-and-answer rounds of the transition evolution data after the data rewrite are greater than or equal to the question-and-answer round threshold, and the transition evolution data after the data rewrite will be determined as deep evolution data.
[0040] Among them, the training sample sequence includes the training sample set A i And training sample set A i+1 , training sample set A in the training sample sequence i In the training sample set A i+1 Before;
[0041] In the training sample sequence, input training samples are sequentially obtained from the sorted training sample set, and the model parameters of the initial question-answering model are adjusted sequentially through the sequentially obtained input training samples until the initial question-answering model training converges to obtain the target question-answering model, including:
[0042] In the jth round of iterative training, in the training sample set A i Randomly obtain B first input training samples from the training set, adjust the model parameters of the initial question answering model through B first input training samples, and continue to use the training sample set A i+1 Randomly obtain C second input training samples in the training sample set, and adjust the model parameters of the initial question answering model through the C second input training samples until the model parameters of the initial question answering model are adjusted through the input training samples in the Nth training sample set in the training sample sequence, and the jth round of iterative training is determined to be completed; B and C are positive integers; j is a positive integer;
[0043] If the jth round of iterative training has been completed and the initial question-answering model meets the model convergence conditions, the initial question-answering model that meets the model convergence conditions will be determined as the target question-answering model;
[0044] If the j-th round of iterative training has been completed and the initial question-answering model does not meet the model convergence conditions, then through the training sample sequence, in the j+1-th round of iterative training, the model parameters of the initial question-answering model will continue to be adjusted in sequence through the input training samples obtained in sequence until the initial question-answering model training converges and the target question-answering model is obtained.
[0045] In one aspect, an embodiment of the present application provides a data processing device, including:
[0046] A category classification module is used to obtain M question-answer pair data and a question-answer pair category information set, classify the M question-answer pair data based on the question-answer pair category information set, and obtain category labels corresponding to the M question-answer pair data; the question-answer pair category information set includes the category labels; M is a positive integer;
[0047] A sample generation module is configured to generate N training sample sets and sample set class labels corresponding to the N training sample sets based on the M question-answer pairs and the class labels corresponding to the M question-answer pairs; N is a positive integer; and an input training sample in each training sample set is associated with at least one question-answer pair from the M question-answer pairs.
[0048] A sorting processing module is used to sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence; the learning difficulty refers to the difficulty of the initial question-answering model training to learn input training samples with the sample set category labels;
[0049] The model training module is used to obtain input training samples from the sorted training sample set in sequence in the training sample sequence, and adjust the model parameters of the initial question-answering model in sequence through the input training samples obtained in sequence until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions.
[0050] In one possible implementation, the category classification module is configured to classify the M question-answer pair data into categories based on the question-answer pair category information set. When obtaining the category labels corresponding to the M question-answer pair data, the module is configured to perform the following operations:
[0051] Input M question-answer pairs and their category information set into the sample partitioning model. In the sample partitioning model, keyword vectors are generated based on the category information set, and feature extraction is performed on the M question-answer pairs to obtain feature word vectors corresponding to each of the M question-answer pairs.
[0052] Based on the feature word vector and keyword vector, determine the category labels corresponding to the M question-answer pairs.
[0053] In one possible implementation, the category classification module is used to extract features from M question-answer pairs. When obtaining feature word vectors corresponding to the M question-answer pairs, the module is specifically used to perform the following operations:
[0054] Filter the M question-answer pair data for business stop words, and perform text segmentation on the filtered M question-answer pair data to obtain P question-answer text segments; P is a positive integer greater than or equal to M;
[0055] Part-of-speech tagging is performed on each of the P question-answer text segments to obtain P question-answer part-of-speech information. Based on the P question-answer text segments and the P question-answer part-of-speech information, feature word vectors corresponding to each of the M question-answer pair data are generated.
[0056] In one possible implementation, the N training sample sets include a training sample set from the first stage, a training sample set from the second stage, and a training sample set from the third stage; the sample generation module is configured to generate the N training sample sets and the sample set category labels corresponding to the N training sample sets based on the M question-answer pair data and the category labels corresponding to the M question-answer pair data, and specifically to perform the following operations:
[0057] Obtain basic question-answer pair data from the M question-answer pair data, divide the basic question-answer pair data into the training sample set of the first stage, and determine the class label corresponding to the basic question-answer pair data as the sample set class label corresponding to the training sample set of the first stage;
[0058] Obtain progressive question-answer pair data from the M question-answer pair data, determine the progressive question-answer pair data as the training sample set for the second stage, generate progressive thinking chain sample category labels, and determine the progressive thinking chain sample category labels as the sample set category labels corresponding to the training sample set for the second stage; the learning difficulty of the progressive thinking chain sample category labels is greater than the learning difficulty of the category labels corresponding to the basic question-answer pair data;
[0059] A first multi-round question-answer pair data is obtained from M question-answer pair data, a second multi-round question-answer pair data is generated based on the basic question-answer pair data, the first multi-round question-answer pair data and the second multi-round question-answer pair data are determined as the training sample set of the third stage, a multi-round category label is generated, and the multi-round category label is determined as the sample set category label corresponding to the training sample set of the third stage; the learning difficulty corresponding to the multi-round category label is greater than the learning difficulty of the progressive thinking chain sample category label.
[0060] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M; the training sample set of the first stage also includes a short dialogue training sample set; and the sample generation module is further configured to perform the following operations:
[0061] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0062] According to the context information, S interrelated basic question-answer pair data are spliced into short dialogue training samples, the short dialogue training samples are added to the short dialogue training sample set, and the short dialogue label is determined as the sample set category label corresponding to the short dialogue training sample set; S is a positive integer less than or equal to H.
[0063] In one possible implementation, when the sample generation module is used to generate the second multiple rounds of question-answer pair data based on the basic question-answer pair data, it is specifically used to perform the following operations:
[0064] The basic question-answer pair data is spliced to obtain the spliced multi-round question-answer pair data;
[0065] Reconstruct the basic question-answer pair data to obtain reconstructed multi-round question-answer pair data;
[0066] The spliced multi-round question-answer pair data and the reconstructed multi-round question-answer pair data are determined as the second multi-round question-answer pair data.
[0067] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The sample generation module is used to perform splicing processing on the basic question-answer pair data to obtain spliced multi-round question-answer pair data, specifically performing the following operations:
[0068] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0069] Obtain Q basic question-answer pair data based on context information, and splice the Q basic question-answer pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic relevance between the Q basic question-answer pair data is greater than a first relevance threshold;
[0070] Obtain T basic question-answer pair data based on context information, and splice the T basic question-answer pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic relevance between the T basic question-answer pair data is less than the second relevance threshold, and the first relevance threshold is greater than or equal to the second relevance threshold;
[0071] The following conversation data and random conversation data are determined as spliced multi-round question-answer pair data.
[0072] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The sample generation module is used to reconstruct the basic question-answer pair data, and when obtaining the reconstructed multi-round question-answer pair data, specifically to perform the following operations:
[0073] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category labels contained in the question-answer pair category information set; the deep evolutionary data is associated with the evolutionary category labels;
[0074] Obtain context information corresponding to H basic question-answer pair data, and rewrite the mutually related basic question-answer pair data into breadth-evolved data based on the context information; the context information is used to indicate the semantic relevance between the H basic question-answer pair data;
[0075] Deep evolution data and broad evolution data are determined as reconstructed multi-round question-answer pair data.
[0076] In one possible implementation, when the sample generation module is used to rewrite the basic question-answer pair data into the deeply evolved data based on the evolutionary constraint information, it is specifically used to perform the following operations:
[0077] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into transition evolution data;
[0078] If the question-and-answer rounds of the transition evolution data are less than the question-and-answer round threshold, the transition evolution data will continue to be rewritten based on the evolution constraint information until the question-and-answer rounds of the transition evolution data after the data rewrite are greater than or equal to the question-and-answer round threshold, and the transition evolution data after the data rewrite will be determined as deep evolution data.
[0079] In a possible implementation, the training sample sequence includes a training sample set A i And training sample set A i+1 , training sample set A in the training sample sequence i In the training sample set A i+1 Before; the model training module is used to obtain input training samples from the sorted training sample set in the training sample sequence, and adjust the model parameters of the initial question-answering model in sequence through the input training samples obtained in sequence until the initial question-answering model training converges and the target question-answering model is obtained. Specifically, it is used to perform the following operations:
[0080] In the jth round of iterative training, in the training sample set A i Randomly obtain B first input training samples from the training set, adjust the model parameters of the initial question answering model through B first input training samples, and continue to use the training sample set A i+1 Randomly obtain C second input training samples in the training sample set, and adjust the model parameters of the initial question answering model through the C second input training samples until the model parameters of the initial question answering model are adjusted through the input training samples in the Nth training sample set in the training sample sequence, and the jth round of iterative training is determined to be completed; B and C are positive integers; j is a positive integer;
[0081] If the jth round of iterative training has been completed and the initial question-answering model meets the model convergence conditions, the initial question-answering model that meets the model convergence conditions will be determined as the target question-answering model;
[0082] If the j-th round of iterative training has been completed and the initial question-answering model does not meet the model convergence conditions, then through the training sample sequence, in the j+1-th round of iterative training, the model parameters of the initial question-answering model will continue to be adjusted in sequence through the input training samples obtained in sequence until the initial question-answering model training converges and the target question-answering model is obtained.
[0083] An embodiment of the present application provides a computer device, including: a processor, a memory, and a network interface;
[0084] The processor is connected to the memory and the network interface, wherein the network interface is used to provide data communication functions, and the memory is used to store computer programs. When the computer program is executed by the processor, the computer device executes the method provided in the embodiment of the present application.
[0085] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.
[0086] In one aspect, an embodiment of the present application provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in the embodiment of the present application.
[0087] The embodiment of the present application divides the question-answer pair data into categories through the question-answer pair category information set, and obtains the category labels corresponding to each question-answer pair data. Then, through the question-answer pair data and the category labels corresponding to the question-answer pair data, a training sample set and a sample set category label corresponding to the training sample set are generated. By sorting the training sample sets by learning difficulty, the model can first learn the features of the training sample sets with lower learning difficulty during the model training process, and then continue training with the training sample sets with higher learning difficulty, gradually helping the model learn complex features, thereby avoiding overfitting of the model in the training samples with higher difficulty, generalizing more complex and diverse question-answering scenarios, and improving the accuracy of the trained model. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0089] FIG1 is a schematic diagram of the structure of a blockchain network provided in an embodiment of the present application;
[0090] FIG2 is a first schematic diagram of a data processing scenario provided by an embodiment of the present application;
[0091] FIG3 is a flow chart of a data processing method according to an embodiment of the present application;
[0092] FIG4a is a second flow chart of a data processing method provided in an embodiment of the present application;
[0093] FIG4b is a second schematic diagram of a data processing scenario provided by an embodiment of the present application;
[0094] FIG5 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0095] FIG6 is a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0096] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0097] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0098] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0099] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye in identifying and measuring objects, and then further processes images to make them more suitable for human observation or transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and smart transportation.
[0100] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0101] Deep learning (DL) is the process of learning the inherent patterns and representational hierarchies of sample data. The information gained during this learning process is highly helpful in interpreting data such as text, images, and sounds. Its ultimate goal is to enable machines to acquire the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition that far surpass previous technologies.
[0102] The solutions provided in the embodiments of this application involve deep learning technology of artificial intelligence, which is specifically illustrated by the following embodiments.
[0103] Please refer to Figure 1, which is a schematic diagram of a network architecture provided by an embodiment of the present application. As shown in Figure 1, the network architecture may include a business server 100 and a terminal device cluster, and the terminal device cluster may include terminal device 10a, terminal device 10b, ..., terminal device 10n, wherein any terminal device in the terminal device cluster may have a communication connection with the business server 100, for example, there is a communication connection between terminal device 10a and the business server 100, and there is a communication connection between terminal device 10b and the business server 100, wherein the above-mentioned communication connection does not limit the connection method, and may be directly or indirectly connected by wired communication, or directly or indirectly connected by wireless communication, or by other methods, and the present application does not make any restrictions here.
[0104] Specifically, each terminal device in the terminal cluster shown in FIG1 may send a category classification request to the service server 100 . The category classification request may include question-answer pair data and a question-answer pair category information set.
[0105] Among them, question-answer data can be text data used to train large language models, including a series of questions and corresponding answers. Question-answer data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0106] The question-answer pair category information set may be an information set for labeling question-answer pair data, and may include multiple category labels. The category labels may be divided according to dimensions such as the learning difficulty of the question and answer, the field and subject of the question, and the type and length of the answer.
[0107] Taking terminal device 10a as an example, service server 100 can obtain a category classification request sent by terminal device 10a and, using a sample classification model in service server 100, classify the question-answer pair data and the question-answer pair category information set in the category classification request into categories, thereby obtaining category labels corresponding to the question-answer pair data. These category labels belong to the question-answer pair category information set. The sample classification model can be a converged large language model, which can use the question-answer pair category information set as a textual prompt for the question-answer pair data, classify the question-answer pair data into categories, and obtain corresponding category labels.
[0108] The business server 100 can generate a training sample set and sample category labels corresponding to the training sample set using the question-answer pair data and the category labels corresponding to the question-answer pair data, and sort the training sample set according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence.
[0109] It can be understood that by sorting the training sample set by learning difficulty, the model can be guaranteed to learn the features of training samples with lower learning difficulty during the model training process, and then gradually help the model learn complex features through training samples with higher learning difficulty, thereby achieving the same accuracy as using the entire training sample sequence. This can avoid overfitting of the model in training samples with higher difficulty, generalize more complex and diverse question-and-answer scenarios, and improve the efficiency of model training and the accuracy of the model.
[0110] The service server 100 can send a training sample sequence to the terminal device 10a. The terminal device 10a can sequentially obtain input training samples from the sorted training sample set within the training sample sequence and sequentially adjust the model parameters of the initial question-answering model using the sequentially obtained input training samples until the initial question-answering model in the terminal device 10a converges and a target question-answering model is obtained. The target question-answering model can be used to generate question-answering service results based on business questions.
[0111] It is understandable that the business server 100 and the terminal device 10a may also be integrated into one server, which may be called a model server, for example.
[0112] Please refer to Figure 2, which is a schematic diagram of a data processing scenario provided by an embodiment of the present application. As shown in Figure 2, the terminal device 10a can send a category division request to the service server 100, and the category division request can include question-answer pair data and question-answer pair category information set.
[0113] Among them, question-answer data can be text data used to train large language models, including a series of questions and corresponding answers. Question-answer data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0114] The question-answer pair category information set may be an information set for labeling question-answer pair data, and may include multiple category labels. The category labels may be divided according to dimensions such as the learning difficulty of the question and answer, the field and subject of the question, and the type and length of the answer.
[0115] Taking the terminal device 10a as an example, the business server 100 can obtain the category classification request sent by the terminal device 10a, and through the sample classification model in the business server 100, classify the question and answer pair data and the question and answer pair category information set in the category classification request to obtain the category label corresponding to the question and answer pair data.
[0116] The sample classification model in the business server 100 can use the question-answer pair category information set as a text prompt for the question-answer pair data, and classify the question-answer pair data. The classification steps can be: "The following is the category system: question-answer pair category information set. Please classify the following text "{question-answer pair data 1, question-answer pair data 2, ..., question-answer pair data N}" into the most relevant category tag in the above question-answer pair category information set, and output the result to the following structure {"category":result} in result'". The organization form of the result "result" can be {"question-answer pair data 1": NLP (Natural Language Processing) basic category tag, "question-answer pair data 2": copy generation category tag, ..., "question-answer pair data N": multi-round dialogue category tag}.
[0117] The business server 100 can obtain basic question-answer pair data (also known as conventional natural corpus) from the question-answer pair data, and divide the basic question-answer pair data into the training sample set of the first stage (basic question-answer pair data can include NLP basic question-answer pair data, copywriting generation question-answer pair data, domain-specific question-answer pair data, short multi-round dialogue data and logical reasoning question-answer pair data).
[0118] Among them, NLP basic question-and-answer data can be conversations such as regular translation tasks, text summarization, text rewriting, and part-of-speech tagging. NLP basic question-and-answer data focuses on covering text data with objective measurement indicators to judge the results. Copy generation question-and-answer data can be text data for copy generation question-and-answer data, which can be generation-type conversations such as writing comments, writing press releases, and brainstorming. Copy generation question-and-answer data is used to cover text data with open answers (the quality of answers will vary depending on the evaluator). The text data of domain professional question-and-answer data can be conversations involving keywords in various professional fields such as law, medicine, and physics. Short multi-round conversation data can be question-and-answer data with multiple question-and-answer rounds. Logical reasoning question-and-answer data can be reasoning conversations such as sorting comparison, comprehensive analysis, and paradox problems.
[0119] The category labels of the basic question-answer pair data may include NLP basic category labels (corresponding to the NLP basic question-answer pair data in the question-answer pair data, i.e., the NLP basic sample shown in FIG2 ), copy generation category labels (corresponding to the copy generation question-answer pair data in the question-answer pair data, i.e., the copy generation sample shown in FIG2 ), domain professional category labels (corresponding to the domain professional question-answer pair data in the question-answer pair data, i.e., the N domain professional sample shown in FIG2 ) and logical reasoning category labels (corresponding to the logical reasoning question-answer pair data in the question-answer pair data, i.e., the logical reasoning sample shown in FIG2 ). The business server 100 may generate a sample set category label corresponding to the training sample set of the first phase. For example, the training sample set of the first phase may include The training sample set consists of NLP basic question and answer pair data, the training sample set consists of copy generation question and answer pair data, the training sample set consists of domain professional question and answer pair data, and the training sample set consists of logical reasoning question and answer pair data. The sample set category labels corresponding to the training sample set in the first stage can include the NLP basic sample set category labels corresponding to the training sample set consisting of NLP basic question and answer pair data, the copy generation sample set category labels corresponding to the training sample set consisting of copy generation question and answer pair data, the domain professional sample set category labels corresponding to the training sample set consisting of domain professional question and answer pair data, and the logical reasoning sample set category labels corresponding to the training sample set consisting of logical reasoning question and answer pair data.
[0120] The business server 100 obtains progressive question-answer pair data (also referred to as progressive thought chain corpus) from the basic question-answer pair samples. For example, the progressive question-answer pair data may be obtained by gradually breaking down the steps of answering the answers in the basic question-answer pair data into multiple step rounds. The progressive question-answer pair data is determined as the second-stage training sample set (which may include progressive thought chain samples), and a progressive thought chain sample category label is generated. The progressive thought chain sample category label is determined as the sample set category label corresponding to the second-stage training sample set. For example, the progressive thought chain sample set category label may be determined as the sample set category label corresponding to the progressive thought chain sample.
[0121] The business server 100 can obtain the first multi-round question and answer pair data from the question and answer pair data (for example, the original data form can be question and answer pair data of multi-round dialogues). The business server 100 can perform data processing on the basic question and answer pair data through the evolutionary algorithm (it can be data splicing and data rewriting of the basic question and answer pair data) to obtain the second multi-round question and answer pair data (also referred to as multi-round evolutionary corpus). The business server 100 can determine the first multi-round question and answer pair data and the second multi-round question and answer pair data as the training sample set of the third stage. In addition, it can also generate a multi-round category label and determine the multi-round category label as the sample set category label corresponding to the training sample set of the third stage, for example, it can be the evolutionary multi-round sample set category label (corresponding to the medium evolutionary multi-round sample shown in Figure 2), the complex evolutionary multi-round sample set category label (corresponding to the complex evolutionary multi-round sample shown in Figure 2) and the GAtt sample set category label (corresponding to the GAtt sample shown in Figure 2).
[0122] The service server 100 may sort the training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a sorted training sample sequence, wherein the training sample sets in the training sample sequence may be arranged in order of learning difficulty from easy to difficult.
[0123] The service server 100 can send a training sample sequence to the terminal device 10a. The terminal device 10a can sequentially obtain input training samples from the sorted training sample set within the training sample sequence and sequentially adjust the model parameters of the initial question-answering model using the sequentially obtained input training samples until the initial question-answering model in the terminal device 10a converges and a target question-answering model is obtained. For example, the target question-answering model can be used to generate question-answering service results based on business questions.
[0124] In a certain iterative training of the initial question-answering model, the terminal device 10a can first obtain input training sample 1 from the training sample set corresponding to the first stage, adjust the model parameters of the initial question-answering model by input training sample 1, and then obtain input training sample 2 from the training sample set corresponding to the second stage, adjust the model parameters of the initial question-answering model by input training sample 2, and then obtain input training sample 3 from the training sample set corresponding to the third stage, adjust the model parameters of the initial question-answering model by input training sample 3, thereby realizing the sequential adjustment of the model parameters of the initial question-answering model by the input training samples obtained in sequence.
[0125] It is understood that the fine-tuning algorithm based on curriculum learning proposed in the embodiments of this application can improve the generalization ability of the target question-answering model and bring a better conversation experience. Based on these improvements, the product experience in the following different areas can be significantly improved:
[0126] 1. Customer Service: Targeted question-answering models can be used to create chatbots that can understand and answer customer questions, process refund and return requests, provide product information, etc. In addition, chatbots can also be used for telephone service, automatically answering calls and handling simple requests.
[0127] 2. Content creation: Targeted question answering models can generate various types of content, including blog posts, news reports, social media posts, etc. Targeted question answering models can generate content based on given topics or keywords, and can also provide writing suggestions, such as different ways to express a sentence or possible plots for a story.
[0128] 3. Education: Targeted question-answering models can be used to create personalized learning resources, such as generating exercises based on students’ learning progress and comprehension. They can also be used for online tutoring to understand and answer students’ questions, providing detailed explanations and feedback.
[0129] 4. Translation and language learning: Targeted question answering models can be used for real-time translation, understanding, and translating text in various languages. They can also be used for language learning, providing grammar and pronunciation suggestions to help learners understand and use new languages.
[0130] 5. Health Consultation: The targeted question-answering model can be used to provide basic health and medical consultation, such as explaining symptoms, providing advice on healthy living, and answering questions about medications. However, this is not a substitute for professional medical advice, but rather a supplementary resource.
[0131] 6. Personal Assistants: Targeted question-answering models can be used to create intelligent personal assistants that can understand and answer a variety of questions, such as checking the weather, setting reminders, and managing schedules. Targeted question-answering models can also be used to send emails and messages, perform online shopping, and more.
[0132] 7. Games: Targeted question-answering models can be used to create richer and deeper gaming experiences. For example, they can generate dialogue, drive non-player characters (NPCs), and create complex storylines. Furthermore, targeted question-answering models can be used to understand and respond to player input, providing a dynamic gaming experience.
[0133] The supervised fine-tuning of a large language model was performed using the training data arrangement method based on curriculum learning proposed in the embodiments of this application. Compared with the data splicing method without curriculum learning, the LLM (Large Language Model) trained by this invention showed improvements in the five major dimensions of NLP basic capabilities, multi-round dialogue capabilities, domain application capabilities, reasoning capabilities, and text generation capabilities. Among them, the reasoning and multi-round dialogue capabilities were relatively improved, with the effect increased by more than 10%.
[0134] The embodiment of the present application divides the question-answer pair data into categories through the question-answer pair category information set, and obtains the category labels corresponding to the question-answer pair data. Then, through the question-answer pair data and the category labels corresponding to the question-answer pair data, a training sample set and a sample set category label corresponding to the training sample set are generated. The number of training samples can be increased and the diversity of training samples can be increased. By sorting the training sample set by learning difficulty, it can be ensured that the model learns the features of training samples with lower learning difficulty during model training, and then gradually helps the model learn complex features through training samples with higher learning difficulty, thereby achieving the same accuracy as using the entire training sample sequence, avoiding overfitting of the model in training samples with higher difficulty, generalizing more complex and diverse question-answering scenarios, and improving the efficiency of model training and the accuracy of the model.
[0135] Please refer to Figure 3, which is a flowchart of a data processing method provided in an embodiment of the present application. The data processing method can be executed by a computer device, which can be a model server integrated with the business server 100 and the terminal device 10a shown in Figure 1. The following description will take the data processing method executed by a computer device as an example. The data processing method can include at least the following steps S101-S104:
[0136] Step S101: obtain M question-answer pair data and a question-answer pair category information set, classify the M question-answer pair data into categories based on the question-answer pair category information set, and obtain category labels corresponding to the M question-answer pair data; the question-answer pair category information set includes the category labels; M is a positive integer;
[0137] Specifically, the computer device can obtain question-answer pair data and question-answer pair category information sets.
[0138] Among them, question-answer data can be text data used to train large language models, including a series of questions and corresponding answers. Question-answer data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0139] The question-answer category information set can be an information set for labeling question-answer data, which can be divided into multiple category labels. The category labels can be divided according to the learning difficulty of the question and answer, the field of the question (such as the copywriting field or the professional field of a certain discipline) and the theme, the answer type (such as the logical reasoning type or the incremental thinking type) and the length (such as single-round answer or multi-round answer), etc.
[0140] The computer device can classify the question and answer pair data into categories using the question and answer pair category information set to obtain category labels corresponding to the question and answer pair data.
[0141] For example, the above-mentioned question-answer pair category information set can be used to classify "{question-answer pair data 1, question-answer pair data 2, ..., question-answer pair data N}" into the most relevant category label, and the division result can be {"question-answer pair data 1": NLP basic category label, "question-answer pair data 2": copywriting generation category label, ..., "question-answer pair data N": multi-round dialogue category label}.
[0142] Category tags corresponding to question-answer data may include NLP basic category tags, copywriting generation category tags, domain professional category tags, logical reasoning category tags, progressive thinking chain category tags, and multi-round dialogue category tags.
[0143] The text data for basic NLP question-and-answer data can include conversations related to common translation tasks, text summarization, text rewriting, and part-of-speech tagging. Basic NLP question-and-answer data focuses on text data with objective metrics to determine the results. For example, basic NLP question-and-answer data can be in the form of text: "Question: Classify a given text. Please classify the following sentence as positive, negative, or neutral: This movie is very good. Answer: This movie is very good, positive."
[0144] The text data for generating question-and-answer data can include generative conversations such as writing reviews, press releases, and brainstorming. This data is used to cover text data with open-ended responses (the quality of responses varies depending on the evaluator). For example, the text for generating question-and-answer data could be something like, "Question: I want to write a one-sentence review of a science fiction novel I like. What would you say? Title: "Science Fiction A"; Author: Author A. Answer: Author A's "Science Fiction A," with its unique science fiction concepts, worldview that transcends time and space, and magnificent imagination, presents us with a breathtaking cosmic epic."
[0145] The text data for domain-specific question-and-answer data can include conversations involving keywords from various professional fields, such as law, medicine, and physics. For example, domain-specific question-and-answer data can be in the form of text such as, "Question: Please find the following medications that are diuretics: A. Amoxicillin; B. Dexamethasone; C. Furosemide; D. Aspirin. Answer: The diuretic is: C. Furosemide."
[0146] The text data of logical reasoning question and answer data can be reasoning dialogues such as sorting comparison, comprehensive analysis, and paradox problems. For example, the logical reasoning question and answer data can be "Question: A unit organizes quality development training for employees, and employees can voluntarily sign up to participate. Lao Zhang met a new employee Xiao Li and talked about this. Lao Zhang reminded Xiao Li, "The unit organizes quality development training, hurry up and sign up." Xiao Li said, "I haven't finished the work on hand, so I don't need to sign up." Which of the following except one can be the assumption contained in Xiao Li's answer? A. If the work is handled, I have to sign up for quality development training; B. As long as my work is not handled, I don't have to participate in quality development training; C. Anyone who signs up for quality development training has finished the work; D. Only those who have finished the work can sign up for quality development training. Answer: First, let's analyze Xiao Li's answer: "I haven't finished the work on hand, so I don't need to sign up." Next, let's analyze the relationship between each option and Xiao Li's answer. Xiao Li based on "work is not handled" The conclusion is "no need to report". In order for this conclusion to be valid, it is necessary to find the relationship between "not processed" and "no need to report", and it is necessary to get "no need to report" from "not processed". Options B, C and D can all be expressed as: "The work is not processed, do not sign up", these options can all be used as bridging premises, but this question requires a negative answer, so BCD are excluded. Option A can be expressed as: "No need to report, it is not processed", the direction of the bridge here is opposite, so option A is not the premise of Xiao Li's answer. According to the requirements of the question, the correct answer is A." in text form.
[0147] Among them, regarding the contents of the progressive thinking chain question and answer pair data and the multi-round dialogue question and answer pair data, please refer to the specific description of steps S203 to S205 corresponding to the embodiment of Figure 4a below.
[0148] It can be understood that the text data in the question-and-answer data can cross-cover multiple fields or topics, and there can be multiple related category tags. When the computer device categorizes the question-and-answer data, it can classify the question-and-answer data into the most relevant D category tags. The D category tags can be three category tags, or the degree of correlation between the question-and-answer data and the category tag is greater than a certain preset threshold, and the question-and-answer data can be classified into the category tag. The embodiment of the present application does not limit this.
[0149] Step S102: Based on the M question-answer pairs and the class labels corresponding to the M question-answer pairs, generate N training sample sets and the sample set class labels corresponding to the N training sample sets; N is a positive integer;
[0150] Specifically, the computer device may generate a training sample set and a sample set category label corresponding to each training sample set based on the question-answer pair data and the category label corresponding to each question-answer pair data. The input training sample in each training sample set is associated with at least one question-answer pair data from the M question-answer pair data.
[0151] The computer device can determine a single round of question-and-answer data (a dialogue round is text data with one question and one answer) of a certain basic field or topic as the basic question-and-answer data. The basic field or topic can be an NLP basic dialogue, a copy generation dialogue, a domain professional dialogue, and a logical reasoning dialogue. The embodiments of this application do not limit this.
[0152] The computer device can divide the basic question-answer pair data into the training sample set of the first stage, that is, the input training samples (also referred to as training samples) in the training sample set of the first stage are the basic question-answer pair data, and the sample set category labels corresponding to the training sample set of the first stage are generated through the category labels of the basic question-answer pair data. The category labels of the basic question-answer pair data can be the same as the category labels of the sample set. Question-answer pair data with a dialogue turn number below a certain preset threshold can also be divided into the training sample set of the first stage. For example, the training sample set corresponding to a short multi-round dialogue with a dialogue turn number less than or equal to 8 rounds can be divided into the training sample set of the first stage. The training sample set of the first stage can include a training sample set of NLP basic samples, a training sample set of copywriting samples, a training sample set of domain professional samples, a training sample set of short multi-round dialogues, and a training sample set of logical reasoning samples.
[0153] The computer device can obtain progressive question-answer pair data from the basic question-answer pair samples. For example, the basic question-answer pair data can be formatted, or the steps of answering the answers in the basic question-answer pair data can be gradually broken down into multiple step rounds to obtain progressive question-answer pair data. The computer device can determine the progressive question-answer pair data as the training sample set for the second stage, that is, the input training samples (also referred to as training samples) in the training sample set for the second stage are the progressive question-answer pair data obtained after corresponding processing of the basic question-answer pair data, and then directly generate the sample set category label corresponding to the training sample set for the second stage. Therefore, the input training samples in the training sample set for the second stage all have the sample set category label.
[0154] The computer device can obtain the first multiple rounds of question and answer pair data from the question and answer pair data (for example, the original data form can be question and answer pair data of multiple rounds of dialogues), and the computer device can perform data processing on the basic question and answer pair data through an evolutionary algorithm, for example, it can perform data splicing and data rewriting on the basic question and answer pair data to obtain the second multiple rounds of question and answer pair data. The computer device can determine the first multiple rounds of question and answer pair data and the second multiple rounds of question and answer pair data as the training sample set of the third stage, that is, the input training samples (also referred to as training samples) in the training sample set of the third stage are the directly obtained first multiple rounds of question and answer pair data and the second multiple rounds of question and answer pair data obtained after corresponding processing of the basic question and answer pair data, and then directly generate the sample set category label corresponding to the training sample set of the third stage. Therefore, the input training samples in the training sample set of the third stage all have the sample set category label.
[0155] Step S103: sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence;
[0156] Specifically, the computer device can sort the training sample sets from easy to difficult according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence. The training sample sequence can include multiple training sample sets, and the sorting order can be {training sample set of the first stage, training sample set of the second stage, training sample set of the third stage}. Learning difficulty can also be understood as the difficulty of the initial question-answering model training to learn input training samples with sample set category labels. If the initial question-answering model has difficulty learning the question-answering intent and logic of input training samples with a certain sample set category label, it means that the learning difficulty of this type of input training sample is high, that is, the learning difficulty of this type of sample set category label is relatively high.
[0157] The learning difficulty can be determined based on factors such as the complexity of the training sample's data composition, the magnitude of the loss value generated during training, the answer error rate, and the amount, diversity, and quality of the training sample data. For example, for a training sample set with a multi-round dialogue category label, the input training samples contained therein are all complex data from multi-round dialogues, meaning the data composition is highly complex. Therefore, it can be determined that this type of training sample set has a high learning difficulty. For another example, for a training sample set with a basic NLP category label, when training the input training samples contained therein, the loss value generated is generally small, and the predicted answer error rate is low, indicating that this type of training sample set has a low learning difficulty.
[0158] In the training sample sets of each stage, the computer device can also sort the training sample sets according to the learning difficulty corresponding to the sample set category labels. For example, in the training sample sets of the first stage, the sorting order can be {training sample set of NLP basic samples, training sample set of copywriting generation samples, training sample set of domain professional samples, training sample set of short multi-round dialogues, and training sample set of logical reasoning samples}.
[0159] Step S104: In the training sample sequence, input training samples are obtained in sequence from the sorted training sample set, and the model parameters of the initial question-answering model are adjusted in sequence through the input training samples obtained in sequence until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions.
[0160] Specifically, in the training sample sequence, the computer device can sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-answering model through the sequentially obtained input training samples.
[0161] For example, the training sample sequence can be used as the input dataset for an epoch (training cycle) of the initial question-answering model training. The training sample sets corresponding to each stage in the training sample sequence can be a batch (training batch) of the initial question-answering model training. An epoch is the process of inputting the training sample sequence into the initial question-answering model to complete a forward calculation and backpropagation. An epoch can also be called an iterative training.
[0162] In an epoch, the computer device can obtain input training samples from the training sample set of the first stage, the training sample set of the second stage, and the training sample set of the third stage in sequence. For example, it can obtain input training sample 1 from the training sample set of the first stage, adjust the model parameters of the initial question-answering model by input training sample 1, and complete a training batch of the initial question-answering model training, and then obtain input training sample 2 from the training sample set of the second stage, adjust the model parameters of the initial question-answering model by input training sample 2, and then obtain input training sample 3 from the training sample set of the third stage, adjust the model parameters of the initial question-answering model by input training sample 3, and complete a training cycle of the initial question-answering model training.
[0163] If the initial question-answering model meets the convergence conditions indicated by the model evaluation indicators in this training cycle (iterative training), the trained and converged initial question-answering model can be determined as the target question-answering model. The target question-answering model can be used to generate question-answering service results based on business questions. The target question-answering model can also be used as a general question-answering model that has been fine-tuned through a training sample sequence. The general question-answering model can then be trained using question-answering texts from specific fields or topics, so that the trained question-answering model can generate better question-answering service results for business questions in specific fields or topics.
[0164] Among them, the convergence condition indicated by the model evaluation index can be a performance metric of the model, such as accuracy, precision, recall, precision-recall, etc., which is not limited in this application.
[0165] The embodiment of the present application divides the question-answer pair data into categories through the question-answer pair category information set, and obtains the category labels corresponding to the question-answer pair data. Then, through the question-answer pair data and the category labels corresponding to the question-answer pair data, a training sample set and a sample set category label corresponding to the training sample set are generated. Since the training sample set can include input training samples expanded based on the question-answer pair data, the number of training samples can be increased and the diversity of training samples can be increased. By sorting the training sample set by learning difficulty, it can be ensured that the model learns the features of training samples with lower learning difficulty during the model training process, and then gradually helps the model learn complex features through training samples with higher learning difficulty, thereby avoiding overfitting of the model in training samples with higher difficulty, generalizing more complex and diverse question-answering scenarios, and improving the efficiency of model training and the accuracy of the model.
[0166] Please refer to Figure 4a, which is a second flow chart of a data processing method provided in an embodiment of the present application. The data processing method can be executed by a computer device, which can be a model server integrated with the business server 100 and the terminal device 10a shown in Figure 1. The following description will take the execution of the data processing method by a computer device as an example. The data processing method can include at least the following steps S201-S207:
[0167] Step S201, obtaining M question-answer pair data and question-answer pair category information sets;
[0168] For details, please refer to the specific content of step S101 corresponding to the embodiment of Figure 3 above, and the embodiment of this application will not be described in detail here.
[0169] In step S202, the M question-answer pair data and the question-answer pair category information set are input into the sample partitioning model. In the sample partitioning model, a keyword vector is generated based on the question-answer pair category information set, and features are extracted from the M question-answer pair data to obtain the feature word vectors corresponding to the M question-answer pair data respectively; based on the feature word vectors and the keyword vectors, the category labels corresponding to the M question-answer pair data are determined.
[0170] Specifically, the computer device can input the question-answer pair data and the question-answer pair category information set into the sample segmentation model.
[0171] Among them, the sample division model can be a converged large language model, which can use the question-answer pair category information set as a text prompt for the question-answer pair data, divide the question-answer pair data into categories, and obtain corresponding category labels.
[0172] In the sample partitioning model, keyword vectors can be generated based on the question-answer pair category information set, and features can be extracted from the question-answer pair data to obtain the feature word vectors corresponding to the question-answer pair data.
[0173] The feature extraction process of the question-answer pair data by the computer device can be: filtering business stop words from the M question-answer pair data, performing text segmentation on the filtered M question-answer pair data to obtain P question-answer text fragments; P is a positive integer greater than or equal to M; performing part-of-speech tagging on the P question-answer text fragments to obtain P question-answer part-of-speech information, and generating feature word vectors corresponding to the M question-answer pair data based on the P question-answer text fragments and the P question-answer part-of-speech information.
[0174] Specifically, the computer device can filter out business stop words from the Q&A pair data. Business stop words refer to words or punctuation marks that are common in natural language texts but have no specific meaning in semantic analysis. For example, in Chinese texts, business stop words can include: "de", "le", "shi", "zai", "you", etc. In English texts, business stop words can include: "the", "a", "an", "in", etc.
[0175] The computer device can perform text segmentation (also known as word segmentation) on the training text set after removing business stop words. Text segmentation can be achieved through rule matching or probability matching to obtain several Q&A text fragments corresponding to each Q&A pair data.
[0176] The computer device can perform part-of-speech tagging on the Q&A text fragments to obtain the corresponding Q&A part-of-speech information. Part-of-speech tagging can be a grammatical part-of-speech classification of the Q&A text fragments. For example, it can be classified into nouns, verbs, adjectives, adverbs, auxiliary words, etc. By performing part-of-speech tagging, the grammatical structure and semantic information of the sentence can be better understood, thereby improving the accuracy and efficiency of natural language processing tasks. The part-of-speech tagging method can be a rule-based method, a statistic-based method, a deep learning-based method, etc., and the embodiments of this application do not limit this here.
[0177] Based on the Q&A text fragments and the corresponding Q&A part-of-speech information, the computer device generates a feature word vector for the Q&A pair data. The computer device can determine the class target tags corresponding to the Q&A pair data respectively through the matching degree between the feature word vector and the keyword vector.
[0178] Step S203: Obtain the basic Q&A pair data from the M Q&A pair data, divide the basic Q&A pair data into the training sample set of the first stage, and determine the class target tag corresponding to the basic Q&A pair data as the sample set class target tag corresponding to the training sample set of the first stage;
[0179] Specifically, the computer device can determine the single-round Q&A pair data (text data with one question and one answer in the conversation round) of a certain basic domain or topic as the basic Q&A pair data. The basic domain or topic can be NLP basic conversations, copywriting generation conversations, domain professional conversations, and logical reasoning conversations, and the embodiments of this application do not limit this here.
[0180] The computer device can divide the basic Q&A pair data into the training sample set of the first stage and generate the sample set class target tag corresponding to the training sample set of the first stage through the class target tag of the basic Q&A pair data.
[0181] Optionally, the computer device can also divide the question and answer pair data of short conversations into the training sample set of the first stage, and the specific process can be: obtaining context information of the basic question and answer pair data; the context information is used to indicate the semantic correlation between the basic question and answer pair data; according to the context information, S mutually related basic question and answer pair data are spliced into short conversation training samples, the short conversation training samples are added to the short conversation training sample set, and the short conversation label is determined as the sample set category label corresponding to the short conversation training sample set; S is a positive integer.
[0182] Specifically, contextual information can be used to indicate the semantic correlation between basic question-answer pair data. It can be the text content about the text and task that the model needs to understand (for example, in sentiment analysis tasks, it can be a series of sentences with sentiment labels, such as "This product is very easy to use" and "This product is very bad", etc.), contextual relationships (for example, in named entity recognition tasks, it can be the relationship between entities in the text, such as keywords, place names, time, etc.), semantic information (for example, in part-of-speech tagging tasks, it can be the part of speech of each word), and other related information.
[0183] The computer device can use context information to combine multiple interrelated basic question-answer pairs into short conversation training samples. Short conversation training samples can be question-answer pairs with a conversation turn length of 8 or less, where one question and one answer constitute one turn.
[0184] A short conversation training sample could be: "Question 1: What is Actor C's zodiac sign? Answer 1: Actor C is a Libra. Question 2: You're awesome! You know so much. Answer 2: I also know that Actor C won the Best Actress award at the Asian Film Awards! Question 3: Actor C is my idol, and her acting is superb. Answer 3: She was truly born for acting, and she has also won the Best Actress award at the Chinese Film Media Awards. Question 4: I really like the movies she acted in. Answer 4: I recommend watching "Li Mi's Guess" starring Actor C. Actor C dedicated great passion and sincerity to this film, bringing the audience into the situation of "loving until death", which makes people think about many things. A city, about love, about gains and losses. Question 5: How is the reputation of this movie? Answer 5: The reputation of this movie is generally good. "Movie C" is rated 8.1 points (out of 10 points) and has been reviewed by more than 370,000 people. Many viewers think that the film has a compact plot, sincere emotions, and successful character development. Actor C's performance, in particular, has received high praise. At the same time, the film's focus on real life and exploration of human nature have also been well received. Of course, some viewers feel that the plot of the film is a bit too bizarre or the ending is a bit abrupt. In general, "Movie C" is a movie worth watching. Question 6: Then I'll go see it now. Answer 6: "Okay."
[0185] The computer device can add the short conversation training sample to the short conversation training sample set, determine the short conversation label as the sample set category label corresponding to the short conversation training sample set, and divide the short conversation training sample set into the training sample set of the first stage.
[0186] It can be understood that the training sample set of the first stage can include the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain professional samples, the training sample set of short multi-round dialogues, and the training sample set of logical reasoning samples.
[0187] Step S204: Obtain progressive question-answer pair data from the M question-answer pair data, determine the progressive question-answer pair data as the training sample set for the second stage, generate progressive thinking chain sample category labels, and determine the progressive thinking chain sample category labels as the sample set category labels corresponding to the training sample set for the second stage; the learning difficulty of the progressive thinking chain sample category labels is greater than the learning difficulty of the category labels corresponding to the basic question-answer pair data;
[0188] Specifically, the computer device obtains progressive question-answer pair data from the basic question-answer pair samples. For example, the basic question-answer pair data can be formatted, and the steps of solving the answer in the basic question-answer pair data can be gradually broken down into multiple step rounds to obtain progressive question-answer pair data.
[0189] For example, question-and-answer data B may be “Question: Xiao Ming has 10 chicken legs at home. He wants to divide them among 3 people. How many chicken legs does each person take? How many chicken legs are left? Answer: We need to find a number so that the 10 chicken legs can be divided evenly among 3 people. Divide 10 chicken legs by 3 people, and we get 10 / 3=3 with a remainder of 1. Therefore, each person can get 3 chicken legs. In the end, there is 1 chicken leg left. The answer is 3 for each person, and 1 chicken leg is left. The computer device can gradually decompose the steps of question-and-answer data B into multiple step rounds to obtain decomposition steps. The decomposition steps may be “Question 1: Xiao Ming has 10 chicken legs at home. He wants to divide them among 3 people. How many chicken legs does each person take? How many chicken legs are left? Answer 1: Step 1: We need to find a number so that the 10 chicken legs can be divided evenly among 3 people. Question 2: Xiao Ming has 10 chicken legs at home. He wants to divide them among 3 people. How many chicken legs does each person take? How many chicken legs are left? Answer 2: Step 1: We need to find a number that allows 10 chicken legs to be divided evenly between 3 people. Step 2: Divide 10 chicken legs by 3 people, and we get 10 / 3 = 3 with a remainder of 1. Question 3: Xiao Ming has 10 chicken legs. He wants to divide them among 3 people. How many chicken legs does each person take? How many chicken legs are left over? Answer 3: Step 1: We need to find a number that allows 10 chicken legs to be divided evenly between 3 people. Step 2: Divide 10 chicken legs by 3 people, and we get 10 / 3 = 3 with a remainder of 1. Step 3: Therefore, each person gets 3 chicken legs. Question 4: Xiao Ming has 10 chicken legs. He wants to divide them among 3 people. How many chicken legs does each person take? How many chicken legs are left over? Answer 4: Step 1: We need to find a number so that 10 chicken legs can be divided equally among 3 people; Step 2: Divide 10 chicken legs by 3 people, and we get 10 / 3 = 3 with a remainder of 1; Step 3: Therefore, each person can get 3 chicken legs; Step 4: At the end, there is 1 chicken leg left. The answer is 3 chicken legs per person, and 1 chicken leg is left.
[0190] It can be understood that the disassembly step can include 4 question and answer rounds. The computer device can ignore question 2, question 3, and question 4, and determine question 1, answer 1, answer 2, answer 3, and answer 4 as progressive question and answer pair data B. The specific method of step disassembly can be through Stepwise formatting, which is not limited in this embodiment of the present application.
[0191] The computer device can determine the progressive question-answer pair data as the training sample set of the second stage (the training sample set of the second stage may include progressive thinking chain samples), generate a progressive thinking chain sample category label, and determine the sample set category label corresponding to the training sample set of the second stage as the progressive thinking chain sample category label.
[0192] Understandably, progressive question answering requires analyzing multiple conditions in the question and gradually reasoning to increase accuracy. To this end, we collect samples of progressive thought chains for these questions, including a series of intermediate reasoning steps. Compared to using a single calculation, this format is more clear and standardized. Templated and standardized responses also stimulate the model's thinking when encountering similar questions, namely, how to flexibly apply a vast amount of prior knowledge and integrate information to accurately answer questions.
[0193] Step S205: obtain the first multi-round question-answer pair data from the M question-answer pair data, generate the second multi-round question-answer pair data based on the basic question-answer pair data, determine the first multi-round question-answer pair data and the second multi-round question-answer pair data as the training sample set of the third stage, generate multi-round category labels, and determine the multi-round category labels as the sample set category labels corresponding to the training sample set of the third stage; the learning difficulty corresponding to the multi-round category labels is greater than the learning difficulty of the progressive thinking chain sample category labels.
[0194] Specifically, the computer device may obtain a first plurality of rounds of question-answer pair data from the question-answer pair data. The first plurality of rounds of question-answer pair data may be original question-answer pair data having more than eight question-answer rounds.
[0195] The computer device can generate second multi-round question and answer pair data through the basic question and answer pair data, and the process can be: splicing the basic question and answer pair data to obtain spliced multi-round question and answer pair data; reconstructing the basic question and answer pair data to obtain reconstructed multi-round question and answer pair data; determining the spliced multi-round question and answer pair data and the reconstructed multi-round question and answer pair data as the second multi-round question and answer pair data.
[0196] Specifically, the computer device may process the basic question-answer pair data through an evolutionary algorithm, for example, it may perform data splicing and data rewriting on the basic question-answer pair data to obtain a second round of question-answer pair data.
[0197] The process of data splicing of basic question and answer pair data by a computer device can be: obtaining context information corresponding to H basic question and answer pair data; the context information is used to indicate the semantic correlation between the H basic question and answer pair data; obtaining Q basic question and answer pair data according to the context information, and splicing the Q basic question and answer pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation between the Q basic question and answer pair data is greater than a first correlation threshold; obtaining T basic question and answer pair data according to the context information, and splicing the T basic question and answer pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation between the T basic question and answer pair data is less than a second correlation threshold, and the first correlation threshold is greater than or equal to the second correlation threshold; and determining the follow-up dialogue data and the random dialogue data as spliced multi-round question and answer pair data.
[0198] Specifically, the computer device can obtain Q basic question-answer pair data through context information and splice the Q basic question-answer pair data into follow-up dialogue data. Q is a positive integer less than or equal to H, and Q can be a positive integer greater than 8. The semantic relevance between the Q basic question-answer pair data is greater than a first relevance threshold. The first relevance threshold can indicate that the question-answer pair data have the same attention keywords, that is, the topics between the two rounds of dialogue are related. For example, the computer device can splice 14 basic question-answer pair data into a follow-up dialogue data through context information, and the number of dialogue rounds in the follow-up dialogue data can be 14.
[0199] For example, the following dialogue data can be: "Question 1: From now on you need to play the role of E all the time, and I hope you can act like him. I hope you can answer and respond in the tone, manner and vocabulary of character E. Do not write any explanation, just answer like character E. You must understand all the knowledge of character E. After that, please be sure to answer in accordance with the personality, preferences, speaking style and language of the character you are playing. Where did you first meet Harry Potter and Ron Weasley? Answer 1: I first met character E on the train, and then the three of us became friends. Question 2: How did you view character E in that situation? Answer 2: When I first met character E, I must admit that I was full of curiosity about him. After all, his name is already a household name. However, seeing him, and Getting to know him is a different matter. I didn't immediately warm to him. In fact, I initially thought he might be a bit of a megalomaniac, but as I got to know him, I realized that Character E is a brave, selfless, and kind friend. Every risk he takes is to protect the people he loves. He doesn't exploit his fame; instead, he hopes for a peaceful life. So, I slowly began to respect and admire Character E. Question 3: What qualities do you particularly like about Character E as a friend? Answer 3: I've always been impressed by Character E's determination and courage. No matter what difficulties he faces, he never gives up. He is loyal, upright, and always willing to do the right thing for others. (Omitted the 10 middle rounds of dialogue) Question 14: Who are you? Answer 14: I am Character E.
[0200] The computer device can obtain T basic question-answer pair data through context information and splice the T basic question-answer pair data into random dialogue data.
[0201] Here, T is a positive integer less than or equal to H, and T can be greater than 8. The semantic relevance between the T basic question-answer pairs is less than the second relevance threshold, and the first relevance threshold is greater than or equal to the second relevance threshold. That is, the semantic relevance of the basic question-answer pairs required for splicing in the follow-up conversation data is greater than that of the basic question-answer pairs required in the random conversation data. In the random conversation data, the computer device can control the content type of the spliced samples to gradually diversify and gradually increase the length of the conversation. The basic question-answer pairs can be similar conversations with large topic jumps. This process can also randomly splice consecutive similar basic question-answer pairs, which can be used to increase the anti-interference ability of the model conversation.
[0202] For example, a computer device may combine 12 basic question-answer pair data into a random conversation data through context information, and the number of conversation rounds of the random conversation data may be 12 rounds.
[0203] Random conversation data may be "Question 1: Please arrange the following actions in the correct order: open the door, take off shoes, hang up the coat, and wash hands. Answer 1: This order may vary from person to person, but it may usually be like this: 1. Open the door; 2. Take off shoes; 3. Hang up the coat; 4. Wash hands. Question 2: What is a dynamic agent? What are its applications? Answer 2: A dynamic agent is a proxy class that is dynamically generated at runtime: Applications of dynamic agents include data query, backend of a test framework, Java annotation object acquisition, etc. Question 3: Which has a larger volume, a 2-ton monitor or a 2-ton feather? Answer 3: The density of a feather is less than that of a monitor. Therefore, under the same weight, the volume of a feather is larger than that of a monitor. Therefore, the volume of a 2-ton feather is larger. (Omit the middle 8 rounds of conversation) Question 12: Which has a larger volume, a 7-kilogram laptop or 7-kilogram toilet paper? Answer 12: The density of toilet paper is less than that of a laptop. Therefore, under the same weight, the volume of toilet paper is larger than that of a laptop. Therefore, the volume of 7-kilogram toilet paper is larger."
[0204] The computer device can determine the following dialogue data (also known as complex evolution multi-round samples) and the random dialogue data (also known as medium evolution multi-round samples) as spliced multi-round question and answer pair data, divide the spliced multi-round question and answer pair data into training samples corresponding to the third stage, determine the sample set category label of the following dialogue data as the complex evolution sample set category label, and determine the sample set category label of the random dialogue data as the medium evolution sample set category label.
[0205] The process of data reconstruction of basic question and answer pair data (the number of basic question and answer pair data can be H, where H is a positive integer) by a computer device can be: based on evolutionary constraint information, rewriting the H basic question and answer pair data into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category labels contained in the question and answer pair category information set; the deep evolutionary data is associated with the evolutionary category labels; obtaining contextual information corresponding to the H basic question and answer pair data, and rewriting the mutually related basic question and answer pair data into broad evolutionary data based on the contextual information; the contextual information is used to indicate the semantic correlation between the H basic question and answer pair data; and determining the deep evolutionary data and the broad evolutionary data as reconstructed multi-round question and answer pair data.
[0206] Specifically, the computer device can rewrite the basic question-answer pair data into deep evolution data based on the evolution constraint information.
[0207] Among them, the evolution constraint information is used to indicate the evolution category labels contained in the question-answer pair category information set. For example, a computer device can perform deep data rewriting on the domain-specific question-answer pair data F and the copywriting-generated question-answer pair data G to obtain deep evolution data K.
[0208] The question-answer data F can be "Question: Please find out which of the following drugs are diuretics: A. Amoxicillin; B. Dexamethasone; C. Furosemide; D. Aspirin. Answer: The diuretic is: C. Furosemide.", and the question-answer data G can be "Question: What are the medicinal effects of diuretics? Answer: They have diuretic effects, lower blood pressure, improve heart failure, and treat renal insufficiency."
[0209] The deep evolution data K obtained by deep data rewriting can be "Question 1: What are the diuretics? Answer 1: Furosemide. Question 2: What are the medicinal effects of furosemide? Answer 2: Furosemide has diuretic effects, lowers blood pressure, improves heart failure, and treats renal insufficiency."
[0210] Optionally, when the question and answer rounds of the deep evolution data K obtained by data rewriting are less than the question and answer round threshold, the computer device can continue to rewrite the deep evolution data K. The computer device can rewrite the basic question and answer data into transition evolution data based on the evolution constraint information. If the question and answer rounds of the transition evolution data are less than the question and answer round threshold, the transition evolution data will continue to be rewritten based on the evolution constraint information until the question and answer rounds of the transition evolution data after the data rewrite are greater than or equal to the question and answer round threshold, and the transition evolution data after the data rewrite is determined to be deep evolution data.
[0211] The computer device can obtain contextual information corresponding to the basic H question-answer pairs and use this contextual information to rewrite the interrelated basic question-answer pairs into breadth-evolved data. For example, domain-specific question-answer pair data F and copywriting-generated question-answer pair data G can be rewritten into breadth-evolved data L.
[0212] The breadth evolution data L obtained by rewriting the breadth data can be "Question 1: What are the medicinal effects of amoxicillin, dexamethasone, furosemide and aspirin? Answer 1: The medicinal effect of amoxicillin is to treat bacterial infections,..., the medicinal effect of furosemide is to have a diuretic effect, lower blood pressure, improve heart failure, and treat renal insufficiency."
[0213] The computer device can determine the deep evolution data and the broad evolution data as reconstructed multi-round question-answer pair data (also referred to as GAtt samples), divide the reconstructed multi-round question-answer pair data into the training sample set of the third stage, and determine the sample set category label of the reconstructed multi-round question-answer pair data as the GAtt sample set category label.
[0214] As you can understand, breadth evolution aims to enhance the topic coverage, skill coverage, and overall dataset diversity of question-answering data. Deep evolution gradually increases the difficulty, ensuring that the generated instructions are challenging. The deeply evolved data obtained through deep evolution and the broadly evolved data obtained through breadth evolution can optimize model performance and robustness during training.
[0215] Step S206: sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence;
[0216] Specifically, the computer device can sort the training sample set according to the learning difficulty corresponding to the sample set category label to obtain a training sample sequence. The sorting order of the training sample sequence can be {training sample set of the first stage: {training sample set of NLP basic samples, training sample set of copywriting generation samples, training sample set of domain professional samples, training sample set of short multi-round dialogues, training sample set of logical reasoning samples}, training sample set of the second stage: {training sample set of progressive thinking chain samples}, training sample set of the third stage: {training sample set of medium evolution multi-round samples, training sample set of complex evolution multi-round samples, training sample set of GAtt samples}.
[0217] Optionally, when short multi-round dialogue samples are composed of more complex question-answer pair data, the learning difficulty of short multi-round dialogue samples is greater than that of logical reasoning samples. The course arrangement of question-answer pair data can adopt a course arrangement method that combines coarse-grained and fine-grained methods. Overall, it can be divided into four stages of course learning. The first stage is regular single-round sample learning, the second stage is simple (fewer rounds) multi-round dialogue sample learning, the third stage is progressive thinking chain sample learning, and the fourth stage is complex (more rounds) multi-round dialogue sample learning. In the multi-round dialogue field of the second and fourth stages, the difficulty of multi-round dialogue samples can be defined from three dimensions, namely topic richness, length, and depth. According to the measurement dimension of the difficulty of multi-round dialogue, management is carried out from the source of data collection.
[0218] In the above-mentioned course arrangement method combining coarse-grained and fine-grained methods, the three-stage training sample sequence can be further subdivided into four stages, that is, the short multi-round dialogue samples in the first stage of the three-stage training sample sequence can be used as the training sample set of the second stage alone, and the short multi-round dialogue samples can be determined as the training stage after the first stage, so as to obtain the four-stage training sample sequence. The sorting order of the four-stage training sample sequence can be {training sample set of the first stage: {training sample set of NLP basic samples, training sample set of copywriting generation samples, training sample set of domain professional samples, training sample set of logical reasoning samples}, training sample set of the second stage: {training sample set of short multi-round dialogues}, training sample set of the third stage: {training sample set of progressive thinking chain samples}, training sample set of the fourth stage: {training sample set of medium evolution multi-round samples, training sample set of complex evolution multi-round samples, training sample set of GAtt samples}. In the iterative training process below, the order of obtaining input training samples from the training sample sequence divided into four stages can be the training sample set of the first stage, the training sample set of the second stage, the training sample set of the third stage and the training sample set of the fourth stage.
[0219] Step S207: In the training sample sequence, input training samples are obtained in sequence from the sorted training sample set, and the model parameters of the initial question-answering model are adjusted in sequence through the input training samples obtained in sequence until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions.
[0220] Specifically, in the training sample sequence, the computer device can sequentially obtain input training samples from the sorted training sample set, and adjust the model parameters of the initial question-answering model in sequence through the sequentially obtained input training samples, thereby performing several rounds of iterative training.
[0221] For ease of understanding, the jth round of iterative training is used as an example. The process of the jth round of iterative training can be: in the jth round of iterative training, in the training sample set A i Randomly obtain B first input training samples from the training set, adjust the model parameters of the initial question answering model through B first input training samples, and continue to use the training sample set A i+1 C second input training samples are randomly obtained in the training sample sequence, and the model parameters of the initial question and answer model are adjusted through the C second input training samples until the model parameters of the initial question and answer model are adjusted through the input training samples in the Nth training sample set in the training sample sequence, and the j-th round of iterative training is determined to be completed; B and C are positive integers; j is a positive integer; if the j-th round of iterative training is completed and the initial question and answer model meets the model convergence condition, the initial question and answer model that meets the model convergence condition is determined as the target question and answer model; if the j-th round of iterative training is completed and the initial question and answer model does not meet the model convergence condition, then through the training sample sequence, in the j+1-th round of iterative training, the model parameters of the initial question and answer model are continued to be adjusted in sequence through the input training samples obtained in sequence until the target question and answer model is obtained after the initial question and answer model training converges.
[0222] Specifically, the training process of the initial question-answering model may include multiple iterative training (epochs, also called training cycles), and each iterative training may include multiple training batches.
[0223] In an epoch, the computer device can obtain input training samples from the training sample set of the first stage, the training sample set of the second stage, and the training sample set of the third stage in sequence. For example, it can obtain input training sample 1 from the training sample set of the first stage, adjust the model parameters of the initial question-answering model by input training sample 1, and complete a training batch of the initial question-answering model training, and then obtain input training sample 2 from the training sample set of the second stage, adjust the model parameters of the initial question-answering model by input training sample 2, and then obtain input training sample 3 from the training sample set of the third stage, adjust the model parameters of the initial question-answering model by input training sample 3, and complete a training cycle of the initial question-answering model training.
[0224] For ease of understanding, taking the jth round of iterative training as an example, the training sample sequence may include the training sample set A i And training sample set A i+1 , training sample set A in the training sample sequence i In the training sample set A i+1 Before.
[0225] In the jth round of iterative training, the computer device can iRandomly obtain B first input training samples from the training set, adjust the model parameters of the initial question answering model through B first input training samples, and continue to use the training sample set A i+1 C second input training samples are randomly obtained, and the model parameters of the initial question-answering model are adjusted through the C second input training samples. When the model parameters of the initial question-answering model are adjusted through the input training samples in the Nth training sample set in the training sample sequence, it is determined that the jth round of iterative training is completed.
[0226] Training sample set A i It can be a training sample set of NLP basic samples, training sample set A i+1 It can be a training sample set of copywriting generation samples. When the initial question-answering model finishes adjusting the model parameters through the input training samples in the training sample set of GAtt samples, it is determined that the jth round of iterative training is completed.
[0227] If the jth round of iterative training is complete and the initial question-answering model meets the model convergence criteria, the initial question-answering model that meets the model convergence criteria is determined as the target question-answering model. The target question-answering model can be used as a sample partitioning model that has been fine-tuned using the training sample sequence, and can be used to perform category division for model training in specific fields or topics.
[0228] Among them, the convergence condition indicated by the model evaluation index can be a performance metric of the model, such as accuracy, precision, recall rate, recall rate, etc., which is not limited in this application.
[0229] If the j-th round of iterative training has been completed and the initial question-answering model does not meet the model convergence conditions, then through the training sample sequence, in the j+1-th round of iterative training, continue to use the input training samples obtained in sequence, for example, in the training sample sequence, continue to obtain input training samples in the training sample set of NLP basic samples, input training samples in the training sample set of copywriting generation samples, ..., input training samples in the training sample set of GAtt samples, and adjust the model parameters of the initial question-answering model in sequence until the target question-answering model is obtained after the initial question-answering model training converges.
[0230] Please also refer to Figure 4b, which is a second schematic diagram of a data processing scenario provided in an embodiment of the present application. As shown in Figure 4b, the computer device can obtain training samples for training and iteratively train the initial question-answering model until the initial question-answering model meets the model convergence conditions.
[0231] In the i-th round of iterative training, the computer device can randomly obtain the input training samples of the first stage from the training sample set of the first stage through the training sample sequence, and adjust the model parameters of the initial question-answering model through the input training samples of the first stage.
[0232] Among them, the training sample sets of the first stage include the training sample sets of NLP basic samples, the training sample sets of copywriting generation samples, the training sample sets of domain professional samples, the training sample sets of short multi-round dialogues, and the training sample sets of logical reasoning samples. The computer device can randomly obtain input training sample 1 from the training sample set of NLP basic samples in sequence, adjust the model parameters of the initial question-answering model by inputting training sample 1, and then randomly obtain input training sample 2 from the training sample set of copywriting generation samples, adjust the model parameters of the initial question-answering model by inputting training sample 2, ..., finally, randomly obtain input training sample 5 from the training sample set of logical reasoning samples, adjust the model parameters of the initial question-answering model by inputting training sample 5, and complete the model parameter adjustment of the initial question-answering model in the first stage.
[0233] Then, through the training sample sequence, the input training samples of the second stage are randomly obtained from the training sample set of the second stage, and the model parameters of the initial question-answering model are adjusted through the input training samples of the second stage. The input training sample 6 can be randomly obtained from the training sample set of the progressive thinking chain sample, and the model parameters of the initial question-answering model are adjusted through the input training sample 6 to complete the model parameter adjustment of the initial question-answering model in the second stage.
[0234] Then, through the training sample sequence, the input training samples of the third stage are randomly obtained from the training sample set of the third stage, and the model parameters of the initial question-answering model are adjusted through the input training samples of the second stage.
[0235] The training sample set for the third stage may include a training sample set for medium-evolution multi-round samples, a training sample set for complex-evolution multi-round samples, and a training sample set for GAtt samples. The computer device may sequentially randomly obtain input training sample 7 from the training sample set for medium-evolution multi-round samples, adjust the model parameters of the initial question-answering model using the input training sample 7, then randomly obtain input training sample 8 from the training sample set for complex-evolution multi-round samples, adjust the model parameters of the initial question-answering model using the input training sample 8, and finally randomly obtain input training sample 9 from the training sample set for GAtt samples, adjust the model parameters of the initial question-answering model using the input training sample 9, thereby completing the model parameter adjustment of the initial question-answering model in the third stage.
[0236] When the initial question-answering model completes the model parameter adjustment of the first, second and third stages, it can be determined that the initial question-answering model has completed the i-th round of iterative training. The computer device can determine whether the initial question-answering model meets the model convergence conditions. If the model convergence conditions are not met, the i+1-th round of iterative training is performed (which can be continuing to obtain input training sample 10, input training sample 11,..., input training sample 18). Through the training sample sequence, new input training samples are re-obtained in the training sample set of the first stage, and the initial question-answering model is iteratively trained until the initial question-answering model meets the model convergence conditions.
[0237] It can be understood that in different rounds of iterative training, for example, the input training sample 1 obtained from the training sample set of the NLP basic sample in the i-th iterative training and the input training sample 10 obtained from the training sample set of the NLP basic sample in the i+1-th iterative training, the order of samples obtained in different rounds of iterative training can be randomly distributed, for example, the input training sample 1 can include training sample A, training sample B, and training sample C. The input training sample 2 can include training sample B, training sample C, and training sample A. Not all samples may be sampled in one round of iterative training, and this embodiment of the present application does not limit this.
[0238] Optionally, when obtaining input training samples from the training sample sequence in different iterative training (training cycles), the sorting order of the training sample sets in each training stage of the training sample sequence can also be randomly shuffled. For example, the training sample sequence of the i-th round of iterative training can be {training sample set of the first stage: {training sample set of NLP basic samples, training sample set of copywriting generation samples, training sample set of domain professional samples, training sample set of logical reasoning samples, training sample set of short multi-round dialogues}, training sample set of the second stage: {training sample set of progressive thinking chain samples}, training sample set of the third stage: {training sample set of medium evolution multi-round samples, training sample set of complex evolution multi-round samples, training sample set of GAtt samples}. For example, the training sample sequence for the i+1th round of iterative training can be {training sample set of the first stage: {training sample set of logical reasoning samples, training sample set of domain professional samples, training sample set of copywriting generation samples, training sample set of NLP basic samples, training sample set of short multi-round dialogues}, training sample set of the second stage: {training sample set of progressive thinking chain samples}, training sample set of the third stage: {training sample set of GAtt samples, training sample set of medium evolution multi-round samples, training sample set of complex evolution multi-round samples}.
[0239] In the embodiment of the present application, the question-answer pair data is divided into categories through the question-answer pair category information set to obtain the category labels corresponding to the question-answer pair data. Then, the training sample set and the sample set category labels corresponding to the training sample set are generated through the question-answer pair data and the category labels corresponding to the question-answer pair data. Since the training sample set can include input training samples that are expanded based on the question-answer pair data, the number of training samples can be increased and the diversity of training samples can be increased. By sorting the training sample set by learning difficulty, it is possible to ensure that the model learns the features of training samples with lower learning difficulty during model training, and then gradually help the model learn complex features through training samples with higher learning difficulty, thereby avoiding overfitting of the model in training samples with higher difficulty, generalizing more complex and diverse question-answer scenarios, and improving the efficiency of model training and the accuracy of the model. On the other hand, through data splicing and data rewriting, the topic coverage, skill coverage and diversity of the overall data set of the question-answer pair data can be enhanced, and the model performance and robustness can be optimized during the training process.
[0240] Please refer to Figure 5, which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in Figure 5, the data processing device 1 includes a category classification module 510, a sample generation module 520, a sorting processing module 530 and a model training module 540.
[0241] The category division module 510 is used to obtain M question-answer pair data and a question-answer pair category information set, and classify the M question-answer pair data based on the question-answer pair category information set to obtain category labels corresponding to the M question-answer pair data; the question-answer pair category information set includes category labels; M is a positive integer; the specific functions of the category division module 510 can be found in the specific description of step S101 of the embodiment corresponding to Figure 3 above, and will not be repeated here.
[0242] The sample generation module 520 is used to generate N training sample sets and sample set category labels corresponding to the N training sample sets based on M question-answer pair data and the category labels corresponding to the M question-answer pair data; N is a positive integer; the input training samples in each training sample set are associated with at least one question-answer pair data in the M question-answer pair data; the specific functions of the sample generation module 520 can be found in the specific description of step S102 of the embodiment corresponding to Figure 3 above, and will not be repeated here.
[0243] The sorting processing module 530 is used to sort the N training sample sets according to the learning difficulty corresponding to the sample set category label to obtain a training sample sequence; the specific function of the sorting processing module 530 can be found in the specific description of step S103 of the embodiment corresponding to Figure 3 above, and will not be repeated here.
[0244] The model training module 540 is used to sequentially obtain input training samples from the sorted training sample set in the training sample sequence, and sequentially adjust the model parameters of the initial question-answering model using the sequentially obtained input training samples until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions; the learning difficulty refers to the difficulty of the initial question-answering model training learning input training samples with the sample set category label. The specific functions of the model training module 540 can be found in the specific description of step S104 of the embodiment corresponding to Figure 3 above, and will not be repeated here.
[0245] In one possible implementation, the category classification module 510 is configured to classify the M question-answer pair data into categories based on the question-answer pair category information set. When obtaining the category labels corresponding to the M question-answer pair data, the module is configured to perform the following operations:
[0246] Input M question-answer pairs and their category information set into the sample partitioning model. In the sample partitioning model, keyword vectors are generated based on the category information set, and feature extraction is performed on the M question-answer pairs to obtain feature word vectors corresponding to each of the M question-answer pairs.
[0247] Based on the feature word vector and keyword vector, determine the category labels corresponding to the M question-answer pairs.
[0248] In one possible implementation, the category classification module 510 is configured to perform feature extraction on the M question-answer pairs to obtain feature word vectors corresponding to the M question-answer pairs, and specifically to perform the following operations:
[0249] Filter the M question-answer pair data for business stop words, and perform text segmentation on the filtered M question-answer pair data to obtain P question-answer text segments; P is a positive integer greater than or equal to M;
[0250] Part-of-speech tagging is performed on each of the P question-answer text segments to obtain P question-answer part-of-speech information. Based on the P question-answer text segments and the P question-answer part-of-speech information, feature word vectors corresponding to each of the M question-answer pair data are generated.
[0251] In one possible implementation, the N training sample sets include a training sample set from the first stage, a training sample set from the second stage, and a training sample set from the third stage; the sample generation module 520 is configured to generate the N training sample sets and the sample set category labels corresponding to the N training sample sets based on the M question-answer pair data and the category labels corresponding to the M question-answer pair data, and specifically to perform the following operations:
[0252] Obtain basic question-answer pair data from the M question-answer pair data, divide the basic question-answer pair data into the training sample set of the first stage, and determine the class label corresponding to the basic question-answer pair data as the sample set class label corresponding to the training sample set of the first stage;
[0253] Obtain progressive question-answer pair data from the M question-answer pair data, determine the progressive question-answer pair data as the training sample set for the second stage, generate progressive thinking chain sample category labels, and determine the progressive thinking chain sample category labels as the sample set category labels corresponding to the training sample set for the second stage; the learning difficulty of the progressive thinking chain sample category labels is greater than the learning difficulty of the category labels corresponding to the basic question-answer pair data;
[0254] A first multi-round question-answer pair data is obtained from M question-answer pair data, a second multi-round question-answer pair data is generated based on the basic question-answer pair data, the first multi-round question-answer pair data and the second multi-round question-answer pair data are determined as the training sample set of the third stage, a multi-round category label is generated, and the multi-round category label is determined as the sample set category label corresponding to the training sample set of the third stage; the learning difficulty corresponding to the multi-round category label is greater than the learning difficulty of the progressive thinking chain sample category label.
[0255] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M; the training sample set of the first stage also includes a short dialogue training sample set; and the sample generation module 520 is further configured to perform the following operations:
[0256] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0257] According to the context information, S interrelated basic question-answer pair data are spliced into short dialogue training samples, the short dialogue training samples are added to the short dialogue training sample set, and the short dialogue label is determined as the sample set category label corresponding to the short dialogue training sample set; S is a positive integer less than or equal to H.
[0258] In one possible implementation, when the sample generation module 520 is configured to generate the second multiple rounds of question-answer pair data based on the basic question-answer pair data, it is specifically configured to perform the following operations:
[0259] The basic question-answer pair data is spliced to obtain the spliced multi-round question-answer pair data;
[0260] Reconstruct the basic question-answer pair data to obtain reconstructed multi-round question-answer pair data;
[0261] The spliced multi-round question-answer pair data and the reconstructed multi-round question-answer pair data are determined as the second multi-round question-answer pair data.
[0262] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The sample generation module 520 is configured to perform splicing processing on the basic question-answer pair data to obtain the spliced multiple-round question-answer pair data, specifically to perform the following operations:
[0263] Obtain context information corresponding to H basic question-answer pairs; the context information is used to indicate the semantic relevance between the H basic question-answer pairs;
[0264] Obtain Q basic question-answer pair data based on context information, and splice the Q basic question-answer pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic relevance between the Q basic question-answer pair data is greater than a first relevance threshold;
[0265] Obtain T basic question-answer pair data based on context information, and splice the T basic question-answer pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic relevance between the T basic question-answer pair data is less than the second relevance threshold, and the first relevance threshold is greater than or equal to the second relevance threshold;
[0266] The following conversation data and random conversation data are determined as spliced multi-round question-answer pair data.
[0267] In one possible implementation, the number of basic question-answer pair data is H, where H is a positive integer less than or equal to M. The sample generation module 520 is configured to reconstruct the basic question-answer pair data to obtain the reconstructed multi-round question-answer pair data, specifically performing the following operations:
[0268] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category labels contained in the question-answer pair category information set; the deep evolutionary data is associated with the evolutionary category labels;
[0269] Obtain context information corresponding to H basic question-answer pair data, and rewrite the mutually related basic question-answer pair data into breadth-evolved data based on the context information; the context information is used to indicate the semantic relevance between the H basic question-answer pair data;
[0270] Deep evolution data and broad evolution data are determined as reconstructed multi-round question-answer pair data.
[0271] In a possible implementation, when the sample generation module 520 is used to rewrite the basic question-answer pair data into the deeply evolved data based on the evolution constraint information, it is specifically used to perform the following operations:
[0272] Based on the evolutionary constraint information, the basic question-answer pair data is rewritten into transition evolution data;
[0273] If the question-and-answer rounds of the transition evolution data are less than the question-and-answer round threshold, the transition evolution data will continue to be rewritten based on the evolution constraint information until the question-and-answer rounds of the transition evolution data after the data rewrite are greater than or equal to the question-and-answer round threshold, and the transition evolution data after the data rewrite will be determined as deep evolution data.
[0274] In a possible implementation, the training sample sequence includes a training sample set A i And training sample set A i+1 , training sample set A in the training sample sequence i In the training sample set A i+1 Before; the model training module 540 is used to obtain input training samples from the sorted training sample set in sequence in the training sample sequence, and adjust the model parameters of the initial question-answering model in sequence through the input training samples obtained in sequence until the initial question-answering model training converges to obtain the target question-answering model, specifically for performing the following operations:
[0275] In the jth round of iterative training, in the training sample set A i Randomly obtain B first input training samples from the training set, adjust the model parameters of the initial question answering model through B first input training samples, and continue to use the training sample set A i+1 Randomly obtain C second input training samples in the training sample set, and adjust the model parameters of the initial question answering model through the C second input training samples until the model parameters of the initial question answering model are adjusted through the input training samples in the Nth training sample set in the training sample sequence, and the jth round of iterative training is determined to be completed; B and C are positive integers; j is a positive integer;
[0276] If the jth round of iterative training has been completed and the initial question-answering model meets the model convergence conditions, the initial question-answering model that meets the model convergence conditions will be determined as the target question-answering model;
[0277] If the j-th round of iterative training has been completed and the initial question-answering model does not meet the model convergence conditions, then through the training sample sequence, in the j+1-th round of iterative training, the model parameters of the initial question-answering model will continue to be adjusted in sequence through the input training samples obtained in sequence until the initial question-answering model training converges and the target question-answering model is obtained.
[0278] Please refer to Figure 6, which is a structural diagram of a computer device provided in an embodiment of the present application. As shown in Figure 6, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may also be at least one storage device located away from the aforementioned processor 1001. As shown in Figure 6, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module and a device control application.
[0279] In the computer device 1000 shown in FIG6 , the network interface 1004 can provide a network communication element; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0280] Obtain M question-answer pair data and a question-answer pair category information set, classify the M question-answer pair data into categories based on the question-answer pair category information set, and obtain category labels corresponding to the M question-answer pair data; the question-answer pair category information set includes the category labels; M is a positive integer;
[0281] Based on the M question-answer pairs and the class labels corresponding to the M question-answer pairs, generate N training sample sets and the sample set class labels corresponding to the N training sample sets; N is a positive integer; the input training sample in each training sample set is associated with at least one question-answer pair in the M question-answer pairs;
[0282] Sort the N training sample sets by the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence; the learning difficulty refers to the difficulty of the initial question-answering model training to learn the input training samples with the sample set category labels;
[0283] In the training sample sequence, input training samples are obtained in sequence from the sorted training sample set, and the model parameters of the initial question-answering model are adjusted in sequence through the input training samples obtained in sequence until the target question-answering model is obtained after the initial question-answering model training converges; the target question-answering model is used to generate question-answering service results based on business questions.
[0284] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method in any of the embodiments corresponding to Figures 3 and 4a above, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated.
[0285] In addition, it should be noted that the embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program. When the processor executes the computer program, it can perform the description of the data processing method in any of the embodiments corresponding to Figures 3 and 4a above. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0286] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been displayed or is about to be displayed.
[0287] In addition, it should be noted that embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in either of the embodiments corresponding to FIG. 3 and FIG. 4a above.
[0288] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0289] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example in terms of network elements. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described network elements for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0290] The method and related apparatus provided in the embodiment of the present application are described with reference to the method flow chart and / or structural diagram provided in the embodiment of the present application, and specifically can be implemented by computer program instructions for each process and / or box of the method flow chart and / or structural diagram, and the combination of the process and / or box in the flow chart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable device to produce a machine, so that the instructions executed by the processor of the computer or other programmable device produce a device for implementing the function specified in one process or multiple processes of the flow chart and / or one box or multiple boxes of the structural diagram. These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, and the instruction device implements the function specified in one process or multiple processes of the flow chart and / or one box or multiple boxes of the structural diagram. These computer program instructions can also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the structural diagram.
[0291] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0292] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0293] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain M question-and-answer pair data and a question-and-answer pair category information set, perform category division on the M question-and-answer pair data based on the question-and-answer pair category information set, and obtain category labels corresponding to the M question-and-answer pair data respectively; The question-and-answer pair category information set includes the category labels; M is a positive integer; Based on the M question-and-answer pair data and the category labels corresponding to the M question-and-answer pair data respectively, generate N training sample sets and sample set category labels corresponding to the N training sample sets respectively; N is a positive integer; Each input training sample in each training sample set is associated with at least one of the M question-and-answer pair data; Sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels, and obtain a training sample sequence; the learning difficulty refers to the difficulty of the initial question-and-answer model training to learn the input training samples with the sample set category labels; In the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model converges to obtain a target question-and-answer model; the target question-and-answer model is used to generate a question-and-answer service result through a business question.
2. The method according to claim 1, wherein The performing category division on the M question-and-answer pair data based on the question-and-answer pair category information set to obtain category labels corresponding to the M question-and-answer pair data respectively includes: Input the M question-and-answer pair data and the question-and-answer pair category information set into a sample division model, and in the sample division model, generate keyword vectors based on the question-and-answer pair category information set; Extract features from the M question-and-answer pair data to obtain feature word vectors corresponding to the M question-and-answer pair data respectively; Based on the feature word vectors and the keyword vectors, determine category labels corresponding to the M question-and-answer pair data respectively.
3. The method according to claim 1 or 2, characterized in that, The extracting features from the M question-and-answer pair data to obtain feature word vectors corresponding to the M question-and-answer pair data respectively includes: Filter business stop words from the M question-and-answer pair data, perform text division on the filtered M question-and-answer pair data to obtain P question-and-answer text segments; P is a positive integer greater than or equal to M; Perform part-of-speech tagging on the P question-and-answer text segments respectively to obtain P question-and-answer part-of-speech information, and generate feature word vectors corresponding to the M question-and-answer pair data respectively based on the P question-and-answer text segments and the P question-and-answer part-of-speech information.
4. The method according to any one of claims 1 to 3, characterized in that, The N training sample sets include a training sample set in the first stage, a training sample set in the second stage, and a training sample set in the third stage; The generating N training sample sets and sample set category labels corresponding to the N training sample sets respectively based on the M question-and-answer pair data and the category labels corresponding to the M question-and-answer pair data respectively includes: Obtain basic question-and-answer pair data from the M question-and-answer pair data, divide the basic question-and-answer pair data into the training sample set in the first stage, and determine the category label corresponding to the basic question-and-answer pair data as the sample set category label corresponding to the training sample set in the first stage; Obtain progressive Q&A pair data from the M Q&A pair data, determine the progressive Q&A pair data as the training sample set for the second stage, generate progressive thinking chain sample class labels, and determine the progressive thinking chain sample class labels as the sample set class labels corresponding to the training sample set for the second stage; the learning difficulty of the progressive thinking chain sample class labels is greater than the learning difficulty of the class labels corresponding to the basic Q&A pair data. Obtain the first multi-turn Q&A pair data from the M Q&A pair data, generate the second multi-turn Q&A pair data according to the basic Q&A pair data, determine the first multi-turn Q&A pair data and the second multi-turn Q&A pair data as the training sample set for the third stage, generate multi-turn class labels, and determine the multi-turn class labels as the sample set class labels corresponding to the training sample set for the third stage; the learning difficulty corresponding to the multi-turn class labels is greater than the learning difficulty of the progressive thinking chain sample class labels.
5. The method according to any one of claims 1 to 4, characterized in that, The number of the basic Q&A pair data is H, and H is a positive integer less than or equal to M; the training sample set for the first stage further includes a short dialogue training sample set; the method further includes: Obtain the context information corresponding to the H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data. Splice S mutually related basic Q&A pair data into short dialogue training samples according to the context information, add the short dialogue training samples to the short dialogue training sample set, and determine the short dialogue labels as the sample set class labels corresponding to the short dialogue training sample set; S is a positive integer less than or equal to H.
6. The method according to any one of claims 1 to 5, characterized in that The generating the second multi-turn Q&A pair data according to the basic Q&A pair data includes: Perform splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data. Perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data. Determine the spliced multi-turn Q&A pair data and the reconstructed multi-turn Q&A pair data as the second multi-turn Q&A pair data.
7. The method according to any one of claims 1 to 6, characterized in that, The number of the basic Q&A pair data is H, and H is a positive integer less than or equal to M; the performing splicing processing on the basic Q&A pair data to obtain the spliced multi-turn Q&A pair data includes: Obtain the context information corresponding to the H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data. Obtain Q basic Q&A pair data according to the context information, and splice the Q basic Q&A pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation threshold. Obtain T basic Q&A pair data according to the context information, and splice the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation threshold, and the first correlation threshold is greater than or equal to the second correlation threshold. Determine the follow-up dialogue data and the random dialogue data as the spliced multi-turn Q&A pair data.
8. The method according to any one of claims 1 to 7, characterized in that, The number of the basic Q&A pair data is H, where H is a positive integer less than or equal to M; the data reconstruction of the basic Q&A pair data to obtain the reconstructed multi-turn Q&A pair data includes: Based on the evolutionary constraint information, rewriting the basic Q&A pair data into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary class tags included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary class tags; Obtaining the context information corresponding to H basic Q&A pair data, and rewriting the mutually related basic Q&A pair data into broad evolutionary data according to the context information; the context information is used to indicate the semantic association degree between the H basic Q&A pair data; Determining the deep evolutionary data and the broad evolutionary data as the reconstructed multi-turn Q&A pair data.
9. The method according to any one of claims 1 to 8, characterized in that, The rewriting of the basic Q&A pair data into deep evolutionary data based on the evolutionary constraint information includes: Based on the evolutionary constraint information, rewriting the basic Q&A pair data into transitional evolutionary data; If the number of Q&A turns of the transitional evolutionary data is less than the Q&A turn threshold, then based on the evolutionary constraint information, continue to rewrite the transitional evolutionary data until the number of Q&A turns of the rewritten transitional evolutionary data is greater than or equal to the Q&A turn threshold, and determine the rewritten transitional evolutionary data as the deep evolutionary data.
10. The method according to any one of claims 1 to 9, characterized in that The training sample sequence includes training sample set A i and training sample set A i+1 , in the training sample sequence, training sample set A i is located before training sample set A i+1 ; In the training sample sequence, sequentially obtaining input training samples from the sorted training sample set, and sequentially adjusting the model parameters of the initial Q&A model through the sequentially obtained input training samples until the target Q&A model is obtained after the training of the initial Q&A model converges, including: In the j-th round of iterative training, among the training sample set A i randomly obtain B first input training samples, and adjust the model parameters of the initial Q&A model through the B first input training samples. Then continue to randomly obtain C second input training samples in the training sample set A i+1 and adjust the model parameters of the initial Q&A model through the C second input training samples. Until when the model parameter adjustment of the initial Q&A model is completed by using the input training samples in the N-th training sample set in the training sample sequence, it is determined that the j-th round of iterative training is completed; B and C are positive integers; j is a positive integer; If the j-th round of iterative training is completed and the initial Q&A model meets the model convergence condition, then determine the initial Q&A model that meets the model convergence condition as the target Q&A model; If the j-th round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then in the (j + 1)-th round of iterative training through the training sample sequence, continue to sequentially adjust the model parameters of the initial Q&A model through the sequentially obtained input training samples until the target Q&A model is obtained after the training of the initial Q&A model converges.
11. A data processing device, characterized in that, Including: A category division module, configured to obtain M Q&A pair data and a Q&A pair category information set, and perform category division on the M Q&A pair data based on the Q&A pair category information set to obtain the class tags corresponding to the M Q&A pair data respectively; The Q&A pair category information set includes the class tags; M is a positive integer; A sample generation module, configured to generate N training sample sets and the sample set class tags corresponding to the N training sample sets respectively based on the M Q&A pair data and the class tags corresponding to the M Q&A pair data respectively; N is a positive integer; The input training samples in each training sample set are associated with at least one of the M Q&A pair data; A sorting processing module, configured to sort the N training sample sets according to the learning difficulty corresponding to the target label of the sample set class, so as to obtain a training sample sequence; the learning difficulty refers to the difficulty of training and learning the input training samples with the target label of the sample set class by the initial question and answer model. A model training module, configured to sequentially obtain input training samples from the sorted training sample sets in the training sample sequence, and sequentially adjust the model parameters of the initial question and answer model through the sequentially obtained input training samples until the initial question and answer model converges after training to obtain a target question and answer model; the target question and answer model is used to generate a question and answer service result through a business question.
12. A computer device, characterized in that, Comprising: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-10.
14. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium and is adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-10.
Citation Information
Patent Citations
Text classification model training method and device, text classification method and device and storage medium
CN112131366A
Question and answer processing method and device and question and answer processing model training method and device
CN116662495A
Model training method and device, question answering method and device, equipment and medium
CN116894080A
Questionnaire survey method and device, questionnaire survey model training method and device, equipment, medium and product
CN117312530A
Shared network learning for machine learning enabled text classification
US20230169362A1
Cited By
Method and device for acquiring industry multi-modal data, and electronic equipment
CN121413750A