Data processing method, device and equipment and readable storage medium
Through question and answer classification and sorting of category information sets, the training sample sets are solved, and the training efficiency and accuracy problems of large language models in the supervised fine-tuning stage are achieved, achieving more efficient model training and better question-and-answer service capabilities.
Patent Information
- Application Number
- CN202410024726.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
During the supervised fine-tuning stage, existing large language models have mixed training of samples of different learning difficulties, resulting in over-tuning parameters to fit complex samples, and cannot effectively learn simple sample features, reducing training efficiency and model accuracy.
Through Q&A, the category information set is classified, the training sample set is generated and sorted by learning difficulty, and the model parameters are gradually adjusted until converges, ensuring that the model first learns the characteristics of low-difficulty samples and then gradually learns high-difficulty samples.
This improves the efficiency and accuracy of model training, avoids overfitting, and enhances the generalization ability of the model in complex and diverse question-and-answer scenarios.
Smart Images

Figure CN120277178A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, and readable storage medium. Background Art
[0002] In the supervised fine-tuning stage of existing large language models, samples with different learning difficulties are trained in the same batch. The huge samples will be mixed with simple and complex samples. When the model does not have a certain understanding ability, directly adjusting the model parameters through complex samples will cause the model to over-adjust the parameters to fit the complex samples, resulting in the inability to learn the features of simple samples, and thus unable to be applied to simple or normal samples, reducing the efficiency of model training, and also causing sub-optimal model training and low model accuracy, affecting the generation ability of the model. Summary of the Invention
[0003] Embodiments of the present application provide a data processing method, apparatus, device, and readable storage medium, which can improve the training efficiency and model accuracy of the target question-and-answer model.
[0004] On the one hand, an embodiment of the present application provides a data processing method, including:
[0005] Obtain M question-and-answer pair data and a question-and-answer pair category information set, and perform category division on the M question-and-answer pair data based on the question-and-answer pair category information set to obtain class labels corresponding to the M question-and-answer pair data respectively; the question-and-answer pair category information set includes class labels; M is a positive integer;
[0006] Generate N training sample sets and sample set class labels corresponding to the N training sample sets respectively based on the M question-and-answer pair data and the class labels corresponding to the M question-and-answer pair data respectively; N is a positive integer;
[0007] Sort the N training sample sets according to the learning difficulty corresponding to the sample set class labels to obtain a training sample sequence;
[0008] In the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model converges after training to obtain a target question-and-answer model; the target question-and-answer model is used to generate a question-and-answer service result through a business question.
[0009] Among them, performing category division on the M question-and-answer pair data based on the question-and-answer pair category information set to obtain class labels corresponding to the M question-and-answer pair data respectively includes:
[0010] Input M pairs of Q&A data and the Q&A category information set into the sample division model. In the sample division model, generate keyword vectors based on the Q&A category information set, extract features from the M pairs of Q&A data, and obtain the feature word vectors corresponding to the M pairs of Q&A data respectively.
[0011] Based on the feature word vectors and the keyword vectors, determine the class labels corresponding to the M pairs of Q&A data respectively.
[0012] Among them, extracting features from the M pairs of Q&A data to obtain the feature word vectors of the M pairs of Q&A data includes:
[0013] Filter the business stop words from the M pairs of Q&A data, and perform text division on the filtered M pairs of Q&A data to obtain P Q&A text segments; P is a positive integer greater than or equal to M.
[0014] Perform part-of-speech tagging on the P Q&A text segments to obtain P Q&A part-of-speech information, and generate the feature word vectors of the M pairs of Q&A data based on the P Q&A text segments and the P Q&A part-of-speech information.
[0015] Among them, the N training sample sets include the training sample set in the first stage, the training sample set in the second stage, and the training sample set in the third stage.
[0016] Based on the M pairs of Q&A data and the class labels corresponding to the M pairs of Q&A data respectively, generate N training sample sets and the sample set class labels corresponding to the N training sample sets respectively, including:
[0017] Obtain the basic Q&A data from the M pairs of Q&A data, divide the basic Q&A data into the training sample set in the first stage, and determine the class label corresponding to the basic Q&A data as the sample set class label corresponding to the training sample set in the first stage.
[0018] Obtain the progressive Q&A data from the M pairs of Q&A data, determine the progressive Q&A data as the training sample set in the second stage, and determine the class label corresponding to the progressive Q&A data as the sample set class label corresponding to the training sample set in the second stage; the learning difficulty of the class label corresponding to the progressive Q&A data is greater than the learning difficulty of the class label corresponding to the basic Q&A data.
[0019] Obtain the first multi-turn Q&A data from the M pairs of Q&A data, generate the second multi-turn Q&A data according to the basic Q&A data, determine the first multi-turn Q&A data and the second multi-turn Q&A data as the training sample set in the third stage, generate the multi-turn class label, and determine the multi-turn class label as the sample set class label corresponding to the training sample set in the third stage; the learning difficulty of the multi-turn class label is greater than the learning difficulty of the class label corresponding to the progressive Q&A data.
[0020] Among them, the training sample set in the first stage further includes a short dialogue training sample set; the method further includes:
[0021] Obtain the context information of the basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data;
[0022] According to the context information, splice S mutually related basic Q&A pair data into short dialogue training samples, add the short dialogue training samples to the short dialogue training sample set, and determine the short dialogue label as the sample set class target label corresponding to the short dialogue training sample set; S is a positive integer.
[0023] Among them, generating the second multi-turn Q&A pair data according to the basic Q&A pair data includes:
[0024] Perform splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data;
[0025] Perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data;
[0026] Determine the spliced multi-turn Q&A pair data and the reconstructed multi-turn Q&A pair data as the second multi-turn Q&A pair data.
[0027] Among them, the number of basic Q&A pair data is H, and H is a positive integer less than or equal to M; performing splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data includes:
[0028] Obtain the context information corresponding to H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data;
[0029] According to the context information, obtain Q basic Q&A pair data, and splice the Q basic Q&A pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation degree threshold;
[0030] According to the context information, obtain T basic Q&A pair data, and splice the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation degree threshold, and the first correlation degree threshold is greater than or equal to the second correlation degree threshold;
[0031] Determine the follow-up dialogue data and the random dialogue data as the spliced multi-turn Q&A pair data.
[0032] Among them, performing data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data includes:
[0033] Based on evolutionary constraint information, rewrite the basic Q&A pair data into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category tags included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary category tags;
[0034] Obtain the context information of the basic Q&A pair data, and rewrite the interrelated basic Q&A pair data into broad evolutionary data according to the context information; the context information is used to indicate the semantic association degree between the basic Q&A pair data;
[0035] Determine the deep evolutionary data and the broad evolutionary data as the reconstructed multi-round Q&A pair data.
[0036] Among them, based on the evolutionary constraint information, rewriting the basic Q&A pair data into deep evolutionary data includes:
[0037] Based on the evolutionary constraint information, rewrite the basic Q&A pair data into transitional evolutionary data;
[0038] If the number of Q&A rounds of the transitional evolutionary data is less than the Q&A round threshold, then based on the evolutionary constraint information, continue to rewrite the transitional evolutionary data until the number of Q&A rounds of the rewritten transitional evolutionary data is greater than or equal to the Q&A round threshold, and determine the rewritten transitional evolutionary data as the deep evolutionary data.
[0039] Among them, the training sample sequence includes training sample set A i and training sample set A i+1 , in the training sample sequence, training sample set A i is located before training sample set A i+1 ;
[0040] In the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and adjust the model parameters of the initial Q&A model in turn through the sequentially obtained input training samples until the target Q&A model is obtained after the initial Q&A model converges in training, including:
[0041] In the j-th round of iterative training, randomly obtain B first input training samples in training sample set A i , adjust the model parameters of the initial Q&A model through the B first input training samples, and continue to randomly obtain C second input training samples in training sample set A i+1 , adjust the model parameters of the initial Q&A model through the C second input training samples until when the adjustment of the model parameters of the initial Q&A model is completed through the input training samples in the N-th training sample set in the training sample sequence, it is determined that the j-th round of iterative training is completed; B and C are positive integers; j is a positive integer;
[0042] If the j-th round of iterative training is completed and the initial Q&A model meets the model convergence condition, the initial Q&A model that meets the model convergence condition is determined as the target Q&A model;
[0043] If the j-th round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then in the (j + 1)-th round of iterative training, through the training sample sequence, the model parameters of the initial Q&A model are sequentially adjusted by the input training samples obtained in order until the target Q&A model is obtained after the initial Q&A model converges in training.
[0044] One aspect of the embodiments of the present application provides a data processing device, including:
[0045] A category division module, configured to obtain M Q&A pair data and a Q&A pair category information set, and perform category division on the M Q&A pair data based on the Q&A pair category information set to obtain class labels corresponding to the M Q&A pair data respectively; the Q&A pair category information set includes class labels; M is a positive integer;
[0046] A sample generation module, configured to generate N training sample sets and sample set class labels corresponding to the N training sample sets respectively based on the M Q&A pair data and the class labels corresponding to the M Q&A pair data respectively; N is a positive integer;
[0047] A sorting processing module, configured to sort the N training sample sets through the learning difficulty corresponding to the sample set class labels to obtain a training sample sequence;
[0048] A model training module, configured to sequentially obtain input training samples from the sorted training sample sets in the training sample sequence, and sequentially adjust the model parameters of the initial Q&A model by the input training samples obtained in order until the target Q&A model is obtained after the initial Q&A model converges in training; the target Q&A model is used to generate a Q&A service result through a business question.
[0049] In a possible implementation manner, when the category division module is configured to perform category division on the M Q&A pair data based on the Q&A pair category information set to obtain class labels corresponding to the M Q&A pair data respectively, it is specifically configured to perform the following operations:
[0050] Input the M Q&A pair data and the Q&A pair category information set into a sample division model. In the sample division model, generate keyword vectors based on the Q&A pair category information set, extract features from the M Q&A pair data, and obtain feature word vectors corresponding to the M Q&A pair data respectively;
[0051] Determine the class labels corresponding to the M Q&A pair data respectively based on the feature word vectors and the keyword vectors.
[0052] In a possible implementation, when the category division module is used to extract features from M question-and-answer pair data to obtain the feature word vectors of the M question-and-answer pair data, it is specifically used to perform the following operations:
[0053] Filter out business stop words from the M question-and-answer pair data, and perform text division on the filtered M question-and-answer pair data to obtain P question-and-answer text segments; P is a positive integer greater than or equal to M;
[0054] Perform part-of-speech tagging on the P question-and-answer text segments to obtain P question-and-answer part-of-speech information, and generate the feature word vectors of the M question-and-answer pair data based on the P question-and-answer text segments and the P question-and-answer part-of-speech information.
[0055] In a possible implementation, the N training sample sets include the training sample set in the first stage, the training sample set in the second stage, and the training sample set in the third stage; when the sample generation module is used to generate the N training sample sets and the sample set class target tags respectively corresponding to the N training sample sets based on the M question-and-answer pair data and the class target tags respectively corresponding to the M question-and-answer pair data, it is specifically used to perform the following operations:
[0056] Obtain the basic question-and-answer pair data from the M question-and-answer pair data, divide the basic question-and-answer pair data into the training sample set in the first stage, and determine the class target tag corresponding to the basic question-and-answer pair data as the sample set class target tag corresponding to the training sample set in the first stage;
[0057] Obtain the progressive question-and-answer pair data from the M question-and-answer pair data, determine the progressive question-and-answer pair data as the training sample set in the second stage, and determine the class target tag corresponding to the progressive question-and-answer pair data as the sample set class target tag corresponding to the training sample set in the second stage; the learning difficulty of the class target tag corresponding to the progressive question-and-answer pair data is greater than the learning difficulty of the class target tag corresponding to the basic question-and-answer pair data;
[0058] Obtain the first multi-turn question-and-answer pair data from the M question-and-answer pair data, generate the second multi-turn question-and-answer pair data according to the basic question-and-answer pair data, determine the first multi-turn question-and-answer pair data and the second multi-turn question-and-answer pair data as the training sample set in the third stage, generate the multi-turn class target tag, and determine the multi-turn class target tag as the sample set class target tag corresponding to the training sample set in the third stage; the learning difficulty of the multi-turn class target tag is greater than the learning difficulty of the class target tag corresponding to the progressive question-and-answer pair data.
[0059] In a possible implementation, the training sample set in the first stage further includes a short conversation training sample set; the sample generation module is further used to perform the following operations:
[0060] Obtain the context information of the basic question-and-answer pair data; the context information is used to indicate the semantic correlation degree between the basic question-and-answer pair data;
[0061] S interrelated basic Q&A pair data are concatenated into short dialogue training samples according to context information, the short dialogue training samples are added to the short dialogue training sample set, and the short dialogue labels are determined as the sample set class target labels corresponding to the short dialogue training sample set; S is a positive integer.
[0062] In a possible implementation, when the sample generation module is used to generate second multi-turn Q&A pair data based on the basic Q&A pair data, it is specifically used to perform the following operations:
[0063] Perform concatenation processing on the basic Q&A pair data to obtain concatenated multi-turn Q&A pair data;
[0064] Perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data;
[0065] Determine the concatenated multi-turn Q&A pair data and the reconstructed multi-turn Q&A pair data as the second multi-turn Q&A pair data.
[0066] In a possible implementation, the number of basic Q&A pair data is H, and H is a positive integer less than or equal to M; when the sample generation module is used to perform concatenation processing on the basic Q&A pair data to obtain concatenated multi-turn Q&A pair data, it is specifically used to perform the following operations:
[0067] Obtain the context information corresponding to H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data;
[0068] Obtain Q basic Q&A pair data according to the context information, and concatenate the Q basic Q&A pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation degree threshold;
[0069] Obtain T basic Q&A pair data according to the context information, and concatenate the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation degree threshold, and the first correlation degree threshold is greater than or equal to the second correlation degree threshold;
[0070] Determine the follow-up dialogue data and the random dialogue data as the concatenated multi-turn Q&A pair data.
[0071] In a possible implementation, when the sample generation module is used to perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data, it is specifically used to perform the following operations:
[0072] Based on the evolutionary constraint information, rewrite the basic Q&A pair data as deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary class target labels included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary class target labels;
[0073] Obtain the context information of the basic Q&A pair data, and rewrite the interrelated basic Q&A pair data as breadth-evolved data according to the context information; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data;
[0074] Determine the depth-evolved data and the breadth-evolved data as the reconstructed multi-round Q&A pair data.
[0075] In a possible implementation manner, when the sample generation module is used to rewrite the basic Q&A pair data as depth-evolved data based on the evolutionary constraint information, it is specifically used to perform the following operations:
[0076] Rewrite the basic Q&A pair data as transitional-evolved data based on the evolutionary constraint information;
[0077] If the number of Q&A rounds of the transitional-evolved data is less than the Q&A round threshold, continue to rewrite the transitional-evolved data based on the evolutionary constraint information until the number of Q&A rounds of the rewritten transitional-evolved data is greater than or equal to the Q&A round threshold, and then determine the rewritten transitional-evolved data as the depth-evolved data.
[0078] In a possible implementation manner, the training sample sequence includes training sample set A i and training sample set A i+1 , in the training sample sequence, training sample set A i is located before training sample set A i+1 ; the model training module is used to sequentially obtain input training samples from the sorted training sample sets in the training sample sequence, and sequentially adjust the model parameters of the initial Q&A model through the obtained input training samples until the target Q&A model is obtained after the initial Q&A model training converges. Specifically, it is used to perform the following operations:
[0079] In the j-th round of iterative training, randomly obtain B first input training samples in training sample set A i , and adjust the model parameters of the initial Q&A model through the B first input training samples. Then continue to randomly obtain C second input training samples in training sample set A i+1 , and adjust the model parameters of the initial Q&A model through the C second input training samples until the adjustment of the model parameters of the initial Q&A model is completed by using the input training samples in the N-th training sample set in the training sample sequence, and then determine that the j-th round of iterative training is completed; B and C are positive integers; j is a positive integer;
[0080] If the j-th round of iterative training is completed and the initial Q&A model meets the model convergence condition, then determine the initial Q&A model that meets the model convergence condition as the target Q&A model;
[0081] If the j-th round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then in the (j + 1)-th round of iterative training, through the training sample sequence, the model parameters of the initial Q&A model are sequentially adjusted by the input training samples obtained in order until the target Q&A model is obtained after the initial Q&A model training converges.
[0082] One aspect of the embodiments of the present application provides a computer device, including: a processor, a memory, and a network interface;
[0083] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, and the memory is used to store computer programs. When the computer programs are executed by the processor, the computer device executes the methods provided by the embodiments of the present application.
[0084] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is suitable for being loaded and executed by a processor so that a computer device having the processor executes the methods provided by the embodiments of the present application.
[0085] One aspect of the embodiments of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the methods provided by the embodiments of the present application.
[0086] In the embodiments of the present application, through the Q&A pair category information set, the Q&A pair data is categorized to obtain the class labels corresponding to the Q&A pair data respectively. Then, through the Q&A pair data and the class labels corresponding to the Q&A pair data respectively, a training sample set and the sample set class labels corresponding to the training sample set are generated. The number of training samples can be increased, and the diversity of the training samples can be increased. By sorting the training sample set according to the learning difficulty, during the model training process, it can be ensured that the model learns the features of the training samples with lower learning difficulty, and then trained by the training samples with higher learning difficulty, gradually helping the model learn complex features, so as to achieve the same accuracy as using the entire training sample sequence, avoiding overfitting of the model in the training samples with higher difficulty, generalizing more complex and diverse Q&A scenarios, and improving the efficiency of model training and the accuracy of the model. Description of the Drawings
[0087] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0088] Figure 1 It is a schematic structural diagram of a blockchain network provided by an embodiment of the present application;
[0089] Figure 2 It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 1 ;
[0090] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;
[0091] Figure 4a It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;
[0092] Figure 4b It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 2 ;
[0093] Figure 5 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0094] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0095] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0096] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable them to have functions such as perception, reasoning, and decision-making.
[0097] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0098] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and intelligent transportation.
[0099] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0100] Deep Learning (DL) is to learn the internal laws and representation levels of sample data. The information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analytical learning like humans and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.
[0101] The solution provided in the embodiments of this application relates to the deep learning technology of artificial intelligence, and will be specifically described through the following embodiments.
[0102] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a network architecture provided in the embodiments of this application. As Figure 1 shown, the network architecture may include a business server 100 and a cluster of terminal devices. The cluster of terminal devices may include terminal devices 10a, 10b,..., 10n. Among them, any terminal device in the cluster of terminal devices may have a communication connection with the business server 100. For example, there is a communication connection between terminal device 10a and business server 100, and there is a communication connection between terminal device 10b and business server 100. Among them, the above communication connection does not limit the connection method, and can be directly or indirectly connected through wired communication methods, or can be directly or indirectly connected through wireless communication methods, or can be connected through other methods, which are not limited in this application.
[0103] Specifically, as Figure 1 shown, each terminal device in the terminal cluster can send a category division request to the business server 100. The category division request may include question-and-answer pair data and a question-and-answer pair category information set.
[0104] Among them, the question-and-answer pair data may be text data for training a large language model, including a series of questions and corresponding answers. The question-and-answer pair data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0105] The question-and-answer pair category information set may be an information set for annotating the question-and-answer pair data, and may include multiple category labels. The category labels can be divided according to dimensions such as the learning difficulty of the question and answer, the field and topic of the question, the answer type and length, etc.
[0106] Taking terminal device 10a as an example, the business server 100 can obtain the category division request sent by terminal device 10a, and through the sample division model in the business server 100, perform category division on the question-and-answer pair data and the question-and-answer pair category information set in the category division request to obtain the category label corresponding to the question-and-answer pair data.
[0107] The business server 100 can generate a training sample set and sample class target tags corresponding to the training sample set respectively through question-and-answer pair data and class target tags corresponding to the question-and-answer pair data, and sort the training sample set according to the learning difficulty corresponding to the sample set class target tags to obtain a training sample sequence. Among them, the sample division model can be a converged large language model, which can use the question-and-answer pair category information set as the text prompt of the question-and-answer pair data to divide the question-and-answer pair data into corresponding class target tags.
[0108] It can be understood that by sorting the training sample set according to the learning difficulty, during the model training process, the model can be guaranteed to learn the features of training samples with lower learning difficulty, and then trained with training samples with higher learning difficulty to gradually help the model learn complex features, so as to achieve the same accuracy as using the entire training sample sequence, avoid the model from overfitting in training samples with higher difficulty, generalize more complex and diverse question-and-answer scenarios, and improve the efficiency of model training and the accuracy of the model.
[0109] The business server 100 can send the training sample sequence to the terminal device 10a. The terminal device 10a can, in the training sample sequence, sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-and-answer model through the input training samples obtained in sequence until the initial question-and-answer model in the terminal device 10a converges after training to obtain a target question-and-answer model. The target question-and-answer model can be used to generate a question-and-answer service result through a business question.
[0110] It can be understood that the business server 100 and the terminal device 10a can also be integrated into a single server, for example, it can be called a model server.
[0111] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a data processing scenario provided by an embodiment of the present application. Figure 1 . As Figure 2 shown, the terminal device 10a can send a category division request to the business server 100. The category division request can include question-and-answer pair data and a question-and-answer pair category information set.
[0112] Among them, the question-and-answer pair data can be text data for training a large language model, including a series of questions and corresponding answers. The question-and-answer pair data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0113] The question-and-answer pair category information set can be an information set for annotating the question-and-answer pair data, and can include multiple class target tags. The class target tags can be divided according to dimensions such as the learning difficulty of the question and answer, the field and topic of the question, the answer type and length, etc.
[0114] Taking the terminal device 10a as an example, the service server 100 can obtain the category division request sent by the terminal device 10a, and through the sample division model in the service server 100, perform category division on the Q&A pair data and the Q&A pair category information set in the category division request to obtain the category label corresponding to the Q&A pair data.
[0115] The sample division model in the service server 100 can use the Q&A pair category information set as the text prompt for the Q&A pair data and perform category division on the Q&A pair data. Its classification steps can be: "The following is the category system: Q&A pair category information set. Please classify the following text "{Q&A pair data 1, Q&A pair data 2,..., Q&A pair data N}" into the most relevant category label in the above Q&A pair category information set, and output the result to the result in the following structure {"category": result}", and the organization form of the result "result" can be {"Q&A pair data 1": NLP (Natural Language Processing) basic category label, "Q&A pair data 2": copywriting generation category label,..., "Q&A pair data N": multi-round dialogue category label}.
[0116] The service server 100 can obtain basic Q&A pair data (which can also be called a regular natural corpus) from the Q&A pair data, and divide the basic Q&A pair data into the training sample set of the first stage (the basic Q&A pair data can include NLP basic Q&A pair data, copywriting generation Q&A pair data, domain professional Q&A pair data, short multi-round dialogue data, and logical reasoning Q&A pair data).
[0117] Among them, the NLP basic Q&A pair data can be regular translation tasks, text summarization, text rewriting, part-of-speech tagging and other conversations. The NLP basic Q&A pair data focuses on covering text data with objective measurement indicators to judge the results. The copywriting generation Q&A pair data can be text data for copywriting generation Q&A pair data, such as writing comments, writing press releases, brainstorming and other generation-type conversations. The copywriting generation Q&A pair data is used to cover text data with open-ended answers (the quality of the answers will vary depending on the evaluator). The text data of the domain professional Q&A pair data can be conversations involving keywords in various professional fields such as law, medicine, and physics. The short multi-round dialogue data can be Q&A pair data with multiple Q&A rounds. The logical reasoning Q&A pair data can be reasoning-type conversations such as sorting comparison, comprehensive analysis, and paradox problems.
[0118] The category labels of the basic Q&A pair data can include NLP basic category labels (corresponding to the NLP basic Q&A pair data in the Q&A pair data, that is Figure 2The NLP basic samples shown), the copywriting generation type target tags (corresponding to the copywriting generation Q&A pair data in the Q&A pair data, that is Figure 2 The copywriting generation samples shown), the domain professional type target tags (corresponding to the domain professional Q&A pair data in the Q&A pair data, that is Figure 2 The N domain professional samples shown) and the logical reasoning type target tags (corresponding to the logical reasoning Q&A pair data in the Q&A pair data, that is Figure 2 The logical reasoning samples shown), the service server 100 can generate the sample set type target tags corresponding to the training sample set in the first stage. For example, the training sample set in the first stage can include the training sample set composed of NLP basic Q&A pair data, the training sample set composed of copywriting generation Q&A pair data, the training sample set composed of domain professional Q&A pair data, and the training sample set composed of logical reasoning Q&A pair data. Among them, the sample set type target tags corresponding to the training sample set in the first stage can include the NLP basic sample set type target tags corresponding to the training sample set composed of NLP basic Q&A pair data, the copywriting generation sample set type target tags corresponding to the training sample set composed of copywriting generation Q&A pair data, the domain professional sample set type target tags corresponding to the training sample set composed of domain professional Q&A pair data, and the logical reasoning sample set type target tags corresponding to the training sample set composed of logical reasoning Q&A pair data.
[0119] The service server 100 obtains progressive Q&A pair data (which can also be called progressive thought chain corpus) in the basic Q&A pair samples. For example, it can be to gradually disassemble the steps of the answers in the basic Q&A pair data into multiple step rounds to obtain progressive Q&A pair data. The progressive Q&A pair data is determined as the training sample set in the second stage (which can include progressive thought chain samples), and the type target tags corresponding to the progressive Q&A pair data are determined as the sample set type target tags corresponding to the training sample set in the second stage. For example, it can be to determine the progressive thought chain sample set type target tags as the sample set type target tags corresponding to the progressive thought chain samples.
[0120] The service server 100 can obtain the first multi-round Q&A pair data in the Q&A pair data (for example, the Q&A pair data in the form of multi-round conversations in the original data). The service server 100 can perform data processing on the basic Q&A pair data through an evolutionary algorithm (which can be data splicing and data rewriting of the basic Q&A pair data) to obtain the second multi-round Q&A pair data (which can also be called multi-round evolutionary corpus). The service server 100 can determine the first multi-round Q&A pair data and the second multi-round Q&A pair data as the training sample set in the third stage. In addition, it can also generate multi-round type target tags and determine the multi-round type target tags as the sample set type target tags corresponding to the training sample set in the third stage. For example, it can be the evolutionary multi-round sample set type target tags (corresponding to the medium evolutionary multi-round samples such as Figure 2 The medium evolutionary multi-round samples shown), the complex evolutionary multi-round sample set type target tags (corresponding to such asFigure 2 the complex evolutionary multi-round samples shown) and the class labels of the GAtt sample set (corresponding to the GAtt samples shown as Figure 2 the GAtt samples shown).
[0121] The service server 100 can sort the training sample set according to the learning difficulty corresponding to the class labels of the sample set, and obtain the sorted training sample sequence. Among them, the training sample sets in the training sample sequence can be arranged in ascending order of learning difficulty.
[0122] The service server 100 can send the training sample sequence to the terminal device 10a. The terminal device 10a can, in the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model in the terminal device 10a converges after training to obtain the target question-and-answer model. For example, the target question-and-answer model can be used to generate question-and-answer service results through business questions.
[0123] In a certain iterative training of the initial question-and-answer model training, the terminal device 10a can first obtain the input training sample 1 from the training sample set corresponding to the first stage, adjust the model parameters of the initial question-and-answer model through the input training sample 1, then obtain the input training sample 2 from the training sample set corresponding to the second stage, adjust the model parameters of the initial question-and-answer model through the input training sample 2, and then obtain the input training sample 3 from the training sample set corresponding to the third stage, and adjust the model parameters of the initial question-and-answer model through the input training sample 3, so as to sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples.
[0124] It can be understood that the fine-tuning algorithm based on curriculum learning proposed in the embodiments of the present application can improve the generalization ability of the target question-and-answer model and can bring a better conversation experience. Based on the improvement of these capabilities, the product experiences in the following different fields can be significantly improved:
[0125] 1. Customer service: The target question-and-answer model can be used to create a chatbot. The chatbot can understand and answer customers' questions, process refund and return requests, provide product information, etc. In addition, the chatbot can also be used for telephone services, automatically answering calls and processing simple requests.
[0126] 2. Content creation: The target question-and-answer model can generate various types of content, including blog articles, news reports, social media posts, etc. The target question-and-answer model can generate content according to a given topic or keyword, or provide writing suggestions, such as providing different expressions of sentences, or providing possible plots of stories.
[0127] 3. Education: Targeted Q&A models can be used to create personalized learning resources, such as generating practice questions based on students' learning progress and comprehension abilities. They can also be used for online tutoring, understanding and answering students' questions, and providing detailed explanations and feedback.
[0128] 4. Translation and language learning: Targeted Q&A models can be used for real-time translation, understanding and translating texts in various languages. They can also be used for language learning, providing grammar and pronunciation suggestions to help learners understand and use new languages.
[0129] 5. Health consultation: Targeted Q&A models can be used to provide basic health and medical consultations, such as explaining symptoms, providing advice on a healthy lifestyle, answering questions about medications, etc. However, this cannot replace professional medical advice but serves as a supplementary resource.
[0130] 6. Personal assistants: Targeted Q&A models can be used to create intelligent personal assistants that can understand and answer various questions, such as querying the weather, setting reminders, managing schedules, etc. Targeted Q&A models can also be used to send emails and messages, conduct online shopping, etc.
[0131] 7. Games: Targeted Q&A models can be used to create richer and more in-depth gaming experiences. For example, they can generate conversations, drive non-player characters (NPCs), create complex storylines, etc. In addition, targeted Q&A models can also be used to understand and respond to players' inputs, providing dynamic gaming experiences.
[0132] The large language model is supervised and fine-tuned using the training data orchestration method based on curriculum learning proposed in the embodiments of this application. Compared with the data splicing method without using curriculum learning, the LLM (Large Language Model) trained by the present invention has improvements in the evaluations of five major dimensions: NLP basic capabilities, multi-turn dialogue capabilities, domain application capabilities, reasoning capabilities, and text generation capabilities. Among them, the reasoning and multi-turn dialogue capabilities have relatively significant improvements, with an improvement effect of more than 10%.
[0133] In the embodiments of the present application, the category information set of question-answer pairs is used to classify the question-answer pair data, and the category tags corresponding to the question-answer pair data are obtained. Then, based on the question-answer pair data and the category tags corresponding to the question-answer pair data respectively, a training sample set and the sample set category tags corresponding to the training sample set are generated. This can increase the number of training samples and enhance the diversity of the training samples. By sorting the training sample set according to the learning difficulty, during the process of model training, it can ensure that the model learns the features of the training samples with lower learning difficulty, and then through training with the training samples with higher learning difficulty, gradually help the model learn complex features, so as to achieve the same accuracy as using the entire training sample sequence, avoid overfitting of the model in the training samples with higher difficulty, generalize more complex and diverse question-answer scenarios, and improve the efficiency of model training and the accuracy of the model.
[0134] Please refer to Figure 3 , Figure 3 which is a schematic flow chart of a data processing method provided by the embodiments of the present application. Figure 1 This data processing method can be executed by a computer device, and the computer device can be a model server integrated with the service server 100 and the terminal device 10a as shown in Figure 1 . Hereinafter, it will be described by taking this data processing method being executed by a computer device as an example. Among them, this data processing method can at least include the following steps S101-S104:
[0135] Step S101, obtain M question-answer pair data and the question-answer pair category information set, and classify the M question-answer pair data based on the question-answer pair category information set to obtain the category tags corresponding to the M question-answer pair data respectively; the question-answer pair category information set includes category tags; M is a positive integer;
[0136] Specifically, the computer device can obtain the question-answer pair data and the question-answer pair category information set.
[0137] Among them, the question-answer pair data can be text data for training a large language model, including a series of questions and corresponding answers. The question-answer pair data can be obtained by manual annotation or automatic extraction, and can cover different fields and topics.
[0138] The question-answer pair category information set can be an information set for annotating the question-answer pair data, which can be divided into multiple category tags. The division of the category tags can be based on the learning difficulty of the question and answer, the field and topic of the question, the answer type and length, etc.
[0139] The computer device can classify the question-answer pair data based on the question-answer pair category information set to obtain the category tags corresponding to the question-answer pair data respectively.
[0140] For example, it can be to classify "{Q&A pair data 1, Q&A pair data 2,..., Q&A pair data N}" into the most relevant category tags through the above Q&A pair category information set, and the classification result can be {"Q&A pair data 1": NLP basic category tag, "Q&A pair data 2": copywriting generation category tag,..., "Q&A pair data N": multi-turn dialogue category tag}.
[0141] The category tags corresponding to the Q&A pair data can include NLP basic category tags, copywriting generation category tags, domain professional category tags, logical reasoning category tags, progressive thinking chain category tags, and multi-turn dialogue category tags.
[0142] Among them, the text data of NLP basic Q&A pair data can be conventional translation tasks, text summarization, text rewriting, part-of-speech tagging, etc. NLP basic Q&A pair data focuses on covering text data with objective measurement indicators to judge the results. For example, the NLP basic Q&A pair data can be in the text form of "Question: Classify the specified text. Please classify the following sentence as positive, negative, or neutral: This movie is very good. Answer: This movie is very good, positive."
[0143] The text data of copywriting generation Q&A pair data can be generation-based conversations such as writing reviews, writing press releases, brainstorming, etc. Copywriting generation Q&A pair data is used to cover text data with open-ended answers (the quality of the answers may vary depending on the evaluator). For example, the copywriting generation Q&A pair data can be in the text form of "Question: I want to write a one-sentence review for a favorite science fiction novel. How do you think it should be written? Book title: 'Science Fiction Novel A'; Author: Author A. Answer: Author A's 'Science Fiction Novel A' presents us with a breathtaking cosmic epic with its unique science fiction concepts, world view spanning time and space, and magnificent imagination."
[0144] The text data of domain professional Q&A pair data can be conversations involving keywords in various professional fields such as law, medicine, physics, etc. For example, the domain professional Q&A pair data can be in the text form of "Question: Please find out the diuretics from the following drugs: A. Amoxicillin; B. Dexamethasone; C. Furosemide; D. Aspirin. Answer: The diuretic is: C. Furosemide."
[0145] The text data of logical reasoning Q&A pair data can be reasoning-based dialogues such as sorting and comparison, comprehensive analysis, paradox problems, etc. For example, the logical reasoning Q&A pair data can be "Question: A unit organizes employees to participate in quality development training, and employees can sign up voluntarily. Lao Zhang met the new employee Li and talked about this matter. Lao Zhang reminded Li and said, 'The unit organizes quality development training. Hurry up and sign up.' Li said, 'I haven't finished dealing with my current work, so I don't need to sign up.' Except for which of the following, all can be the assumptions included in Li's answer? A. If the work is done, then sign up for the quality development training; B. As long as I haven't finished dealing with my work, I don't have to participate in the quality development training; C. Those who sign up for the quality development training have all finished dealing with their work; D. Only those who have finished dealing with their work sign up for the quality development training. Answer: First, we analyze Li's answer: 'I haven't finished dealing with my current work, so I don't need to sign up.' Next, we analyze the relationship between each option and Li's answer. Li bases on 'haven't finished dealing with the work' to get the conclusion 'don't need to sign up'. To make this conclusion hold, we need to find the relationship connecting 'haven't finished dealing with' and 'don't need to sign up', and we need to get 'don't need to sign up' from 'haven't finished dealing with'. Options B, C, and D can all be expressed as: 'haven't finished dealing with the work, don't sign up'. These options can all be used as bridging premises, but this question requires choosing the non-conforming ones, so we exclude BCD. And option A can be expressed as: 'don't need to sign up, haven't finished dealing with'. Here, the direction of bridging is opposite, so option A is not the premise of Li's answer. According to the requirements of the question, the correct answer is A." in text form.
[0146] Among them, the content about progressive thinking chain Q&A pair data and multi-round dialogue Q&A pair data can be seen below Figure 4a Specific descriptions of steps S203 to S205 corresponding to the embodiment.
[0147] It can be understood that the text data in the Q&A pair data can cross-cover multiple fields or topics, and there can be multiple related category tags. When the computer device classifies the Q&A pair data, it can classify the Q&A pair data into the most relevant D category tags. The D category tags can be three category tags, or if the degree of association between the Q&A pair data and the category tag is greater than a certain preset threshold, the Q&A pair data can be classified into this category tag. This application embodiment does not make restrictions here.
[0148] Step S102, based on M Q&A pair data and the category tags corresponding to the M Q&A pair data respectively, generate N training sample sets and the sample set category tags corresponding to the N training sample sets respectively; N is a positive integer;
[0149] Specifically, the computer device can generate training sample sets and the sample set category tags corresponding to the training sample sets respectively based on the Q&A pair data and the category tags corresponding to the Q&A pair data respectively.
[0150] The computer device can determine the single-round Q&A pair data of a certain basic field or theme (text data with one question and one answer in each dialogue turn) as the basic Q&A pair data. The basic field or theme can be NLP basic dialogue, copywriting generation dialogue, domain professional dialogue, and logical reasoning dialogue, which are not limited in the embodiments of the present application.
[0151] The computer device can divide the basic Q&A pair data into the training sample set of the first stage, and generate the sample set class target label corresponding to the training sample set of the first stage through the class target label of the basic Q&A pair data. It can also divide the Q&A pair data with the dialogue turn lower than a certain preset threshold into the training sample set of the first stage. For example, the training sample set corresponding to the short multi-round dialogue with the dialogue turn less than or equal to 8 turns can be divided into the training sample set of the first stage. The training sample set of the first stage can include the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain professional samples, the training sample set of short multi-round dialogues, and the training sample set of logical reasoning samples.
[0152] The computer device can obtain progressive Q&A pair data from the basic Q&A pair samples. For example, it can format the data of the basic Q&A pair data, or gradually disassemble the steps of the answer solution in the basic Q&A pair data into multiple step turns to obtain the progressive Q&A pair data. The computer device can determine the progressive Q&A pair data as the training sample set of the second stage and directly generate the sample set class target label corresponding to the training sample set of the second stage.
[0153] The computer device can obtain the first multi-round Q&A pair data from the Q&A pair data (for example, the Q&A pair data in the form of multi-round dialogue in the original data). The computer device can perform data processing on the basic Q&A pair data through an evolutionary algorithm. For example, it can splice and rewrite the data of the basic Q&A pair data to obtain the second multi-round Q&A pair data. The computer device can determine the first multi-round Q&A pair data and the second multi-round Q&A pair data as the training sample set of the third stage and directly generate the sample set class target label corresponding to the training sample set of the third stage.
[0154] Step S103: Sort the N training sample sets according to the learning difficulty corresponding to the sample set class target label to obtain a training sample sequence;
[0155] Specifically, the computer device can sort the training sample sets according to the learning difficulty corresponding to the sample set class target label from easy to difficult to obtain a training sample sequence. The training sample sequence can include multiple training sample sets, and the sorting order can be {the training sample set of the first stage, the training sample set of the second stage, the training sample set of the third stage} in sequence.
[0156] Among them, the learning difficulty can be the complexity of the data composition of the training samples, and the degree of adjustment of the model parameters during training, which may be related to factors such as the data volume, diversity, and quality of the training samples.
[0157] In the training sample set of each stage, the computer device can also sort the training sample set according to the learning difficulty corresponding to the sample set class label. For example, in the training sample set of the first stage, the sorting order can be, in sequence, {the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain professional samples, the training sample set of short multi-turn dialogues, the training sample set of logical reasoning samples}.
[0158] Step S104, in the training sample sequence, sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model converges after training to obtain a target question-and-answer model; the target question-and-answer model is used to generate a question-and-answer service result through a business question.
[0159] Specifically, in the training sample sequence, the computer device can sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples.
[0160] For example, the training sample sequence can be used as the dataset input for one epoch (training cycle) of training the initial question-and-answer model, and the training sample sets corresponding to each stage in the training sample sequence can be one batch (training batch) of training the initial question-and-answer model. One epoch can be a process of completing one forward calculation and backpropagation by inputting the training sample sequence into the initial question-and-answer model, and one epoch can also be called one iteration training.
[0161] In one epoch, the computer device can sequentially obtain input training samples from the training sample set of the first stage, the training sample set of the second stage, and the training sample set of the third stage. For example, it can obtain input training sample 1 from the training sample set of the first stage, adjust the model parameters of the initial question-and-answer model through input training sample 1, that is, complete one training batch of training the initial question-and-answer model, then obtain input training sample 2 from the training sample set of the second stage, adjust the model parameters of the initial question-and-answer model through input training sample 2, and then obtain input training sample 3 from the training sample set of the third stage, adjust the model parameters of the initial question-and-answer model through input training sample 3, that is, complete one training cycle of training the initial question-and-answer model.
[0162] If the initial Q&A model meets the convergence conditions indicated by the model evaluation metrics in this training cycle (iterative training), the initial Q&A model after convergence can be determined as the target Q&A model. The target Q&A model can be used to generate Q&A service results through business questions. The target Q&A model can also be used as a general Q&A model that has been fine-tuned through a training sample sequence. Subsequently, the general Q&A model can be trained with Q&A texts in a specific domain or topic, so that the trained Q&A model can generate better Q&A service results for business questions in a specific domain or topic.
[0163] Among them, the convergence conditions indicated by the model evaluation metrics can be performance metrics of the model. For example, they can be accuracy, precision, recall, precision-recall, etc. This application does not limit this here.
[0164] In the embodiments of this application, through the Q&A pair category information set, the Q&A pair data is classified by category to obtain the category labels corresponding to the Q&A pair data respectively. Then, through the Q&A pair data and the category labels corresponding to the Q&A pair data respectively, a training sample set and the sample set category labels corresponding to the training sample set are generated. The number of training samples can be increased, and the diversity of training samples can be increased. By sorting the training sample set according to the learning difficulty, during the process of model training, it can be ensured that the model learns the features of training samples with lower learning difficulty, and then through training with training samples with higher learning difficulty, it gradually helps the model learn complex features, so as to achieve the same accuracy as using the entire training sample sequence, and it can avoid the model from overfitting in training samples with higher difficulty, and can generalize more complex and diverse Q&A scenarios, improving the efficiency of model training and the accuracy of the model.
[0165] Please refer to Figure 4a , Figure 4a which is a schematic flowchart of a data processing method provided by the embodiments of this application Figure 2 , and this data processing method can be executed by a computer device, and the computer device can be a model server integrated by the business server 100 and the terminal device 10a as shown in Figure 1 . The following will take this data processing method being executed by a computer device as an example for illustration. Among them, this data processing method can at least include the following steps S201 - step S207:
[0166] Step S201, obtain M Q&A pair data and the Q&A pair category information set;
[0167] Specifically, reference can be made to the specific content of step S101 corresponding to the above Figure 3 embodiment, and the embodiments of this application will not elaborate here.
[0168] Step S202: Input the M question-answer pair data and the question-answer pair category information set into the sample partitioning model. In the sample partitioning model, generate a keyword vector based on the question-answer pair category information set, perform feature extraction on the M question-answer pair data, and obtain the feature word vectors corresponding to the M question-answer pair data respectively; based on the feature word vector and the keyword vector, determine the category labels corresponding to the M question-answer pair data respectively.
[0169] Specifically, the computer device can input the question-answer pair data and the question-answer pair category information set into the sample partitioning model.
[0170] Among them, the sample division model can be a converged large language model, and the question-answer pair category information set can be used as a text prompt for the question-answer pair data, and the question-answer pair data can be divided into categories to obtain corresponding category labels.
[0171] In the sample partitioning model, keyword vectors can be generated based on the question-answer pair category information set, and features can be extracted from the question-answer pair data to obtain feature word vectors corresponding to the question-answer pair data.
[0172] The feature extraction process of the question and answer pair data by the computer device can be: filtering business stop words for the M question and answer pair data, performing text segmentation on the filtered M question and answer pair data to obtain P question and answer text fragments; P is a positive integer greater than or equal to M; performing part-of-speech tagging on the P question and answer text fragments to obtain P question and answer part-of-speech information, and generating feature word vectors for the M question and answer pair data based on the P question and answer text fragments and the P question and answer part-of-speech information.
[0173] Specifically, the computer device can filter business stop words from the question-answer pair data. Business stop words refer to words or punctuation marks that are common in natural language texts but have no specific meaning in semantic analysis. For example, in Chinese texts, business stop words may include: "的", "了", "是", "在", "有", etc. In English texts, business stop words may include: "the", "a", "an", "in", etc.
[0174] The computer device can perform text segmentation (also known as word segmentation) on the training text set that removes business stop words. The text segmentation can obtain several question and answer text segments corresponding to each question and answer pair data through rule matching or probability matching.
[0175] The computer device can perform part-of-speech tagging on the Q&A text segment to obtain the Q&A part-of-speech information corresponding to the Q&A text segment. Part-of-Speech Tagging can be the grammatical part-of-speech classification of the Q&A text segment. For example, it can be divided into parts of speech such as nouns, verbs, adjectives, adverbs, and auxiliary words. By performing part-of-speech tagging, the grammatical structure and semantic information of the sentence can be better understood, thereby improving the accuracy and efficiency of natural language processing tasks. The part-of-speech tagging method can be a rule-based method, a statistic-based method, a deep learning-based method, etc., and the embodiments of the present application do not limit this here.
[0176] The computer device generates a feature word vector of the Q&A pair data based on the Q&A text segment and the Q&A part-of-speech information corresponding to the Q&A text segment. The computer device can determine the class target tags corresponding to the Q&A pair data respectively through the matching degree between the feature word vector and the keyword vector.
[0177] Step S203: Obtain the basic Q&A pair data from the M Q&A pair data, divide the basic Q&A pair data into the training sample set of the first stage, and determine the class target tag corresponding to the basic Q&A pair data as the sample set class target tag corresponding to the training sample set of the first stage;
[0178] Specifically, the computer device can determine the single-round Q&A pair data (text data with one question and one answer for the conversation round) of a certain basic domain or topic as the basic Q&A pair data. The basic domain or topic can be NLP basic conversations, copywriting generation conversations, domain professional conversations, and logical reasoning conversations, and the embodiments of the present application do not limit this here.
[0179] The computer device can divide the basic Q&A pair data into the training sample set of the first stage, and generate the sample set class target tag corresponding to the training sample set of the first stage through the class target tag of the basic Q&A pair data.
[0180] Optionally, the computer device can also divide the Q&A pair data of the short conversation into the training sample set of the first stage. The specific process can be: obtain the context information of the basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data; splice S mutually related basic Q&A pair data into a short conversation training sample according to the context information, add the short conversation training sample to the short conversation training sample set, and determine the short conversation label as the sample set class target tag corresponding to the short conversation training sample set; S is a positive integer.
[0181] Specifically, context information can be used to indicate the semantic correlation degree between basic Q&A pair data, and it can be the text content about the text and task that the model needs to understand (for example, in the sentiment analysis task, it can be a series of sentences with sentiment labels, such as "This product is very easy to use" and "This product is very bad", etc.), context relationships (for example, in the named entity recognition task, it can be the relationships between entities in the text, such as keywords, place names, time, etc.), semantic information (for example, in the part-of-speech tagging task, it can be the part of speech of each word), and other relevant information.
[0182] The computer device can splice a number of interrelated basic Q&A pair data into short dialogue training samples through context information. The short dialogue training samples can be Q&A pair data with the number of dialogue turns less than or equal to 8 turns, and one question and one answer form one turn.
[0183] The short dialogue training samples can be: "Question 1: What is the constellation of actress C? Answer 1: Actress C is a Libra. Question 2: You are so great! You know so much. Answer 2: I also know that actress C is the best actress at the Asian Film Awards! Question 3: Actress C is my idol and her acting skills are super good. Answer 3: She was truly born for acting and has also won the Best Actress at the Huayu Film Media Awards. Question 4: I really like the movies she acted in. Answer 4: I recommend you to watch 'The Equation of Love and Death' starring actress C. In this film, actress C dedicated extremely passionate and sincere performances, bringing the audience into the situation of 'loving until death', making people have many feelings about the city, about love, about gains and losses. Question 5: What is the word-of-mouth of this movie? Answer 5: Generally speaking, the word-of-mouth of this movie is good. The score of 'Film C' is 8.1 points (out of 10), and the number of reviewers exceeds 370,000. Many audiences think that the plot of this movie is compact, the emotions are sincere, and the characters are successfully shaped. Especially the performance of actress C has received high praise. At the same time, the movie's attention to real life and exploration of human nature have also been well received. Of course, some audiences think that the plot of the movie is a bit too strange or the ending is a bit abrupt. Generally speaking, 'Film C' is a movie worth watching. Question 6: Then I'll go and watch it now. Answer 6: Okay".
[0184] The computer device can add the short dialogue training samples to the short dialogue training sample set, determine the short dialogue label as the sample set class target label corresponding to the short dialogue training sample set, and divide the short dialogue training sample set into the training sample set of the first stage.
[0185] It can be understood that the training sample set of the first stage can include the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain professional samples, the training sample set of short multi-turn dialogues, and the training sample set of logical reasoning samples.
[0186] In step S204, progressive Q&A pair data is obtained from the M Q&A pair data. The progressive Q&A pair data is determined as the training sample set for the second stage, and the class target label corresponding to the progressive Q&A pair data is determined as the sample set class target label corresponding to the training sample set for the second stage; the learning difficulty of the class target label corresponding to the progressive Q&A pair data is greater than that of the class target label corresponding to the basic Q&A pair data.
[0187] Specifically, the computer device obtains progressive Q&A pair data from the basic Q&A pair samples. For example, the data formatting of the basic Q&A pair data can be performed, which can be to gradually break down the steps of the answer solution in the basic Q&A pair data into multiple step rounds to obtain the progressive Q&A pair data.
[0188] For example, the Q&A pair data B can be "Question: Xiaoming's family has 10 chicken legs. He wants to distribute them to 3 people. How many chicken legs does each person get? How many chicken legs are left? Answer: We need to find a number such that 10 chicken legs can be evenly distributed among 3 people. Divide 10 chicken legs by 3 people, getting 10 / 3 = 3 remainder 1. So, each person can get 3 chicken legs. Finally, there is 1 chicken leg left. The answer is that each person gets 3 chicken legs and there is 1 chicken leg left." The computer device can gradually break down the steps of the Q&A pair data B into multiple step rounds to obtain the disassembled steps. The disassembled steps can be "Question 1: Xiaoming's family has 10 chicken legs. He wants to distribute them to 3 people. How many chicken legs does each person get? How many chicken legs are left? Answer 1: Step 1: We need to find a number such that 10 chicken legs can be evenly distributed among 3 people. Question 2: Xiaoming's family has 10 chicken legs. He wants to distribute them to 3 people. How many chicken legs does each person get? How many chicken legs are left? Answer 2: Step 1: We need to find a number such that 10 chicken legs can be evenly distributed among 3 people; Step 2: Divide 10 chicken legs by 3 people, getting 10 / 3 = 3 remainder 1. Question 3: Xiaoming's family has 10 chicken legs. He wants to distribute them to 3 people. How many chicken legs does each person get? How many chicken legs are left? Answer 3: Step 1: We need to find a number such that 10 chicken legs can be evenly distributed among 3 people; Step 2: Divide 10 chicken legs by 3 people, getting 10 / 3 = 3 remainder 1; Step 3: So, each person can get 3 chicken legs. Question 4: Xiaoming's family has 10 chicken legs. He wants to distribute them to 3 people. How many chicken legs does each person get? How many chicken legs are left? Answer 4: Step 1: We need to find a number such that 10 chicken legs can be evenly distributed among 3 people; Step 2: Divide 10 chicken legs by 3 people, getting 10 / 3 = 3 remainder 1; Step 3: So, each person can get 3 chicken legs; Step 4: Finally, there is 1 chicken leg left. The answer is that each person gets 3 chicken legs and there is 1 chicken leg left."
[0189] It can be understood that the disassembling step may include 4 question-and-answer rounds. The computer device may ignore questions 2, 3, and 4, and determine questions 1, answer 1, answer 2, answer 3, and answer 4 as progressive question-and-answer pair data B. The specific way of disassembling the steps may be through Stepwise formatting, which is not limited in this embodiment of the present application.
[0190] The computer device may determine the progressive question-and-answer pair data as the training sample set in the second stage (the training sample set in the second stage may include progressive thinking chain samples), and determine the sample set class target label corresponding to the training sample set in the second stage as the progressive thinking chain sample class target label.
[0191] It can be understood that in the answer of the progressive question-and-answer pair data, multiple conditions in the question need to be analyzed, and the derivation is gradually carried out in the answer to increase accuracy. For this, progressive thinking chain samples containing a series of intermediate reasoning steps are obtained for these questions. Compared with using a single computational answer, this form is clearer, and the templatized and standardized responses can also stimulate the thinking mode of the model when encountering such questions, that is, how to flexibly use a large amount of prior knowledge and comprehensively analyze information to accurately answer questions.
[0192] Step S205, obtain the first multi-round question-and-answer pair data from the M question-and-answer pair data, generate the second multi-round question-and-answer pair data according to the basic question-and-answer pair data, determine the first multi-round question-and-answer pair data and the second multi-round question-and-answer pair data as the training sample set in the third stage, generate multi-round class target labels, and determine the multi-round class target labels as the sample set class target labels corresponding to the training sample set in the third stage; the learning difficulty corresponding to the multi-round class target labels is greater than the learning difficulty corresponding to the class target labels of the progressive question-and-answer pair data.
[0193] Specifically, the computer device may obtain the first multi-round question-and-answer pair data from the question-and-answer pair data. The first multi-round question-and-answer pair data may be question-and-answer pair data with the number of question-and-answer rounds in the original question-and-answer pair data greater than 8 rounds.
[0194] The computer device may generate the second multi-round question-and-answer pair data through the basic question-and-answer pair data. The process may be: perform splicing processing on the basic question-and-answer pair data to obtain spliced multi-round question-and-answer pair data; perform data reconstruction on the basic question-and-answer pair data to obtain reconstructed multi-round question-and-answer pair data; determine the spliced multi-round question-and-answer pair data and the reconstructed multi-round question-and-answer pair data as the second multi-round question-and-answer pair data.
[0195] Specifically, the computer device may perform data processing on the basic question-and-answer pair data through an evolutionary algorithm. For example, it may perform data splicing and data rewriting on the basic question-and-answer pair data to obtain the second multi-round question-and-answer pair data.
[0196] The process of a computer device splicing basic Q&A pair data can be as follows: Obtain the context information corresponding to H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data; obtain Q basic Q&A pair data according to the context information, and splice the Q basic Q&A pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation degree threshold; obtain T basic Q&A pair data according to the context information, and splice the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation degree threshold, and the first correlation degree threshold is greater than or equal to the second correlation degree threshold; determine the follow-up dialogue data and the random dialogue data as the spliced multi-round Q&A pair data.
[0197] Specifically, the computer device can obtain Q basic Q&A pair data through the context information and splice the Q basic Q&A pair data into follow-up dialogue data. Among them, Q is a positive integer less than or equal to H, Q can be a positive integer greater than 8, the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation degree threshold, and the first correlation degree threshold can indicate that there are the same attention keywords between the Q&A pair data, that is, the topics between two rounds of conversations are related. For example, the computer device can splice 14 basic Q&A pair data into a follow-up dialogue data through the context information, and the number of dialogue rounds of the follow-up dialogue data can be 14 rounds.
[0198] For example, the following conversation data could be: "Question 1: From now on, you need to always act as Character E. I hope you can act like him. I hope you answer and respond in Character E's tone, manner, and vocabulary. Don't write any explanations, just answer like Character E. You must know all the knowledge about Character E. After that, you must follow the personality, preferences, speaking style, and language of the character you are playing to answer. Where did you first meet Harry Potter and Ron Weasley? Answer 1: I first met Character E on the train, and then the three of us became friends. Question 2: How did you view Character E in that situation? Answer 2: When I first met Character E, I must admit I was full of curiosity about him. After all, his name was already well-known. But seeing him and getting to know him were two different things. I didn't immediately take a liking to him. In fact, I initially thought he might be an arrogant person, but after getting to know him, I realized that Character E was a brave, selfless, and kind friend. Every time he went on an adventure, it was to protect the people he loved. He didn't use his fame; instead, he hoped to have a peaceful life. So, I gradually began to respect and appreciate Character E. Question 3: As your friend, what are the personalities of Character E that you especially like? Answer 3: Character E's determination and courage have always impressed me deeply. No matter what difficulties he faces, he never gives up. He is loyal, upright, and always willing to do the right thing for others. (Omit the middle 10 rounds of conversation) Question 14: Who are you? Answer 14: I am Character E."
[0199] The computer device can obtain T basic Q&A pair data through context information and splice the T basic Q&A pair data into random conversation data.
[0200] Among them, T is a positive integer less than or equal to H, and T can be greater than 8. The semantic correlation degree between the T basic Q&A pair data is less than the second correlation degree threshold, and the first correlation degree threshold is greater than or equal to the second correlation degree threshold, that is, the basic Q&A pair data required for splicing in the follow-up conversation data has a greater semantic correlation degree than the basic Q&A pair data required in the random conversation data. In the random conversation data, the computer device can control the sample content types spliced to become gradually diverse and gradually increase the conversation length. The basic Q&A pair data can be of the same type of conversation with a large topic jump. This process can also randomly splice consecutive basic Q&A pairs of the same type, which can be used to increase the anti-interference ability of the model conversation.
[0201] For example, the computer device can splice 12 basic Q&A pair data into a random conversation data through context information, and the conversation rounds of the random conversation data can be 12 rounds.
[0202] The random conversation data can be "Question 1: Please arrange the following actions in the correct order: open the door, take off shoes, hang up the coat, wash hands. Answer 1: This order may vary from person to person, but usually it might be like this: 1. Open the door; 2. Take off shoes; 3. Hang up the coat; 4. Wash hands. Question 2: What is dynamic proxy? What are its applications? Answer 2: Dynamic proxy is to generate proxy classes dynamically at runtime. The applications of dynamic proxy include data query, the backend of testing frameworks, obtaining Java annotation objects, etc. Question 3: Which has a larger volume, a 2-ton monitor or 2 tons of feathers? Answer 3: The density of feathers is less than that of the monitor. Therefore, for the same weight, the volume of feathers is larger than that of the monitor. So, the volume of 2 tons of feathers is larger. (8 rounds of conversations in the middle are omitted) Question 12: Which has a larger volume, a 7-kilogram laptop or 7 kilograms of toilet paper? Answer 12: The density of toilet paper is less than that of the laptop. So, for the same weight, the volume of toilet paper is larger than that of the laptop. Therefore, the volume of 7 kilograms of toilet paper is larger."
[0203] The computer device can determine the concatenated multi-round Q&A pair data from the follow-up conversation data (which can also be called complex evolutionary multi-round samples) and the random conversation data (which can also be called medium evolutionary multi-round samples), divide the concatenated multi-round Q&A pair data into the training samples corresponding to the third stage, determine the sample set class target label of the follow-up conversation data as the complex evolutionary sample set class target label, and determine the sample set class target label of the random conversation data as the medium evolutionary sample set class target label.
[0204] The process by which the computer device reprocesses the basic Q&A pair data can be as follows: Based on the evolutionary constraint information, rewrite the basic Q&A pair data into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary class target labels included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary class target labels; obtain the context information of the basic Q&A pair data, and rewrite the mutually related basic Q&A pair data into broad evolutionary data according to the context information; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data; determine the deep evolutionary data and the broad evolutionary data as the reconstructed multi-round Q&A pair data.
[0205] Specifically, the computer device can rewrite the basic Q&A pair data into deep evolutionary data based on the evolutionary constraint information.
[0206] Among them, the evolutionary constraint information is used to indicate the evolutionary class target labels included in the Q&A pair category information set. For example, the computer device can perform deep data rewriting on the domain-specific Q&A pair data F and the Q&A pair data G for copywriting generation to obtain the deep evolutionary data K.
[0207] The Q&A pair data F can be "Question: Please find out the diuretics from the following drugs: A. Amoxicillin; B. Dexamethasone; C. Furosemide; D. Aspirin. Answer: The diuretic is: C. Furosemide.", and the Q&A pair data G can be "Question: What are the medicinal effects of diuretics? Answer: It has diuretic effects, reduces blood pressure, improves heart failure, and treats renal insufficiency, etc.".
[0208] The deep evolution data K obtained by performing deep data rewriting can be "Question 1: What are the diuretics? Answer 1: Furosemide. Question 2: What are the medicinal effects of furosemide? Answer 2: Furosemide has diuretic effects, reduces blood pressure, improves heart failure, and treats renal insufficiency, etc.".
[0209] Optionally, when the number of Q&A rounds of the deep evolution data K obtained by data rewriting is less than the Q&A round threshold, the computer device can continue to perform data rewriting on the deep evolution data K. The computer device can rewrite the basic Q&A pair data into transitional evolution data based on the evolution constraint information. If the number of Q&A rounds of the transitional evolution data is less than the Q&A round threshold, then based on the evolution constraint information, continue to perform data rewriting on the transitional evolution data until the number of Q&A rounds of the transitional evolution data after data rewriting is greater than or equal to the Q&A round threshold, and determine the transitional evolution data after data rewriting as the deep evolution data.
[0210] The computer device can obtain the context information of the basic Q&A pair data and rewrite the interrelated basic Q&A pair data into broad evolution data through the context information. For example, the broad evolution data L is obtained by performing broad data rewriting on the Q&A pair data F in the field of expertise and the Q&A pair data G generated by the copywriting.
[0211] The broad evolution data L obtained by performing broad data rewriting can be "Question 1: What are the medicinal effects of Amoxicillin, Dexamethasone, Furosemide, and Aspirin? Answer 1: The medicinal effect of Amoxicillin is used to treat bacterial infections,..., and the medicinal effect of Furosemide is to have diuretic effects, reduce blood pressure, improve heart failure, and treat renal insufficiency, etc.".
[0212] The computer device can determine the deep evolution data and the broad evolution data as the reconstructed multi-round Q&A pair data (which can also be called the GAtt sample), divide the reconstructed multi-round Q&A pair data into the training sample set of the third stage, and determine the sample set class label of the reconstructed multi-round Q&A pair data as the GAtt sample set class label.
[0213] It can be understood that breadth evolution aims to enhance the topic coverage, skill coverage of question-and-answer pair data, and the diversity of the overall dataset. Depth evolution can gradually increase the difficulty to ensure that the generated instructions are challenging. The depth-evolved data obtained through depth evolution and the breadth-evolved data obtained through breadth evolution can optimize the model performance and robustness during the training process.
[0214] Step S206: Sort the N training sample sets according to the learning difficulty corresponding to the class target label of the sample set to obtain a training sample sequence.
[0215] Specifically, the computer device can sort the training sample sets according to the learning difficulty corresponding to the class target label of the sample set to obtain a training sample sequence. The sorting order of the training sample sequence can be successively {training sample sets in the first stage: {training sample sets of NLP basic samples, training sample sets of copywriting generation samples, training sample sets of domain-specific samples, training sample sets of short multi-turn dialogue samples, training sample sets of logical reasoning samples}, training sample sets in the second stage: {training sample sets of progressive thinking chain samples}, training sample sets in the third stage: {training sample sets of medium-evolved multi-turn samples, training sample sets of complex-evolved multi-turn samples, training sample sets of GAtt samples}}.
[0216] Optionally, when the short multi-turn dialogue samples are composed of relatively complex question-and-answer pair data, the learning difficulty of the short multi-turn dialogue samples is greater than that of the logical reasoning samples. The curriculum arrangement of the question-and-answer pair data can adopt a combination of coarse-grained and fine-grained curriculum arrangements. Generally speaking, it can be divided into four-stage curriculum learning. The first stage is the learning of regular single-round samples, the second stage is the learning of simple (with fewer rounds) multi-turn dialogue samples, the third stage is the learning of progressive thinking chain samples, and the fourth stage is the learning of complex (with more rounds) multi-turn dialogue samples. In the multi-turn dialogue fields of the second stage and the fourth stage, the difficulty of the multi-turn dialogue samples can be defined from three dimensions: topic richness, length, and depth. Manage from the source of data collection according to the measurement dimensions of the multi-turn dialogue difficulty.
[0217] In the above-mentioned curriculum arrangement method combining coarse-grained and fine-grained, the training sample sequence of three stages can be further subdivided into four stages, that is, the short multi-turn dialogue samples in the first stage of the training sample sequence of three stages can be separately used as the training sample set of the second stage, and the short multi-turn dialogue samples are determined as the subsequent training stage after the first stage, obtaining a training sample sequence of four stages. The sorting order of the training sample sequence of four stages can be successively {the training sample set of the first stage: {the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain professional samples, the training sample set of logical reasoning samples}, the training sample set of the second stage: {the training sample set of short multi-turn dialogues}, the training sample set of the third stage: {the training sample set of progressive thinking chain samples}, the training sample set of the fourth stage: {the training sample set of medium evolution multi-turn samples, the training sample set of complex evolution multi-turn samples, the training sample set of GAtt samples}}. In the process of iterative training below, the order of obtaining the input training samples from the training sample sequence divided into four stages can be successively the training sample set of the first stage, the training sample set of the second stage, the training sample set of the third stage, and the training sample set of the fourth stage.
[0218] Step S207, in the training sample sequence, sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-answering model through the sequentially obtained input training samples until the target question-answering model is obtained after the initial question-answering model converges; the target question-answering model is used to generate a question-answering service result through a business question.
[0219] Specifically, in the training sample sequence, the computer device can sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial question-answering model through the sequentially obtained input training samples, so as to perform several rounds of iterative training.
[0220] For easy understanding, the j-th round of iterative training is taken as an example for illustration. The process of the j-th round of iterative training can be: in the j-th round of iterative training, randomly obtain B first input training samples in the training sample set A i and adjust the model parameters of the initial question-answering model through the B first input training samples, and continue in the training sample set A i+1Randomly obtain C second input training samples from it, and adjust the model parameters of the initial Q&A model through the C second input training samples until, when the adjustment of the model parameters of the initial Q&A model is completed using the input training samples in the Nth training sample set in the training sample sequence, it is determined that the jth round of iterative training is completed; B and C are positive integers; j is a positive integer; if the jth round of iterative training is completed and the initial Q&A model meets the model convergence condition, then the initial Q&A model that meets the model convergence condition is determined as the target Q&A model; if the jth round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then through the training sample sequence, in the (j + 1)th round of iterative training, continue to sequentially adjust the model parameters of the initial Q&A model using the input training samples obtained in order until the target Q&A model is obtained after the initial Q&A model converges during training.
[0221] Specifically, the training process of the initial Q&A model can include multiple iterative trainings (epochs, also referred to as training cycles), and each iterative training can include multiple training batches (batches).
[0222] In one epoch, the computer device can sequentially obtain input training samples from the training sample set in the first stage, the training sample set in the second stage, and the training sample set in the third stage. For example, it can obtain input training sample 1 from the training sample set in the first stage, and adjust the model parameters of the initial Q&A model through input training sample 1, that is, complete one training batch of the initial Q&A model training. Then obtain input training sample 2 from the training sample set in the second stage, adjust the model parameters of the initial Q&A model through input training sample 2, and then obtain input training sample 3 from the training sample set in the third stage, and adjust the model parameters of the initial Q&A model through input training sample 3, that is, complete one training cycle of the initial Q&A model training.
[0223] For ease of understanding, taking the jth round of iterative training as an example, the training sample sequence can include training sample set A i and training sample set A i+1 , in the training sample sequence, training sample set A i is located before training sample set A i+1 previously.
[0224] In the jth round of iterative training, the computer device can randomly obtain B first input training samples from training sample set A i , and adjust the model parameters of the initial Q&A model through the B first input training samples, and continue to obtain B first input training samples from training sample set A i+1Randomly obtain C second input training samples from it, and adjust the model parameters of the initial Q&A model through the C second input training samples until, when the model parameters of the initial Q&A model are adjusted by the input training samples in the Nth training sample set in the training sample sequence, it is determined that the jth round of iterative training is completed.
[0225] Training sample set A i Can be a training sample set of NLP basic samples, training sample set A i+1 Can be a training sample set of copywriting generation samples. When the model parameters of the initial Q&A model are adjusted by the input training samples in the training sample set of GAtt samples and the adjustment is completed, it is determined that the jth round of iterative training is completed.
[0226] If the jth round of iterative training is completed and the initial Q&A model meets the model convergence condition, then the initial Q&A model that meets the model convergence condition is determined as the target Q&A model. The target Q&A model can be used as a sample partitioning model that has been fine-tuned through the training sample sequence and can perform category partitioning for model training in a specific domain or topic.
[0227] Among them, the convergence condition indicated by the model evaluation index can be a performance metric of the model, such as accuracy, precision, recall, recall rate, etc., which are not limited in this application.
[0228] If the jth round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then through the training sample sequence, in the (j + 1)th round of iterative training, continue to adjust the model parameters of the initial Q&A model in sequence through the input training samples obtained in order, such as continuing to obtain the input training samples in the training sample set of NLP basic samples, the input training samples in the training sample set of copywriting generation samples,..., the input training samples in the training sample set of GAtt samples in the training sample sequence until the target Q&A model is obtained after the initial Q&A model converges.
[0229] Please also refer to Figure 4b , Figure 4b which is a schematic diagram of a data processing scenario provided by an embodiment of this application Figure 2 , such as Figure 4b shown, the computer device can obtain training samples for training and perform iterative training on the initial Q&A model until the initial Q&A model meets the model convergence condition.
[0230] In the iterative training of the ith round, the computer device can randomly obtain the input training samples of the first stage from the training sample sequence in the training sample set of the first stage and adjust the model parameters of the initial Q&A model through the input training samples of the first stage.
[0231] Among them, the training sample set in the first stage includes the training sample set of NLP basic samples, the training sample set of copywriting generation samples, the training sample set of domain-specific samples, the training sample set of short multi-turn dialogues, and the training sample set of logical reasoning samples. The computer device can randomly obtain the input training sample 1 from the training sample set of NLP basic samples in sequence, adjust the model parameters of the initial Q&A model through the input training sample 1, then randomly obtain the input training sample 2 from the training sample set of copywriting generation samples, and adjust the model parameters of the initial Q&A model through the input training sample 2,..., finally randomly obtain the input training sample 5 from the training sample set of logical reasoning samples, and adjust the model parameters of the initial Q&A model through the input training sample 5 to complete the adjustment of the model parameters of the initial Q&A model in the first stage.
[0232] Then, through the training sample sequence, randomly obtain the input training sample in the second stage from the training sample set in the second stage, and adjust the model parameters of the initial Q&A model through the input training sample in the second stage. It can be randomly obtain the input training sample 6 from the training sample set of progressive thinking chain samples, and adjust the model parameters of the initial Q&A model through the input training sample 6 to complete the adjustment of the model parameters of the initial Q&A model in the second stage.
[0233] Then, through the training sample sequence, randomly obtain the input training sample in the third stage from the training sample set in the third stage, and adjust the model parameters of the initial Q&A model through the input training sample in the second stage.
[0234] Among them, the training sample set in the third stage can include the training sample set of medium evolution multi-turn samples, the training sample set of complex evolution multi-turn samples, and the training sample set of GAtt samples. The computer device can randomly obtain the input training sample 7 from the training sample set of medium evolution multi-turn samples in sequence, adjust the model parameters of the initial Q&A model through the input training sample 7, then randomly obtain the input training sample 8 from the training sample set of complex evolution multi-turn samples, and adjust the model parameters of the initial Q&A model through the input training sample 8, and finally randomly obtain the input training sample 9 from the training sample set of GAtt samples, and adjust the model parameters of the initial Q&A model through the input training sample 9 to complete the adjustment of the model parameters of the initial Q&A model in the third stage.
[0235] When the initial Q&A model completes the adjustment of model parameters in the first stage, the second stage, and the third stage, it can be determined that the initial Q&A model completes the i-th round of iterative training. The computer device can determine whether the initial Q&A model meets the model convergence condition. If the model convergence condition is not met, the (i + 1)-th round of iterative training can be performed (it can be to continue to obtain input training samples 10, input training samples 11,..., input training samples 18). Through the training sample sequence, new input training samples are re-obtained from the training sample set in the first stage to perform iterative training on the initial Q&A model until the initial Q&A model meets the model convergence condition.
[0236] It can be understood that in iterative training of different rounds, for example, the input training sample 1 obtained from the training sample set of NLP basic samples in the i-th iterative training is different from the input training sample 10 obtained from the training sample set of NLP basic samples in the (i + 1)-th iterative training. The sample order obtained from iterative training of different rounds can be randomly distributed. For example, the input training sample 1 can include training sample A, training sample B, and training sample C. The input training sample 2 can include training sample B, training sample C, and training sample A. It is also possible not to sample all samples in one round of iterative training, and the embodiments of the present application do not limit this here.
[0237] Optionally, when obtaining input training samples from the training sample sequence in different iterative trainings (training cycles), the sorting order of the training sample sets in each training stage of the training sample sequence can also be randomly disrupted. For example, the training sample sequence in the i-th round of iterative training can be successively {Training sample set in the first stage: {Training sample set of NLP basic samples, Training sample set of copywriting generation samples, Training sample set of domain professional samples, Training sample set of logical reasoning samples, Training sample set of short multi-round dialogues}, Training sample set in the second stage: {Training sample set of progressive thinking chain samples}, Training sample set in the third stage: {Training sample set of medium evolution multi-round samples, Training sample set of complex evolution multi-round samples, Training sample set of GAtt samples}}. For example, the training sample sequence in the (i + 1)-th round of iterative training can be successively {Training sample set in the first stage: {Training sample set of logical reasoning samples, Training sample set of domain professional samples, Training sample set of copywriting generation samples, Training sample set of NLP basic samples, Training sample set of short multi-round dialogues}, Training sample set in the second stage: {Training sample set of progressive thinking chain samples}, Training sample set in the third stage: {Training sample set of GAtt samples, Training sample set of medium evolution multi-round samples, Training sample set of complex evolution multi-round samples}}.
[0238] In the embodiments of the present application, the category information set of question-and-answer pairs is used to classify the question-and-answer pair data into category labels corresponding to the question-and-answer pair data respectively. Then, based on the question-and-answer pair data and the category labels corresponding to the question-and-answer pair data respectively, a training sample set and sample set category labels corresponding to the training sample set are generated. The number of training samples can be increased, and the diversity of training samples can be enhanced. By sorting the training sample set according to the learning difficulty, during the model training process, the model can be ensured to learn the features of training samples with lower learning difficulty, and then trained with training samples with higher learning difficulty, gradually helping the model learn complex features, so as to achieve the same accuracy as using the entire training sample sequence, avoiding overfitting of the model in training samples with higher difficulty, generalizing more complex and diverse question-and-answer scenarios, and improving the efficiency of model training and the accuracy of the model. On the other hand, through data splicing and data rewriting, the theme coverage, skill coverage of the question-and-answer pair data and the diversity of the overall data set can be enhanced, and the model performance and robustness can be optimized during the training process.
[0239] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. As Figure 5 shown, the data processing device 1 includes a category division module 510, a sample generation module 520, a sorting processing module 530, and a model training module 540.
[0240] The category division module 510 is configured to obtain M question-and-answer pair data and a question-and-answer pair category information set, and classify the M question-and-answer pair data based on the question-and-answer pair category information set to obtain category labels corresponding to the M question-and-answer pair data respectively; the question-and-answer pair category information set includes category labels; M is a positive integer; the specific function of the category division module 510 can refer to the specific description of step S101 in the corresponding Figure 3 embodiment above, which will not be elaborated here.
[0241] The sample generation module 520 is configured to generate N training sample sets and sample set category labels corresponding to the N training sample sets respectively based on the M question-and-answer pair data and the category labels corresponding to the M question-and-answer pair data respectively; N is a positive integer; the specific function of the sample generation module 520 can refer to the specific description of step S102 in the corresponding Figure 3 embodiment above, which will not be elaborated here.
[0242] The sorting processing module 530 is configured to sort the N training sample sets through the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence; the specific function of the sorting processing module 530 can refer to the specific description of step S103 in the corresponding Figure 3 embodiment above, which will not be elaborated here.
[0243] The model training module 540 is used to sequentially obtain input training samples from the sorted training sample set in the training sample sequence, and sequentially adjust the model parameters of the initial Q&A model through the input training samples obtained in sequence until the target Q&A model is obtained after the training of the initial Q&A model converges; the target Q&A model is used to generate Q&A service results through business questions. For the specific functions of the model training module 540, reference can be made to the specific description of step S104 in the corresponding embodiment above, which will not be elaborated here. Figure 3 For the specific description of step S104 in the corresponding embodiment, it will not be elaborated here.
[0244] In a possible implementation manner, when the category division module 510 is used to perform category division on M Q&A pair data based on the Q&A pair category information set to obtain the category labels corresponding to the M Q&A pair data respectively, it is specifically used to perform the following operations:
[0245] Input the M Q&A pair data and the Q&A pair category information set into the sample division model. In the sample division model, generate keyword vectors based on the Q&A pair category information set, extract features from the M Q&A pair data, and obtain the feature word vectors corresponding to the M Q&A pair data respectively;
[0246] Based on the feature word vectors and the keyword vectors, determine the category labels corresponding to the M Q&A pair data respectively.
[0247] In a possible implementation manner, when the category division module 510 is used to extract features from the M Q&A pair data to obtain the feature word vectors of the M Q&A pair data respectively, it is specifically used to perform the following operations:
[0248] Filter the business stop words from the M Q&A pair data, and perform text division on the filtered M Q&A pair data to obtain P Q&A text segments; P is a positive integer greater than or equal to M;
[0249] Perform part-of-speech tagging on the P Q&A text segments to obtain P Q&A part-of-speech information, and generate the feature word vectors of the M Q&A pair data based on the P Q&A text segments and the P Q&A part-of-speech information.
[0250] In a possible implementation manner, the N training sample sets include the training sample set in the first stage, the training sample set in the second stage, and the training sample set in the third stage; when the sample generation module 520 is used to generate the N training sample sets and the sample set category labels corresponding to the N training sample sets respectively based on the M Q&A pair data and the category labels corresponding to the M Q&A pair data respectively, it is specifically used to perform the following operations:
[0251] Obtain the basic Q&A pair data from the M Q&A pair data, divide the basic Q&A pair data into the training sample set in the first stage, and determine the category label corresponding to the basic Q&A pair data as the sample set category label corresponding to the training sample set in the first stage;
[0252] Obtain progressive Q&A pair data from the M Q&A pair data, determine the progressive Q&A pair data as the training sample set for the second stage, and determine the class target label corresponding to the progressive Q&A pair data as the sample set class target label corresponding to the training sample set for the second stage; the learning difficulty of the class target label corresponding to the progressive Q&A pair data is greater than that of the class target label corresponding to the basic Q&A pair data.
[0253] Obtain the first multi-turn Q&A pair data from the M Q&A pair data, generate the second multi-turn Q&A pair data based on the basic Q&A pair data, determine the first multi-turn Q&A pair data and the second multi-turn Q&A pair data as the training sample set for the third stage, generate multi-turn class target labels, and determine the multi-turn class target labels as the sample set class target labels corresponding to the training sample set for the third stage; the learning difficulty of the multi-turn class target labels is greater than that of the class target label corresponding to the progressive Q&A pair data.
[0254] In a possible implementation, the training sample set for the first stage further includes a short dialogue training sample set; the sample generation module 520 is further configured to perform the following operations:
[0255] Obtain the context information of the basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data.
[0256] According to the context information, splice S mutually related basic Q&A pair data into short dialogue training samples, add the short dialogue training samples to the short dialogue training sample set, and determine the short dialogue label as the sample set class target label corresponding to the short dialogue training sample set; S is a positive integer.
[0257] In a possible implementation, when the sample generation module 520 is used to generate the second multi-turn Q&A pair data based on the basic Q&A pair data, it is specifically configured to perform the following operations:
[0258] Perform splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data;
[0259] Perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data;
[0260] Determine the spliced multi-turn Q&A pair data and the reconstructed multi-turn Q&A pair data as the second multi-turn Q&A pair data.
[0261] In a possible implementation, the number of basic Q&A pair data is H, where H is a positive integer less than or equal to M; when the sample generation module 520 is used to perform splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data, it is specifically configured to perform the following operations:
[0262] Obtain the context information corresponding to H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data;
[0263] Obtain Q basic Q&A pair data according to the context information, and splice the Q basic Q&A pair data into following dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation degree threshold;
[0264] Obtain T basic Q&A pair data according to the context information, and splice the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation degree threshold, and the first correlation degree threshold is greater than or equal to the second correlation degree threshold;
[0265] Determine the following dialogue data and the random dialogue data as the spliced multi-round Q&A pair data.
[0266] In a possible implementation manner, when the sample generation module 520 is used to perform data reconstruction on the basic Q&A pair data to obtain the reconstructed multi-round Q&A pair data, it is specifically used to perform the following operations:
[0267] Based on the evolutionary constraint information, rewrite the basic Q&A pair data into deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary category tags included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary category tags;
[0268] Obtain the context information of the basic Q&A pair data, and rewrite the mutually related basic Q&A pair data into breadth evolutionary data according to the context information; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data;
[0269] Determine the deep evolutionary data and the breadth evolutionary data as the reconstructed multi-round Q&A pair data.
[0270] In a possible implementation manner, when the sample generation module 520 is used to rewrite the basic Q&A pair data into deep evolutionary data based on the evolutionary constraint information, it is specifically used to perform the following operations:
[0271] Based on the evolutionary constraint information, rewrite the basic Q&A pair data into transitional evolutionary data;
[0272] If the number of Q&A rounds of the transitional evolutionary data is less than the Q&A round threshold, then based on the evolutionary constraint information, continue to perform data rewriting on the transitional evolutionary data until the number of Q&A rounds of the rewritten transitional evolutionary data is greater than or equal to the Q&A round threshold, and determine the rewritten transitional evolutionary data as the deep evolutionary data.
[0273] In a possible implementation manner, the training sample sequence includes the training sample set A iand the training sample set A i+1 In the training sample sequence, the training sample set A i is located before the training sample set A i+1 Before; the model training module 540 is used to sequentially obtain input training samples from the sorted training sample sets in the training sample sequence, and sequentially adjust the model parameters of the initial question-and-answer model through the input training samples obtained in sequence until the target question-and-answer model is obtained after the initial question-and-answer model training converges. Specifically, it is used to perform the following operations:
[0274] In the j-th round of iterative training, randomly obtain B first input training samples in the training sample set A i and adjust the model parameters of the initial question-and-answer model through the B first input training samples. Then continue to randomly obtain C second input training samples in the training sample set A i+1 and adjust the model parameters of the initial question-and-answer model through the C second input training samples until the adjustment of the model parameters of the initial question-and-answer model by the input training samples in the N-th training sample set in the training sample sequence is completed, and it is determined that the j-th round of iterative training is completed; B and C are positive integers; j is a positive integer;
[0275] If the j-th round of iterative training is completed and the initial question-and-answer model meets the model convergence condition, then the initial question-and-answer model that meets the model convergence condition is determined as the target question-and-answer model;
[0276] If the j-th round of iterative training is completed and the initial question-and-answer model does not meet the model convergence condition, then through the training sample sequence, in the (j + 1)-th round of iterative training, continue to sequentially adjust the model parameters of the initial question-and-answer model through the input training samples obtained in sequence until the target question-and-answer model is obtained after the initial question-and-answer model training converges.
[0277] Please refer to Figure 6 , Figure 6 is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 6As shown in the figure, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 6 shown, in the memory 1005, as a computer-readable storage medium, there may be included an operating system, a network communication module, a user interface module, and a device control application program.
[0278] In the Figure 6 computer device 1000 as shown in the figure, the network interface 1004 can provide network communication network elements; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to achieve:
[0279] Obtain M pairs of question-and-answer data and a question-and-answer category information set, perform category division on the M pairs of question-and-answer data based on the question-and-answer category information set, and obtain category labels corresponding to the M pairs of question-and-answer data respectively; the question-and-answer category information set includes category labels; M is a positive integer;
[0280] Based on the M pairs of question-and-answer data and the category labels corresponding to the M pairs of question-and-answer data respectively, generate N training sample sets and sample set category labels corresponding to the N training sample sets respectively; N is a positive integer;
[0281] Sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence;
[0282] In the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model converges after training to obtain a target question-and-answer model; the target question-and-answer model is used to generate a question-and-answer service result through a business question.
[0283] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the foregoing Figure 3 andFigure 4a The description of the data processing method in any corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either.
[0284] In addition, it should be noted here that: The embodiments of the present application also provide a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium. When the above-mentioned processor executes the above-mentioned computer program, it can execute the content described in any one of the previous Figure 3 and Figure 4a corresponding embodiments of the data processing method. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0285] The above-mentioned computer-readable storage medium may be the data processing device provided in any of the previous embodiments or the internal storage unit of the above-mentioned computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been displayed or will be displayed.
[0286] In addition, it should be noted here that: The embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in any one of the previous Figure 3 and Figure 4a corresponding embodiments.
[0287] In the description, claims, and drawings of the embodiments of the present application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products, or equipment.
[0288] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described in terms of network elements in the above description. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described network elements for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0289] The methods and related devices provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in a process Figure 1 a process or multiple processes and / or structural schematic Figure 1 a block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in a process Figure 1 a process or multiple processes and / or structural schematic Figure 1 a block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a process Figure 1 a process or multiple processes and / or structural schematic a block or multiple blocks.
[0290] The steps in the method of the embodiment of the present application can be adjusted, combined and deleted according to actual needs.
[0291] The modules in the device of the embodiment of the present application can be combined, divided and deleted according to actual needs.
[0292] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain M question-and-answer pair data and a question-and-answer pair category information set, perform category division on the M question-and-answer pair data based on the question-and-answer pair category information set, and obtain category labels corresponding to the M question-and-answer pair data respectively; The question-and-answer pair category information set includes the category labels; M is a positive integer; Based on the M question-and-answer pair data and the category labels corresponding to the M question-and-answer pair data respectively, generate N training sample sets and sample set category labels corresponding to the N training sample sets respectively; N is a positive integer; Sort the N training sample sets according to the learning difficulty corresponding to the sample set category labels to obtain a training sample sequence; In the training sample sequence, sequentially obtain input training samples from the sorted training sample sets, and sequentially adjust the model parameters of the initial question-and-answer model through the sequentially obtained input training samples until the initial question-and-answer model converges after training to obtain a target question-and-answer model; the target question-and-answer model is used to generate a question-and-answer service result through a business question.
2. The method according to claim 1, wherein The performing category division on the M question-and-answer pair data based on the question-and-answer pair category information set to obtain category labels corresponding to the M question-and-answer pair data respectively includes: Input the M question-and-answer pair data and the question-and-answer pair category information set into a sample division model. In the sample division model, generate keyword vectors based on the question-and-answer pair category information set, perform feature extraction on the M question-and-answer pair data, and obtain feature word vectors corresponding to the M question-and-answer pair data respectively; Based on the feature word vectors and the keyword vectors, determine the category labels corresponding to the M question-and-answer pair data respectively.
3. The method according to claim 2, wherein The performing feature extraction on the M question-and-answer pair data to obtain the feature word vectors of the M question-and-answer pair data includes: Filter business stop words from the M question-and-answer pair data, perform text division on the filtered M question-and-answer pair data to obtain P question-and-answer text segments; P is a positive integer greater than or equal to M; Perform part-of-speech tagging on the P question-and-answer text segments to obtain P question-and-answer part-of-speech information, and generate the feature word vectors of the M question-and-answer pair data based on the P question-and-answer text segments and the P question-and-answer part-of-speech information.
4. The method according to claim 1, wherein The N training sample sets include a first-stage training sample set, a second-stage training sample set, and a third-stage training sample set; The generating N training sample sets and sample set category labels corresponding to the N training sample sets respectively based on the M question-and-answer pair data and the category labels corresponding to the M question-and-answer pair data respectively includes: Obtain basic question-and-answer pair data from the M question-and-answer pair data, divide the basic question-and-answer pair data into the first-stage training sample set, and determine the category label corresponding to the basic question-and-answer pair data as the sample set category label corresponding to the first-stage training sample set; Obtain progressive Q&A pair data from the M Q&A pair data, determine the progressive Q&A pair data as the training sample set for the second stage, and determine the class target label corresponding to the progressive Q&A pair data as the sample set class target label corresponding to the training sample set for the second stage; the learning difficulty of the class target label corresponding to the progressive Q&A pair data is greater than that of the class target label corresponding to the basic Q&A pair data. Obtain the first multi-turn Q&A pair data from the M Q&A pair data, generate the second multi-turn Q&A pair data according to the basic Q&A pair data, determine the first multi-turn Q&A pair data and the second multi-turn Q&A pair data as the training sample set for the third stage, generate multi-turn class target labels, and determine the multi-turn class target labels as the sample set class target labels corresponding to the training sample set for the third stage; the learning difficulty of the multi-turn class target labels is greater than that of the class target label corresponding to the progressive Q&A pair data.
5. The method according to claim 4, wherein The training sample set for the first stage further includes a short dialogue training sample set; the method further includes: Obtain the context information of the basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data. According to the context information, splice S mutually related basic Q&A pair data into short dialogue training samples, add the short dialogue training samples to the short dialogue training sample set, and determine the short dialogue label as the sample set class target label corresponding to the short dialogue training sample set; S is a positive integer.
6. The method according to claim 4, wherein The generating the second multi-turn Q&A pair data according to the basic Q&A pair data includes: Perform splicing processing on the basic Q&A pair data to obtain spliced multi-turn Q&A pair data. Perform data reconstruction on the basic Q&A pair data to obtain reconstructed multi-turn Q&A pair data. Determine the spliced multi-turn Q&A pair data and the reconstructed multi-turn Q&A pair data as the second multi-turn Q&A pair data.
7. The method according to claim 6, characterized in that, The number of the basic Q&A pair data is H, and H is a positive integer less than or equal to M; the performing splicing processing on the basic Q&A pair data to obtain the spliced multi-turn Q&A pair data includes: Obtain the context information corresponding to H basic Q&A pair data; the context information is used to indicate the semantic correlation degree between the H basic Q&A pair data. Obtain Q basic Q&A pair data according to the context information, and splice the Q basic Q&A pair data into follow-up dialogue data; Q is a positive integer less than or equal to H; the semantic correlation degree between the Q basic Q&A pair data is greater than the first correlation threshold. Obtain T basic Q&A pair data according to the context information, and splice the T basic Q&A pair data into random dialogue data; T is a positive integer less than or equal to H; the semantic correlation degree between the T basic Q&A pair data is less than the second correlation threshold, and the first correlation threshold is greater than or equal to the second correlation threshold. Determine the follow-up dialogue data and the random dialogue data as the spliced multi-turn Q&A pair data.
8. The method according to claim 6, wherein The performing data reconstruction on the basic Q&A pair data to obtain the reconstructed multi-turn Q&A pair data includes: Based on the evolutionary constraint information, rewrite the basic Q&A pair data into the deep evolutionary data; the evolutionary constraint information is used to indicate the evolutionary class tags included in the Q&A pair category information set; the deep evolutionary data is associated with the evolutionary class tags; Obtain the context information of the basic Q&A pair data, and rewrite the mutually related basic Q&A pair data into the broad evolutionary data according to the context information; the context information is used to indicate the semantic correlation degree between the basic Q&A pair data; Determine the deep evolutionary data and the broad evolutionary data as the reconstructed multi-round Q&A pair data.
9. The method according to claim 8, characterized in that The rewriting of the basic Q&A pair data into deep evolutionary data based on the evolutionary constraint information includes: Based on the evolutionary constraint information, rewrite the basic Q&A pair data into transitional evolutionary data; If the number of Q&A rounds of the transitional evolutionary data is less than the Q&A round threshold, then based on the evolutionary constraint information, continue to rewrite the transitional evolutionary data until the number of Q&A rounds of the rewritten transitional evolutionary data is greater than or equal to the Q&A round threshold, and determine the rewritten transitional evolutionary data as the deep evolutionary data.
10. The method according to claim 1, characterized in that, The training sample sequence includes training sample set A i and training sample set A i+1 , in the training sample sequence, training sample set A i is located before training sample set A i+1 ; In the training sample sequence, sequentially obtain input training samples from the sorted training sample set, and sequentially adjust the model parameters of the initial Q&A model through the sequentially obtained input training samples until the initial Q&A model converges after training to obtain the target Q&A model, including: In the j-th round of iterative training, among the training sample set A i randomly obtain B first input training samples, and adjust the model parameters of the initial Q&A model through the B first input training samples. Then continue to randomly obtain C second input training samples in the training sample set A i+1 and adjust the model parameters of the initial Q&A model through the C second input training samples. Until when the adjustment of the model parameters of the initial Q&A model is completed by using the input training samples in the N-th training sample set in the training sample sequence, it is determined that the j-th round of iterative training is completed; B and C are positive integers; j is a positive integer; If the j-th round of iterative training is completed and the initial Q&A model meets the model convergence condition, then determine the initial Q&A model that meets the model convergence condition as the target Q&A model; If the j-th round of iterative training is completed and the initial Q&A model does not meet the model convergence condition, then in the (j + 1)-th round of iterative training through the training sample sequence, continue to sequentially adjust the model parameters of the initial Q&A model through the sequentially obtained input training samples until the initial Q&A model converges after training to obtain the target Q&A model.
11. A data processing device, characterized in that, Including: A category division module, configured to obtain M Q&A pair data and a Q&A pair category information set, and perform category division on the M Q&A pair data based on the Q&A pair category information set to obtain the class tags corresponding to the M Q&A pair data respectively; The Q&A pair category information set includes the class tags; M is a positive integer; A sample generation module, configured to generate N training sample sets and the sample set class tags corresponding to the N training sample sets respectively based on the M Q&A pair data and the class tags corresponding to the M Q&A pair data respectively; N is a positive integer; A sorting processing module, configured to sort the N training sample sets through the learning difficulty corresponding to the sample set class tags to obtain a training sample sequence; A model training module, configured to sequentially obtain input training samples from the sorted training sample set in the training sample sequence, and sequentially adjust the model parameters of the initial Q&A model through the sequentially obtained input training samples until the initial Q&A model converges after training to obtain the target Q&A model; the target Q&A model is used to generate a Q&A service result through a business question.
12. A computer device, characterized in that, Including: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method described in any one of claims 1-10.
14. A computer program product, characterized in that, The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium and is suitable for being read and executed by a processor so that a computer device having the processor executes the method described in any one of claims 1-10.