Voice data processing method and device based on large model, storage medium and electronic device
By amplifying and classifying the marked data sets, high-quality positive and negative sample data sets are generated, and the weighted loss function is used to optimize model training, the problem of low data quality in the existing data amplification methods is solved, and the text processing accuracy of the model is improved.
Patent Information
- Application Number
- CN202510768691.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The quality of speech data generated by existing data amplification methods is relatively low, especially in complex semantic tasks. The generated sample data is inconsistent with the original data intention, which affects the accuracy of model training.
By amplifying the marked data sets, classifying the data sets into positive and negative samples using a parsed model, generating preference-optimized data sets, and using weighted loss function optimization model training, combining sample weights and original loss function to generate high-quality amplified data.
Improves the diversity and quality of data, reduces training offsets caused by low-quality samples, and improves the accuracy and intent understanding of the model when processing new input text data.
Smart Images

Figure CN120279893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech data processing. Specifically, it relates to a method, apparatus, storage medium, and electronic device for speech data processing based on a large model. Background Art
[0002] Currently, in the related fields of using large models to process speech data, the processing ability of large models depends on high-quality training data. However, the collection of speech data in the real world is often limited by time and cost, resulting in insufficient data volume available for training large models. To address this challenge, data processing means such as data augmentation have emerged. Data augmentation refers to generating additional training samples based on the original data set to increase richer training data, thereby improving the generalization ability and accuracy of the model. Existing data augmentation methods are mainly divided into two categories. The first category of methods focuses on surface modification of the original training data, such as adding random noise, back-translation, etc. to increase the diversity of the training set. However, this type of method relies on predefined rules or limited synonym libraries and is difficult to break out of the framework of existing data, resulting in deficiencies in data diversity. The second category of methods, such as using generative language models like generative adversarial networks to generate new sample data similar to the distribution of the original data, but the data generated by this type of method is not stable in quality. For example, in tasks with highly complex semantics, it will generate sample data with correct grammar but inconsistent semantic intentions with the original data, reducing the accuracy of the model using these sample data for training.
[0003] Therefore, in the related art, there is a technical problem that the data quality obtained by existing data augmentation methods is relatively low.
[0004] In view of the technical problem that the data quality obtained by existing data augmentation methods is relatively low in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, storage medium, and electronic device for speech data processing based on a large model, so as to at least solve the technical problem that the data quality obtained by existing data augmentation methods is relatively low in the related art.
[0006] According to an embodiment of the present application, there is provided a method for processing speech data based on a large model, including: performing data augmentation operations on the labeled data of the labeled data set to obtain an augmented training data set, wherein the labeled data is collected by a sound pickup device; using an analysis model to classify the augmented training data set into a positive sample data set and a negative sample data set, and generating a preference optimization data set by using the positive sample data set and the negative sample data set. A preference optimization sample in the preference optimization data set includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; generating a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the analysis model, obtaining a data augmentation model that completes a preset training target by using the weighted loss function, and using the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented data set.
[0007] In an exemplary embodiment, before using the analysis model to classify the augmented training data set into a positive sample data set and a negative sample data set, the method further includes: generating current labeled data based on the current labeling operation of the initial data by the target object, and generating the labeled data set according to the historical labeled data provided by the target object and the current labeled data, wherein the labeled data set at least includes the question sentence of the target object, the field labeled by the target object, the intention labeled by the target object, and the slot corresponding to the entity labeled by the target object; using the fine-tuning instruction of the target object to instruct the initial large language model to perform an identification operation on the labeled data, and outputting an identification result according to the output requirements of the instruction fine-tuning data in the fine-tuning instruction. The identification result at least includes: labeled field, labeled intention, labeled entity, and labeled slot.
[0008] In an exemplary embodiment, data augmentation operations are performed on the labeled data of the labeled dataset to obtain an augmented training dataset, including: selecting seed corpus data to be augmented from the labeled data, and determining the seed annotation domain, seed annotation intention, and target intention corresponding to the seed corpus according to the seed corpus data; using a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task of transforming from the seed annotation intention to the target intention on the seed corpus within the seed annotation domain, to obtain an augmented training dataset. The preset augmentation prompt statement at least includes the task information required for the initial large language model to perform the data augmentation task, and the task information at least includes task description information, role description information, task examples, and the input seed corpus. The role description information is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the seed annotation intention, and the new corpus has the target intention.
[0009] In an exemplary embodiment, using a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task of transforming from the seed annotation intention to the target intention on the seed corpus within the seed annotation domain, to obtain an augmented training dataset, including: for an annotation intention set belonging to the same seed annotation domain, screening the first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions in the annotation intention set, generating an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention, and a second seed annotation intention similar or opposite to the first seed annotation intention, and the intention pair is expressed as , is the first seed annotation intention, is the second seed annotation intention, and n is a positive integer; determining a first augmentation prompt statement from the preset extension prompt statement, where the first role description information in the first augmentation prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the first seed annotation intention; using the first augmentation prompt statement to instruct the initial large language model to perform a data augmentation operation with a similar intention for the seed corpus to obtain an augmented training dataset.
[0010] In an exemplary embodiment, a preset amplification prompt statement is used to instruct the initial large language model to perform a data amplification task of converting the seed corpus according to the seed annotation intention to the target intention within the seed annotation field, so as to obtain an amplified training data set, including: screening other annotation intentions except the first seed annotation intention from the set of annotation intentions, where the other annotation intentions only have similar intentions; determining a second expansion prompt statement from the preset expansion prompt statement, where the second role description information in the second expansion prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to the other annotation intentions; using the second expansion prompt statement to instruct the initial large language model to perform the seed corpus of a data amplification operation with similar intentions to obtain an amplified training data set.
[0011] In an exemplary embodiment, a preset amplification prompt statement is used to instruct the initial large language model to perform a data amplification task of converting the seed corpus according to the seed annotation intention to the target intention within the seed annotation field, so as to obtain an amplified training data set, including: for a set of annotation intentions belonging to the same seed annotation field, screening a first seed annotation intention with similar or opposite intentions from the seed annotation intentions in the set of annotation intentions to generate an intention pair corresponding to the first seed annotation intention, where the intention pair includes the first seed annotation intention and a second seed annotation intention similar or opposite to the first seed annotation intention, and the intention pair is expressed as , is the first seed annotation intention, is the second seed annotation intention, and n is a positive integer; determining a third expansion prompt statement from the preset expansion prompt statement, where the third role description information in the third expansion prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the second seed annotation intention; using the third expansion prompt statement to instruct the initial large language model to perform the seed corpus of a data amplification operation with opposite intentions to obtain an amplified training data set.
[0012] In an exemplary embodiment, the annotated data set is expressed as , is the seed corpus to be amplified, is the seed annotation field, is the seed annotation intention, and the amplified training data set is expressed as , as the target intention, , where i, M, and N are positive integers.
[0013] In an exemplary embodiment, using a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset includes: inputting the data of the augmented training dataset into the parsing model, so that the parsing model performs classification processing on the augmented training data, and outputs the predicted domain and predicted intention ; according to and to determine the positive sample dataset and the negative sample dataset, and the verification result includes the seed annotation domain and the first verification result of the predicted domain and the seed annotation intention and the second verification result of the predicted intention .
[0014] In an exemplary embodiment, determining the positive sample dataset according to the verification result of and includes: when it is determined that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention Figure 1 , storing the seed corpus corresponding to the seed annotation domain into the positive sample dataset according to the first mapping relationship , and the first mapping relationship represents a relationship with the corpus quadruple of the seed corpus to be augmented in the annotation data as the key and the positive sample dataset as the value, is the seed corpus to be augmented, is the seed annotation domain, is the seed annotation intention, is the target intention.
[0015] In an exemplary embodiment, determining the negative sample dataset according to the verification result of and includes: when it is determined that any one of the first verification result indicating that the seed annotation domain is consistent with the predicted domain and the second verification result indicating that the seed annotation intention is consistent with the predicted intention Figure 1 does not hold, storing the seed corpus corresponding to the seed annotation domain into the negative sample dataset according to the second mapping relationship , and the second mapping relationship represents a relationship with as the key and the negative sample dataset as the value.
[0016] In an exemplary embodiment, the method further includes: inputting the augmented training data set into the parsing model to enable the parsing model to identify the augmented training data and obtain the slot relationship corresponding to the augmented training data; replacing the slot relationship with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence pattern of the augmented training data to be the same as that of the seed corpus to be augmented, determining the updated augmented training data as positive sample candidate data, and determining the positive sample data set according to the positive sample candidate data.
[0017] In an exemplary embodiment, generating a preference optimization data set by using the positive sample data set and the negative sample data set, including: for a negative sample , collecting k positive samples from the positive sample data set according to a preset sample ratio of 1:k , , ; using the prompt word, the negative sample and the k positive samples to generate a preference optimization sample, and determining the preference optimization data set according to the sample set of all preference optimization samples, where the prompt word is generated by a corpus quadruple and a preset lead-in.
[0018] In an exemplary embodiment, collecting k positive samples from the positive sample data set according to a preset sample ratio of 1:k , including: calculating the cosine similarity between the negative sample and the k positive samples according to the vector encoding results of the negative sample and the k positive samples , and calculating the similarity probability distribution result of the negative sample with respect to the positive sample data set based on the cosine similarity; sampling the positive sample data set by using the similarity probability distribution result to obtain the k positive samples corresponding to the negative sample .
[0019] In an exemplary embodiment, calculating the cosine similarity between the negative sample and the k positive samples through the following formula:
[0020] ;
[0021] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
[0022] In an exemplary embodiment, calculating a similarity probability distribution result of the negative samples with respect to the positive sample dataset based on cosine similarity includes: calculating the similarity probability distribution result of the negative samples with respect to the positive sample dataset using a similarity probability distribution function of the negative samples with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0023] ;
[0024] constant ; the similarity probability distribution result is expressed as .
[0025] In an exemplary embodiment, the method further includes: representing a preference optimization sample in the form of a triple, then the preference optimization dataset is expressed as follows:
[0026] , denotes a prompt word.
[0027] In an exemplary embodiment, before generating a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, the method further includes: determining a sensitivity quantization index of the parsing model according to a perplexity parameter; normalizing a data subset in the preference optimization dataset using the sensitivity quantization index, and determining the result of the normalization as the sample weights, where the preference optimization samples in the data subset include prompt words with the same intention or prompt words with the opposite intention.
[0028] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , S represents a set input sequence of the parsing model, x represents an input prompt word for prompting the generation of a similar intention, denotes an input prompt word for prompting the generation of an opposite intention, y represents a response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the parsing model generating the response sequence y when x is the set input sequence, represents when taking as the set input sequence, the perplexity of the parsing model generating the response sequence y.
[0029] In an exemplary embodiment, normalizing a data subset in the preference optimization dataset using the sensitivity quantization index, and determining the result of the normalization as the sample weights, includes: obtaining a target quantization index generated according to the sensitivity quantization index , the target quantization index indicates , the sum of the sensitivity quantization indices for the prompts with the same intention or the prompts with the opposite intention ; and determining the result of normalizing with respect to the data subset as the sample weight, and the result of the normalization is expressed as follows: .
[0030] In an exemplary embodiment, obtaining a data augmentation model that completes the training objective using the weighted loss function includes: iteratively training a policy model with the minimization of the weighted loss function as the training objective, and stopping training the policy model and determining the policy model as the data augmentation model when the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is the smallest. The loss value of the weighted loss function represents the deviation degree between the policy model and the reference model.
[0031] According to another aspect of the embodiments of the present application, there is also provided a speech data processing device based on a large model, including: a first augmentation module for performing data augmentation operations on the labeled data of the labeled data set to obtain an augmented training data set; an output module for using an analysis model to classify the augmented training data set into a positive sample data set and a negative sample data set, generating a preference optimization data set using the positive sample data set and the negative sample data set. One preference optimization sample in the preference optimization data set includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; a second augmentation module for generating a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the analysis model, obtaining a data augmentation model that completes a preset training objective using the weighted loss function, and using the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented data set.
[0032] According to yet another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the above-mentioned speech data processing method based on a large model when running.
[0033] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned speech data processing method based on a large model.
[0034] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the above-mentioned speech data processing method based on a large model through the computer program.
[0035] In the embodiments of the present application, data augmentation operations are performed on the labeled data of the labeled data set to obtain an augmented training data set; an analysis model is used to classify the augmented training data set into a positive sample data set and a negative sample data set, and a preference optimization data set is generated by using the positive sample data set and the negative sample data set. A preference optimization sample in the preference optimization data set includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; a weighted loss function is generated according to the sample weight of the preference optimization sample and the original loss function of the analysis model, and a data augmentation model that completes a preset training target by using the weighted loss function is obtained, and the data augmentation model is used to perform data augmentation on newly input text data to obtain a target augmented data set; by adopting the above technical solution, data diversity can be improved by performing data augmentation on the labeled data. The augmented training data set is classified by the analysis model to obtain a positive sample data set and a negative sample data set, and then a preference optimization data set is constructed based on the relevance between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weight of the preference optimization sample in the preference optimization data set and the original loss function of the analysis model, the quality of different preference optimization samples can be reasonably quantified in a weighted manner, and a data augmentation model that completes a preset training target by using the weighted loss function is obtained, and the data augmentation model is used to complete the data augmentation of the newly input text data. By training the model with a weighted loss function that combines sample weights, the training deviation caused by low-quality samples can be reduced, and the model performance can be further optimized, enabling the model to more accurately understand and generate the corpus with the required intention when processing newly input text data, solving the technical problem of low data quality obtained by existing data augmentation methods, thereby improving the quality of the augmented data and also enhancing the accuracy of the model for text processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0038] Figure 1 is a schematic diagram of the hardware environment of a speech data processing method based on a large model according to an embodiment of the present application;
[0039] Figure 2 is a flowchart of a speech data processing method based on a large model according to an embodiment of the present application;
[0040] Figure 3 is a schematic diagram of a speech data processing method based on a large model according to an embodiment of the present application;
[0041] Figure 4 is a structural block diagram of a speech data processing device based on a large model according to an embodiment of the present application. Detailed implementation manners
[0042] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0044] According to one aspect of the embodiments of the present application, a speech data processing method based on a large model is provided. The speech data processing method based on a large model is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above-mentioned speech data processing method based on a large model can be applied to, for example, Figure 1 the hardware environment composed of the terminal device 102 and the server 104 as shown in Figure 1As shown, the server 104 is connected to the terminal device 102 through a network and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.
[0045] The above network can include, but is not limited to, at least one of the following: wired network, wireless network. The above wired network can include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart drying rack, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification equipment, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.
[0046] In this embodiment, a method for processing voice data based on a large model is provided, which is applied to the above terminal device. Figure 2 It is a flowchart of the method for processing voice data based on a large model according to an embodiment of the present application. The process includes the following steps:
[0047] Step S202: Perform a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set; wherein, the labeled data is collected by a sound pickup device.
[0048] Optionally, the labeled data can also be understood as the text-form data obtained by performing text conversion on the voice data collected by the sound pickup device. The labeled data includes, for example, the labeled field corresponding to the seed corpus, the labeling intention, the slots of the labeled entity, etc.
[0049] Step S204: Use the parsing model to classify the augmented training data set into a positive sample data set and a negative sample data set, and generate a preference optimization data set by using the positive sample data set and the negative sample data set. One preference optimization sample in the preference optimization data set includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on an initial large language model.
[0050] Step S206: Generate a weighted loss function based on the sample weights of the preference-optimized samples and the original loss function of the parsing model, obtain a data augmentation model that achieves a preset training objective using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset.
[0051] Through the above steps, perform data augmentation operations on the labeled data of the labeled dataset to obtain an augmented training dataset; use the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generate a preference-optimized dataset using the positive sample dataset and the negative sample dataset. A preference-optimized sample in the preference-optimized dataset includes a prompt word, negative sample data, and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on an initial large language model; generate a weighted loss function based on the sample weights of the preference-optimized samples and the original loss function of the parsing model, obtain a data augmentation model that achieves a preset training objective using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset; adopting the above technical solution, by performing data augmentation on the labeled data, data diversity can be improved. Use the parsing model to classify the augmented training dataset to obtain a positive sample dataset and a negative sample dataset, and then construct a preference-optimized dataset based on the correlation between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weights of the preference-optimized samples in the preference-optimized dataset and the original loss function of the parsing model, the quality of different preference-optimized samples can be reasonably quantified in a weighted manner, and obtain a data augmentation model that achieves a preset training objective using the weighted loss function, and use the data augmentation model to complete the data augmentation of newly input text data. By training the model with a weighted loss function that combines sample weights, the training deviation caused by low-quality samples can be reduced, the model performance can be further optimized, so that when the model processes newly input text data, it can more accurately understand and generate the corpus with the required intention, solve the technical problem of the low data quality obtained by existing data augmentation methods, thereby improving the quality of the augmented data and also enhancing the accuracy of the model for text processing.
[0052] In an exemplary embodiment, before classifying the augmented training dataset into a positive sample dataset and a negative sample dataset using the parsing model, further, current labeled data can be generated based on the current annotation operation of the initial data by the target object, and the labeled dataset can be generated according to the historical labeled data provided by the target object and the current labeled data, where the labeled dataset at least includes the question sentences of the target object, the fields labeled by the target object, the intents labeled by the target object, and the slots corresponding to the entities labeled by the target object; use the fine-tuning instruction of the target object to instruct the initial large language model to perform an identification operation on the labeled data, and output an identification result according to the output requirements of the instruction fine-tuning data in the fine-tuning instruction, where the identification result at least includes: labeled field, labeled intent, labeled entity, and labeled slot.
[0053] Optionally, the labeled field, labeled intent, labeled entity, and labeled slot in the above identification result respectively correspond to the field labeled by the target object, the intent labeled by the target object, the entity labeled by the target object, and the slot corresponding to the entity labeled by the target object, that is, the labeled data represents the data that is expected to be recognized by the initial large language model.
[0054] Optionally, the target object can be a person or an entity company. For example, if the target object is an online retail company, the question sentences involve aspects such as product query, order management, and customer service. Through the fine-tuning instruction, it can guide the expected initial large model to focus on the learning of these specific fields and intents, so as to generate labeled data that better meets the business requirements.
[0055] In an alternative embodiment, the large model fine-tuning process is described in combination with the following embodiments. The large model fine-tuning process is the process of using the fine-tuning instruction of the target object to instruct the initial large language model to perform an identification operation on the labeled data and output an identification result according to the output requirements of the instruction fine-tuning data in the fine-tuning instruction. Specifically, the LLaMA2-base-7B model is used as the initial large language model, and a seed dataset is obtained through manual construction and online data annotation. Then, the LLaMA2-base-7B model is fine-tuned using the seed dataset to train the model's ability to generate correct field intents and slots under given instructions. The fine-tuned model should have the ability to identify fields, intents, and slots, and be able to understand and execute specific business rules, such as semantic parsing tasks in the home appliance control scenario, providing a basis for subsequent data augmentation.
[0056] In an exemplary embodiment, a data augmentation operation is performed on the annotated data of an annotated data set to obtain an augmented training data set, including: selecting seed corpus data to be augmented from the annotated data, and determining the seed annotation field, seed annotation intent and target intent corresponding to the seed corpus according to the seed corpus data; using a preset augmentation prompt sentence to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus in the seed annotation field, to obtain an augmented training data set, the preset augmentation prompt sentence at least includes the task information required for the initial large language model to perform the data augmentation task, the task information at least includes task description information, role description information, task examples, and input seed corpus, the role description information is used to set the initial large language model to generate a new corpus according to the seed corpus and the seed annotation intent, and the new corpus has the target intent. This embodiment makes it easier for the large model to understand and perform complex intent conversion tasks by clearly giving the task description information, role description information, task examples and input seed corpus in the task information of the preset augmentation prompt sentence, greatly enriching the diversity of training data, thereby improving the generalization ability of the model.
[0057] It can be understood that, in order to determine the seed annotation domain, seed annotation intention and target intention corresponding to the seed corpus according to the seed corpus data, each seed corpus is actually obtained from the seed corpus data, and the seed annotation domain, seed annotation intention and target intention of each seed corpus are determined in turn. If the seed corpus is "I want to buy a pair of sneakers", the seed annotation intention is "buy goods", and the target intention is "return policy", then the preset augmentation prompt sentence is, for example: "Please generate a sentence asking about the return policy based on the following sentence", and a description of the relationship between "buy goods" and "return policy" is attached.
[0058] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus within the seed annotation domain, so as to obtain an augmented training data set, including: for a set of annotation intents belonging to the same seed annotation domain, a first seed annotation intent having a similar intent or an opposite intent is selected from the seed annotation intents of the annotation intent set, and an intent pair corresponding to the first seed annotation intent is generated, wherein the intent pair includes the first seed annotation intent and a second seed annotation intent that is similar to or opposite to the first seed annotation intent, and the intent pair is represented as , Label the first seed with intent, For the second seed annotation intention, where n is a positive integer; determine a first extended prompt statement from the preset extended prompt statements, and the first role description information in the first extended prompt is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention, and the target intention of the new corpus is similar to the first seed annotation intention; use the first extended prompt statement to instruct the initial large language model to target of the seed corpus to perform a data augmentation operation for similar intentions to obtain an augmented training data set. In this embodiment, symmetric intentions (i.e., similar intentions or opposite intentions) are represented by intention pairs. By introducing specific prompt statements, this embodiment guides the large language model to generate more relevant corpus in a given domain and intention, thereby enriching the training data set and improving the performance of the model on specific tasks.
[0059] Optionally, the first seed annotation intention represents the original intention , and the second seed annotation intention represents the symmetric intention of the original intention . Taking intention pairs as an example, for example, intention pairs include "open - close, increase temperature - decrease temperature, start - pause", then "open, increase temperature, start" correspond to the first seed annotation intention, "close, decrease temperature, pause" correspond to the second seed annotation intention, and the target intention of the new corpus is similar to the first seed annotation intention, and the target intention is, for example, "open, warm up, start".
[0060] In the above embodiment, by combining role description information and seed corpus, the model can learn a wider range of expression ways and intention changes, and solve the problem of data scarcity. For example, in a customer service dialogue system, by augmenting dialogue instances similar to or opposite to "query order status" (such as "cancel order"), the model's ability to understand and process such complex scenarios can be enhanced, improving the user experience.
[0061] In an exemplary embodiment, use the preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task from the seed annotation intention to the target intention on the seed corpus within the seed annotation domain, to obtain an augmented training data set, including: screening other annotation intentions except the first seed annotation intention from the set of annotation intentions, and the other annotation intentions only have similar intentions; determining a second extended prompt statement from the preset extended prompt statements, and the second role description information in the second extended prompt is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to the other annotation intentions; using the second extended prompt statement to instruct the initial large language model to target of the seed corpus Perform data augmentation operations on similar intents to obtain an augmented training dataset. In this embodiment, for the original intents where there are no symmetric intents in the annotation intent set, it is set that the initial large language model generates new corpus based on the seed corpus and these similar intents, ensuring that the target intent of the new corpus is similar to the original intent, avoiding generating corpus that is irrelevant or contrary to the original intent, and ensuring the quality of the augmented data.
[0062] Through the precise screening and the guidance of the role description information in the above embodiment, the model can create more variants while maintaining the domain and intent Figure 1 consistency, effectively increasing the diversity and coverage of the training data. For example, in an intelligent recommendation system, by augmenting user queries similar to "view product details", such as "understand product features", the model can be trained more comprehensively, enabling it to make more accurate recommendations when facing diverse user needs.
[0063] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus within the seed annotation domain, to obtain an augmented training dataset, including: for the annotation intent set belonging to the same seed annotation domain, screening the first seed annotation intent with similar or opposite intents from the seed annotation intents of the annotation intent set, generating an intent pair corresponding to the first seed annotation intent, the intent pair including the first seed annotation intent, and a second seed annotation intent similar or opposite to the first seed annotation intent, the intent pair is expressed as , is the first seed annotation intent, is the second seed annotation intent, and n is a positive integer; determining a third augmentation prompt statement from the preset extension prompt statement, the third role description information in the third augmentation prompt statement is used to set the function of the initial large language model to generate new corpus based on the seed corpus and the intent pair, and the target intent of the new corpus is similar to the second seed annotation intent; using the third augmentation prompt statement to instruct the initial large language model to perform data augmentation operations with opposite intents on the seed corpus to obtain an augmented training dataset. In this embodiment, for the first seed annotation intent with similar or opposite intents, an intent pair is generated, and then a third augmentation prompt statement is determined from the preset extension prompt statement, and its role description information guides the initial large language model to generate new corpus consistent with the second seed annotation Figure 1 intent, that is, the target intent is opposite to the original intent.
[0064] By clearly differentiating and augmenting corpora with opposite intentions, this embodiment enhances the model's ability to handle adversarial situations, which is crucial for building a dialogue system capable of understanding complex user intentions. For example, in the Q&A system of an online education platform, augmenting instances of "negative answers" and "positive answers" can help the model better distinguish between user questions and statements, improving the accuracy of interactions. Leveraging the generative capabilities of large models, through prompt words and role description information, this embodiment achieves efficient augmentation of existing labeled data, capable of generating not only samples with similar intentions but also samples with opposite intentions, thus constructing a more comprehensive and balanced training dataset.
[0065] Based on the above embodiment, the following implementation process is proposed: Select seed corpora containing specific intentions. For example, "lower the air conditioner temperature a bit". Then, construct a prompt, the role of which is to guide the model to generate text that conforms to the expected intention. For example, the prompt is "You are a voice assistant in an air conditioner control scenario and need to generate a statement with the same intention as the given query but with a different expression."
[0066] Input the seed corpus and the above prompt into the LLaMA2 model, facilitating the model to generate new statements with the same original intention as the seed corpus but with different expressions according to the instructions of the prompt. Figure 1 To ensure the diversity of the generated new statements, the temperature parameter during model generation can be adjusted. Temperature is a parameter that controls the randomness of the generated text. A lower temperature will result in the generated text being more inclined to the most likely output learned by the model, while a higher temperature will make the generated text more random and diverse. In this scenario, reasonably setting the temperature can increase the diversity of the generated statements while maintaining their consistency with the seed corpus.
[0067] Next, use manual inspection or domain-specific rules or an automatically calibrated semantic parsing model after fine-tuning to verify the generated new statements, ensuring that the new statements indeed conform to the expected intention and are reasonable in terms of grammar and context logic. Retain the new statements that pass the verification and express the same intention as the seed corpus but are not exactly the same as part of the augmented dataset. Figure 1 Through the above process, the generative capabilities of the LLaMA2 model and the guiding role of the prompt can be utilized to effectively augment text data with the same intention as the seed corpus but different expressions, thereby enriching the training set and improving the generalization ability and performance of the model in handling actual business scenarios.
[0068]
[0069]
[0070] In an exemplary embodiment, the labeled dataset is represented as , is the seed corpus to be amplified, is the seed annotation domain, is the seed annotation intention, and the amplified training dataset is represented as , is the target intention, , where i, M, and N are positive integers. In this embodiment, by structurally representing the labeled dataset, the composition of the dataset and the data representation after amplification are clearly defined, which helps the model understand and learn the associations between different domains and intentions, and can significantly improve the computing power of the model.
[0071] In an exemplary embodiment, an analysis model is used to classify the amplified training dataset into a positive sample dataset and a negative sample dataset, including: inputting the data of the amplified training dataset into the analysis model, so that the analysis model performs classification processing on the amplified training data and outputs the predicted domain and the predicted intention ; according to and to determine the positive sample dataset and the negative sample dataset, and the verification result includes the first verification result of the seed annotation domain and the predicted domain and the second verification result of the seed annotation intention and the predicted intention . In this embodiment, an analysis model is used to classify the amplified training dataset into a positive sample dataset and a negative sample dataset, including inputting the data of the amplified training dataset into the analysis model, outputting the predicted domain and the predicted intention, and determining the positive and negative sample datasets according to the verification result.
[0072] For example, for the amplified text , it is further input into the analysis model for classification to obtain the predicted domain and the predicted intention , and the predicted domain and the predicted intention are further denoted as .
[0073] The above embodiment uses the classification processing function of the analysis model to automatically screen out positive samples that meet the expected intention and negative samples that do not meet the expectation, reducing the workload of manual inspection and accelerating the speed of data amplification. For example, in a social media monitoring system, a large number of user comments can be quickly identified and classified, and comments that match the intention of "positive brand evaluation" are used as positive samples, while comments that are contrary to it are used as negative samples, which can effectively train the model to identify and filter negative information on the network.
[0074] In an exemplary embodiment, determining the positive sample dataset according to the verification results of and includes: when it is determined that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention Figure 1 , storing the seed corpus corresponding to the seed annotation domain into the positive sample dataset according to the first mapping relationship , where the first mapping relationship represents a relationship with the corpus quadruple of the seed corpus to be amplified in the annotation data as the key and the positive sample dataset as the value, is the seed corpus to be amplified, is the seed annotation domain, is the seed annotation intention, is the target intention. In this embodiment, determining the positive sample dataset according to the verification results of and includes storing the seed corpus into the positive sample dataset when the seed annotation domain is consistent with the predicted domain and the seed annotation intention is consistent with the predicted intention Figure 1 .
[0075] Optionally, maintaining a mapping table to represent the above first mapping relationship, using the corpus quadruple as the key value in the mapping table, and maintaining two lists and respectively to store candidate data that can be used as positive and negative samples. The preference optimization dataset is constructed by verifying whether the predicted domain intention is the same as the expected domain intention. For example, in this embodiment, the seed corpus with the seed annotation domain consistent with the predicted domain and the seed annotation intention consistent with the predicted intention Figure 1 is stored into the positive sample dataset to construct the positive sample dataset, improving the accuracy of the amplified data, effectively filtering out irrelevant or incorrect amplified data, and ensuring that the information learned by the model is of high quality and high relevance.
[0076] In an exemplary embodiment, determining the negative sample dataset according to the verification results of and includes: when it is determined that either the first verification result indicating that the seed annotation domain is consistent with the predicted domain or the second verification result indicating that the seed annotation intention is consistent with the predicted intention Figure 1 does not hold, storing the seed corpus corresponding to the seed annotation domain into the negative sample dataset according to the second mapping relationship , where the second mapping relationship represents with As the key, the negative sample dataset is the value. and The negative sample data set can be determined by the verification result, and the seed annotation field is consistent with the prediction field, and the seed annotation intention is consistent with the prediction intention. Figure 1 When any of the above is not true, the seed corpus is stored in the negative sample dataset. By constructing a negative sample dataset based on corpora with failed intent recognition or incorrect domain division, the model can be improved in a targeted manner to perform better when faced with complex or edge cases. For example, in the voice command recognition system of smart home devices, using "turn off the lights" corpora that are mistakenly recognized as "turn on the lights" commands as negative samples can prompt the model to learn to distinguish similar but opposite commands and improve user experience.
[0077] Optionally, the seed annotation domain is consistent with the predicted domain, and the seed annotation intent is consistent with the predicted intent. Figure 1 The situation, for example, is expressed as .
[0078] The seed annotation domain is consistent with the predicted domain, and the seed annotation intention is consistent with the predicted intention. Figure 1 If any of the above does not hold, it can be expressed as: ; ; .
[0079] In an exemplary embodiment, the following technical solution is further proposed: the augmented training data set is input into the parsing model so that the parsing model can identify the augmented training data and obtain the slot relationship corresponding to the augmented training data; the slot relationship is replaced with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data to the same sentence structure as the seed corpus to be augmented, the updated augmented training data is determined as the positive sample candidate data, and the positive sample data set is determined based on the positive sample candidate data. This embodiment can use the parsing model to identify the slots of the augmented training data, and replace the identified slots with the slot values of the seed corpus, thereby constructing positive sample data with the same sentence structure and adding it to the positive sample candidate set. In this embodiment, the training efficiency and prediction accuracy of the model are improved by unifying the sentence structure of the data, ensuring the structural consistency of the augmented data and the seed corpus. For example, in natural language processing tasks, by adjusting the sentence structure of the augmented data to the same framework as the seed corpus, the model can more easily capture the core of the semantics and avoid understanding bias caused by differences in sentence structure.
[0080] Optionally, the slot replacement method in the above embodiments can be applied, for example, to situations where the expected predictions are in different fields but have the same intention. For example, the expected prediction field is "Oven", the prediction intention is "increase Temperature", the extended corpus is "Make the temperature of the steamer a little higher", and then this corpus is input into the parsing model. With "Steamer" as the prediction field and "increase Temperature" as the prediction intention, the slot is the "(device)" corresponding to the steamer, and the slot relationship is the corresponding relationship between the steamer and "(device)". The corpus marked with the slot is "[(device) steamer] Make the temperature a little higher". At this time, the value of the (device) slot can be replaced with the device name corresponding to the prediction field, such as an oven. Then the replaced corpus is "Make the temperature of the oven a little higher". If the original seed corpus is "Increase the temperature of the steamer" and the seed slot relationship is the corresponding relationship between the steamer and "(device)", then replacing the value of the (device) slot with an oven, the resulting corpus is "Increase the temperature of the oven".
[0081] In an exemplary embodiment, generating a preference optimization dataset using the positive sample dataset and the negative sample dataset includes: for a negative sample , collecting k positive samples from the positive sample dataset according to a preset sample ratio of 1:k , , ; using the prompt word, the negative sample and the k positive samples to generate a preference optimization sample, and determining the preference optimization dataset according to the sample set of all preference optimization samples. The prompt word is generated by a corpus quadruple and a preset lead-in. In this embodiment, generating a preference optimization dataset using the positive sample dataset and the negative sample dataset includes, for each negative sample, collecting positive samples from the positive sample dataset proportionally, using the prompt word, the negative sample and the positive sample to generate a preference optimization sample, and then determining the preference optimization dataset. Through contrastive learning, the ability of the model to identify correct intentions and fields can be strengthened. For example, in the product search system of an online shopping platform, by comparing the user queries "Find the latest mobile phone" and "Find the cheapest mobile phone" (assuming the former is a positive sample and the latter is a negative sample), the model can learn how to distinguish the user's purchase tendency and provide more personalized search results.
[0082] In an exemplary embodiment, collecting k positive samples from the positive sample dataset according to a preset sample ratio of 1:k , includes: calculating the negative sample according to the vector encoding result of the negative sample and the vector encoding results of the k positive samples and k positive samples the cosine similarity between them, and calculate the similarity probability distribution result of the negative sample with respect to the positive sample dataset based on the cosine similarity; sample the positive sample dataset using the similarity probability distribution result to obtain negative samples the corresponding k positive samples .
[0083] Optionally, encode the vectors of the positive samples and the vectors of the negative samples respectively to obtain the vector encoding results of the negative samples and the vector encoding results of the positive samples, and calculate the similarity probability distribution results of each negative sample with respect to the positive sample set, and sample from the positive sample candidates using this distribution result , and then obtain k positive samples , and then k preference samples can be constructed
[0084] In an exemplary embodiment, the cosine similarity between the negative sample and k positive samples is calculated by the following formula:
[0085] ;
[0086] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample
[0087] In an exemplary embodiment, calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset based on the cosine similarity includes: calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset using the similarity probability distribution function of the negative sample with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0088] ;
[0089] constant ; the similarity probability distribution result is expressed as .
[0090] In an exemplary embodiment, the following technical solution is further proposed: representing a preference optimization sample in the form of a triple, then the preference optimization dataset is expressed as follows:
[0091] , represents a prompt
[0092] In an exemplary embodiment, before generating a weighted loss function based on the sample weights of the samples optimized according to preferences and the original loss function of the parsing model, the following technical solution is further proposed: determining a sensitivity quantization index of the parsing model according to a perplexity parameter; normalizing a data subset in the preference optimization dataset by using the sensitivity quantization index, and determining the result of the normalization as the sample weight, where the preference optimization samples in the data subset include cue words with the same intention or cue words with the opposite intention. It should be noted that the sample weight reflects the degree of importance the model attaches to different samples. In this embodiment, by calculating the perplexity difference, the sensitivity of the model to different cue words can be quantified, and then the influence of the samples in the training process can be adjusted. For example, in an intelligent translation system, by analyzing the perplexity difference of the model when processing two cue words of "formal language" and "informal language", the proportion of these two types of corpora in training can be dynamically adjusted, enabling the model to translate smoothly in both formal and informal situations.
[0093] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , S represents the set input sequence of the parsing model, x represents the input cue word for prompting the generation of a similar intention, represents the input cue word for prompting the generation of the opposite intention, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity when the parsing model generates the response sequence y with x as the set input sequence, represents when taking as the set input sequence, the perplexity when the parsing model generates the response sequence y.
[0094] In an exemplary embodiment, normalizing a data subset in the preference optimization dataset by using the sensitivity quantization index and determining the result of the normalization as the sample weight includes: obtaining a target quantization index generated according to the sensitivity quantization index , the target quantization index represents , respectively represents the sum of the sensitivity quantization indexes regarding cue words with the same intention or cue words with the opposite intention ; determining the result of normalizing regarding the data subset as the sample weight, and the result of the normalization is expressed as follows:
[0095] .
[0096] 。
[0097] In an exemplary embodiment, obtaining a data augmentation model that uses the weighted loss function to achieve the training objective includes: iteratively training a policy model with the minimization of the weighted loss function as the training objective, and stopping training the policy model and determining the policy model as the data augmentation model when the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is the smallest. The loss value of the weighted loss function represents the degree of deviation between the policy model and the reference model. In this embodiment, obtaining a data augmentation model that uses the weighted loss function to achieve the training objective includes iteratively training a policy model with the minimization of the weighted loss function as the training objective until the optimal state is reached. By assigning different weights to different samples and introducing the weighted loss function, the learning direction of the model can be more finely controlled during training, avoiding the overfitting phenomenon caused by over-focusing on certain samples. For example, in the obstacle recognition module of an autonomous driving system, by assigning higher weights to samples of key obstacles such as "pedestrians", "bicycles", and "cars", it can be ensured that the model has stronger recognition ability for these important categories, thereby improving driving safety.
[0098] It should be noted that in the context of reinforcement learning or preference optimization, the policy model represents the model being trained, and the reference model is the baseline model used to evaluate the performance of the policy model or the original model that has not undergone specific training. In machine learning and deep learning, the loss function is a measure used to measure the difference between the model's prediction result and the actual result. The training objective is to find a set of model parameters that minimize the value of the loss function, so as to obtain a prediction result that is as close as possible to the actual value. During model training, the smaller the loss function, the smaller the degree of deviation between the policy model and the reference model. This means that the prediction error of the model is minimized and the fitting degree of the model to the training data is improved. When the loss function reaches the minimum value, theoretically the model performs best on the training data, but this also needs to be balanced with the risk of overfitting (that is, the model is too dependent on the training data, resulting in a decline in generalization ability).
[0099] In an exemplary embodiment, the original loss function is expressed as follows:
[0100] ;
[0101] represents the input prompt, represents the negative sample, represents the positive sample, is the policy model, is the reference model, is the hyperparameter, is an impulse function, .
[0102] In an exemplary embodiment, the weighted loss function is expressed as follows:
[0103] ;
[0104] represents the input prompt, represents the negative sample, represents the positive sample, is the policy model, is the reference model, is the hyperparameter, is the impulse function.
[0105] To better understand the process of the above-mentioned large model-based speech data processing method, the following further describes the implementation method flow of the above-mentioned large model-based speech data processing in combination with optional embodiments, but is not used to limit the technical solutions of the embodiments of the present application.
[0106] In this embodiment, a large model-based speech data processing method is provided. Figure 3 is a schematic diagram of the large model-based speech data processing method according to the embodiments of the present application, as Figure 3 shown, the specific steps are as follows:
[0107] Step S301: Instructionally fine-tune a semantic parsing model (corresponding parsing model).
[0108] First, for the speech data collected by the sound pickup device, after converting it into text data, an annotated data set is obtained by respectively performing artificial construction and online data annotation on the text data, which specifically includes the user's query, the annotated domain, intent, and slots to be extracted. Then, an instruction fine-tuning data set is constructed, and the LLaMA2-base-7B model is used as the base model for fine-tuning, and then a semantic parsing model is obtained, which has the domain intent classification ability and entity recognition ability that conform to business rules.
[0109] An example of the annotated data is as follows:
[0110] {
[0111] "Now cook the roasted chicken wings for me a little longer": {
[0112] "domain": "recipe",
[0113] "intent": "increaseTime",
[0114] "slots": "Now adjust [(recipe) Roasted Chicken Wings] for a longer time for me"
[0115] };
[0116] "Make the baking time of the oven longer": {
[0117] "domain": "Oven",
[0118] "intent": "increaseTime",
[0119] "slots": "Make the baking time of [(device) oven] longer"
[0120] };
[0121] "Add another quarter of an hour to the washing time of the washing machine": {
[0122] "domain": "Washer",
[0123] "intent": "increaseTime",
[0124] "slots": "Add another quarter of an hour to the washing time of [(device) washing machine]"
[0125] };
[0126] }。
[0127] Among them, the instruction fine-tuning data of the first labeled data example can be expressed as:
[0128] {
[0129] "prompt": "You are a semantic parser in the scenario of intelligent home appliance voice control. You need to identify the domain, intent, and slots from the given query.\nRequirements: \n1. Do not perform any parsing, and only output a Json response with domain, intent, and slots as keys; 2. If there is no domain or intent, set their values to other.\nInput: \"Now adjust the roasted chicken wings for a longer time for me\" \nAnswer: ",
[0130] "target": "{\"domain\": \" recipe \",\"intent\": \"increaseTime\",\"slots\": \"Now adjust [(recipe) Roasted Chicken Wings] for a longer time for me\"}"
[0131] }。
[0132] Taking the LLaMA2-base-7B model as the base model, load the pre-trained weights and define the fine-tuning parameters, such as learning rate, number of fine-tuning epochs, batch size, etc. for fine-tuning settings. Then, load the instruction fine-tuning dataset (e.g., the above-mentioned labeled dataset) into the training process of the model, and train the model to learn to extract domain, intent, and slot information from the query. During the fine-tuning process, the model will adjust its parameters to minimize the loss according to the difference between the prompt and the target. After completing the fine-tuning, save the optimized model weights, which will serve as the basis for subsequent data augmentation and preference optimization. Next, prepare a test dataset that has not been used in training, and use the test dataset to evaluate the accuracy, recall, F1 score, and other metrics of the fine-tuned model on domain, intent, and slot recognition tasks. Through these steps, based on the LLaMA2-base-7B model, an instruction fine-tuning can be used to obtain a model with preliminary semantic parsing capabilities, preparing for further data augmentation and preference optimization.
[0133] Step S302: Agreement graph data augmentation and cross-intent data augmentation based on prompt learning.
[0134] The use of prompt learning techniques can augment the model's generation ability to increase the diversity and colloquial expression of text data. Agreement graph data augmentation refers to generating different expressions under the same intent, that is, generating variants of the seed corpus while keeping the original intent unchanged to cover different expressions under the same intent. Cross-intent data augmentation, on the other hand, generates new text according to similar or opposite intents, that is, performing inversion or symmetry operations on the intent of the seed corpus while keeping the domain unchanged to generate new corpus opposite to the original intent. For example, if the intent of the seed corpus is "increase time", then the goal of the opposite intent augmentation is to generate sentences like "decrease time". This step prompts the model to learn the positive and negative conversion of intents and the diverse expressions of intents within the same domain by setting different prompts.
[0135] Using the labeled dataset as the seed dataset, denoted as: , where is the seed corpus to be augmented, is the domain and intent of this seed corpus. At this stage, to maximize the utilization of the labeled information, the original seed corpus can be augmented towards similar and opposite intents respectively, including the following steps:
[0136] 1. First, for the set of intents under each domain, sort out the intent pairs suitable for corpus conversion. For example: open - close, increase temperature - decrease temperature, start - pause, etc. where there are symmetric intents. And denote it as intentPairs = . For example:
[0137] intentPairs =
[0138] {
[0139] "open": "close",
[0140] "close": "open",
[0141] "startup": "suspend",
[0142] "suspend": "startup",
[0143] "increaseTemp": "decreaseTemp",
[0144] "decreaseTemp": "increaseTemp"
[0145] }}。
[0146] For cases where there are no symmetric intents, such as intents like status query and parameter adjustment, only the corpus augmentation of the consent graph (i.e., similar intents) is considered.
[0147] 2. Construction of prompt words:
[0148] The prompt word Prompt constructed in this application includes three parts: 1) role description and task description, 2) providing examples, and 3) input of the seed corpus. Among them, the role description, task description, example provision, and input of the seed corpus respectively correspond to the task description information, role description information, task examples, and input seed corpus in the above task information.
[0149] Furthermore, task descriptions and examples can be provided specifically for different fields to enhance the diversity and quality after data augmentation. Taking the seed corpus to be augmented as an example, for the field and intent of this corpus, the target intent is , which are three pieces of corpus sampled from the subset with as the field intent.
[0150] For example, the role description is: "You are a corpus generator in the scenario of intelligent home appliance voice control, and you need to amplify 3 pieces of text from the given query and the belonging intent intent to the target intent intent_target."
[0151] The task description is: "You will focus on the field < >, which is mainly used for < >Related voice control. The amplified text is required to be as diverse and colloquial as possible.”
[0152] Task example: “\nFor example: query: < >, domain: < >, intent: < >,intent_target: < >, \n Output: ”。
[0153] Input of the seed corpus: “Based on the above information, \n Input: query: < >, domain: < >,intent: < >, intent_target: < >, \n Output:”。
[0154] The constructed Prompt: “You are a corpus generator in the context of intelligent home appliance voice control. You need to amplify 3 pieces of text for the target intent intent_target according to the given query and the belonging intent intent. You will focus on the domain recipe, which is mainly used for voice control related to recipes. The amplified text is required to be as diverse and colloquial as possible. \n For example: {\"query\": \"Now cook the roasted chicken wings for a longer time\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"},\n Output: "[\"Don't cook the roasted chicken wings for so long\", \"Don't spend too much time cooking the chicken wings\", \"Reduce the cooking time of the roasted chicken wings for me\"]). Based on the above information, \n Input: {\"query\": \"Now cook the roasted chicken wings for a longer time\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"}, \n Output:”。
[0155] Input the constructed prompt (corresponding to the preset extended prompt statement) into the original LLaMA2-base model to generate data for each seed corpus towards the same intent and symmetric intent. Among them, when the target intent When it is the same as the original intent, it means agreeing to the data augmentation of the graph. When the target intent is involved, it means cross-intent data augmentation. In addition, a relatively large temperature parameter can be set to improve the diversity of the generated data.
[0156] Step S303: Construct a preference optimization dataset.
[0157] Construct a preference optimization dataset based on the augmented data obtained in step S302. In this step, the augmented data is predicted using the fine-tuned model and classified into positive samples (consistent with business preferences) and negative samples (inconsistent with business preferences). Then, by calculating the similarity distribution, the samples that are closest to the negative samples but correctly classified are selected from the positive sample candidate set through online sampling as positive samples to construct preference sample pairs. Constructing a preference optimization dataset is to provide the positive and negative sample pairs required for model training in the subsequent preference optimization process to optimize the consistency between the text generated by the model and the actual business expectations. Specifically:
[0158] 1) Combine the model fine-tuned in step S301 to collect candidate positive and negative samples.
[0159] 2) Based on online sampling of the similarity distribution, construct appropriate preference sample pairs.
[0160] For the sake of easy expression, this application records the text expected to be generated by the model as positive samples, that is, chosen samples. The text that is not expected to be generated by the model is recorded as negative samples, that is, rejected samples. The expected positive samples are more in line with the business preferences and expectations than the negative samples.
[0161] For the seed dataset in one of the seed corpora , assuming the target intent is , then the domain intents of the data expected to be augmented are respectively , and the text set after augmentation in step S302 is denoted as .
[0162] Furthermore, step S303 further includes:
[0163] Step S3031: Collection of the positive sample candidate set (corresponding to the positive sample dataset) and the negative sample candidate set (corresponding to the negative sample dataset).
[0164] First, maintain a mapping table with the quadruple as the key. Under this key, two lists need to be maintained, which are used to store candidate data that can be used as positive and negative samples respectively.
[0165] For the amplified text , input it into the semantic parsing model fine-tuned in step 1 for classification processing, and record the predicted domain intention as .
[0166] Furthermore, whether the predicted domain intention is the same as the expected domain intention includes the following four cases:
[0167] 1. : The domain intention that meets the target, which is consistent with the actual business expectation.
[0168] 2. : Different domains but the same intention.
[0169] 3. : The same domain but different intentions.
[0170] 4. : Different domains and different intentions.
[0171] Based on the differences between the predicted results and the expected domain intention reflected in the above 4 cases, filter out the positive and negative sample candidate sets, and then construct a preference optimization data set.
[0172] Those belonging to case 1 : Can be used as positive samples and added to the positive sample candidate set .
[0173] Those belonging to case 2 : Can be used as negative samples and added to the negative sample candidate set . In addition, it can be input into the semantic parsing model in step S301 again to obtain the result of slot recognition, and replace the recognized slot with the slot value of the corresponding seed corpus, then the chosen data with the same sentence pattern can be constructed and added to the positive sample candidate set .
[0174] Those belonging to cases 3 and 4 : Can both be used as negative samples and added to the negative sample candidate set .
[0175] Step S3032: Online sampling based on similarity distribution.
[0176] Through S3031, a positive sample candidate set with the quadruple as the key and a negative sample candidate set are obtained. Next, it is necessary to further construct paired positive and negative samples as preference samples. Here, an online sampling method based on similarity distribution is adopted:
[0177] Suppose for each negative sample , it is necessary to select the most suitable k positive samples from in the ratio of 1:k, and use them to construct k positive and negative sample pairs. To improve the sample training efficiency of the model in the preference optimization stage, samples that are semantically similar to the negative samples but have different labels (i.e., domains or intents) should be selected as positive examples as much as possible, so as to provide the model with supervision information on the fine-grained differences between positive and negative examples. The specific steps are as follows:
[0178] 1) Vector encoding of positive and negative samples: Denote the language model to be optimized in the preference optimization stage as ; traverse each negative sample , and input the negative sample and each positive sample in the positive sample candidate set into the model for encoding, obtain the vector output by the last layer and perform average pooling, so as to use it as the vector representation of positive and negative samples, denoted as vectors and respectively, .
[0179] 2) Calculate similarity: The cosine similarity between the negative sample and each positive sample can be expressed as:
[0180] .
[0181] 3) Calculate the similarity probability distribution of with respect to the positive sample set: Here, softmax normalization is used, and a constant is added to each item to maintain the numerical stability. The similarity probability distribution function is specifically expressed by the following formula:
[0182] .
[0183] 4) Sampling: In each iteration of the preference optimization stage, for each negative sample , calculate the probability distribution of the similarity of the negative sample with respect to the positive sample set according to the above formula, and sample from the positive sample candidate set based on this distribution, and then obtain k positive samples , .
[0184] 5) Construct k preference samples: Each sample in the preference optimization dataset consists of three parts: prompt, chosen, and rejected. Specifically, through each quadruple and the pre-written instruction (corresponding to the preset lead), the following prompt can be constructed:
[0185] ""<instruction>.\nInput: {"query": <q i >, "domain": <D i >, "intent": <I i >, "intent_target": < >} \nAnswer: "."
[0186] Next, for each negative sample , according to the positive samples sampled in step 4) positive samples , then preference samples can be constructed, and each preference sample is as follows:
[0187] {
[0188] "prompt": " <instruction>。\nInput: {"query": <q i >, "domain": <D i >, "intent": <I i >, "intent_target": < >} \nAnswer: ",
[0189] "chosen": "< >",
[0190] "rejected": "< >"
[0191] }。
[0192] Based on the above sampling strategy, among the positive sample candidates, the positive samples that are more semantically similar to the negative samples have a greater probability of being selected. The preference samples constructed based on this provide the model with supervised information on the fine-grained differences between positive and negative examples during the preference optimization stage.
[0193] For the convenience of subsequent presentation, the prompt constructed by is denoted as , then each preference sample can be represented by the triple . The constructed preference sample set can be represented as:
[0194] .
[0195] Step S304: Preference optimization. Specifically, it includes the following steps.
[0196] Step S3041: Construct the original DPO loss function.
[0197] First, the classic Direct Preference Optimization (DPO) loss function can be expressed as the following formula.
[0198] .
[0199] Among them, x is the input prompt, q r is the rejected sample, q c is the chosen sample, π θ is the policy model, is the language model to be optimized, π ref is the reference model, β is a hyperparameter with a value in [0,1], and the parameters of this model do not participate in the update during training. An un-finetuned base model can be selected. is the sigmoid function, that is 。
[0200] In preference optimization, the loss function is designed to measure the quality of the output of the policy model relative to the reference model. For example, Direct Preference Optimization (DPO, loss function) encourages the policy model to generate better responses by comparing the different outputs of the policy model and the reference model for the same input. If the loss function is designed properly, one of the effects of minimizing the loss function during training is to reduce the deviation between the policy model and the reference model, that is, the decisions or outputs of the policy model are closer to or better than the reference model.
[0201] In the scenario of improving the DPO loss function, the hyperparameter β is used to control the degree of deviation between the policy model and the reference model. The smaller the value of β, the more the output of the policy model tends to the output of the reference model. By adjusting β, the relationship between model exploration (generating new data or improving output) and exploitation (maintaining output consistent with the reference model) can be balanced, avoiding the policy model deviating too much from the reference model, resulting in a decline in output quality or non-compliance with business preferences.
[0202] Step S3042: Analyze the problems existing in the preference optimization dataset.
[0203] Since this application includes consent diagrams and cross-intent data generation, in the preference optimization dataset, there may be a situation where "the samples that are positive and negative examples respectively in the consent diagram generation task are negative and positive examples in the cross-intent generation task", and the difference between these two preference samples is only the different target intents in the prompt. If the model is not sensitive to the information of the target intent intent_target in the prompt, then the sample pairs in the above situation will mislead the model and it is difficult to align with the classification preferences of the actual business.
[0204] An example of the prompt is shown below.
[0205] {
[0206] "prompt": "You are a corpus generator in the scenario of intelligent home appliance voice control. You need to amplify 1 piece of text from the given query and the belonging intent intent_target to the target intent intent_target. You will focus on the domain recipe, which is mainly used for voice control related to recipes. The amplified text should be as diverse and colloquial as possible.\nInput: {\"query\": \"Now cook the roasted chicken wings for a longer time for me\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"} \nAnswer: ",
[0207] "chosen": "I want to cook the chicken wings for a shorter time",
[0208] "rejected": "I want to cook the chicken wings for a longer time"
[0209] }。
[0210] {
[0211] "prompt": "You are a corpus generator in the scenario of intelligent home appliance voice control. You need to amplify 1 piece of text from the given query and the belonging intent intent_target to the target intent intent_target. You will focus on the domain recipe, which is mainly used for voice control related to recipes. The amplified text should be as diverse and colloquial as possible.\nInput: {\"query\": \"Now cook the roasted chicken wings for a longer time for me\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"increaseTime\"} \nAnswer: ";
[0212] "chosen": "I want to cook the chicken wings for a longer time",
[0213] "rejected": "I want to cook the chicken wings for a shorter time"
[0214] }。
[0215] Step S3043: Construct the weighted loss of DPO for preference optimization.
[0216] To solve the above problems, in this section, first define the statistic , which is used to measure the sensitivity or instruction compliance degree of the model to two given prompts when they output the same response. Then, sample weights are constructed for the preference optimization dataset, and the original preference optimization loss is improved to finally obtain the weighted loss function. Specifically:
[0217] (1) Define the statistic .
[0218] In the language model system, perplexity can be used to measure the uncertainty of the model in generating a given sequence or to evaluate the fitting degree of the model to a given sequence. Specifically, given a sequence , assuming that the conditional probability of the language model generating the i-th word is , then the perplexity of generating sequence S can be calculated as:
[0219] .
[0220] Based on perplexity, the following statistic PPL diff can be defined:
[0221] .
[0222] Among them, and respectively represent the two input prompts, represents the response sequence generated by the model . and respectively represent the perplexity of the model generating the response sequence and when using as input. Here, the absolute value of the difference between the two is defined as PPL diff .
[0223] By defining the statistic PPL diff , it is used to evaluate the response sensitivity of the model to different prompts, that is, the difference in the model's compliance with different instructions. When PPL diff is close to 0, it indicates that there is no significant difference in the perplexity when generating with these two prompts as input respectively, that is, the model is not sensitive to the difference between these two prompts when generating the same response. On the contrary, when PPL diff is relatively large, it indicates that the model can well comply with these two different prompt instructions, that is, the model is more sensitive to the difference between these two prompts when generating the same response. Therefore, PPL diff It can measure the compliance or sensitivity of the model to two different prompts when outputting the same response. Through PPL diff Assign weights to the samples in the preference optimization dataset, with lower weights for low-sensitivity samples and higher weights for high-sensitivity samples.
[0224] (2)Construct sample weights for the preference optimization dataset to obtain the weighted DPO loss.
[0225] For the preference optimization dataset , each preference sample consists of a prompt , a chosen sample , and a rejected sample . Among them, the prompt is determined by the quadruple . Assume that in the intent pairs in this field, the symmetric intent of is . Then when the target intent is, it is data augmentation for the consent graph; when the target intent is, it is data augmentation across intents; for the sake of easy expression, the prompts for the consent graph and across intents are respectively , and can be specifically expressed as:
[0226] .
[0227] .
[0228] Denote the preference data subset with or as the prompt as . Further, define the statistic , used to represent the sum of the PPL of the positive and negative sample pairs with respect to and diff respectively, and can be specifically expressed as:
[0229] .
[0230] Among them, when is small, it indicates that both the perplexity of generating positive samples and the perplexity of generating negative samples are small, that is, the model is not sensitive to the difference between and ; while when is large, it indicates that at least one of the perplexities of generating positive samples and generating negative samples is large, proving that the model can at least capture the difference between these two prompts.
[0231] Therefore, it is hoped that the weight of the larger preference sample can be increased, and the weight of the smaller preference sample can be decreased. In this way, the interference caused by some samples to the preference optimization of the model can be alleviated as much as possible.
[0232] To represent it as a weight, the subset of preference data can be normalized, specifically expressed as:
[0233] .
[0234] Among them, for prompt and , its is the same. The normalized value is used as the sample weight in the preference optimization loss function, and then the weighted DPO loss function can be obtained:
[0235] .
[0236] The goal of training is to minimize the above loss function. In actual training, optimizers such as AdamW can be selected to update the parameter θ. For the hyperparameter β, a value between 0.1 - 0.5 can be selected to control the deviation degree θ between the policy model π ref and the reference model π.
[0237] In this step, by using the weighted DPO loss function, the model adjustment can be more effectively guided, the training deviation caused by low-quality samples can be reduced, and the data generated by the model can be ensured to be more in line with the business preferences.
[0238] Based on the above steps, this application proposes a data generation method based on large model preference optimization. This method constructs appropriate preference samples based on the sampling strategy of similarity distribution, and optimizes the model preference based on the improved weighted DPO, which improves the diversity and quality of the amplified data while ensuring consistency with the semantic results expected by the business. Specifically, this application first fine-tunes a semantic parsing model for subsequent domain and intent recognition. Then, prompts are constructed to generate the initial amplified data to ensure the quantity and diversity of the data. Further, a sampling strategy based on similarity distribution is used to construct preference samples, and the weights are assigned to different preference samples based on the difference in perplexity. Finally, the weighted DPO loss function is obtained to optimize the model preference. By using the optimized model of this application, data more aligned with the actual business preferences can be generated.
[0239] It can be understood that the data augmentation methods for augmenting the labeled data also include traditional methods such as synonym replacement, NER entity replacement, back translation, random noise injection, syntactic tree enhancement, etc., and methods of constructing prompts by splicing labels with the original text based on language models such as GPT-2 and performing fine-tuning to directly learn to generate data with such labels.
[0240] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present application.
[0241] Figure 4 is a structural block diagram of a speech data processing device based on a large model according to an embodiment of the present application; as Figure 4 shown, it includes:
[0242] A first augmentation module 42, configured to perform data augmentation operations on the labeled data of the labeled data set to obtain an augmented training data set, where the labeled data is collected by a sound pickup device;
[0243] An output module 44, configured to use a parsing model to classify the augmented training data set into a positive sample data set and a negative sample data set, and generate a preference optimization data set by using the positive sample data set and the negative sample data set. A preference optimization sample in the preference optimization data set includes a prompt word, negative sample data, and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on an initial large language model;
[0244] A second augmentation module 46, configured to generate a weighted loss function according to the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data augmentation model that completes a preset training target by using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented data set.
[0245] Using the above device, perform data augmentation operations on the labeled data of the labeled dataset to obtain an augmented training dataset; use an analysis model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generate a preference optimization dataset using the positive sample dataset and the negative sample dataset. One preference optimization sample in the preference optimization dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; generate a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the analysis model, obtain a data augmentation model that completes a preset training objective using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset; adopting the above technical solution, by performing data augmentation on the labeled data, data diversity can be improved. Use the analysis model to classify the augmented training dataset to obtain a positive sample dataset and a negative sample dataset, and then construct a preference optimization dataset based on the relevance between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weight of the preference optimization sample in the preference optimization dataset and the original loss function of the analysis model, the quality of different preference optimization samples can be reasonably quantified in a weighted manner, and a data augmentation model that completes a preset training objective using the weighted loss function is obtained, and the data augmentation model is used to complete the data augmentation of newly input text data. By training the model with a weighted loss function that combines sample weights, the training deviation caused by low-quality samples can be reduced, the model performance can be further optimized, so that when the model processes newly input text data, it can more accurately understand and generate the corpus of the required intention, solve the technical problem of low data quality obtained by existing data augmentation methods, and thus improve the quality of the augmented data and also improve the accuracy of the model for text processing.
[0246] In an exemplary embodiment, before using the analysis model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, the output module is further configured to: generate current labeled data based on the current labeling operation of the target object on the initial data, and generate the labeled dataset according to the historical labeled data provided by the target object and the current labeled data, where the labeled dataset at least includes the question sentence of the target object, the field labeled by the target object, the intention labeled by the target object, and the slot corresponding to the entity labeled by the target object; use the fine-tuning instruction of the target object to instruct the initial large language model to perform a recognition operation on the labeled data, and output a recognition result according to the output requirements of the instruction fine-tuning data in the fine-tuning instruction, where the recognition result at least includes: labeled field, labeled intention, labeled entity, and labeled slot.
[0247] In an exemplary embodiment, the first amplification module is further configured to: select seed corpus data to be amplified from the labeled data, and determine the seed annotation field, the seed annotation intention, and the target intention corresponding to the seed corpus according to the seed corpus data; use a preset amplification prompt statement to instruct the initial large language model to perform a data amplification task from the seed annotation intention to the target intention on the seed corpus within the seed annotation field, and obtain an amplified training data set. The preset amplification prompt statement at least includes the task information required for the initial large language model to perform the data amplification task, and the task information at least includes task description information, role description information, task examples, and the input seed corpus. The role description information is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the seed annotation intention, and the new corpus has the target intention.
[0248] In an exemplary embodiment, the first amplification module is further configured to: for a set of annotation intentions belonging to the same seed annotation field, screen out the first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions in the set of annotation intentions, and generate an intention pair corresponding to the first seed annotation intention. The intention pair includes the first seed annotation intention, and a second seed annotation intention that is similar or opposite to the first seed annotation intention. The intention pair is expressed as , is the first seed annotation intention, is the second seed annotation intention, and n is a positive integer; determine a first extended prompt statement from the preset extended prompt statement. The first role description information in the first extended prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the first seed annotation intention; use the first extended prompt statement to instruct the initial large language model to perform a data amplification operation with a similar intention for the seed corpus to obtain an amplified training data set.
[0249] In an exemplary embodiment, the first amplification module is further configured to: screen out other annotation intentions except the first seed annotation intention from the set of annotation intentions, and the other annotation intentions only have similar intentions; determine a second extended prompt statement from the preset extended prompt statement. The second role description information in the second extended prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to the other annotation intentions; use the second extended prompt statement to instruct the initial large language model to perform a for the seed corpus Perform data augmentation operations on similar intents to obtain an augmented training dataset.
[0250] In an exemplary embodiment, the first augmentation module is further configured to: for a set of labeled intents belonging to the same seed annotation domain, screen out the first seed labeled intents with similar or opposite intents from the seed labeled intents in the set of labeled intents, generate intent pairs corresponding to the first seed labeled intents, the intent pairs including the first seed labeled intents, and the second seed labeled intents that are similar or opposite to the first seed labeled intents, and the intent pairs are represented as , is the first seed labeled intent, is the second seed labeled intent, and n is a positive integer; determine a third extended prompt statement from the preset extended prompt statements, where the third role description information in the third extended prompt statement is used to set the function of the initial large language model to generate new corpus based on the seed corpus and the intent pairs, and the target intent of the new corpus is similar to the second seed labeled intent; use the third extended prompt statement to instruct the initial large language model to perform data augmentation operations on the seed corpus for opposite intents to obtain an augmented training dataset.
[0251] In an exemplary embodiment, the labeled dataset is represented as , is the seed corpus to be augmented, is the seed annotation domain, is the seed labeled intent, and the augmented training dataset is represented as , is the target intent, , and i, M, and N are positive integers.
[0252] In an exemplary embodiment, the output module is further configured to: input the data of the augmented training dataset into the parsing model, so that the parsing model performs classification processing on the augmented training data and outputs the predicted domain and the predicted intent ; determine the positive sample dataset and the negative sample dataset according to the and verification results, and the verification results include the first verification result of the seed annotation domain and the predicted domain and the second verification result of the seed labeled intent and the predicted intent .
[0253] In an exemplary embodiment, the output module is further configured to: upon determining that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention Figure 1 If the seed corpus corresponding to the seed annotation field is consistent with the first mapping relationship, the seed corpus corresponding to the seed annotation field is stored in the positive sample data set The first mapping relationship represents the corpus quadruple of the seed corpus to be amplified in the annotated data. The relationship where is the key and the positive sample dataset is the value. is the seed corpus to be amplified, annotating a field for said seed, annotate the seed with intent, For the purpose.
[0254] In an exemplary embodiment, the output module is further configured to: after determining that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention, Figure 1 If any of the above is not true, the seed corpus corresponding to the seed annotation field is stored in the negative sample data set according to the second mapping relationship. , the second mapping relationship is expressed as The relationship with the negative sample dataset as the key and the negative sample dataset as the value.
[0255] In an exemplary embodiment, the output module is also used to: input the augmented training data set into the parsing model so that the parsing model recognizes the augmented training data and obtains the slot relationship corresponding to the augmented training data; replace the slot relationship with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data to the same sentence structure as the seed corpus to be augmented, determine the updated augmented training data as positive sample candidate data, and determine the positive sample data set based on the positive sample candidate data.
[0256] In an exemplary embodiment, the output module is further configured to: for negative samples , collect k positive samples from the positive sample data set according to the preset sample ratio of 1:k , , ; Using the prompt word, the negative sample and the k positive samples Generate a preference optimization sample, and determine the preference optimization data set according to the sample set of all preference optimization samples, wherein the prompt word is composed of a corpus quadruple And preset introduction generation.
[0257] In an exemplary embodiment, the output module is further configured to: according to the vector encoding result of the negative sample and the vector encoding results of k positive samples calculate the negative sample and the k positive samples calculate the cosine similarity between them, and calculate the similarity probability distribution result of the negative sample with respect to the positive sample dataset based on the cosine similarity; sample the positive sample dataset using the similarity probability distribution result to obtain the k positive samples corresponding to the negative sample .
[0258] In an exemplary embodiment, the output module is further configured to calculate the cosine similarity between the negative sample and the k positive samples through the following formula:
[0259] ;
[0260] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
[0261] In an exemplary embodiment, the output module is further configured to: calculate the similarity probability distribution result of the negative sample with respect to the positive sample dataset using the similarity probability distribution function of the negative sample with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0262] ;
[0263] constant ; the similarity probability distribution result is expressed as .
[0264] In an exemplary embodiment, the output module is further configured to: represent a preference optimization sample in the form of a triple, then the preference optimization dataset is expressed as follows:
[0265] , represents the prompt.
[0266] In an exemplary embodiment, before generating the weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the parsing model, the second amplification module is further configured to: determine the sensitivity quantization index of the parsing model according to the perplexity parameter; use the sensitivity quantization index to normalize the data subset in the preference optimization dataset, and determine the result of the normalization as the sample weight, and the preference optimization samples in the data subset include prompt words with the same intention or prompt words with the opposite intention.
[0267] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , where S represents the set input sequence of the parsing model, x represents the prompt word input for prompting the generation of similar intents, represents the prompt word input for prompting the generation of opposite intents, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity when the parsing model generates the response sequence y with x as the set input sequence, represents when using as the set input sequence, the perplexity when the parsing model generates the response sequence y.
[0268] In an exemplary embodiment, the second amplification module is further configured to: obtain a target quantization index generated according to the sensitivity quantization index , the target quantization index represents , respectively regarding the prompt words of the same intent or the prompt words of the opposite intent the sum of the index values of the sensitivity quantization indexes; and determine the result of normalizing regarding the data subset as the sample weight, and the result of the normalization is expressed as follows: .
[0269] In an exemplary embodiment, the second amplification module is further configured to: perform iterative training on the policy model with the minimization of the weighted loss function as the training objective, and stop training the policy model and determine the policy model as the data amplification model when the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is the smallest, where the loss value of the weighted loss function represents the deviation degree between the policy model and the reference model.
[0270] An embodiment of the present application further provides a storage medium, which includes a stored program, and the above program executes the method of any one of the above when running.
[0271] Optionally, in this embodiment, the above storage medium may be configured to store program codes for performing the following steps:
[0272] S1, perform data amplification operations on the labeled data of the labeled data set to obtain an amplified training data set, where the labeled data is collected by a sound pickup device;
[0273] S2. Use the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generate a preference optimization dataset by using the positive sample dataset and the negative sample dataset. One preference optimization sample in the preference optimization dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on an initial large language model.
[0274] S3. Generate a weighted loss function according to the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data augmentation model that achieves the preset training objective by using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset.
[0275] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0276] Optionally, the above electronic device may further include a transmission device and an input / output device. The transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0277] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0278] S1. Perform a data augmentation operation on the labeled data of the labeled dataset to obtain an augmented training dataset, where the labeled data is collected by a sound pickup device.
[0279] S2. Use the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generate a preference optimization dataset by using the positive sample dataset and the negative sample dataset. One preference optimization sample in the preference optimization dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on an initial large language model.
[0280] S3. Generate a weighted loss function according to the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data augmentation model that achieves the preset training objective by using the weighted loss function, and use the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset.
[0281] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs.
[0282] Optionally, an embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments.
[0283] Optionally, an embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments.
[0284] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0285] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the present application is not limited to any specific combination of hardware and software.
[0286] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.< / instruction> < / instruction>
Claims
1. A method for processing speech data based on a large model, characterized in that, Including: Performing data augmentation operations on the labeled data of the labeled dataset to obtain an augmented training dataset, where the labeled data is collected by a sound pickup device; Using an analysis model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generating a preference optimization dataset by using the positive sample dataset and the negative sample dataset. One preference optimization sample in the preference optimization dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; Generating a weighted loss function according to the sample weights of the preference optimization samples and the original loss function of the analysis model, obtaining a data augmentation model that completes a preset training target by using the weighted loss function, and using the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset.
2. The method for processing voice data based on a large model according to claim 1, wherein, Before using the analysis model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, the method further includes: Generating current labeled data based on the current labeling operation of the initial data by a target object, and generating the labeled dataset according to the historical labeled data provided by the target object and the current labeled data. The labeled dataset at least includes the question sentence of the target object, the field labeled by the target object, the intention labeled by the target object, and the slot corresponding to the entity labeled by the target object; Using the fine-tuning instruction of the target object to instruct the initial large language model to perform an identification operation on the labeled data, and outputting an identification result according to the output requirement of the instruction fine-tuning data in the fine-tuning instruction. The identification result at least includes: labeled field, labeled intention, labeled entity, and labeled slot; 3. The method for processing speech data based on a large model according to claim 1, wherein Performing data augmentation operations on the labeled data of the labeled dataset to obtain an augmented training dataset, including: Selecting seed corpus data to be augmented from the labeled data, and determining the seed annotation field, seed annotation intention, and target intention corresponding to the seed corpus according to the seed corpus data; Using a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task of the seed corpus from the seed annotation intention to the target intention in the seed annotation field to obtain an augmented training dataset. The preset augmentation prompt statement at least includes the task information required for the initial large language model to perform the data augmentation task. The task information at least includes task description information, role description information, task examples, the input seed corpus, and the role description information is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the seed annotation intention, and the new corpus has the target intention.
4. The method for processing voice data based on a large model according to claim 3, wherein Using a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task of the seed corpus from the seed annotation intention to the target intention in the seed annotation field to obtain an augmented training dataset, including: For a set of annotation intentions belonging to the same seed annotation field, screen the first seed annotation intentions with similar or opposite intentions from the seed annotation intentions in the set of annotation intentions, generate intention pairs corresponding to the first seed annotation intentions, where the intention pairs include the first seed annotation intentions, and second seed annotation intentions that are similar or opposite to the first seed annotation intentions, and the intention pairs are represented as , is the first seed annotation intention, is the second seed annotation intention, and n is a positive integer; Determine a first extended prompt statement from the preset extended prompt statements. The first role description information in the first extended prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention. The target intention of the new corpus is similar to the first seed annotation intention. Use the first extended prompt statement to instruct the initial large language model to for the seed corpus of perform a data augmentation operation for similar intents to obtain an augmented training data set.
5. The method for processing speech data based on a large model according to claim 4, wherein Use the preset amplification prompt statement to instruct the initial large language model to perform a data amplification task from the seed annotation intention to the target intention on the seed corpus within the seed annotation field, and obtain an amplified training data set, including: Screen out other annotation intentions except the first seed annotation intention from the set of annotation intentions. The other annotation intentions only have similar intentions. Determine a second extended prompt statement from the preset extended prompt statements. The second role description information in the second extended prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the other annotation intentions. The target intention of the new corpus is similar to the other annotation intentions. Use the second extended prompt statement to instruct the initial large language model to target of the seed corpus Perform a data augmentation operation for similar intents to obtain an augmented training data set.
6. The method for processing speech data based on a large model according to claim 3, wherein, Use the preset amplification prompt statement to instruct the initial large language model to perform a data amplification task from the seed annotation intention to the target intention on the seed corpus within the seed annotation field, and obtain an amplified training data set, including: For a set of annotation intentions belonging to the same seed annotation field, screen the first seed annotation intentions with similar or opposite intentions from the seed annotation intentions of the set of annotation intentions, generate an intention pair corresponding to the first seed annotation intention, the intention pair includes the first seed annotation intention, and a second seed annotation intention similar or opposite to the first seed annotation intention, and the intention pair is expressed as , is the first seed annotation intention, is the second seed annotation intention, and n is a positive integer; Determine a third extended prompt statement from the preset extended prompt statements. The third role description information in the third extended prompt statement is used to set the function of the initial large language model to generate new corpus according to the seed corpus and the intention. The target intention of the new corpus is similar to the second seed annotation intention. Use the third extended prompt statement to instruct the initial large language model to seed corpus perform data augmentation operations with the opposite intention to obtain an augmented training data set.
7. The method for processing speech data based on a large model according to claim 3, wherein The labeled dataset is denoted as , is the seed corpus to be amplified, is the seed annotation domain, is the seed annotation intention, and the amplified training dataset is denoted as , is the target intention, , where i and N are positive integers.
8. The method for processing speech data based on a large model according to claim 1, wherein Use an analysis model to classify the amplified training data set into a positive sample data set and a negative sample data set, including: Input the data of the augmented training dataset into the parsing model, so that the parsing model classifies the augmented training data and outputs the predicted field and predicted intent ; Determine the positive sample dataset and the negative sample dataset according to and The verification result, the verification result includes the seed annotation field and the predicted field The first verification result and the seed annotation intention and the predicted intention The second verification result.
9. The method for processing speech data based on a large model according to claim 8, wherein, Determining the positive sample data set according to the verification results of and includes: When it is determined that the first verification result indicates that the seed annotation field is consistent with the prediction field, and the second verification result indicates that the seed annotation intention is consistent with the prediction intention, store the seed corpus corresponding to the seed annotation field in the positive sample data set according to the first mapping relationship , the first mapping relationship represents a relationship with the quadruple of the corpus of the seed corpus to be amplified in the annotation data as the key and the positive sample data set as the value, is the seed corpus to be amplified, is the seed annotation field, is the seed annotation intention, is the target intention.
10. The method for processing voice data based on a large model according to claim 8, wherein, Determining the negative sample data set according to the verification results of and includes: In the case where it is determined that either the first verification result indicating that the seed annotation domain is consistent with the prediction domain or the second verification result indicating that the seed annotation intention is consistent with the prediction intention does not hold, the seed corpus corresponding to the seed annotation domain is stored in the negative sample data set according to the second mapping relationship. , the second mapping relationship represents a relationship with as the key and the negative sample data set as the value.
11. The method for processing voice data based on a large model according to claim 8, wherein, The method further includes: Input the amplified training data set into the analysis model so that the analysis model identifies the amplified training data and obtains the slot relationship corresponding to the amplified training data. Replace the slot relationship with the seed slot relationship corresponding to the seed corpus to be amplified to update the sentence pattern of the amplified training data to be the same as that of the seed corpus to be amplified. Determine the updated amplified training data as positive sample candidate data, and determine the positive sample data set according to the positive sample candidate data.
12. The method for processing speech data based on a large model according to claim 1, wherein Generate a preference optimization data set using the positive sample data set and the negative sample data set, including: For negative samples , according to the preset sample ratio, collect k positive samples from the positive sample dataset , , ; Using the said prompt, the negative sample and the k positive samples generate a preference optimization sample, and determine the preference optimization data set according to the sample set of all preference optimization samples. The prompt is generated by the corpus quadruple and a preset lead-in.
13. The method for processing speech data based on a large model according to claim 12, wherein, According to the preset sample ratio, k positive samples are collected from the positive sample dataset , including: According to the vector encoding results of the negative samples and the vector encoding results of k positive samples calculate the negative samples and the k positive samples to obtain the cosine similarity therebetween, and calculate the similarity probability distribution result of the negative samples with respect to the positive sample dataset based on the cosine similarity; Sample the positive sample dataset using the similarity probability distribution result to obtain negative samples The corresponding k positive samples .
14. The method for processing speech data based on a large model according to claim 13, wherein Calculate the negative samples using the following formula and k positive samples to calculate the cosine similarity between them: ; The vector encoding result for negative samples, is the vector encoding result for positive samples.
15. The method for processing voice data based on a large model according to claim 13, wherein, Calculate the similarity probability distribution result of the negative sample with respect to the positive sample data set based on cosine similarity, including: Use the similarity probability distribution function of the negative sample with respect to the positive sample data set to calculate the similarity probability distribution result of the negative sample with respect to the positive sample data set. The similarity probability distribution function is expressed as follows: ; Constant ; The similarity probability distribution result is expressed as .
16. The method for processing voice data based on a large model according to claim 12, wherein, The method further includes: Represent a preference optimization sample in the form of a triple, then the preference optimization data set is expressed as follows: , indicates a prompt word.
17. The method for processing voice data based on a large model according to claim 1, wherein Before generating a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the analysis model, the method further includes: Determine the sensitivity quantization index of the analysis model according to the perplexity parameter. Normalize the data subset in the preference optimization dataset using the sensitivity quantization index, and determine the result of the normalization as the sample weight. The preference optimization samples in the data subset include prompt words with the same intention or prompt words with the opposite intention.
18. The method for processing voice data based on a large model according to claim 17, wherein When the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as , , where S represents the set input sequence of the parsing model, x represents the prompt word input for prompting the generation of a similar intention, represents the prompt word input for prompting the generation of an opposite intention, and y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the parsing model generating the response sequence y when x is the set input sequence, represents when taking as the set input sequence, the perplexity of the parsing model generating the response sequence y.
19. The method for processing speech data based on a large model according to claim 17, wherein, Normalize the data subset in the preference optimization dataset using the sensitivity quantization index, and determine the result of the normalization as the sample weight, including: Obtain a target quantization index generated according to the sensitivity quantization index , the target quantization index represents , the sum of the quantization indices of the prompting texts for the same intention or the prompting texts for the opposite intention; The result of normalizing with respect to the data subset is determined as the sample weight, and the result of the normalization is expressed as follows: 。 20. The method for processing speech data based on a large model according to claim 1, wherein Obtain a data augmentation model that uses the weighted loss function to complete the training objective, including: Iteratively train the policy model with the minimization of the weighted loss function as the training objective. When the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is the smallest, stop training the policy model and determine the policy model as the data augmentation model. The loss value of the weighted loss function represents the deviation degree between the policy model and the reference model.
21. A speech data processing device based on a large model, characterized in that, Including: A first augmentation module for performing data augmentation operations on the labeled data in the labeled dataset to obtain an augmented training dataset; An output module for using an analysis model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generating a preference optimization dataset using the positive sample dataset and the negative sample dataset. A preference optimization sample in the preference optimization dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The analysis model is a model obtained by performing recognition training on an initial large language model; A second augmentation module for generating a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the analysis model, obtaining a data augmentation model that uses the weighted loss function to complete the preset training objective, and using the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented dataset.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program that, when running, executes the method described in any one of claims 1 to 20.
23. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 20 through the computer program.
24. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1 to 20.
Citation Information
Patent Citations
Voice data amplification method and system
CN108922518A
Positive and negative sample data balancing method in factory PCB defect detection
CN111126433A
Text classification method based on few samples
CN112765359A
Small sample intention recognition method based on data enhancement
CN115964486A
Voice augmentation method, related method, device, equipment and storage medium
CN118136034A
Cited By
Training data synthesis method of voice assistant model, voice assistant system and computer equipment
CN120783729A
Large model training method and device based on sample difficulty dynamic perception
CN121637058A