Speech data processing method, device, storage medium and electronic device based on large model
By performing data augmentation and classification on labeled datasets, generating preference-optimized datasets, and optimizing model training using weighted loss functions, the problem of low data quality in existing data augmentation methods is solved, and the text processing accuracy of the model is improved.
Patent Information
- Application Number
- CN202510768691.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The data quality generated by existing data augmentation methods is relatively low. Especially in semantically complex tasks, the generated sample data is inconsistent with the semantic intent of the original data, resulting in reduced model training accuracy.
By performing data augmentation on the labeled dataset, the analytical model is used to classify the augmented training dataset into positive samples and negative samples to generate a preference-optimized dataset. A weighted loss function is generated based on the sample weights and the original loss function of the analytical model to construct a data augmentation model and optimize the training process.
It improves data quality and model accuracy, reduces training deviation caused by low-quality samples, and improves the model's understanding and generation capabilities when processing new input text data.
Smart Images

Figure CN120279893B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech data processing, and in particular to a speech data processing method, device, storage medium and electronic device based on a large model. Background Art
[0002] Currently, the processing capabilities of large models for speech data rely on high-quality training data. However, real-world speech data collection is often limited by time and cost, resulting in insufficient data for training large models. To address this challenge, data processing methods such as data augmentation have emerged. Data augmentation involves generating additional training samples from the original dataset to enrich the training data, thereby improving the model's generalization and accuracy. Existing data augmentation methods fall into two main categories. The first category focuses on surface modification of the original training data, such as increasing the diversity of the training set through random noise addition and back-translation. However, these methods rely on predefined rules or limited synonym libraries, making it difficult to transcend the existing data framework and thus lacking data diversity. The second category, such as using generative language models like generative adversarial networks, generates new sample data with a distribution similar to the original data. However, the quality of the data generated by these methods is inconsistent. For example, in highly semantically complex tasks, they can generate grammatically correct samples that are inconsistent with the semantic intent of the original data, reducing the accuracy of model training using these samples.
[0003] Therefore, in the related art, there exists a technical problem that the data quality obtained by the existing data augmentation method is relatively low.
[0004] In the related technologies, there is no effective solution to the technical problem that the data quality obtained by the existing data augmentation methods is relatively low. Summary of the Invention
[0005] The embodiments of the present application provide a large-model-based voice data processing method, device, storage medium, and electronic device to at least solve the technical problem of low data quality obtained by existing data amplification methods in related technologies.
[0006] According to one embodiment of the embodiments of the present application, a speech data processing method based on a large model is provided, comprising: performing a data augmentation operation on the labeled data of a labeled data set to obtain an augmented training data set, wherein the labeled data is collected by a sound pickup device; using a parsing model to classify the augmented training data set into a positive sample data set and a negative sample data set, and using the positive sample data set and the negative sample data set to generate a preference optimization data set, wherein a preference optimization sample in the preference optimization data set includes a prompt word, negative sample data and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model obtained by performing recognition training on an initial large language model; generating a weighted loss function according to the sample weight of the preference optimization sample and the original loss function of the parsing model, obtaining a data augmentation model that uses the weighted loss function to complete a preset training target, and using the data augmentation model to perform data augmentation on newly input text data to obtain a target augmented data set.
[0007] In an exemplary embodiment, before using the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, the method further includes: generating current labeled data based on the current labeling operation of the target object on the initial data, and generating the labeled dataset based on the historical labeled data provided by the target object and the current labeled data, wherein the labeled dataset at least includes the question sentence of the target object, the labeled domain of the target object, the labeled intention of the target object, and the slots corresponding to the labeled entities of the target object; using the fine-tuning instruction of the target object to instruct the initial large language model to perform a recognition operation on the labeled data, and outputting a recognition result according to the output requirements of the fine-tuning data in the fine-tuning instruction, wherein the recognition result at least includes: the labeled domain, the labeled intention, the labeled entity, and the labeled slot.
[0008] In an exemplary embodiment, a data augmentation operation is performed on the labeled data of a labeled data set to obtain an augmented training data set, including: selecting seed corpus data to be augmented from the labeled data, and determining the seed annotation field, seed annotation intent and target intent corresponding to the seed corpus based on the seed corpus data; using a preset augmentation prompt sentence to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus within the seed annotation field to obtain an augmented training data set, the preset augmentation prompt sentence at least includes task information required for the initial large language model to perform the data augmentation task, the task information at least includes task description information, role description information, task examples, and input seed corpus, the role description information is used to set the function of the initial large language model to generate a new corpus based on the seed corpus and the seed annotation intent, and the new corpus has the target intent.
[0009] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention within the seed annotation domain, so as to obtain an augmented training data set, including: for a set of annotation intentions belonging to the same seed annotation domain, screening a first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions of the annotation intention set, generating an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar to or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a first extended prompt sentence is determined from the preset extended prompt sentence, the first role description information in the first extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the first seed marking intention; the first extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
[0010] In an exemplary embodiment, a preset augmentation prompt sentence is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention in the seed annotation domain, so as to obtain an augmented training data set, including: screening other annotation intentions except the first seed annotation intention from the annotation intention set, and the other annotation intentions only have similar intentions; determining a second extended prompt sentence from the preset extended prompt sentence, and the second role description information in the second extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to the other annotation intentions; using the second extended prompt sentence to instruct the initial large language model to target Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
[0011] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention within the seed annotation domain, so as to obtain an augmented training data set, including: for a set of annotation intentions belonging to the same seed annotation domain, screening a first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions of the annotation intention set, generating an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar to or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a third extended prompt sentence is determined from the preset extended prompt sentence, and the third role description information in the third extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the second seed marking intention; the third extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation operations with the opposite intention to obtain an augmented training dataset.
[0012] In an exemplary embodiment, the labeled dataset is represented as , is the seed corpus to be amplified, Label the field for the seed, The intention of the seed is labeled and the training dataset is expanded as , For the stated purpose, , i, M and N are positive integers.
[0013] In an exemplary embodiment, the augmented training data set is classified into a positive sample data set and a negative sample data set using a parsing model, including: inputting data of the augmented training data set into the parsing model so that the parsing model classifies the augmented training data and outputs a prediction domain corresponding to the augmented training data. and predictive intent ;according to and The positive sample data set and the negative sample data set are determined by the verification result, and the verification result includes the seed annotation field and the predicted areas First verification results and seed annotation intentions and the predicted intent The second verification result.
[0014] In one exemplary embodiment, according to and The positive sample data set is determined based on the verification result of the first verification result, including: determining that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention. Figure 1 If the seed corpus is consistent, the seed corpus corresponding to the seed annotation field is stored in the positive sample data set according to the first mapping relationship. The first mapping relationship represents the corpus quadruple of the seed corpus to be amplified in the annotated data. The relationship between the key and the positive sample data set as the value, is the seed corpus to be amplified, Label the field for the seed, labeling the seed with intent, For the purpose.
[0015] In one exemplary embodiment, according to and The negative sample data set is determined based on the verification result of the first verification result, including: determining that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention. Figure 1 If any of the above is not true, the seed corpus corresponding to the seed annotation field is stored in the negative sample dataset according to the second mapping relationship. , the second mapping relationship is represented by A relationship where the key is the negative sample dataset and the value is the negative sample dataset.
[0016] In an exemplary embodiment, the method further includes: inputting the augmented training data set into the parsing model so that the parsing model recognizes the augmented training data and obtains the slot relationship corresponding to the augmented training data; replacing the slot relationship with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data into the same sentence structure as the seed corpus to be augmented, determining the updated augmented training data as positive sample candidate data, and determining the positive sample data set based on the positive sample candidate data.
[0017] In an exemplary embodiment, the preference optimization dataset is generated by using the positive sample dataset and the negative sample dataset, including: , collect k positive samples from the positive sample data set according to the preset sample ratio of 1:k , , ; Using the prompt word, the negative sample and the k positive samples Generate a preference optimization sample, and determine the preference optimization data set based on the sample set of all preference optimization samples, the prompt word is composed of corpus quadruple And preset introduction generation.
[0018] In an exemplary embodiment, k positive samples are collected from the positive sample data set according to a preset sample ratio of 1:k. , including: according to the vector encoding results of negative samples and k positive samples The vector encoding result is used to calculate the negative sample and k positive samples The cosine similarity between the two datasets is calculated, and the similarity probability distribution result of the negative sample with respect to the positive sample dataset is calculated based on the cosine similarity; the positive sample dataset is sampled using the similarity probability distribution result to obtain the negative sample dataset. The corresponding k positive samples .
[0019] In an exemplary embodiment, negative samples are calculated by the following formula and k positive samples The cosine similarity between:
[0020] ;
[0021] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
[0022] In an exemplary embodiment, calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset based on cosine similarity includes: calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset using a similarity probability distribution function of the negative sample with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0023] ;
[0024] constant ; The similarity probability distribution result is expressed as .
[0025] In an exemplary embodiment, the method further includes: representing a preference optimization sample in the form of a triple, and the preference optimization dataset is represented as follows:
[0026] , Indicates a prompt word.
[0027] In an exemplary embodiment, before generating a weighted loss function based on the sample weights of the preference optimized samples and the original loss function of the parsing model, the method further includes: determining a sensitivity quantization index of the parsing model based on a perplexity parameter; normalizing a data subset in the preference optimized data set using the sensitivity quantization index, and determining the result of the normalization as the sample weight, wherein the preference optimized samples in the data subset contain prompts with the same intention or prompts with opposite intentions.
[0028] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , S represents the set input sequence of the parsing model, x represents the input prompt word used to prompt the generation of similar intentions, represents the prompt word input for generating the opposite intention, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the response sequence y generated by the analytical model when x is the set input sequence, Indicates that when When the set input sequence is given, the parsing model generates the perplexity of the response sequence y.
[0029] In an exemplary embodiment, the data subset in the preference optimization data set is normalized using the sensitivity quantization index, and the normalization result is determined as the sample weight, including: obtaining a target quantitative index generated according to the sensitivity quantization index; , the target quantitative index represents , Prompts for the same intention Or a hint of the opposite intention The sum of the sensitivity quantification indicators of About Data Subsets The normalized result is determined as the sample weight, and the normalized result is expressed as follows: .
[0030] In an exemplary embodiment, a data augmentation model that uses the weighted loss function to complete the training objectives is obtained, including: iteratively training the strategy model with minimization of the weighted loss function as the training objective, stopping training the strategy model when a preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is minimized, and determining the strategy model as the data augmentation model, and the loss value of the weighted loss function represents the degree of deviation between the strategy model and the reference model.
[0031] According to another aspect of an embodiment of the present application, a speech data processing device based on a large model is also provided, including: a first amplification module, used to perform data amplification operations on the labeled data of an annotated data set to obtain an amplified training data set; an output module, used to use a parsing model to classify the amplified training data set into a positive sample data set and a negative sample data set, and use the positive sample data set and the negative sample data set to generate a preference optimization data set, wherein a preference optimization sample in the preference optimization data set includes a prompt word, negative sample data and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model after recognition training of the initial large language model; a second amplification module, used to generate a weighted loss function based on the sample weight of the preference optimization sample and the original loss function of the parsing model, obtain a data amplification model that uses the weighted loss function to complete the preset training target, and use the data amplification model to perform data amplification on the newly input text data to obtain a target amplified data set.
[0032] According to another aspect of the embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned large model-based speech data processing method when running.
[0033] The present application also provides a computer program product, including a computer program, which implements the above-mentioned large model-based speech data processing method when executed by a processor.
[0034] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the large model-based speech data processing method through the computer program.
[0035] In an embodiment of the present application, a data augmentation operation is performed on the labeled data of the labeled data set to obtain an augmented training data set; the augmented training data set is classified into a positive sample data set and a negative sample data set using a parsing model, and a preference optimization data set is generated using the positive sample data set and the negative sample data set, wherein a preference optimization sample in the preference optimization data set includes a prompt word, negative sample data and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model obtained by performing recognition training on the initial large language model; a weighted loss function is generated according to the sample weight of the preference optimization sample and the original loss function of the parsing model, and a data augmentation model for completing the preset training target using the weighted loss function is obtained, and the data augmentation model is used to perform data augmentation on the newly input text data to obtain a target augmented data set; by adopting the above technical solution, data diversity can be improved by performing data augmentation on the labeled data, and the utilization of The analytical model is used to classify and expand the training data set to obtain a positive sample data set and a negative sample data set, and then a preference optimization data set is constructed based on the correlation between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weights of the preference optimization samples in the preference optimization data set and the original loss function of the analytical model, the quality of different preference optimization samples can be reasonably quantified in a weighted manner, and a data expansion model that uses the weighted loss function to complete the preset training goals is obtained, and the data expansion model is used to complete the data expansion of the newly input text data. By training the model with the weighted loss function combined with the sample weights, the training offset caused by low-quality samples can be reduced, and the model performance is further optimized, so that the model can more accurately understand and generate the required intention corpus when processing the newly input text data, solving the technical problem of low data quality obtained by the existing data expansion method, thereby improving the quality of the expanded data and the accuracy of the model in text processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0038] Figure 1 This is a schematic diagram of the hardware environment of a large-model-based voice data processing method according to an embodiment of the present application;
[0039] Figure 2 is a flowchart of a method for processing speech data based on a large model according to an embodiment of the present application;
[0040] Figure 3 is a schematic diagram of a large model-based voice data processing method according to an embodiment of the present application;
[0041] Figure 4 This is a structural block diagram of a speech data processing device based on a large model according to an embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0044] According to one aspect of the embodiment of the present application, a method for processing speech data based on a large model is provided. The method for processing speech data based on a large model is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residence (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for processing speech data based on a large model can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0045] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may include, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projector, a smart TV, a smart clothes drying rack, smart curtains, a smart audio / video system, a smart socket, a smart speaker, a smart fresh air system, smart kitchen and bathroom equipment, smart bathroom equipment, a smart robot vacuum, a smart window cleaning robot, a smart robot mop, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen appliance, a smart purifier, a smart water dispenser, a smart door lock, and the like.
[0046] In this embodiment, a method for processing speech data based on a large model is provided, which is applied to the above-mentioned terminal device. Figure 2 4 is a flow chart of a method for processing speech data based on a large model according to an embodiment of the present application, the flow includes the following steps:
[0047] Step S202: performing a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set; wherein the labeled data is collected by a sound pickup device;
[0048] Alternatively, the annotation data may be understood as text data obtained by converting the voice data collected by the sound pickup device into text. The annotation data may include, for example, the annotation domain corresponding to the seed corpus, the annotation intent, the slot of the annotation entity, and the like.
[0049] Step S204: Using a parsing model, the augmented training dataset is classified into a positive sample dataset and a negative sample dataset, and a preference-optimized dataset is generated using the positive sample dataset and the negative sample dataset. A preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples. The parsing model is a model obtained by performing recognition training on the initial large language model.
[0050] Step S206: Generate a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data augmentation model that uses the weighted loss function to complete the preset training target, and use the data augmentation model to perform data augmentation on the newly input text data to obtain a target augmented data set.
[0051] Through the above steps, data augmentation operation is performed on the labeled data of the labeled data set to obtain an augmented training data set; the analytical model is used to classify the augmented training data set into a positive sample data set and a negative sample data set, and the positive sample data set and the negative sample data set are used to generate a preference optimization data set, wherein a preference optimization sample in the preference optimization data set includes a prompt word, negative sample data and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples, and the analytical model is a model obtained by performing recognition training on the initial large language model; a weighted loss function is generated according to the sample weight of the preference optimization sample and the original loss function of the analytical model, and a data augmentation model that uses the weighted loss function to complete the preset training target is obtained, and the data augmentation model is used to perform data augmentation on the newly input text data to obtain a target augmented data set; by adopting the above technical solution, data diversity can be improved by performing data augmentation on the labeled data, and the target amplified data set can be obtained by utilizing the above technical solution. The parsing model classifies and amplifies the training data set to obtain a positive sample data set and a negative sample data set, and then constructs a preference optimization data set based on the correlation between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weights of the preference optimization samples in the preference optimization data set and the original loss function of the parsing model, the quality of different preference optimization samples can be reasonably quantified in a weighted manner, and a data amplification model that uses the weighted loss function to complete the preset training goals is obtained, and the data amplification model is used to complete the data amplification of the newly input text data. By training the model with the weighted loss function combined with the sample weights, the training offset caused by low-quality samples can be reduced, and the model performance is further optimized, so that the model can more accurately understand and generate the required intention corpus when processing the newly input text data, solving the technical problem of low data quality obtained by the existing data amplification method, thereby improving the amplified data quality and improving the accuracy of the model in text processing.
[0052] In an exemplary embodiment, before using the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, further, current labeled data can be generated based on the current labeling operation of the target object on the initial data, and the labeled dataset can be generated based on the historical labeled data provided by the target object and the current labeled data, wherein the labeled dataset at least includes the question sentence of the target object, the labeled domain of the target object, the labeled intention of the target object, and the slots corresponding to the labeled entities of the target object; the fine-tuning instruction of the target object is used to instruct the initial large language model to perform a recognition operation on the labeled data, and the recognition result is output according to the output requirements of the fine-tuning data in the fine-tuning instruction, and the recognition result at least includes: the labeled domain, the labeled intention, the labeled entity, and the labeled slot.
[0053] Optionally, the annotated domain, annotated intention, annotated entity and annotated slot in the above recognition results respectively correspond to the annotated domain of the target object, the annotated intention of the target object, the annotated entity of the target object and the slot corresponding to the annotated entity of the target object, that is, the annotated data represents the data that the initial large language model is expected to recognize.
[0054] Optionally, the target object can be a person or a physical company. For example, if the target object is an online retail company, the questions may involve product inquiries, order management, customer service, and other aspects. By fine-tuning instructions, the initial large model can be guided to focus on learning these specific areas and intents, thereby generating labeled data that better meets business needs.
[0055] In an optional embodiment, the large model fine-tuning process is described in conjunction with the following embodiments. The large model fine-tuning process is the process of using the fine-tuning instruction of the target object to instruct the initial large language model to perform recognition operations on the labeled data, and outputting the recognition results in accordance with the output requirements of the instruction fine-tuning data in the fine-tuning instruction. Specifically, the LLaMA2-base-7B model is used as the initial large language model, and a seed data set is obtained through manual construction and online data annotation. The seed data set is then used to fine-tune the LLaMA2-base-7B model to train the model to generate the correct domain intent and slots under given instructions. The fine-tuned model should have the ability to identify domains, intents, and slots, and be able to understand and execute specific business rules, such as semantic parsing tasks in home appliance control scenarios, to provide a basis for subsequent data amplification.
[0056] In an exemplary embodiment, a data augmentation operation is performed on the labeled data of an annotated data set to obtain an augmented training data set, including: selecting seed corpus data to be augmented from the labeled data, and determining the seed annotation field, seed annotation intent, and target intent corresponding to the seed corpus based on the seed corpus data; using a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus within the seed annotation field, thereby obtaining an augmented training data set, wherein the preset augmentation prompt statement at least includes the task information required for the initial large language model to perform the data augmentation task, and the task information at least includes task description information, role description information, task examples, and input seed corpus, and the role description information is used to set the function of the initial large language model to generate a new corpus based on the seed corpus and the seed annotation intent, and the new corpus has the target intent. This embodiment facilitates the large model to understand and perform complex intent conversion tasks by clearly providing the task description information, role description information, task examples, and input seed corpus in the task information of the preset augmentation prompt statement, greatly enriching the diversity of training data, thereby improving the generalization ability of the model.
[0057] It is understandable that determining the seed annotation domain, seed annotation intent, and target intent corresponding to the seed corpus based on the seed corpus data actually involves obtaining each seed corpus from the seed corpus data and determining the seed annotation domain, seed annotation intent, and target intent for each seed corpus in turn. If the seed corpus is "I want to buy a pair of sneakers," the seed annotation intent is "purchase goods," and the target intent is "return policy," then the preset augmentation prompt sentence could be, for example, "Please generate a sentence asking about the return policy based on the following sentence," along with a description of the relationship between "purchase goods" and "return policy."
[0058] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention within the seed annotation domain, so as to obtain an augmented training data set, including: for a set of annotation intentions belonging to the same seed annotation domain, screening a first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions of the annotation intention set, generating an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar to or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a first extended prompt sentence is determined from the preset extended prompt sentence, the first role description information in the first extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the first seed marking intention; the first extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation for similar intents to obtain an augmented training dataset. In this embodiment, symmetrical intents (i.e., similar intents or opposite intents) are represented through intent pairs. This embodiment introduces specific prompts to guide the large language model to generate more relevant corpus within a given domain and intent, thereby enriching the training dataset and improving the model's performance on specific tasks.
[0059] Optionally, the first seed annotation intent represents the original intent , the second sub-labeling intention is to express the symmetric intention of the original intention . Take the intent pair as an example. For example, the intent pair includes "open-close, increase the temperature-decrease the temperature, start-pause", then "open, increase the temperature, start" corresponds to the first seed labeling intent, "close, decrease the temperature, pause" corresponds to the second seed labeling intent, and the target intent of the new corpus is similar to the first seed labeling intent, and the target intent is, for example, "open, heat up, start".
[0060] By combining role descriptions with seed data, the above-mentioned examples enable the model to learn a wider range of expressions and intent variations, addressing data scarcity. For example, in a customer service conversation system, by augmenting conversation examples with similar or opposite scenarios like "querying order status" (e.g., "cancelling an order"), the model's ability to understand and handle these complex scenarios can be enhanced, improving the user experience.
[0061] In an exemplary embodiment, a preset augmentation prompt sentence is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention in the seed annotation domain, so as to obtain an augmented training data set, including: screening other annotation intentions except the first seed annotation intention from the annotation intention set, and the other annotation intentions only have similar intentions; determining a second extended prompt sentence from the preset extended prompt sentence, and the second role description information in the second extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to the other annotation intentions; using the second extended prompt sentence to instruct the initial large language model to target Seed corpus Perform data augmentation for similar intents to obtain an augmented training dataset. In this embodiment, for original intents for which the annotated intent set does not have symmetric intents, an initial large language model is set to generate new corpus based on the seed corpus and these similar intents. This ensures that the target intent of the new corpus is similar to the original intent, avoids generating corpus that is irrelevant or contradictory to the original intent, and ensures the quality of the augmented data.
[0062] The above embodiment can maintain the domain and meaning through the guidance of accurate screening and role description information. Figure 1 While maintaining consistency, creating more variations effectively increases the diversity and coverage of training data. For example, in an intelligent recommendation system, by expanding user queries similar to "view product details," such as "learn about product features," the model can be trained more comprehensively, enabling it to make more accurate recommendations based on diverse user needs.
[0063] In an exemplary embodiment, a preset augmentation prompt statement is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention within the seed annotation domain, so as to obtain an augmented training data set, including: for a set of annotation intentions belonging to the same seed annotation domain, screening a first seed annotation intention with a similar intention or an opposite intention from the seed annotation intentions of the annotation intention set, generating an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar to or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a third extended prompt sentence is determined from the preset extended prompt sentence, and the third role description information in the third extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the second seed marking intention; the third extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation operations on opposite intentions to obtain augmented training data sets. This embodiment generates intention pairs for the first seed annotation intentions with similar or opposite intentions, and then determines the third extended prompt sentence from the preset extended prompt sentence, whose role description information guides the initial large language model to generate the second seed annotation attention sentence. Figure 1 The target intention is opposite to the original intention.
[0064] By clearly distinguishing and amplifying corpora with opposite intentions, this embodiment enhances the model's ability to handle opposing situations, which is crucial for building a dialogue system that can understand complex user intentions. For example, in the question-and-answer system of an online education platform, amplifying instances of "negative answers" and "positive answers" can help the model better distinguish between user questions and statements, and improve the accuracy of interactions. By leveraging the generation capabilities of large models, this embodiment achieves efficient amplification of existing annotated data through prompt words and role description information, and can not only generate samples with similar intentions, but also generate samples with opposite intentions, thereby building a more comprehensive and balanced training data set.
[0065] Based on the above example, the following implementation process is proposed: select a seed corpus containing a specific intent, for example, "Lower the air conditioner temperature." Then, construct a prompt to guide the model to generate text that meets the expected intent. For example, the prompt could be, "You are a voice assistant in an air conditioning control scenario. You need to generate a sentence with the same intent as the given query, but with a different expression."
[0066] The seed corpus and the above prompt are input into the LLaMA2 model together, so that the model can generate the original meaning of the seed corpus according to the prompt instruction. Figure 1 New sentences that are similar but express different meanings.
[0067] To ensure that the generated new sentences are diverse, you can adjust the temperature parameter when the model is generated. Temperature is a parameter that controls the randomness of the generated text. A lower temperature will cause the generated text to be more inclined to the most likely output learned by the model, while a higher temperature will make the generated text more random and diverse. In this scenario, setting the temperature appropriately can increase the diversity of the generated sentences while maintaining their consistency with the seed corpus. Figure 1 Consistency.
[0068] Next, the generated new sentences are automatically verified using manual inspection, domain rules, or a fine-tuned semantic parsing model to ensure that they meet the expected intent and are grammatically and contextually reasonable. New sentences that pass verification and express the same intent as the seed corpus, but are not identical, are retained as part of the augmented dataset.
[0069] Through the above process, we can use the generative capabilities of the LLaMA2 model and the guiding role of prompts to effectively amplify text data that has the same intent as the seed corpus but different expressions, thereby enriching the training set and improving the model's generalization ability and performance when handling actual business scenarios.
[0070] In an exemplary embodiment, the labeled dataset is represented as , is the seed corpus to be amplified, Label the field for the seed, The intention of the seed is labeled and the training dataset is expanded as , For the stated purpose, , i, M, and N are positive integers. This embodiment clearly defines the composition of the dataset and the representation of the augmented data by structured representation of the annotated dataset, which helps the model understand and learn the relationship between different fields and intentions, and can significantly improve the computing power of the model.
[0071] In an exemplary embodiment, the augmented training data set is classified into a positive sample data set and a negative sample data set using a parsing model, including: inputting data of the augmented training data set into the parsing model so that the parsing model classifies the augmented training data and outputs a prediction domain corresponding to the augmented training data. and predictive intent ;according to and The positive sample data set and the negative sample data set are determined by the verification result, and the verification result includes the seed annotation field and the predicted areas First verification results and seed annotation intentions and the predicted intent The second verification result of this embodiment uses a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, including inputting data of the augmented training dataset into the parsing model, outputting a predicted field and a predicted intent, and determining the positive and negative sample datasets based on the verification result.
[0072] For example, for the expanded text , further input into the analytical model for classification, and obtain the prediction field and predictive intent , further record the prediction domain and prediction intention as .
[0073] The above-described embodiment uses the classification processing function of the analytical model to automatically filter out positive samples that meet the expected intent and negative samples that do not, reducing the workload of manual inspection and accelerating data expansion. For example, in a social media monitoring system, it can quickly identify and classify a large number of user comments, using comments that align with the intent of "positive brand evaluation" as positive samples and comments that contradict this intent as negative samples. This can effectively train the model to identify and filter negative information on the Internet.
[0074] In one exemplary embodiment, according to and The positive sample data set is determined based on the verification result of the first verification result, including: determining that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention. Figure 1 If the seed corpus is consistent, the seed corpus corresponding to the seed annotation field is stored in the positive sample data set according to the first mapping relationship. The first mapping relationship represents the corpus quadruple of the seed corpus to be amplified in the annotated data. The relationship between the key and the positive sample data set as the value, is the seed corpus to be amplified, Label the field for the seed, labeling the seed with intent, This embodiment is based on the and The positive sample data set is determined based on the verification result, including that the seed annotation field is consistent with the predicted field, and the seed annotation intention is consistent with the predicted intention. Figure 1 If the results are consistent, the seed corpus is stored in the positive sample dataset.
[0075] Optionally, the first mapping relationship is represented by maintaining a mapping table, in which the corpus quadruple For key values, maintain two lists based on the key value and They are used to store candidate data that can be used as positive samples and negative samples respectively. The preference optimization dataset is constructed by verifying whether the predicted domain intention is the same as the expected domain intention. For example, in this embodiment, the seed annotation domain is consistent with the predicted domain and the seed annotation intention is consistent with the predicted intention. Figure 1 The consistent seed corpus is stored in the positive sample dataset to construct the positive sample dataset, which improves the accuracy of the amplified data and can effectively filter out irrelevant or erroneous amplified data, ensuring that the model learns high-quality and highly relevant information.
[0076] In one exemplary embodiment, according to and The negative sample data set is determined based on the verification result of the first verification result, including: determining that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention. Figure 1 If any of the above is not true, the seed corpus corresponding to the seed annotation field is stored in the negative sample dataset according to the second mapping relationship. , the second mapping relationship is represented by The relationship between the key and the negative sample dataset is the value. and The verification result of the negative sample dataset is determined to be consistent with the seed annotation field and the prediction field, and the seed annotation intention is consistent with the prediction intention. Figure 1 If any of the above conditions are not met, the seed corpus is stored in a negative sample dataset. By constructing a negative sample dataset based on corpora where intent recognition fails or domain division is incorrect, the model can be improved in a targeted manner, performing better in complex or edge cases. For example, in a voice command recognition system for smart home devices, using "turn off the lights" corpora that are mistakenly recognized as "turn on the lights" commands as negative samples can help the model learn to distinguish between similar but opposite commands, improving the user experience.
[0077] Optionally, the seed annotation field is consistent with the predicted field, and the seed annotation intention is consistent with the predicted intention. Figure 1 The situation, for example, is expressed as .
[0078] The seed annotation field is consistent with the predicted field, and the seed annotation intention is consistent with the predicted intention. Figure 1 If any of the above conditions does not hold, it can be expressed as: ; ; .
[0079] In an exemplary embodiment, the following technical solution is further proposed: the augmented training data set is input into the parsing model so that the parsing model can identify the augmented training data and obtain the slot relationship corresponding to the augmented training data; the slot relationship is replaced with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data to the same sentence structure as the seed corpus to be augmented, the updated augmented training data is determined as the positive sample candidate data, and the positive sample data set is determined based on the positive sample candidate data. This embodiment can use the parsing model to identify the slots of the augmented training data, and replace the identified slots with the slot values of the seed corpus, thereby constructing positive sample data with the same sentence structure and adding it to the positive sample candidate set. This embodiment improves model training efficiency and prediction accuracy by unifying the sentence structure of the data, ensuring structural consistency between the augmented data and the seed corpus. For example, in natural language processing tasks, by adjusting the sentence structure of the augmented data to the same framework as the seed corpus, the model can more easily capture the core semantics and avoid misunderstandings caused by differences in sentence structure.
[0080] Optionally, the slot replacement method in the above embodiment can be applied to situations where different domains are expected to be predicted but the intent is the same. For example, the expected prediction domain is "Oven", the prediction intent is "increase Temperature", and the expanded corpus is "The steamer temperature should be higher". Then, this corpus is input into the parsing model, with "Steamer" as the prediction domain and "increase Temperature" as the prediction intent. The slot is "(device)" corresponding to the steamer, and the slot relationship is the corresponding relationship between the steamer and "(device)". The corpus marked with the slot is "[(device) steamer] The temperature should be higher". At this time, the value of the (device) slot can be replaced with the device name corresponding to the prediction domain, such as oven. The replaced corpus is "The oven temperature should be higher". For example, if the original seed corpus is "Increase the steamer temperature" and the seed slot relationship is the corresponding relationship between steamer and "(device)", then the value of the (device) slot is replaced with oven, and the resulting corpus is "Increase the oven temperature".
[0081] In an exemplary embodiment, the preference optimization dataset is generated by using the positive sample dataset and the negative sample dataset, including: , collect k positive samples from the positive sample data set according to the preset sample ratio of 1:k , , ; Using the prompt word, the negative sample and the k positive samples Generate a preference optimization sample, and determine the preference optimization data set based on the sample set of all preference optimization samples, the prompt word is composed of corpus quadruple and a preset introduction is generated. This embodiment uses the positive sample data set and the negative sample data set to generate a preference optimization data set, including for each negative sample, collecting positive samples in proportion from the positive sample data set, using prompt words, negative samples and positive samples to generate preference optimization samples, and then determining the preference optimization data set. Comparative learning can enhance the model's ability to identify correct intent and domain. For example, in the product search system of an online shopping platform, by comparing user queries "find the latest mobile phone" and "find the cheapest mobile phone" (assuming the former is a positive sample and the latter is a negative sample), the model can learn how to distinguish users' purchasing tendencies and provide more personalized search results.
[0082] In an exemplary embodiment, k positive samples are collected from the positive sample dataset according to a preset sample ratio of 1:k. , including: according to the vector encoding results of negative samples and k positive samples The vector encoding result is used to calculate the negative sample and k positive samples The cosine similarity between the two datasets is calculated, and the similarity probability distribution result of the negative sample with respect to the positive sample dataset is calculated based on the cosine similarity; the positive sample dataset is sampled using the similarity probability distribution result to obtain the negative sample dataset. The corresponding k positive samples .
[0083] Optionally, the vectors of the positive samples and the negative samples are encoded respectively to obtain the vector encoding results of the negative samples and the vector encoding results of the positive samples, and the similarity probability distribution result of each negative sample with respect to the positive sample set is calculated, and this distribution result is used to sample from the positive sample candidate set. , and then get k positive samples , and then k preference samples can be constructed.
[0084] In an exemplary embodiment, negative samples are calculated by the following formula and k positive samples The cosine similarity between:
[0085] ;
[0086] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
[0087] In an exemplary embodiment, calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset based on cosine similarity includes: calculating the similarity probability distribution result of the negative sample with respect to the positive sample dataset using a similarity probability distribution function of the negative sample with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0088] ;
[0089] constant ; The similarity probability distribution result is expressed as .
[0090] In an exemplary embodiment, the following technical solution is further proposed: a preference optimization sample is represented in the form of a triple, and the preference optimization dataset is represented as follows:
[0091] , Indicates a prompt word.
[0092] In an exemplary embodiment, before generating a weighted loss function based on the sample weights of the preference-optimized samples and the original loss function of the parsing model, the following technical solution is further proposed: determining a sensitivity quantification index of the parsing model based on a perplexity parameter; using the sensitivity quantification index to normalize a data subset in the preference-optimized dataset, and determining the normalization result as the sample weight, wherein the preference-optimized samples in the data subset contain prompts with the same intention or prompts with opposite intentions. It should be noted that the sample weight reflects the degree of importance the model attaches to different samples. This embodiment can quantify the model's sensitivity to different prompts by calculating the perplexity difference, and then adjust the influence of the sample in the training process. For example, in an intelligent translation system, by analyzing the perplexity difference when the model processes the two prompts "formal language" and "informal language", the proportion of these two corpora in training can be dynamically adjusted, so that the model can translate fluently in both formal and informal occasions.
[0093] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , S represents the set input sequence of the parsing model, x represents the input prompt word used to prompt the generation of similar intentions, represents the prompt word input for generating the opposite intention, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the response sequence y generated by the analytical model when x is the set input sequence, Indicates that when When the set input sequence is given, the parsing model generates the perplexity of the response sequence y.
[0094] In an exemplary embodiment, the data subset in the preference optimization data set is normalized using the sensitivity quantization index, and the normalization result is determined as the sample weight, including: obtaining a target quantitative index generated according to the sensitivity quantization index; , the target quantitative index represents , Prompts for the same intention Or a hint of the opposite intention The sum of the sensitivity quantification indicators of About Data Subsets The normalized result is determined as the sample weight, and the normalized result is expressed as follows:
[0095] .
[0096] .
[0097] In one exemplary embodiment, obtaining a data augmentation model that achieves a training objective using the weighted loss function includes iteratively training a policy model with minimizing the weighted loss function as the training objective. Upon reaching a predetermined number of iterative training iterations and the minimum value of the weighted loss function in the current iteration, training the policy model is stopped and the policy model is determined as the data augmentation model. The loss value of the weighted loss function represents the degree of deviation between the policy model and a reference model. In this embodiment, obtaining a data augmentation model that achieves a training objective using the weighted loss function includes iteratively training the policy model with minimizing the weighted loss function as the training objective until an optimal state is reached. By assigning different weights to different samples and introducing a weighted loss function, the learning direction of the model can be more precisely controlled during training, avoiding overfitting caused by excessive focus on certain samples. For example, in an obstacle recognition module of an autonomous driving system, assigning higher weights to samples of key obstacles such as "pedestrians," "bicycles," and "cars" can ensure that the model has stronger recognition capabilities for these important categories, thereby improving driving safety.
[0098] It's important to note that in the context of reinforcement learning or preference optimization, the policy model refers to the model being trained, while the reference model is the baseline model used to evaluate the policy model's performance, or a raw model that hasn't been specifically trained. In machine learning and deep learning, the loss function is a metric used to measure the difference between a model's predictions and actual results. The training goal is to find a set of model parameters that minimizes the loss function, thereby producing predictions that are as close to the actual values as possible. During model training, the smaller the loss function, the smaller the deviation between the policy model and the reference model. This means the model's prediction error is minimized, improving the model's fit to the training data. When the loss function reaches its minimum, the model theoretically performs best on the training data, but this must be balanced against the risk of overfitting (i.e., the model becomes overly dependent on the training data, resulting in reduced generalization ability).
[0099] In an exemplary embodiment, the original loss function is expressed as follows:
[0100] ;
[0101] Indicates the input prompt. represents negative samples, represents a positive sample, For the strategy model, is the reference model, is a hyperparameter, is the impulse function, .
[0102] In an exemplary embodiment, the weighted loss function is expressed as follows:
[0103] ;
[0104] Indicates the input prompt. represents negative samples, represents a positive sample, For the strategy model, is the reference model, is a hyperparameter, is the impulse function.
[0105] In order to better understand the process of the above-mentioned large-model-based voice data processing method, the implementation method flow of the above-mentioned large-model-based voice data processing is described below in combination with optional embodiments, but it is not used to limit the technical solution of the embodiments of this application.
[0106] In this embodiment, a method for processing speech data based on a large model is provided. Figure 3 is a schematic diagram of a method for processing speech data based on a large model according to an embodiment of the present application, such as Figure 3 As shown, the specific steps are as follows:
[0107] Step S301: instructing fine-tuning of a semantic parsing model (corresponding to a parsing model).
[0108] First, the voice data collected by the audio pickup device is converted into text data. This text data is then manually constructed and annotated online to create a labeled dataset. This dataset includes the user's query, the annotated domain, the intent, and the slots to be extracted. Next, a command fine-tuning dataset is constructed and fine-tuned using the LLaMA2-base-7B model as the base model. This results in a semantic parsing model that can classify domain intents and recognize entities according to business rules.
[0109] The following are examples of labeled data:
[0110] {
[0111] "Now grill the chicken wings for me for a little longer": {
[0112] "domain": "recipe",
[0113] "intent": "increaseTime",
[0114] "slots": "Now cook [(recipe) grilled chicken wings] for me a little longer."
[0115] };
[0116] "Please bake in the oven for a longer time": {
[0117] "domain": "Oven",
[0118] "intent": "increaseTime",
[0119] "slots": "[(device) oven] bakes for a little longer."
[0120] };
[0121] "Add another quarter of an hour for the washing machine to wash": {
[0122] "domain": "Washer",
[0123] "intent": "increaseTime",
[0124] "slots": "Add another quarter of an hour for washing machine washing time"
[0125] };
[0126] }.
[0127] Among them, the instruction fine-tuning data of the first labeled data example can be expressed as:
[0128] {
[0129] "prompt": "You are a semantic parser for smart home appliance voice control. You need to identify the domain, intent, and slot from a given query.\nRequirements:\n1. No parsing is required; simply output a JSON response with domain, intent, and slots as keys. 2. If there is no domain or intent, set its value to other.\nInput: \"Now, grill the chicken wings for me for a little longer\" \nAnswer: ",
[0130] "target": "{\"domain\": \" recipe \",\"intent\": \"increaseTime\",\"slots\": \"Now cook [(recipe) Grilled Chicken Wings] for a little longer\"}"
[0131] }.
[0132] Using the LLaMA2-base-7B model as the base model, load the pretrained weights and define fine-tuning parameters such as the learning rate, number of fine-tuning epochs, and batch size. Then, load the fine-tuning dataset (e.g., the labeled dataset described above) into the model's training process. The model is trained to extract domain, intent, and slot information from the query. During fine-tuning, the model adjusts its parameters to minimize loss based on the difference between the prompt and target. After fine-tuning, save the optimized model weights, which will serve as the basis for subsequent data augmentation and preference optimization. Next, prepare a test dataset not used in training and use it to evaluate the fine-tuned model's performance on domain, intent, and slot recognition tasks, including precision, recall, and F1 score. Through these steps, a model with preliminary semantic parsing capabilities can be obtained through fine-tuning based on the LLaMA2-base-7B model, paving the way for further data augmentation and preference optimization.
[0133] Step S302: Consent graph data augmentation and cross-intent data augmentation based on prompt learning.
[0134] The use of prompt learning technology can be used to amplify the model's generation capabilities to increase the diversity and colloquial expressions of text data. Consensus graph data amplification refers to the generation of different expressions under the same intent, that is, while keeping the original intent unchanged, variants of the seed corpus are generated to cover different expressions under the same intent. Cross-intent data amplification, on the other hand, is the generation of new text based on similar or opposite intentions, that is, while keeping the domain unchanged, the intent of the seed corpus is reversed or symmetrically operated to generate new corpus that is opposite to the original intent. For example, if the intent of the seed corpus is "increase time", then the goal of opposite intent amplification is to generate a sentence of "reduce time". This step, by setting different prompts, prompts the model to learn the positive and negative conversion of intent, as well as the diverse expressions of intent within the same domain.
[0135] The labeled dataset is used as the seed dataset, which is denoted as: ,in is the seed corpus to be amplified, This is the domain and intent of the seed corpus. At this stage, in order to maximize the use of the annotated information, the original seed corpus can be expanded towards similar intents and opposite intents respectively, including the following steps:
[0136] 1. First, for each domain's intent set, identify the intent pairs that are suitable for corpus conversion. For example, there are symmetrical intents such as turn on / off, increase / decrease the temperature, start / pause, etc. These pairs are recorded as intentPairs = .For example:
[0137] intentPairs=
[0138] {
[0139] "open": "close",
[0140] "close": "open",
[0141] "startup": "suspend",
[0142] "suspend": "startup",
[0143] "increaseTemp": "decreaseTemp",
[0144] "decreaseTemp": "increaseTemp"
[0145] }.
[0146] For cases where there is no symmetric intent, such as status query, parameter adjustment, etc., only the corpus expansion of the same graph (i.e., similar intent) is considered.
[0147] 2. Construction of prompt words:
[0148] The prompt word Prompt constructed in this application consists of three parts: 1) role description and task description, 2) provided examples, and 3) seed corpus input. The role description, task description, provided examples, and seed corpus input correspond to the task description, role description, task examples, and seed corpus input in the task information above, respectively.
[0149] Furthermore, task descriptions and examples can be provided for different fields to improve the diversity and quality of data after amplification. For example, is the domain and intent of the corpus, and the target intent is , It is from Three pieces of data sampled from the subset of domain intent.
[0150] For example, the role description is: "You are a corpus generator in the smart home appliance voice control scenario. You need to amplify 3 texts to the target intent intent_target based on the given query and the corresponding intent."
[0151] The task description is: "You will focus on the field < >, this field is mainly used for < >Related voice control. The augmented text is required to be as diverse and colloquial as possible."
[0152] An example task is: "\nFor example: query: < >,domain: < >, intent: < >,intent_target: < >, \n Output: ”.
[0153] Seed corpus input: "Based on the above information,\n Input: query: < >,domain: < >,intent: < >, intent_target: < >, \n Output:".
[0154] The constructed prompt is: "You are a corpus generator for the smart home appliance voice control scenario. You need to amplify three texts towards the target intent_target based on a given query and intent. You will focus on the recipe domain, which is primarily used for recipe-related voice control. The amplified text should be as diverse and colloquial as possible." For example: {\"query\": \"Roast the chicken wings for a little longer now\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"},\n Output: "[\"You don't need to roast the chicken wings for that long\", \"You don't need to roast the chicken wings for too long\", \"Please reduce the roasting time for the chicken wings."] Based on the above information,\n Input: {\"query\": \"Now bake the chicken wings for a little longer\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"}, \n Output:".
[0155] The constructed prompt (corresponding to the preset extended prompt sentence) is input into the original LLaMA2-base model, and data is generated for each seed corpus respectively to the same graph and symmetric intent. If the target intent is the same as the original intent, the data augmentation of the graph is agreed. In addition, a larger temperature parameter can be set to increase the diversity of generated data.
[0156] Step S303: Construct a preference optimization data set.
[0157] A preference-optimized dataset is constructed based on the augmented data obtained in step S302. In this step, the augmented data is predicted using the fine-tuned model and classified into positive samples (consistent with the business preference) and negative samples (inconsistent with the business preference). Then, by calculating the similarity distribution, online sampling is used to select the samples from the positive sample candidate set that are closest to the negative samples but correctly classified as positive samples to construct preference sample pairs. The purpose of constructing the preference-optimized dataset is to provide the positive and negative sample pairs required for model training in the subsequent preference optimization process, thereby optimizing the consistency of the model-generated text with actual business expectations. Specifically:
[0158] 1) Combine the model fine-tuned in step S301 to collect candidate positive and negative samples.
[0159] 2) Based on online sampling of similarity distribution, appropriate preference sample pairs are constructed.
[0160] For ease of presentation, this application denotes text generated by the desired model as positive samples, or chosen samples. Text generated by the undesirable model is denotes negative samples, or rejected samples. Positive samples are expected to be more consistent with business preferences and expectations than negative samples.
[0161] For the seed dataset A seed corpus in , assuming the target intention is , then the domain intentions of the data expected to be expanded are , the text set after the amplification in step S302 is recorded as .
[0162] Furthermore, step S303 further includes:
[0163] Step S3031: Collection of positive sample candidate sets (corresponding to positive sample data sets) and negative sample candidate sets (corresponding to negative sample data sets).
[0164] First, with the quad Maintain a mapping table for the key. Under this key, two lists need to be maintained , which are used to store candidate data that can be used as positive samples and negative samples respectively.
[0165] For the expanded text , input it into the semantic parsing model fine-tuned in step 1 for classification processing, and record the predicted domain intention as .
[0166] Furthermore, whether the predicted domain intent is the same as the expected domain intent includes the following four cases:
[0167] 1. : It is consistent with the domain intent of the target and consistent with actual business expectations.
[0168] 2. : Different fields but agree on FIG.
[0169] 3. : Same field but different intentions.
[0170] 4. : Different fields and different intentions.
[0171] Based on the differences between the prediction results reflected in the above four situations and the expected domain intentions, the positive and negative sample candidate sets are screened out, and then the preference optimization dataset is constructed.
[0172] Belong to case 1 : Can be used as a positive sample and added to the positive sample candidate set middle.
[0173] Belong to case 2 : Can be used as negative samples and added to the negative sample candidate set In addition, the semantic parsing model of step S301 can be input again to obtain the slot recognition result, and the slot recognized can be replaced with the slot value of the seed corpus to construct the chosen data of the same sentence and add it to the positive sample candidate set. middle.
[0174] Belong to situations 3 and 4 : Both can be used as negative samples and added to the negative sample candidate set middle.
[0175] Step S3032: Online sampling based on similarity distribution.
[0176] Through S3031, we get the four-tuple Positive sample candidate set for key and negative sample candidate sets Next, we need to further construct paired positive and negative samples as preference samples. Here we use an online sampling method based on similarity distribution:
[0177] Assume that for each negative sample , need to be in the ratio of 1:k from The k most appropriate positive samples are selected from the dataset to construct k positive and negative sample pairs. To improve the sample training efficiency of the model during the preference optimization phase, samples with similar semantics to the negative samples but different labels (i.e., domain or intent) should be selected as positive samples to provide the model with fine-grained supervision information on the differences between positive and negative samples. The specific steps are as follows:
[0178] 1) Vector encoding of positive and negative samples: The language model to be optimized in the preference optimization phase is ; Traverse each negative sample , respectively, the negative samples , and each positive sample in the positive sample candidate set Input to model Encode, obtain the vector output by the last layer and perform average pooling to use it as the vector representation of positive and negative samples, respectively recorded as vectors and , .
[0179] 2) Calculating Similarity: Negative Samples With each positive sample The cosine similarity can be expressed as:
[0180] .
[0181] 3) Calculation Regarding the similarity probability distribution of the positive sample set: softmax normalization is used here, and a constant is added to each item , in order to maintain the stability of the value. The similarity probability distribution function is specifically expressed by the following formula:
[0182] .
[0183] 4) Sampling: In each iteration of the preference optimization phase, for each negative sample , according to the above formula, the probability distribution of the similarity of negative samples to the positive sample set is calculated, and based on this distribution, samples are sampled from the positive sample candidate set , and then get k positive samples , .
[0184] 5) Construct k preference samples: Each sample in the preference optimization dataset consists of three parts: prompt, chosen, and rejected. Specifically, through each four-tuple And the pre-written instruction (corresponding to the preset introduction), you can construct the following prompt:
[0185] "<instruction>。\nInput: {\"query\": i >,\"domain\": <D i >,\"intent\": i >, \"intent_target\":< >} \nAnswer: ".
[0186] Next, for each negative sample , according to step 4) Positive samples , and then we can construct Preference samples, each preference sample is as follows:
[0187] {
[0188] "prompt": " <instruction>。\nInput: {\"query\": i >,\"domain\":< D i >,\"intent\": < I i >, \"intent_target\":< >} \nAnswer: ",
[0189] "chosen": "< >",
[0190] "rejected": "< >"
[0191] }.
[0192] Based on the above sampling strategy, in the positive sample candidate set, compared with the negative sample Positive samples with more similar semantics have a greater probability of being selected. The preference samples constructed based on this provide the model with supervision information on the fine-grained differences between positive and negative examples during the preference optimization stage.
[0193] To facilitate subsequent description, The constructed prompt is recorded as , then each preference sample can be represented by a triple The constructed preference sample set can be expressed as:
[0194] .
[0195] Step S304: Preference optimization, which specifically includes the following steps.
[0196] Step S3041: Construct the original DPO loss function.
[0197] First, the classic Direct Preference Optimization (DPO) loss function can be expressed as the following formula.
[0198] .
[0199] Among them, x is the input prompt, q r is the rejected sample, q c is the chosen sample, π θ is the strategy model, is the language model to be optimized, π ref is the reference model, β is a hyperparameter with a value of [0,1]. The parameters of this model are not updated during the training process, and the untuned base model can be selected. is the sigmoid function, that is .
[0200] In preference optimization, a loss function is designed to measure the performance of a policy model's output relative to a reference model. For example, Direct Preference Optimization (DPO) compares the policy model's output to a reference model for the same input, encouraging the policy model to generate a more optimal response. If the loss function is designed appropriately, minimizing the loss function during training will, in part, reduce the deviation between the policy model and the reference model, meaning that the policy model's decisions or outputs are closer to or better than those of the reference model.
[0201] When improving the DPO loss function, the hyperparameter β controls the degree of deviation between the policy model and the reference model. A smaller β value means the policy model's output is more inclined toward the reference model's output. By adjusting β, we can balance model exploration (generating new data or improving output) with exploitation (maintaining output consistent with the reference model), preventing the policy model from deviating too much from the reference model, resulting in poor output quality or failure to meet business preferences.
[0202] Step S3042: Analyze problems existing in the preference optimization data set.
[0203] Because this application involves generating both consent graphs and cross-intent data, the preference optimization dataset may contain examples where the positive and negative examples in the consent graph generation task are reversed, resulting in examples being used as negative and positive examples in the cross-intent generation task. These two preference examples differ only in the target intent in the prompt. If the model is insensitive to the target intent (intent_target) in the prompt, these example pairs will mislead the model and make it difficult to align with the actual business classification preferences.
[0204] An example of Prompt is shown below.
[0205] {
[0206] "prompt": "You are a corpus generator for the smart home appliance voice control scenario. You need to amplify a text to the target intent_target based on a given query and intent. You will focus on the recipe domain, which is primarily used for recipe-related voice control. The amplified text must be as diverse and colloquial as possible.\nInput: {\"query\": \"Roast the chicken wings for a little longer now\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"decreaseTime\"} \nAnswer: ",
[0207] "chosen": "I want to bake the chicken wings for a shorter time",
[0208] "rejected": "I want to bake the chicken wings longer."
[0209] }.
[0210] {
[0211] "prompt": "You are a corpus generator for the smart home appliance voice control scenario. You need to amplify a text to the target intent_target based on a given query and intent. You will focus on the recipe domain, which is primarily used for recipe-related voice control. The amplified text must be as diverse and colloquial as possible.\nInput: {\"query\": \"Roast the chicken wings for a little longer now\",\"domain\": \"recipe\",\"intent\": \"increaseTime\", \"intent_target\": \"increaseTime\"} \nAnswer: ";
[0212] "chosen": "I want to bake the chicken wings longer",
[0213] "rejected": "I want to bake the chicken wings for a shorter time."
[0214] }.
[0215] Step S3043: Construct a weighted DPO loss to perform preference optimization.
[0216] In order to solve the above problems, in this section, we first define the statistic , which is used to measure the model's sensitivity to two prompts or the degree of compliance with instructions when given the same response from two prompts. Next, we construct sample weights for the preference optimization dataset and improve the original preference optimization loss to finally obtain a weighted loss function. Specifically:
[0217] (1) Define statistics .
[0218] In the language model system, perplexity can be used to measure the uncertainty of the model generating a given sequence, or to evaluate the degree of fit of the model for a given sequence. Specifically, given a sequence , assuming that the conditional probability of the language model generating the i-th word is , then the perplexity of generating sequence S is It can be calculated as:
[0219] .
[0220] Based on the perplexity, the following statistics PPL can be defined diff :
[0221] .
[0222] in, and Represents two input prompts, Represents the response sequence generated by the model . and Respectively indicate when and When input is , the model generates a response sequence The absolute value of the difference between the two is defined as PPL diff .
[0223] By defining the statistic PPL diff , used to evaluate the model's response sensitivity to different prompts, that is, the difference in the model's compliance with different instructions. diff When it is close to 0, it means that the two prompts are used as input to generate There is no significant difference in the perplexity when , that is, the model is not sensitive to the difference between the two prompts when generating the same response. In contrast, when PPL diff When it is large, it means that the model can follow these two different prompt instructions well, that is, the model is more sensitive to the difference between the two prompts when generating the same response. diff It can measure the compliance or sensitivity of the model to two different prompts when outputting the same response. diff Assign weights to samples in the preference optimization dataset, with low-sensitivity samples having lower weights and high-sensitivity samples having higher weights.
[0224] (2) Construct sample weights for the preference optimization dataset and obtain the weighted DPO loss.
[0225] For preference optimization dataset , each preference sample is represented by prompt 、chosen sample , rejected samples Composition. Among them prompt By quad Decision. Assume that in the intentPairs field, The symmetry intention is , then when the target intention When the target intention is , it is cross-intent data expansion; for the convenience of expression, the prompts of the consent graph and cross-intent are respectively and , which can be specifically expressed as:
[0226] .
[0227] .
[0228] Record or The preference data subset for prompt is . Further, define the statistic , used to represent positive and negative sample pairs About and PPL diff The sum can be expressed as:
[0229] .
[0230] Among them, when When it is small, it means that the perplexity of generating positive samples and the perplexity of generating negative samples are both small, that is, the model is and is not sensitive to the difference; When it is larger, it means that the perplexity of at least one of the positive and negative samples is larger, which proves that the model can at least capture the difference between the two prompts.
[0231] Therefore, we hope to increase The larger the weight of the preferred sample, the smaller the The weight of the smaller preferred samples is used to minimize the interference of some samples on the model preference optimization.
[0232] To express it as a weight, About preference data subsets Normalization is performed, specifically expressed as:
[0233] .
[0234] Among them, for prompt and ,That The normalized value is used as the sample weight in the preference optimization loss function, and then the weighted DPO loss function can be obtained:
[0235] .
[0236] The goal of training is to minimize the above loss function. In actual training, you can choose an optimizer such as AdamW to update the parameter θ. For the hyperparameter β, you can choose between 0.1-0.5 to control the strategy model π θ With the reference model π ref The degree of deviation between them.
[0237] This step uses the weighted DPO loss function to more effectively guide model adjustments, reduce training deviations caused by low-quality samples, and ensure that the data generated by the model is more in line with business preferences.
[0238] Based on the above steps, the present application proposes a data generation method based on large model preference optimization. This method constructs suitable preference samples based on the sampling strategy of similarity distribution, and optimizes the model's preferences based on the improved weighted DPO, ensuring consistency with the semantic results expected by the business while improving the diversity and quality of the amplified data. Specifically, the present application first fine-tunes a semantic parsing model for subsequent domain and intent recognition. Then, prompt words are constructed to generate the initial amplified data to ensure the quantity and diversity of the data. Further, a sampling strategy based on similarity distribution is used to construct preference samples, and weights are assigned to different preference samples based on the difference in perplexity, and finally a weighted DPO loss function is obtained to optimize the model's preferences. By using the model optimized by this application, data that is more aligned with actual business preferences can be generated.
[0239] It is understandable that data augmentation methods for annotated data also include: traditional methods such as synonym replacement, NER entity replacement, back translation, random noise injection, and syntax tree enhancement. Based on language models such as GPT-2, labels and original text are spliced to construct prompts, and fine-tuned to directly learn the method of generating data for the label.
[0240] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0241] Figure 4 is a structural block diagram of a speech data processing device based on a large model according to an embodiment of the present application; Figure 4 As shown, including:
[0242] A first amplification module 42 is configured to perform a data amplification operation on the labeled data of the labeled data set to obtain an amplified training data set, wherein the labeled data is collected by a sound pickup device;
[0243] Output module 44 is configured to classify the augmented training dataset into a positive sample dataset and a negative sample dataset using a parsing model, and generate a preference-optimized dataset using the positive sample dataset and the negative sample dataset, wherein a preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model obtained by performing recognition training on the initial large language model;
[0244] The second amplification module 46 is used to generate a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data amplification model that uses the weighted loss function to complete the preset training target, and use the data amplification model to perform data amplification on the newly input text data to obtain a target amplified data set.
[0245] By means of the above-mentioned device, data augmentation operation is performed on the labeled data of the labeled data set to obtain an augmented training data set; the analytical model is used to classify the augmented training data set into a positive sample data set and a negative sample data set, and the positive sample data set and the negative sample data set are used to generate a preference optimization data set, wherein a preference optimization sample in the preference optimization data set includes a prompt word, negative sample data and positive sample data generated based on the labeled data, and each negative sample corresponds to multiple positive samples, and the analytical model is a model obtained by performing recognition training on the initial large language model; a weighted loss function is generated according to the sample weight of the preference optimization sample and the original loss function of the analytical model, and a data augmentation model is obtained to complete the preset training target by using the weighted loss function, and the data augmentation model is used to perform data augmentation on the newly input text data to obtain a target augmented data set; by adopting the above-mentioned technical solution, data diversity can be improved by performing data augmentation on the labeled data, and the target amplified data set can be obtained by utilizing the above-mentioned technical solution. The parsing model classifies and amplifies the training data set to obtain a positive sample data set and a negative sample data set, and then constructs a preference optimization data set based on the correlation between each negative sample and multiple positive samples. By generating a weighted loss function based on the sample weights of the preference optimization samples in the preference optimization data set and the original loss function of the parsing model, the quality of different preference optimization samples can be reasonably quantified in a weighted manner, and a data amplification model that uses the weighted loss function to complete the preset training goals is obtained, and the data amplification model is used to complete the data amplification of the newly input text data. By training the model with the weighted loss function combined with the sample weights, the training offset caused by low-quality samples can be reduced, and the model performance is further optimized, so that the model can more accurately understand and generate the required intention corpus when processing the newly input text data, solving the technical problem of low data quality obtained by the existing data amplification method, thereby improving the amplified data quality and improving the accuracy of the model in text processing.
[0246] In an exemplary embodiment, before using the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, the output module is further used to: generate current labeled data based on the current labeling operation of the target object on the initial data, and generate the labeled dataset based on the historical labeled data provided by the target object and the current labeled data, wherein the labeled dataset at least includes the question sentence of the target object, the labeled domain of the target object, the labeled intention of the target object, and the slots corresponding to the labeled entities of the target object; use the fine-tuning instruction of the target object to instruct the initial large language model to perform a recognition operation on the labeled data, and output the recognition result according to the output requirements of the fine-tuning data in the fine-tuning instruction, and the recognition result at least includes: the labeled domain, the labeled intention, the labeled entity, and the labeled slot.
[0247] In an exemplary embodiment, the first augmentation module is also used to: select seed corpus data to be augmented from the annotation data, and determine the seed annotation field, seed annotation intent and target intent corresponding to the seed corpus based on the seed corpus data; use a preset augmentation prompt statement to instruct the initial large language model to perform a data augmentation task from the seed annotation intent to the target intent on the seed corpus within the seed annotation field, to obtain an augmented training data set, the preset augmentation prompt statement at least includes the task information required for the initial large language model to perform the data augmentation task, the task information at least includes task description information, role description information, task examples, and input seed corpus, the role description information is used to set the function of the initial large language model to generate a new corpus based on the seed corpus and the seed annotation intent, and the new corpus has the target intent.
[0248] In an exemplary embodiment, the first amplification module is further used to: for a set of annotation intentions belonging to the same seed annotation field, screen out a first seed annotation intention with similar or opposite intentions from the seed annotation intentions of the annotation intention set, generate an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a first extended prompt sentence is determined from the preset extended prompt sentence, the first role description information in the first extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the first seed marking intention; the first extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
[0249] In an exemplary embodiment, the first expansion module is also used to: filter out other annotation intentions other than the first seed annotation intention from the annotation intention set, and the other annotation intentions only have similar intentions; determine a second extended prompt sentence from the preset extended prompt sentence, and the second role description information in the second extended prompt sentence is used to set the function of the initial large language model to generate a new corpus based on the seed corpus and the other annotation intentions, and the target intention of the new corpus is similar to that of the other annotation intentions; use the second extended prompt sentence to instruct the initial large language model to target Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
[0250] In an exemplary embodiment, the first amplification module is further used to: for a set of annotation intentions belonging to the same seed annotation field, screen out a first seed annotation intention with similar or opposite intentions from the seed annotation intentions of the annotation intention set, generate an intention pair corresponding to the first seed annotation intention, the intention pair including the first seed annotation intention and a second seed annotation intention that is similar or opposite to the first seed annotation intention, the intention pair being represented as , Label the first seed with intent, is the second seed marking intention, n is a positive integer; a third extended prompt sentence is determined from the preset extended prompt sentence, and the third role description information in the third extended prompt sentence is used to set the function of the initial large language model to generate a new corpus according to the seed corpus and the intention pair, and the target intention of the new corpus is similar to the second seed marking intention; the third extended prompt sentence is used to instruct the initial large language model to target Seed corpus Perform data augmentation operations with the opposite intention to obtain an augmented training dataset.
[0251] In an exemplary embodiment, the labeled dataset is represented as , is the seed corpus to be amplified, Label the field for the seed, The intention of the seed is labeled and the training dataset is expanded as , For the stated purpose, , i, M and N are positive integers.
[0252] In an exemplary embodiment, the output module is further configured to: input the data of the augmented training data set into the analytical model so that the analytical model performs classification processing on the augmented training data and outputs the predicted domain corresponding to the augmented training data. and predictive intent ;according to and The positive sample data set and the negative sample data set are determined by the verification result, and the verification result includes the seed annotation field and the predicted areas First verification results and seed annotation intentions and the predicted intent The second verification result.
[0253] In an exemplary embodiment, the output module is further configured to: upon determining that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention, Figure 1 If the seed corpus is consistent, the seed corpus corresponding to the seed annotation field is stored in the positive sample data set according to the first mapping relationship. The first mapping relationship represents the corpus quadruple of the seed corpus to be amplified in the annotated data. The relationship between the key and the positive sample data set as the value, is the seed corpus to be amplified, Label the field for the seed, labeling the seed with intent, For the purpose.
[0254] In an exemplary embodiment, the output module is further configured to: determine that the first verification result indicates that the seed annotation domain is consistent with the predicted domain, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention. Figure 1 If any of the above is not true, the seed corpus corresponding to the seed annotation field is stored in the negative sample dataset according to the second mapping relationship. , the second mapping relationship is represented by A relationship where the key is the negative sample dataset and the value is the negative sample dataset.
[0255] In an exemplary embodiment, the output module is further used to: input the augmented training data set into the parsing model so that the parsing model recognizes the augmented training data and obtains the slot relationship corresponding to the augmented training data; replace the slot relationship with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data to the same sentence structure as the seed corpus to be augmented, determine the updated augmented training data as positive sample candidate data, and determine the positive sample data set based on the positive sample candidate data.
[0256] In an exemplary embodiment, the output module is further configured to: , collect k positive samples from the positive sample data set according to the preset sample ratio of 1:k , , ; Using the prompt word, the negative sample and the k positive samples Generate a preference optimization sample, and determine the preference optimization data set based on the sample set of all preference optimization samples, the prompt word is composed of corpus quadruple And preset introduction generation.
[0257] In an exemplary embodiment, the output module is further configured to: generate the vector encoding result of the negative sample and k positive samples according to the vector encoding result of the negative sample and k positive samples. The vector encoding result is used to calculate the negative sample and k positive samples The cosine similarity between the two datasets is calculated, and the similarity probability distribution result of the negative sample with respect to the positive sample dataset is calculated based on the cosine similarity; the positive sample dataset is sampled using the similarity probability distribution result to obtain the negative sample dataset. The corresponding k positive samples .
[0258] In an exemplary embodiment, the output module is further configured to calculate negative samples using the following formula: and k positive samples The cosine similarity between:
[0259] ;
[0260] is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
[0261] In an exemplary embodiment, the output module is further configured to: calculate a similarity probability distribution result of the negative sample with respect to the positive sample dataset using a similarity probability distribution function of the negative sample with respect to the positive sample dataset; the similarity probability distribution function is expressed as follows:
[0262] ;
[0263] constant ; The similarity probability distribution result is expressed as .
[0264] In an exemplary embodiment, the output module is further configured to represent a preference optimization sample in the form of a triple, and the preference optimization dataset is represented as follows:
[0265] , Indicates a prompt word.
[0266] In an exemplary embodiment, before generating a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, the second amplification module is also used to: determine a sensitivity quantization index of the parsing model based on the perplexity parameter; use the sensitivity quantization index to normalize a data subset in the preference optimization data set, and determine the result of the normalization as the sample weight, wherein the preference optimization samples in the data subset contain prompts with the same intention or prompts with opposite intentions.
[0267] In an exemplary embodiment, when the perplexity parameter is expressed as PPL(S), the sensitivity quantization index is expressed as PPL diff , , S represents the set input sequence of the parsing model, x represents the input prompt word used to prompt the generation of similar intentions, represents the prompt word input for generating the opposite intention, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the response sequence y generated by the analytical model when x is the set input sequence, Indicates that when When the set input sequence is given, the parsing model generates the perplexity of the response sequence y.
[0268] In an exemplary embodiment, the second amplification module is further configured to: obtain a target quantitative index generated according to the sensitivity quantitative index , the target quantitative index represents , Prompts for the same intention Or a hint of the opposite intention The sum of the sensitivity quantification indicators of About Data Subsets The normalized result is determined as the sample weight, and the normalized result is expressed as follows: .
[0269] In an exemplary embodiment, the second augmentation module is also used to iteratively train the strategy model with the minimization of the weighted loss function as the training goal. When the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is the smallest, the training of the strategy model is stopped, and the strategy model is determined as the data augmentation model. The loss value of the weighted loss function represents the degree of deviation between the strategy model and the reference model.
[0270] An embodiment of the present application further provides a storage medium, which includes a stored program, and when the program is run, any of the above methods is executed.
[0271] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:
[0272] S1, performing a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set, wherein the labeled data is collected by a sound pickup device;
[0273] S2, using a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and using the positive sample dataset and the negative sample dataset to generate a preference-optimized dataset, wherein a preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model obtained by performing recognition training on the initial large language model;
[0274] S3, generating a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, obtaining a data augmentation model that uses the weighted loss function to complete the preset training target, and using the data augmentation model to perform data augmentation on the newly input text data to obtain a target augmented data set.
[0275] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0276] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0277] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0278] S1, performing a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set, wherein the labeled data is collected by a sound pickup device;
[0279] S2, using a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and using the positive sample dataset and the negative sample dataset to generate a preference-optimized dataset, wherein a preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the labeled data, negative sample data, and positive sample data, and each negative sample corresponds to multiple positive samples, and the parsing model is a model obtained by performing recognition training on the initial large language model;
[0280] S3, generating a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, obtaining a data augmentation model that uses the weighted loss function to complete the preset training target, and using the data augmentation model to perform data augmentation on the newly input text data to obtain a target augmented data set.
[0281] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.
[0282] Optionally, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps in any method embodiment are implemented.
[0283] Optionally, an embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any method embodiment are implemented.
[0284] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0285] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0286] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.< / instruction> < / instruction>
Claims
1. A speech data processing method based on a large model, characterized in that: include: Performing a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set, wherein the labeled data is collected by a sound pickup device; Using a parsing model, the augmented training dataset is classified into a positive sample dataset and a negative sample dataset, and a preference-optimized dataset is generated using the positive sample dataset and the negative sample dataset, wherein a preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the labeled data, a negative sample, and a positive sample, and each negative sample corresponds to multiple positive samples, the positive sample represents text that is expected to be generated by the parsing model, and the negative sample represents a sample that is not expected to be generated by the parsing model, and the parsing model is a model obtained by performing recognition training on an initial large language model; A weighted loss function is generated according to the sample weights of the preference optimization samples and the original loss function of the parsing model, a data augmentation model is obtained for completing the preset training target using the weighted loss function, and the data augmentation model is used to perform data augmentation on the newly input text data to obtain a target augmented data set.
2. The method for processing speech data based on a large model according to claim 1, characterized in that: Before using the parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, the method further includes: Generate current labeled data based on the target object's current labeling operation on the initial data, and generate the labeled dataset based on the historical labeled data provided by the target object and the current labeled data, wherein the labeled dataset includes at least the target object's question sentence, the target object's labeled domain, the target object's labeled intent, and slots corresponding to the target object's labeled entities; The fine-tuning instruction of the target object is used to instruct the initial large language model to perform a recognition operation on the labeled data, and output a recognition result according to the output requirements of the fine-tuning data instructed in the fine-tuning instruction, wherein the recognition result includes at least: a labeled field, a labeled intention, a labeled entity, and a labeled slot.
3. The method for processing speech data based on a large model according to claim 1, characterized in that: Perform data augmentation on the labeled data of the labeled dataset to obtain an augmented training dataset, including: Selecting seed corpus data to be amplified from the labeled data, and determining a seed annotation field, a seed annotation intent, and a target intent corresponding to the seed corpus according to the seed corpus data; A preset extended prompt statement is used to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention within the seed annotation domain, so as to obtain an augmented training data set. The preset extended prompt statement includes at least task information required for the initial large language model to perform the data augmentation task. The task information includes at least task description information, role description information, task examples, and the input seed corpus. The role description information is used to set the function of the initial large language model to generate a new corpus based on the seed corpus and the seed annotation intention, and the new corpus has the target intention.
4. The method for processing speech data based on a large model according to claim 3, characterized in that: Using a preset extended prompt sentence to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention in the seed annotation domain, to obtain an augmented training data set, including: For a set of annotation intentions belonging to the same seed annotation field, a first seed annotation intention with similar or opposite intentions is screened from the seed annotation intentions of the annotation intention set, and an intention pair corresponding to the first seed annotation intention is generated. The intention pair includes the first seed annotation intention and a second seed annotation intention that is similar or opposite to the first seed annotation intention. The intention pair is represented as , Label the first seed with intent, Label the intention for the second seed, where n is a positive integer; Determining a first extended prompt sentence from the preset extended prompt sentences, wherein the first role description information in the first extended prompt sentence is used to set a function of the initial large language model to generate a new corpus based on the seed corpus and the intent pair, wherein the target intent of the new corpus is similar to the first seed labeled intent; The first extended prompt sentence is used to instruct the initial large language model to Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
5. The method for processing speech data based on a large model according to claim 4, characterized in that: Using a preset extended prompt sentence to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention in the seed annotation domain, to obtain an augmented training data set, including: Filtering other labeling intentions except the first seed labeling intention from the labeling intention set, wherein the other labeling intentions only have similar intentions; Determining a second extended prompt sentence from the preset extended prompt sentence, wherein the second role description information in the second extended prompt sentence is used to set a function of the initial large language model to generate a new corpus based on the seed corpus and the other annotation intents, wherein the target intent of the new corpus is similar to the other annotation intents; The second extended prompt sentence is used to instruct the initial large language model to Seed corpus Perform data augmentation operations with similar intentions to obtain an augmented training dataset.
6. The method for processing speech data based on a large model according to claim 3, characterized in that: Using a preset extended prompt sentence to instruct the initial large language model to perform a data augmentation task on the seed corpus according to the seed annotation intention to the target intention in the seed annotation domain, to obtain an augmented training data set, including: For a set of annotation intentions belonging to the same seed annotation field, a first seed annotation intention with similar or opposite intentions is screened from the seed annotation intentions of the annotation intention set, and an intention pair corresponding to the first seed annotation intention is generated. The intention pair includes the first seed annotation intention and a second seed annotation intention that is similar or opposite to the first seed annotation intention. The intention pair is represented as , Label the first seed with intent, Label the intention for the second seed, where n is a positive integer; Determining a third extended prompt sentence from the preset extended prompt sentence, wherein the third role description information in the third extended prompt sentence is used to set a function of the initial large language model to generate a new corpus based on the seed corpus and the intent pair, wherein the target intent of the new corpus is similar to the second seed labeled intent; The third extended prompt sentence is used to instruct the initial large language model to Seed corpus Perform data augmentation operations with the opposite intention to obtain an augmented training dataset.
7. The method for processing speech data based on a large model according to claim 3, characterized in that: The labeled dataset is represented as , is the seed corpus to be amplified, Label the field for the seed, The intention of the seed is labeled and the training dataset is expanded as , For the stated purpose, , For the expanded text, is the expanded text set, i, M and N are positive integers.
8. The method for processing speech data based on a large model according to claim 1, characterized in that: Using a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset includes: The data of the augmented training data set is input into the analytical model so that the analytical model classifies the augmented training data and outputs the prediction domain corresponding to the augmented training data. and predictive intent ; according to and The positive sample data set and the negative sample data set are determined by the verification result, and the verification result includes the seed annotation field and the predicted areas First verification results and seed annotation intentions and the predicted intent The second verification result.
9. The method for processing speech data based on a large model according to claim 8, characterized in that: According to and The positive sample data set is determined by the verification result, including: When it is determined that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention, the seed corpus corresponding to the seed annotation field is stored in the positive sample dataset according to the first mapping relationship. The first mapping relationship represents the corpus quadruple of the seed corpus to be amplified in the annotated data. The relationship between the key and the positive sample data set as the value, is the seed corpus to be amplified, Label the field for the seed, labeling the seed with intent, For the purpose.
10. The method for processing speech data based on a large model according to claim 8, characterized in that: According to and The negative sample dataset is determined by the verification result, including: When it is determined that the first verification result indicates that the seed annotation field is consistent with the predicted field, and the second verification result indicates that the seed annotation intention is consistent with the predicted intention, the seed corpus corresponding to the seed annotation field is stored in the negative sample dataset according to the second mapping relationship. , the second mapping relationship is represented by A relationship where the key is the negative sample dataset and the value is the negative sample dataset.
11. The method for processing speech data based on a large model according to claim 8, characterized in that: The method further comprises: Inputting the augmented training data set into the analytical model so that the analytical model recognizes the augmented training data and obtains the slot relationship corresponding to the augmented training data; The slot relationship is replaced with the seed slot relationship corresponding to the seed corpus to be augmented to update the sentence structure of the augmented training data to the same sentence structure as the seed corpus to be augmented, the updated augmented training data is determined as the positive sample candidate data, and the positive sample data set is determined based on the positive sample candidate data.
12. The method for processing speech data based on a large model according to claim 1, wherein: Generating a preference optimization dataset using the positive sample dataset and the negative sample dataset includes: For negative samples ,according to Collect k positive samples from the positive sample data set with a preset sample ratio , , ; Using the prompt word, the negative sample and the k positive samples Generate a preference optimization sample, and determine the preference optimization data set based on the sample set of all preference optimization samples, the prompt word is composed of corpus quadruple And preset introduction generation.
13. The method for processing speech data based on a large model according to claim 12, characterized in that: according to Collect k positive samples from the positive sample data set with a preset sample ratio ,include: According to the vector encoding results of negative samples and k positive samples The vector encoding result is used to calculate the negative sample and k positive samples and calculating the similarity probability distribution result of the negative sample with respect to the positive sample data set based on the cosine similarity; The positive sample data set is sampled using the similarity probability distribution result to obtain negative samples The corresponding k positive samples .
14. The method for processing speech data based on a large model according to claim 13, characterized in that: The negative samples are calculated by the following formula and k positive samples The cosine similarity between: ; is the vector encoding result of the negative sample, is the vector encoding result of the positive sample.
15. The method for processing speech data based on a large model according to claim 13, characterized in that: Calculating a similarity probability distribution result of the negative sample with respect to the positive sample data set based on cosine similarity includes: Calculate the similarity probability distribution result of the negative sample with respect to the positive sample data set using the similarity probability distribution function of the negative sample with respect to the positive sample data set; The similarity probability distribution function is expressed as follows: ; constant ; The similarity probability distribution result is expressed as .
16. The method for processing speech data based on a large model according to claim 12, wherein: The method further comprises: A preference optimization sample is represented in the form of a triple, and the preference optimization dataset is represented as follows: , Indicates a prompt word.
17. The method for processing speech data based on a large model according to claim 1, wherein: Before generating a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the analytical model, the method further includes: Determining a sensitivity quantification index of the analytical model according to the perplexity parameter; The data subset in the preference optimization data set is normalized using the sensitivity quantification index, and the normalization result is determined as the sample weight, and the preference optimization samples in the data subset include prompts with the same intention or prompts with opposite intentions.
18. The method for processing speech data based on a large model according to claim 17, characterized in that: When the perplexity parameter is expressed as PPL(S), the sensitivity quantification index is expressed as , , S represents the set input sequence of the parsing model, x represents the input prompt word used to prompt the generation of similar intentions, represents the prompt word input for generating the opposite intention, y represents the response sequence generated by the parsing model based on the set input sequence, represents the perplexity of the response sequence y generated by the analytical model when x is the set input sequence, Indicates that when When the set input sequence is given, the parsing model generates the perplexity of the response sequence y.
19. The method for processing speech data based on a large model according to claim 17, wherein: Normalizing a data subset in the preference optimization data set using the sensitivity quantification index, and determining the normalization result as the sample weight, including: Obtain the target quantitative index generated according to the sensitivity quantitative index , the target quantitative index represents , Prompts for the same intention Or a hint of the opposite intention The sum of the sensitivity quantification indicators; Will About Data Subsets The normalized result is determined as the sample weight, and the normalized result is expressed as follows: 。 20. The method for processing speech data based on a large model according to claim 1, characterized in that: Obtaining a data augmentation model that uses the weighted loss function to achieve a training objective, including: The strategy model is iteratively trained with the minimization of the weighted loss function as the training goal. When the preset number of iterative training times is reached and the function value of the weighted loss function in the current iteration round is minimum, the training of the strategy model is stopped, and the strategy model is determined as the data augmentation model. The loss value of the weighted loss function represents the degree of deviation between the strategy model and the reference model.
21. A speech data processing device based on a large model, characterized in that: include: A first augmentation module is used to perform a data augmentation operation on the labeled data of the labeled data set to obtain an augmented training data set; an output module, configured to use a parsing model to classify the augmented training dataset into a positive sample dataset and a negative sample dataset, and generate a preference-optimized dataset using the positive sample dataset and the negative sample dataset, wherein a preference-optimized sample in the preference-optimized dataset includes a prompt word generated based on the annotated data, a negative sample, and a positive sample, and each negative sample corresponds to multiple positive samples, the positive sample represents text that is expected to be generated by the parsing model, and the negative sample represents a sample that is not expected to be generated by the parsing model, and the parsing model is a model obtained by performing recognition training on an initial large language model; The second amplification module is used to generate a weighted loss function based on the sample weights of the preference optimization samples and the original loss function of the parsing model, obtain a data amplification model that uses the weighted loss function to complete the preset training objectives, and use the data amplification model to amplify the newly input text data to obtain a target amplified data set.
22. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, and when the program is executed, the method according to any one of claims 1 to 20 is executed.
23. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 20 through the computer program.
24. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 20 is implemented.
Citation Information
Patent Citations
Voice data amplification method and system
CN108922518A
Voice augmentation method, related method, device, equipment and storage medium
CN118136034A