Data cleaning method and device based on self-supervised learning, medium and product
Through the data cleaning method of self-supervised learning, the large language model is trained using unlabeled data, which solves the problems of manual labeling dependence and insufficient adaptability in existing technologies, and realizes efficient and automated data cleaning and error recognition.
Patent Information
- Application Number
- CN202511159973.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-24
AI Technical Summary
Existing data cleaning technologies rely on manual labeling and intervention, consume a lot of manpower costs, and lack adaptive capabilities, resulting in low cleaning efficiency and limited adaptability.
A self-supervised learning method is adopted to pre-train the large language model using unlabeled historical conversation data. Through dynamic context perception, task stratification and context comparative learning, the trained large language model is obtained. Secondary training is performed using conversation data containing erroneous information to generate a target large language model for automated cleaning.
It reduces manual labeling costs, improves data cleaning efficiency and adaptability, and can automatically identify and process various error messages to adapt to data cleaning needs in different scenarios.
Smart Images

Figure CN120832473A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and in particular to a data cleaning method and device based on self-supervised learning, a medium and a product. BACKGROUND
[0002] In data-driven applications, data quality directly determines system performance and decision accuracy, but the data collection process is often accompanied by noise, missing values, inconsistency and other problems, which adversely affect subsequent data analysis and model training. The existing data cleaning related technologies have the following obvious limitations: on the one hand, excessive reliance on manual annotation and intervention, many methods need to rely on manual adjustment and optimization of the model or rely on a large amount of labeled data, which not only consumes a lot of manpower and time cost, but also limits the application and expansion speed of the model in a specific field; on the other hand, the feedback optimization mechanism lacks self-adaptive ability, when the cleaning effect is not good, it is difficult to realize automatic feedback and self-adaptive adjustment, which restricts the efficiency and applicability of data cleaning. SUMMARY
[0003] At least one embodiment of the present application provides a data cleaning method, device, medium and product based on self-supervised learning, to solve the problems of low efficiency and limited adaptability of data cleaning in the prior art.
[0004] To solve the above technical problems, the present application is implemented as follows:
[0005] In a first aspect, the embodiments of the present application provide a data cleaning method based on self-supervised learning, comprising:
[0006] Based on a self-supervised learning training task, using unannotated historical dialogue data, a pre-set large language model is initially pre-trained to obtain a trained large language model;
[0007] Using the historical dialogue data, dialogue data containing error information is determined;
[0008] According to the dialogue data containing error information, the trained large language model is secondarily trained to obtain a target large language model after training; the target large language model is used for automatic cleaning processing of dialogue data.
[0009] Optionally, based on a self-supervised learning training task, using unannotated historical dialogue data, a pre-set large language model is initially pre-trained to obtain a trained large language model, comprising:
[0010] The self-supervised learning training task is determined; the self-supervised learning training task includes dynamic context awareness, task layering and context contrast learning;
[0011] According to the historical dialogue data, the preset large language model is pre-trained by dynamic context awareness, task layering and context contrast learning, and a trained large language model is obtained.
[0012] Optionally, the self-supervised learning training task includes pre-training the preset large language model by dynamic context awareness, task layering and context contrast learning in the case of dynamic context awareness, and obtaining a trained large language model.
[0013] At the current time point, the historical dialogue data is used to obtain current input information input into a preset computing model and a context state corresponding to the current input information.
[0014] According to the current input information and the context state, a predicted context probability of the next time point is obtained.
[0015] According to the context probability of the next time point, a context total loss value of the next time point is obtained.
[0016] According to the context total loss value of the next time point, the large language model is trained by dynamic context awareness self-supervised learning, and a large language model trained by dynamic context awareness self-supervised learning is obtained.
[0017] Optionally, the self-supervised learning training task includes task layering, and according to the historical dialogue data, the preset large language model is pre-trained by task layering, and a trained large language model is obtained, including:
[0018] The historical dialogue data is used to obtain a task level label used for training and an input feature corresponding to the task level label.
[0019] According to the task level label and the input feature, the input feature corresponding to the current input task level is determined.
[0020] According to the input feature corresponding to the current input task level, a loss function for task layering self-supervised learning is constructed.
[0021] Based on the loss function for task layering self-supervised learning, the large language model is trained by task layering self-supervised learning, and a large language model trained by task layering self-supervised learning is obtained.
[0022] Optionally, the self-supervised learning training task includes context contrast learning, and according to the historical dialogue data, the preset large language model is pre-trained by context contrast learning, and a trained large language model is obtained, including:
[0023] The historical dialogue data is used to obtain a first similarity between an original context for training and a correct subsequent dialogue, and a second similarity between the original context and a noise-generated invalid dialogue;
[0024] According to the first similarity and the second similarity, a loss function of self-supervised learning for context contrast learning is constructed;
[0025] Based on the loss function of self-supervised learning for context contrast learning, the large language model is trained by self-supervised learning for context contrast learning, to obtain the large language model after the self-supervised learning for context contrast learning.
[0026] Optionally, according to the dialogue data containing error information, the trained large language model is further trained to obtain a target large language model after training, including:
[0027] A target self-supervised task for further training is constructed; the target self-supervised task includes a noise contrast learning task, a context consistency detection task, and a self-supervised mask language model correction task;
[0028] A loss function of the target self-supervised task is constructed;
[0029] According to the loss function of the target self-supervised task and the dialogue data containing error information, the trained large language model is further trained to obtain a target large language model after training.
[0030] Optionally, according to the loss function of the target self-supervised task and the dialogue data containing error information, the trained large language model is further trained to obtain a target large language model after training, including:
[0031] According to the loss function of the target self-supervised task and the dialogue data containing error information, a total loss containing context consistency detection and noise contrast learning is determined;
[0032] According to the total loss, it is determined whether there is a target sentence with semantic or grammatical error;
[0033] If there is a target sentence with semantic or grammatical error, the mask language model is used to mask and predict the correct word or phrase of the target sentence, and the error-corrected dialogue data is re-input into the large language model for verification until the dialogue data passes the verification of all detection tasks of the large language model.
[0034] In a second aspect, the embodiments of the present application provide a data cleaning device based on self-supervised learning, including:
[0035] The first processing module is used to perform initial pre-training on a preset large language model based on a self-supervised learning training task using unlabeled historical conversation data to obtain a trained large language model;
[0036] A first determining module is used to determine the conversation data containing error information by using the historical conversation data;
[0037] The second processing module is used to perform secondary training on the trained large language model based on the dialogue data containing error information to obtain a trained target large language model; the target large language model is used for automatic cleaning of the dialogue data.
[0038] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0039] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method described in any one of the first aspects.
[0040] Compared with the prior art, the data cleaning method, device, medium and product based on self-supervised learning provided by the embodiments of the present application use unlabeled data for initial pre-training, without the need for manual labeling, thereby reducing the cost of initial data preparation; the target large language model obtained after the secondary training can automatically process conversation data cleaning, replacing traditional manual or semi-automatic methods, greatly reducing manpower input, and improving cleaning efficiency; the initial pre-training allows the model to enable the trained large language model to learn general language rules, and the secondary training of the trained large language model focuses on data containing erroneous information to obtain the trained target large language model, so that the target model can grasp various error features in a targeted manner. Compared with traditional fixed rules or templates, the target large language model of the present application can adapt to the error patterns of different scenarios through learning. When facing new types of errors, it only needs to update the training data to adjust, and it is more adaptable. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0042] Figure 1 A flowchart of a data cleaning method based on self-supervised learning provided in an embodiment of the present application;
[0043] Figure 2A structural schematic diagram of a data cleaning device based on self-supervised learning is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0044] The terms "first", "second", and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second" are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, scenario one: including A and not including B; scenario two: including B and not including A; scenario three: including A and B. The character " / " generally represents that the objects before and after are in an "or" relationship.
[0045] The term "indication" in the present application can be a direct indication (or explicit indication) or an indirect indication (or implicit indication). Among them, the direct indication can be understood as that the sender explicitly informs the receiver of specific information, operations to be performed or requested results, etc. in the sent indication; the indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or judges and determines the operation to be performed or the requested result according to the judgment result.
[0046] As described in the background, the data cleaning method in the prior art is one of over-reliance on manual annotation and intervention, and the current many data cleaning methods rely on manual intervention or a large number of labeled data to adjust and optimize the model. This mode not only consumes a lot of human cost and time, but also limits the application and expansion speed of the model in a specific field; the second is that the feedback optimization mechanism lacks self-adaptive ability, and the existing model data cleaning method lacks automatic feedback and self-adaptive adjustment mechanism when the cleaning effect is not ideal. The above two cases seriously affect the cleaning effect of data cleaning. To solve at least one of the above problems, the embodiments of the present application provide a data cleaning method, device, medium and product based on self-supervised learning, which can reduce or avoid the occurrence of the above situation and improve the adaptability and cleaning efficiency of data cleaning.
[0047] Please refer to Figure 1 The embodiments of the present application provide a data cleaning method based on self-supervised learning, which comprises:
[0048] Step 11, based on self-supervised learning training task, using unlabeled historical dialogue data, pre-training the preset large language model for the first time, and obtaining the trained large language model.
[0049] In the embodiments of the present application, unannotated historical dialogue data such as customer service dialogue and user chat records are used without manual annotation of right or wrong or classification. Through a self-supervised learning task, a pre-set large language model is initially trained to obtain a trained large language model, and the model learns the general rules of dialogue (such as language structure, semantic association, and common expression habits) from a large amount of unannotated data to have a basic dialogue understanding ability, laying a foundation for subsequent targeted processing of error information. Without manual annotation of data, the early-stage cost is reduced and the training efficiency is improved.
[0050] Here, the pre-set large language model can be a large language model commonly used in a customer service system, such as a large language model based on a Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pretrained Transformer).
[0051] Step 12, determining dialogue data containing error information by using the historical dialogue data.
[0052] In step 11 of the present application, a large amount of unannotated historical dialogue data is received, and the unannotated historical dialogue data contains interactive records of users and customer service. The input dialogue data is represented by a dialogue sequence D, specifically: D = {S1, S2, … S n}, where each S i represents a dialogue turn, and the data can contain some noise, unreasonable sentences or words. The historical dialogue data is used to determine dialogue data containing error information, and the error information can include at least one of the following error types:
[0053] Semantic error: the meaning of a sentence is not coherent or does not match the context.
[0054] Syntax error: the structure of a sentence does not conform to language specifications.
[0055] Task error: a reply that does not match the user's intention or task.
[0056] Context inconsistency error: information in a dialogue is contradictory, causing the answer to be disjointed from the context.
[0057] Here, the pre-set error type rule can be used for preliminary screening to determine dialogue data containing error information; or the large language model after initial training can be used for auxiliary identification, such as using the model obtained in step 11 to preliminarily judge the dialogue, marking the samples predicted by the model as possible errors, and then verifying and confirming a small amount of manual verification to reduce the workload of pure manual screening.
[0058] The application screens the dialogue data containing error information from the historical dialogue data used in step 11 as error samples for subsequent secondary training, focuses on the error samples, provides targeted data for secondary training, and enables the model to accurately identify such errors subsequently.
[0059] In step 13, the dialogue data containing error information is used to perform secondary training on the trained large language model to obtain a target large language model after training; and the target large language model is used for automatic cleaning of dialogue data.
[0060] It should be noted that since the model learns general rules from massive data and is trained for error samples, it can adapt to various error types (without the need to design rules for each scenario) when facing dialogue data in different scenarios (such as e-commerce customer service and medical consultation), thereby solving the problem of limited adaptability of traditional methods. The target large language model is applied to automatically clean dialogue data, and the trained target large language model can be directly used for automatic cleaning of dialogue data.
[0061] In the embodiment of the application, the dialogue data containing error information determined in step 12 is used as a training sample to perform secondary training on the large language model obtained in step 11, and finally a target large language model is obtained, so that the target large language model learns the characteristics of error information such as error types and common error patterns on the basis of general dialogue understanding ability, and has the ability to accurately identify and process error information in dialogue.
[0062] The method of the application uses the process of unlabeled data pre-training, error sample screening, and targeted secondary training, so that the model can efficiently utilize massive data and reduce manual work, accurately process error information, and finally realize automatic cleaning of dialogue data, greatly improving efficiency and enhancing adaptability to different scenarios.
[0063] Optionally, the task design of secondary training can include:
[0064] Design an error identification task: let the model classify the input dialogue segment and determine whether it contains error information;
[0065] Design an error correction task: let the model output the corrected correct content for the dialogue containing errors.
[0066] In the secondary training process, the loss of the predicted result of the preset calculation model and the real error label (or correct content) is used to continuously adjust the model parameters and strengthen the sensitivity and processing ability of the model to error information.
[0067] Optionally, step 11 described above includes:
[0068] determine the self-supervised learning training task; the self-supervised learning training task comprises dynamic context perception, task layering, and context contrast learning;
[0069] According to the historical dialogue data, the preset large language model is subjected to initial pre-training of dynamic context perception, task layering, and context contrast learning, and a trained large language model is obtained.
[0070] In the embodiment of the application, this step is achieved by designing three targeted self-supervised learning subtasks, so that the large language model learns the dynamic rules, hierarchical structure, and semantic association of the dialogue from the unannotated historical dialogue data, laying a foundation for subsequent processing of the dialogue data. The self-supervised learning training task determined in the application can include three subtasks, which respectively strengthen the model's understanding ability of the dialogue data from different dimensions; wherein the dynamic context perception can focus on the features of the dynamic change of the context in the dialogue over time or turns, such as the anaphora relationship, topic shift, semantic progression, etc.; the task layering can decompose the dialogue understanding into subtasks at different levels, such as basic semantic recognition, logical relationship judgment, etc., so that the model masters the dialogue rules layer by layer; the context contrast learning can strengthen the model's ability to distinguish the semantic similarity and difference by comparing similar or greatly different dialogue segments.
[0071] The historical dialogue data is used to train the preset large language model, and the three subtasks are trained, such as dynamic context perception training, task layering training, and context contrast learning training. Through the cooperative training of the above three subtasks, the preset large language model autonomously learns from the unannotated historical dialogue data the dynamic context association rules of the dialogue, the hierarchical understanding ability of the dialogue information, and the ability to distinguish the semantic similarity and difference.
[0072] At this time, the trained large language model has a strong general dialogue understanding foundation, providing a reliable model base for subsequent secondary training for error information processing.
[0073] It should be noted that the historical dialogue data of the customer service system is used to pre-train the large language model, and the trained large language model is obtained, which can specifically include: obtaining input data. The input data is the historical dialogue data in the customer service system. Assuming that the dialogue sequence is D = {S1, S2, S3, …, S n} where each S i represents a dialogue turn between the customer service and the user.
[0074] Pre-training task design: In the self-supervised learning of intelligent customer service, the following pre-training tasks are usually selected: Masked Language Model (MLM): randomly select some words in the input to be masked, and let the model predict these masked words. The loss function is:
[0075]
[0076] wherein, L MLM represents the loss function of the masked language model, measuring the accuracy of the model in correcting the masked words in the sentence; M represents the set of masked words, indicating the words that need to be corrected in the dialogue sentence; w i represents the i-th masked word, and the target is that the model can correctly predict this word; X M represents the context of the sentence after masking the words M; P(w i | X M ) represents the probability of the model predicting the word w i given the masked context.
[0077] Next Sentence Prediction (NSP): the model inputs two sentences in a dialogue and predicts whether the second sentence is a reasonable subsequent sentence of the first sentence. The loss function is:
[0078]
[0079] wherein, L nsp represents the total loss value of the next sentence prediction, used to measure the prediction error of the model for the coherence of each pair of sentences, and the smaller the value, the more accurate the prediction of the model; N represents the number of samples, i.e. the number of dialogue sequences or sentence pairs, used to calculate the prediction loss of each sentence pair; P(S i+1 | S i ) represents the probability of the model predicting the next sentence given the previous sentence.
[0080] Optionally, the self-supervised learning training task includes initial pre-training of the preset large language model in a dynamic context-aware manner, task layering, and context contrast learning, to obtain the trained large language model, including:
[0081] At the current time point, the historical dialogue data is used to obtain the current input information input into the preset calculation model and the context state corresponding to the current input information;
[0082] According to the current input information and the context state, the predicted context probability of the next time point is obtained;
[0083] According to the context probability of the next time point, the context total loss value of the next time point is obtained;
[0084] According to the context total loss value of the next time point, the large language model is subjected to self-supervised learning training in a dynamic context-aware manner, to obtain the large language model subjected to self-supervised learning training in a dynamic context-aware manner.
[0085] In the embodiments of the present application, at the current time point, such as the tth round of dialogue, the current input information is extracted from the historical dialogue data, denoted as X t , the current input information X t , that is, the dialogue content of the tth round; the current context state is denoted as C t , the context state C t may represent the overall dialogue context up to the tth round, including the topic, user intent, semantic association and other comprehensive information of the historical round, rather than isolated word splicing. Through the current input information and the context state, the model provides the current dialogue content and the existing context background as the basis for predicting the next round of context changes.
[0086] The initial model predicts the context state at the next time point (t+1th round) based on the current input information and the context state, which can be denoted as C t+1 , C t+1 is used to represent that the model needs to predict the context state at the next time point; for example, in a consultation scenario, if the current input information X t is “what to do if the product is broken”, C t shows that the user previously consulted “purchase process”, then the model needs to predict that the next round of context state C t+1 is more likely to be “after-sales service” rather than continuing “purchase process”, and outputs the probability of this state. Different from the traditional MLM, which only predicts the masked words, here the prediction is the “evolution of the overall context state”, which is directly related to the user intent and topic shift, and is more suitable for the dynamics of multi-round dialogue.
[0087] According to the predicted probability P(C t+1 |C t , X t ), the total loss value is calculated through a preset loss function, and if the prediction is wrong, the P value is low, and the loss will increase. Therefore, the loss function essentially punishes incorrect predictions and rewards correct predictions, forcing the model to learn the real evolution law of the context state.
[0088] According to the context loss value at the next time point, the large language model is subjected to dynamic context-aware self-supervised learning training, and a large language model subjected to dynamic context-aware self-supervised learning training is obtained. The training process captures the self-supervised task of dynamic evolution of the context state, rather than simply predicting words, and combines a targeted loss function to make the model truly understand the logical flow and dynamic dependency of multi-round dialogue, solving the problem that the traditional method cannot handle “intent transfer and topic switching”. It provides key context understanding capabilities for subsequent identification of incorrect information in the dialogue, such as the customer service answering the wrong position after topic transfer, and realizes automatic data cleaning.
[0089] Specifically, conversations in intelligent customer service often have strong contextual dependencies, especially in multi-turn conversations. Although the traditional masked language model (MLM) can capture the relationship at the lexical level, it cannot well capture the dynamic context of multi-turn conversations. Therefore, a dynamic context-aware task model needs to not only predict the masked words, but also infer changes in customer intent and topic shifts in the conversation sequence. This application designs a specific loss function for the self-supervised task as:
[0090]
[0091] Among them, C t represents the context state of the conversation at time t; X t is the current dialogue input; the model needs to predict the context state C at the next time point t+1 ;L DC Represents the total context loss value at the next time point; P(C t+1 |C t , X t ) represents the probability of predicting the context at the next time point given the current dialogue input.
[0092] Optionally, when the self-supervised learning training task includes task stratification, performing initial task stratification pre-training on a preset large language model based on the historical conversation data to obtain the trained large language model includes:
[0093] Using the historical conversation data, obtaining task-level labels for training and input features corresponding to the task-level labels;
[0094] Determining input features corresponding to the current input task level according to the task level label and the input features;
[0095] Constructing a loss function for self-supervised learning of task hierarchies based on input features corresponding to the current input task hierarchy;
[0096] Based on the loss function of the task-layered self-supervised learning, the large language model is trained by task-layered self-supervised learning to obtain a large language model after the task-layered self-supervised learning training.
[0097] In the embodiments of the present application, the task level labels and corresponding input features are automatically extracted from historical dialogue data. The task level labels are used to divide the task into multiple levels according to the complexity of dialogue understanding, and each level corresponds to a label. For example, it is divided into basic level, intermediate level and high level. The basic level label includes entity recognition and keyword extraction. Entity recognition is used to identify specific information such as product name, time and amount in the dialogue. Keyword extraction is used to extract core words such as refund and delivery. The intermediate level label can perform intent judgment and logical relationship identification. The intent judgment can be used to judge whether the user is "consulting", "complaining" or "suggesting" logic. The high level label can perform topic classification, such as judging whether the dialogue belongs to logistics problem, after-sales problem or product function. The input features are used to represent the corresponding dialogue data features matched for each level task, such as word sequence and part-of-speech tagging of dialogue text, semantic vector of single-turn dialogue, interaction relationship between turns, overall semantic vector of multi-turn dialogue, topic switching node identification, etc.
[0098] For each task level, the features specific to this level are selected from all input features, and the interference features of other levels are excluded. For the target of each task level, a corresponding sub-loss function is designed, and then a total loss function is formed by weighted combination. Through the hierarchical loss function, the learning error of the model at each level is quantified, and the model can be effectively optimized at each link such as basic information recognition, intermediate logical understanding and high-level topic grasping. Based on the loss function for task layering training, the pre-set large language model can be iteratively trained to minimize the total loss function of task layering, and the model parameters are continuously adjusted through back propagation. During the training process, the model first masters the basic ability of entity recognition and keyword extraction, and then learns to judge user intent and identify logical relationships based on the basic ability. Finally, it learns to summarize the dialogue topic and analyze the overall sentiment at the high level, forming progressive ability from local details to global understanding.
[0099] Specifically, there are many different tasks in the intelligent customer service system, such as question answering, order processing, after-sales service, etc. During the task layering self-supervised learning training process of the large language model, the self-supervised learning model can automatically detect the features of different tasks and learn them specifically. By dividing different tasks into different levels and designing self-supervised tasks at different levels, the model can learn the specific features of each type of task. The specific task layering loss function is:
[0100]
[0101] Where, T k is the task level label; X k is the input feature corresponding to the task level. In this way, the model can distinguish between different types of tasks in the customer service system during the training process and optimize them specifically; Ltask represents the loss function of task hierarchy based self-supervised learning; P(T k |X k ) represents the input feature corresponding to the given input task hierarchy, and the probability of the model task hierarchy forming a label.
[0102] The task hierarchy self-supervised training defines the task by layers, matches the exclusive features, and designs the hierarchical loss, so that the model understands the multi-dimensional information of the dialogue layer by layer from simple to complex, overcomes the limitation of the traditional model which treats all information in the same way, and makes it more accurate to handle the details and overall logic in the dialogue, which facilitates the subsequent automatic cleaning of the dialogue data.
[0103] Optionally, in the case of context contrastive learning, the self-supervised learning training task includes initial pre-training of a preset large language model according to the historical dialogue data, and obtaining the trained large language model, including:
[0104] Using the historical dialogue data, a first similarity between an original context used for training and a correct subsequent dialogue is obtained, and a second similarity between the original context and a disturbance generated invalid dialogue is obtained;
[0105] According to the first similarity and the second similarity, a loss function of self-supervised learning for context contrastive learning is constructed;
[0106] Based on the loss function of self-supervised learning for context contrastive learning, the self-supervised learning training of the large language model is performed, and the large language model after the self-supervised learning training of the context contrastive learning is obtained.
[0107] In the embodiments of the present application, the historical dialogue data often contains a large amount of noise and invalid information. In order to better learn valuable dialogue information, the context contrastive learning (CCL) task generates a disturbed invalid dialogue for each dialogue sequence. The original context C is extracted from the historical dialogue data, such as the user's multi-round questions, and two types of subsequent dialogues are generated: the correct subsequent dialogue C + , which is an effective dialogue consistent with the logic of the original context and semantically related; and the disturbance generated invalid dialogue C - , which is generated by disturbing the correct subsequent. The similarity is calculated: the model calculates the first semantic similarity score of the original context C and the correct subsequent dialogue C + , denoted as The score should be higher in the ideal case, because the two are related; the model calculates the semantic similarity score of the original context C and the invalid dialogue C - , denoted as Ideally, the score should be low, as the two are unrelated. The model learns a more accurate semantic representation by distinguishing between valid and invalid conversations. The loss function of the context contrast learning is:
[0108]
[0109] where C is the original context, C + is the correct continuation of the context, C - is the perturbed invalid conversation, and the model needs to learn the correct context conversation sequence and distinguish the invalid context. L CCL represents the loss function of the context contrast learning, represents the similarity score of the conversation sequence C and the positive sample C + .
[0110] If the model can accurately identify the correct continuation, the numerator is much larger than the denominator, and the overall score is close to 1, and the loss value L CCL after log is small; if the model confuses the correct continuation with the invalid conversation, the denominator increases, and the overall score is close to 0, and the loss value L CCL after log increases. Therefore, the loss function rewards correct differentiation and punishes confusion, forcing the model to learn the semantic correlation rules of valid conversations and the noise characteristics of invalid conversations.
[0111] The context contrast learning of the present application constructs valid and invalid conversation pairs, calculates semantic similarity, and designs a contrast loss function process to let the model focus on valuable conversation information in noisy data and strengthen its ability to distinguish semantic relevance. This training enables the model to more accurately identify incorrect information caused by context mismatch when processing conversation data, such as customer service replies unrelated to user questions, providing key semantic filtering capabilities for automated conversation data cleaning.
[0112] In the present application, the input data obtained is used to train a large language model based on the above designed training task, to obtain a trained large language model, thereby improving the understanding ability of the large language model for conversation data of a customer service system.
[0113] Optionally, the trained large language model is further trained based on the conversation data containing incorrect information to obtain a trained target large language model, including:
[0114] Constructing target self-supervised tasks for secondary training; the target self-supervised tasks include: Noise Contrastive Learning (NCL), Contextual Consistency Detection (CCD), and a self-supervised Masked Language Model (MLM for Correction) task;
[0115] Constructing a loss function for the target self-supervised task;
[0116] The trained large language model is trained a second time based on the loss function of the target self-supervised task and the dialogue data containing error information to obtain a trained target large language model.
[0117] In this embodiment, a large language model undergoes secondary training using conversation data containing erroneous information. The overall process includes data input, noise contrastive learning, context consistency detection, data correction, and iterative optimization. A target self-supervised task is constructed for secondary training. The target self-supervised task includes noise contrastive learning, context consistency detection, and self-supervised masked language model correction.
[0118] Optionally, in constructing the loss function of the target self-supervised task, a noise contrast learning task is constructed, including: designing a model to learn to distinguish noisy data from normal data by comparing noisy conversations with clean conversations. For a given conversation sequence D, noise is randomly introduced:
[0119]
[0120] Among them, L NCL represents the loss function of noise contrast learning, which measures the performance of the model in distinguishing clean data from noisy data; D represents the original dialogue data sequence; D + represents a clean conversation data sequence, which is used as a positive sample; D - represents a sequence of conversation data with noise, which is used as a negative sample; f(D,D + ) represents the dialogue sequence D and the positive sample D + The similarity score of , usually uses the similarity measurement method between vectors, such as inner product or cosine similarity. f(D,D - ): dialogue sequence D and negative sample D - The similarity score of .
[0121] Optionally, the context consistency detection can be used to detect context errors in the dialogue, and the model learns to judge the consistency of the information before and after in the dialogue. In constructing the loss function of the target self-supervised task, the specific loss function of the context consistency detection is:
[0122] L CCD = -logP(C t+1 |C t , X t );
[0123] wherein L CCD represents the loss function of the context consistency detection, which measures whether the model's prediction of the context is consistent with the continuity of the dialogue; C t represents the state of the current dialogue context, which represents the dialogue content, customer intent and context features at time step t; C t+1 represents the context state of the next time step, which represents the model's prediction of the context of the next time step; X t represents a specific input sequence in the dialogue (the dialogue content at the current time step); P(C t+1 |C t , X t ) represents the probability of predicting the next context C t given the current context C t and the input X t+1 .
[0124] Optionally, the self-supervised mask language model correction task is to correct the detected grammatical or semantic errors, and a self-supervised mask language model MLM can be used. When the model detects an error in a sentence, it will automatically mask part of the sentence and let the model predict the correct word or phrase through the context. In constructing the loss function of the target self-supervised task, the loss function of the self-supervised mask language model correction task is constructed as:
[0125]
[0126] wherein L MLM represents the loss function of the mask language model, which measures the accuracy of the model in correcting the masked words in the sentence; M represents the set of masked words, i.e., the words in the dialogue sentence that need to be corrected; w i represents the i-th masked word, and the goal is for the model to correctly predict this word; X M represents the sentence context after masking the words M; P(w i |X M ) represents the probability of the model predicting the word w i given the masked context.
[0127] Further, based on the loss function corresponding to the training task designed above, i.e., the loss function of the target self-supervised task, the large language model after training is further trained using the obtained dialogue data containing error information, to obtain a target large language model after secondary training, so as to improve the data error detection and cleaning capability of the large language model on the dialogue data of the customer service system.
[0128] Optionally, the secondary training of the large language model after training based on the loss function of the target self-supervised task and the dialogue data containing error information comprises:
[0129] determining a total loss containing context consistency detection and noise contrast learning based on the loss function of the target self-supervised task and the dialogue data containing error information;
[0130] determining whether there is a target sentence with semantic or grammatical error based on the total loss;
[0131] If there is a target sentence with semantic or grammatical error, the target sentence is masked and the correct words or phrases are predicted by using a mask language model, and the dialogue data after error correction is re-input into the large language model for verification until the dialogue data passes the verification of all detection tasks of the large language model.
[0132] In the embodiments of the present application, the context consistency detection L CCD and the noise contrast learning L NCL can automatically identify the context error and noise in the dialogue. The formula for determining the total loss containing context consistency detection and noise contrast learning based on the loss function of the target self-supervised task and the dialogue data containing error information is as follows:
[0133] L total = L CCD + λ·L NCL ;
[0134] wherein, L total represents the total loss function containing context consistency detection and noise contrast learning; λ represents a balance coefficient for controlling the weight between context consistency detection and noise contrast learning; L CCD represents the context consistency detection loss; and L NCL represents the noise contrast learning loss.
[0135] Further, determining whether there is a target sentence with semantic or grammatical error based on the total loss can realize data error correction. Once an error is detected, the system triggers a correction mechanism, and the specific steps are as follows: (1) for a sentence with semantic or grammatical error, the model will mask the sentence by using a mask language model LMLM Masking sentences and predicting the correct words or phrases. (2) For contextually inconsistent errors, the model adjusts the current dialogue response to be consistent with previous dialogues through a context prediction mechanism. After error correction, the system inputs the corrected dialogue data into the model for verification. If further errors are detected, the model will continue to correct until the dialogue data passes all detection tasks. After completing the training of the target large language model, the trained target large language model is used to automatically clean the dialogue data of the customer service system.
[0136] The application designs an automatic feedback optimization mechanism that dynamically adjusts the loss function weight and optimizes the parameters according to the quality of the cleaned data. This method automatically feeds back the evaluation results to the model, thereby improving the adaptability of the model, quickly and automatically optimizing the data cleaning effect, avoiding the dependence on human intervention, and improving the cleaning efficiency and reliability. An automatic data detection and correction mechanism is also designed, which can automatically detect and correct data errors during self-supervised learning. Through multi-level error detection and automatic correction mechanism, the accuracy of data cleaning can be significantly improved, and the learning and overfitting of the model to bad data can be reduced, optimizing the data quality.
[0137] In summary, compared to the traditional data cleaning task which relies on a large amount of manually annotated data, the self-supervised design of the present application does not require manual annotation and efficiently learns from unannotated data. Through the design of mask prediction, sentence prediction, self-supervised contrastive learning and other tasks, the model can automatically learn useful representations from data. Through the design of multi-task self-supervised learning, the model can not only clean single noisy data, but also optimize for coherence and semantic errors in multi-turn dialogues, further improving accuracy. By jointly training multiple self-supervised tasks, the model can understand the dialogue structure from different perspectives, including the relationship between words and sentences, and the coherence of the context. Multi-task joint optimization can improve the performance of the model in data cleaning, especially in automatically discovering and processing noisy data.
[0138] Compared with the prior art, the present application has the following advantages:
[0139] (1) Reducing the cost of manual annotation: Since self-supervised learning does not rely on a large amount of annotated data, the model can train on a large amount of unannotated data, thereby reducing the manual involvement in data cleaning work.
[0140] (2) Automatic processing of complex data cleaning tasks: Through the design of multi-task self-supervised learning, the model can not only clean single noisy data, but also optimize for coherence and semantic errors in multi-turn dialogues, further improving accuracy.
[0141] (3) Progressive optimization capability: The self-supervised model can continuously receive new data input and pre-train through the self-supervised task, constantly improve the depth of understanding of a specific field, and gradually improve the efficiency and effect of data cleaning.
[0142] The above introduces various methods of embodiments of the present application. The following will further provide a device for implementing the above method.
[0143] Please refer to Figure 2 The embodiments of the present application also provide a data cleaning device based on self-supervised learning, comprising:
[0144] The first processing module 21 is configured to perform initial pre-training on a preset large language model based on a self-supervised learning training task, using unannotated historical dialogue data, to obtain a trained large language model.
[0145] The first determination module 22 is configured to determine dialogue data containing error information using the historical dialogue data.
[0146] The second processing module 23 is configured to perform secondary training on the trained large language model based on the dialogue data containing error information, to obtain a target large language model after training; the target large language model is used for automatic cleaning processing of dialogue data.
[0147] Optionally, the first processing module 21 described above comprises:
[0148] The first determination unit is configured to determine the self-supervised learning training task; the self-supervised learning training task comprises dynamic context perception, task layering, and context contrast learning.
[0149] The first processing unit is configured to perform initial pre-training of dynamic context perception, task layering, and context contrast learning on a preset large language model based on the historical dialogue data, to obtain a trained large language model.
[0150] Optionally, in the case that the self-supervised learning training task comprises dynamic context perception, the first processing unit is specifically configured to:
[0151] At the current time point, the historical dialogue data is used to obtain current input information input into a preset calculation model and a context state corresponding to the current input information;
[0152] According to the current input information and the context state, a predicted context probability of the next time point is obtained.
[0153] According to the context probability of the next time point, a context total loss value of the next time point is obtained.
[0154] According to the context total loss value of the next time point, the large language model is subjected to dynamic context-aware self-supervised learning training to obtain the large language model subjected to dynamic context-aware self-supervised learning training.
[0155] Optionally, in the case that the self-supervised learning training task includes task layering, the first processing unit is specifically configured to:
[0156] Using the historical dialogue data, obtain a task layering label used for training and input features corresponding to the task layering label;
[0157] According to the task layering label and the input features, determine input features corresponding to a current input task layering;
[0158] According to the input features corresponding to the current input task layering, construct a loss function for task layering self-supervised learning;
[0159] Based on the loss function for task layering self-supervised learning, the large language model is subjected to task layering self-supervised learning training to obtain the large language model subjected to task layering self-supervised learning training.
[0160] Optionally, in the case that the self-supervised learning training task includes context contrast learning, the first processing unit is specifically configured to:
[0161] Using the historical dialogue data, obtain a first similarity between an original context and correct subsequent dialogue used for training and a second similarity between the original context and invalid dialogue generated by disturbance;
[0162] According to the first similarity and the second similarity, construct a loss function for context contrast learning self-supervised learning;
[0163] Based on the loss function for context contrast learning self-supervised learning, the large language model is subjected to context contrast learning self-supervised learning training to obtain the large language model subjected to context contrast learning self-supervised learning training.
[0164] Optionally, the second processing module 23 includes:
[0165] A first construction unit is configured to construct a target self-supervised task subjected to secondary training; the target self-supervised task includes a noise contrast learning task, a context consistency detection task and a self-supervised mask language model correction task;
[0166] A second construction unit is configured to construct a loss function of the target self-supervised task;
[0167] The second processing unit is configured to perform secondary training on the trained large language model according to a loss function of the target self-supervised task and the dialogue data containing the error information, to obtain a target large language model after training.
[0168] Optionally, the second processing unit is specifically configured to:
[0169] determine an overall loss containing context consistency detection and noise contrast learning according to the loss function of the target self-supervised task and the dialogue data containing the error information;
[0170] determine whether there is a target sentence with semantic or grammatical error according to the overall loss;
[0171] If there is a target sentence with semantic or grammatical error, a mask language model is used to mask and predict the correct words or phrases of the target sentence, and the dialogue data after error correction is re-input to the large language model for verification until the dialogue data passes the verification of all detection tasks of the large language model.
[0172] It should be noted that the device in this embodiment corresponds to the above-mentioned device applied to the data cleaning method based on self-supervised learning. The implementation manners in the above-mentioned embodiments are applicable to the embodiments of the device and can achieve the same technical effects. The above-mentioned device provided in the embodiments of the present application can realize all method steps realized by the above-mentioned method embodiments and can achieve the same technical effects. Here, the same parts and beneficial effects in the embodiments of the present application as the method embodiments will not be described in detail.
[0173] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement each process of the above-mentioned data cleaning method based on self-supervised learning, and can achieve the same technical effects. To avoid repetition, this will not be described here. The computer readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0174] The embodiments of the present application also provide a computer program product, which includes computer instructions. The computer instructions are executed by a processor to implement each process of the above-mentioned data cleaning method based on self-supervised learning, and can achieve the same technical effects. To avoid repetition, this will not be described here.
[0175] It should be noted that, in the present document, the terms "comprises / comprising" or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0176] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and the necessary general hardware platform, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a part of the prior art that makes a contribution. The computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0177] The embodiments of the present application are described above in combination with the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope of protection of the claims.
Claims
1. A data cleaning method based on self-supervised learning, characterized in that, The method comprises the following steps: training a preset large language model based on a self-supervised learning training task and using unannotated historical dialogue data to obtain a trained large language model; using the historical dialogue data to determine dialogue data containing error information; training the trained large language model based on the dialogue data containing error information to obtain a target large language model for automatic cleaning processing of dialogue data.
2. The method of claim 1, wherein, The method comprises the following steps: determining the self-supervised learning training task; the self-supervised learning training task comprises dynamic context awareness, task layering, and context contrast learning; training a preset large language model based on a self-supervised learning training task and using unannotated historical dialogue data to obtain a trained large language model.
3. The method of claim 2, wherein, In the case that the self-supervised learning training task comprises dynamic context awareness, the method comprises the following steps: at a current time point, using the historical dialogue data to obtain current input information input into a preset computing model and a context state corresponding to the current input information; obtaining a predicted context probability of a next time point based on the current input information and the context state; obtaining a context total loss value of the next time point based on the context probability of the next time point; training the large language model based on the context total loss value of the next time point to obtain a large language model trained based on dynamic context awareness.
4. The method of claim 2, wherein, In the case that the self-supervised learning training task comprises task layering, the method comprises the following steps: using the historical dialogue data to obtain a task layering label used for training and input features corresponding to the task layering label; determining input features corresponding to a current input task layer based on the task layering label and the input features; constructing a loss function of self-supervised learning for task layering based on the input features corresponding to the current input task layer; training the large language model based on the loss function of self-supervised learning for task layering to obtain a large language model trained based on self-supervised learning for task layering.
5. The method of claim 2, wherein, In the case that the self-supervised learning training task comprises context contrast learning, the method comprises the following steps: using the historical dialogue data to obtain a first similarity between an original context and correct subsequent dialogue used for training and a second similarity between the original context and invalid dialogue generated by disturbance; constructing a loss function of self-supervised learning for context contrast learning according to the first similarity and the second similarity; training the large language model by self-supervised learning of context contrast learning based on the loss function of self-supervised learning of context contrast learning, to obtain the large language model after training by self-supervised learning of context contrast learning.
6. The method of claim 1, wherein, According to the dialogue data containing error information, the trained large language model is further trained to obtain a target large language model after training, comprising: constructing a target self-supervised task for further training; the target self-supervised task comprises a noise contrast learning task, a context consistency detection task and a self-supervised mask language model correction task; constructing a loss function of the target self-supervised task; training the large language model according to the loss function of the target self-supervised task and the dialogue data containing error information, to obtain a target large language model after training.
7. The method of claim 6, wherein, According to the loss function of the target self-supervised task and the dialogue data containing error information, the trained large language model is further trained to obtain a target large language model after training, comprising: determining a total loss comprising context consistency detection and noise contrast learning according to the loss function of the target self-supervised task and the dialogue data containing error information; determining whether there is a target sentence with semantic or grammatical errors according to the total loss; if there is a target sentence with semantic or grammatical errors, using a mask language model to mask and predict the correct words or phrases of the target sentence, and re-inputting the error-corrected dialogue data into the large language model for verification until the dialogue data passes the verification of all detection tasks of the large language model.
8. A data cleaning device based on self-supervised learning, characterized by, comprising: a first processing module configured to train a preset large language model by self-supervised learning based on a self-supervised learning training task and using unlabeled historical dialogue data, to obtain a trained large language model; a first determining module configured to determine dialogue data containing error information using the historical dialogue data; a second processing module configured to further train the trained large language model according to the dialogue data containing error information, to obtain a target large language model after training; the target large language model is used for automatic cleaning processing of dialogue data.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.
10. A computer program product, characterised in that, The computer program comprises computer instructions, and the computer instructions are executed by the processor to implement the steps of the method of any one of claims 1 to 7.
Citation Information
Cited By
Training method and device of dialogue interactive German intelligent writing feedback large model
CN121388122A