Language model training method and device, equipment, storage medium and product

By identifying and processing biased content in large language models, using probability distribution of semantic text to train sample pairs, the generalization ability and accuracy problems of model are solved, and more efficient deep semantic understanding and prediction are achieved.

CN120256631APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410009503.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

Smart Images

  • Figure CN120256631A_ABST
    Figure CN120256631A_ABST
Patent Text Reader

Abstract

The invention provides a language model training method and device, equipment, a storage medium and a product, and belongs to the technical field of artificial intelligence. The method comprises the steps of obtaining multiple groups of sample pairs; for each group of sample pairs, inputting the sample pairs and a preset text into a first language model, and obtaining probability distribution of the preset text by taking the sample pairs as prompt information through the first language model, the preset text being a semantic-free text; on the basis of the probability distributions corresponding to the multiple groups of sample pairs, determining prejudice contents in the sample texts of the multiple groups of sample pairs; the multiple sets of sample pairs are processed, multiple sets of processed sample pairs are obtained, and sample texts of the multiple sets of processed sample pairs do not comprise prejudice content; and training the first language model based on the processed multiple groups of sample pairs. According to the method, the generalization and the robustness of the first language model are improved, that is, the performance of the first language model is improved, so that the accuracy of the first language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method, device, equipment, storage medium and product of a language model. Background Art

[0002] Large Language Models (LLMs) have demonstrated powerful semantic representation and reasoning capabilities in various natural language understanding tasks. However, LLMs are affected by the biased content in the training samples and rely on some simple words or sentence patterns for reasoning during inference. For any input text, as long as a certain vocabulary exists in the input text, the LLMs will classify the input text into a certain category, rather than truly understanding the semantics of the input text for category classification. Obviously, this will reduce the generalization ability of the LLMs and further reduce the accuracy of the results predicted by the LLMs. Summary of the Invention

[0003] Embodiments of this application provide a training method, device, equipment, storage medium and product of a language model, which improve the generalization and robustness of the first language model, thereby being able to improve the accuracy of the results predicted by the first language model. The technical solutions are as follows:

[0004] On the one hand, a training method of a language model is provided. The method includes:

[0005] Obtain multiple groups of sample pairs, where each group of sample pairs includes a sample text and the labeled category of the sample text;

[0006] For each group of sample pairs, input the sample pair and a preset text into the first language model, and through the first language model, use the sample pair as a prompt message to obtain the probability distribution of the preset text. The preset text is a text without semantics, and the probability distribution is used to indicate the probabilities that the preset text belongs to multiple categories respectively;

[0007] Based on the probability distributions respectively corresponding to the multiple groups of sample pairs, determine the biased content in the sample texts of the multiple groups of sample pairs. The biased content is the text content in the sample text that has a partial correlation with the labeled category;

[0008] Process the multiple groups of sample pairs to obtain processed multiple groups of sample pairs, and the sample texts of the processed multiple groups of sample pairs do not include the biased content;

[0009] Train the first language model based on the processed multiple groups of sample pairs.

[0010] On the other hand, a training device of a language model is provided. The device includes:

[0011] An acquisition module, configured to acquire multiple groups of sample pairs, each group of sample pairs including a sample text and an annotation category of the sample text;

[0012] An input / output module, configured to, for each group of sample pairs, input the sample pair and a preset text into a first language model, and through the first language model, using the sample pair as a prompt message, obtain a probability distribution of the preset text, where the preset text is a text without semantics, and the probability distribution is used to indicate the probabilities of the preset text belonging to multiple categories respectively;

[0013] A determination module, configured to determine bias content in the sample texts of the multiple groups of sample pairs based on the probability distributions respectively corresponding to the multiple groups of sample pairs, where the bias content is text content in the sample text that has a biased correlation with the annotation category;

[0014] A processing module, configured to process the multiple groups of sample pairs to obtain processed multiple groups of sample pairs, where the sample texts of the processed multiple groups of sample pairs do not include the bias content;

[0015] A training module, configured to train the first language model based on the processed multiple groups of sample pairs.

[0016] In some embodiments, the determination module is configured to:

[0017] For each group of sample pairs, determine a bias value of the sample pair based on the probability distribution corresponding to the sample pair and a uniform distribution, where the bias value is used to indicate the possibility of bias content existing in the sample text of the sample pair;

[0018] Determine the bias content in the sample texts of the multiple groups of sample pairs based on the bias values of the multiple groups of sample pairs respectively.

[0019] In some embodiments, the determination module is configured to:

[0020] Based on the bias values of the multiple groups of sample pairs respectively, determine multiple groups of target sample pairs from the multiple groups of sample pairs, where the target sample pairs are sample pairs with the bias values ranked in the top preset positions;

[0021] Determine the bias content in the sample texts of the multiple groups of sample pairs based on the multiple groups of target sample pairs.

[0022] In some embodiments, the determination module is configured to:

[0023] Determine the bias content in the sample texts of the multiple groups of target sample pairs, and use the bias content in the sample texts of the multiple groups of target sample pairs as the bias content in the sample texts of the multiple groups of sample pairs.

[0024] In some embodiments, the determining module is configured to:

[0025] For each set of target sample pairs, through a second language model, determine the analysis results of the sample text of the target sample pair in multiple dimensions, and based on the analysis results of the sample text of the target sample pair in the multiple dimensions, obtain the biased content in the sample text of the target sample pair, where the multiple dimensions include at least one of morphology, syntax, semantics, and sentiment.

[0026] In some embodiments, the determining module is configured to:

[0027] Determine the relative entropy between the probability distribution corresponding to the sample pair and the uniform distribution, and use the relative entropy as the bias value of the sample pair.

[0028] In some embodiments, the processing module is configured to:

[0029] For each set of sample pairs, through a second language model, delete the biased content in the sample text of the sample pair or replace the biased content based on the target text content, where the target text content is the text content other than the biased content.

[0030] On the other hand, a computer device is provided, which includes a processor and a memory. The memory is used to store at least one program, and the at least one program is loaded and executed by the processor to implement the training method of the language model in the embodiments of the present application.

[0031] On the other hand, a computer-readable storage medium is provided, in which at least one program is stored, and the at least one program is loaded and executed by a processor to implement the training method of the language model in the embodiments of the present application.

[0032] On the other hand, a computer program product is provided, which includes at least one program. The at least one program is stored in a computer-readable storage medium. A processor of a computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device executes the training method of the language model in any of the above implementation manners.

[0033] The embodiments of the present application provide a method for training a language model. For each group of sample pairs, using the sample pairs as prompt information, the probability distribution of a preset text is predicted through a first language model. Since the preset text is a text without semantics, if there is no biased content in the sample text, the predicted probability distribution should be close to a uniform distribution. Therefore, based on this probability distribution, it is possible to indicate the presence of biased content in the sample text, and further, the biased content in the sample text can be determined based on this probability distribution. Then, the biased content in the sample text is processed, and the first language model is trained based on the sample pairs after processing the biased content. Training the first language model based on the sample pairs after processing the biased content can reduce the dependence of the first language model on the biased content, enabling the first language model to make predictions based on deep semantic understanding rather than biased content, thereby improving the generalization and robustness of the first language model, that is, improving the performance of the first language model, and thus improving the accuracy of the first language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application;

[0036] Figure 2 is a flowchart of a method for training a language model provided by the embodiments of the present application;

[0037] Figure 3 is a flowchart of another method for training a language model provided by the embodiments of the present application;

[0038] Figure 4 is a block diagram of a device for training a language model provided by the embodiments of the present application;

[0039] Figure 5 is a block diagram of a terminal provided by the embodiments of the present application;

[0040] Figure 6 is a block diagram of a server provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0042] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or chronological dependence between "first", "second", and "nth", nor are the quantity and execution order limited.

[0043] In this application, the term "at least one" means one or more, and the meaning of "multiple" is two or more.

[0044] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, multiple groups of sample pairs involved in this application are obtained under full authorization.

[0045] Hereinafter, the professional terms involved in this application will be introduced:

[0046] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called the large model or the basic model, and after fine-tuning, it can be widely applied to downstream tasks in various directions of artificial intelligence. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0047] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0048] Natural Language Processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0049] The following introduces the implementation environment related to this application:

[0050] The training method of the language model provided by the embodiments of this application can be executed by a computer device, which can be provided as a server or a terminal. The following introduces the schematic diagram of the implementation environment of the training method of the language model provided by the embodiments of this application.

[0051] See Figure 1 , Figure 1A schematic diagram of an implementation environment for a method of training a language model provided by an embodiment of the present application. This implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this. In some embodiments, the server 102 is used to train a language model, and the trained language model is used to perform natural language understanding tasks, such as classification tasks, translation tasks, dialogue tasks, etc. In some embodiments, a target application is installed on the terminal 101, and this target application is used to perform natural language understanding tasks through the language model. In some embodiments, the trained language model is embedded in the terminal 101, and the terminal 101 performs natural language understanding tasks through the language model. In other embodiments, the terminal 101 performs natural language understanding tasks through the language model on the server 102.

[0052] In some embodiments, the terminal 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc., but is not limited thereto. In some embodiments, the server 102 is an independent server or can also be a server cluster or a distributed system composed of multiple servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing service, and the terminal 101 undertakes the main computing work; or, a distributed computing architecture is adopted between the server 102 and the terminal 101 for collaborative computing.

[0053] See Figure 2 , Figure 2 A flowchart of a method for training a language model provided by an embodiment of the present application. This method includes the following steps.

[0054] 201. The computer device obtains multiple groups of sample pairs, and each group of sample pairs includes a sample text and an annotation category of the sample text.

[0055] In the embodiments of the present application, the multiple groups of sample pairs can be sample pairs obtained from corpora in the open domain, and the open domain includes corpora in multiple fields. Correspondingly, the multiple groups of sample pairs are used to train the first language model to perform natural language understanding tasks in multiple fields.

[0056] The multiple groups of sample pairs can also be sample pairs obtained from the corpus in the target domain. Correspondingly, the multiple groups of sample pairs are used to train the first language model to perform natural language understanding tasks on the downstream tasks corresponding to the target domain. For example, all the multiple groups of sample pairs are sample pairs in the text classification field and are used to train the first language model to perform text classification tasks.

[0057] 202. For each group of sample pairs, the computer device inputs the sample pair and the preset text into the first language model. Through the first language model, using the sample pair as the prompt information, a probability distribution of the preset text is obtained. The preset text is a text without semantics, and the probability distribution is used to indicate the probabilities that the preset text belongs to multiple categories respectively.

[0058] In the embodiments of the present application, the preset text is a text without semantics, and the text content of the preset text can be set and changed as needed. For example, the preset text is a special character without semantics, such as the special character [N / A].

[0059] In the embodiments of the present application, the first large language model is used to determine the probabilities of the preset text on multiple categories under the prompting of the sample text and the labeled categories. That is, the first language model uses the sample text and the labeled categories as context information to predict the category of the preset text.

[0060] 203. The computer device determines the biased content in the sample texts of the multiple groups of sample pairs based on the probability distributions respectively corresponding to the multiple groups of sample pairs. The biased content is the text content in the sample text that has a partial correlation with the labeled category.

[0061] In the embodiments of the present application, since the preset text is a text without semantics, if there is no biased content in the sample text, the predicted probability distribution should be close to a uniform distribution. Therefore, based on this probability distribution, it can indicate the situation of the existence of biased content in the sample text, that is, it reflects the possibility of the existence of biased content in the sample text.

[0062] In the embodiments of the present application, the biased content can be at least one of words, sentence patterns, or any other content. The biased content in the sample text is also the shortcut structure in the sample text. When the first language model infers the input text, it will rely on the biased content in the input text for simple inference to obtain the category to which the input text belongs. For example, as long as a certain biased content is included in the input text, it will be classified into a certain category, rather than obtaining the category through in-depth semantic understanding of the input text. That is, when the first language model outputs the category, it mainly relies on the biased content in the input text for shallow and quick prediction. Predicting in this way by relying on superficial content has low accuracy.

[0063] 204. The computer device processes multiple groups of sample pairs to obtain the processed multiple groups of sample pairs, and the sample texts of the processed multiple groups of sample pairs do not include biased content.

[0064] In the embodiments of the present application, the computer device deletes or replaces the biased content in the sample text, so that the sample text no longer includes biased content.

[0065] 205. The computer device trains the first language model based on the processed multiple groups of sample pairs.

[0066] It should be noted that the computer device iteratively executes the above steps 202-205 through multiple groups of sample pairs to iteratively train the first language model until the preset requirements are met. The preset requirements can be that the number of iterations reaches a preset number, or the probability distribution obtained based on the first language model converges, or the gap between the probability distribution and the uniform distribution is less than a preset gap, etc., which are not specifically limited here.

[0067] The embodiments of the present application provide a method for training a language model. For each group of sample pairs, using the sample pair as the prompt information, the probability distribution of the preset text is predicted through the first language model. Since the preset text is a text without semantics, if there is no biased content in the sample text, the predicted probability distribution should be close to the uniform distribution. Therefore, based on this probability distribution, the situation of the existence of biased content in the sample text can be indicated, and then the biased content in the sample text can be determined based on this probability distribution. Then, the biased content in the sample text is processed, and the first language model is trained based on the sample pairs after processing the biased content. Training the first language model based on the sample pairs after processing the biased content can reduce the dependence of the first language model on the biased content, so that the first language model makes predictions depending on deep semantic understanding rather than biased content, thereby improving the generalization and robustness of the first language model, that is, improving the performance of the first language model, and thus improving the accuracy of the first language model.

[0068] The above Figure 2 is the basic process of the method for training a language model. Next, based on Figure 3 the method for training a language model will be further introduced. Refer to Figure 3 , Figure 3 which is a flowchart of a method for training a language model provided by the embodiments of the present application. This method is executed by a computer device and includes the following steps.

[0069] 301. The computer device obtains multiple groups of sample pairs, and each group of sample pairs includes a sample text and a labeled category.

[0070] In the embodiments of the present application, each group of sample pairs is a binary tuple and can be represented as (x i ,yi )。x i represents the sample text in the i-th sample pair, y i represents the annotation category of the sample text in the i-th sample pair. The sample text can be a word, a sentence, or a paragraph composed of multiple sentences, etc., which is not specifically limited here.

[0071] For example, $y i \in\mathcal{Y}$ represents the annotation category of the sample text in the i-th sample pair, $\mathcal{Y}$ represents the predefined category space, and the category space can be set and changed as needed. For example, the category space includes positive sentiment categories and negative sentiment categories.

[0072] In the embodiments of the present application, the training data set composed of multiple groups of sample pairs can be expressed as D = {(x1, y1), (x2, y2), …, (x N , y N )}. Among them, D represents the training data set, (x1, y1), (x2, y2), (x N , y N ) etc. respectively represent a group of sample pairs. N represents the number of multiple groups of sample pairs.

[0073] 302. For each group of sample pairs in the computer device, the sample pair and the preset text are input into the first language model. Through the first language model, with the sample pair as the prompt information, the probability distribution of the preset text is obtained. The preset text is a text without semantics, and the probability distribution is used to indicate the probabilities that the preset text belongs to multiple categories respectively.

[0074] In the embodiments of the present application, the sample pair is used to construct a prompt template. And the output bias of the first language model, that is, the probability distribution, is actually partly affected by different prompt template construction methods. And the construction of the prompt template is affected by various factors such as the selection of annotated samples and the arrangement order of samples. Therefore, it is very important to evaluate the bias degree of the sample text under the given different prompt templates. And in this embodiment, for each group of sample pairs, it is incorporated into the input template of the first language model as context knowledge, and a preset text without any semantics is inserted after it as the final input, which improves the convenience.

[0075] It should be noted that for a preset text without any semantics, the probability distribution obtained by the first language model should be as evenly distributed as possible. Therefore, the greater the gap between the probability distribution and the uniform distribution, the greater the bias of the corresponding sample text. That is, when the preset text does not contain any semantics, the probability distribution obtained based on a high-quality prompt should be close to the uniform distribution, that is, there should be no prediction bias. Therefore, the probability distribution obtained in the above manner can effectively represent the presence of biased content in the sample text.

[0076] Optionally, the probability distribution is expressed as representing the preset text The probabilities that the preset text belongs to M categories respectively under the prompting effect of the prompt information l.

[0077] 303. For each group of sample pairs, the computer device determines the bias value of the sample pair based on the probability distribution corresponding to the sample pair and the uniform distribution. The bias value is used to indicate the possibility of the presence of biased content in the sample text of the sample pair.

[0078] In the embodiments of the present application, the greater the bias value, the greater the possibility of the presence of biased content in the sample text. Optionally, the bias value can also be used to indicate the amount of biased content present in the sample text. The greater the bias value, the more biased content there is in the sample text.

[0079] In some embodiments, the computer device determines the relative entropy between the probability distribution corresponding to the sample pair and the uniform distribution, and uses the relative entropy as the bias value of the sample pair. The determination process of the relative entropy is shown in the following formula (1).

[0080]

[0081] Wherein, represents the KL (Kullback-Leibler Divergence) distance, that is, the relative entropy. The greater the KL distance, the greater the difference between the two probability distributions. U(i) represents the uniform distribution, represents the probability distribution. i represents the i-th category, and ∑ represents summation. Among them, the probabilities that the sample text belongs to multiple categories in the uniform distribution are the same. The probability distribution and the uniform distribution are each a vector, and the relative entropy is also used to measure the distribution similarity between these two vectors.

[0082] In the embodiments of the present application, the relative entropy is used as the bias value, expressed as Bias(t p ) represents the bias value. Correspondingly, the bias value is also used to indicate the gap between the probability distribution and the uniform distribution, and thus is used to measure the partial correlation between the sample text and the labeled category of the sample text.

[0083] In this embodiment, since the relative entropy between the probability distribution and the uniform distribution can indicate the gap between the probability distribution and the uniform distribution, and the greater the difference between the probability distribution and the uniform distribution, the greater the possibility that there is biased content in the sample text. Therefore, the relative entropy is used as the bias value, which improves the determination efficiency and convenience of the bias value on the basis of ensuring the reliability of the bias value.

[0084] In other embodiments, the computer device may also determine the cosine distance between the probability distribution and the uniform distribution, and use the cosine distance as the bias value.

[0085] 304. The computer device determines the biased content in the sample text of multiple pairs of samples based on the bias values of each pair of samples.

[0086] In some embodiments, the process by which the computer device determines the biased content in the sample text of multiple pairs of samples based on the bias values of each pair of samples includes the following steps: The computer device determines multiple pairs of target samples from multiple pairs of samples based on the bias values of each pair of samples. The target sample pair is the sample pair whose bias value ranks among the top preset positions; based on the multiple pairs of target samples, the biased content in the sample text of the multiple pairs of samples is determined.

[0087] Optionally, the sample set composed of multiple pairs of target samples can be expressed as S = {(x1, y1), (x2, y2), …, (x k , y k )}, where (x1, y1), (x2, y2), (x k , y k ) respectively represent a pair of target samples, and k represents the number of target sample pairs, that is, the number corresponding to the preset position.

[0088] Among them, the computer device sorts multiple pairs of samples based on the bias value. The greater the bias value, the higher the ranking. Then, the sample pair with the bias value ranked among the top preset positions is determined as the target sample pair. In other embodiments, the computer device uses the sample pair with a bias value greater than the bias value threshold as the target sample pair.

[0089] In the embodiments of the present application, multiple sample texts with the largest bias values are found, and thus the biased content in the training samples can be more comprehensively discovered, providing a basis for subsequent removal of biased content. In this embodiment, the greater the bias value, the greater the possibility that there is biased content in the sample text or the more biased content exists. By screening out these sample pairs with large bias values, that is, finding the sample pairs most likely to have biased content, it is then convenient to find the biased content in the sample pairs. This not only improves the efficiency without having to search for biased content based on all sample pairs, but also ensures that the biased content can be comprehensively searched for.

[0090] In some embodiments, the process by which the above computer device determines the biased content in the sample texts of multiple groups of target sample pairs includes the following steps: The computer device determines the biased content in the sample texts of multiple groups of target sample pairs, and uses the biased content in the sample texts of multiple groups of target sample pairs as the biased content in the sample texts of multiple groups of sample pairs.

[0091] In this embodiment, the biased content of multiple groups of target sample pairs is directly used as the biased content of multiple sample pairs. Since the bias value of the target sample pair is large, the possibility of having biased content is high. Furthermore, the accuracy of the biased content found from the target sample pair is high. Then, the biased content in multiple groups of target sample pairs is used as the overall biased content, which improves the efficiency on the basis of ensuring the accuracy.

[0092] In some embodiments, the process by which the above computer device determines the biased content in the sample texts of multiple groups of target sample pairs includes the following steps: For each group of target sample pairs, the computer device determines the analysis results of the sample text of the target sample pair in multiple dimensions through a second language model. Based on the analysis results of the sample text of the target sample pair in multiple dimensions, the biased content in the sample text of the target sample pair is obtained. The multiple dimensions include at least one of morphology, syntax, semantics, and sentiment.

[0093] In the embodiments of the present application, the first language model and the second language model can be the same language model or different language models. The first language model is the language model to be trained, and the second language model is the language model for extracting biased content and modifying samples.

[0094] Optionally, the computer device inputs the target sample pair into the second language model, instructs the second language model to analyze the sample text of the target sample pair in multiple dimensions, and instructs the second language model to determine which words, sentence patterns, etc. in the sample text have a strong partial correlation with the annotation category and may thus become biased content.

[0095] In the embodiments of the present application, the biased content of each sample pair can be represented as s i = M(x i , y i ), i = 1, …, k. Wherein, s i represents the biased content in the i-th sample pair M.

[0096] In the embodiments of the present application, the lexical analysis of the sample text includes part-of-speech tagging of the sample text, judgment of semantic roles, etc. The syntactic analysis of the sample text includes analysis of the sentence structure, the dependency relationship between words, etc. The semantic analysis of the sample text includes extraction of the sentence theme, keywords, etc. The sentiment analysis of the sample text includes determination of the sentiment tendency of the text.

[0097] In this embodiment, through the above steps 303-304, the process of determining the biased content in the sample text of multiple groups of sample pairs based on the corresponding probability distributions of the multiple groups of sample pairs is realized. In this embodiment, an index for determining a bias value is based on the probability distribution. Since this bias value is determined based on the probability distribution and the uniform distribution, it can indicate the possibility of the existence of a bias value in the sample text. Furthermore, based on this bias value, the biased content is determined, realizing the determination of the biased content in a quantified manner through the index, improving the reliability and convenience of determining the biased content.

[0098] It should be noted that the above steps 303-304 are only an optional implementation manner for realizing the process of determining the biased content in the sample text of multiple groups of sample pairs based on the corresponding probability distributions of the multiple groups of sample pairs. The computer device can also implement this process through other optional implementation manners. For example, the computer device directly determines the biased content in the sample text of each group of sample pairs through the second large language model.

[0099] 305. The computer device processes multiple groups of sample pairs to obtain the processed multiple groups of sample pairs, and the sample text of the processed multiple groups of sample pairs does not include biased content.

[0100] In some embodiments, the process of processing the multiple groups of sample pairs to obtain the processed multiple groups of sample pairs includes the following steps: for each group of sample pairs, the computer device deletes the biased content in the sample text of the sample pair or replaces the biased content based on the target text content through the second language model, and the target text content is the text content other than the biased content.

[0101] Among them, the computer device inputs each group of sample pairs into the second language model, instructing the second language model to delete the biased content in the sample text or replace the biased content based on the target text content according to the determined biased content. Among them, the target text content can be set and changed as needed. For example, the target text content can be any text content without semantics. Again, the target text content is the text content related to the biased content and without semantics.

[0102] In the embodiments of the present application, if the second language model is used to analyze the text and extract the biased content, the second language model can directly process the biased content in the target sample pair based on the extracted biased content in the target sample pair.

[0103] In this embodiment, the biased content in the sample text is processed by automatically deleting or replacing it through a language model, which can effectively remove the biased content in the sample text and improve the efficiency.

[0104] In the embodiment of the present application, the computer device determines the biased content through multiple groups of target sample pairs.

[0105] In some embodiments, the computer device only processes the biased content in the target sample pairs. Optionally, for each group of target sample pairs, the biased content determined from the sample text passing through the target sample pair is processed. Or, for each group of target sample pairs, not only the biased content determined from the sample text passing through the target sample pair is processed, but also the biased content determined from the sample text by other target sample pairs in the sample text is processed. In this embodiment, since the bias value of the target sample pair is large and the existing biased content has a great impact on the first language model, processing the biased content in the target sample pair not only ensures the effective processing of the biased content in multiple groups of sample pairs but also improves the processing efficiency.

[0106] In other embodiments, the computer device, based on the biased content determined through multiple groups of target sample pairs, for any sample pair, finds and processes any biased content existing in the sample text of the sample pair. In this embodiment, the biased content is processed for each sample pair separately, improving comprehensiveness. Then, by training the first language model with the sample pairs from which the biased content has been comprehensively removed, the dependence of the first language model on the biased content can be further reduced, and the generalization and robustness of the first language model can be improved.

[0107] In the embodiment of the present application, according to the determined bias structure, a new enhanced sample text is generated through the second language model to reduce the dependence of the first language model on the biased content. Among them, for each sample text including biased content, the sample text after processing the biased content can be represented as x i ′ = M(x i , y i , s i ). The sample set composed of the sample pairs after processing the biased content can be represented as S ′ = {(x1 ′ , y1), (x2 ′ , y2), …, (x ′ k , y k )}.

[0108] 306. The computer device trains the first language model based on the processed multiple groups of sample pairs.

[0109] It should be noted that among the multiple groups of sample pairs, there are sample pairs containing biased content and sample pairs without biased content. Among them, the multiple groups of sample pairs for processing biased content include the sample pairs obtained by processing the biased content in step 305, and also include the sample pairs that originally do not have biased content.

[0110] In some other embodiments, after the computer device determines the target sample pairs among the multiple groups of sample pairs, it can directly perform model training based on the sample pairs other than the target sample pairs among the multiple groups of sample pairs, without performing the steps of determining and processing the biased content.

[0111] Through the method provided in the embodiments of the present application, potential biased content can be automatically found from the given sample text, and the corresponding biased content in the sample text can be further automatically eliminated through the large language model, so that the first language model relies on the deep semantics of the input text rather than the biased content for prediction, thereby improving the generalization ability of the first language model.

[0112] In the embodiments of the present application, the computer device performs iterative training on the first language model based on multiple groups of sample pairs. In some embodiments, the computer device re-executes steps 302-305 on the trained first language model based on the multiple groups of sample pairs to obtain multiple groups of sample pairs for processing the biased content again, and then executes step 306 for training. In some other embodiments, the computer device re-executes steps 302-305 on the trained first language model based on the multiple groups of sample pairs after processing the biased content, processes the biased content again based on the multiple groups of sample pairs for processing the biased content, and then executes step 306 for training.

[0113] In the embodiments of the present application, the computer device performs iterative training on the first language model based on multiple groups of sample pairs until a preset requirement is met. Among them, meeting the preset requirement can be that the number of iterations reaches the preset number of times, or the bias value obtained based on the probability distribution reaches the preset threshold, or the gap between the bias value of the trained first language model and the training value of the first language model before training reaches the preset gap, that is, the bias value obtained by the trained first language model is significantly improved compared to the first language model before training.

[0114] In the embodiments of the present application, the improvement of the bias value may be caused by the update of the sample pair or the update of the first language model. Therefore, the method of controlling variables can be used to determine the reason for the improvement. For example, the first language model before training and the first language model after training are respectively used to test on the same sample set, and the changes in the bias value are compared. The sample set can be a sample set composed of multiple sample pairs before processing the biased content, or a sample set composed of multiple sample pairs after processing the biased content. If the first language model after training obtains a smaller bias value on the same sample set, it indicates that the improvement of the bias value is caused by the update of the first language model. Similarly, multiple sample pairs before processing the biased content and multiple sample pairs after processing the biased content can be used to test on the same first language model. The first language model can be the first language model before training, or any first language model after training. If multiple sample pairs after processing the biased content obtain a smaller bias value on the first language model, it indicates that the improvement of the bias value is caused by the update of the sample pair. Similarly, it can also be determined whether the improvement of the bias value is caused by the update of the sample pair and the update of the first language model together.

[0115] In this embodiment, taking the example of detecting whether the first language model is effectively improved by comparing the bias values of multiple sample pairs before processing the biased content on the first language model before and after training is described. Among them, the bias value comparison can compare the mean value or the sum value of the bias values of multiple sample pairs.

[0116] The method provided by the embodiments of the present application is the first to examine the inherent shortcut learning tendency of large language models in context learning without parameter updates. Through the method provided by the embodiments of the present application, the way in which large language models naturally process prompt information can be understood more deeply.

[0117] It should be noted that shortcut learning is an important obstacle to improving the generalization ability of neural network models. In multiple natural language understanding tasks, it has been found that the fine-tuned language model is vulnerable to the influence of biased content in the training samples and relies on some simple words or sentence patterns for reasoning, resulting in poor generalization ability of the language model. Therefore, how to efficiently find samples with potential biased content has become a key issue. The embodiments of the present application provide a method combining biased content screening and large language model analysis, which can automatically find potential biased content from the given samples and further process the corresponding biased content through automated data augmentation, thereby improving the generalization ability of the language model.

[0118] An embodiment of the present application provides a method for training a language model. For each group of sample pairs, using the sample pair as a prompt message, the probability distribution of a preset text is predicted through a first language model. Since the preset text is a text without semantics, if there is no biased content in the sample text, the predicted probability distribution should be close to a uniform distribution. Therefore, based on this probability distribution, the situation of the existence of biased content in the sample text can be indicated, and then the biased content in the sample text can be determined based on this probability distribution. Then, the biased content in the sample text is processed, and the first language model is trained based on the sample pairs after processing the biased content. Training the first language model based on the sample pairs after processing the biased content can reduce the dependence of the first language model on the biased content, so that the first language model makes predictions based on deep semantic understanding rather than biased content, thereby improving the generalization and robustness of the first language model, that is, improving the performance of the first language model, and thus improving the accuracy of the first language model.

[0119] Figure 4 is a block diagram of a training device for a language model according to an embodiment of the present application. Refer to Figure 4 , the device includes:

[0120] An acquisition module 401, configured to acquire multiple groups of sample pairs, where each group of sample pairs includes a sample text and an annotation category of the sample text;

[0121] An input / output module 402, configured to, for each group of sample pairs, input the sample pair and a preset text into a first language model, and through the first language model, using the sample pair as a prompt message, obtain the probability distribution of the preset text. The preset text is a text without semantics, and the probability distribution is used to indicate the probabilities that the preset text belongs to multiple categories respectively;

[0122] A determination module 403, configured to determine the biased content in the sample texts of multiple groups of sample pairs based on the probability distributions respectively corresponding to the multiple groups of sample pairs. The biased content is the text content in the sample text that has a partial correlation with the annotation category;

[0123] A processing module 404, configured to process multiple groups of sample pairs to obtain multiple groups of processed sample pairs, and the sample texts of the multiple groups of processed sample pairs do not include biased content;

[0124] A training module 405, configured to train the first language model based on the multiple groups of processed sample pairs.

[0125] In some embodiments, the determination module 403 is configured to:

[0126] For each group of sample pairs, determine a bias value based on the probability distribution corresponding to the sample pair and the uniform distribution. The bias value is used to indicate the possibility of the existence of biased content in the sample text of the sample pair;

[0127] Based on the bias values of multiple groups of samples with respect to each other, determine the biased content in the sample texts of the multiple groups of sample pairs.

[0128] In some embodiments, the determining module 403 is configured to:

[0129] Based on the bias values of multiple groups of samples with respect to each other, determine multiple groups of target sample pairs from the multiple groups of sample pairs, where the target sample pairs are the sample pairs with the bias values ranked in the top preset positions;

[0130] Based on the multiple groups of target sample pairs, determine the biased content in the sample texts of the multiple groups of sample pairs.

[0131] In some embodiments, the determining module 403 is configured to:

[0132] Determine the biased content in the sample texts of the multiple groups of target sample pairs, and use the biased content in the sample texts of the multiple groups of target sample pairs as the biased content in the sample texts of the multiple groups of sample pairs.

[0133] In some embodiments, the determining module 403 is configured to:

[0134] For each group of target sample pairs, through a second language model, determine the analysis results of the sample text of the target sample pair in multiple dimensions, and based on the analysis results of the sample text of the target sample pair in multiple dimensions, obtain the biased content in the sample text of the target sample pair, where the multiple dimensions include at least one of morphology, syntax, semantics, and sentiment.

[0135] In some embodiments, the determining module 403 is configured to:

[0136] Determine the relative entropy between the probability distribution corresponding to the sample pair and the uniform distribution, and use the relative entropy as the bias value of the sample pair.

[0137] In some embodiments, the processing module 404 is configured to:

[0138] For each group of sample pairs, through a second language model, delete the biased content in the sample text of the sample pair or replace the biased content based on the target text content, where the target text content is the text content other than the biased content.

[0139] An embodiment of the present application provides a training device for a language model. For each group of sample pairs, using the sample pair as a prompt message, the probability distribution of a preset text is predicted through a first language model. Since the preset text is a text without semantics, if there is no biased content in the sample text, the predicted probability distribution should be close to a uniform distribution. Therefore, based on this probability distribution, the situation of the existence of biased content in the sample text can be indicated, and then the biased content in the sample text can be determined based on this probability distribution. Then, the biased content in the sample text is processed, and the first language model is trained based on the sample pairs after processing the biased content. Training the first language model based on the sample pairs after processing the biased content can reduce the dependence of the first language model on the biased content, so that the first language model makes predictions based on deep semantic understanding rather than biased content, thereby improving the generalization and robustness of the first language model, that is, improving the performance of the first language model, and thus improving the accuracy of the first language model.

[0140] In an embodiment of the present application, the computer device can be a terminal or a server. When the computer device is a terminal, the terminal is used as the execution subject to implement the technical solution provided by the embodiment of the present application; when the computer device is a server, the server is used as the execution subject to implement the technical solution provided by the embodiment of the present application; or, the technical solution provided by the present application is implemented through the interaction between the terminal and the server. The embodiment of the present application does not limit this.

[0141] Figure 5 The structural block diagram of a terminal 500 provided by an exemplary embodiment of the present application is shown.

[0142] Generally, the terminal 500 includes: a processor 501 and a memory 502.

[0143] The processor 501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0144] The memory 502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 502 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 502 is used to store at least one program code, and the at least one program code is used to be executed by the processor 501 to implement the training method of the language model provided in the method embodiments of the present application.

[0145] In some embodiments, the terminal 500 may further optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, the memory 502, and the peripheral device interface 503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 503 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, and a power supply 508.

[0146] The peripheral device interface 503 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0147] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 504 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 504 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0148] The display screen 505 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 505 is a touch display screen, the display screen 505 also has the ability to collect touch signals on or above the surface of the display screen 505. The touch signals can be input to the processor 501 as control signals for processing. At this time, the display screen 505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 505, which is provided on the front panel of the terminal 500; in other embodiments, there may be at least two display screens 505, which are respectively provided on different surfaces of the terminal 500 or are in a foldable design; in other embodiments, the display screen 505 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 500. Even, the display screen 505 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 505 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0149] The camera module 506 is used to collect images or videos. Optionally, the camera module 506 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 506 may also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0150] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 501 for processing, or input to the radio frequency circuit 504 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 507 may further include a headphone jack.

[0151] The power supply 508 is used to supply power to each component in the terminal 500. The power supply 508 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 508 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0152] In some embodiments, the terminal 500 further includes one or more sensors 509. The one or more sensors 509 include but are not limited to: an acceleration sensor 510, a gyroscope sensor 511, a pressure sensor 512, an optical sensor 513, and a proximity sensor 514.

[0153] The acceleration sensor 510 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 500. For example, the acceleration sensor 510 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 501 can control the display screen 505 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 510. The acceleration sensor 510 can also be used for collecting game or user's motion data.

[0154] The gyroscope sensor 511 can detect the body direction and rotation angle of the terminal 500. The gyroscope sensor 511 can cooperate with the acceleration sensor 510 to collect the 3D actions of the user on the terminal 500. According to the data collected by the gyroscope sensor 511, the processor 501 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0155] The pressure sensor 512 may be disposed on the side frame of the terminal 500 and / or the lower layer of the display screen 505. When the pressure sensor 512 is disposed on the side frame of the terminal 500, it can detect the holding signal of the user on the terminal 500, and the processor 501 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 512. When the pressure sensor 512 is disposed on the lower layer of the display screen 505, the processor 501 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 505. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0156] The optical sensor 513 is used to collect the ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 according to the ambient light intensity collected by the optical sensor 513. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is decreased. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera module 506 according to the ambient light intensity collected by the optical sensor 513.

[0157] The proximity sensor 514, also known as a distance sensor, is usually disposed on the front panel of the terminal 500. The proximity sensor 514 is used to collect the distance between the user and the front of the terminal 500. In one embodiment, when the proximity sensor 514 detects that the distance between the user and the front of the terminal 500 is gradually decreasing, the processor 501 controls the display screen 505 to switch from the lit state to the off state; when the proximity sensor 514 detects that the distance between the user and the front of the terminal 500 is gradually increasing, the processor 501 controls the display screen 505 to switch from the off state to the lit state.

[0158] Those skilled in the art can understand that Figure 5 the structure shown in does not constitute a limitation on the terminal 500, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0159] Figure 6It is a schematic structural diagram of a server provided according to an embodiment of the present application. The server 600 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 601 and one or more memories 602. Among them, the memory 602 is used to store executable program codes, and the processor 601 is configured to execute the above executable program codes to implement the training method of the language model provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.

[0160] The embodiment of the present application also provides a computer-readable storage medium, in which at least one segment of program is stored. The at least one segment of program is loaded and executed by a processor to implement the training method of the language model in any of the above implementation manners.

[0161] The embodiment of the present application also provides a computer program product, which includes at least one segment of program. The at least one segment of program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one segment of program from the computer-readable storage medium, and the processor executes the at least one segment of program, so that the computer device executes the training method of the language model in any of the above implementation manners.

[0162] In some embodiments, the computer program product involved in the embodiment of the present application can be deployed to be executed on a computer device, or on multiple computer devices located at one place. Or, it can be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.

[0163] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a language model, characterized in that, The method includes: Obtaining multiple groups of sample pairs, where each group of sample pairs includes a sample text and the labeled category of the sample text; For each group of sample pairs, inputting the sample pair and a preset text into a first language model, and through the first language model, using the sample pair as prompt information to obtain the probability distribution of the preset text, where the preset text is a text without semantics, and the probability distribution is used to indicate the probabilities of the preset text belonging to multiple categories respectively; Based on the probability distributions respectively corresponding to the multiple groups of sample pairs, determining the biased content in the sample texts of the multiple groups of sample pairs, where the biased content is the text content in the sample text that has a biased correlation with the labeled category; Processing the multiple groups of sample pairs to obtain processed multiple groups of sample pairs, where the sample texts of the processed multiple groups of sample pairs do not include the biased content; Training the first language model based on the processed multiple groups of sample pairs.

2. The method according to claim 1, characterized in that, The determining the biased content in the sample texts of the multiple groups of sample pairs based on the probability distributions respectively corresponding to the multiple groups of sample pairs includes: For each group of sample pairs, determining the bias value of the sample pair based on the probability distribution corresponding to the sample pair and the uniform distribution, where the bias value is used to indicate the possibility of the existence of biased content in the sample text of the sample pair; Based on the bias values of the multiple groups of sample pairs respectively, determining the biased content in the sample texts of the multiple groups of sample pairs.

3. The method according to claim 2, characterized in that, The determining the biased content in the sample texts of the multiple groups of sample pairs based on the bias values of the multiple groups of sample pairs respectively includes: Based on the bias values of the multiple groups of sample pairs respectively, determining multiple groups of target sample pairs from the multiple groups of sample pairs, where the target sample pairs are the sample pairs whose bias values are ranked in the top preset positions; Based on the multiple groups of target sample pairs, determining the biased content in the sample texts of the multiple groups of sample pairs.

4. The method according to claim 3, characterized in that The determining the biased content in the sample texts of the multiple groups of sample pairs based on the multiple groups of target sample pairs includes: Determining the biased content in the sample texts of the multiple groups of target sample pairs, and taking the biased content in the sample texts of the multiple groups of target sample pairs as the biased content in the sample texts of the multiple groups of sample pairs.

5. The method according to claim 4, characterized in that, The determining the biased content in the sample texts of the multiple groups of target sample pairs includes: For each group of target sample pairs, through a second language model, determining the analysis results of the sample text of the target sample pair in multiple dimensions, and based on the analysis results of the sample text of the target sample pair in the multiple dimensions, obtaining the biased content in the sample text of the target sample pair, where the multiple dimensions include at least one of morphology, syntax, semantics, and sentiment.

6. The method according to claim 2, wherein The determining the bias value of the sample pair based on the probability distribution corresponding to the sample pair and the uniform distribution includes: Determining the relative entropy between the probability distribution corresponding to the sample pair and the uniform distribution, and taking the relative entropy as the bias value of the sample pair.

7. The method according to claim 1, wherein The processing the multiple groups of sample pairs to obtain processed multiple groups of sample pairs includes: For each group of sample pairs, through a second language model, bias content in the sample text of the sample pair is deleted or the bias content is replaced based on the target text content, where the target text content is the text content other than the bias content.

8. A training device for a language model, characterized in that, The device includes: an acquisition module, configured to acquire multiple groups of sample pairs, each group of sample pairs including a sample text and an annotation category of the sample text; an input / output module, configured to, for each group of sample pairs, input the sample pair and a preset text into a first language model, and through the first language model, use the sample pair as a prompt message to obtain a probability distribution of the preset text, where the preset text is a text without semantics, and the probability distribution is used to indicate the probabilities of the preset text belonging to multiple categories respectively; a determination module, configured to determine bias content in the sample text of the multiple groups of sample pairs based on the probability distributions respectively corresponding to the multiple groups of sample pairs, where the bias content is text content in the sample text that has a biased correlation with the annotation category; a processing module, configured to process the multiple groups of sample pairs to obtain processed multiple groups of sample pairs, where the sample texts of the processed multiple groups of sample pairs do not include the bias content; a training module, configured to train the first language model based on the processed multiple groups of sample pairs.

9. A computer device, characterized in that, The computer device includes a processor and a memory, where the memory is used to store at least one segment of program, and the at least one segment of program is loaded and executed by the processor to perform the training method of the language model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one segment of program, and the at least one segment of program is used to perform the training method of the language model according to any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes at least one segment of program, the at least one segment of program is stored in the computer-readable storage medium, the processor of the computer device reads the at least one segment of program from the computer-readable storage medium, and the processor executes the at least one segment of program, so that the computer device performs the training method of the language model according to any one of claims 1 to 7.