Model training method and device, electronic equipment, computer readable storage medium and computer program product

By differentially processing the parameter matrix of the language model and combining training samples and loss functions, the gradient and update degree of parameter values ​​are dynamically adjusted, which solves the problem of long training time and poor performance of language models, and achieves optimization of model performance and improvement of adaptability.

CN120975133APending Publication Date: 2025-11-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511079694.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies suffer from long training times and poor training results in language model training, especially when performing question answering tasks in vertical domains, making it difficult to simultaneously maintain the initial performance of the model and optimize training results.

Method used

By obtaining the parameter matrix of the language model, determining the difference between each parameter value and the matrix mean, calculating the update degree parameter, and combining the training samples and the loss function, the gradient and update degree of the parameter values ​​are dynamically adjusted to achieve differentiated training of the language model.

Benefits of technology

The training effect was optimized while retaining the initial performance of the model, improving the model's adaptability and accuracy in vertical domain tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975133A_ABST
    Figure CN120975133A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining a first parameter matrix of a first language model; for the first parameter matrix, based on a first difference value between each first parameter value in the first parameter matrix and a matrix mean value of the first parameter matrix, determining a first update degree parameter of each first parameter value in the first parameter matrix; determining a first loss of the first language model based on the first training sample, and determining a gradient of each first parameter value based on the first loss; and based on the gradient of each first parameter value and the first updating degree parameter of each first parameter value, updating each first parameter value in at least one first parameter matrix of the first language model to obtain a second language model. According to the method and the device, the training degree of the model parameters can be processed differently, so that the training effect can be optimized, and the initial performance of the model can be reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] In the field of machine learning, language models gradually acquire the ability to perform certain tasks through training. The training process of language models involves updating parameters. Usually, in order to optimize the training effect, the language model is updated in stages, which leads to a longer training time. Moreover, since the language model performs question answering in vertical domains, the requirements for training effect are high. The training effect obtained by the above method is often unsatisfactory. Summary of the Invention

[0003] This application provides a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can differentiate the training level of model parameters, thereby not only optimizing the training effect but also preserving the initial performance of the model.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a model training method, including:

[0006] Obtain at least one first parameter matrix of the first language model;

[0007] For each of the first parameter matrices, a first update degree parameter is determined for each first parameter value in the first parameter matrix based on the first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix.

[0008] Based on the first training sample, determine the first loss of the first language model, and determine the gradient of each of the first parameter values ​​based on the first loss;

[0009] Based on the gradient of each first parameter value and the first update degree parameter of each first parameter value, each first parameter value in at least one first parameter matrix of the first language model is updated to obtain a second language model.

[0010] This application also provides a model training method, including:

[0011] Generate a first text sample corresponding to the first task, and obtain a second text sample corresponding to the second task and associated with the first labeled data;

[0012] The first text sample is subjected to forward inference corresponding to the first task by a third language model to obtain a second prediction result corresponding to the first text sample, wherein the third language model is a model trained based on the first task;

[0013] The first text sample is subjected to forward inference corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample, and the second text sample is subjected to forward inference corresponding to the second task to obtain a first prediction result corresponding to the second text sample.

[0014] Based on the first prediction result and the second prediction result corresponding to the first text sample, a second loss corresponding to the first text sample is determined, and based on the first prediction result and the first annotation data corresponding to the second text sample, a third loss corresponding to the second text sample is determined.

[0015] The first language model is updated based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the second language model.

[0016] This application provides a model training apparatus, including:

[0017] The acquisition module is used to acquire at least one first parameter matrix of the first language model;

[0018] The first determining module is used to determine, for each of the first parameter matrices, a first update degree parameter for each first parameter value in the first parameter matrix based on a first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix;

[0019] The first determining module is further configured to determine a first loss of the first language model based on the first training samples, and to determine the gradient of each of the first parameter values ​​based on the first loss;

[0020] A first training module is used to update each of the first parameter values ​​in at least one first parameter matrix of the first language model based on the gradient of each of the first parameter values ​​and a first update degree parameter of each of the first parameter values, so as to obtain a second language model.

[0021] This application embodiment also provides a model training apparatus, including:

[0022] The generation module is used to generate a first text sample corresponding to the first task and to obtain a second text sample corresponding to the second task and associated with the first annotation data.

[0023] The reasoning module is used to perform forward reasoning on the first text sample corresponding to the first task using a third language model to obtain a second prediction result corresponding to the first text sample, wherein the third language model is a model trained based on the first task.

[0024] The reasoning module is further configured to perform forward reasoning on the first text sample corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample, and to perform forward reasoning on the second text sample to obtain a first prediction result corresponding to the second text sample.

[0025] The second determining module is used to determine a second loss corresponding to the first text sample based on the first prediction result and the second prediction result corresponding to the first text sample, and to determine a third loss corresponding to the second text sample based on the first prediction result and the first annotation data.

[0026] The second training module is used to perform fusion processing based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the first loss of the first language model, and to update the first language model based on the first loss of the first language model to obtain the second language model.

[0027] This application provides an electronic device, including:

[0028] Memory is used to store executable instructions or computer programs.

[0029] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the model training method provided in the embodiments of this application.

[0030] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, they implement the model training method provided in this application.

[0031] This application provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the model training method provided in this application.

[0032] The embodiments of this application have the following beneficial effects:

[0033] First, obtaining at least one first parameter matrix of the first language model clarifies the trainable parameters of the language model, laying the foundation for subsequent differentiated updates. Then, for each first parameter matrix, an update degree parameter is determined based on the first difference between each first parameter value and the matrix mean of the first parameter matrix. Different update degree parameters can be applied to different first parameter values, thereby controlling the learning degree of different first parameter values ​​and helping to maintain the initial performance of the first language model. Next, a first loss of the first language model is determined based on the training samples, and the gradient of each first parameter value is determined based on the first loss. Finally, the parameters of the first language model are updated based on the gradients of the first parameter values ​​and the corresponding update degree parameters, resulting in a second language model. By combining gradients and update degree parameters, differentiated processing of model parameter training and preservation is achieved, enabling precise learning of model parameters. This not only optimizes the training effect but also preserves the initial performance of the model, ultimately improving the overall training effect of the model. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the architecture of the model training system provided in the embodiments of this application;

[0035] Figure 2A This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0036] Figure 2B This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0037] Figure 3 This is a schematic diagram of the first process of the model training method provided in the embodiments of this application;

[0038] Figure 4 This is a schematic diagram of the second process of the model training method provided in the embodiments of this application;

[0039] Figure 5 This is a schematic diagram of the third process of the model training method provided in the embodiments of this application;

[0040] Figure 6 This is a schematic diagram of the fourth process of the model training method provided in the embodiments of this application;

[0041] Figure 7 This is a schematic diagram of the fifth process of the model training method provided in the embodiments of this application;

[0042] Figure 8 This is a schematic diagram of the sixth process of the model training method provided in the embodiments of this application;

[0043] Figure 9This is a schematic diagram of the seventh process of the model training method provided in the embodiments of this application;

[0044] Figure 10 This is a schematic diagram of the eighth process of the model training method provided in the embodiments of this application;

[0045] Figure 11 This is a service flow framework diagram of the model training method provided in the embodiments of this application;

[0046] Figure 12 This is a schematic diagram of the interface for an information summary task provided in an embodiment of this application;

[0047] Figure 13 This is a flowchart illustrating the training process of the model training method provided in this application embodiment;

[0048] Figure 14 This is a schematic diagram of data preparation for the model training method provided in the embodiments of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0051] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0052] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0053] In the following description, the terms “first, second, ...” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second, ...” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0055] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0056] 1) Incremental learning: Incremental learning refers to the continuous learning and adaptation of a model throughout its working life, integrating new knowledge while retaining previously learned information to prevent catastrophic forgetting.

[0057] 2) Catastrophic Forgetting: Catastrophic forgetting refers to the phenomenon that a model rapidly loses previously learned knowledge when continuously learning new tasks, resulting in a significant decline in the performance of old tasks.

[0058] 3) The Plasticity-Plasticity Dilemma: This refers to the difficulty a model faces in simultaneously maintaining a balance between its ability to adapt to new knowledge (plasticity) and its ability to retain old knowledge (stability) during continuous learning. Plasticity refers to a model's ability to quickly absorb new data and adjust parameters to adapt to new tasks, manifested as efficient learning of unknown patterns; stability refers to a model's ability to solidify existing knowledge, resist parameter drift, and ensure that the performance of mastered tasks does not degrade.

[0059] 4) Vertical Domain Large Model: A vertical domain large model (also known as a vertical domain large model) is an artificial intelligence model that focuses on a specific industry or domain. Through in-depth training and optimization of domain-specific data, it achieves the ability to accurately process complex tasks in that domain.

[0060] Incremental learning refers to a model continuously learning and adapting throughout its working life, integrating new knowledge while retaining previously learned information to prevent catastrophic forgetting. The main challenges of incremental learning include: 1) Catastrophic forgetting: Catastrophic forgetting is one of the core challenges of lifelong learning, as the introduction of new information may overwrite previously learned content; 2) The plasticity-stability dilemma: Finding a balance between a model's learning ability and stability directly affects its ability to acquire new knowledge and retain its broad generalizability; 3) High computational costs: Full fine-tuning of large language models incurs very high computational costs; 4) Unavailability of model parameters or pre-training data: Due to privacy, proprietary restrictions, or commercial licensing, raw training data or model parameters are often unavailable for further improvement.

[0061] Due to different application objectives, large-scale vertical domain models often perform poorly when faced with new, untrained requirements. For example, large-scale models for script comprehension and script generation (used for creating or generating a text description) often struggle to accurately summarize information from film and television dramas. This is because the summaries focus on summarizing actual data and have high accuracy requirements, while script creation and comprehension involve fewer numbers and no data consistency requirements. Furthermore, due to data privacy and other reasons, historical training data is unavailable, leading to stability-plasticity dilemmas and difficulties in incremental learning when the model performs tasks, such as the unavailability of old training data.

[0062] In related technologies, incremental learning of models is often achieved through the following methods:

[0063] (1) Collect data from all tasks and train the model using all task data (including task data for incremental learning). However, this method is difficult to implement in reality because old task data is often unavailable (due to reasons such as confidentiality or data loss).

[0064] (2) Adapter tuning involves adding an adapter to the original model during training, which is equivalent to expanding the model parameters to improve the model's performance on new tasks. However, since the model structure has been changed, it will cause irreversible damage to the generalizability of the model on other downstream tasks; and by adding an adapter, the characteristics of the old task during inference are changed, so the correctness of the model on the old task cannot be guaranteed, thus causing the problem of forgetting the old task.

[0065] (3) A fine-tuning method based on prefix optimization of prompt words (Prefix tuning) constructs a segment of task-related virtual tokens as a prefix before the input tokens. During training, only the parameters of the prefix part are updated, while the parameters of other parts of the pre-trained language model (PLM) remain fixed. By specifically learning task keywords, task keywords are bound to the data. This method fine-tunes the prefix part and can adapt to new tasks. However, due to the limited training parameters, it is difficult to guarantee the training effect for new tasks; and it cannot overcome the illusion of a large model, making it difficult to guarantee the accuracy of data understanding and expression.

[0066] Based on this, embodiments of this application provide a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can differentiate the training level of model parameters, thereby not only optimizing the training effect but also preserving the initial performance of the model. The following describes exemplary applications of the electronic devices provided in embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers.

[0067] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the model training system provided in the embodiments of this application. To support a model training application, such as... Figure 1 As shown, the model training system 100 includes: a server 200, a network 300, and a terminal 400. The terminal 400 is connected to the server 200 through the network 300. The network 300 can be a local area network (LAN), a wide area network (WAN), or a combination of both.

[0068] In some embodiments, a first language model is deployed on server 200. Server 200 trains the first language model using the model training method provided in this application embodiment to obtain a second language model. In response to terminal 400 receiving a prompt word, which prompts the second language model to perform a data analysis task, terminal 400 sends the prompt word to server 200 via network 300. Server 200 performs the data analysis task based on the prompt word and returns the data analysis results to terminal 400 via network 300, where the data analysis results are displayed. The model training method provided in this application embodiment can be executed by the terminal or server, or it can be executed collaboratively by the terminal and server (e.g., the terminal sends training samples to the server for the server to train the model).

[0069] The model training method provided in this application can be applied to various language model training scenarios, such as the following scenarios:

[0070] 1) In a translation scenario, assuming the first language model performs a Chinese-French translation task and the second language model performs a Chinese-English translation task, the model training method provided in this application is used to train the first language model using the first training samples to obtain the second language model. Since the training degree of the model parameters is differentiated, not only can the training effect be optimized, but the initial performance of the model can also be preserved. Therefore, after inputting the prompt words corresponding to the Chinese-English translation task into the second language model, the second language model can output the Chinese-English translation result. In addition, after inputting the prompt words corresponding to the Chinese-French translation task into the second language model, the second language model can also output the Chinese-French translation result.

[0071] 2) In text analysis, assuming the first language model performs a script analysis task and the second language model performs a consultation analysis task, the model training method provided in this application uses the first training sample to train the first language model to obtain the second language model. Since the training degree of the model parameters is differentiated, not only can the training effect be optimized, but the initial performance of the model can also be preserved. Therefore, after inputting the prompt words corresponding to the consultation analysis task into the second language model, the second language model can output the consultation analysis result. In addition, after inputting the prompt words corresponding to the script analysis task into the second language model, the second language model can also output the script analysis result.

[0072] In other embodiments, the model training method provided in this application can be implemented independently by a terminal or a server. The terminal or server calls the first language model and trains it using the model training method provided in this application to obtain a second language model. The terminal or server then calls the second language model to perform corresponding tasks, such as information analysis tasks, summary generation tasks, etc.

[0073] Here, the server can be a single server. In this case, the model training method provided in this application embodiment can be implemented by the same server. The server can also be a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. In this case, the model training method provided in this application embodiment can be implemented by different servers, and this application embodiment does not limit this.

[0074] The structure of the electronic device for implementing the model training method provided in the embodiments of this application will be further described below. Taking electronic device 500 as a server as an example, see... Figure 2A , Figure 2A This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2A The illustrated electronic device 500 includes at least one processor 510, a memory 540, and at least one network interface 520. The various components in the electronic device 500 are coupled together via a bus system 530. It is understood that the bus system 530 is used to implement communication between these components. In addition to a data bus, the bus system 530 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2A The general labeled all buses as Bus System 530.

[0075] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0076] The memory 540 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 540 may optionally include one or more storage devices physically located away from the processor 510.

[0077] The memory 540 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 540 described in this application embodiment is intended to include any suitable type of memory.

[0078] In some embodiments, memory 540 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0079] Operating system 541 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks;

[0080] The network communication module 542 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0081] In some embodiments, the model training apparatus provided in this application can be implemented in software. Figure 2A A model training device 543 stored in memory 540 is shown. This device can be software in the form of programs and plugins, including the following software modules: an acquisition module 5431, a first determination module 5432, a first training module 5433, a transfer module 5434, and a first correction module 5435. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. It should be noted that... Figure 2A For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the model training device 543, which may only include the acquisition module 5431, the first determination module 5432, and the first training module 5433. The functions of each module will be explained below.

[0082] The structure of the electronic device for implementing the model training method provided in the embodiments of this application will be further described below, taking electronic device 600 as a server as an example, see below. Figure 2B , Figure 2B This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2B The electronic device shown includes a model training device 643, which can be software in the form of programs and plug-ins, including the following software modules: a generation module 6431, an inference module 6432, a second determination module 6433, a second training module 6434, a third training module 6435, and a second correction module 6436. The functions of each module will be described below. It should be noted that in Figure 2B For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the model training device 643, which may only include the generation module 6431, the inference module 6432, the second determination module 6433, and the second training module 6434. The functions of each module will be explained below.

[0083] It should be noted that, Figure 2B The illustrated electronic device 600 includes a processor 610, a network interface 620, a bus system 630, a memory 640, an operating system 641, and a network communication module 642, all of which are related to... Figure 2A The corresponding modules contained therein have the same structure and the same function, and will not be described again in the embodiments of this application.

[0084] In some embodiments, the apparatus provided in this application can be implemented in hardware. For example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the model training method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0085] The model training method provided in this application will be specifically described below with reference to the exemplary application and implementation of the server provided in the embodiments of this application, taking the server as the execution subject.

[0086] See Figure 3 , Figure 3 This is a schematic diagram of the first process of the model training method provided in the embodiments of this application, which will be combined with Figure 3 The steps shown are explained.

[0087] In step 101, at least one first parameter matrix of the first language model is obtained.

[0088] It's important to note that the first language model can be an intermediate model used in the process of training a third language model based on a second task to obtain the second language model, or the first language model itself can be the third language model, trained specifically for the first task. The first task refers to the inherent capabilities of the third language model, such as script comprehension. The second task is the task that the third language model needs to learn incrementally, such as data summarization. The first parameter matrix refers to the parameter matrix used for discriminative learning during model training. This matrix can include various types such as weight matrices, bias vectors, and attention matrices. The framework's tools or libraries can be used to load the first language model, and the framework's API can be used to iterate through all layers of the first language model, extracting the corresponding first parameter matrix for each layer. Specifically, the attention matrix and weight matrix of each self-attention module in the last layer of the Transformer model can be extracted as the first parameter matrix.

[0089] In some embodiments, see Figure 4 , Figure 4 This is a schematic diagram of the second process of the model training method provided in the embodiments of this application, as shown below. Figure 4 As shown, during execution Figure 3 Before step 101 shown, you can also perform... Figure 4 Steps 106 to 110 shown are used to obtain the first language model, which will be combined with... Figure 4 The steps shown are explained.

[0090] In step 106, at least one second parameter matrix of the third language model is obtained.

[0091] Here, the third language model is trained based on the first task. The first task refers to the task capabilities inherent in the third language model itself, such as the ability to perform script comprehension tasks.

[0092] It should be noted that the second parameter matrix refers to the parameter matrix used for discriminative learning during model training. This parameter matrix can include various types such as weight matrices, bias vectors, and attention matrices. The framework's tools or libraries can be used to load a third-language model, and the framework's API can be used to iterate through all layers of the third-language model, extracting the corresponding second parameter matrix for each layer. Specifically, the attention matrix and weight matrix of each self-attention module in the last layer of the Transformer model can be extracted as the second parameter matrix.

[0093] In step 107, for each of the second parameter matrices of the third language model, a second update degree parameter is determined for each of the second parameter values ​​in the second parameter matrix based on the second difference between each second parameter value in the second parameter matrix and the matrix mean of the second parameter matrix.

[0094] It should be noted that the update degree parameter corresponding to the parameter value in the parameter matrix is ​​used to control the update magnitude of the model parameters in each iteration, thereby ensuring the differentiated processing of model parameter training and maintenance.

[0095] In some embodiments, the determination of the second update degree parameter for each second parameter value in the second parameter matrix based on the second difference between each second parameter value in the second parameter matrix and the matrix mean of the second parameter matrix in step 107 above can be achieved as follows: Calculate the standard deviation of all second parameter values ​​in the second parameter matrix to obtain the standard deviation of the second parameter matrix; determine a second threshold positively correlated with the standard deviation of the second parameter matrix; determine the parameter type of each second parameter value in the second parameter matrix based on the relationship between the second difference corresponding to each second parameter value in the second parameter matrix and the second threshold; and determine the second update degree parameter corresponding to each second parameter value in the second parameter matrix based on the parameter type of each second parameter value in the second parameter matrix. Thus, by analyzing and classifying each parameter value separately according to the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, personalized parameter updates can be achieved, ensuring differentiated processing of model parameter training and maintenance, thereby improving the overall training effect of the subsequent model.

[0096] As an example, suppose the second parameter matrix is ​​a 3×3 matrix, containing the second parameter values ​​[[0.1,0.2,0.3],[0.4,0.5,0.6],[0.7,0.8,0.9]]. The standard deviation of the second parameter matrix is ​​calculated to be 0.15. The second threshold can be 1.5 times the standard deviation, assuming a second threshold of 0.225. Calculate the second difference for each second parameter value. For example, the second difference for the second parameter value 0.5 is 0.04, and its absolute value is less than the second threshold of 0.3. Therefore, the second parameter value type can be determined as "normal activation". The second update level parameter corresponding to the second parameter value 0.5 can be 0.1. This second update level parameter can be pre-set, and each parameter type will be matched with a corresponding second update level parameter. Repeating the above process for each second parameter value in the second parameter matrix yields the second update level parameter corresponding to each second parameter value in the second parameter matrix.

[0097] In some embodiments, the method of determining the parameter type of each second parameter value in the second parameter matrix based on the relationship between the second difference corresponding to each second parameter value in the second parameter matrix and the second threshold can be implemented as follows: For each second parameter value in the second parameter matrix, the following processing is performed: when the absolute value of the second difference corresponding to the second parameter value is not greater than the second threshold, the parameter type of the second parameter value is determined to be normal activation type; when the absolute value of the second difference corresponding to the second parameter value is greater than the second threshold, and the second difference corresponding to the second parameter value is positive, the parameter type of the second parameter value is determined to be high activation type; when the absolute value of the second difference corresponding to the second parameter value is greater than the second threshold, and the second difference corresponding to the second parameter value is negative, the parameter type of the second parameter value is determined to be low activation type. Thus, by analyzing and classifying each parameter value separately based on the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, parameters can be managed more precisely, thereby achieving personalized updates for different types of parameters and improving the overall training effect of subsequent models.

[0098] It should be noted that high activation type, low activation type, and normal activation type refer to the classification of the magnitude of parameter changes during model training. High activation type indicates that the parameter changes significantly during the current task training process; low activation type indicates that the parameter changes relatively little during the current task training process; and normal activation type indicates that the parameter changes within the normal range during the current task training process.

[0099] Continuing with the above example, assuming the second difference corresponding to the second parameter value of 0.5 is 0.04, that is, the absolute value of the second difference is less than the second threshold of 0.225 and is a positive number, then the parameter type of the second parameter value of 0.5 is normal activation type; assuming the second difference corresponding to the second parameter value of 0.8 is 0.35, that is, the absolute value of the second difference is greater than the second threshold of 0.225 and is a positive number, then the parameter type of the second parameter value of 0.8 is high activation type; assuming the second difference corresponding to the second parameter value of 0.2 is -0.26, that is, the absolute value of the second difference is greater than the second threshold of 0.225 and is a negative number, then the parameter type of the second parameter value of 0.2 is low activation type.

[0100] In step 108, the first loss of the third language model is determined based on the second training samples corresponding to the second task.

[0101] In some embodiments, step 108 described above can be implemented as follows: forward inference is performed on the second training sample using the third language model to obtain a first prediction result corresponding to the second training sample; based on the first prediction result of the second training sample and the first labeled data corresponding to the second training sample, a first loss of the third language model is determined. Thus, the model loss is determined based on the prediction result of the model's forward inference of the training sample, and the model parameters are adjusted based on the loss, thereby improving the overall training effect of the subsequent model.

[0102] As an example, the second training sample corresponding to the second task can be the fourth text sample that only includes the second task, or it can include the fifth text sample that corresponds to the first task and the fourth text sample that corresponds to the second task. Therefore, the first prediction result can be the first prediction result of the fourth text sample and the first prediction result of the fifth text sample, or it can be only the first prediction result of the fourth text sample.

[0103] As an example, if the second training sample only includes the fourth text sample corresponding to the second task, the difference between the first prediction result of the fourth text sample and the labeled data corresponding to the fourth text sample can be characterized by cross-entropy loss, that is, cross-entropy loss is used as the first loss of the third language model. If the second training sample also includes the fifth text sample corresponding to the first task, since the supervision signal corresponding to the fifth text sample is also generated by the third language model to predict the fifth text sample, there is actually no loss for the fifth text sample at this time. Therefore, the first loss of the third language model is still cross-entropy loss.

[0104] In step 109, the gradient of each of the second parameter values ​​is determined based on the first loss of the third language model.

[0105] As an example, the server determines the first loss based on the cross-entropy loss function, backpropagates the first loss from the output layer of the third language model, and backpropagates the first loss layer by layer. When the first loss is passed to each layer, the gradient (i.e., the partial derivative of the loss function with respect to the parameter values ​​of each layer) is solved in combination with the passed first loss, so as to obtain the gradient of each second parameter value.

[0106] In step 110, based on the gradient of each second parameter value and the second update degree parameter of each second parameter value, each second parameter value in at least one second parameter matrix of the third language model is updated to obtain the first language model.

[0107] In some embodiments, step 110 described above can be implemented as follows: obtaining a preset learning rate for the third language model; for each second parameter value in at least one second parameter matrix of the third language model, performing the following processing: multiplying the preset learning rate of the third language model, the gradient of the second parameter value, and the second update degree parameter of the second parameter value to obtain a second multiplication result; updating the second parameter value based on the second multiplication result corresponding to the second parameter value. Thus, through the synergistic effect of the learning rate and the update degree parameter, the parameter update intensity can be dynamically balanced, effectively improving the training effect of the model.

[0108] As an example, suppose the preset learning rate of the third language model is 0.001, the gradient of the second parameter value of 0.5 is -0.3, and the second update degree parameter of the second parameter value of 0.5 is 0.9. Multiplying the three together, we can get the second multiplication result -0.00027. Then, we fuse the second multiplication result -0.00027 with the second parameter value of 0.5 to get 0.49973, which is used as the updated second parameter value, thus completing the update of the second parameter value.

[0109] Thus, by obtaining at least one second parameter matrix from the third language model, the second parameter values ​​that need to be differentiated in the third language model can be clearly identified. Subsequently, for the second parameter matrix, the difference between each second parameter value and the matrix mean of the second parameter matrix is ​​calculated to determine the second update degree parameter for each second parameter value. Then, based on the second training samples corresponding to the second task, the first loss of the third language model is determined, and the gradient of each second parameter value is determined according to the first loss. Thus, the parameters of the third language model are updated according to the gradient of each second parameter value and the second update degree parameter corresponding to the second parameter value, resulting in the first language model. By combining the gradient information of the second parameter values ​​with the second update degree parameter, the discriminative processing of model parameter training and preservation can be improved, and the important parameters of the model can be accurately corrected, thereby improving the overall training effect of the model.

[0110] In step 102, for each of the first parameter matrices, a first update degree parameter is determined for each first parameter value in the first parameter matrix based on the first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix.

[0111] It should be noted that the specific implementation of step 102 is similar to that of step 107, and you can refer to the implementation of step 107 for details.

[0112] In some embodiments, step 102 described above can be implemented as follows: Calculate the standard deviation of all first parameter values ​​in the first parameter matrix to obtain the standard deviation of the first parameter matrix; determine a first threshold positively correlated with the standard deviation; determine the parameter type of each first parameter value in the first parameter matrix based on the relationship between the first difference corresponding to each first parameter value in the first parameter matrix and the first threshold; and determine a first update degree parameter corresponding to each first parameter value in the first parameter matrix based on the parameter type of each first parameter value in the first parameter matrix. Thus, by analyzing and classifying each parameter value individually based on the standard deviation of the parameter matrix and the differences between each parameter value in the parameter matrix, personalized parameter updates can be achieved, ensuring differentiated processing of model parameter training and maintenance, thereby improving the overall training effect of the subsequent model.

[0113] As an example, suppose the first parameter matrix is ​​a 2×2 matrix, containing first parameter values ​​of [[0.1,0.2],[0.3,0.4]]. The standard deviation of the first parameter matrix is ​​calculated to be 0.1118. Assuming the first threshold is twice the standard deviation, the first threshold is 0.2236. Calculate the first difference for each first parameter value. For example, the first difference for the first parameter value 0.3 is 0.1, which is less than the first threshold of 0.2236. Therefore, the first parameter value type can be determined as "normal activation". Different parameter types correspond to different first update degree parameters. Here, the first update degree parameter for the normal activation parameter type is 0.1, and the first update degree parameter for the first parameter value 0.3 is also 0.1. Repeating the above process for each first parameter value in the first parameter matrix yields the first update degree parameter for each first parameter value in the first parameter matrix.

[0114] In some embodiments, the method of determining the parameter type of each first parameter value in the first parameter matrix based on the relationship between the first difference corresponding to each first parameter value in the first parameter matrix and the first threshold can be implemented as follows: For each first parameter value in the first parameter matrix, the following processing is performed: when the absolute value of the first difference corresponding to the first parameter value is not greater than the first threshold, the parameter type of the first parameter value is determined to be a normal activation type; when the absolute value of the first difference corresponding to the first parameter value is greater than the first threshold and the first difference corresponding to the first parameter value is positive, the parameter type of the first parameter value is determined to be a high activation type; when the absolute value of the first difference corresponding to the first parameter value is greater than the first threshold and the first difference corresponding to the first parameter value is negative, the parameter type of the first parameter value is determined to be a low activation type. Thus, by analyzing and classifying each parameter value separately based on the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, parameters can be managed more precisely, thereby achieving personalized updates for different types of parameters and improving the overall training effect of subsequent models.

[0115] It should be noted that the specific implementation of determining the parameter type of each first parameter value in the first parameter matrix based on the relationship between the first difference corresponding to each first parameter value in the first parameter matrix and the first threshold in step 102 is similar to the implementation of determining the parameter type of each second parameter value in the second parameter matrix based on the relationship between the second difference corresponding to each second parameter value in the second parameter matrix and the second threshold in step 107. For details, please refer to the implementation of determining the parameter type of each second parameter value in the second parameter matrix based on the relationship between the second difference corresponding to each second parameter value in the second parameter matrix and the second threshold in step 107, which will not be repeated here.

[0116] In step 103, the first loss of the first language model is determined based on the first training samples.

[0117] It should be noted that the first training sample can correspond to both the first and second tasks, meaning the first training sample includes both training samples corresponding to the first task and training samples corresponding to the second task; or it can correspond only to the second task, meaning the first training sample only includes training samples corresponding to the second task.

[0118] In some embodiments, see Figure 5 See Figure 5 , Figure 5 This is a schematic diagram of the third process of the model training method provided in the embodiments of this application, as shown below. Figure 5 As shown, Figure 3 Step 103 shown can be achieved through... Figure 5 Steps 1031 to 1032 shown are implemented, and will be combined with Figure 5 The steps shown are explained.

[0119] In step 1031, the first training sample is forward-inferred using the first language model to obtain a first prediction result corresponding to the first training sample.

[0120] In some embodiments, the first training sample includes a first text sample corresponding to a first task and a second text sample corresponding to a second task and associated with the first labeled data. Step 1031 described above can be implemented as follows: the first language model performs forward inference on the first text sample corresponding to the first task to obtain a first prediction result corresponding to the first text sample; the first language model performs forward inference on the second text sample corresponding to the second task to obtain a first prediction result corresponding to the second text sample. Thus, by using a single language model to process two different tasks simultaneously, it can be ensured that the language model trained subsequently has the ability to analyze and process both tasks, reducing the hardware cost of deploying the model.

[0121] As an example, the first task is an intent classification task, and the second task is a sentiment classification task. The first language model performs the intent classification task based on the first text sample and outputs the first prediction result of the first text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the intent classification task (e.g., "The customer wants to inquire about the way to change the ticket"). The first language model performs the sentiment classification task based on the second text sample and outputs the first prediction result of the second text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the sentiment classification task (e.g., "The customer is currently angry").

[0122] In some embodiments, the first training sample includes a second text sample corresponding to the second task and associated with first labeled data; step 1031 described above can be implemented by performing forward inference on the second text sample corresponding to the second task using the first language model to obtain a first prediction result corresponding to the second text sample. Thus, by directly performing forward inference on the second task using the first language model, the subsequently obtained model can focus on the second task, achieving the best performance in the execution of the second task.

[0123] As an example, suppose the second task is a sentiment classification task. By using the first language model to perform forward inference on the second text sample (e.g., "The customer said I don't like rainy days, please analyze the customer's mood at this moment"), we can obtain the first prediction result for the corresponding second text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the sentiment classification task (e.g., "The customer's mood at this moment may be depressed").

[0124] In step 1032, the first loss of the first language model is determined based on the first prediction result and the first labeled data corresponding to the first training sample.

[0125] As an example, the first training sample may be a second text sample that includes only the second task, or it may include both the first text sample corresponding to the first task and the second text sample corresponding to the second task. Therefore, the first prediction result may be the first prediction result of the first text sample and the first prediction result of the second text sample, or it may be only the first prediction result of the fourth text sample.

[0126] In some embodiments, the first training sample includes a first text sample corresponding to a first task and a second text sample corresponding to a second task and associated with the first labeled data; see also Figure 6 See Figure 6 , Figure 6 This is a schematic diagram of the fourth process of the model training method provided in the embodiments of this application, as shown below. Figure 6 As shown, Figure 5 Step 1032 shown can be achieved through... Figure 6 Steps 10321 to 10324 shown are implemented by combining Figure 5 The steps shown are explained.

[0127] In step 10321, the first text sample is subjected to forward inference corresponding to the first task through the third language model to obtain the second prediction result corresponding to the first text sample.

[0128] It should be noted that the third language model is trained based on the first task; that is, the third language model is a vertical domain model specifically trained for the first task.

[0129] As an example, the first text sample corresponding to the first task is input into the third language model for prediction processing, and a second prediction result corresponding to the first text sample can be obtained, which can be used as a supervision signal.

[0130] In step 10322, based on the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample, the second loss corresponding to the first text sample is determined.

[0131] It should be noted that the first prediction result corresponding to the first text sample is obtained by the first language model predicting the first text sample, and the second prediction result corresponding to the first text sample is obtained by the third language model predicting the first text sample. The first language model is obtained by the third language model based on the training process of the second task. By comparing the difference between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample, it can be determined whether the first task capability of the first language model has changed compared to the third language model.

[0132] In some embodiments, step 10322 described above can be implemented as follows: determining a first ratio between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample; performing logarithmic calculation on the first ratio to obtain a first logarithmic calculation result; multiplying the first logarithmic calculation result with the first prediction result corresponding to the first text sample to obtain a second loss corresponding to the first text sample. Thus, the prediction result of the third language model is used to supervise the prediction ability of the first language model for the first task, ensuring the prediction ability of the first language model for the first task is maintained while the first language model is being trained for the second task.

[0133] As an example, assuming the first prediction result for the first text sample X is P(X) and the second prediction result for the first text sample X is Q(X), we can obtain the first ratio between the first prediction result P(X) and the second prediction result Q(X) for the first text sample X: Subsequently, the first ratio By performing logarithmic calculations, we can obtain the result of the first logarithm. Finally, the result of the first logarithm calculation Multiplying the first prediction result P(X) corresponding to the first text sample X with the second loss Loss2 corresponding to the first text sample yields the second loss Loss2.

[0134] In step 10323, based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample, the third loss corresponding to the second text sample is determined.

[0135] It should be noted that the third loss can be determined by calculating a specific loss function, such as the cross-entropy loss function, the mean squared error loss function, or the mean absolute error loss function.

[0136] In some embodiments, the first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position. Step 10323 described above can be implemented as follows: for each character position, based on the first annotation data, obtain the target candidate character corresponding to the character position from among multiple candidate characters; based on the first prediction result of the character position, determine the first prediction probability of the target candidate character corresponding to the character position; perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result; fuse the second logarithmic calculation results corresponding to multiple character positions to obtain a third loss corresponding to the second text sample. Thus, by constructing a loss function through logarithmic fusion of character probabilities, long-distance dependencies can be effectively captured, thereby improving the training effect of the model.

[0137] As an example, taking the cross-entropy loss function as an example, the third loss Loss3 can be calculated using formula (1).

[0138]

[0139] Among them, y ij This represents the label (usually 1) of the target candidate character corresponding to the j-th character position in the i-th second text sample. is the first predicted probability of the target candidate character corresponding to the j-th character position in the i-th second text sample, n is the number of second text samples, and m is the number of characters corresponding to the answer of the second text sample.

[0140] In step 10324, the second loss corresponding to the first text sample and the third loss corresponding to the second text sample are fused to obtain the first loss of the first language model.

[0141] Thus, through steps 10321 to 10324, the model can learn the second task without forgetting the first task, thereby improving the training effect of the model and enhancing its generalization ability in subsequent application stages.

[0142] It should be noted that the second loss corresponding to the first text sample and the third loss corresponding to the second text sample can be fused using a linear weighted method; a constrained fusion method (i.e., adding a certain loss as a regularization term to the main loss) can also be used; and a phased scheduling method (i.e., using the second loss in phase 1 and the third loss in phase 2 according to the training progress) can also be used. The specific method can be determined according to actual needs, and no specific limitations are made here.

[0143] In some embodiments, step 10324 described above can be implemented as follows: determining a second ratio between the first number of the first text samples and the second number of the second text samples; obtaining the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples, wherein the ratio between the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples is the same as the second ratio; and performing a weighted summation of the second loss corresponding to the first text samples and the third loss corresponding to the second text samples based on the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples to obtain the first loss of the first language model. In this way, by determining the loss weights through the number of samples, the proportion of different training sample numbers is reflected in the loss, thereby improving the training effect of the model.

[0144] As an example, suppose the second ratio between the first number of the first text sample and the second number of the second text sample is 1. It can be seen that the ratio between the weight of the second loss corresponding to the first text sample and the weight of the third loss corresponding to the second text sample is also... This allows us to obtain the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample (for example, the weight of the second loss is...). The weight of the third loss is 1); then, based on the weights of the second loss corresponding to the first text sample and the third loss corresponding to the second text sample, the second loss Loss2 and the third loss Loss3 corresponding to the second text sample are weighted and summed to obtain the first loss Loss1 of the first language model.

[0145] In summary, by using a first language model to predict training samples, the prediction results for the corresponding training samples are obtained. Based on the prediction results and the labeled data corresponding to the training samples, the loss is calculated, forming a unidirectional gradient propagation path, thereby achieving model optimization and improving the performance of the trained model.

[0146] In some embodiments, the first training sample includes a second text sample corresponding to the second task and associated with first labeled data; step 1032 described above can be implemented as follows: based on the first prediction result corresponding to the second text sample and the first labeled data corresponding to the second text sample, a third loss corresponding to the second text sample is determined; based on the third loss, a first loss of the first language model is determined (i.e., the third loss is directly used as the first loss). Thus, the first language model is fine-tuned on the second task through the third loss, thereby achieving transfer learning, and the third loss can be directly backpropagated to all layers of the first language model, achieving end-to-end optimization.

[0147] It should be noted that the implementation of determining the third loss corresponding to the second text sample based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample is similar to the implementation of step 10323 above. For details, please refer to the implementation of step 10323.

[0148] In some embodiments, the first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position. The above steps, based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample, determine the third loss corresponding to the second text sample. This can be achieved as follows: For each character position, based on the first annotation data, obtain the target candidate character corresponding to the character position from multiple candidate characters; based on the first prediction result of the character position, determine the first prediction probability of the target candidate character corresponding to the character position; perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result; fuse the second logarithmic calculation results corresponding to multiple character positions to obtain the third loss corresponding to the second text sample. Thus, by constructing the loss function through logarithmic fusion of character probabilities, long-distance dependencies can be effectively captured, thereby improving the training effect of the model.

[0149] As an example, taking the cross-entropy loss function as an example, the third loss Loss3 can be calculated using formula (2).

[0150]

[0151] Among them, y ij This represents the label (usually 1) of the target candidate character corresponding to the j-th character position in the i-th second text sample. is the first predicted probability of the target candidate character corresponding to the j-th character position in the i-th second text sample, n is the number of second text samples, and m is the number of characters corresponding to the answer of the second text sample.

[0152] In step 104, the gradient of each of the first parameter values ​​is determined based on the first loss.

[0153] It should be noted that the specific implementation of step 104 is similar to the implementation of step 109 above, and you can refer to the implementation of step 109 for details.

[0154] As an example, the server determines the first loss based on the cross-entropy loss function, backpropagates the first loss from the output layer of the first language model, and backpropagates the first loss layer by layer. When the first loss is passed to each layer, the gradient (i.e., the partial derivative of the loss function with respect to the parameter values ​​of each layer) is solved in combination with the passed first loss, so as to obtain the gradient of each first parameter value.

[0155] In step 105, based on the gradient of each first parameter value and the first update degree parameter of each first parameter value, each first parameter value in at least one first parameter matrix of the first language model is updated to obtain a second language model.

[0156] In some embodiments, the step 105 described above, which updates each first parameter value in at least one first parameter matrix of the first language model based on the gradient of each first parameter value and the first update degree parameter of each first parameter value, can be implemented as follows: Obtain a preset learning rate for the first language model; for each first parameter value in at least one first parameter matrix of the first language model, perform the following processing: multiply the preset learning rate, the gradient of the first parameter value, and the first update degree parameter of the first parameter value to obtain a first multiplication result; update the first parameter value based on the first multiplication result corresponding to the first parameter value. Thus, through the synergistic effect of the learning rate and the update degree parameter, the parameter update intensity can be dynamically balanced, effectively improving the training effect of the model.

[0157] It should be noted that the specific implementation of step 105 is similar to the implementation of step 110 above, and you can refer to the implementation of step 110 for details.

[0158] As an example, suppose the preset learning rate of the first language model is 0.001, the gradient of the first parameter value of 0.5 is -0.3, and the first update degree parameter of the first parameter value of 0.5 is 0.9. Multiplying the three together, we can get the first multiplication result -0.00027. Then, we fuse the first multiplication result -0.00027 with the first parameter value of 0.5 to get 0.49973, which is used as the updated first parameter value, thus completing the update of the first parameter value.

[0159] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the fifth process of the model training method provided in the embodiments of this application, as shown below. Figure 7 As shown, during execution Figure 3 After step 105 shown, the following steps can also be performed: Figure 7 Steps 111 to 112 shown are implemented by combining Figure 7 The steps shown are explained.

[0160] In step 111, the first text is forward-reasoned using the second language model to obtain a third prediction result corresponding to the first text.

[0161] As an example, suppose in a data analysis task, a first text T is input into a second language model for predictive analysis, and a third prediction result Q corresponding to the first text T is obtained.

[0162] In step 112, the numerical content in the third prediction result is corrected to obtain the fourth prediction result corresponding to the first text.

[0163] Following the example above, the numerical content contained in the third prediction result is determined; the numerical content is obtained from online data, and the numerical content in the third prediction result is replaced with the numerical content from the obtained online data, thereby obtaining the fourth prediction result corresponding to the first text T.

[0164] In conclusion, conducting online data accuracy verification on the prediction results of the trained model can effectively improve the correctness of data analysis tasks.

[0165] See Figure 8 , Figure 8 This is a schematic diagram of the sixth process of the model training method provided in the embodiments of this application, which will be combined with Figure 8 The steps shown are explained.

[0166] In step 201, a first text sample corresponding to the first task is generated.

[0167] It should be noted that text samples for the corresponding task can be manually constructed according to task requirements; alternatively, pre-set task-related text templates can be used to dynamically generate text samples by filling in keywords; existing data can be used to generate text samples by decomposing and recombining it; or a large model can be used to dynamically generate text data for the corresponding task according to requirements. The specific method can be determined based on the actual situation, and no specific limitations are made here.

[0168] As an example, taking large model generation as an example, corresponding task instructions can be constructed, such as "generate 20 texts containing financial keywords as prompt words". The task instructions are input into the large language model to obtain relevant text data (prompt words) that match the task instructions.

[0169] In some embodiments, before performing step 201, the following processing may also be performed: obtaining at least one third parameter matrix of the first language model; for each third parameter matrix, determining a third update degree parameter for each third parameter value in the third parameter matrix based on the third difference between each third parameter value in the third parameter matrix and the matrix mean of the third parameter matrix. In this way, the trainable parameters of the language model can be clearly defined, and it can be determined that each third parameter value can adopt a corresponding update degree parameter, thereby controlling the learning degree of different third parameter values ​​and helping to maintain the initial performance of the first language model.

[0170] It's important to note that the third parameter matrix refers to the parameter matrix used for discriminative learning during model training. This matrix can include various types such as weight matrices, bias vectors, and attention matrices. The framework's tools or libraries can be used to load a third-language model, and the framework's API can be used to iterate through all layers of the model, extracting the corresponding third parameter matrix for each layer. Specifically, the attention matrix and weight matrix of each self-attention module in the last layer of the Transformer model can be extracted as the third parameter matrix. The update degree parameters corresponding to the parameter values ​​in the parameter matrix control the update magnitude of the model parameters in each iteration, thereby ensuring discriminative processing for model parameter training and retention.

[0171] In some embodiments, the aforementioned determination of the third update degree parameter for each third parameter value in the third parameter matrix based on the third difference between each third parameter value in the third parameter matrix and the matrix mean of the third parameter matrix can be achieved as follows: Calculate the standard deviation of all third parameter values ​​in the third parameter matrix to obtain the standard deviation of the third parameter matrix; determine a third threshold positively correlated with the standard deviation of the third parameter matrix; determine the parameter type of each third parameter value in the third parameter matrix based on the relationship between the third difference corresponding to each third parameter value in the third parameter matrix and the third threshold; and determine the third update degree parameter corresponding to each third parameter value in the third parameter matrix based on the parameter type of each third parameter value in the third parameter matrix. Thus, by analyzing and classifying each parameter value individually based on the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, personalized parameter updates can be achieved, ensuring differentiated processing of model parameter training and maintenance, thereby improving the overall training effect of the subsequent model.

[0172] As an example, suppose the third parameter matrix is ​​a 3×3 matrix, containing the third parameter values ​​[[0.1,0.2,0.3],[0.4,0.5,0.6],[0.7,0.8,0.9]]. The standard deviation of the third parameter matrix is ​​calculated to be 0.15. The third threshold can be 1.5 times the standard deviation, assuming a third threshold of 0.225. Calculate the third difference for each third parameter value. For example, the third difference for a third parameter value of 0.5 is 0.04, and its absolute value is less than the third threshold of 0.3. Therefore, the third parameter value type can be determined as "normal activation". The third update level parameter corresponding to the third parameter value of 0.5 can be 0.1. This third update level parameter can be pre-set, and each parameter type will be matched with a corresponding third update level parameter. Repeating the above process for each third parameter value in the third parameter matrix yields the third update level parameter corresponding to each third parameter value in the third parameter matrix.

[0173] In some embodiments, the method of determining the parameter type of each third parameter value in the third parameter matrix based on the relationship between the third difference corresponding to each third parameter value and the third threshold can be implemented as follows: For each third parameter value in the third parameter matrix, the following processing is performed: when the absolute value of the third difference corresponding to the third parameter value is not greater than the third threshold, the parameter type of the third parameter value is determined to be normal activation type; when the absolute value of the third difference corresponding to the third parameter value is greater than the third threshold, and the third difference corresponding to the third parameter value is positive, the parameter type of the third parameter value is determined to be high activation type; when the absolute value of the third difference corresponding to the third parameter value is greater than the third threshold, and the third difference corresponding to the third parameter value is negative, the parameter type of the third parameter value is determined to be low activation type. Thus, by analyzing and classifying each parameter value separately based on the standard deviation of the parameter matrix and the differences between each parameter value in the parameter matrix, parameters can be managed more precisely, thereby achieving personalized updates for different types of parameters and improving the overall training effect of subsequent models.

[0174] It should be noted that high activation type, low activation type, and normal activation type refer to the classification of the magnitude of parameter changes during model training. High activation type indicates that the parameter changes significantly during the current task training process; low activation type indicates that the parameter changes relatively little during the current task training process; and normal activation type indicates that the parameter changes within the normal range during the current task training process.

[0175] Continuing with the above example, assuming the third difference corresponding to the third parameter value of 0.5 is 0.04, that is, the absolute value of the third difference is less than the third threshold of 0.225 and is a positive number, then the parameter type of the third parameter value of 0.5 is normal activation type; assuming the third difference corresponding to the third parameter value of 0.8 is 0.35, that is, the absolute value of the third difference is greater than the third threshold of 0.225 and is a positive number, then the parameter type of the third parameter value of 0.8 is high activation type; assuming the third difference corresponding to the third parameter value of 0.2 is -0.26, that is, the absolute value of the third difference is greater than the third threshold of 0.225 and is a negative number, then the parameter type of the third parameter value of 0.2 is low activation type.

[0176] In some embodiments, see Figure 9 , Figure 9 This is a schematic diagram of the seventh process of the model training method provided in the embodiments of this application, as shown below. Figure 9 As shown, during execution Figure 8 Before step 201 shown, the following steps can also be performed: Figure 9 Steps 209 to 212 shown are implemented by combining Figure 9 The steps shown are explained.

[0177] In step 209, a third text sample corresponding to the second task and associated with the second annotation data is obtained.

[0178] As an example, assuming the second task is a sentiment classification task, the third text sample could be "The customer says I like sunny days, please analyze the customer's mood at this moment," and the second labeled data associated with the third text sample could be "The customer's mood at this moment is happy."

[0179] In some embodiments, before performing step 209, the following processing may also be performed: obtaining at least one fourth parameter matrix of the third language model; for each of the fourth parameter matrices of the third language model, determining a fourth update degree parameter for each fourth parameter value in the fourth parameter matrix based on the fourth difference between each fourth parameter value in the fourth parameter matrix and the matrix mean of the fourth parameter matrix. In this way, the fourth parameter values ​​in the third language model that require differentiated training can be clearly identified, and the fourth update degree parameter for each fourth parameter value can be determined through the difference between each fourth parameter value and the fourth parameter matrix. This can improve the discriminative processing of subsequent model parameter training and maintenance, thereby improving the overall training effect of the model.

[0180] It's important to note that the fourth parameter matrix refers to the parameter matrix used for discriminative learning during model training. This matrix can include various types such as weight matrices, bias vectors, and attention matrices. The framework's tools or libraries can be used to load a third-language model, and the framework's API can be used to iterate through all layers of the model, extracting the corresponding fourth parameter matrix for each layer. Specifically, the attention matrix and weight matrix of each self-attention module in the last layer of the Transformer model can be extracted as the fourth parameter matrix. The update degree parameters corresponding to the parameter values ​​in the parameter matrix control the update magnitude of the model parameters in each iteration, thereby ensuring discriminative processing for model parameter training and retention.

[0181] In some embodiments, the aforementioned determination of the fourth update degree parameter for each fourth parameter value in the fourth parameter matrix based on the fourth difference between each fourth parameter value and the matrix mean of the fourth parameter matrix can be achieved as follows: Calculate the standard deviation of all fourth parameter values ​​in the fourth parameter matrix to obtain the standard deviation of the fourth parameter matrix; determine a fourth threshold positively correlated with the standard deviation of the fourth parameter matrix; determine the parameter type of each fourth parameter value in the fourth parameter matrix based on the relationship between the fourth difference corresponding to each fourth parameter value and the fourth threshold; and determine the fourth update degree parameter corresponding to each fourth parameter value in the fourth parameter matrix based on the parameter type of each fourth parameter value. Thus, by analyzing and classifying each parameter value individually based on the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, personalized parameter updates can be achieved, ensuring differentiated processing of model parameter training and maintenance, thereby improving the overall training effect of the subsequent model.

[0182] As an example, suppose the fourth parameter matrix is ​​a 3×3 matrix, containing the fourth parameter values ​​[[0.1,0.2,0.3],[0.4,0.5,0.6],[0.7,0.8,0.9]]. The standard deviation of the fourth parameter matrix is ​​calculated to be 0.15. The fourth threshold can be 1.5 times the standard deviation; assuming the fourth threshold is 0.225, calculate the fourth difference for each fourth parameter value. For example, the fourth difference for a fourth parameter value of 0.5 is 0.04, and its absolute value is less than the fourth threshold of 0.3. Therefore, the fourth parameter value type can be determined as "normal activation". The fourth update level parameter corresponding to the fourth parameter value of 0.5 can be 0.1. This fourth update level parameter can be pre-set, and each parameter type will be matched with a corresponding fourth update level parameter. Repeating the above process for each fourth parameter value in the fourth parameter matrix yields the fourth update level parameter corresponding to each fourth parameter value in the fourth parameter matrix.

[0183] In some embodiments, the method of determining the parameter type of each fourth parameter value in the fourth parameter matrix based on the relationship between the fourth difference corresponding to each fourth parameter value and the fourth threshold can be implemented as follows: For each fourth parameter value in the fourth parameter matrix, the following processing is performed: when the absolute value of the fourth difference corresponding to the fourth parameter value is not greater than the fourth threshold, the parameter type of the fourth parameter value is determined to be normal activation type; when the absolute value of the fourth difference corresponding to the fourth parameter value is greater than the fourth threshold, and the fourth difference corresponding to the fourth parameter value is positive, the parameter type of the fourth parameter value is determined to be high activation type; when the absolute value of the fourth difference corresponding to the fourth parameter value is greater than the fourth threshold, and the fourth difference corresponding to the fourth parameter value is negative, the parameter type of the fourth parameter value is determined to be low activation type. Thus, by analyzing and classifying each parameter value separately based on the standard deviation of the parameter matrix and the difference between each parameter value in the parameter matrix, parameters can be managed more precisely, thereby achieving personalized updates for different types of parameters and improving the overall training effect of subsequent models.

[0184] It should be noted that high activation type, low activation type, and normal activation type refer to the classification of the magnitude of parameter changes during model training. High activation type indicates that the parameter changes significantly during the current task training process; low activation type indicates that the parameter changes relatively little during the current task training process; and normal activation type indicates that the parameter changes within the normal range during the current task training process.

[0185] Continuing with the above example, assuming the fourth difference corresponding to the fourth parameter value of 0.5 is 0.04, that is, the absolute value of the fourth difference is less than the fourth threshold of 0.225 and is a positive number, then the parameter type of the fourth parameter value of 0.5 is normal activation type; assuming the fourth difference corresponding to the fourth parameter value of 0.8 is 0.35, that is, the absolute value of the fourth difference is greater than the fourth threshold of 0.225 and is a positive number, then the parameter type of the fourth parameter value of 0.8 is high activation type; assuming the fourth difference corresponding to the fourth parameter value of 0.2 is -0.26, that is, the absolute value of the fourth difference is greater than the fourth threshold of 0.225 and is a negative number, then the parameter type of the fourth parameter value of 0.2 is low activation type.

[0186] In step 210, the third language model is used to perform forward inference on the third text sample corresponding to the second task to obtain a first prediction result corresponding to the third text sample.

[0187] It should be noted that the specific implementation of step 210 is similar to the implementation of step 1031 above, which uses the first language model to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample. Specifically, you can refer to the implementation of step 1031 above, which uses the first language model to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample. The only difference is that the model used is a third language model instead of a first language model, and the data processed is a third text sample instead of a second text sample.

[0188] As an example, suppose the second task is a sentiment classification task. The third language model performs forward inference on the third text sample (e.g., "The customer says I don't like rainy days, please analyze the customer's mood at this moment") to obtain the first prediction result of the corresponding third text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the sentiment classification task (e.g., "The customer's mood at this moment may be depressed").

[0189] In step 211, based on the first prediction result corresponding to the third text sample and the second annotation data, the third loss corresponding to the third text sample is determined.

[0190] It should be noted that the third loss can be determined by calculating a specific loss function, such as the cross-entropy loss function, the mean squared error loss function, or the mean absolute error loss function.

[0191] In some embodiments, the first prediction result corresponding to the third text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position. Step 211 described above can be implemented as follows: for each character position, based on the first annotation data, obtain the target candidate character corresponding to the character position from among multiple candidate characters; based on the first prediction result of the character position, determine the first prediction probability of the target candidate character corresponding to the character position; perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a fourth logarithmic calculation result; fuse the fourth logarithmic calculation results corresponding to multiple character positions to obtain the third loss corresponding to the third text sample. Thus, by constructing a loss function through logarithmic fusion of character probabilities, long-distance dependencies can be effectively captured, thereby improving the training effect of the model.

[0192] It should be noted that the specific implementation of step 211 is similar to the implementation of step 10323 above. For details, please refer to the implementation of step 10323 above.

[0193] As an example, taking the cross-entropy loss function as an example, the third loss Loss3 can be calculated using formula (3).

[0194]

[0195] Among them, y ij This represents the label (usually 1) of the target candidate character corresponding to the j-th character position in the i-th third text sample. is the first predicted probability of the target candidate character corresponding to the j-th character position in the i-th third text sample, n is the number of third text samples, and m is the number of characters corresponding to the answer of the third text sample.

[0196] In step 212, the third language model is updated based on the third loss corresponding to the third text sample to obtain the first language model.

[0197] In practical applications, the third loss based on the third text sample is backpropagated in the third language model, and the model parameters of the third language model are updated during the propagation process.

[0198] Here's an explanation of backpropagation: The third text sample is input into the input layer of the third language model, passes through the hidden layer, and finally reaches the output layer to output the result, obtaining the first prediction result of the third text sample. This is the forward propagation process of the third language model. Since there is an error between the first prediction result of the third text sample output by the third language model and the second labeled data associated with the third text sample, a third loss is calculated between the first prediction result of the third text sample and the second labeled data associated with the third text sample. This third loss is then backpropagated from the output layer to the hidden layer until it reaches the input layer. During the backpropagation process, the values ​​of the model parameters are adjusted according to the third loss. This process is iterated until convergence.

[0199] The server backpropagates the third loss from the output layer of the third language model, layer by layer. When the third loss reaches each layer, it combines the propagated third loss to solve for the gradient (that is, the partial derivative of the third loss with respect to the parameters of each layer), and updates the corresponding gradient values ​​of the parameters of each layer.

[0200] In some embodiments, step 212 described above can also be implemented as follows: determining the gradient of each of the fourth parameter values ​​based on the third loss of the third language model; updating each of the fourth parameter values ​​in at least one fourth parameter matrix of the third language model based on the gradient of each of the fourth parameter values ​​and the fourth update degree parameter corresponding to each of the fourth parameter values, to obtain the first language model. Thus, by determining the gradient of each of the fourth parameter values ​​based on the third loss, and updating the parameters of the third language model based on the gradient of each of the fourth parameter values ​​and the corresponding fourth update degree parameter, the first language model is obtained. By combining the gradient information of the fourth parameter values ​​with the fourth update degree parameter, the discriminative processing of model parameter training and maintenance can be improved, and important parameters of the model can be accurately corrected, thereby improving the overall training effect of the model.

[0201] As an example, the server determines the third loss based on the cross-entropy loss function, backpropagates the third loss from the output layer of the third language model, and backpropagates the third loss layer by layer. When the third loss is passed to each layer, the gradient (i.e., the partial derivative of the loss function with respect to the parameter values ​​of each layer) is solved in combination with the passed third loss, so as to obtain the gradient of each fourth parameter value. The parameters of each layer are updated with the corresponding gradient values, thus obtaining the first language model.

[0202] In some embodiments, updating each fourth parameter value in at least one fourth parameter matrix of the third language model based on the gradient of each fourth parameter value and the fourth update degree parameter of each fourth parameter value to obtain the first language model can be achieved in the following way: obtaining a preset learning rate of the third language model; for each fourth parameter value in at least one fourth parameter matrix of the third language model, performing the following processing respectively: multiplying the preset learning rate of the third language model, the gradient of the fourth parameter value, and the fourth update degree parameter of the fourth parameter value to obtain a multiplication result; updating the fourth parameter value based on the multiplication result corresponding to the fourth parameter value. In this way, through the synergistic effect of the learning rate and the update degree parameter, the parameter update intensity can be dynamically balanced, which can effectively improve the training effect of the model.

[0203] As an example, suppose the preset learning rate of the third language model is 0.001, the gradient of the fourth parameter value of 0.5 is -0.3, and the fourth update level parameter of the fourth parameter value of 0.5 is 0.9. Multiplying the three together, we get a multiplication result of -0.00027. Then, we fuse the multiplication result of -0.00027 with the fourth parameter value of 0.5 to get 0.49973, which is used as the updated fourth parameter value, thus completing the update of the fourth parameter value.

[0204] In summary, by training using only third text samples that are strongly related to the second task, interference from irrelevant tasks can be avoided, allowing model parameter updates to focus more on the specific patterns of the second task. Furthermore, by calculating the third loss using the second labeled data, the model's error on the second task can be directly optimized, making it more reliable than unsupervised or weakly supervised methods.

[0205] In step 202, a second text sample corresponding to the second task and associated with the first labeled data is obtained.

[0206] It should be noted that the specific implementation of step 202 is similar to that of step 209 above, and can be referred to the implementation of step 209 above.

[0207] As an example, assuming the second task is a sentiment classification task, the second text sample could be "The customer says I like sunny days, please analyze the customer's mood at this moment," and the second labeled data associated with the second text sample could be "The customer's mood at this moment is happy."

[0208] In step 203, the first text sample is subjected to forward inference corresponding to the first task through the third language model to obtain the second prediction result corresponding to the first text sample.

[0209] Here, the third language model is a model trained based on the first task, that is, the third language model is a vertical domain model specifically trained for the first task.

[0210] It should be noted that the specific implementation of step 203 is similar to the implementation of step 10321 above. For details, please refer to the implementation of step 10321 above.

[0211] As an example, the first text sample corresponding to the first task is input into the third language model for prediction processing, and a second prediction result corresponding to the first text sample can be obtained, which can be used as a supervision signal.

[0212] In step 204, the first text sample is subjected to forward inference corresponding to the first task through the first language model to obtain the first prediction result corresponding to the first text sample.

[0213] It should be noted that the specific implementation of step 204 is similar to the implementation of step 1031 above, which uses the first language model to perform forward inference on the first text sample corresponding to the first task to obtain the first prediction result corresponding to the first text sample. For details, please refer to the implementation of step 1031 above, which uses the first language model to perform forward inference on the first text sample corresponding to the first task to obtain the first prediction result corresponding to the first text sample.

[0214] As an example, the first task is an intent classification task. The first language model performs the intent classification task based on the first text sample and outputs the first prediction result of the first text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the intent classification task (e.g., "The customer wants to inquire about the way to change the ticket at this time").

[0215] In step 205, forward inference corresponding to the second task is performed on the second text sample to obtain a first prediction result corresponding to the second text sample.

[0216] It should be noted that the specific implementation of step 205 is similar to the implementation of step 1031 above, which uses the first language model to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample. For details, please refer to the implementation of step 1031 above, which uses the first language model to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample.

[0217] As an example, the second task is a sentiment classification task. The first language model performs the sentiment classification task based on the second text sample and outputs the first prediction result of the second text sample. Here, the first prediction result can be the predicted probability of each character in the predicted response of the sentiment classification task (e.g., "The customer is currently angry").

[0218] In step 206, a second loss corresponding to the first text sample is determined based on the first prediction result and the second prediction result corresponding to the first text sample.

[0219] It should be noted that the specific implementation of step 206 is similar to the implementation of step 10322 above. For details, please refer to the implementation of step 10322 above.

[0220] In some embodiments, step 206 described above can be implemented as follows: determining a first ratio between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample; performing logarithmic calculation on the first ratio to obtain a first logarithmic calculation result; multiplying the first logarithmic calculation result with the first prediction result corresponding to the first text sample to obtain a second loss corresponding to the first text sample. In this way, the prediction result of the third language model is used to supervise the prediction ability of the first language model for the first task, ensuring the prediction ability of the first language model for the first task is maintained while the first language model is being trained for the second task.

[0221] As an example, assuming the first prediction result for the first text sample X is P(X) and the second prediction result for the first text sample X is Q(X), we can obtain the first ratio between the first prediction result P(X) and the second prediction result Q(X) for the first text sample X: Subsequently, the first ratio By performing logarithmic calculations, we can obtain the result of the first logarithm. Finally, the result of the first logarithm calculation Multiplying the first prediction result P(X) corresponding to the first text sample X with the second loss Loss2 corresponding to the first text sample yields the second loss Loss2.

[0222] In step 207, a third loss is determined based on the first prediction result corresponding to the second text sample and the first annotation data.

[0223] It should be noted that the method for determining the third loss can be calculated using a specific loss function, such as the cross-entropy loss function, the mean squared error loss function, the mean absolute error loss function, etc. The specific implementation of step 207 is similar to the implementation of step 10323 above, and you can refer to the implementation of step 10323 above for details.

[0224] In some embodiments, the first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position. Step 207 described above can be implemented as follows: for each character position, obtain the target candidate character corresponding to the character position from multiple candidate characters based on the first annotation data; determine the first prediction probability of the target candidate character corresponding to the character position based on the first prediction result of the character position; perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result; fuse the second logarithmic calculation results corresponding to multiple character positions to obtain a third loss corresponding to the second text sample. Thus, by constructing a loss function through logarithmic fusion of character probabilities, long-distance dependencies can be effectively captured, thereby improving the training effect of the model.

[0225] As an example, taking the cross-entropy loss function as an example, the third loss Loss3 can be calculated using formula (4).

[0226]

[0227] Among them, y ijThis represents the label (usually 1) of the target candidate character corresponding to the j-th character position in the i-th second text sample. is the first predicted probability of the target candidate character corresponding to the j-th character position in the i-th second text sample, n is the number of second text samples, and m is the number of characters corresponding to the answer of the second text sample.

[0228] In step 208, the first language model is updated based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the second language model.

[0229] In some embodiments, step 208 described above can be implemented as follows: determining a second ratio between the first number of the first text samples and the second number of the second text samples; obtaining the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples, wherein the ratio between the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples is the same as the second ratio; performing a weighted summation of the second loss corresponding to the first text samples and the third loss corresponding to the second text samples based on the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples to obtain the first loss of the first language model; updating the first language model based on the first loss of the first language model to obtain the second language model. Thus, by determining the loss weights through the number of samples, the proportion of different training sample numbers is reflected in the loss, which can improve the training effect of the model.

[0230] As an example, suppose the second ratio between the first number of the first text sample and the second number of the second text sample is 1. It can be seen that the ratio between the weight of the second loss corresponding to the first text sample and the weight of the third loss corresponding to the second text sample is also... This allows us to obtain the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample (for example, the weight of the second loss is...). The weight of the third loss is 1); then, based on the weights of the second loss corresponding to the first text sample and the third loss corresponding to the second text sample, the second loss Loss2 and the third loss Loss3 corresponding to the second text sample are weighted and summed to obtain the first loss Loss1 of the first language model. Then, the server backpropagates the first loss from the output layer of the first language model, layer by layer. When the first loss is passed to each layer, the gradient (i.e., the partial derivative of the loss function with respect to the parameter values ​​of each layer) is solved in combination with the passed first loss, so as to obtain the gradient of each first parameter value. Finally, the gradient values ​​of the parameters of each layer are updated, and the above update process is iterated until convergence, so as to obtain the second language model.

[0231] Thus, through steps 201 to 208, the model can learn the second task without forgetting the first task, thereby improving the training effect of the model and enhancing its generalization ability in subsequent application stages.

[0232] In some embodiments, updating the first language model based on the first loss of the first language model to obtain the second language model can also be implemented by: determining the gradient of each of the third parameter values ​​based on the first loss; updating each of the third parameter values ​​in at least one third parameter matrix of the third language model based on the gradient of each of the third parameter values ​​and a third update degree parameter for each of the third parameter values ​​to obtain the second language model. Thus, by combining the gradient with the update degree parameter, differentiated processing of model parameter training and preservation is achieved, enabling precise learning of the model parameters. This not only optimizes the training effect but also preserves the initial performance of the model, ultimately improving the overall training effect of the model.

[0233] In some embodiments, updating each third parameter value in at least one third parameter matrix of the first language model based on the gradient of each third parameter value and the third update degree parameter of each third parameter value can be achieved in the following way: obtaining a preset learning rate of the first language model; for each third parameter value in at least one third parameter matrix of the first language model, performing the following processing: multiplying the preset learning rate, the gradient of the third parameter value, and the third update degree parameter of the third parameter value to obtain a multiplication result; updating the third parameter value based on the multiplication result corresponding to the third parameter value. Thus, through the synergistic effect of the learning rate and the update degree parameter, the parameter update intensity can be dynamically balanced, effectively improving the training effect of the model.

[0234] As an example, suppose the preset learning rate of the first language model is 0.001, the gradient of the third parameter value of 0.5 is -0.3, and the third update level parameter of the third parameter value of 0.5 is 0.9. Multiplying the three together, we get a multiplication result of -0.00027. Then, we fuse the multiplication result of -0.00027 with the third parameter value of 0.5 to get 0.49973, which is used as the updated third parameter value, thus completing the update of the third parameter value.

[0235] In some embodiments, see Figure 10 , Figure 10 This is a schematic diagram of the eighth process of the model training method provided in the embodiments of this application, as shown below. Figure 10 As shown, during execution Figure 8 After step 208 shown, the following steps can also be performed: Figure 10 Steps 213 to 214 shown are implemented, and will be combined with Figure 10 The steps shown are explained.

[0236] In step 213, the first text is forward-reasoned using the second language model to obtain the fifth prediction result corresponding to the first text.

[0237] It should be noted that the specific implementation of step 213 is similar to that of step 111 above. For details, please refer to the implementation of step 111.

[0238] As an example, suppose in a data analysis task, a first text T is input into a second language model for predictive analysis, and a third prediction result Q corresponding to the first text T is obtained.

[0239] In step 214, the numerical content in the fifth prediction result is corrected to obtain the sixth prediction result corresponding to the first text.

[0240] It should be noted that the specific implementation of step 214 is similar to the implementation of step 112 above. For details, please refer to the implementation of step 112.

[0241] Following the example above, the numerical content contained in the third prediction result is determined; the numerical content is obtained from online data, and the numerical content in the third prediction result is replaced with the numerical content from the obtained online data, thereby obtaining the fourth prediction result corresponding to the first text T.

[0242] Therefore, conducting online data accuracy verification on the prediction results of the trained model can effectively improve the correctness of data analysis tasks.

[0243] The following describes an exemplary application of the embodiments of this application in a real-world application scenario. This exemplary application describes the specific implementation process of the model training method in a script understanding and script generation scenario.

[0244] Incremental learning refers to a model continuously learning and adapting throughout its working life, integrating new knowledge while retaining previously learned information to prevent catastrophic forgetting. The main challenges of incremental learning include: 1) Catastrophic forgetting: Catastrophic forgetting is one of the core challenges of lifelong learning, as the introduction of new information may overwrite previously learned content; 2) The plasticity-stability dilemma: Finding a balance between a model's learning ability and stability directly affects its ability to acquire new knowledge and retain its broad generalizability; 3) High computational costs: Full fine-tuning of large language models incurs very high computational costs; 4) Unavailability of model parameters or pre-training data: Due to privacy, proprietary restrictions, or commercial licensing, raw training data or model parameters are often unavailable for further improvement.

[0245] Due to different application objectives, large-scale vertical domain models often perform poorly when faced with new, untrained requirements. For example, large-scale models for script comprehension and script generation (used for creating or generating a text description) tend to summarize content based on actual data, requiring high accuracy. However, script creation and comprehension involve fewer numbers and no data consistency requirements, making it difficult for them to accurately summarize film and television information. Furthermore, due to data privacy and other reasons, historical training data is unavailable, leading to stability-plasticity dilemmas and the unavailability of old training data when the model performs incremental learning.

[0246] In related technologies, incremental learning of models is often achieved through the following methods:

[0247] (1) Collect data from all tasks and train the model using all task data (including task data for incremental learning). However, this method is difficult to implement in reality because task data is often unavailable (due to reasons such as confidentiality or data loss).

[0248] (2) Adapter tuning involves adding an adapter to the original model during training to expand the model parameters and improve the model's performance on new tasks. However, since the model structure has been changed, it will cause irreversible damage to the generalizability of the model on other downstream tasks; and by adding an adapter, the characteristics of the old task during inference are changed, so the correctness of the model on the old task cannot be guaranteed, thus causing the problem of forgetting the old task.

[0249] (3) Prefix tuning, based on prompt word prefix optimization, constructs a set of task-related virtual tokens as a prefix before the input token. During training, only the parameters of the prefix are updated, while the parameters of other parts of the pre-trained language model (PLM) remain fixed. By specifically learning task keywords, task keywords are bound to the data. However, while this method can fine-tune the prefix and adapt to new tasks, it is difficult to guarantee the training effect for new tasks due to the limited number of training parameters; furthermore, it cannot overcome the illusion of a large model, making it difficult to guarantee the accuracy of data understanding and representation.

[0250] Based on this, this application designs a method for replaying pseudo-old task training data and using partial parameter and feature regularization to retain old knowledge, as well as a method for customized parameter control learning for new tasks. It constructs pseudo-old task training data using unlabeled data and identifies the low-activation and high-activation parameters of the original model under the old task. During incremental learning, the high-activation parameters are retained for the old task under the pseudo-old task training data. Simultaneously, a task embedding unique to the new task is set to activate the low-activation parameter part in the original model through the task embedding, and the low-activation parameters are used to carry over the learning of the new task. Furthermore, during model application, this application performs online data correction while forward propagating reasoning to output the answer to the question, thereby improving the accuracy of the data output. Finally, without changing the model structure, incremental learning can support both existing applications and new data summary and analysis applications. The model obtained through incremental learning has a simplified model structure, is quick to deploy, and can be directly reused to execute old tasks.

[0251] In some embodiments, see Figure 11 , Figure 11 This is a service flow framework diagram of the model training method provided in the embodiments of this application, such as... Figure 11As shown, for the data input of Task 1 to the corresponding position in the Task 1 data input / output interface, the Task 1 scheduling module calls the model to perform prediction processing on the Task 1 data, obtains the corresponding output result, and returns the output result to the Task 1 input / output interface for display; for the data input of Task 2 to the corresponding position in the Task 2 data input / output interface, the Task 2 scheduling module calls the model to perform prediction processing on the Task 2 data, obtains the corresponding output result, and returns the output result to the Task 2 input / output interface for display; for the data crawling operation triggered by Task 3, Internet data 1, 2, and 3 are obtained, the Task 3 scheduling module calls the model to perform prediction processing on the crawled Internet data 1, 2, and 3, obtains the corresponding output result, and returns the output result to the Task 3 interface. It should be noted that Task 1 and Task 2 are old tasks (corresponding to the first task mentioned above), and Task 3 is a new task (corresponding to the second task mentioned above). A single model can support the services of all three tasks simultaneously. That is, based on the original model (which supports Task 1 and Task 2), incremental learning is used to directly update the model in the original model service and integrate the new task (i.e., Task 3, corresponding to the first task mentioned above), so that the trained model can support the services of all three tasks at the same time.

[0252] In some embodiments, see Figure 12 , Figure 12 This is a schematic diagram of the interface for an information summary task provided in an embodiment of this application, such as... Figure 12 As shown, based on the information summarization task posed by the user, the language model analyzes and summarizes the crawled internet data to obtain the corresponding output summary answer, i.e. Figure 12 "News Summary for Month xx: The trending keyword 'xx' on the xx platform has attracted 200 million views and 5 million discussions, showing an upward trend. Its negative impact on the completion rate of 'x Waist' needs to be monitored."

[0253] In some embodiments, see Figure 13 , Figure 13 This is a flowchart illustrating the training process of the model training method provided in this application embodiment, as shown below. Figure 13As shown, due to the difficulty in obtaining the original training data of the old tasks in the original model, an unlabeled old task training dataset is collected. During the collection of pseudo-old task training data, it is necessary to collect the original model's inference information. Here, inference information refers to the inference information obtained by inputting pseudo samples (corresponding to the first text sample mentioned above) into the original model and inferring it (corresponding to the second prediction result of the first text sample mentioned above). That is, to obtain the business application inference data of the old tasks of the original model (corresponding to the first text sample of the first task mentioned above), and input the business application inference data of the old tasks of the original model (corresponding to the first text sample of the first task mentioned above) into the original model. The prediction results (corresponding to the third language model mentioned above) are obtained in the model (corresponding to the second prediction results corresponding to the first text sample in the first task mentioned above). The unlabeled old task training data and the predictions are combined to obtain pseudo old task training data (corresponding to the combination of the first text sample and the second prediction results corresponding to the first text sample in the first task mentioned above). This is used to align the original model with the old task processing results of the original model in the incremental learning process. Next, different activation parameters are statistically analyzed. By calculating the parameter values ​​(first parameter value or second parameter value) of each parameter matrix in the last layer of the original model's Transformer, the parameter class corresponding to the parameter value is determined. The model is configured with different activation parameters, including low activation parameters, high activation parameters, and normal activation parameters, and different update degree parameters (first update degree parameter or second update degree parameter) are set according to different parameter types. Subsequently, during incremental model learning, the unlabeled old task training data and predictions (the combination of the first text sample and the second prediction result corresponding to the first text sample in the first task mentioned above) and the incremental task data (the combination of the second text sample and the associated first labeled data in the second task mentioned above) are mixed in a certain proportion and input into the model (corresponding to the third language model or the first language model mentioned above). The model is based on the unlabeled old task training data. The model performs prediction alignment (corresponding to the first task mentioned above), which involves aligning the predictions between the model to be trained and the model for the old task that has already been trained. The old task alignment loss (corresponding to the second loss mentioned above) is calculated for the prediction alignment. At the same time, the model performs new task supervised learning based on incremental task data (corresponding to the second task mentioned above), and the new task supervised learning loss (corresponding to the third loss mentioned above) is calculated for the new task supervised learning. The old task alignment loss and the new task supervised loss are weighted and summed with certain weights to obtain the total loss of model training (corresponding to the first loss mentioned above). When updating the model gradient, the parameters of the model to be trained are updated by backpropagation in combination with the update degree parameter.

[0254] It should be noted that the basic model in this application adopts the Qwen2.5-32B model (models with more parameters, such as Qwen2.5-72B, QwQ-32B, etc., can also be used). This model uses a dictionary length of 152064, mapping each character in the dictionary to a 1×5120 embedding. A 64-layer decoding layer is then used as the main structure of the model, followed by a normalization layer. The output embedding is then predicted by a classification head (Im_head) structure to obtain its probability among the 152064 characters. The specific structure of this basic model is consistent with the original Qwen2.5 model architecture. During model training, the embedding, which is a mapping from characters in the input text to dictionary representations, does not require training. Instead, the ModuleList needs to be trained. Here, ModuleList is a dedicated container class for dynamically managing sub-modules, which can organize multiple neural network layers (such as linear layers, convolutional layers, etc.) into a list structure and ensure that the parameters of these neural network layers can be automatically identified and optimized.

[0255] As an example, the structure of the Qwen2.5-32B model in this application can be represented in the following form:

[0256]

[0257] In some embodiments, see Figure 14 , Figure 14 This is a schematic diagram of data preparation for the model training method provided in the embodiments of this application, as shown below. Figure 14As shown, the first step is to collect unlabeled old task training datasets. During the collection of pseudo-old task training data, it is necessary to collect original model inference information, that is, to obtain the business application inference data of the original model's old tasks (corresponding to the first text sample of the first task mentioned above). These data do not have labeled correct answers and can be used as unlabeled data for the old tasks. Since it is difficult to obtain the original training data of the old tasks in the original model, the task can be reproduced through the business application inference data of the old tasks. The inference prediction results of the old task application inference data are used to align with the training process of the model's old tasks (corresponding to the first task mentioned above). It should be noted that pseudo-old task training data refers to fabricated old task training data. Since it is not the real original training data of the old tasks in the original model, but rather alignment data used to maintain the model's ability to perform old tasks during incremental learning. The business application inference data of the old task in the original model (corresponding to the first text sample of the first task mentioned above) is input into the original model (corresponding to the third language model mentioned above). The embedding vector predicted by the last layer of the original model (corresponding to the second prediction result of the first text sample of the first task mentioned above) is used to obtain the unlabeled old task training data and prediction, that is, to obtain the pseudo old task training data (the combination of the first text sample of the first task mentioned above and the second prediction result of the first text sample).

[0258] As an example, the training data for pseudo-old tasks (the combination of the first text sample corresponding to the first task mentioned above and the second prediction result corresponding to the first text sample) can be represented in the format of task description + specific business data, where "+" indicates string concatenation. For example, for the script comprehension task: "Please generate a plot summary of a scene from a film or television script. The script content is: xxxx", where xxxx is the specific script content. For example, the script content of the first scene of script xx includes the following: "Time: 1960s / 70s, 1990s. Location: City A, specific location not limited." For display purposes. All buildings and vehicles in the play, including buildings, stages, and trams, do not need to appear as physical objects. The stage does not use real scenery; the background can be projected. There is a small square table and two chairs in one corner of the stage. The stage has ample open space for moving large props up and down easily. Above the stage (which could be a second-level platform or a lifting device), ten-year-old Abao and six-year-old Betty sit side-by-side. Betty holds onto Abao tightly, their hair flying. All around is a white expanse, with the sound of wind mixed with the whistling of boats on the Huangpu River. Young Abao: "Sweetie, come down. Grandma said we're not allowed to climb on the roof." Betty: "Let me look again! Grandma from Shaoxing is the worst!"

[0259] In some embodiments, during the collection of new task training data (a combination of the first text sample corresponding to the first task mentioned above and the second prediction result corresponding to the first text sample), new task training data for incremental learning is obtained. The data format of the new task training data is the same as that of the pseudo-old task training data, that is, it can be represented as task description + specific business data.

[0260] As an example, the task description for the new task training data could be: "Today's ranking data, including: xx drama series hot list, xx channel list. Please analyze today's ranking data: 1. Summarize today's ranking data, analyzing changes and trends from the perspective of TV projects or sectors, within 20 words. 2. Provide a summary: Analyze the changes and trends of today's dramas on the rankings from the perspective of TV projects or sectors, paying attention to inflection points, etc. The data cited in the analysis must be consistent with the given ranking data for today. Give conclusions in points, within 100-200 words, and must be in Markdown format. Return in JSON format, with the following style: {"brief":xx,"summary":xxx}, where brief is the summary, summary is the analysis, and the value of summary must be in Markdown format. Return the result directly, without using "json" or the word "json". The following is the ranking data: xxxx". The data in the rankings could be: "Top 2010 Rankings (Updated: xx / xx / 20xx): 1. xx Life, Market Share: 16.4%, Broadcast Platform: xx, Rating: S+; 2. xx Bamboo Pavilion, Market Share: 13.8%, Broadcast Platform: xx, Rating: S+; Top 2010 Popular Rankings (Updated: xx / xx / 20xx): 1. xx Bamboo Pavilion, Starring: Liu xx, Zhang xx, Wu xx, Rating: No rating available, Popularity: 99.9, Genre: Ancient Romance; 2. xx Meeting You, Starring: Liu xx, Hu xx, Rating: No rating available, Popularity: 90, Genre: Ancient Romance."

[0261] In some embodiments, see continue to see Figure 14The original model (corresponding to the third language model mentioned earlier) is statistically analyzed using different activation parameters (corresponding to the first or second parameter matrices mentioned earlier). In the last Transformer layer (the last layer in the 64-layer structure), the Q / K / V matrices of each self-attention module (parameters of each head, corresponding to the first or second parameter matrices mentioned earlier) and the parameters of the output projection matrix feature head (each parameter is a parameter matrix, corresponding to the first or second parameter matrices mentioned earlier) are calculated. For each parameter, the mean and variance of all values ​​in that parameter matrix are calculated to determine the activation type of each parameter value, including normal activation parameters, high activation parameters, and low activation parameters. When a parameter value is within the range of the mean of the parameter matrix ± 2 * the standard deviation of the parameter matrix, it is determined as a normal activation parameter; when a parameter value is greater than the mean of the parameter matrix + 2 * the standard deviation of the parameter matrix, it is determined as a high activation parameter; and when a parameter value is less than the mean of the parameter matrix - 2 * the standard deviation of the parameter matrix, it is determined as a low activation parameter. For each parameter value, a new update degree matrix `control` with the same dimensions as the parameter matrix is ​​created (corresponding to the matrix composed of update degree parameters for each parameter value mentioned above). In the parameter matrix, if a parameter value at a certain position is a high-activation or normal-activation parameter, the corresponding value in the update degree matrix is ​​set to 0.1; if a parameter value at a certain position is a low-activation parameter, the corresponding value in the update degree matrix is ​​set to 0.9. Here, the value setting of the corresponding position in the update degree matrix can be adjusted according to the high and low activation learning degree required for different incremental tasks, and no specific limitation is made here. Finally, in the subsequent model training process, this parameter is the secondary weight for gradient update. Generally, the parameter uses the gradient update method w = w0 + r × grad w0 r is the learning rate, w is the parameter matrix after gradient update, w0 is the parameter matrix before gradient update, and grad... w0 Let w be the gradient value corresponding to the parameter matrix w0; after adding the quadratic weights, the parameter update method can be w = w0 + r × grad. w0 ×control w0 control w0 This is the update degree matrix corresponding to the parameter matrix w0 (the matrix composed of the update degree parameters for each parameter value mentioned above). That is, when updating the parameter matrix w0, it is not done according to r×grad. w0 Instead of updating, it updates with a gradient of 0.1x or 0.9x based on the activation type corresponding to each parameter value in the parameter matrix.

[0262] In some embodiments, during the model training process of this application, firstly, the parameters trained on the old task are obtained to initialize the model (corresponding to the third language model mentioned above); then, the task embedding is initialized. For example, for the task "#Daily Hot Topic Information Summary#", ​​the mapping vector from "#Daily Hot Topic Information Summary#" to the dictionary embedding is 12*D_emb. This mapping vector is a trainable parameter for task embedding, where 12 represents the sequence length or number of elements, and D_emb represents the embedding dimension of each unit. Here, in application, when the preposition of the problem is the same as "#Daily Hot Topic Information Summary#", Then, the mapping vector of the original dictionary can be replaced by the trained 12*D_emb. In addition, the model is fine-tuned with all parameters, that is, all parameters of the model, including the problem, need to be fine-tuned. Finally, the model is incrementally learned (corresponding to the second task above). In the incremental learning process, m unlearned task samples are randomly selected from N task samples and batch training is performed. The above training is repeated until each task sample has been traversed once, which is to complete one epoch. Multiple epochs are performed until the average epoch loss (i.e., the average of the batch loss in this epoch) no longer decreases in a certain epoch.

[0263] Here, for each batch of training, the training data for the current batch is first obtained. The ratio of the new task training data (the combination of the first text sample and the second prediction result corresponding to the first text sample in the previous first task) to the pseudo-old task training data (the combination of the first text sample and the second prediction result corresponding to the first text sample in the previous first task) is 1:k (e.g., k=3, etc.). Next, the training data is input into the model to be trained (corresponding to the first language model or the third language model in the previous section), and the model prediction result is obtained through forward propagation. Subsequently, the new task supervision loss (i.e., the new task training data) is calculated for the new task training data. Figure 13 The new task supervision learning (corresponding to the third loss mentioned above) is used to calculate the old task alignment loss for the pseudo-old task training data (the alignment loss here is achieved through the prediction alignment between the two models, that is, the prediction alignment between the model to be trained and the already trained old task model, corresponding to the second loss mentioned above). The new task supervision loss (corresponding to the third loss mentioned above) and the old task alignment loss (corresponding to the second loss mentioned above) are then fused to obtain the total loss (corresponding to the first loss mentioned above). After that, the total loss is backpropagated in the network to calculate the gradient of each parameter of the network. Finally, the parameters of the model are updated according to the gradient of each parameter and the update degree matrix.

[0264] In some embodiments, the goal of the old task alignment loss (corresponding to the second loss mentioned above) is to ensure that the prediction of the model to be trained (corresponding to the first language model or the third language model mentioned above) on the pseudo old data is consistent with the prediction of the original model (corresponding to the third language model mentioned above) in each learning process. The prediction distribution consistency loss is used as a measure. Taking the KL loss (Kullback-Leibler divergence) as an example, the old task alignment loss Los2 can be calculated by formula (5).

[0265]

[0266] Where X is the range of input x in the pseudo-old task training data, P(x) is the prediction result of the new model input x (corresponding to the first prediction result of the first text sample above), and Q(x) is the prediction result of the original model input x (corresponding to the second prediction result of the first text sample above).

[0267] In some embodiments, the new task supervision loss (corresponding to the third loss mentioned above) adopts the cross-entropy loss of text generation supervision, and the new task supervision loss Loss3 can be calculated by formula (6).

[0268]

[0269] Among them, y ij This represents the label (usually 1) of the target candidate character corresponding to the j-th character position in the i-th second text sample. is the first predicted probability of the target candidate character corresponding to the j-th character position in the i-th second text sample, n is the number of second text samples, and m is the number of characters corresponding to the answer of the second text sample.

[0270] It should be noted that the loss of the large model is the classification loss for predicting each word (the concept of a class comes from each word in the dictionary, and each word can be considered as a class). For a word in the text that serves as supervision information, the prediction is considered correct if the model's one-hot prediction matches that word; otherwise, it is incorrect. The one-hot prediction of the correct word in the dictionary serves as supervision information, with only the word's position in the dictionary being 1 and the others being 0. For the one-hot supervision information of the above words, cross-entropy is used as the loss for plot prediction, i.e., the new task supervision loss.

[0271] In some embodiments, after calculating the old task alignment loss Loss2 (corresponding to the second loss mentioned above) and the new task supervision loss Loss3 (corresponding to the third loss mentioned above), the total training loss Loss1 (corresponding to the first loss mentioned above) is obtained by weighted summation according to formula (7).

[0272] Loss1 = a * Loss2 + Loss3 (7)

[0273] Where 'a' is a weight parameter, and the value of 'a' is the ratio of the amount of data in the old task training data (corresponding to the first number of the first text samples in the previous text) to the amount of data in the new task training data (corresponding to the second number of the second text samples in the previous text), so as to ensure that the proportion of new knowledge brought about by the data ratio is reflected in the loss.

[0274] In some embodiments, the model training method provided in this application can improve the accuracy of the model in data analysis tasks, increasing the data accuracy from 85% to 95%, and improving the correctness of the overall analysis conclusions (e.g., the judgment description of trend upward, trend maintenance, etc.). However, there are still a few cases of inaccurate data during the model inference process. Therefore, a post-processing process can be added to the model application for data verification. The main implementation process includes: First, for the input business data x (corresponding to the first text above), obtain the corresponding real data mapping, such as {Yesterday's _xx Hot List_《X-Mirror》_Popularity: 6.5, Yesterday's _xx Hot List_《X-Mirror》_Rank: 5, Yesterday's _xx_《X-Mirror》_Market Share: 17%, Yesterday's _xx_《X-Mirror》_Rank: 4, ...}; Next, input the business data x into the trained model to generate the question-answering output result y (corresponding to the third prediction result of the first text above); Subsequently, use deep learning... seekR1-32B determines the format of the value to be confirmed for the output result y based on the business situation: "Given a text, please replace specific data with <time_xx_hot_list_project_popularity>, <time_xx_hot_list_project_ranking>, <time_xx_project_market share>, <time_xx_project_ranking>. The following is the text content: y"—For other data given values, you can refer to this and list them in a finite closed set manner." For example, regarding the result "**Analysis Results:** 1. **Platform Competition Differentiation**: xx occupies two of the xx Top 3 spots with 'xx Crossing' and 'xx Born,' while xx focuses on the suspense and urban genres with 'x Mirror' and 'x Family.' 2. **Positive Correlation Between Reputation and Popularity**: In the xx list, 'xx Crossing' leads in popularity with a score of 8.6, and high-reputation dramas are more likely to achieve cross-border dissemination." "After correction, we can obtain: **Analysis Results:** 1. **Platform Competition Differentiation:** xx occupies two of the top 3 spots in xx with 'xx Crossing' and 'xx Born,' while xx focuses on the suspense and urban genres with 'x Mirror' and 'x Family.' 2. **Positive Correlation Between Reputation and Popularity:** In the xx rankings, 'xx Crossing' leads in popularity with a score of <yesterday_xx Hot List_'xx Crossing'_Popularity>, indicating that highly acclaimed dramas are more likely to achieve wider dissemination." Finally, for the output of the confirmation value format, find the content corresponding to the format "<>", obtain the real value from the real data mapping, and replace the data of "<>" with the real value to complete the data correction, obtaining the corrected question and answer output result (corresponding to the fourth prediction result of the first text above).

[0275] In summary, this application improves the discriminative processing of model parameter training and preservation by pre-determining some learnable parameters (e.g., low activation parameters) and non-learnable parameters (e.g., high activation parameters) in specific layers of the model and treating them differently. Furthermore, it constructs a batch of pseudo-old task training data using unlabeled data to ensure the consistency of the model's performance on old tasks during training. Simultaneously, it controls the non-learnable parameters to minimize changes to ensure the model's performance on old tasks. In addition, online data accuracy verification during the model application phase effectively improves the correctness of the generated results.

[0276] The following description continues to illustrate the exemplary structure of the model training device 543 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2A As shown, the software modules stored in the model training device 543 in the memory 540 may include: an acquisition module 5431, a first determination module 5432, and a first training module 5433.

[0277] The acquisition module 5431 is used to acquire at least one first parameter matrix of the first language model; the first determination module 5432 is used to determine a first update degree parameter for each first parameter value in the first parameter matrix based on a first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix; the first determination module 5432 is further used to determine a first loss of the first language model based on the first training samples, and determine the gradient of each first parameter value based on the first loss; the first training module 5433 is used to update each first parameter value in at least one first parameter matrix of the first language model based on the gradient of each first parameter value and the first update degree parameter of each first parameter value, to obtain a second language model.

[0278] In some embodiments, the first determining module 5432 is further configured to calculate the standard deviation of all first parameter values ​​in the first parameter matrix to obtain the standard deviation of the first parameter matrix; determine a first threshold positively correlated with the standard deviation; determine the parameter type of each first parameter value in the first parameter matrix based on the relationship between the first difference corresponding to each first parameter value in the first parameter matrix and the first threshold; and determine a first update degree parameter corresponding to each first parameter value in the first parameter matrix based on the parameter type of each first parameter value in the first parameter matrix.

[0279] In some embodiments, the model training apparatus 543 further includes a transfer module 5434.

[0280] The transfer module 5434 is used to obtain at least one second parameter matrix of a third language model, wherein the third language model is trained based on a first task; for each second parameter matrix of the third language model, a second update degree parameter is determined for each second parameter value in the second parameter matrix based on a second difference between each second parameter value in the second parameter matrix and the matrix mean of the second parameter matrix; a first loss of the third language model is determined based on a second training sample corresponding to the second task, and a gradient of each second parameter value is determined based on the first loss of the third language model; and each second parameter value in at least one second parameter matrix of the third language model is updated based on the gradient of each second parameter value and the second update degree parameter of each second parameter value to obtain the first language model.

[0281] In some embodiments, the first determining module 5432 is further configured to perform forward inference on the first training sample through the first language model to obtain a first prediction result corresponding to the first training sample; and determine a first loss of the first language model based on the first prediction result and the first labeled data corresponding to the first training sample.

[0282] In some embodiments, the first training sample includes a first text sample corresponding to a first task and a second text sample corresponding to a second task and associated with first annotation data; the first determining module 5432 is further configured to perform forward inference on the first text sample corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample; perform forward inference on the second text sample corresponding to the second task using the first language model to obtain a first prediction result corresponding to the second text sample; the first determining module 5432 is further configured to perform forward inference on the first text sample corresponding to the first task using a third language model to obtain a second prediction result corresponding to the first text sample; determine a second loss corresponding to the first text sample based on the first prediction result and the second prediction result; determine a third loss corresponding to the second text sample based on the first prediction result and the first annotation data corresponding to the second text sample; and fuse the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain a first loss of the first language model.

[0283] In some embodiments, the first determining module 5432 is further configured to determine a first ratio between a first prediction result corresponding to the first text sample and a second prediction result corresponding to the first text sample; perform logarithmic calculation on the first ratio to obtain a first logarithmic calculation result; and multiply the first logarithmic calculation result with the first prediction result corresponding to the first text sample to obtain a second loss corresponding to the first text sample.

[0284] In some embodiments, the first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position; the first determining module 5432 is further configured to, for each character position, obtain the target candidate character corresponding to the character position from a plurality of candidate characters based on the first annotation data, determine the first prediction probability of the target candidate character corresponding to the character position based on the first prediction result of the character position, perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result; and perform fusion processing on the second logarithmic calculation results corresponding to a plurality of character positions to obtain a third loss corresponding to the second text sample.

[0285] In some embodiments, the first determining module 5432 is further configured to determine a second ratio between the first number of the first text samples and the second number of the second text samples; obtain the weight of the second loss corresponding to the first text sample and the weight of the third loss corresponding to the second text sample, wherein the ratio between the weight of the second loss corresponding to the first text sample and the weight of the third loss corresponding to the second text sample is the same as the second ratio; and perform a weighted summation of the second loss corresponding to the first text sample and the third loss corresponding to the second text sample based on the weight of the second loss corresponding to the first text sample and the weight of the third loss corresponding to the second text sample to obtain the first loss of the first language model.

[0286] In some embodiments, the first training sample includes a second text sample corresponding to the second task and associated with first labeled data; the first determining module 5432 is further configured to perform forward inference on the second text sample corresponding to the second task through the first language model to obtain a first prediction result corresponding to the second text sample; the first determining module 5432 is further configured to determine a third loss corresponding to the second text sample based on the first prediction result corresponding to the second text sample and the first labeled data corresponding to the second text sample; and determine a first loss of the first language model based on the third loss.

[0287] In some embodiments, the first training module 5433 is further configured to obtain a preset learning rate of the first language model; and for each first parameter value in at least one first parameter matrix of the first language model, perform the following processing respectively: multiply the preset learning rate, the gradient of the first parameter value and the first update degree parameter of the first parameter value to obtain a first multiplication result; and update the first parameter value based on the first multiplication result corresponding to the first parameter value.

[0288] In some embodiments, the model training device 543 further includes a first correction module 5435.

[0289] The first correction module 5435 is used to perform forward reasoning on the first text through the second language model to obtain a third prediction result corresponding to the first text; and to correct the numerical content in the third prediction result to obtain a fourth prediction result corresponding to the first text.

[0290] In some embodiments, such as Figure 2B As shown, the software modules stored in the model training device 643 in the memory 640 may include: a generation module 6431, an inference module 6432, a second determination module 6433, and a second training module 6434.

[0291] The generation module 6431 is used to generate a first text sample corresponding to a first task and obtain a second text sample corresponding to a second task and associated with first labeled data; the inference module 6432 is used to perform forward inference on the first text sample corresponding to the first task using a third language model to obtain a second prediction result corresponding to the first text sample, wherein the third language model is a model trained based on the first task; the inference module 6432 is also used to perform forward inference on the first text sample corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample, and perform forward inference on the second text sample corresponding to the second task to obtain a first prediction result corresponding to the second text sample; the second determination module 6433 is used to determine a second loss corresponding to the first text sample based on the first prediction result and the second prediction result corresponding to the first text sample, and determine a third loss corresponding to the second text sample based on the first prediction result and the first labeled data; the second training module 6434 is used to perform fusion processing based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain a first loss of the first language model, and update the first language model based on the first loss of the first language model to obtain a second language model.

[0292] In some embodiments, the second determining module 6433 is further configured to determine a first ratio between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample; perform logarithmic calculation on the first ratio to obtain a first logarithmic calculation result; and multiply the first logarithmic calculation result with the first prediction result corresponding to the first text sample to obtain a second loss corresponding to the first text sample.

[0293] In some embodiments, the first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position; the second determining module 6433 is further configured to, for each character position, obtain the target candidate character corresponding to the character position from the plurality of candidate characters based on the first annotation data, determine the first prediction probability of the target candidate character corresponding to the character position based on the first prediction result of the character position, perform logarithmic calculation on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result; and perform fusion processing on the second logarithmic calculation results corresponding to the plurality of character positions to obtain a third loss corresponding to the second text sample.

[0294] In some embodiments, the second training module 6434 is further configured to: determine a second ratio between the first number of the first text samples and the second number of the second text samples; obtain the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples, wherein the ratio between the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples is the same as the second ratio; perform a weighted summation of the second loss corresponding to the first text samples and the third loss corresponding to the second text samples based on the weights of the second loss corresponding to the first text samples and the weights of the third loss corresponding to the second text samples to obtain the first loss of the first language model; and update the first language model based on the first loss of the first language model to obtain the second language model.

[0295] In some embodiments, the model training device 643 further includes a third training module 6435.

[0296] The third training module 6435 is used to acquire a third text sample corresponding to the second task and associated with second labeled data; perform forward inference on the third text sample corresponding to the second task using a third language model to obtain a first prediction result corresponding to the third text sample; determine a third loss corresponding to the third text sample based on the first prediction result corresponding to the second text sample and the second labeled data; and update the third language model based on the third loss corresponding to the third text sample to obtain the first language model.

[0297] In some embodiments, the model training device 643 further includes a second correction module 6436.

[0298] The second correction module 6436 is used to perform forward reasoning on the first text through the second language model to obtain a fifth prediction result corresponding to the first text; and to correct the numerical content in the fifth prediction result to obtain a sixth prediction result corresponding to the first text.

[0299] It should be noted that the description of the apparatus in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments, therefore, it will not be repeated. For any technical details not covered in the model training apparatus provided in this application, please refer to... Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 ,or Figure 10 The meaning is understood in accordance with the description of any of the accompanying drawings.

[0300] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the model training method described above in this application.

[0301] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the model training method provided in this application. For example, ... Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 ,or Figure 10 The model training method is shown.

[0302] In some embodiments, the computer-readable storage medium may be a memory such as ferroelectric random access memory (FRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); or it may be a device that includes one or any combination of the above-mentioned memories.

[0303] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0304] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0305] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0306] In summary, the embodiments of this application first obtain at least one first parameter matrix of the first language model, which clarifies the trainable parameters of the language model and lays the foundation for subsequent differentiated updates. Then, for each first parameter matrix, an update degree parameter for each first parameter value is determined based on the first difference between each first parameter value and the matrix mean of the first parameter matrix. Different update degree parameters can be applied to different first parameter values, thereby controlling the learning degree of different first parameter values ​​and helping to maintain the initial performance of the first language model. Next, a first loss of the first language model is determined based on the training samples, and the gradient of each first parameter value is determined based on the first loss. Finally, the parameters of the first language model are updated based on the gradient of the first parameter value and the update degree parameter corresponding to the first parameter value, resulting in a second language model. By combining the gradient with the update degree parameter, differentiated processing of model parameter training and preservation is achieved, enabling precise learning of the model parameters. This not only optimizes the training effect but also preserves the initial performance of the model, ultimately improving the overall training effect of the model.

[0307] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A model training method, characterized in that, The method includes: Obtain at least one first parameter matrix of the first language model; For each of the first parameter matrices, a first update degree parameter is determined for each first parameter value in the first parameter matrix based on the first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix. Based on the first training sample, determine the first loss of the first language model, and determine the gradient of each of the first parameter values ​​based on the first loss; Based on the gradient of each first parameter value and the first update degree parameter of each first parameter value, each first parameter value in at least one first parameter matrix of the first language model is updated to obtain a second language model.

2. The method according to claim 1, characterized in that, The step of determining the first update degree parameter for each first parameter value in the first parameter matrix based on the first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix includes: Calculate the standard deviation of all first parameter values ​​in the first parameter matrix to obtain the standard deviation of the first parameter matrix; Determine a first threshold that is positively correlated with the standard deviation; Based on the relationship between the first difference corresponding to each first parameter value in the first parameter matrix and the first threshold, the parameter type of each first parameter value in the first parameter matrix is ​​determined; Based on the parameter type of each first parameter value in the first parameter matrix, determine the first update degree parameter corresponding to each first parameter value in the first parameter matrix.

3. The method according to claim 1, characterized in that, The method further includes: Obtain at least one second parameter matrix of a third language model, wherein the third language model is trained based on a first task; For each of the second parameter matrices of the third language model, a second update degree parameter is determined for each second parameter value in the second parameter matrix based on the second difference between each second parameter value in the second parameter matrix and the matrix mean of the second parameter matrix; Based on the second training samples corresponding to the second task, the first loss of the third language model is determined, and the gradient of each second parameter value is determined based on the first loss of the third language model. Based on the gradient of each second parameter value and the second update degree parameter of each second parameter value, the first language model is obtained by updating each second parameter value in at least one second parameter matrix of the third language model.

4. The method according to claim 1, characterized in that, The step of determining the first loss of the first language model based on the first training samples includes: The first language model is used to perform forward reasoning on the first training sample to obtain the first prediction result corresponding to the first training sample. Based on the first prediction result and the first labeled data corresponding to the first training sample, the first loss of the first language model is determined.

5. The method according to claim 4, characterized in that, The first training samples include a first text sample corresponding to the first task and a second text sample corresponding to the second task and associated with the first labeled data; The step of performing forward inference on the first training sample using the first language model to obtain a first prediction result corresponding to the first training sample includes: The first language model is used to perform forward inference on the first text sample corresponding to the first task to obtain the first prediction result corresponding to the first text sample. The first language model is used to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample. The step of determining the first loss of the first language model based on the first prediction result and the first labeled data corresponding to the first training sample includes: By performing forward inference on the first text sample corresponding to the first task using a third language model, a second prediction result corresponding to the first text sample is obtained. Based on the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample, the second loss corresponding to the first text sample is determined. Based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample, the third loss corresponding to the second text sample is determined; The second loss corresponding to the first text sample and the third loss corresponding to the second text sample are fused together to obtain the first loss of the first language model.

6. The method according to claim 5, characterized in that, The step of determining the second loss corresponding to the first text sample based on the first prediction result corresponding to the first text sample includes: Determine a first ratio between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample; Perform logarithmic calculation on the first ratio to obtain the first logarithmic calculation result; The first logarithmic calculation result is multiplied by the first prediction result corresponding to the first text sample to obtain the second loss corresponding to the first text sample.

7. The method according to claim 5, characterized in that, The first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position; The step of determining the third loss corresponding to the second text sample based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample includes: For each character position, a target candidate character corresponding to the character position is obtained from multiple candidate characters based on the first annotation data. Based on the first prediction result of the character position, a first prediction probability of the target candidate character corresponding to the character position is determined. Logarithmic calculation is performed on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result. The second logarithm calculation results corresponding to multiple character positions are fused to obtain the third loss corresponding to the second text sample.

8. The method according to claim 5, characterized in that, The step of fusing the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the first loss of the first language model includes: Determine a second ratio between a first number of the first text samples and a second number of the second text samples; Obtain the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample, wherein the ratio between the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample is the same as the second ratio. Based on the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample, the second loss corresponding to the first text sample and the third loss corresponding to the second text sample are weighted and summed to obtain the first loss of the first language model.

9. The method according to claim 1, characterized in that, The step of updating each first parameter value in at least one first parameter matrix of the first language model based on the gradient of each first parameter value and the first update degree parameter of each first parameter value includes: Obtain the preset learning rate of the first language model; For each value of the first parameter in at least one first parameter matrix of the first language model, the following processing is performed respectively: The preset learning rate, the gradient of the first parameter value, and the first update degree parameter of the first parameter value are multiplied together to obtain the first multiplication result; The first parameter value is updated based on the first multiplication result corresponding to the first parameter value.

10. The method according to claim 4, characterized in that, The first training sample includes a second text sample that corresponds to the second task and is associated with the first labeled data; The step of performing forward inference on the first training sample using the first language model to obtain a first prediction result corresponding to the first training sample includes: The first language model is used to perform forward inference on the second text sample corresponding to the second task to obtain the first prediction result corresponding to the second text sample. The step of determining the first loss of the first language model based on the first prediction result and the first labeled data corresponding to the first training sample includes: Based on the first prediction result corresponding to the second text sample and the first annotation data corresponding to the second text sample, the third loss corresponding to the second text sample is determined; Based on the third loss, the first loss of the first language model is determined.

11. A model training method, characterized in that, The method includes: Generate a first text sample corresponding to the first task, and obtain a second text sample corresponding to the second task and associated with the first labeled data; The first text sample is subjected to forward inference corresponding to the first task by a third language model to obtain a second prediction result corresponding to the first text sample, wherein the third language model is a model trained based on the first task; The first text sample is subjected to forward inference corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample, and the second text sample is subjected to forward inference corresponding to the second task to obtain a first prediction result corresponding to the second text sample. Based on the first prediction result and the second prediction result corresponding to the first text sample, a second loss corresponding to the first text sample is determined, and based on the first prediction result and the first annotation data corresponding to the second text sample, a third loss corresponding to the second text sample is determined. The first language model is updated based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the second language model.

12. The method according to claim 11, characterized in that, The step of determining the second loss corresponding to the first text sample based on the first prediction result and the second prediction result corresponding to the first text sample includes: Determine a first ratio between the first prediction result corresponding to the first text sample and the second prediction result corresponding to the first text sample; Perform logarithmic calculation on the first ratio to obtain the first logarithmic calculation result; The first logarithmic calculation result is multiplied by the first prediction result corresponding to the first text sample to obtain the second loss corresponding to the first text sample.

13. The method according to claim 11, characterized in that, The first prediction result corresponding to the second text sample includes the first prediction result for each character position, and the first prediction result for each character position includes the first prediction probability for each candidate character corresponding to the character position; The step of determining the third loss corresponding to the second text sample based on the first prediction result and the first annotation data includes: For each character position, a target candidate character corresponding to the character position is obtained from multiple candidate characters based on the first annotation data. Based on the first prediction result of the character position, a first prediction probability of the target candidate character corresponding to the character position is determined. Logarithmic calculation is performed on the first prediction probability corresponding to the target candidate character to obtain a second logarithmic calculation result. The second logarithm calculation results corresponding to multiple character positions are fused to obtain the third loss corresponding to the second text sample.

14. The method according to claim 11, characterized in that, The step of updating the first language model based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the second language model includes: Determine a second ratio between a first number of the first text samples and a second number of the second text samples; Obtain the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample, wherein the ratio between the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample is the same as the second ratio. Based on the weights of the second loss corresponding to the first text sample and the weights of the third loss corresponding to the second text sample, the second loss corresponding to the first text sample and the third loss corresponding to the second text sample are weighted and summed to obtain the first loss of the first language model. The first language model is updated based on the first loss of the first language model to obtain the second language model.

15. The method according to claim 11, characterized in that, The method further includes: Obtain the third text sample that corresponds to the second task and is associated with the second labeled data; The third language model is used to perform forward inference on the third text sample corresponding to the second task to obtain a first prediction result corresponding to the third text sample; Based on the first prediction result and the second annotation data corresponding to the third text sample, a third loss corresponding to the third text sample is determined; The third language model is updated based on the third loss corresponding to the third text sample to obtain the first language model.

16. A model training device, characterized in that, The device includes: The acquisition module is used to acquire at least one first parameter matrix of the first language model; The first determining module is used to determine, for each of the first parameter matrices, a first update degree parameter for each first parameter value in the first parameter matrix based on a first difference between each first parameter value in the first parameter matrix and the matrix mean of the first parameter matrix; The first determining module is further configured to determine a first loss of the first language model based on the first training samples, and to determine the gradient of each of the first parameter values ​​based on the first loss; A first training module is used to update each of the first parameter values ​​in at least one first parameter matrix of the first language model based on the gradient of each first parameter value and a first update degree parameter of each first parameter value, to obtain a second language model.

17. A model training device, characterized in that, The device includes: The generation module is used to generate a first text sample corresponding to the first task and to obtain a second text sample corresponding to the second task and associated with the first annotation data. The reasoning module is used to perform forward reasoning on the first text sample corresponding to the first task using a third language model to obtain a second prediction result corresponding to the first text sample, wherein the third language model is a model trained based on the first task. The reasoning module is further configured to perform forward reasoning on the first text sample corresponding to the first task using the first language model to obtain a first prediction result corresponding to the first text sample, and to perform forward reasoning on the second text sample to obtain a first prediction result corresponding to the second text sample. The second determining module is used to determine a second loss corresponding to the first text sample based on the first prediction result and the second prediction result corresponding to the first text sample, and to determine a third loss corresponding to the second text sample based on the first prediction result and the first annotation data. The second training module is used to perform fusion processing based on the second loss corresponding to the first text sample and the third loss corresponding to the second text sample to obtain the first loss of the first language model, and to update the first language model based on the first loss of the first language model to obtain the second language model.

18. An electronic device, characterized in that, include: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the model training method according to any one of claims 1 to 15.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the model training method according to any one of claims 1 to 15 is implemented.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the model training method according to any one of claims 1 to 15 is implemented.