Machine learning model training method and device, electronic equipment, computer readable storage medium and computer program product
By performing phased iterative training and parameter fusion on the machine learning model, the problem of balancing the learning performance of historical and new tasks in existing technologies is solved, the stability and resource utilization efficiency of the model in multi-task training are improved, and the adaptability of the model in complex scenarios is enhanced.
Patent Information
- Application Number
- CN202511197537.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies struggle to achieve an optimal balance between maintaining learning performance for historical and new tasks during continuous learning, leading to catastrophic forgetting problems and increased computational resources and training time.
By performing the first stage of iterative training on the first machine learning model, determining the parameter changes and loss value, adjusting the number of iterations, performing the second stage of iterative training, and fusing the model parameters, a fourth machine learning model is formed, which can adapt to various prediction tasks.
It improves the stability of machine learning models during multi-task training and the efficiency of computing resource utilization, reduces performance fluctuations caused by task switching, and enhances the adaptability and stability of models in complex scenarios.
Smart Images

Figure CN121094014A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a machine learning model training method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] Continual learning aims to enable a model to continuously accumulate knowledge from non-stationary data. To address the catastrophic forgetting problem in continual learning, related technologies can alleviate catastrophic forgetting for trained tasks or promote knowledge transfer for newly trained tasks through regularization constraints, historical data playback, and adjustment of model architecture, but these methods require storage of a large amount of historical data, occupy a large amount of computing resources, and especially when dealing with large-scale data and complex tasks, the training time is significantly increased, and it is difficult to achieve an optimal balance between maintaining the learning performance of historical tasks and new tasks. SUMMARY
[0003] The embodiments of the present application provide a machine learning model training method and device, electronic equipment, computer readable storage medium and computer program product, which can improve the stability of the model in the multi-task training process.
[0004] The technical solutions of the embodiments of the present application are as follows:
[0005] The embodiments of the present application provide a machine learning model training method, which comprises:
[0006] The first machine learning model is trained based on the first sample data to obtain a second machine learning model through first-stage iterative training of the first machine learning model based on first sample data and second sample data of a second prediction task, wherein the first machine learning model is trained based on the first sample data;
[0007] The first loss value and the parameter change amount of the second machine learning model are determined;
[0008] The number of iterations to be performed on the second machine learning model is determined according to the parameter change amount;
[0009] The second machine learning model is iteratively trained based on the first loss value and the number of iterations to be performed on the second machine learning model through the first sample data and the second sample data to obtain a third machine learning model;
[0010] The third machine learning model and the second machine learning model are fused to obtain a fourth machine learning model, wherein the fourth machine learning model is used to perform at least one of the first prediction task and the second prediction task.
[0011] In the scheme, the first weight of the first parameter variation and the second weight of the second parameter variation are determined according to the first performance index and the second performance index, including:
[0012] determining a sum of the first performance index and the second performance index;
[0013] taking a ratio of the first performance index to the sum as the first weight of the first parameter variation;
[0014] taking a ratio of the second performance index to the sum as the second weight of the second parameter variation.
[0015] An embodiment of the present application provides a training device of a machine learning model, and the device comprises:
[0016] a first training module configured to perform first-stage iterative training on a first machine learning model based on first sample data of a first prediction task and second sample data of a second prediction task, and obtain a second machine learning model, wherein the first machine learning model is trained based on the first sample data;
[0017] a determination module configured to determine a first loss value and a parameter variation of the second machine learning model, and determine a to-be-iterated number of times of the second machine learning model according to the parameter variation;
[0018] a second training module configured to perform second-stage iterative training on the second machine learning model based on the first loss value and the to-be-iterated number of times of the second machine learning model through the first sample data and the second sample data, and obtain a third machine learning model;
[0019] a fusion module configured to perform model parameter fusion on the third machine learning model and the second machine learning model, and obtain a fourth machine learning model, wherein the fourth machine learning model is used to perform at least one of the first prediction task and the second prediction task.
[0020] An embodiment of the present application provides an electronic device, and the electronic device comprises:
[0021] a memory configured to store computer executable instructions or computer programs;
[0022] a processor configured to execute the computer executable instructions or computer programs stored in the memory, and implement the training method of the machine learning model provided in the embodiments of the present application.
[0023] An embodiment of the present application provides a computer readable storage medium, which stores computer programs or computer executable instructions, and is used to implement the training method of the machine learning model provided in the embodiments of the present application when executed by a processor.
[0024] The embodiment of the present application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the training method of the machine learning model provided by the embodiment of the present application.
[0025] The embodiment of the present application has the following beneficial effects:
[0026] After the first machine learning model is trained by the first sample data of the first prediction task, the first machine learning model is iteratively trained in the first stage by the first sample data of the first prediction task and the second sample data of the second prediction task, and the second machine learning model is obtained, and the number of iterations to be performed is determined according to the parameter variation of the second machine learning model. In the first stage, the model is iteratively trained by the data of two different tasks, and the parameter variation can be used as a basis for adjusting the training intensity. This dynamic adjustment mechanism based on the parameter variation can accurately control the training intensity, avoid the waste of computing resources caused by overtraining, and prevent the performance of the model from being affected by insufficient training.
[0027] After the number of iterations to be performed is determined, the second machine learning model is iteratively trained in the second stage based on the first loss value and the number of iterations to be performed by the first sample data and the second sample data. This can make full use of existing data resources, avoid data redundancy and resource waste in single task training, make the obtained third machine learning model take into account the performance of the two prediction tasks, improve the stability of the machine learning model in the multi-task training process, more efficiently use computing resources, and reduce the computing overhead caused by task switching. The model parameters of the third machine learning model and the second machine learning model are fused, so that the fourth machine learning model obtained finally can make full use of the training results of the two stages, better adapt to multiple prediction tasks, perform more balanced when processing different tasks, reduce the performance fluctuation caused by task switching, and thus significantly improve the overall adaptability and stability of the computer system in complex and variable application scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is an architecture schematic diagram of the training system 100 of the machine learning model provided by the embodiment of the present application;
[0029] Figure 2 is a structure schematic diagram of the server 200 provided by the embodiment of the present application;
[0030] Figure 3A is a first flow schematic diagram of the training method of the machine learning model provided by the embodiment of the present application;
[0031] Figure 3BFIG. 2 is a second flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0032] Figure 3C FIG. 3 is a third flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0033] Figure 3D FIG. 4 is a fourth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0034] Figure 3E FIG. 5 is a fifth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0035] Figure 3F FIG. 6 is a sixth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0036] Figure 3G FIG. 7 is a seventh flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0037] Figure 3H FIG. 8 is an eighth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0038] Figure 3I FIG. 9 is a ninth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0039] Figure 3J FIG. 10 is a tenth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0040] Figure 3K FIG. 11 is an eleventh flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0041] Figure 3L FIG. 12 is a twelfth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0042] Figure 3M FIG. 13 is a thirteenth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0043] Figure 3N FIG. 14 is a fourteenth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0044] Figure 3O FIG. 15 is a fifteenth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0045] Figure 3PFig. 16 is a sixteenth flowchart illustrating a method for training a machine learning model according to an embodiment of the present application;
[0046] Figure 4 Fig. 17 is a diagram illustrating a task state visualization according to an embodiment of the present application;
[0047] Figure 5 Fig. 18 is a model fusion timing change diagram according to an embodiment of the present application;
[0048] Figure 6A Fig. 19 is a diagram illustrating a model single round fusion according to an embodiment of the present application;
[0049] Figure 6B Fig. 20 is a diagram illustrating a model multi-round fusion according to an embodiment of the present application;
[0050] Figure 6C Fig. 21 is a diagram illustrating an adaptive iterative fusion according to an embodiment of the present application;
[0051] Figure 7 Fig. 22 is a diagram illustrating a learning signal change according to an embodiment of the present application;
[0052] Figure 8 Fig. 23 is a diagram illustrating a forgetting signal change according to an embodiment of the present application;
[0053] Figure 9 Fig. 24 is a diagram illustrating a comparison of historical losses of different methods according to an embodiment of the present application;
[0054] Figure 10 Fig. 25 is a diagram illustrating a first change curve according to an embodiment of the present application;
[0055] Figure 11 Fig. 26 is a diagram illustrating a second change curve according to an embodiment of the present application.
[0056] It should be noted that the above-mentioned "first" and "second" are only used to distinguish different schemes, and do not represent the degree of superiority or priority in the implementation process. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the accompanying drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by a person of ordinary skill in the art without making creative labor fall within the scope of protection of the present application.
[0058] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0059] In the following description, the terms "first / second / third" are merely used to distinguish similar objects, and do not represent a specific order of the objects. Understandably, the "first / second / third" can be interchanged in a specific order or sequence as permitted, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0060] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0061] As no specific description is given, at least one of the following description refers to one or more cases, and "a plurality of" can refer to two or more cases.
[0062] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by a person skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0063] The relevant data collection process in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and within the scope of authorization of laws and regulations and the personal information subject, carry out subsequent data use and processing behavior.
[0064] Before further detailing the embodiments of the present application, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.
[0065] 1) Prediction task: The prediction task refers to a specific task that a machine learning model needs to complete in the training process, for example, the types of prediction tasks can include image classification, text classification, dialogue question and answer, text matching, text generation, etc. Taking text classification as an example, the prediction task can be news topic classification, customer review sentiment analysis, etc.
[0066] 2) Prediction data: The output result generated by the machine learning model according to the input sample data.
[0067] 3) Loss value: It is an index for measuring the difference between the prediction result output by the machine learning model for the sample data and the true label corresponding to the sample data.
[0068] 4) Model parameter fusion: refers to a process of merging parameters corresponding to machine learning models obtained in different iterations respectively to construct a new machine learning model in the iterative training process of the machine learning model.
[0069] 5) Parameter change amount: refers to the change amount of the parameter of the machine learning model between the fusion results of adjacent two model parameter fusions, which reflects the updating degree of the machine learning model in different fusion stages.
[0070] 6) Change index parameter: calculated based on the plurality of sub-parameter change amounts included in the parameter change amount, and is an index for measuring the parameter change trend of the machine learning model in the iterative training process.
[0071] 7) In response to: used to represent the condition or state on which the operation is dependent, when the dependent condition or state is met, the one or more operations performed can be real-time or have a set delay; in the absence of special instructions, there is no restriction on the execution order of the multiple operations performed.
[0072] 8) Graphical interface: an interface for displaying data statistics of the machine learning model in the iterative training process. For example, a graphical user interface (Graphical User Interface, GUI) display, such as an augmented reality (Augmented Reality, AR) interface, a virtual reality (Virtual Reality, VR) interface, a voice user interface (Voice User Interface, VUI), an interactive projection interface (using projection technology to display information on a plane), an eye movement detection interface (an interface controlled by detecting the user's visual line), a holographic interface (a three-dimensional holographic image formed by holographic projection technology, without wearing special glasses to see a stereoscopic image), a multi-modal interface (an interactive interface combining multiple interaction modes such as touch, vision, hearing, etc.), a brain-machine interface (Brain-Machine Interface, BMI) interface, etc.
[0073] The related art can alleviate catastrophic forgetting or promote knowledge transfer through regularization constraints, historical data playback, and adjustment of model architecture, but these methods still have limitations, that is, it is difficult to achieve an optimal balance between maintaining the learning performance of historical tasks and new tasks.
[0074] Based on the above analysis, the applicant finds that the training method of the machine learning model in the related art cannot achieve an optimal balance between maintaining old knowledge memory and new knowledge learning performance, and improve the stability of the model in the multi-task training process. In view of the above problems, the embodiment of the present application provides a machine learning model training method, which can improve the stability of the model in the multi-task training process.
[0075] In view of the above problems, the embodiment of the present application provides a machine learning model training method, device, electronic device, computer readable storage medium and computer program product, which can improve the stability of the model in the multi-task training process. The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart television, a vehicle-mounted terminal and various types of terminals, and can also be implemented as a server. The following will illustrate the exemplary application when the electronic device is implemented as a terminal or a server.
[0076] Referring to Figure 1 , Figure 1 is an architecture schematic diagram of the machine learning model training system 100 provided by the embodiment of the present application. In order to support the training application of one machine learning model, the terminal 400 (exemplarily shows a graphical interface 411) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0077] Taking a medical scenario as an example, the first prediction task can be to predict whether a patient has a certain disease (classification task) according to the sample data such as the symptoms and examination results of the patient; and the second prediction task is to recommend the best treatment plan for the diagnosed patient (multi-classification task). The terminal 400 is used to train the first machine learning model in two stages based on the sample data of the two prediction tasks, to obtain a third machine learning model; and the third machine learning model and the second machine learning model are fused in model parameters to obtain a fourth machine learning model, so that the fourth machine learning model can learn the diagnosis of the disease and the recommendation of the treatment plan at the same time, and improve the overall medical decision-making efficiency.
[0078] Taking a script analysis scenario as an example, the first prediction task can be to analyze character relationships and plot details in the script; and the second prediction task is to analyze the emotional types of characters in the script (such as anger, sadness, etc. of the characters). The terminal 400 is configured to train the first machine learning model in two stages based on sample data of the two prediction tasks (such as text of partial paragraphs of the script), to obtain a third machine learning model; and to perform model parameter fusion on the third machine learning model and the second machine learning model, to obtain a fourth machine learning model, so that the fourth machine learning model can simultaneously learn the abilities of character relationship analysis and emotional type judgment, and improve the analysis accuracy of character relationships and plot details in the script, and emotional types of characters.
[0079] Taking a natural language processing scenario as an example, the first prediction task can be to classify text (such as news classification, document classification, etc.); and the second prediction task is to perform sentiment analysis on the text (such as positive, negative, and neutral sentiment). The terminal 400 is configured to train the first machine learning model in two stages based on sample data of the two prediction tasks (such as news articles, user comments, etc.), to obtain a third machine learning model; and to perform model parameter fusion on the third machine learning model and the second machine learning model, to obtain a fourth machine learning model, so that the fourth machine learning model can simultaneously learn text classification and sentiment analysis, and improve the generalization ability and accuracy of the machine learning model.
[0080] Taking an image recognition scenario as an example, the first prediction task can be to detect objects in an image (such as vehicles, pedestrians, etc.); and the second prediction task is to identify attributes of the detected objects (such as color and brand of a vehicle). The terminal 400 is configured to train the first machine learning model in two stages based on sample data of the two prediction tasks (such as video frames captured by a camera, street view images, etc.), to obtain a third machine learning model; and to perform model parameter fusion on the third machine learning model and the second machine learning model, to obtain a fourth machine learning model, so that the fourth machine learning model can simultaneously learn object detection and attribute recognition, and improve the overall recognition efficiency and accuracy.
[0081] Taking an intelligent transportation scenario as an example, the first prediction task can be to predict traffic flow (a regression task); and the second prediction task is to predict the probability of occurrence of traffic accidents (a classification task). The terminal 400 is configured to train the first machine learning model in two stages based on sample data of the two prediction tasks, to obtain a third machine learning model; and to perform model parameter fusion on the third machine learning model and the second machine learning model, to obtain a fourth machine learning model, so that the fourth machine learning model can simultaneously learn traffic flow and accident prediction, and improve the efficiency and safety of traffic management.
[0082] Taking an e-commerce scenario as an example, the first prediction task can be to predict whether a user will purchase a certain product (a classification task), and the second prediction task is to recommend products for the user (a multi-classification task). The terminal 400 is configured to train the first machine learning model in two stages based on sample data (such as the user's historical purchase records, browsing behavior, product features, etc.) of the two prediction tasks, to obtain a third machine learning model; and to fuse model parameters of the third machine learning model and the second machine learning model to obtain a fourth machine learning model, so that the fourth machine learning model can learn both the user's purchase records and product recommendations, improving user experience and sales conversion rate.
[0083] The terminal 400 is configured to perform first-stage iterative training of the first machine learning model based on sample data corresponding to the two different prediction tasks, to obtain a second machine learning model; to determine the number of iterations to be performed based on a parameter change of the second machine learning model; to perform second-stage iterative training of the second machine learning model based on the first loss value of the second machine learning model and the number of iterations to be performed, to obtain a third machine learning model; to fuse model parameters of the third machine learning model and the second machine learning model to obtain a fourth machine learning model; and to display the training state of the machine learning model in the iterative training process on the graphical interface 411.
[0084] In some embodiments, the terminal 400 is configured to send first sample data of the first prediction task and second sample data of the second prediction task to the server 200, and the server 200 is configured to perform first-stage iterative training of the first machine learning model based on sample data corresponding to the two prediction tasks, to obtain a second machine learning model; to determine the number of iterations to be performed based on a parameter change of the second machine learning model; to perform second-stage iterative training of the second machine learning model based on the first loss value of the second machine learning model and the number of iterations to be performed, to obtain a third machine learning model; to fuse model parameters of the third machine learning model and the second machine learning model to obtain a fourth machine learning model; and to display the training state of the machine learning model in the iterative training process on the graphical interface 411, send it to the terminal 400, and display the training state in the training process in real time on the graphical interface 411.
[0085] In some embodiments, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the embodiments of the present application.
[0086] Referring to Figure 2 , Figure 2 is a structural schematic diagram of the server 200 provided by the embodiments of the present application, Figure 2 The server 200 shown in the figure includes at least one processor 210, a memory 230, and at least one network interface 220. The various components in the server 200 are coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between the components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 240 in the figure. Figure 2
[0087] The processor 210 can be an integrated circuit chip with signal processing capability, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor.
[0088] The memory 230 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disk drives, etc. The memory 230 can optionally include one or more storage devices that are physically located away from the processor 210.
[0089] The memory 230 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.
[0090] In some embodiments, the memory 230 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, which are exemplarily illustrated below.
[0091] The operating system 231 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks;
[0092] The network communication module 232 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), and the like;
[0093] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in a software manner, Figure 2 A training apparatus 233 of a machine learning model stored in the memory 230 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a first training module 2331, a determination module 2332, a second training module 2333, and a fusion module 2334, which are logical, and thus can be combined or further split according to the functions implemented. The functions of each module will be described below.
[0094] In some embodiments, the terminal or server can implement the machine learning model training method provided by the embodiments of the present application by running various computer executable instructions or computer programs. For example, the computer executable instructions can be microprogram level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; can be a native (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run; or can be a small program that can be embedded into any APP, that is, a program that only needs to be downloaded into a browser environment to run. In summary, the above computer executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins.
[0095] Next, the machine learning model training method provided by the embodiments of the present application is described. As described earlier, the electronic device implementing the machine learning model training method of the embodiments of the present application can be a terminal or a server, or a combination of the two. Therefore, the execution subject of each step will not be repeated in the following description.
[0096] The exemplary application and implementation of the terminal provided by the embodiments of the present application will be described to illustrate the training method of the machine learning model provided by the embodiments of the present application.
[0097] Referring to Figure 3A , Figure 3A is a first flowchart of the training method of the machine learning model provided by the embodiments of the present application, and the server is taken as an example of an execution subject to illustrate the steps shown in combination with Figure 3A .
[0098] In step 101, the first machine learning model is iteratively trained in the first stage by the first sample data of the first prediction task and the second sample data of the second prediction task to obtain the second machine learning model, wherein the first machine learning model is trained based on the first sample data.
[0099] Here, the first prediction task is different from the second prediction task, and the first sample data is different from the second sample data. For example, the types of prediction tasks can include image classification, text classification, dialogue question and answer, text matching, text generation, etc. Taking text classification as an example, the prediction task can be news topic classification, customer review sentiment analysis, etc.
[0100] Taking the scenario of script content analysis as an example, the first prediction task can be to analyze the relationship between characters in the script (i.e., a character relationship classification task, which can include friendly, neutral, hostile, etc.); the second prediction task is to analyze the emotional type of the character in the script (i.e., an emotion classification task, which can include anger, sadness, curiosity, surprise, fear, happiness, etc.). For example, the first sample data is "character A defeats character B", and the corresponding first label is "hostile"; the second sample data is "character A: where is this?", and the corresponding second label is "curious".
[0101] In some embodiments, the first machine learning model is obtained by training a machine learning model to be trained based on the first sample data. The machine learning model to be trained can be trained in the following manner to obtain the first machine learning model: first, the machine learning model to be trained is called to predict the first sample data to obtain prediction data; then, according to the difference between the prediction data and the first label of the first sample data, a preset loss function is used to determine a loss value, and according to the loss value, the parameters of the pre-trained machine learning model are updated to obtain the first machine learning model.
[0102] As an example, the machine learning model can be a Support Vector Machine (SVM), a Random Forest, a K-Nearest Neighbors (KNN), a Multilayer Perceptron (MLP), a Convolutional Neural Network (CNN), a Gradient Boosting Machine (GBM), a Deep Neural Network (DNN), a Long Short-Term Memory (LSTM), or the like. The loss function used to determine the loss value can be a Mean Squared Error (MSE), a Root Mean Squared Error (RMSE), a Mean Absolute Error (MAE), a Cross-Entropy Loss, or the like.
[0103] In some embodiments, referring to Figure 3B , Figure 3B FIG. 2 is a second flowchart of a method for training a machine learning model according to an embodiment of the present application. Figure 3A Step 101 of FIG. 1 can be implemented by steps 1011 to 1016 of FIG. 2, which are described in detail below. Figure 3B
[0104] In step 1011, a first prediction task is performed by calling a first machine learning model with first sample data to obtain first prediction data.
[0105] In some embodiments, the first sample data can be first extracted to obtain first data features. For example, for time series data, the year, month, day, hour, etc. in the timestamp can be extracted as features; for image data, the edge, texture, etc. can be extracted as features. Then, the first data features are input into the first machine learning model to perform the first prediction task and obtain the first prediction data.
[0106] Here, the structure of the first machine learning model can include an input layer, a hidden layer, and an output layer. The first data features are input into the input layer, the first data features are linearly processed by the hidden layer to obtain first linear features, the first linear features are activated by an activation function to obtain first output features of the hidden layer, and finally, the first output features of the hidden layer are normalized by the output layer to obtain the first prediction data.
[0107] Taking the scene of analyzing the content of a script as an example, the first sample data is "character A repelled character B", and the first label corresponding to the first sample data is "hostile". First, the first sample data can be subjected to word segmentation processing, and a feature extraction method can be used to map each word to a low-dimensional vector space. The word embedding vectors obtained by mapping each word are fused to obtain a first data feature.
[0108] For example, the word embedding vector corresponding to "character A" is [0.1, 0.2, 0.3], the word embedding vector corresponding to "repelled" is [0.4, 0.5, 0.6], and the word embedding vector corresponding to "character B" is [0.7, 0.8, 0.9]. The word embedding vectors of the entire sentence can be concatenated or averaged to obtain the first data feature. For example, the average of the word embedding vectors of the above three words is [0.4, 0.5, 0.6], and the first data feature x1 is [0.4, 0.5, 0.6].
[0109] Secondly, the first data feature [0.4, 0.5, 0.6] is input into the input layer of the first machine learning model. Assuming that the hidden layer has 3 neurons, the weight matrix is The bias vector is b = [0.1, 0.2, 0.3], and the first linear feature z1 can be expressed as
[0110] Here, if the activation function is a rectified linear unit (ReLU), the first output feature of the hidden layer is Assuming that the output layer has 2 neurons (corresponding to "hostile" and "friendly" two categories respectively), the weight matrix is The bias vector is c1 = [0.1, 0.2]. The linear feature of the output layer is
[0111] Finally, the linear feature of the output layer can be normalized using the Softmax function to convert the output into a probability distribution, that is, Where the first element 0.24 represents the probability of "hostile", and the second element 0.76 represents the probability of "friendly". Therefore, the first predicted data is "friendly".
[0112] In step 1012, a second loss value is determined according to the difference between the first predicted data and the first label of the first sample data.
[0113] In some embodiments, the first prediction data can be encoded first to obtain a first prediction vector, and the first label can be encoded to obtain a first label vector. The distance between the first prediction vector and the first label vector can be determined by a distance measurement method as a difference between the first prediction data and the first label of the first sample data. According to the difference, a second loss value can be determined by using a preset loss function.
[0114] As an example, the distance measurement method can be Euclidean distance (ED), Hamming distance (HD), Manhattan distance (MD), or Pearson correlation coefficient. The loss function used to determine the loss value can be a mean square error, a root mean square error, a mean absolute error, a cross-entropy loss, or the like.
[0115] Taking the Euclidean distance as an example, if the first prediction vector is represented as (x1, y1) and the first label vector is represented as (x2, y2), the distance between the first prediction vector and the first label vector can be represented as
[0116] Taking the example of the above step 1011, the first prediction data is "friendly", the first label is "hostile", and if one-hot encoding is used, the first prediction vector obtained by encoding the first prediction data is [0.24, 0.76], and the first label vector is represented as (1, 0). The difference between the first prediction data and the first label of the first sample data is 1.073. The second loss value determined by using the cross-entropy loss can be represented as where y i is the first label vector, y′ i is the first prediction vector, and the corresponding vectors are substituted to obtain the second loss value of 1.427.
[0117] In step 1013, the first machine learning model is called by the second sample data to perform a second prediction task to obtain second prediction data.
[0118] In some embodiments, the second data feature can be extracted from the second sample data first. Then, the second data feature can be used as an input of the first machine learning model to perform the second prediction task to obtain the second prediction data.
[0119] Here, in a case where the structure of the first machine learning model includes an input layer, a hidden layer, and an output layer, the second data feature is taken as input of the input layer, the second data feature is linearly processed through the hidden layer to obtain a second linear feature, and the second linear feature is activated by using an activation function to obtain a second output feature of the hidden layer. Finally, the second output feature of the hidden layer is normalized by the output layer to obtain second prediction data.
[0120] Taking the scenario of analyzing the content of a script as an example, the second sample data is "Role A: Where is this?", and the corresponding second label is "curiosity". First, the second sample data can be subjected to word segmentation processing, and a feature extraction method is used to map each word to a low-dimensional vector space. The word embedding vectors obtained by mapping each word are fused to obtain the second data feature.
[0121] For example, the word embedding vector corresponding to "Role A" obtained by word segmentation of the second sample data is [0.1, 0.2, 0.3], the word embedding vector corresponding to "this" is [0.4, 0.5, 0.6], the word embedding vector corresponding to "is" is [0.7, 0.8, 0.9], the word embedding vector corresponding to "where" is [1.0, 1.1, 1.2], and the word embedding vector corresponding to "?" is [1.3, 1.4, 1.5]. The word embedding vectors of the entire sentence can be averaged to obtain the second data feature x2 as [0.7, 0.8, 0.9].
[0122] Secondly, the second data feature [0.7, 0.8, 0.9] is input to the input layer of the first machine learning model, and the second linear feature z2 can be represented as z2 = W·x2 + b. If the activation function is a rectified linear function, the second output feature of the hidden layer is Assuming that the output layer has 2 neurons (corresponding to the "curiosity" and "other" two categories respectively), the linear feature of the output layer is
[0123] Finally, the Softmax function can be used to normalize the linear feature of the output layer to convert the output into a probability distribution, that is, Wherein, the first element 0.27 represents the probability of "curiosity", and the second element 0.73 represents the probability of "other". Therefore, the second prediction data is "other".
[0124] In step 1014, a third loss value is determined according to the difference between the second prediction data and the second label of the second sample data.
[0125] In some embodiments, the second prediction data can be encoded to obtain a second prediction vector, the second label can be encoded to obtain a second label vector, a distance between the first prediction vector and the second label vector can be determined by a distance measurement method as a difference between the second prediction data and the second label of the second sample data, and a third loss value can be determined by using a preset loss function according to the difference.
[0126] As an example, the distance measurement method can be Euclidean distance, Hamming distance, Manhattan distance or Pearson correlation coefficient. The loss function used to determine the loss value can be a mean square error, a root mean square error, a mean absolute error, a cross-entropy loss or the like.
[0127] Taking the example of the above step 1013, the second sample data is "Role A: Where is this?", the corresponding second label is "curiosity", if using one-hot encoding, the first prediction vector obtained by encoding the second prediction data is [0.24, 0.76], and the first label vector is represented as (1, 0), then the difference between the first prediction data and the first label of the first sample data is 1.073. The second loss value determined by using the cross-entropy loss can be represented as where y i is the first label vector, y′ i is the first prediction vector, and the corresponding vectors are substituted to obtain the second loss value of 1.427.
[0128] In step 1015, the second loss value and the third loss value are weighted and summed to obtain a fourth loss value.
[0129] In some embodiments, a first loss weight preset for the second loss value and a second loss weight preset for the third loss value can be obtained, and the second loss value and the third loss value are weighted and summed to obtain the fourth loss value.
[0130] As an example, the first loss value is a, the first loss weight is w1, the second loss value is b, and the second loss weight is w2. The fourth loss value can be represented as a*w1+b*w2.
[0131] In step 1016, the parameters of the first machine learning model are updated based on the fourth loss value to obtain a second machine learning model.
[0132] In some embodiments, based on the fourth loss value, the parameters of the first machine learning model can be updated by using a back propagation algorithm to obtain the second machine learning model. Specifically, the fourth loss value is passed from the output layer to the input layer, so that the first machine learning model can learn the difference between the first predicted data and the first label, and the difference between the second predicted data and the second label; then, according to the fourth loss value, the gradient of each parameter in the first machine learning model is calculated, and the parameters of the first machine learning model are updated according to the gradient of the parameters and the preset learning rate by using an optimization algorithm, so as to reduce the fourth loss value, until the fourth loss value of the first machine learning model converges or reaches the preset number of training rounds.
[0133] The embodiments of the present application determine the second loss value by comparing the difference between the first predicted data of the first machine learning model performing the first prediction task and the first label of the first sample data, determine the third loss value by comparing the difference between the second predicted data of the first machine learning model performing the second prediction task and the second label of the second sample data, and then determine the fourth loss value according to the second loss value and the third loss value. Based on the fourth loss value, the parameters of the first machine learning model are updated to obtain the second machine learning model. The difference between the predicted data of the first machine learning model for different prediction tasks and the corresponding label is quantified by the loss value, and the loss values of the two tasks are combined to form a comprehensive loss value to update the model parameters, so that the obtained second machine learning model can better optimize on the two prediction tasks.
[0134] In the script content analysis scene, the machine learning model is trained by using the first sample data and the second sample data, so that the machine learning model can learn the relevant features of the role relationship and the emotion type at the same time. Through multi-task learning, the underlying feature extraction layer (such as the word embedding layer and the hidden layer) can be shared, the number of parameters of the machine learning model and the calculation complexity in the training process can be reduced, the machine learning model can more efficiently utilize the computing resources in the training and inference stages, the calculation efficiency is significantly improved, and better performance can be achieved on the two tasks. The machine learning model trained can understand the script from multiple perspectives and better adapt to different types of script content. For example, for a script containing complex emotions and role relationships, the machine learning model can analyze more comprehensively.
[0135] With reference to Figure 3A , the step 101 is continued to be described.
[0136] In step 102, the first loss value of the second machine learning model and the parameter change amount are determined.
[0137] In some embodiments, with reference to Figure 3C , Figure 3CFIG. 3 is a third flowchart of a method for training a machine learning model according to an embodiment of the present application. Figure 3A Step 102 of FIG. 1 can be implemented by steps 1021-1023 of FIG. 2. Figure 3C Steps 1021-1023 of FIG. 2 are described in detail as follows.
[0138] In step 1021, a second machine learning model is called to perform a first prediction task on the first sample data to obtain third prediction data.
[0139] In some embodiments, the first sample data can be first extracted to obtain first data features. For example, for time series data, the year, month, day, hour, etc. in the timestamp can be extracted as features; for image data, edge, texture, etc. can be extracted as features. Then, the first data features are taken as the input of the second machine learning model to perform the first prediction task to obtain the third prediction data.
[0140] Here, the structure of the second machine learning model can include an input layer, a hidden layer, and an output layer. The first data features are taken as the input of the input layer of the second machine learning model, the first data features are linearly processed by the hidden layer of the second machine learning model to obtain third linear features, the third linear features are activated by an activation function to obtain third output features of the hidden layer of the second machine learning model, and finally, the third output features of the hidden layer are normalized by the output layer to obtain the third prediction data.
[0141] Taking a scene analysis of a script content as an example, the first sample data is "character A repels character B", and the corresponding first label is "hostile". The first data features extracted from the first sample data can be input into the input layer of the second machine learning model, the first data features are linearly processed by the hidden layer of the second machine learning model to obtain third linear features, the third linear features are activated by an activation function to obtain third output features of the hidden layer of the second machine learning model, and finally, the third output features of the hidden layer are normalized by the output layer to obtain the third prediction data, for example, the third prediction data is "friendly". For specific implementation, refer to the calculation process of step 1011 described above, which will not be repeated here.
[0142] In step 1022, a first loss value of the second machine learning model is determined according to the difference between the third prediction data and the first label of the first sample data.
[0143] In some embodiments, the third prediction data can be encoded to obtain a third prediction vector, the first label can be encoded to obtain a first label vector, a distance between the third prediction vector and the first label vector can be determined as a difference between the third prediction data and the first label of the first sample data by a distance measurement method, and a first loss value of the second machine learning model can be determined according to the difference by using a preset loss function.
[0144] As an example, the distance measurement method can be Euclidean distance, Hamming distance, Manhattan distance, or Pearson correlation coefficient.
[0145] In step 1023, a variation of the parameters of the second machine learning model relative to the parameters of the first machine learning model is determined as a parameter variation of the second machine learning model.
[0146] In some embodiments, the parameters of the machine learning model refer to variables that can be adjusted in the machine learning model, which determine the prediction ability of the machine learning model, and different machine learning models have different parameter types. For example, the parameters of a linear regression model include coefficients and intercepts; the parameters of a neural network include weights and biases between neurons; and the parameters of a decision tree include the structure of the tree and the split points of the nodes.
[0147] As an example, the first machine learning model and the second machine learning model are linear regression models, the parameters include weights and biases, and if the mathematical expression of the linear regression model is y = w1x1 + w2x2 + … + w n x n + b, where y is the output, w1, w2, …, w n is the weight, x1, x2, …, x n is the input, and b is the bias. In the case of input x1 and x2, the weight bias b (1) of the first machine learning model is 1; the weight bias b (2) of the second machine learning model is 1.5; and the variation of the parameters of the second machine learning model relative to the weight parameter w1 of the first machine learning model is 0.5, the variation of the weight parameter w2 is 0.5, and the variation of the bias parameter b is 0.5.
[0148] In this embodiment, the second machine learning model is invoked to perform a first prediction task using first sample data to obtain third prediction data. Based on the difference between the first label of the third prediction data and the first sample data, the first loss value of the second machine learning model is determined, which can accurately evaluate the performance of the second machine learning model on the first prediction task. The change in the parameters of the second machine learning model relative to the parameters of the first machine learning model is used as the change in the parameters of the second machine learning model, which can intuitively reflect the degree of parameter adjustment of the machine learning model in the multi-task learning process.
[0149] See also Figure 3A The following will be an explanation following step 102 above.
[0150] In step 103, the number of iterations to be performed on the second machine learning model is determined based on the amount of parameter change.
[0151] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the training method for the machine learning model provided in the embodiments of this application. Figure 3A Step 103 can be achieved through Figure 3D Steps 1031 to 1034 are implemented, and the details are explained below.
[0152] In step 1031, if multiple model parameter fusion operations are performed during the iterative training in the first stage, the parameter change between the parameter fusion results of two adjacent model parameter fusions is determined.
[0153] Here, the parameter change between the parameter fusion results of two adjacent model parameter fusions is the difference between the parameters of the machine learning model after the i-th iteration training and the parameters of the machine learning model after the j-th iteration training. Model parameter fusion represents the fusion of the machine learning model after the i-th iteration training and the machine learning model after the j-th iteration training to obtain the initial machine learning model after the (j+1)-th iteration training, where i and j are both positive integers and i≠j.
[0154] As an example, during the iterative training of the initial machine learning model, if the first model parameter fusion occurs after the 5th iteration, the initial machine learning model and the machine learning model after the 5th iteration are fused, and the fused machine learning model is used as the initial machine learning model for the 6th iteration. If the second model parameter fusion occurs after the 10th iteration, the initial machine learning model after the 6th iteration is fused, and the machine learning model after the 10th iteration is used as the initial machine learning model for the 11th iteration.
[0155] In step 1032, the change index parameter is determined based on the parameter change amount.
[0156] In some embodiments, referring to Figure 3E , Figure 3E FIG. 5 is a fifth flow diagram of a method for training a machine learning model according to an embodiment of the present application. Figure 3D Step 1032 of FIG. 10 can be implemented by Figure 3E Steps 10321-10322 of FIG. 10 are described in detail as follows.
[0157] In step 10321, the sum of the absolute values of each sub-parameter change amount in the parameter change amount is determined.
[0158] In some embodiments, for different machine learning models, the parameter change amount can include multiple sub-parameter change amounts. For example, for a linear regression model, the multiple sub-parameter change amounts included in the parameter change amount can include a sub-parameter change amount corresponding to a weight parameter and a sub-parameter change amount corresponding to a bias parameter.
[0159] Continuing with the above example, the change amount of the weight parameter w1 of the second machine learning model relative to the first machine learning model is 0.5, the change amount of the weight parameter w2 is 0.5, and the change amount of the bias parameter b is 0.5, i.e., the sub-parameter change amount corresponding to the weight parameter w1 is 0.5, the sub-parameter change amount corresponding to the weight parameter w2 is 0.5, and the sub-parameter change amount corresponding to the bias parameter b is 0.5. Then, the sum of the absolute values of each sub-parameter change amount in the parameter change amount is 1.5.
[0160] In step 10322, the ratio of the sum to the number of iterations to be performed is determined as the change index parameter.
[0161] As an example, if the number of iterations to be performed is 10 and the sum of the absolute values of each sub-parameter change amount in the parameter change amount is 1.5, then the change index parameter is 0.15.
[0162] The sum of the absolute values of each sub-parameter change amount in the parameter change amount is determined according to the embodiments of the present application. The ratio of the sum to the number of iterations to be performed is determined as the change index parameter, which can accurately quantify the degree of change of the machine learning model parameters, thereby facilitating the timely discovery of abnormal changes in the model parameters, avoiding overfitting or underfitting, and improving the stability of the training.
[0163] Continuing to refer to Figure 3D , the step 1032 is described in continuation.
[0164] In step 1033, the parameter change trend is determined according to the change index parameter between each adjacent two model parameter fusions.
[0165] In some embodiments, referring toFigure 3F , Figure 3F is a sixth flowchart of a method for training a machine learning model according to an embodiment of the present application. Figure 3D Step 1033 of the method can be implemented by Figure 3F Steps 10331 to 10335 of the method can be implemented as follows.
[0166] In step 10331, the change indicator parameter obtained by each model parameter fusion is fitted into a first change curve according to the order of model parameter fusion, where the ordinate of the first change curve represents the change indicator parameter between adjacent model parameter fusions, and the abscissa represents the iteration number.
[0167] In some embodiments, during the process of multiple model parameter fusions, each model parameter fusion produces a change indicator parameter for measuring the degree of parameter change between adjacent model parameter fusions. In order to more intuitively observe and analyze the evolution trend of these change indicator parameters with the fusion process, these change indicator parameters can be arranged according to the order of model parameter fusion operation, and a first change curve is fitted. The ordinate clearly represents the change indicator parameter between adjacent model parameter fusions, which can reflect the key information such as the change amplitude or direction of model parameters at each iteration in the model parameter fusion process; the abscissa represents the iteration number, that is, the sequential number of model parameter fusion operation, and the entire process of model parameter fusion can be clearly detected by the iteration number.
[0168] As an example, refer to Figure 10 , Figure 10 is a schematic diagram of a first change curve according to an embodiment of the present application. As shown in Figure 10 the abscissa of the first change curve is the iteration number, and the ordinate is the change indicator parameter. The ordinate of the value of each point in the first change curve represents the change indicator parameter of the current model parameter fusion compared with the previous model parameter fusion, and the abscissa represents the iteration number of the second machine learning model corresponding to the current model parameter fusion.
[0169] By drawing the first change curve, the dynamic change of the change indicator parameter in the model parameter fusion process can be intuitively observed, so that the effect and trend of the model parameter fusion can be better understood, and a powerful reference basis is provided for the optimization and adjustment of the iteration number of the subsequent second machine learning model.
[0170] In step 10332, the first change curve is sampled multiple times by a sliding window, and the number of curves belonging to the rising section and the number of curves belonging to the falling section are counted.
[0171] In some embodiments, the first change curve diagram can be sampled multiple times by a preset length of a sliding window, and the number of curves belonging to the rising section and the number of curves belonging to the falling section obtained by sampling are counted.
[0172] As an example, continuing to refer to Figure 10 As shown in the left graph, in the sampling range of the sliding window 1001, the number of rising sections is 3, and the number of falling sections is 2; as shown in the right graph, in the sampling range of the sliding window 1001, the number of rising sections is 2, and the number of falling sections is 3. Figure 10 Figure 10 As shown in the left graph, in the sampling range of the sliding window 1001, the number of rising sections is 3, and the number of falling sections is 2; as shown in the right graph, in the sampling range of the sliding window 1001, the number of rising sections is 2, and the number of falling sections is 3.
[0173] In step 10333, in response to the number of rising sections being greater than the number of falling sections, it is determined that the parameter change trend is an upward trend.
[0174] As an example, continuing to refer to Figure 10 , in Figure 10 In the left graph, the number of rising sections in the sampling range of the sliding window 1001 is 3, and the number of falling sections is 2, so the parameter change trend is an upward trend.
[0175] In step 10334, in response to the number of rising sections being less than the number of falling sections, it is determined that the parameter change trend is a downward trend.
[0176] As an example, continuing to refer to Figure 10 , in Figure 10 In the right graph, the number of rising sections in the sampling range of the sliding window 1001 is 2, and the number of falling sections is 3, so the parameter change trend is a downward trend.
[0177] In step 10335, in response to the number of rising sections being equal to the number of falling sections, it is determined that the parameter change trend is no change.
[0178] As an example, if the number of rising sections and the number of falling sections are both 2, it is determined that the parameter change trend is no change.
[0179] According to the model parameter fusion change index parameter, the first change curve diagram is fitted, the first change curve diagram is sampled multiple times by a sliding window, the number of curves belonging to the rising section and the number of curves belonging to the falling section obtained by sampling are counted, and the parameter change trend is determined according to the number. The abnormal change of the model parameter can be found in time, overfitting or underfitting can be avoided, the stability of training can be improved, the training strategy can be dynamically adjusted, the stability of the model can be improved, the model parameter can be optimized, and the generalization ability and adaptability of the model can be improved.
[0180] Continuing to refer to Figure 3D , the step 1033 is continued.
[0181] In step 1034, according to the preset adjustment strategy of the parameter change trend, the interval iteration number of the last two model parameter fusions in the multiple model parameter fusions is updated to obtain the to-be-iterated number of the second machine learning model, wherein the interval iteration number of the last two model parameter fusions is the difference between the number of iteration training performed before the last model parameter fusion and the number of iteration training performed before the second last model parameter fusion.
[0182] As an example of the interval iteration number of the last two model parameter fusions in the multiple model parameter fusions, if the number of iteration training performed before the last model parameter fusion is 30 times, and the number of iteration training performed before the second last model parameter fusion is 20 times, then the interval iteration number of the last two model parameter fusions is 10 times.
[0183] Here, different preset adjustment strategies are associated with different parameter change trends, which will be described in detail below.
[0184] In some embodiments, referring to Figure 3G , Figure 3G is a seventh flowchart of the training method of the machine learning model provided in the embodiments of the present application. Figure 3D Step 1034 of the method 1000 can be implemented by steps 10341 to 10343 of the method 1000, which will be described in detail below. Figure 3G
[0185] In step 10341, in response to the parameter change trend being an upward trend, the ratio of the interval iteration number to the preset step length adjustment factor is determined, and the maximum value of the ratio and the preset minimum iteration number is taken as the to-be-iterated number of the second machine learning model.
[0186] Here, when the parameter change trend is an upward trend, if the interval iteration number is represented as S b , the preset step length adjustment factor is represented as γ learn , and the preset minimum iteration number is represented as S min , then the ratio of the interval iteration number to the preset step length adjustment factor can be represented as S b / γ learn , and the maximum value of the ratio and the preset minimum iteration number can be represented as max(S min ,S b / γ learn ), that is, the to-be-iterated number of the second machine learning model is max(S min ,S b / γ learn .
[0187] As an example, if the interval iteration number is denoted as 20, the preset step length adjustment factor is denoted as 2, and the preset minimum iteration number is denoted as 5, the ratio of the interval iteration number to the preset step length adjustment factor is 10, and the maximum value of the ratio and the preset minimum iteration number is 10, that is, the iteration number of the second machine learning model to be performed is 10.
[0188] In step 10342, in response to the parameter change trend being a downward trend, a product of the interval iteration number and the preset step length adjustment factor is determined, and the minimum value of the product and the preset maximum iteration number is determined as the iteration number to be performed.
[0189] Here, when the parameter change trend is a downward trend, if the interval iteration number is denoted as S b , the preset step length adjustment factor is denoted as γ learn , and the preset maximum iteration number is denoted as S max , the product of the interval iteration number and the preset step length adjustment factor can be denoted as S b *γ learn , and the minimum value of the product and the preset maximum iteration number can be denoted as min(S max , S b *γ learn ), that is, the iteration number of the second machine learning model to be performed is min(S max , S b *γ learn ).
[0190] As an example, if the interval iteration number is denoted as 20, the preset step length adjustment factor is denoted as 2, and the preset maximum iteration number is denoted as 128, the product of the interval iteration number and the preset step length adjustment factor is 40, and the minimum value of the product and the preset maximum iteration number is 40, that is, the iteration number of the second machine learning model to be performed is 40.
[0191] In step 10343, in response to the parameter change trend being no change, the interval iteration number is determined as the iteration number to be performed.
[0192] As an example, when the parameter change trend is no change, if the interval iteration number is denoted as S b , the interval iteration number S b is determined as the iteration number to be performed.
[0193] Continuing to refer to Figure 3A , the step 103 is continued.
[0194] In step 104, the second machine learning model is iteratively trained in the second stage based on the first loss value and the iteration number to be performed through the first sample data and the second sample data, and a third machine learning model is obtained.
[0195] In some embodiments, referring to Figure 3H , Figure 3H is the eighth flowchart of the method for training the machine learning model provided in the embodiments of the present application. Figure 3A Step 104 of the method can perform second-stage iterative training on the second machine learning model based on the first sample data, wherein after each iteration of the second stage, the method performs Figure 3H Steps 1041A to 1043A of the method are implemented, which are described in detail below.
[0196] In step 1041A, a first over-limit count value is determined, wherein after any iteration of the second stage, if the first loss value of the second machine learning model after the iteration is greater than the preset loss threshold, the first over-limit count value is incremented by 1.
[0197] In some embodiments, during the process of performing iterative training on the second machine learning model based on the number of iterations to be performed in the second stage, after each iteration of the second machine learning model based on the first sample data, the method can compare the first loss value of the second machine learning model with the preset loss threshold, and in the case where the first loss value is greater than the preset loss threshold, the first over-limit count value is incremented by 1. Here, the initial value of the first over-limit count value is 0, and the increment of the first over-limit count value indicates that the first loss value of the second machine learning model after the current iteration is greater than the preset loss threshold.
[0198] As an example, if the first loss value of the second machine learning model obtained after the fifth iteration of the second machine learning model is greater than the preset loss threshold, the first over-limit count value is 1; the first loss value of the second machine learning model obtained after the sixth iteration of the second machine learning model is less than the preset loss threshold, the first over-limit count value is still 1, and the next iteration is continued until the first loss value of the second machine learning model obtained after the tenth iteration of the second machine learning model is again greater than the preset loss threshold, and the first over-limit count value is updated to 2.
[0199] In step 1042A, if the first over-limit count value is greater than 0 and less than a preset iteration threshold, the number of iterations to be performed is taken as the target number of iterations.
[0200] As an example, if the preset iteration threshold is 3, and during the process of performing iterative training on the second machine learning model based on the number of iterations to be performed in the second stage, the first over-limit count value is 1, the number of iterations to be performed is taken as the target number of iterations, i.e. the number of iterations to be performed on the second machine learning model is not adjusted, and after performing the number of iterations to be performed on the second machine learning model based on the second sample data, the third machine learning model can be obtained.
[0201] In step 1043A, the second machine learning model is iteratively trained by the second sample data for a target number of iterations to obtain a third machine learning model.
[0202] In the case that the first over-limit count value is greater than 0 and less than a preset threshold value, the second machine learning model is iteratively trained by the second sample data for a to-be-iterated number of iterations to obtain a third machine learning model.
[0203] In some embodiments, referring to Figure 3L , Figure 3L FIG. 12 is a twelfth flowchart of a method for training a machine learning model according to an embodiment of the present application. Figure 3H Step 1043A of FIG. 10 can be implemented by iteratively training the second machine learning model by the second sample data for a target number of iterations, and in each iteration, performing Figure 3L Steps 10431 to 10433 of FIG. 10 can be implemented as follows.
[0204] In step 10431, the initialized machine learning model of the n th iteration is called to predict the second sample data to obtain n th prediction data.
[0205] In some embodiments, in the process of iteratively training the second machine learning model by the second sample data for a target number of iterations, the machine learning model after each iteration is taken as the initialized machine learning model of the next iteration. When n is 1, the initialized machine learning model of the first iteration is the second machine learning model.
[0206] In step 10432, an n th loss value is determined according to the difference between the n th prediction data and the second label of the second sample data.
[0207] In some embodiments, the n th prediction data can be encoded to obtain an n th prediction vector, and the second label of the second sample data can be encoded to obtain a second label vector. The distance between the n th prediction vector and the second label vector can be determined by a distance measurement method, and taken as the difference between the n th prediction data and the second label of the second sample data. According to the difference, the n th loss value can be determined by using a preset loss function.
[0208] As an example, the distance measurement method can be Euclidean distance, Hamming distance, Manhattan distance or Pearson correlation coefficient. Taking Euclidean distance as an example, if the n th prediction vector is represented as (x n , y n ), and the second label vector is represented as (x2, y2), then the distance between the n th prediction vector and the second label vector can be represented as
[0209] In step 10433, the parameters of the initialized machine learning model trained in the nth iteration are updated according to the nth loss value to obtain an updated machine learning model, wherein when n is 1, the initial machine learning model is the second machine learning model, n≤N, n and N are both positive integers, N is the target iteration number, and when n=N, the updated machine learning model is the third machine learning model.
[0210] In some embodiments, based on the nth loss value, the parameters of the initialized machine learning model trained in the nth iteration can be updated using a back propagation algorithm to obtain an updated machine learning model. Specifically, the nth loss value is passed from the output layer to the input layer, so that the initialized machine learning model trained in the nth iteration can learn the difference between the nth predicted data and the second label; then, according to the nth loss value, the gradient of each parameter in the initialized machine learning model trained in the nth iteration is calculated, and the parameters of the initialized machine learning model trained in the nth iteration are updated using an optimization algorithm according to the gradient of the parameters and a preset learning rate, so as to reduce the nth loss value, until the nth loss value converges or reaches a preset training number of rounds.
[0211] As an example, first, the second machine learning model is taken as the initialized machine learning model trained in the first iteration, the second machine learning model is called to predict the second sample data to obtain the first predicted data; the first loss value is determined according to the difference between the first predicted data and the second label of the second sample data; and the parameters of the second machine learning model are updated according to the first loss value to obtain the updated machine learning model trained in the first iteration. Then, the updated machine learning model trained in the first iteration is taken as the initialized machine learning model trained in the second iteration, the updated machine learning model trained in the first iteration is called to predict the second sample data to obtain the second predicted data; the second loss value is determined according to the difference between the second predicted data and the second label of the second sample data; and the parameters of the updated machine learning model trained in the first iteration are updated according to the second loss value to obtain the updated machine learning model trained in the second iteration, and the next training is continued. Until the Nth iteration training is performed to obtain the updated machine learning model trained in the Nth iteration as the third machine learning model.
[0212] The second machine learning model is iteratively trained for the target number of iterations by the second sample data. In each iteration process, the loss value is determined according to the difference between the prediction data and the second label, and then the machine learning model is updated according to the loss value, and finally the third machine learning model is obtained. Through the iterative training of the second machine learning model on the second sample data for the target number of iterations, the adaptability of the machine learning model can be significantly improved, the parameters of the machine learning model can be optimized, the convergence speed and stability of the machine learning model can be improved, the generalization ability of the machine learning model can be improved, and the waste of computing resources can be reduced.
[0213] In some embodiments, referring to Figure 3I , Figure 3I FIG. 9 is a ninth flow diagram of a method for training a machine learning model according to an embodiment of the present application. Figure 3A Step 104 of the method can further include iteratively training the second machine learning model on the first sample data for a second stage, wherein after each iteration training of the second stage, the method further includes Figure 3I Steps 1041B to 1043B of the method are implemented as follows, which are described in detail below.
[0214] In step 1041B, a second over-limit count value is determined, wherein after any iteration training of the second stage, if the first loss value of the second machine learning model after the iteration training is greater than the preset loss threshold, the second over-limit count value is incremented by 1.
[0215] In some embodiments, during the iteration training of the second machine learning model based on the number of iterations to be iterated in the second stage, after each iteration training of the second machine learning model based on the first sample data, the first loss value of the second machine learning model can be compared with the preset loss threshold, and in the case that the first loss value is greater than the preset loss threshold, the second over-limit count value is incremented by 1. Here, the initial value of the second over-limit count value is 0, and the accumulation of the second over-limit count value indicates that the first loss value of the second machine learning model after the current iteration training is greater than the preset loss threshold.
[0216] As an example, if the first loss value of the second machine learning model obtained after the fifth iteration training of the second machine learning model is greater than the preset loss threshold, the second over-limit count value is 1; the first loss value of the second machine learning model obtained after the sixth iteration training of the second machine learning model is less than the preset loss threshold, the second over-limit count value is still 1, and the next iteration training is continued until the first loss value of the second machine learning model obtained after the tenth iteration training of the second machine learning model is again greater than the preset loss threshold, and the second over-limit count value is updated to 2.
[0217] In step 1042B, if the second over-limit count value is greater than or equal to the preset count threshold, then the number of iterations corresponding to the preset count threshold is taken as the target number of iterations.
[0218] As an example, if the preset counting threshold is 3, in the second stage of iterative training of the second machine learning model based on the number of iterations to be performed, after the 19th iteration of training of the second machine learning model, the cumulative second over-limit count value is still 2, and after the 20th iteration of training of the second machine learning model, the cumulative second over-limit count value is 3. Then, 20 times is taken as the target number of iterations.
[0219] In step 1043B, the second machine learning model is iteratively trained with the second sample data for the target number of iterations to obtain the third machine learning model.
[0220] Here, during the training of the second machine learning model using the second sample data for the target number of iterations, the machine learning model trained in each iteration is used as the initial machine learning model for the next iteration. The specific implementation process is detailed in steps 10431 to 10433 above and will not be repeated here.
[0221] In some embodiments, see Figure 3J , Figure 3J This is a schematic diagram of the tenth process of the training method for the machine learning model provided in the embodiments of this application. Figure 3A Step 104 can also be executed Figure 3J Steps 1041C to 1043C are implemented, and the details are explained below.
[0222] In step 1041C, the maximum number of iterations in the second stage is initialized, wherein the maximum number of iterations is greater than the number of iterations to be performed.
[0223] Here, the maximum number of iterations in the second stage is a preset number of iterations, for example, the maximum number of iterations could be 40.
[0224] In some embodiments, the second machine learning model is subjected to a second phase of iterative training using the first sample data, wherein after each iteration of training in the second phase, the following steps 1042C to 1043C may be performed.
[0225] In step 1042C, if the number of iterations is less than the maximum number of iterations, and the first loss value of the second machine learning model after the current iteration training is less than the preset loss threshold, then continue to the next iteration training.
[0226] As an example, in the second-stage iterative training of the second machine learning model by the first sample data, if the iteration number is 20, the maximum iteration number is 40, and the first loss value of the second machine learning model after the 20th iteration training (i.e., the current iteration training) is less than the preset loss threshold, the 21st iteration training is continued.
[0227] In step 1043C, if the first loss value of the second machine learning model after the current iteration training is greater than or equal to the preset loss threshold, and the current iteration number is less than the maximum iteration number, the current iteration number is taken as the target iteration number; the second machine learning model is iteratively trained by the second sample data for the target iteration number to obtain a third machine learning model.
[0228] As an example, if the maximum iteration number is 40, the first loss value of the second machine learning model after the 20th iteration training (i.e., the current iteration training) is greater than or equal to the preset loss threshold, and 20 is less than the maximum iteration number 40, the current iteration number (i.e., 20) is taken as the target iteration number.
[0229] Here, in the training process of the second machine learning model by the second sample data for the target iteration number, the machine learning model after each iteration training is taken as the initialization machine learning model of the next iteration training. For specific implementation process, refer to steps 10431 to 10433 described above, which will not be repeated here.
[0230] In some embodiments, referring to Figure 3K , Figure 3K is the eleventh flowchart of the training method of the machine learning model provided by the embodiments of the present application. Figure 3A The step 104 of the method can also be implemented by executing Figure 3K steps 1041D to 1043D of the method, which will be described in detail below.
[0231] In step 1041D, the maximum iteration number of the second stage is initialized, wherein the maximum iteration number is greater than the iteration number to be iterated.
[0232] Here, the maximum iteration number of the second stage is a preset iteration number, for example, the maximum iteration number can be 40.
[0233] In some embodiments, the second machine learning model is iteratively trained in the second stage by the first sample data, wherein after each iteration training in the second stage, the following steps 1042D to 1043D can be performed.
[0234] In step 1042D, if the number of iterations is less than the maximum number of iterations, and the first loss value of the second machine learning model after the current iteration training is less than the preset loss threshold, the next iteration training is continued.
[0235] As an example, in the second stage of iteration training of the second machine learning model by the first sample data, if the number of iterations is 20, the maximum number of iterations is 40, and the first loss value of the second machine learning model after the 20th iteration training (i.e. the current iteration training) of the second machine learning model is less than the preset loss threshold, the 21st iteration training is continued.
[0236] In step 1043D, if the first loss value of the second machine learning model after the current iteration training is less than the preset loss threshold, and the current number of iterations is the maximum number of iterations, the maximum number of iterations is taken as the target number of iterations; the second machine learning model is trained by the second sample data for the target number of iterations to obtain a third machine learning model.
[0237] As an example, if the maximum number of iterations is 40, and the first loss value of the second machine learning model after the 40th iteration training (i.e. the current iteration training) of the second machine learning model is less than the preset loss threshold, the maximum number of iterations (i.e. 20) is taken as the target number of iterations.
[0238] Here, in the training process of the second machine learning model by the second sample data for the target number of iterations, the machine learning model after each iteration training is taken as the initialization machine learning model of the next iteration training. For specific implementation process, refer to steps 10431 to 10433 described above, which will not be repeated here.
[0239] The embodiments of the present application update the interval number of iterations between the last two model parameter fusions in the multiple model parameter fusions according to the preset adjustment strategy of the parameter variation trend, obtain the number of iterations to be performed of the second machine learning model, dynamically adjust the number of iterations by analyzing the parameter variation trend, make the training process of the machine learning model more flexible and efficient, make the machine learning model reduce the number of iterations when the parameter changes quickly, avoid unnecessary calculation; increase the number of iterations when the parameter changes slowly, speed up the convergence, so that the machine learning model can be optimized according to the parameter variation trend at different stages, improve the overall performance of the machine learning model, better cope with parameter variation, avoid unstable training caused by too fast or too slow parameter variation, reduce unnecessary calculation in the training process, improve the utilization rate of computing resources, and reduce the training cost.
[0240] Continuing to refer to Figure 3A , the step 104 is continued.
[0241] In step 105, the model parameters of the third machine learning model and the second machine learning model are fused to obtain a fourth machine learning model, wherein the fourth machine learning model is used to perform at least one of the first prediction task and the second prediction task.
[0242] Here, the fourth machine learning model can be used to predict the target data of the first prediction task to obtain the prediction result of the first prediction task, or it can be used to predict the target data of the second prediction task to obtain the prediction result of the second prediction task.
[0243] In some embodiments, see Figure 3M , Figure 3M This is a schematic diagram of the thirteenth step of the training method for the machine learning model provided in the embodiments of this application. Figure 3A Step 105, "fusing the model parameters of the third and second machine learning models to obtain the fourth machine learning model," can be achieved by executing... Figure 3M Steps 1051 to 1056 are implemented, and the details are explained below.
[0244] In step 1051, the first parameter change of the parameters of the third machine learning model relative to the parameters of the second machine learning model is determined.
[0245] In some embodiments, the parameters of a machine learning model refer to the adjustable variables in the machine learning model that determine its predictive ability. Different machine learning models have different parameter types. For example, the parameters of a linear regression model include weights and biases; the parameters of a neural network include the weights and biases between neurons; and the parameters of a decision tree include the tree structure and the split points of the nodes.
[0246] As an example, the second and third machine learning models are linear regression models, with parameters including weights and biases. If the mathematical expression of the linear regression model is y = w1x1 + w2x2 + ... + w n x n +b, where y is the output, w1, w2, ..., w n The weights are x1, x2, ..., x. n Here, 'b' is the input and 'b' is the bias. Given inputs x1 and x2, the weights of the third machine learning model... bias b (3) =2; Weights of the second machine learning model bias b (2) =1.5; then the change in the parameters of the third machine learning model relative to the weight parameter w1 of the second machine learning model is 1.5, the change in the weight parameter w2 is 1, and the change in the bias parameter b is 0.5.
[0247] In step 1052, the third machine learning model is fine-tuned using the first sample data for a preset number of iterations to obtain the fine-tuned third machine learning model.
[0248] In some embodiments, step 1052 above can be implemented in the following ways: calling the third machine learning model to perform the first prediction task through the first sample data to obtain the fourth prediction data; determining the loss value of the third machine learning model based on the difference between the first label of the fourth prediction data and the first sample data; updating the parameters of the third machine learning model based on the loss value to obtain the fine-tuned third machine learning model.
[0249] Specifically, the process of calling the third machine learning model to perform the first prediction task using the first sample data to obtain the fourth prediction data can be achieved as follows: First data features are extracted from the first sample data. These first data features are then used as input to the input layer of the third machine learning model. The first data features are linearly processed through the hidden layer of the third machine learning model to obtain the fourth linear features. An activation function is then used to activate the fourth linear features, resulting in the fourth output features of the hidden layer of the third machine learning model. Finally, the fourth output features of the hidden layer are normalized through the output layer to obtain the fourth prediction data.
[0250] In step 1053, the amount of change of the parameters of the fine-tuned third machine learning model relative to the second parameter of the third machine learning model is determined.
[0251] In some embodiments, the specific implementation of determining the change in the parameters of the fine-tuned third machine learning model relative to the second parameter of the third machine learning model can be found in the implementation process of "determining the change in the parameters of the third machine learning model relative to the first parameter of the second machine learning model" in step 1051 above, and will not be repeated here.
[0252] In step 1054, the first weight of the change in the first parameter and the second weight of the change in the second parameter are determined.
[0253] In some embodiments, see Figure 3N , Figure 3N This is a schematic diagram of the fourteenth step of the training method for the machine learning model provided in this application embodiment. (The above...) Figure 3M Step 1054 can be executed Figure 3N Steps 10541 to 10543 are implemented, and the details are explained below.
[0254] In step 10541, a first performance index is determined based on the changing trend of the parameters of the third machine learning model during the fine-tuning process of a preset number of iterations. The first performance index is used to characterize the performance of the third machine learning model in performing the second prediction task.
[0255] In some embodiments, see Figure 3O , Figure 3O This is the fifteenth flowchart illustrating the training method for the machine learning model provided in this application embodiment. (The above...) Figure 3N Step 10541 can be executed Figure 3O Steps 201 to 203 are implemented, and the details are explained below.
[0256] In step 201, based on the parameter change trend of the third machine learning model during the fine-tuning process of the preset number of iterations, a second change curve is fitted, wherein the vertical axis of the second change curve represents the parameters after each iteration, and the horizontal axis of the second change curve represents the number of iterations.
[0257] In some embodiments, during the fine-tuning of the third machine learning model for a preset number of iterations (e.g., 20), the parameters of the third machine learning model after each iteration can be recorded. To more intuitively observe and analyze the evolution trend of the parameters during the iterative training process, these parameters can be arranged according to the order of the iterative training to fit a second change curve. The vertical axis explicitly represents the parameters of the third machine learning model after the current iteration; the horizontal axis represents the number of iterations, which is the sequential number of the third machine learning model in the iterative training process. The number of iterations clearly shows the entire process of the third machine learning model's iterative training within the preset number of iterations.
[0258] As an example, see Figure 11 , Figure 11 This is a schematic diagram of the second variation curve provided in the embodiments of this application. For example... Figure 11 In the second change curve shown, the horizontal axis represents the number of iterations, and the vertical axis represents the parameters. The vertical axis of the value of each point in the second change curve represents the parameters of the third machine learning model after the current iteration number, and the horizontal axis represents the current iteration number of the third machine learning model.
[0259] In step 202, the second change curve is sampled multiple times using a sliding window, and the number of curves that belong to the rising segment and the number of curves that belong to the falling segment are counted.
[0260] In some embodiments, the second change curve can be sampled multiple times using a sliding window of a preset length, and the number of curves that belong to the rising segment and the number of curves that belong to the falling segment can be counted.
[0261] As an example, see further. Figure 11 ,like Figure 11As shown in the left figure, within the sampling range of sliding window 1002, the number of rising segments is 3 and the number of falling segments is 2; Figure 11 As shown in the figure on the right, within the sampling range of sliding window 1002, the number of rising segments is 2 and the number of falling segments is 3.
[0262] In step 203, the ratio of the number of rising segments to the total number is used as the first performance indicator, where the total number is the sum of the number of rising segments and the number of falling segments.
[0263] As an example, see further. Figure 11 In the right-hand diagram, the number of rising segments is 2 and the number of falling segments is 3. Therefore, the ratio of the number of rising segments to the total number is 0.4, which means the first performance indicator is 0.5.
[0264] This application embodiment records the parameters of the machine learning model after each iteration, plots a curve showing the change of parameters with the number of iterations, selects a fixed-size sliding window, slides along the curve, records whether the curve segment within the window is an ascending segment or a descending segment each time, counts the number, and calculates the ratio of the number of ascending segments to the total number as a quantitative indicator of the machine learning model's performance. With lower computational cost (only requiring parameter recording and sliding window statistics), it achieves more accurate convergence judgment, more efficient fine-tuning optimization, and accurately evaluates the parameter change trend of the machine learning model during the fine-tuning process, thereby quantifying the performance of the machine learning model and reducing the computational cost during model training.
[0265] See also Figure 3N The following will be an explanation following step 10541 above.
[0266] In step 10542, a second performance index is determined based on the fifth loss value of the third machine learning model during the fine-tuning process of a preset number of iterations. The second performance index is used to characterize the performance of the third machine learning model in performing the first prediction task.
[0267] In some embodiments, see Figure 3P , Figure 3P This is the sixteenth flowchart illustrating the training method for the machine learning model provided in this application embodiment. (The above...) Figure 3N Step 10542 can be executed Figure 3P Steps 301 to 302 are implemented, and the details are explained below.
[0268] In step 301, a third over-limit count value is determined. If, after any iteration of training in the fine-tuning process of a preset number of iterations, the fifth loss value of the third machine learning model after iterative training is greater than a preset loss threshold, the third over-limit count value is incremented by 1.
[0269] In some embodiments, during the preset number of iterations of training of the third machine learning model, after each iteration of training, the fifth loss value of the third machine learning model after the current iteration of training is compared with the preset loss threshold value, and in the case that the fifth loss value is greater than the preset loss threshold value, the third overflow count value is incremented by 1, where the initial value of the third overflow count value is 0.
[0270] For example, if the fifth loss value of the third machine learning model obtained after the fifth iteration of training of the third machine learning model is greater than the preset loss threshold value, the third overflow count value is 1; if the fifth loss value of the third machine learning model obtained after the sixth iteration of training of the third machine learning model is less than the preset loss threshold value, the third overflow count value is still 1, and the next iteration of training is continued until the fifth loss value of the third machine learning model obtained after the tenth iteration of training of the third machine learning model is again greater than the preset loss threshold value, and the third overflow count value is updated to 2.
[0271] In step 302, the ratio of the third overflow count value to the preset count threshold value is taken as the second performance indicator.
[0272] For example, if the preset number of iterations of iteration of training of the third machine learning model is 20 times, the preset technology threshold value is 3, and the number of times that the fifth loss value of the third machine learning model after iteration of training is greater than the preset loss threshold value during the 20 iterations of training is 6 times, i.e., the third overflow count value is 6, then the ratio of the third overflow count value to the preset count threshold value is 2, i.e., the second performance indicator is 2.
[0273] The embodiments of the present application can more accurately and intuitively reflect the stability of the model in the training process by determining the number of times (i.e., the third overflow count value) that the fifth loss value of the third machine learning model after iteration of training is greater than the preset loss threshold value, and taking the ratio of the third overflow count value to the preset count threshold value as the second performance indicator. The second performance indicator in the form of a ratio can normalize the overflow count and the preset count threshold value, so that the performance indicators in different machine learning models or different training scenarios are comparable, facilitating horizontal comparison between multiple models, and thus more accurately evaluating the performance of the machine learning model.
[0274] Continuing to refer to Figure 3N , the step 10542 is explained.
[0275] In step 10543, according to the first performance indicator and the second performance indicator, a first weight of the first parameter variation and a second weight of the second parameter variation are determined.
[0276] In some embodiments, the step 10543 can be implemented by determining a sum of the first performance indicator and the second performance indicator; taking a ratio of the first performance indicator to the sum as the first weight of the first parameter variation; and taking a ratio of the second performance indicator to the sum as the second weight of the second parameter variation.
[0277] Here, if the first performance indicator is denoted as P new , and the second performance indicator is denoted as P past , the sum of the first performance indicator and the second performance indicator can be denoted as P new + P past ; the ratio of the first performance indicator to the sum can be denoted as i.e., the first weight a1 of the first parameter variation can be denoted as The ratio of the second performance indicator to the sum can be denoted as i.e., the second weight a2 of the second parameter variation can be denoted as
[0278] Taking the above example, if the first performance indicator is 0.5 and the second performance indicator is 2, the sum of the first performance indicator and the second performance indicator is 2.5, the ratio of the first performance indicator to the sum can be denoted as 0.2, i.e., the first weight of the first parameter variation is 0.2; the ratio of the second performance indicator to the sum can be denoted as 0.8, i.e., the second weight of the second parameter variation is 0.8.
[0279] Continuing to refer to Figure 3M , the step 1054 is explained as follows.
[0280] In step 1055, a sum of a product of the parameter of the second machine learning model, the first parameter variation and the first weight, and a product of the second parameter variation and the second weight is determined.
[0281] Here, if the parameter of the second machine learning model is denoted as the first parameter variation is denoted as the second parameter variation is denoted as the first weight is denoted as a1, and the second weight is denoted as a2, the sum of the product of the parameter of the second machine learning model, the first parameter variation and the first weight, and the product of the second parameter variation and the second weight can be denoted as
[0282] In step 1056, a fourth machine learning model is constructed based on the sum.
[0283] In some embodiments, the parameters of the second machine learning model, the product of the first parameter variation and the first weight, and the product of the second parameter variation and the second weight can be added as the parameters of the fourth machine learning model to construct the fourth machine learning model.
[0284] The embodiments of the present application determine the first parameter variation of the parameters of the third machine learning model relative to the parameters of the second machine learning model, and determine the second parameter variation of the parameters of the fine-tuned third machine learning model relative to the third machine learning model, determine the first weight of the first parameter variation and the second weight of the second parameter variation, and then construct the fourth machine learning model according to the parameter variations and the corresponding weights. By calculating the parameter variations and the corresponding weights, the parameters of the machine learning models at different stages can be adjusted more finely, so that the machine learning model can achieve a good balance on multiple tasks. At the same time, by fusing the parameters of the machine learning models at different stages, the machine learning model can better cope with parameter changes, the parameter update in the training process is more smooth, the shock in the training process is reduced, and the stability of the training is improved.
[0285] In summary, after the first machine learning model is trained by the first sample data of the first prediction task, the first machine learning model is iteratively trained in the first stage by the first sample data of the first prediction task and the second sample data of the second prediction task to obtain the second machine learning model. According to the parameter variation of the second machine learning model, the number of iterations to be performed is determined. In the first stage, the model is iteratively trained by data of two different tasks, and the parameter variation can be used as a basis for adjusting the training intensity to avoid overtraining or insufficient training.
[0286] After the number of iterations to be performed is determined, the second machine learning model is iteratively trained in the second stage based on the first loss value and the number of iterations to be performed by the first sample data and the second sample data, so that the obtained third machine learning model can take into account the performance of the two prediction tasks, and the stability of the machine learning model in the multi-task training process is improved. The model parameters of the third machine learning model and the second machine learning model are fused, so that the fourth machine learning model obtained finally can make full use of the training results of the two stages, better adapt to multiple prediction tasks, and perform more balancedly when processing different tasks, thereby improving the overall adaptability of the machine learning model.
[0287] In the script content analysis scene, the first machine learning model is iteratively trained in the first stage by the first sample data of the role relationship classification task and the second sample data of the sentiment classification task to obtain a second machine learning model. According to the parameter variation of the second machine learning model, the number of iterations to be performed is determined, which can avoid over-training or under-training, avoid waste of computing power caused by repeated iterations, and ensure that the model fully learns the basic features of the two tasks, and does not cause low accuracy of sentiment classification or relationship classification due to under-training.
[0288] For example, after the 20th iteration of the machine learning model, the machine learning model has mastered the prediction ability of the role relationship classification task and the sentiment type classification task, and still continues to iterate, which will cause waste of computing resources and training time. If the number of iterations is insufficient, the machine learning model only preliminarily masters the judgment logic of the role relationship (such as "action word→relationship"), but does not fully learn the features of the sentiment classification (such as "dialogue tone, interrogative word→emotion"), which leads to low accuracy in subsequent analysis of the emotional type of the role in the script, and affects the performance of the sentiment task.
[0289] After the number of iterations to be performed is determined, the second machine learning model is iteratively trained in the second stage based on the first loss value and the number of iterations to be performed by using the first sample data and the second sample data, so that the machine learning model can train the sentiment classification task on the basis of reducing the loss of the role relationship classification task, and simultaneously process the training of the two tasks, without allocating independent computing power for the two tasks, so as to simultaneously optimize the "role relationship loss" and the "sentiment classification loss", thereby improving the overall computing power utilization rate; the third machine learning model obtained can take into account the performance of the two prediction tasks, and improves the stability of the machine learning model in the multi-task training process.
[0290] The model parameters of the third machine learning model and the second machine learning model are fused, so that the fourth machine learning model obtained finally can fully utilize the training results of the two stages and simultaneously process the role relationship classification and the sentiment classification task, only needs to load one model parameter when performing task prediction, reduces the delay of data loading, shortens the task processing time, and performs more balanced in processing different tasks, improves the overall adaptability and stability of the machine learning model, and meets the reliable operation requirements in complex scenes (such as real-time script analysis and multi-user concurrent query).
[0291] In the following, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0292] Continuous learning aims to enable the model to continuously accumulate knowledge from non-stationary data. To address the problem of catastrophic forgetting in continuous learning, the training method of the machine learning model in the related art includes: a regularization-based method, which avoids catastrophic forgetting by adding a regularization loss term to constrain the parameter change of the model; a history data playback-based method, which implements playback by generating pseudo samples of historical tasks to help the model maintain memory of old tasks when learning new tasks; a model architecture-based method, which avoids interference between different tasks by designing network modules specifically for each task; by activating different sub-models to cope with different tasks, the computing efficiency and task processing capability are improved; meta-learning-based continuous learning, which introduces a meta-learning strategy based on multi-task learning, enabling the model to adapt to changes in different tasks, thereby improving knowledge sharing and transfer ability; by distilling the learning knowledge of new tasks into the model of old tasks, the knowledge transfer between tasks is promoted; by identifying the importance of parameters, the fusion of knowledge unique and knowledge shared areas is realized; the parameter space of different tasks is forced to be orthogonal, thereby effectively avoiding forgetting; the selection module is used to combine different modules according to task relevance, and the knowledge transfer is promoted; adaptive selection weighting parameters can be used to fuse the models before and after training; through the strategy of multiple rounds of fusion, the intermediate state in the training process is utilized, and the effect is further improved.
[0293] In summary, the related art provides effective means for alleviating catastrophic forgetting or promoting knowledge transfer through regularization constraints, historical data playback, and adjusting model architecture. However, these methods still have limitations, i.e., it is difficult to achieve an optimal balance between maintaining old knowledge memory and new knowledge learning performance. To address the problem of catastrophic forgetting in continuous learning, model fusion methods have gradually become the mainstream solution due to their strong knowledge transfer ability, and mainly focus on two technical paths: one is single-round fusion method, which usually fuses the models before and after training, and uses global or fine-grained weights for fusion; as an example, see Figure 6A , Figure 6A is a schematic diagram of the single-round model fusion provided by the embodiments of the present application, as shown in Figure 6A , the model is trained for multiple iterations, and only after the last iteration, the model obtained after the last iteration is fused with the initialized model. The second is a multi-round fusion method, which enhances performance by merging model parameters multiple times during training, which can gradually fuse knowledge in multiple training steps. As an example, see Figure 6B , Figure 6Bis a schematic diagram of model multi-round fusion provided by an embodiment of the present application. During the process of model multi-round fusion, after each 3-round iteration training, the model of the current iteration training is fused with the initialized model, and the fused model is used as the initialized model of the next iteration training. However, these methods still have the following deficiencies:
[0294] 1. Fixed fusion opportunity: The single-round fusion method only fuses the trained model, and cannot utilize the valuable intermediate state generated during the model training process. The multi-round fusion method cannot dynamically adjust the fusion opportunity according to the training process, which may cause the model to be fused too early or too frequently during the new task learning stage, thereby affecting the learning efficiency of new knowledge and causing forgetting of historical knowledge.
[0295] 2. Lack of dynamic adjustment mechanism: The multi-round fusion method can increase the merging times, but the interval or frequency of merging is usually preset and fixed, which cannot effectively respond to the dynamic needs of the model during the learning of new tasks, and limits the performance improvement space.
[0296] 3. Unable to balance the performance of new and old knowledge: Most existing methods cannot find an optimal balance point between the learning of new tasks and the retention of old task knowledge, which still causes the problem of catastrophic forgetting.
[0297] 4. Lack of adaptability of fusion strategy: Although the multi-round fusion method can enhance performance through multiple merging, how to dynamically set the fusion weight according to the real-time state of the model is still a core problem to be solved. In related technologies, the continuous learning method based on multi-round model fusion usually adopts a fixed fusion interval, but this method cannot effectively balance the performance of the learning of new tasks and the forgetting of old tasks.
[0298] The training method of the machine learning model provided in the embodiments of the present application proposes an innovative adaptive multi-round iteration model fusion scheme. By introducing learning signals and forgetting signals, the change of the training trajectory in the model training process is dynamically detected, the learning of new knowledge and the forgetting of old knowledge are quantified respectively, and the timing and frequency of model fusion are adaptively adjusted according to the indications of different signals. The timing and frequency of fusion are adaptively adjusted according to the learning process of different tasks, avoiding the limitations of fixed fusion intervals. By adaptively controlling model fusion, the influence of redundant fusion operations on the performance of new tasks is avoided, while the necessary fusion is enhanced to avoid the forgetting of historical knowledge, ensuring that the learning of new tasks and the retention of old task knowledge can achieve the optimal balance. The learning and memory of new and old knowledge can be effectively balanced while dynamically adjusting the frequency and timing of model fusion. Through iterative fusion, not only the stability of the model in the continuous learning process is improved, but also the generalization ability of the model is enhanced, which shows excellent performance on multiple benchmark datasets, further improving the performance of the continuous learning model.
[0299] In summary, the adaptive iterative model fusion method proposed in the embodiments of the present application can achieve the optimal balance of new and old knowledge, improve the stability and training efficiency of the model, and is suitable for large-scale continuous learning tasks. It aims to solve the catastrophic forgetting problem of large language models in the continuous learning process and has significant advantages in solving catastrophic forgetting and improving knowledge transfer performance.
[0300] Specifically, the change amount of the model parameters (the change amount of the parameters in the model training process using new sample data for training) is used to represent the learning signal, and the loss value of the historical task (the difference between the output result of the model for the historical sample data and the label in the training process of the new task) is used to represent the forgetting signal. Through the fusion controller, the interval and frequency of model fusion are dynamically adjusted according to the real-time change of the signals, ensuring that the fusion frequency is increased in the rapid learning stage of the new task to prevent catastrophic forgetting, and the fusion frequency is reduced in the stable training stage to improve the training efficiency and model stability. Finally, through the model fusion module based on knowledge replay, the knowledge of the new task and the historical task is effectively fused by using the replay technology, the learning performance of the new task is improved by using the weighted fusion strategy, and the stability of the historical task knowledge is maintained.
[0301] The embodiments of the present application provide an intuitive graphical user interface (Graphical User Interface, GUI) to help users understand the model training process. As an example, refer to Figure 4 , Figure 4 is a schematic diagram of task state visualization provided by the embodiments of the present application, as shown in Figure 4As shown, the progress of learning new knowledge and the retention of historical knowledge in the new task are displayed, as well as the performance change after the fusion of new and old knowledge (i.e., new knowledge and historical knowledge), wherein, Figure 4 The left side is a single round model parameter fusion, Figure 4 The right side is a double round model parameter fusion. Next, referring to Figure 5 , Figure 5 is a model fusion timing change diagram provided by an embodiment of the present application, as shown in Figure 5 The upper two graphs show the changes of the forgetting signal and the learning signal during the training process of the model in two different tasks (task 1 in the left graph and task 2 in the right graph), and the horizontal coordinates of the lower two graphs are the number of iterations, and the vertical coordinates are the fusion interval, which shows the real-time fusion interval during the training process of the two different tasks, reflecting the timing and frequency of the model parameter fusion operation.
[0302] The training method of the machine learning model provided by the embodiments of the present application is not only suitable for traditional natural language processing tasks (such as text classification, dialogue question and answer, etc.), but also can be flexibly adapted to continuous learning tasks in different business scenarios, such as medical text processing, script content analysis, etc.
[0303] The training method of the machine learning model provided by the embodiments of the present application supports automatic policy switching, and automatically adjusts the model fusion strategy during the training process according to different data distribution or task requirements. For example, when the model faces a new task, the system can automatically switch to a strategy based on adaptive iterative fusion; and when a specific task has higher requirements for speed or memory, the system will automatically adjust to a lightweight fusion strategy, such as only fusing part of the model parameters, and all strategy evolution processes will be displayed through a graphical interface, and the user can view and adjust at any time.
[0304] In addition, through adaptive iterative model fusion, the fusion timing is dynamically adjusted in combination with the learning signal and the forgetting signal, and a knowledge fusion module based on historical data playback is integrated, so that the knowledge of new and old tasks can be efficiently fused. In addition, a dynamic weight adjustment algorithm is integrated in the system, and appropriate task vector fusion weights are automatically allocated according to the importance of the task.
[0305] In summary, the training method of the machine learning model provided by the embodiments of the present application realizes a closed-loop process from training data analysis, real-time monitoring of learning signals and forgetting signals, adaptive fusion strategy opening, a knowledge fusion module based on historical data playback, real-time evaluation and visualization of the training process. And it has high visualization, controllability and scalability, can significantly improve the performance of large language models in the continuous learning process, especially in achieving a better balance between retaining historical knowledge and improving the learning effect of new tasks.
[0306] The embodiment of the present application proposes a training framework based on adaptive iterative model fusion, aiming to solve the problem of catastrophic forgetting of large language models in continuous learning tasks, especially suitable for multi-task learning scenarios and knowledge injection tasks. The framework dynamically detects the trajectory changes in the model training process, and uses the proposed learning signal (corresponding to the parameter change amount in the above) and forgetting signal (corresponding to the loss value greater than the preset loss threshold in the above) to adaptively adjust the timing and frequency of model fusion, see Figure 6C , Figure 6C is the schematic diagram of adaptive iterative fusion provided by the embodiment of the present application, as shown in Figure 6C , the nail 1 and the nail 2 represent that when the learning signal and the forgetting signal are activated at the same time, the first model parameter fusion is triggered; the nail 3 represents that when the learning signal is activated and the forgetting signal is not activated, the model parameter fusion is not triggered.
[0307] The training framework based on adaptive iterative model fusion has the characteristics of lightweight model structure, significant performance improvement, strong integrability, etc. The core technical architecture includes the following two technical components: a fusion controller guided by training trajectory and a knowledge fusion module based on data playback, which will be described in detail below.
[0308] In some embodiments, the fusion controller guided by training trajectory is used to adaptively adjust the timing and frequency of model fusion according to the learning state of the model on new tasks and the forgetting situation on historical tasks. By combining the learning signal and the forgetting signal, a dynamic task fusion strategy is realized.
[0309] First, the core function of the learning signal is to dynamically adjust the fusion interval according to the current training state of the model to adapt to the learning of new knowledge.
[0310] In some embodiments, the change amount of model parameters at two different training times can be used to represent the learning signal. Assuming that the bth fusion occurs in the jth iteration process of the model, and the interval from the (b-1)th fusion is S b (corresponding to the number of iterations to be iterated in the above), the task vector τ b (corresponding to the parameter change amount in the above) of the model between the two consecutive fusions (i.e. from the (b-1)th to the bth) can be represented as formula (1).
[0311]
[0312] Where θ j is the model state after the jth iteration, is the model state before the (b-1)th fusion.
[0313] In some embodiments, the task vector τ bSumming up the absolute values of all elements in the task vector τ b and calculating the ratio of the sum to the fusion interval S b , an index Λ measuring the state of the model in the new task learning state is obtained b .
[0314]
[0315] wherein, is the i-th element in the task vector τ b . Here, the ratio of the sum to the fusion interval S b is calculated to normalize the sum of the absolute values of all elements in the task vector τ b , so that fusion intervals of different lengths can be compared fairly.
[0316] In some embodiments, the state of the model in learning new knowledge can be evaluated based on the index Λ. By comparing the current index Λ b-1 with the index Λ b of the last fusion, the trend of parameter changes can be observed, and accordingly the fusion interval can be adjusted from S b+1 .
[0317] Here, if only the trend between two consecutive values is considered, it may lead to the learning signal being too sensitive to short-term fluctuations. To solve this problem, the technique of sliding window can be used to analyze the trend of parameter changes between multiple historical points, so as to capture a more reliable overall learning trajectory. Specifically, a list of historical values H = [Λ1, Λ2, …, Λ b-1 ] can be maintained, and according to a given sliding window length L w , the trend changes between consecutive terms are compared, i.e. Λ b and Λ b-1 , Λ b-1 and Λ b-2 , and so on, until and , so that according to the comparison results, the adjustment strategy of the fusion interval in different cases is determined:
[0318] Case 1 (case1): If the upward trend dominates, it indicates a rapid learning phase of new knowledge, and the fusion interval can be reduced based on the magnitude of parameter changes, so as to avoid the forgetting of historical knowledge caused by the influx of a large amount of new knowledge. As an example, see Figure 7 , Figure 7 is a schematic diagram of the change of the learning signal provided by the embodiments of the present application. As shown in Figure 7 , the horizontal and vertical axes are the iteration times, and the vertical axis is the parameter change amount, Figure 7The left graph shows the parameter variation amount of the model in the fast learning stage of task 1 with the iteration number.
[0319] Case 2 (case 2): If the upward or downward trend is not obvious, the current interval is kept unchanged.
[0320] Case 3 (case 3): If the downward trend is dominant, indicating that the model is in the slow convergence stage of new knowledge, the fusion interval is increased, thereby avoiding redundant fusion operations and affecting the learning of new knowledge. As an example, see Figure 7 , Figure 7 The right graph shows the parameter variation amount of the model in the slow convergence stage of task 2 with the iteration number.
[0321] Therefore, the adjustment strategy of the fusion interval between the b th fusion and the b+1 th fusion is defined as formula (3).
[0322]
[0323] where S min and S max represent the preset minimum iteration number and the preset maximum iteration number, for example, S min is 2 iterations, and S max is 128 iterations; γ learn is a step adjustment factor, which can be a preset value (such as 2), used to control the adjustment speed of the fusion interval. In the early stage of training, the model will have a cold start stage, during which no adjustment is made. The length of the cold start stage is equal to the length of the sliding window, which facilitates the collection of sufficient data for analysis, and the initial fusion interval is set to S init .
[0324] In some embodiments, if only the learning signal is relied on, that is, the model adjusts the fusion interval according to the learning state of new knowledge, this will make the model ignore the forgetting of historical knowledge, resulting in a suboptimal fusion strategy. To solve this problem, the forgetting signal is further integrated to help the controller adjust the fusion strategy while considering both new and old knowledge. As an example, see Figure 8 , Figure 8 is a variation diagram of the forgetting signal provided by the embodiments of the present application, as shown in Figure 8 the left graph shows the variation of the historical loss with the iteration number in the training process of the model performing the prediction task 1; and the right graph shows the variation of the historical loss with the iteration number in the training process of the model performing the prediction task 2. The embodiments of the present application can trigger the model parameter fusion in advance or delay the model parameter fusion according to the number of times of the forgetting signal, thereby optimizing the fusion interval of the model parameter fusion.
[0325] Here, the forgetting signal is used to represent the loss change of the historical data (i.e., the difference between the predicted data obtained by using the model to predict the historical data and the labels of the historical data) in the new task training process. Specifically, in each model iteration process, a batch of historical data is sampled from the memory buffer and combined with the data of the current task to be fed to the model. It should be noted that the historical data is only used for loss calculation and does not participate in gradient update. When the loss of the historical task exceeds a predetermined threshold, the forgetting signal is activated.
[0326] In some embodiments, the average loss of the historical data of the first 2 / 3xS b+1 may be determined first, for example, the average loss of the historical task of the first 20 iterations is calculated when the fusion interval of the fusion from the bth fusion to the b+1th fusion is 30 iterations; and then the product of the average loss and an adjustment factor γ forget is taken as the threshold δ b+1 . Here, the adjustment factor γ forget may be a preset value, for example, γ forget may be 2. If the historical loss exceeds the threshold in the subsequent training process, the forgetting signal is triggered, and the number of activations is accumulated, and the number of activations of the forgetting signal can be represented as F(b+1) = F(b+1) + 1.
[0327] In some embodiments, if the forgetting signal has been activated multiple times (for example, F(b+1) ≥ Fmax) before the planned fusion interval S b+1 , the model fusion is triggered in advance to prevent further forgetting, that is, the actual fusion interval is S' b+1 < S b+1 . Conversely, if the forgetting signal has not been activated when the predetermined fusion interval S b+1 is reached, the fusion can be delayed so that the model continues to focus on learning new knowledge. At this time, the fusion is triggered in the following two cases: one is that the forgetting signal is activated (S' b+1 > S b+1 ); the other is that the number of iterations reaches the upper limit (i.e., S' b+1 = 2xS b+1 ); for example, the planned fusion interval S b+1 is 10 iterations, and the forgetting signal has not been activated at the 20th time, then the fusion is directly performed. That is, the fusion controller provided in the embodiments of the present application can dynamically balance the learning and forgetting signals to optimize the acquisition of new knowledge while minimizing forgetting, thereby realizing an adaptive fusion strategy.
[0328] After determining the fusion strategy according to the above method, the actual task knowledge fusion can be performed through the knowledge fusion module. Assuming that the bth fusion occurs at the jth training iteration, the parameter change amount τ newbIt can be defined as formula (4).
[0329]
[0330] Where, θ j This refers to the state of the model after the b-th fusion in the j-th training iteration. This is the state of the model after the (b-1)th fusion.
[0331] Next, historical data will be used to analyze θ. j Perform S′ b After two fine-tuning steps, the updated model state θ is obtained. j(M) At this point, the task vector of historical knowledge... It can be defined as formula (5).
[0332]
[0333] Finally, formula (6) can be used to calculate the parameter changes of the new knowledge. The task vector τ of historical knowledge pastb Weighted fusion is performed to update the model parameters and achieve model fusion.
[0334]
[0335] Where α1 is the parameter change amount of the new knowledge. The fusion weights, where α2 is the task vector of historical knowledge. The fusion weights can be obtained in the following way:
[0336] First, for τ new It can assess the proportion of the learning signal that shows an upward trend within the sliding window, P. new =L up / L w This represents the model's active learning of new knowledge. Then, for τ... past It can calculate the number of activations of the forgetting signal. The ratio of P to the maximum threshold Fmax past =F(b) / Fmax, representing the degree of forgetting of historical knowledge. Finally, based on P using formula (7). new and P past Normalization is performed to obtain the fusion weights α1 and α2.
[0337]
[0338] As an example, the sliding window length of the learning signal and the maximum threshold F of the forgotten signal can be used. max Set to 3. After fusion is complete, the model training will start from the updated model state. Continue.
[0339] The embodiments of the present application provide a schematic diagram of the effect of alleviating catastrophic forgetting compared with the related art. As an example, see Figure 9 , Figure 9 is a comparative schematic diagram of historical loss of different methods provided by the embodiments of the present application. As shown in Figure 9 , the historical loss points of the training method of the machine learning model provided by the embodiments of the present application and the training method of the related art are shown, and the fitting line obtained by the historical loss points, it can be seen that the training method of the machine learning model provided by the embodiments of the present application causes smaller historical loss than the training method of the related art.
[0340] The adaptive iterative model fusion method proposed in the embodiments of the present application does not need to introduce additional complex mechanisms, has an end-to-end closed-loop characteristic, and can flexibly adjust the fusion time and frequency in continuous learning tasks. This method significantly improves the stability and knowledge retention ability of large language models in multi-task learning by introducing a dynamic adjustment mechanism for learning signals and forgetting signals, which can effectively avoid catastrophic forgetting and promote the learning of new knowledge. And the technical scheme of the present application has high applicability, which is suitable for any scene involving multi-task, such as solving continuous different vertical domain knowledge injection, has a significant performance improvement effect and wide application promotion value.
[0341] The following continues to illustrate an exemplary structure of the implementation of the machine learning model training device 233 provided by the embodiments of the present application as a software module. In some embodiments, as shown in Figure 2 , the software module stored in the machine learning model training device 233 of the storage 230 can include:
[0342] The first training module 2331 is configured to perform first-stage iterative training on the first machine learning model based on the first sample data of the first prediction task and the second sample data of the second prediction task, to obtain a second machine learning model, wherein the first machine learning model is trained based on the first sample data.
[0343] The determination module 2332 is configured to determine a first loss value and a parameter variation of the second machine learning model, and determine a to-be-iterated number of times of the second machine learning model according to the parameter variation.
[0344] The second training module 2333 is configured to perform second-stage iterative training on the second machine learning model based on the first loss value and the to-be-iterated number of times through the first sample data and the second sample data, to obtain a third machine learning model.
[0345] The fusion module 2334 is configured to fuse model parameters of the third machine learning model and the second machine learning model to obtain a fourth machine learning model, where the fourth machine learning model is configured to perform at least one of the first prediction task and the second prediction task.
[0346] In some embodiments, the first training module 2331 is further configured to call the first machine learning model to perform the first prediction task by using the first sample data to obtain first prediction data, determine a second loss value according to a difference between the first prediction data and a first label of the first sample data, call the first machine learning model to perform the second prediction task by using the second sample data to obtain second prediction data, determine a third loss value according to a difference between the second prediction data and a second label of the second sample data, perform weighted summation on the second loss value and the third loss value to obtain a fourth loss value, and update parameters of the first machine learning model based on the fourth loss value to obtain the second machine learning model.
[0347] In some embodiments, the determination module 2332 is further configured to call the second machine learning model to perform the first prediction task by using the first sample data to obtain third prediction data, determine a first loss value of the second machine learning model according to a difference between the third prediction data and the first label of the first sample data, and determine a variation amount of parameters of the second machine learning model relative to the parameters of the first machine learning model as the parameter variation amount of the second machine learning model.
[0348] In some embodiments, the determination module 2332 is further configured to, in the process of the iterative training in the first stage, if the operation of the multiple times of model parameter fusion is performed, determine a parameter variation amount between parameter fusion results of adjacent two times of model parameter fusion, determine a variation index parameter based on the parameter variation amount, determine a parameter variation trend according to the variation index parameter between every adjacent two times of model parameter fusion, and update an interval iteration number of last two times of model parameter fusion in the multiple times of model parameter fusion according to a preset adjustment strategy of the parameter variation trend to obtain the to-be-iterated number of times of the second machine learning model, where the interval iteration number of the last two times of model parameter fusion is a difference between a number of iteration training that has been performed before performing the last time of model parameter fusion and a number of iteration training that has been performed before performing the penultimate time of model parameter fusion.
[0349] In some embodiments, the determination module 2332 is further configured to determine a sum of absolute values of each sub-parameter variation amount in the parameter variation amount, and determine a ratio of the sum to the to-be-iterated number of times as the variation index parameter.
[0350] In some embodiments, the determining module 2332 is further configured to fit the change indicator parameter obtained in each model parameter fusion into a first change curve diagram in the order of the model parameter fusion, where the ordinate of the first change curve diagram represents the change indicator parameter between adjacent two model parameter fusions, and the abscissa of the first change curve diagram represents the iteration number; perform multiple sliding samplings on the first change curve diagram through a sliding window, and count the number of the rising sections and the number of the falling sections in the sampled multiple curves; in response to the number of the rising sections being greater than the number of the falling sections, determine that the parameter change trend is an upward trend; in response to the number of the rising sections being less than the number of the falling sections, determine that the parameter change trend is a downward trend; and in response to the number of the rising sections being equal to the number of the falling sections, determine that the parameter change trend is no change.
[0351] In some embodiments, the determining module 2332 is further configured to, in response to the parameter change trend being the upward trend, determine the ratio of the interval iteration number to the preset step length adjustment factor, and take the maximum value of the ratio and the preset minimum iteration number as the iteration number to be iterated; in response to the parameter change trend being the downward trend, determine the product of the interval iteration number and the preset step length adjustment factor, and take the minimum value of the product and the preset maximum iteration number as the iteration number to be iterated; and in response to the parameter change trend being no change, take the interval iteration number as the iteration number to be iterated.
[0352] In some embodiments, the second training module 2333 is further configured to perform second-stage iterative training on the second machine learning model through the first sample data, and perform the following processing after each iteration training in the second stage: determine a first over-limit count value, wherein after any iteration training in the second stage, if the first loss value of the first machine learning model after the iteration training is greater than the preset loss threshold, the first over-limit count value is incremented by 1; if the first over-limit count value is greater than 0 and less than a preset count threshold, the iteration number to be iterated is taken as a target iteration number; and perform iteration training on the second machine learning model through the second sample data for the target iteration number to obtain a third machine learning model.
[0353] In some embodiments, the second training module 2333 is further configured to perform second-stage iterative training on the second machine learning model through the first sample data, and perform the following processing after each iteration training in the second stage: determine a second over-limit count value, wherein after any iteration training in the second stage, if the first loss value of the first machine learning model after the iteration training is greater than the preset loss threshold, the second over-limit count value is incremented by 1; if the second over-limit count value is greater than or equal to the preset count threshold, the iteration number corresponding to the preset count threshold is taken as a target iteration number; and perform iteration training on the second machine learning model through the second sample data for the target iteration number to obtain a third machine learning model.
[0354] In some embodiments, the second training module 2333 is further configured to initialize a maximum iteration number of the second stage, wherein the maximum iteration number is greater than the iteration number to be performed; perform iteration training on the second machine learning model by using the first sample data, wherein after each iteration training of the second stage, the following processing is performed: if the iteration number performed is less than the maximum iteration number, and the first loss value of the second machine learning model after the current iteration training is less than a preset loss threshold, the next iteration training is continued; if the first loss value of the second machine learning model after the current iteration training is greater than or equal to the preset loss threshold, and the iteration number performed is less than the maximum iteration number, the iteration number performed is taken as a target iteration number; and perform iteration training on the second machine learning model by using the second sample data for the target iteration number, to obtain a third machine learning model.
[0355] In some embodiments, the second training module 2333 is further configured to initialize a maximum iteration number of the second stage, wherein the maximum iteration number is greater than the iteration number to be performed; perform iteration training on the second machine learning model by using the first sample data, wherein after each iteration training of the second stage, the following processing is performed: if the iteration number performed is less than the maximum iteration number, and the first loss value of the second machine learning model after the current iteration training is less than a preset loss threshold, the next iteration training is continued; if the first loss value of the second machine learning model after the current iteration training is less than the preset loss threshold, and the iteration number performed is the maximum iteration number, the maximum iteration number is taken as a target iteration number; and perform iteration training on the second machine learning model by using the second sample data for the target iteration number, to obtain a third machine learning model.
[0356] In some embodiments, the second training module 2333 is further configured to perform training on the second machine learning model by using the second sample data for the target iteration number, and in each iteration training, the following processing is performed: call the initialized machine learning model of the n th iteration training to predict the second sample data, to obtain the n th prediction data; determine the n th loss value according to the difference between the n th prediction data and the second sample label of the second sample data; and update the parameters of the initialized machine learning model of the n th iteration training according to the n th loss value, to obtain an updated machine learning model, wherein when n is 1, the initialized machine learning model is the second machine learning model, n≤N, n and N are both positive integers, and N is the target iteration number; when n=N, the updated machine learning model is the third machine learning model.
[0357] In some embodiments, the fusion module 2334 is further configured to determine a first parameter variation of parameters of the third machine learning model relative to parameters of the second machine learning model, fine-tune the third machine learning model using the first sample data for a preset number of iterations to obtain a fine-tuned third machine learning model, determine a second parameter variation of parameters of the fine-tuned third machine learning model relative to the third machine learning model, determine a first weight of the first parameter variation and a second weight of the second parameter variation, determine a sum of a product of the parameters of the second machine learning model, the first parameter variation and the first weight, and a product of the second parameter variation and the second weight, and construct a fourth machine learning model based on the sum.
[0358] In some embodiments, the fusion module 2334 is further configured to determine a first performance indicator according to a variation trend of the parameters of the third machine learning model in the fine-tuning process for the preset number of iterations, wherein the first performance indicator is used to represent a performance of the third machine learning model in performing the second prediction task, determine a second performance indicator according to the fifth loss value of the third machine learning model in the fine-tuning process for the preset number of iterations, wherein the second performance indicator is used to represent a performance of the third machine learning model in performing the first prediction task, and determine the first weight of the first parameter variation and the second weight of the second parameter variation according to the first performance indicator and the second performance indicator.
[0359] In some embodiments, the fusion module 2334 is further configured to fit a second variation curve according to the variation trend of the parameters of the third machine learning model in the fine-tuning process for the preset number of iterations, wherein the ordinate of the second variation curve represents the parameters after each iteration, and the abscissa of the second variation curve represents the number of iterations, perform multiple sliding samplings on the second variation curve through a sliding window, and count the number of the rising sections and the number of the falling sections in the multiple curves sampled, and take the ratio of the number of the rising sections to the total number as the first performance indicator, wherein the total number is the sum of the number of the rising sections and the number of the falling sections.
[0360] In some embodiments, the fusion module 2334 is further configured to determine a third transfinite count value, wherein after any iteration training in the fine-tuning process for the preset number of iterations, if the fifth loss value of the third machine learning model after the iteration training is greater than a preset loss threshold, the third transfinite count value is incremented by 1, and the ratio of the third transfinite count value to a preset count threshold is taken as the second performance indicator.
[0361] In some embodiments, the fusion module 2334 is further configured to determine a sum of the first performance indicator and the second performance indicator, take the ratio of the first performance indicator to the sum as the first weight of the first parameter variation, and take the ratio of the second performance indicator to the sum as the second weight of the second parameter variation.
[0362] The embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the machine learning model training method provided by the embodiment of the present application.
[0363] The embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions or computer programs. When the computer executable instructions or computer programs are executed by a processor, the processor will execute the machine learning model training method provided by the embodiment of the present application, for example, as shown in the machine learning model training method. Figure 3A The embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions or computer programs. When the computer executable instructions or computer programs are executed by a processor, the processor will execute the machine learning model training method provided by the embodiment of the present application, for example, as shown in the machine learning model training method.
[0364] In some embodiments, the computer readable storage medium can be RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM memory, etc.; or various devices including one or any combination of the above storage medium.
[0365] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.
[0366] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in a part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code parts).
[0367] As an example, the computer executable instructions can be deployed to execute on one electronic device, or on multiple electronic devices located in one place, or on multiple electronic devices distributed in multiple places and interconnected through a communication network.
[0368] To sum up, after the first machine learning model is trained by the first sample data of the first prediction task, the first machine learning model is iteratively trained by the first sample data of the first prediction task and the second sample data of the second prediction task to obtain the second machine learning model in the first stage, and the number of iterations to be performed is determined according to the parameter variation of the second machine learning model. In the first stage, the model is iteratively trained by the data of two different tasks, and the parameter variation can be used as a basis for adjusting the training intensity to avoid overtraining or insufficient training.
[0369] After the number of iterations to be performed is determined, the second machine learning model is iteratively trained by the first sample data and the second sample data based on the first loss value and the number of iterations to be performed in the second stage, so that the third machine learning model obtained can take into account the performance of the two prediction tasks, and the stability of the machine learning model in the multi-task training process is improved. The model parameters of the third machine learning model and the second machine learning model are fused, so that the fourth machine learning model obtained finally can make full use of the training results of the two stages and better adapt to multiple prediction tasks, perform more balanced in processing different tasks, and improve the overall adaptability of the machine learning model.
[0370] The above is only an embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A method for training a machine learning model, the method comprising: The method comprises: performing first-stage iterative training on a first machine learning model by first sample data of a first prediction task and second sample data of a second prediction task to obtain a second machine learning model, wherein the first machine learning model is trained based on the first sample data; determining a first loss value and a parameter change amount of the second machine learning model; determining a to-be-iterated number of times of the second machine learning model according to the parameter change amount; performing second-stage iterative training on the second machine learning model based on the first loss value and the to-be-iterated number of times of times by the first sample data and the second sample data to obtain a third machine learning model; performing model parameter fusion on the third machine learning model and the second machine learning model to obtain a fourth machine learning model, wherein the fourth machine learning model is used to execute at least one of the first prediction task and the second prediction task.
2. The method of claim 1, wherein, The method comprises: calling the first machine learning model to execute the first prediction task by the first sample data to obtain first prediction data; determining a second loss value according to a difference between the first prediction data and a first label of the first sample data; calling the first machine learning model to execute the second prediction task by the second sample data to obtain second prediction data; determining a third loss value according to a difference between the second prediction data and a second label of the second sample data; performing weighted summation on the second loss value and the third loss value to obtain a fourth loss value; updating parameters of the first machine learning model based on the fourth loss value to obtain a second machine learning model.
3. The method of claim 1, wherein, The method comprises: calling the second machine learning model to execute the first prediction task by the first sample data to obtain third prediction data; determining a first loss value of the second machine learning model according to a difference between the third prediction data and a first label of the first sample data; determining a change amount of parameters of the second machine learning model relative to parameters of the first machine learning model as a parameter change amount of the second machine learning model.
4. The method of claim 1, wherein, The method comprises: in the process of the first-stage iterative training, if a plurality of model parameter fusion operations are performed, determining a parameter change amount between parameter fusion results of adjacent two model parameter fusions; determining a change index parameter based on the parameter change amount; determining a parameter change trend according to the change index parameter between each adjacent two model parameter fusions; updating, according to a preset adjustment strategy of the parameter change trend, interval iteration times of last two of the model parameter fusions in the multiple model parameter fusions, to obtain the iteration times to be performed on the second machine learning model, wherein the interval iteration times of the last two of the model parameter fusions are a difference between a number of iteration trainings performed before performing the last model parameter fusion and a number of iteration trainings performed before performing the second last model parameter fusion.
5. The method of claim 4, wherein, The determining the change index parameter based on the parameter change amount comprises: determining a sum of absolute values of each of the parameter change amounts; determining a ratio of the sum to the iteration times to be performed as the change index parameter.
6. The method of claim 5, wherein, The determining the parameter change trend according to the change index parameter between each of adjacent two of the model parameter fusions comprises: fitting, in a sequence of the model parameter fusions, the change index parameter obtained by each of the model parameter fusions into a first change curve, wherein a vertical coordinate in the first change curve represents the change index parameter between the adjacent two of the model parameter fusions, and a horizontal coordinate in the first change curve represents iteration times; statistically counting a number of rising sections and a number of falling sections in a plurality of curves obtained by multiple sliding samplings of the first change curve through a sliding window; in response to the number of the rising sections being greater than the number of the falling sections, determining that the parameter change trend is an upward trend; in response to the number of the rising sections being less than the number of the falling sections, determining that the parameter change trend is a downward trend; in response to the number of the rising sections being equal to the number of the falling sections, determining that the parameter change trend is no change.
7. The method of claim 4, wherein, The updating, according to a preset adjustment strategy of the parameter change trend, interval iteration times of last two of the model parameter fusions in the multiple model parameter fusions, to obtain the iteration times to be performed on the second machine learning model, wherein the interval iteration times of the last two of the model parameter fusions are a difference between a number of iteration trainings performed before performing the last model parameter fusion and a number of iteration trainings performed before performing the second last model parameter fusion. The determining the change index parameter based on the parameter change amount comprises: determining a sum of absolute values of each of the parameter change amounts; determining a ratio of the sum to the iteration times to be performed as the change index parameter.
8. The method according to any one of claims 1 to 7, characterized in that, The determining the parameter change trend according to the change index parameter between each of adjacent two of the model parameter fusions comprises: fitting, in a sequence of the model parameter fusions, the change index parameter obtained by each of the model parameter fusions into a first change curve, wherein a vertical coordinate in the first change curve represents the change index parameter between the adjacent two of the model parameter fusions, and a horizontal coordinate in the first change curve represents iteration times; statistically counting a number of rising sections and a number of falling sections in a plurality of curves obtained by multiple sliding samplings of the first change curve through a sliding window; in response to the number of the rising sections being greater than the number of the falling sections, determining that the parameter change trend is an upward trend; in response to the number of the rising sections being less than the number of the falling sections, determining that the parameter change trend is a downward trend; in response to the number of the rising sections being equal to the number of the falling sections, determining that the parameter change trend is no change. The updating, according to a preset adjustment strategy of the parameter change trend, interval iteration times of last two of the model parameter fusions in the multiple model parameter fusions, to obtain the iteration times to be performed on the second machine learning model, wherein the interval iteration times of the last two of the model parameter fusions are a difference between a number of iteration trainings performed before performing the last model parameter fusion and a number of iteration trainings performed before performing the second last model parameter fusion. The determining the change index parameter based on the parameter change amount comprises: determining a sum of absolute values of each of the parameter change amounts; determining a ratio of the sum to the iteration times to be performed as the change index parameter. The determining the parameter change trend according to the change index parameter between each of adjacent two of the model parameter fusions comprises: fitting, in a sequence of the model parameter fusions, the change index parameter obtained by each of the model parameter fusions into a first change curve, wherein a vertical coordinate in the first change curve represents the change index parameter between the adjacent two of the model parameter fusions, and a horizontal coordinate in the first change curve represents iteration times; statistically counting a number of rising sections and a number of falling sections in a plurality of curves obtained by multiple sliding samplings of the first change curve through a sliding window; in response to the number of the rising sections being greater than the number of the falling sections, determining that the parameter change trend is an upward trend; in response to the number of the rising sections being less than the number of the falling sections, determining that the parameter change trend is a downward trend; in response to the number of the rising sections being equal to the number of the falling sections, determining that the parameter change trend is no change. The updating, according to a preset adjustment strategy of the parameter change trend, interval iteration times of last two of the model parameter fusions in the multiple model parameter fusions, to obtain the iteration times to be performed on the second machine learning model, wherein the interval iteration times of the last two of the model parameter fusions are a difference between a number of iteration trainings performed before performing the last model parameter fusion and a number of iteration trainings performed before performing the second last model parameter fusion. The determining the change index parameter based on the parameter change amount comprises: determining a sum of absolute values of each of the parameter change amounts; determining a ratio of the sum to the iteration times to be performed as the change index parameter. determining a first over-limit count value, wherein after each iteration training in the second stage, if the first loss value of the second machine learning model after the iteration training is greater than a preset loss threshold, the first over-limit count value is added by 1; if the first over-limit count value is greater than 0 and less than a preset count threshold, the number of iterations to be iterated is taken as a target iteration number; performing iteration training on the second machine learning model through the second sample data for the target iteration number to obtain a third machine learning model.
9. The method according to any one of claims 1 to 7, characterized in that, The second stage of iteration training on the second machine learning model based on the first loss value and the number of iterations to be iterated to obtain a third machine learning model comprises: performing iteration training on the second machine learning model through the first sample data, wherein after each iteration training in the second stage, the following processing is performed: determining a second over-limit count value, wherein after each iteration training in the second stage, if the first loss value of the second machine learning model after the iteration training is greater than a preset loss threshold, the second over-limit count value is added by 1; if the second over-limit count value is greater than or equal to a preset count threshold, the iteration number corresponding to the preset count threshold is taken as a target iteration number; performing iteration training on the second machine learning model through the second sample data for the target iteration number to obtain a third machine learning model.
10. The method according to any one of claims 1 to 7, characterized in that, The second stage of iteration training on the second machine learning model based on the first loss value and the number of iterations to be iterated to obtain a third machine learning model comprises: initializing a maximum iteration number of the second stage, wherein the maximum iteration number is greater than the number of iterations to be iterated; performing iteration training on the second machine learning model through the first sample data, wherein after each iteration training in the second stage, the following processing is performed: if the number of iterations is less than the maximum iteration number, and the first loss value of the second machine learning model after the current iteration training is less than a preset loss threshold, the next iteration training is continued; if the first loss value of the second machine learning model after the current iteration training is greater than or equal to the preset loss threshold, and the current iteration number is less than the maximum iteration number, the current iteration number is taken as a target iteration number; performing iteration training on the second machine learning model through the second sample data for the target iteration number to obtain a third machine learning model.
11. The method according to any one of claims 1 to 7, characterized in that, The second stage of iteration training on the second machine learning model based on the first loss value and the number of iterations to be iterated to obtain a third machine learning model comprises: initializing a maximum iteration number of the second stage, wherein the maximum iteration number is greater than the number of iterations to be iterated; performing iteration training on the second machine learning model through the first sample data, wherein after each iteration training in the second stage, the following processing is performed: if the current iteration number is less than the maximum iteration number, and the first loss value of the second machine learning model after the current iteration training is less than a preset loss threshold, then continue the next iteration training; if the first loss value of the second machine learning model after the current iteration training is less than the preset loss threshold, and the current iteration number is the maximum iteration number, then take the maximum iteration number as a target iteration number; perform iteration training on the second machine learning model by using the second sample data for the target iteration number, to obtain a third machine learning model.
12. The method of claim 8, wherein, the iteration training on the second machine learning model for the target iteration number to obtain a third machine learning model, comprises: in each iteration training, the following processing is performed: call the initialized machine learning model of the n th iteration training to predict the second sample data, to obtain n th prediction data; determine an n th loss value according to the difference between the n th prediction data and the second label of the second sample data; update the parameters of the initialized machine learning model of the n th iteration training according to the n th loss value, to obtain an updated machine learning model, wherein when n is 1, the initial machine learning model is the second machine learning model, n≤N, n and N are both positive integers, N is the target iteration number, and when n=N, the updated machine learning model is the third machine learning model.
13. The method according to any one of claims 1 to 7, characterized in that, the model parameter fusion of the third machine learning model and the second machine learning model to obtain a fourth machine learning model, comprises: determine a first parameter change amount of the parameters of the third machine learning model relative to the parameters of the second machine learning model; perform fine-tuning on the third machine learning model by using the first sample data for a preset iteration number, to obtain a fine-tuned third machine learning model; determine a second parameter change amount of the parameters of the fine-tuned third machine learning model relative to the third machine learning model; determine a first weight of the first parameter change amount and a second weight of the second parameter change amount; determine the sum of the product of the parameters of the second machine learning model, the first parameter change amount and the first weight, and the product of the second parameter change amount and the second weight; construct a fourth machine learning model based on the sum.
14. The method of claim 13, wherein, the determination of the first weight of the first parameter change amount and the second weight of the second parameter change amount, comprises: determine a first performance index according to the change trend of the parameters of the third machine learning model in the fine-tuning process for the preset iteration number, wherein the first performance index is used to represent the performance of the third machine learning model in performing the second prediction task; determine a second performance index according to a fifth loss value of the third machine learning model in the fine-tuning process for the preset iteration number, wherein the second performance index is used to represent the performance of the third machine learning model in performing the first prediction task; According to the first performance index and the second performance index, a first weight of the first parameter change amount and a second weight of the second parameter change amount are determined.
15. The method of claim 14, wherein, The first performance index is determined according to the change trend of the parameters of the third machine learning model in the fine-tuning process of the preset number of iterations, including: According to the change trend of the parameters of the third machine learning model in the fine-tuning process of the preset number of iterations, a second change curve is fitted, wherein the ordinate in the second change curve represents the parameter after each iteration, and the abscissa in the second change curve represents the number of iterations. The number of the rising segments and the number of the falling segments in the plurality of curves sampled are counted by sliding the window on the second change curve; The ratio of the number of the rising segments to the total number is taken as the first performance index, wherein the total number is the sum of the number of the rising segments and the number of the falling segments.
16. The method of claim 14, wherein, The second performance index is determined according to the fifth loss value of the third machine learning model in the fine-tuning process of the preset number of iterations, including: A third over-limit count value is determined, wherein after any iteration training in the fine-tuning process of the preset number of iterations, if the fifth loss value of the third machine learning model after the iteration training is greater than a preset loss threshold, the third over-limit count value is incremented by 1; The ratio of the third over-limit count value to a preset count threshold is taken as the second performance index.
17. An apparatus for training a machine learning model, the apparatus comprising: The device includes: A first training module is configured to perform first-stage iteration training on a first machine learning model based on first sample data of a first prediction task and second sample data of a second prediction task, to obtain a second machine learning model, wherein the first machine learning model is trained based on the first sample data; A determination module is configured to determine a first loss value and a parameter change amount of the second machine learning model, and determine a to-be-iterated number of the second machine learning model based on the parameter change amount; A second training module is configured to perform second-stage iteration training on the second machine learning model based on the first loss value and the to-be-iterated number through the first sample data and the second sample data, to obtain a third machine learning model; A fusion module is configured to perform model parameter fusion on the third machine learning model and the second machine learning model, to obtain a fourth machine learning model, wherein the fourth machine learning model is used to perform at least one of the first prediction task and the second prediction task.
18. An electronic device, comprising: The electronic device includes: A memory is configured to store computer executable instructions or computer programs; A processor is configured to execute the computer executable instructions or computer programs stored in the memory, to implement the training method of the machine learning model in any one of claims 1 to 16.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program comprise the steps of claim 18. The computer executable instructions or computer programs are executed by the processor to implement the training method of the machine learning model in any one of claims 1 to 16.
20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, The computer executable instructions or computer programs, when executed by the processor, implement the training method of the machine learning model according to any one of claims 1 to 16.