Model training method and apparatus, terminal device, and storage medium
By combining user static data and user behavior time series data in the training method of a multi-task encoding network in the financial marketing recommendation model, the problem of poor training effect caused by user static data is solved and the prediction accuracy of the model is improved.
Patent Information
- Application Number
- CN202211655700.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-22
AI Technical Summary
The existing financial marketing recommendation models use static user data during training, resulting in poor training performance.
User static data and time-series data based on user behavior are input into a multi-task encoding network, which encodes and predicts data through an embedding layer, a self-attention layer, and a fully connected layer. The model is then optimized using a binary cross-entropy loss function.
This improved the training performance of the model and enhanced the prediction accuracy for tasks related to the temporal sequence of user behavior.
Smart Images

Figure CN115879005B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and in particular to a model training method and device, a terminal device, and a storage medium. BACKGROUND
[0002] Currently, in a financial marketing recommendation scenario, traditional feature engineering is usually used for feature processing, and then the processed features are used as inputs of a model. However, the training data of the current model generally uses static data of a user, such as a user identity (ID) and user statistical data, which leads to poor training effect of the model. SUMMARY
[0003] The present application provides a model training method and device, a terminal device, and a storage medium, which can improve the training effect of a model.
[0004] A first aspect of the present application provides a model training method, comprising:
[0005] inputting user static data in a first sample into the first encoding network to obtain a first encoding result of the first sample; the first sample is any one in a target sample set;
[0006] inputting time series data based on user behavior in the first sample into the second encoding network to obtain a second encoding result of the first sample;
[0007] combining the encoding result of the first sample and the encoding result of the second sample, and inputting the combination into the multi-task encoding network to obtain a multi-task prediction result of the first sample;
[0008] calculating a training loss of the second encoding network and a training loss of the multi-task encoding network;
[0009] optimizing the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0010] A second aspect of the present application provides a model training device, which is applied to a multi-task model, and the multi-task model comprises a first encoding network, a second encoding network, and a multi-task encoding network. The device comprises:
[0011] a first encoding unit, configured to input user static data in a first sample into the first encoding network to obtain a first encoding result of the first sample; the first sample is any one in a target sample set;
[0012] a second encoding unit, configured to input the time series data based on user behavior in the first sample into the second encoding network to obtain a second encoding result of the first sample;
[0013] a prediction unit, configured to input the encoding result of the first sample and the encoding result of the second sample into the multi-task encoding network after combination to obtain a multi-task prediction result of the first sample;
[0014] a loss calculation unit, configured to calculate a training loss of the second encoding network and a training loss of the multi-task encoding network;
[0015] an optimization unit, configured to optimize the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0016] A third aspect of the embodiment of the present application provides a terminal device, comprising a processor and a memory, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to invoke the program instructions to execute the step instructions in the first aspect of the embodiment of the present application.
[0017] A fourth aspect of the embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program for electronic data exchange, the computer program comprises program instructions, and the program instructions make the processor execute the step instructions in the first aspect of the embodiment of the present application when the processor executes the program instructions.
[0018] A fifth aspect of the embodiment of the present application provides a computer program product, wherein the computer program product comprises a computer program, the computer program comprises program instructions, and the program instructions make the processor execute the step instructions in the first aspect of the embodiment of the present application when the processor executes the program instructions.
[0019] The model training method provided in the embodiments of the present application inputs user static data in a first sample into the first encoding network to obtain a first encoding result of the first sample; the first sample is any one in a target sample set; inputs time series data based on user behavior in the first sample into the second encoding network to obtain a second encoding result of the first sample; inputs the encoding result of the first sample and the encoding result of the second sample after combination into the multi-task encoding network to obtain a multi-task prediction result of the first sample; calculates a training loss of the second encoding network and a training loss of the multi-task encoding network; and optimizes the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network. In the embodiments of the present application, when the multi-task model is trained, the sample contains not only user static data but also time series data based on user behavior. Considering that the multi-task prediction result is related to the time sequence of user behavior, the multi-task model is trained based on the time series data based on user behavior, thereby improving the training effect of the model. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0021] Figure 1 is a structural schematic diagram of a multi-task model provided by the embodiments of the present application;
[0022] Figure 2 is a flow schematic diagram of a model training method provided by the embodiments of the present application;
[0023] Figure 3 is a flow schematic diagram of a sample cleaning method provided by the embodiments of the present application;
[0024] Figure 4 is a structural schematic diagram of a model training device provided by the embodiments of the present application;
[0025] Figure 5 is a structural schematic diagram of a terminal device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0026] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.
[0027] The terms "first", "second", and the like in the specification of the present application and the above drawings are used to distinguish different objects, rather than to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product, or device.
[0028] In the present application, the phrase "embodiment" means that the specific features, structures, or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the present application can be combined with other embodiments.
[0029] The terminal device involved in the embodiments of the present application is a device with display capability. It can be a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an on-board unit (OBU), a wearable device (such as a watch, a bracelet, a smart helmet, etc.), a smart home device (a rice cooker, a sound system, a home butler device, etc.), an augmented reality (AR) / virtual reality (VR) device, etc.
[0030] The model training method of the embodiments of the present application, when training the multi-task model, not only contains user static data in the sample, but also contains time series data based on user behavior. Considering that the multi-task prediction result is related to the time sequence of user behavior, the multi-task model is trained based on the time series data of user behavior, thereby improving the training effect of the model. The following will be specifically described.
[0031] Please refer to Figure 1 ,Figure 1 is a structural schematic diagram of a multi-task model provided by an embodiment of the present application. As shown in the figure, the multi-task model comprises a first encoding network, a second encoding network and a multi-task encoding network. Figure 1
[0032] In the embodiment of the present application, the user static data in the first sample can be input into the first encoding network to obtain a first encoding result of the first sample; the time series data based on user behavior in the first sample can be input into the second encoding network to obtain a second encoding result of the first sample; the encoding result of the first sample and the encoding result of the second sample are combined and input into the multi-task encoding network to obtain a multi-task prediction result of the first sample; the training loss of the second encoding network is calculated, the training loss of the multi-task encoding network is calculated; and the multi-task model is optimized based on the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0033] The first encoding network can comprise an embedding layer, and the second encoding network can comprise an embedding layer and an attention layer. The embedding layer in the first encoding network and the embedding layer in the second encoding network do not share parameters. The multi-task encoding network can comprise a plurality of sub-task encoding networks, each of which can comprise an attention layer and a fully connected layer. The attention layer in the second encoding network and the attention layer in each sub-task encoding network do not share parameters. The attention layers in each sub-task encoding network do not share parameters.
[0034] The target sample set can comprise a plurality of samples, each of which can comprise user static data and time series data of user behavior. Each sample can correspond to user static data and time series data of user behavior of a user. For example, each sample can correspond to a user ID, and the user ID corresponding to each sample is different.
[0035] The core components of the multi-task model include an embedding layer, an attention layer and a fully connected layer.
[0036] The embedding layer, also known as an embedding module, is used for vectorization encoding of user static data by the embedding layer in the first encoding network, and vectorization encoding of time series data of user behavior (User App history behavior sequence) by the embedding layer in the second encoding network.
[0037] Self-attention layers, which can also be referred to as self-attention encoding layers, self-attention modules, self-attention encoding modules, etc. The input of the self-attention layer is defined as M, and the output is M'. The self-attention layer can be a multi-head self-attention layer.
[0038] In the second encoding network, the output M of the embedding layer can be taken as the input of the self-attention layer.
[0039] The output can be calculated by the following formula :
[0040] First, perform feature mapping on the input M in the manner of ; obtain , where represents a normalization layer, is a learnable parameter; where, W Q , W K , W v are d-dimensional by d-dimensional matrices.
[0041] Second, perform multi-head attention calculation on the Query, key and Value obtained in the first step, according to the number of heads of the self-attention layer Split in the hidden vector dimension , obtain , where the calculation method of the hth head is:
[0042]
[0043] Third, merge the calculation results of each head, where is a learnable parameter:
[0044]
[0045] Fourth, define a feed-forward layer (FFN) calculation function FFN(X), where is a learnable parameter :
[0046]
[0047] Obtain the output through the normalization layer and residual operation.
[0048]
[0049] Perform multiple feature fusions on M by stacking multiple self-attention layers, and take the output of the last layer as .
[0050] A fully connected layer (Sigmoid DNN Layer), the multi-task encoding network is composed of multiple fully connected layers, and the last layer outputs a prediction value through a function.
[0051] .
[0052] Each sub-task encoding network in the multi-task encoding network can output a prediction value.
[0053] Please refer to Figure 2 , Figure 2 is a flowchart of a model training method provided by an embodiment of the present application. The method can be applied to Figure 1 a multi-task model as shown in the figure. As shown in Figure 2 , the method can include the following steps.
[0054] 201, the terminal device inputs user static data in the first sample into the first encoding network to obtain a first encoding result of the first sample; the first sample is any one in a target sample set.
[0055] 202, the terminal device inputs time series data based on user behavior in the first sample into the second encoding network to obtain a second encoding result of the first sample.
[0056] In the embodiments of the present application, the multi-task model takes two types of data, user static data and time series data of user behavior, as input. The user static data is data that does not contain time series, and the user static data includes but is not limited to user ID features, user statistical features (such as counting the number of APP operations of the user within a period of time, the cumulative use time of the APP, etc.), user demographic attributes, and the like. The time series data of user behavior can also be referred to as user application behavior data (User App history behavior sequence). As long as a series of operations performed by the user in time sequence, it is considered to belong to the time series data of user behavior. For example, a series of click operations of the user within the same application (application, APP). Specifically, taking the promotion of financial advertisements as an example, the series of operations can include: a click operation on a financial advertisement; after performing the click action, performing the execution action of real-name information registration; after performing the real-name information registration, completing the credit operation through the system audit; after completing the credit operation through the system audit, performing the loan operation. As can be seen, the click operation, the real-name information registration, the credit operation, and the loan operation are a series of click operations in time sequence within the same APP, and the time sequence of the four operations cannot be reversed. For example, a series of click operations of the user between several APPs. Specifically, a series of click operations of APP1, APP2, and APP3 clicked by the user in sequence.
[0057] The samples in the target sample set can be cleaned samples to ensure the reliability of the samples. For example, the samples in the training sample set can be cleaned, and the cleaned samples can be used as the target sample set.
[0058] For the time series data of user behavior: , wherein represents the event occurring the th time in the sequence. Each can be embedded and coded (Embedding coding), and each can be unified into a d-dimensional vector space. d is a hyperparameter. Since the input data of the model contains various data with different semantics, data with completely different meanings can be encoded into the same vector space (d-dimensional vector space), so that subsequent model operations can be performed. Unified coding into a d-dimensional vector space can ensure that data with different meanings are encoded into a unified vector space.
[0059] The APP set of each user can be embedded and coded to obtain an embedding matrix , it is assumed that there are N APPs in the APP set, and the index of each APP corresponds to a vector with a dimension of d.
[0060] Optionally, step 201 can include the following steps:
[0061] The terminal device inputs the user static data in the first sample into the first encoding network, and the first encoding network is used to embed and encode the user static data to map to a d-dimensional vector space to obtain a first encoding result of the first sample, and the first encoding result is a d-dimensional encoding tensor, and d is an integer greater than or equal to 2.
[0062] In the embodiment of the application, the user static data can include discrete data and continuous features. For continuous features, a fully connected layer can be used to map the continuous features to a d-dimensional vector space. For discrete data, each discrete data can be embedded and encoded to a d-dimensional vector space. Discrete data belongs to ID features, and ID features can be distinguished by numbering different features. For example, categories can be numbered as 1, 2, 3, and 4. APPs can also be uniquely identified by numbering. For example, WeChat is numbered as 1, and Taobao is numbered as 2.
[0063] Specifically, for discrete data, the first step is to assign an index to each different discrete data, and it is assumed that there are N discrete data; the second step is to create an embedding matrix , and each index corresponds to a vector with a dimension of d.
[0064] For continuous data: each continuous data X e R is encoded into a d-dimensional vector using the following formula, , , are all learnable parameters:
[0065]
[0066] Continuous features, such as the user's use time of the APP (for example, how many seconds are used) and the like. For continuous features, since a continuous value cannot be uniquely numbered, it cannot be directly embedded and encoded. The embodiment of the application can design another way to encode, that is, by encoding data with completely different meanings into the same vector space (d-dimensional vector space), different meanings of data can be operated in the unified vector space (d-dimensional vector space).
[0067] For example, Figure 1As shown, the first encoding network is composed of an embedding layer, which can map all continuous features and discrete features in the user static data into a d-dimensional vector space, obtaining the output of this part: the first encoding result. The first encoding result will be one of the inputs of the multi-task encoding network.
[0068] Optionally, step 201 can include the following steps:
[0069] (11) The terminal device generates input data, positive label data and negative label data based on the user behavior-based time series data in the first sample; wherein the positive label data is the input data shifted one bit forward in the time dimension, and the negative label data is different from the positive label data in each time dimension;
[0070] (12) The terminal device inputs the input data, the positive label data and the negative label data into the second encoding network, and the second encoding network is used to embed and encode the input data, the positive label data and the negative label data respectively to map to a d-dimensional vector space, obtaining a d-dimensional encoding tensor corresponding to the input data, a d-dimensional encoding tensor corresponding to the positive label data and a d-dimensional encoding tensor corresponding to the negative label data, and d is an integer greater than or equal to 2;
[0071] (13) The terminal device performs self-attention encoding on the d-dimensional encoding tensor corresponding to the input data to obtain a second encoding result of the first sample, and the second encoding result is a d-dimensional encoding tensor.
[0072] In the embodiments of the present application, it is assumed that the user behavior-based time series data is: Wherein represents the event occurring for the th time in the sequence. In step (11), the input data of the second encoding network can be constructed as: And the training positive label data of the second encoding network is: The relationship between Seq Input and Seq Pos Label is that Seq Pos Label is Seq Input shifted one bit forward in the time dimension. Meanwhile, the training negative label data of the second encoding network is constructed as: The training negative label data is based on the training positive label data, and only needs to be different from the event of the positive label in the corresponding position through random sampling. For example, .
[0073] In order to enable the second encoding network to learn the user behavior preference, the user behavior preference can help the learning of the subsequent multi-task encoding network. By way of example, the time series data of the user behavior of the user in the past time is four records abcd arranged in time sequence, if the multi-task model can predict the record d when the records abc are given, it is considered that the multi-task model better captures the user behavior preference. The embodiment of the present application can guide the learning of the multi-task model by constructing the input data and the training positive label data by time dimension offset, constructing the input data x and the training positive label data y, expecting the input x = abc and the output y = bcd.
[0074] In step (12), the Seq Input in the first step, and are embedded and encoded, and all input data are unified into a d-dimensional vector space, wherein the encoding of the Seq Input is represented as , the encoding of the Seq Input is represented as , the encoding of the Seq Input is represented as .
[0075] In step (13), the is encoded by a self-attention layer, and the second encoding result is output, denoted as . is one of the inputs of the multi-task encoding network. Meanwhile, the and the and the are used for loss calculation for training the second encoding network, which will be described in detail later.
[0076] 203, the terminal device combines the encoding result of the first sample and the encoding result of the second sample, and inputs the multi-task encoding network to obtain the multi-task prediction result of the first sample.
[0077] In the embodiment of the present application, the encoding result of the first sample and the encoding result of the second sample can be combined (Concatenation) to obtain a merged encoding result (Concat Embedding Tensor), and the merged encoding result is input into the multi-task encoding network to obtain the multi-task prediction result of the first sample. The multi-task prediction result can include the prediction result of each task encoding subnetwork.
[0078] Optionally, the multi-task encoding network includes R task encoding subnetworks connected in sequence. For example, Figure 1As shown, each task encoding subnetwork includes a self-attention layer and a fully connected layer, and the output of the self-attention layer of the task encoding subnetwork of the i-th task in the current order (e.g., the i-th task encoding subnetwork in the current order in FIG. 2) is connected to the input of the self-attention layer of the task encoding subnetwork of the (i+1)-th task in the next order (e.g., the (i+1)-th task encoding subnetwork in the next order in FIG. 2). Figure 1 As shown, each task encoding subnetwork includes a self-attention layer and a fully connected layer, and the output of the self-attention layer of the task encoding subnetwork of the i-th task in the current order (e.g., the i-th task encoding subnetwork in the current order in FIG. 2) is connected to the input of the self-attention layer of the task encoding subnetwork of the (i+1)-th task in the next order (e.g., the (i+1)-th task encoding subnetwork in the next order in FIG. 2). Figure 1 As shown, each task encoding subnetwork includes a self-attention layer and a fully connected layer, and the output of the self-attention layer of the task encoding subnetwork of the i-th task in the current order (e.g., the i-th task encoding subnetwork in the current order in FIG. 2) is connected to the input of the self-attention layer of the task encoding subnetwork of the (i+1)-th task in the next order (e.g., the (i+1)-th task encoding subnetwork in the next order in FIG. 2).
[0079] Step 201 can include the following steps:
[0080] (21) The terminal device combines the encoding result of the first sample and the encoding result of the second sample to obtain a merged encoding result.
[0081] (22) The terminal device inputs the merged encoding result into the self-attention layers of the R task encoding subnetworks, respectively, and the fully connected layers of the R task encoding subnetworks output R prediction results, respectively, where R is an integer greater than or equal to 2.
[0082] In the embodiment of the present application, in step (21), the encoding result of the first sample and the encoding result of the second sample are combined (Concatenation) to obtain a merged encoding result (Concat Embedding Tensor), and the merged encoding result is input into the self-attention layers of the R task encoding subnetworks, and the fully connected layers of the R task encoding subnetworks output R prediction results, respectively.
[0083] The multi-task encoding network includes R task encoding subnetworks connected in sequence, and the size of R is determined according to the number of tasks. For example, if there are R tasks, the multi-task encoding network constructed includes R task encoding subnetworks connected in sequence, one task corresponding to one task encoding subnetwork. Among them, the R tasks have time correlation. Taking R=4 as an example, the R tasks include a task of predicting whether to click, a task of predicting whether to real-name, a task of predicting whether to grant credit, and a task of predicting whether to borrow. Since the four events of whether to click, whether to real-name, whether to grant credit, and whether to borrow have time correlation, only after clicking can the real-name be done, only after real-name can the credit be granted, and only after the credit is granted can the loan be done. If there is no click, the prediction results of the subsequent tasks of whether to real-name, whether to grant credit, and whether to borrow should all be no. Therefore, Figure 1 In the multi-task encoding network, the output of the self-attention layer of the task encoding subnetwork of the i-th task in the current order (e.g., the i-th task encoding subnetwork in the current order in FIG. 2) is connected to the input of the self-attention layer of the task encoding subnetwork of the (i+1)-th task in the next order (e.g., the (i+1)-th task encoding subnetwork in the next order in FIG. 2). Figure 1 In the multi-task encoding network, the output of the self-attention layer of the task encoding subnetwork of the i-th task in the current order (e.g., the i-th task encoding subnetwork in the current order in FIG. 2) is connected to the input of the self-attention layer of the task encoding subnetwork of the (i+1)-th task in the next order (e.g., the (i+1)-th task encoding subnetwork in the next order in FIG. 2). Figure 1The input of the self-attention layer of the i+1th task encoding subnetwork in the sequence can take the prediction result of the task encoding subnetwork in the current sequence as the input of the task encoding subnetwork in the next sequence, which can help the task encoding subnetwork in the next sequence to learn better and improve the training effect of the task encoding subnetwork.
[0084] In step (22), a multi-task network structure is constructed according to the number of tasks, and each task encoding subnetwork shares the merged encoding result as part of the input.
[0085] First, for the first task encoding subnetwork in the R task encoding subnetworks, it can be named as Task1. The Task1 only takes the merged encoding result as the input, obtains the output of the self-attention layer (Attention Layers) in the Task1 through the self-attention layer, and then obtains the prediction value of the Task1 (Target 1 Output) through the fully connected layer (DNN Layers) in the Task1. The loss calculation will be described in detail later.
[0086] For the downstream Task (not the first task encoding subnetwork) in the R task encoding subnetworks, it is named as Task i+1. In addition to taking the merged encoding result as the input, the output of the self-attention layer of the previous task encoding subnetwork (Task i) is also obtained. The output of the self-attention layer of the Task i is concatenated with the merged encoding result to obtain the input of the self-attention layer (Attention Layers) of the Task i+1, and then the prediction value of the Task i+1 (Target i +1 Output) is obtained through the fully connected layer of the Task i+1. The loss calculation will be described in detail later.
[0087] According to the scene, the number of task encoding subnetworks needs to be equal to the number of conversion links in the task scene.
[0088] For example, the execution actions of financial advertisements in different links will form the labels of whether to click, whether to real-name, whether to grant credit, and whether to borrow. When the multi-task model models the targets of whether to click, whether to real-name, whether to grant credit, and whether to borrow at the same time, the number of task encoding subnetworks needs to be 4. The four task encoding subnetworks will give the predictions of whether to click, whether to real-name, whether to grant credit, and whether to borrow.
[0089] 204, the terminal device calculates a training loss of the second encoding network based on the second encoding result, and calculates a training loss of the multi-task encoding network based on the multi-task prediction result.
[0090] In the embodiments of the present application, the binary cross-entropy loss function can be used to calculate the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0091] The loss of each task encoding sub-network in the task encoding network can be calculated based on the multi-task prediction result and the label corresponding to each task, and the losses of all task encoding sub-networks are added to obtain the training loss of the multi-task encoding network.
[0092] Optionally, in step 204, the terminal device calculates the training loss of the second encoding network based on the second encoding result, which can include the following steps:
[0093] The terminal device calculates the training loss of the second encoding network based on the second encoding result of the first sample, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data.
[0094] In the embodiments of the present application, as described above, the training loss of the second encoding network is calculated. The input data is embedded and encoded to obtain the second encoding result: The positive label data can also be encoded to obtain the d-dimensional encoding tensor corresponding to the positive label data: The negative label data can also be encoded to obtain the d-dimensional encoding tensor corresponding to the negative label data: .
[0095] The output of the self-attention layer (Attention Encoding Layer) of the second encoding network is respectively multiplied by and by matrix bit by bit, and then summed in each dimension, and then passed through a classification layer (such as a sigmoid layer) to obtain the relevance of the output data to the positive label data , and the relevance of the output data to the negative label .
[0096]
[0097]
[0098] The binary cross-entropy loss function can be used to make the correlation tend to 1, so that the correlation tends to 0, and the training loss of the second encoding network can be defined as:
[0099]
[0100] Optionally, in step 204, the terminal device calculates the training loss of the multi-task encoding network based on the multi-task prediction result, which can include the following steps.
[0101] The terminal device calculates the loss of each task encoding subnetwork based on the prediction result output by the full connection layer of each task encoding subnetwork and the corresponding label, adds the losses of the R task encoding subnetworks, and obtains the training loss of the multi-task encoding network.
[0102] In the embodiments of the present application, the binary cross-entropy loss function can be used to calculate the training loss of the multi-task encoding network.
[0103] For each task encoding subnetwork, there is a conversion task corresponding to a link. As described above, conversion is divided into two possibilities: successful conversion and non-conversion, and the corresponding labels (labels) are 1 and 0. Therefore, the loss of each task encoding subnetwork is calculated by the cross-entropy loss function, and then the loss values of all task encoding subnetworks are added together as the total loss of the multi-task encoding network .
[0104] In step 205, the loss values obtained by the two parts of the loss function can be added together (i.e., the training loss of the second encoding network and the training loss of the multi-task encoding network are added together), and the total loss of the multi-task model is obtained: Based on the total loss, the multi-task model is optimized. For example, the weight parameters of each layer in the multi-task model are updated.
[0105] The embodiments of the present application train the multi-task model, which can train the multi-task model based on the time series data of user behavior, thereby helping the multi-task model to better train, thereby improving the advertising conversion effect of each link.
[0106] 205, the terminal device optimizes the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0107] In the embodiments of the present application, when training the multi-task model, the samples not only contain user static data, but also contain time series data based on user behavior. Considering that the multi-task prediction result is related to the time sequence of user behavior, the multi-task model is trained based on the time series data of user behavior, thereby improving the training effect of the model.
[0108] The basic process of training of a classification model: collecting positive and negative samples -> training -> model convergence. When the training data is relatively clean, the training of the model can converge in a few rounds. When the training data is not clean and has a large number of noise samples mixed in, the number of times of model training can be doubled, and even a better convergence cannot be obtained. Usually, noise samples belong to samples that are difficult to learn or have large differences from clean samples. The noise samples are relatively difficult to converge in the learning process, and the loss decreases relatively slowly. Noise samples can affect the training effect of the model.
[0109] In order to clean the noise samples in the data set, the current method usually uses an optimized model structure to complete the identification and processing of noise samples, that is, without distinguishing between good and bad samples, all the samples are input into the model, and the identification and processing of good and bad samples are completed in the model, which can easily reduce the efficiency of model training.
[0110] Embodiments of the present application provide a sample cleaning method, please refer to Figure 3 , Figure 3 is a flowchart of a sample cleaning method provided by embodiments of the present application. The method can be applied to Figure 1 the multi-task model shown in the figure. The method can be executed before step 201. As shown in Figure 3 , the sample cleaning method can include the following steps.
[0111] 301, the terminal device acquires a set of training samples to be cleaned.
[0112] The set of training samples to be cleaned is a set of training samples that have not been cleaned.
[0113] 302, the terminal device generates P groups of training sets and P groups of validation sets based on the set of training samples; P is an integer greater than or equal to 2; any two groups of training sets in the P groups of training sets are not completely the same; any two groups of validation sets in the P groups of validation sets are not completely the same; the P groups of training sets and the P groups of validation sets correspond to each other, one group of training sets and the corresponding one group of validation sets form a set of training samples; the number of times that any sample in the set of training samples appears in the P validation sets is N, and N is an integer greater than or equal to 2.
[0114] In embodiments of the present application, generating P groups of training sets and P groups of validation sets can ensure that the number of times that any sample in the set of training samples appears in the P validation sets is N, which facilitates subsequent calculation of the confidence of each sample.
[0115] Optionally, in step 302, the terminal device generates P groups of training sets and P groups of validation sets based on the set of training samples, which can include the following steps:
[0116] (31) The terminal device divides the training sample set into M sample subsets, where M is an integer greater than or equal to 2;
[0117] (32) The terminal device generates M sets of training sets and M sets of verification sets based on the M sample subsets;
[0118] (33) The terminal device shuffles the order of the training sample set and performs the step of dividing the training sample set into M sample subsets;
[0119] (34) When the number of times of performing the step of dividing the training sample set into M sample subsets reaches N, the terminal device determines to generate P sets of training sets and P sets of verification sets, where P = M N.
[0120] 303. The terminal device inputs the P sets of training sets into P sample cleaning models respectively for training, to obtain P trained sample cleaning models; and inputs the P sets of verification sets into the corresponding P trained sample cleaning models respectively, to obtain N verification results of each sample in the training sample set.
[0121] In the embodiments of the present application, the sample cleaning model can be a simple and complex model with fast training speed, for example, a tree model.
[0122] 304. The terminal device determines the confidence of each sample based on the N verification results of each sample in the training sample set.
[0123] Optionally, the verification result includes a judgment value, and the step 304 can include the following steps:
[0124] (41) The terminal device records 1 for the judgment value as true and -1 for the judgment value as false in the N verification results of each sample, adds the N judgment values to obtain the score of each sample, marked as X;
[0125] (42) The terminal device takes X / Y as the confidence of each sample, where when the label value of each sample is true, Y = N, and when the label value of each sample is false, Y = -N.
[0126] 305. The terminal device selects a target sample set with a confidence greater than a set threshold from the training sample set, and the target sample set is used for training of the multi-task model.
[0127] In the embodiments of the present application, the set threshold can be set in advance. The confidence of each sample can be calculated, and the sample with high confidence can be selected according to the confidence of each sample, which can ensure the accuracy of the label of the target sample set, and further improve the training effect of the multi-task model in Figure 2 .
[0128] Machine learning models need a large amount of training data to update model parameters so that the prediction index of the model reaches a good state, therefore, a large amount of user browsing and operation behavior data needs to be collected as training data and a training set is obtained through manual annotation.
[0129] In financial services, for each user, the browsing and operation behavior data of the user on the device in a statistical window (a time interval artificially specified) is counted, and the execution action of the user on the financial advertisement issued by the online financial service in the statistical window is monitored, and these statistical data and execution actions together constitute the training data of the model (wherein the execution actions of the financial advertisement at different links form: whether to click, whether to real name, whether to grant credit, whether to borrow, etc.).
[0130] For such training data, ensuring the accuracy of the label is crucial for model training. However, in the financial scenario, for the same type of financial advertisement repeatedly reached by the same user in different statistical windows, there are often different execution actions for the same label, which leads to the problem of different sample labels for the same sample. Therefore, the present application provides a cleaning method for large-scale machine learning training set label data to solve the problem of interference caused by ambiguous label data to the model. For the same type of financial advertisement repeatedly reached by the same user in different statistical windows, there are often different execution actions (labels) under the premise of the same statistical characteristics (input), which leads to the problem of different labels (labels) corresponding to the same input.
[0131] The cleaning method can include the following steps:
[0132] 1. Obtain a set of training samples to be cleaned;
[0133] Among them, the training sample to be cleaned refers to the training set composed of the browsing and operation statistical data of each user collected by the financial server within a period of time and the execution action of the financial advertisement; wherein the label corresponding to the execution action is usually: whether to click, whether to real name, whether to grant credit, whether to borrow;
[0134] a) So-called click label: the monitoring of the execution action of the click link of the financial advertisement under the browsing and operation statistical data of a certain user, if the click is performed, the click label is marked as true, otherwise as false;
[0135] b) So-called real name label: the monitoring of the execution action of the real name link of the financial advertisement under the browsing and operation statistical data of a certain user. After the click link performs the click action, the execution action of the real name information registration is performed, then the real name label is marked as true, otherwise as false;
[0136] c) So-called credit label: the monitoring of the action of the financial advertisement credit link under the corresponding browsing and operation statistical data of a user. After the click link is clicked and the real-name link is registered, the credit is completed after passing the system audit, and the credit label is marked as true, otherwise as false;
[0137] d) So-called loan label: the monitoring of the action of the financial advertisement loan link under the corresponding browsing and operation statistical data of a user. After the click link is clicked, the real-name link is registered, the credit link is completed, and the loan operation is performed, the loan label is marked as true, otherwise as false;
[0138] Among them, a), b), c), d), the labels of the four links are related: if the label of the previous link is false, the label of the subsequent link must be false.
[0139] 2. Divide the training sample into M parts, and divide the M parts of data into two parts: the first part as the validation set, the second to M parts as the training set, which is the first group. The second part is used as the validation set, and the first and third to M parts are used as the training set, which is the second group. The a-th part is used as the validation set, and the first to a-1 parts and the a+1 to M parts are used as the training set, which is the a-th group. After repeating M times, the last group is: the last M-th part is used as the validation set, and the first to M-1 parts are used as the training set, which is the M-th group. A total of M groups of "training set and validation set" can be obtained.
[0140] Among them, M is a hyperparameter and can be defined artificially. Generally, for large-scale data sets, the number of data in the data set can be divided by M, so that it can be evenly divided into M parts.
[0141] Each group consists of M-1 training sets and 1 validation set. For example, a training set has 100 samples, which is divided into M=4 parts; that is, the data set =【1, 2, 3, 4】, each part has 25 samples. Through the method of the present application, the following four groups of data sets can be obtained:
[0142] The first group: the training set consists of 75 samples from the 【1, 2, 3】 part, and the validation set consists of 25 samples from the 【4】 part;
[0143] The second group: the training set consists of 75 samples from the 【1, 2, 4】 part, and the validation set consists of 25 samples from the 【3】 part;
[0144] The third group: the training set consists of 75 samples from the 【1, 3, 4】 part, and the validation set consists of 25 samples from the 【2】 part;
[0145] Group 4: The training set consists of 75 samples from parts 【4, 2, 3】 and the validation set consists of 25 samples from part 【1】.
[0146] 3. Shuffle the training samples and generate a new set of M groups of training sets and validation sets according to step 2. Shuffle and repeat N times to finally obtain N sets of M groups of training sets and validation sets.
[0147] From the above distribution, it can be found that each sample will only appear once in the validation set. By training each group of data with the model, only after running the first, second, third, and fourth data sets, each sample will have a prediction value from the model. Therefore, through N batches, each sample will have N prediction values from the model, which prepares for the calculation of confidence in the following steps.
[0148] 4. Use the tree model to train each group of training sets in the set, and after training, infer the data of the corresponding validation set to obtain the judgment value of the model on whether to click, whether to real-name, whether to grant credit, and whether to borrow for each data in the validation set (the judgment value options are true or false).
[0149] 5. Through step 4, for each label of each data, there are N judgment values derived from the tree model. At the same time, for each label of each data, there is a manually labeled label value. According to the N judgment values, set a judgment threshold, and when the sample is greater than the threshold, it is considered as a label error noise sample and is removed.
[0150] The judgment method is as follows:
[0151] Set the judgment value true to 1 and the judgment value false to -1. Add the N judgment values to obtain the score of the sample, marked as X.
[0152] Set the reference value of the manually labeled label true to Y = N and the reference value of the manually labeled label false to Y = -N.
[0153] The confidence calculation formula of the sample label is S = X / Y, then S ∈ [-1, 1], S is close to 1, which indicates that the confidence of the sample label is higher; S is closer to 0, which indicates that the confidence of the sample label is lower.
[0154] Set the threshold value and take out the samples with confidence S greater than the set threshold value to form a new training set (i.e., the target sample set).
[0155] In the embodiments of the present application, for the easily confused samples in the training sample set, due to different statistical windows, labeling errors and other reasons, the user data with similar statistical information have different labels. The label noise cleaning problem is preposed, and the noise samples are identified before the data is input into the multi-task model, so as to improve the effect of the multi-task model trained by using the data set and improve the training speed of the multi-task model. At the same time, with the assistance of the tree model, the shortcomings of large data volume, sample label checking and cleaning difficulty are overcome, the efficiency of label noise cleaning of large-scale data set is improved, and the accuracy of data labeling is improved.
[0156] The above mainly introduces the scheme of the embodiments of the present application from the perspective of the execution process of the method. It can be understood that the terminal device includes hardware structure and / or software modules corresponding to the execution of each function in order to implement the above functions. Those skilled in the art should easily realize that, in combination with the unit and algorithm steps of each example described in the embodiments provided herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0157] The embodiments of the present application can divide the functional units of the terminal device according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be realized in the form of hardware or software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical function division. Actual implementation can have another division method.
[0158] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of a model training device provided by the embodiments of the present application. The model training device 400 is applied to a terminal device, and the model training device 400 is applied to a multi-task model. The multi-task model includes a first encoding network, a second encoding network and a multi-task encoding network. The model training device 400 can include a first encoding unit 401, a second encoding unit 402, a prediction unit 403, a loss calculation unit 404 and an optimization unit 405, wherein:
[0159] The first encoding unit 401 is configured to input user static data in the first sample into the first encoding network to obtain a first encoding result of the first sample. The first sample is any one in a target sample set.
[0160] The second encoding unit 402 is configured to input the user behavior-based time series data in the first sample into the second encoding network to obtain a second encoding result of the first sample.
[0161] The prediction unit 403 is configured to input the encoding result of the first sample and the encoding result of the second sample into the multi-task encoding network to obtain a multi-task prediction result of the first sample.
[0162] The loss calculation unit 404 is configured to calculate a training loss of the second encoding network and a training loss of the multi-task encoding network.
[0163] The optimization unit 405 is configured to optimize the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network.
[0164] Optionally, the first encoding unit 401 inputs the user static data in the first sample into the first encoding network to obtain a first encoding result of the first sample, and the method comprises the following steps.
[0165] The first encoding unit 401 inputs the user static data in the first sample into the first encoding network, and the first encoding network is configured to perform embedding encoding on the user static data to map to a d-dimensional vector space to obtain a first encoding result of the first sample, wherein the first encoding result is a d-dimensional encoding tensor, and d is an integer greater than or equal to 2.
[0166] Optionally, the second encoding unit 402 inputs the user behavior-based time series data in the first sample into the second encoding network to obtain a second encoding result of the first sample, and the method comprises the following steps.
[0167] The second encoding unit 402 inputs the user behavior-based time series data in the first sample into the second encoding network to obtain a second encoding result of the first sample, and the method comprises the following steps.
[0168] The second encoding unit 402 inputs the user behavior-based time series data in the first sample into the second encoding network to obtain a second encoding result of the first sample, and the method comprises the following steps.
[0169] The input data corresponding d-dimensional encoding tensor is self-attention encoded to obtain a second encoding result of the first sample, and the second encoding result is a d-dimensional encoding tensor.
[0170] Optionally, the loss calculation unit 404 calculates the training loss of the second encoding network based on the second encoding result, including:
[0171] The training loss of the second encoding network is calculated based on the second encoding result of the first sample, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data.
[0172] Optionally, the multi-task encoding network includes R task encoding sub-networks connected in sequence, each task encoding sub-network includes a self-attention layer and a fully connected layer, and the output of the self-attention layer of the current sequential task encoding sub-network is connected to the input of the self-attention layer of the next sequential task encoding sub-network.
[0173] The prediction unit 403 inputs the combined encoding result of the first sample and the second sample into the multi-task encoding network to obtain a multi-task prediction result of the first sample, including:
[0174] The encoding result of the first sample and the encoding result of the second sample are combined to obtain a merged encoding result.
[0175] The merged encoding result is input into the self-attention layer of the R task encoding sub-networks respectively, and the fully connected layer of the R task encoding sub-networks outputs R prediction results respectively, R being an integer greater than or equal to 2.
[0176] Optionally, the loss calculation unit 404 calculates the training loss of the multi-task encoding network based on the multi-task prediction result, including:
[0177] The loss of each task encoding sub-network is calculated based on the prediction result output by the fully connected layer of each task encoding sub-network and the corresponding label, and the training loss of the multi-task encoding network is obtained by adding the losses of the R task encoding sub-networks.
[0178] Optionally, the model training apparatus 400 can further include a cleaning unit 406.
[0179] The cleaning unit 406 is configured to obtain a training sample set to be cleaned, generate P groups of training sets and P groups of verification sets based on the training sample set, P is an integer greater than or equal to 2, any two groups of training sets in the P groups of training sets are not completely the same, any two groups of verification sets in the P groups of verification sets are not completely the same, the P groups of training sets and the P groups of verification sets correspond to each other, one group of training sets and a corresponding group of verification sets form the training sample set, and the number of times that any sample in the training sample set appears in the P groups of verification sets is N, N is an integer greater than or equal to 2. The P groups of training sets are respectively input into P sample cleaning models for training to obtain P trained sample cleaning models, the P groups of verification sets are respectively input into the P trained sample cleaning models corresponding thereto, N verification results of each sample in the training sample set are obtained, the confidence of each sample is determined based on the N verification results of each sample in the training sample set, and a target sample set with a confidence greater than a set threshold is selected from the training sample set, and the target sample set is used for training of a multi-task model.
[0180] Optionally, the cleaning unit 406 generates the P groups of training sets and the P groups of verification sets based on the training sample set, and the method comprises the following steps.
[0181] The training sample set is divided into M sample subsets, and M is an integer greater than or equal to 2.
[0182] M groups of training sets and M groups of verification sets are generated based on the M sample subsets.
[0183] The order of the training sample set is shuffled, and the step of dividing the training sample set into M sample subsets is performed.
[0184] When the number of times that the step of dividing the training sample set into M sample subsets is performed reaches N times, it is determined that the P groups of training sets and the P groups of verification sets are generated, and P=M N.
[0185] Optionally, the verification result comprises a judgment value, and the cleaning unit 406 determines the confidence of each sample based on the N verification results of each sample in the training sample set, and the method comprises the following steps.
[0186] The judgment value of each sample is recorded as 1 when the judgment value is true, and the judgment value is recorded as -1 when the judgment value is false, the N judgment values are added to obtain the score of each sample, and the score is marked as X.
[0187] X / Y is taken as the confidence of each sample, wherein when the label value of each sample is true, Y=N, and when the label value of each sample is false, Y=-N.
[0188] The first encoding unit 401, the second encoding unit 402, the prediction unit 403, the loss calculation unit 404, the optimization unit 405, and the cleaning unit 406 in the embodiments of the present application can be processors in a terminal device.
[0189] Figure 4 For specific implementation of the model training apparatus 400 shown, refer to Figure 2 or the method embodiments shown in FIG. 3, which will not be described herein.
[0190] In the embodiments of the present application, when training the multi-task model, the samples contain not only user static data but also time series data based on user behavior. Considering that the multi-task prediction result is related to the time sequence of user behavior, the multi-task model is trained based on the time series data of user behavior, thereby improving the training effect of the model.
[0191] For specific implementation of the model training apparatus 400 shown, refer to Figure 5 , Figure 5 is a structural schematic diagram of a terminal device provided by an embodiment of the present application, as Figure 5 shown, the terminal device 500 includes a processor 501 and a memory 502, and the processor 501 and the memory 502 can be connected to each other through a communication bus 503. The communication bus 503 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 503 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The memory 502 is used to store a computer program, and the computer program includes program instructions. The processor 501 is configured to invoke the program instructions. The above program includes program instructions for executing Figure 2 or Figure 3 the steps in the method shown in the figure.
[0192] The memory 502 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.
[0193] In the embodiments of the present application, when the multi-task model is trained, the samples contain not only user static data but also time series data based on user behavior. Considering that the multi-task prediction result is related to the time sequence of user behavior, the multi-task model is trained based on the time series data of user behavior, so as to improve the training effect of the model.
[0194] The embodiments of the present application also provide a computer readable storage medium, wherein the computer readable storage medium stores a computer program for electronic data exchange, and the computer program causes the computer to execute part or all steps of any one of the model training methods described in the above method embodiments.
[0195] It should be noted that, for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0196] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0197] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely an example, and the division can be other division manners. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0198] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0199] In addition, the functional units in each embodiment of the application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software program module.
[0200] The integrated unit, if realized in the form of a software program module and sold or used as an independent product, can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal device or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0201] A person of ordinary skill in the art can understand that all or part of the steps of the various methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.
[0202] The above has carried out the detailed introduction to the embodiment of the application, the principle and implementation mode of the application have been described by applying specific examples in this paper, the above embodiment explanation is only for helping understanding the method of the application and its core idea; at the same time, for the general technical personnel in the art, according to the idea of the application, there will be changes in specific implementation mode and application range, and the above-mentioned, the content of the specification should not be understood as the limitation of the application.
Claims
1. A model training method, characterized in that, The method is applied to a multi-task model, which includes a first encoding network, a second encoding network, and a multi-task encoding network, wherein the multi-task encoding network includes R sequentially connected task encoding sub-networks; the method includes: The user static data in the first sample is input into the first encoding network to obtain the first encoding result of the first sample; the first sample is any one in the target sample set; the user static data includes: user ID features, user statistical features, and user demographic attributes; The time-series data based on user behavior in the first sample is input into the second encoding network to obtain the second encoding result of the first sample; the time-series data includes a series of operations performed by the user in chronological order; The first encoding result and the second encoding result are combined and input into the multi-task encoding network to obtain the multi-task prediction result of the first sample; the multi-task prediction result includes the prediction results of the R task encoding sub-networks, the prediction results of the R task encoding sub-networks are the prediction results of user behavior, and the prediction results of the R task encoding sub-networks have time correlation. The training loss of the second encoding network is calculated based on the second encoding result, and the training loss of the multi-task encoding network is calculated based on the multi-task prediction result. The multi-task model is optimized based on the training loss of the second encoding network and the training loss of the multi-task encoding network. The step of inputting the time-series data based on user behavior from the first sample into the second encoding network to obtain the second encoding result of the first sample includes: Input data, positive label data, and negative label data are generated based on the time-series data of user behavior in the first sample; wherein, the positive label data is the input data shifted forward one position in the time dimension, and the negative label data is different from the positive label data in each time dimension; The input data, the positive label data, and the negative label data are input into the second encoding network. The second encoding network is used to embed and encode the input data, the positive label data, and the negative label data respectively, so as to map them to a d-dimensional vector space, and obtain the d-dimensional encoding tensor corresponding to the input data, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data, where d is an integer greater than or equal to 2. The input data is self-attention encoded into a d-dimensional encoding tensor to obtain the second encoding result of the first sample. The second encoding result is a d-dimensional encoding tensor. The step of calculating the training loss of the second encoding network based on the second encoding result includes: The training loss of the second encoding network is calculated based on the second encoding result of the first sample, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data.
2. The method according to claim 1, characterized in that, The step of inputting the user static data from the first sample into the first encoding network to obtain the first encoding result of the first sample includes: The user static data in the first sample is input into the first encoding network. The first encoding network is used to embed and encode the user static data to map it to a d-dimensional vector space, thereby obtaining the first encoding result of the first sample. The first encoding result is a d-dimensional encoding tensor, where d is an integer greater than or equal to 2.
3. The method according to claim 1, characterized in that, Each task coding subnetwork includes a self-attention layer and a fully connected layer. The output of the self-attention layer of the current task coding subnetwork is connected to the input of the self-attention layer of the next task coding subnetwork. The step of combining the first encoding result and the second encoding result and inputting them into the multi-task encoding network to obtain the multi-task prediction result of the first sample includes: The first encoding result and the second encoding result are combined to obtain the merged encoding result; The merged encoding results are respectively input into the self-attention layers of the R task encoding sub-networks, and the fully connected layers of the R task encoding sub-networks respectively output R prediction results, where R is an integer greater than or equal to 2.
4. The method according to claim 3, characterized in that, The calculation of the training loss of the multi-task coding network based on the multi-task prediction results includes: The loss of each task coding subnetwork is calculated based on the prediction results output by the fully connected layer of each task coding subnetwork and the corresponding label. The losses of the R task coding subnetworks are added together to obtain the training loss of the multi-task coding network.
5. The method according to any one of claims 1 to 4, characterized in that, Before inputting the user static data from the first sample into the first encoding network, the method further includes: Obtain the set of training samples to be cleaned; Based on the training sample set, P training sets and P validation sets are generated; P is an integer greater than or equal to 2; no two training sets in the P training sets are completely identical; no two validation sets in the P validation sets are completely identical; the P training sets and the P validation sets are in one-to-one correspondence, and a training set and its corresponding validation set constitute the training sample set; any sample in the training sample set appears N times in the P validation sets, where N is an integer greater than or equal to 2. The P training sets are input into the P sample cleaning models respectively for training, resulting in P trained sample cleaning models; the P validation sets are input into the corresponding P trained sample cleaning models respectively, resulting in N validation results for each sample in the training sample set. The confidence level of each sample is determined based on the N validation results of each sample in the training sample set; A target sample set with a confidence level greater than a set threshold is selected from the training sample set, and the target sample set is used for training the multi-task model.
6. The method according to claim 5, characterized in that, The step of generating P training sets and P validation sets based on the training sample set includes: The training sample set is divided into M sample subsets, where M is an integer greater than or equal to 2; Based on the M sample subsets, generate M training sets and M validation sets; Disorder the training sample set and perform the step of dividing the training sample set into M sample subsets; When the step of dividing the training sample set into M sample subsets is performed N times, P training sets and P validation sets are generated, where P=M. N.
7. The method according to claim 5, characterized in that, The verification result includes a judgment value. The determination of the confidence level for each sample based on the N verification results for each sample in the training sample set includes: For each sample, a true value is recorded as 1 and a false value as -1. The N values are summed to obtain the score of each sample, which is marked as X. X / Y is used as the confidence level for each sample, where Y=N when the label value of each sample is true and Y=-N when the label value of each sample is false.
8. A model training device, characterized in that, The apparatus is applied to a multi-task model, the multi-task model comprising: a first encoding network, a second encoding network, and a multi-task encoding network, the multi-task encoding network comprising R sequentially connected task encoding sub-networks; the apparatus comprises: The first encoding unit is used to input the user static data in the first sample into the first encoding network to obtain the first encoding result of the first sample; the first sample is any one in the target sample set; the user static data includes: user ID features, user statistical features, and user demographic attributes. The second encoding unit is used to input the time-series data based on user behavior from the first sample into the second encoding network to obtain the second encoding result of the first sample; the time-series data includes a series of operations performed by the user in a chronological order. The prediction unit is used to combine the first encoding result and the second encoding result and input them into the multi-task encoding network to obtain the multi-task prediction result of the first sample; the multi-task prediction result includes the prediction results of the R task encoding sub-networks, the prediction results of the R task encoding sub-networks are time-related, and the prediction results of the R task encoding sub-networks are the prediction results of user behavior. The loss calculation unit is used to calculate the training loss of the second encoding network and the training loss of the multi-task encoding network. An optimization unit is used to optimize the multi-task model based on the training loss of the second encoding network and the training loss of the multi-task encoding network. The second encoding unit inputs the user behavior-based time-series data from the first sample into the second encoding network to obtain the second encoding result of the first sample, including: generating input data, positive label data, and negative label data based on the user behavior-based time-series data from the first sample; wherein, the positive label data is the input data shifted forward one position in the time dimension, and the negative label data is different from the positive label data in each time dimension; inputting the input data, the positive label data, and the negative label data into the second encoding network, the second encoding network is used to embed and encode the input data, the positive label data, and the negative label data respectively to map them to a d-dimensional vector space, to obtain the d-dimensional encoding tensor corresponding to the input data, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data, where d is an integer greater than or equal to 2; performing self-attention encoding on the d-dimensional encoding tensor corresponding to the input data to obtain the second encoding result of the first sample, the second encoding result being a d-dimensional encoding tensor; The loss calculation unit calculates the training loss of the second encoding network based on the second encoding result, including: calculating the training loss of the second encoding network based on the second encoding result of the first sample, the d-dimensional encoding tensor corresponding to the positive label data, and the d-dimensional encoding tensor corresponding to the negative label data.
9. A terminal device, characterized in that, The device includes a processor and a memory, the memory being used to store a computer program, the computer program including program instructions, and the processor being configured to invoke the program instructions to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Advertisement click rate estimation method based on user preferences
CN112700274A
Task prediction method and device, equipment and storage medium
CN113822439A
Multi-task model training method, multi-task prediction method, related device and medium
CN115049108A