Model processing method, device, storage medium and electronic equipment

By copying and randomly retracting the training data of the initial neural network model, combined with self-supervised learning, the model instability caused by the differences in offline and online data distribution is solved, and the robustness and prediction effect of the model are improved.

CN115496250BActive Publication Date: 2025-09-02ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210581616.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-09-02
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

In daily transaction scenarios, machine learning models have unstable prediction effects due to the difference in data distribution between offline training data and online transaction data, and the model processing effect is poor.

Method used

By copying the first training data of the initial neural network model, generating the second training data, and inputting the two into the model for training, random retraction processing and self-supervised learning, the model training process is optimized and the model's robustness is improved.

Benefits of technology

It realizes good adaptability to changes in online training data during model training, improves the prediction effect and stability after the model is launched, reduces dependence on label data, and improves model processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496250B_ABST
    Figure CN115496250B_ABST
Patent Text Reader

Abstract

This specification discloses a model processing method, device, storage medium and electronic device, wherein the method includes: performing data copy processing on first training data for an initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information to obtain a trained target neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model processing method, device, storage medium, and electronic device. Background Art

[0002] In everyday business scenarios, content recommendations for users' corresponding clients are often involved, such as displaying information streams, recommending information, and precisely delivering content to specific users. Currently, machine learning-based neural network models are often used in these scenarios. Typically, machine learning is used offline to collect user data over a period of time to train the corresponding business model, and then the trained model is applied online to daily business scenarios. Summary of the Invention

[0003] This specification provides a model processing method, device, storage medium, and electronic device. The technical solution is as follows:

[0004] In a first aspect, this specification provides a model processing method, the method comprising:

[0005] Performing data replication processing on the first training data of the initial neural network model to obtain second training data;

[0006] Inputting the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data;

[0007] The initial neural network model is trained based on the first representation information and the second representation information to obtain a target neural network model.

[0008] In a second aspect, this specification provides a model processing device, the device comprising:

[0009] A data replication module is used to perform data replication processing on the first training data for the initial neural network model to obtain second training data;

[0010] a model processing module, configured to input the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data;

[0011] A model training module is used to perform model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model.

[0012] In a third aspect, this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0013] In a fourth aspect, this specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0014] The beneficial effects of the technical solutions provided by some embodiments of the present application include at least:

[0015] In one or more embodiments of the present specification, by performing data copy processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved. The target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is a schematic diagram of a model training system scenario provided in this manual;

[0018] Figure 2 It is a flowchart of a model processing method provided in this manual;

[0019] Figure 3 It is a flowchart of a model processing method provided in this manual;

[0020] Figure 4 This is a schematic diagram of a random fade-out processing scenario provided in this manual;

[0021] Figure 5This is a schematic diagram of another random retirement processing scenario provided in this manual;

[0022] Figure 6 This is a schematic diagram of a model training scenario provided in this manual;

[0023] Figure 7 It is a flowchart of a model processing method provided in this manual;

[0024] Figure 8 It is a structural diagram of a model processing device provided in this specification;

[0025] Figure 9 This is a structural diagram of a model processing module provided in this specification;

[0026] Figure 10 This is a schematic diagram of the structure of an information determination unit provided in this specification;

[0027] Figure 11 This is a structural diagram of a model training module provided in this manual;

[0028] Figure 12 This is a structural diagram of another model training device provided in this manual;

[0029] Figure 13 This is a structural diagram of an electronic device provided in this manual. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in this specification to clearly and completely describe the technical solutions in this specification. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0032] In the daily transaction scenarios of related technologies, machine learning usually has data distribution differences between offline labeled training data and online transaction data. By collecting user data over a period of time as offline data and annotating it with corresponding labels, this part of labeled training data is used for machine learning to train the corresponding model. However, in the online prediction stage after the model is launched, it is found that there is a huge difference in the data distribution (such as user distribution) between offline training and online services, resulting in unstable model prediction results and poor model processing results in daily transaction scenarios.

[0033] See Figure 1 , is a schematic diagram of a scenario of a model training system provided in this specification. Figure 1 As shown, the model training system may include at least a client cluster and a service platform 100 .

[0034] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.

[0035] Each client in the client cluster can be an electronic device with communication capabilities, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic devices in 5G network or future evolution network, etc.

[0036] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device with strong computing capabilities; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be symmetrically composed, wherein each server has equivalent functions and status in the transaction link, and each server can provide services to the outside world independently. The independent service can be understood as not requiring the assistance of other servers.

[0037] In one or more embodiments of the present specification, the service platform 100 may establish a communication connection with at least one client in the client cluster, and complete data interaction during the model training process based on the communication connection, such as online transaction data interaction. For example, the service platform 100 may recommend content to the client based on the target neural network model obtained based on the model processing method of the present specification; for example, the service platform 100 may obtain training data from the client, such as the first training data.

[0038] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, wherein the network can be a wireless network or a wired network, the wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network, and the wired network includes but is not limited to Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data exchanged over the network (such as a target compressed package). In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.

[0039] The model training system embodiments provided in this specification and the model processing methods described in one or more embodiments are based on the same concept. The execution entity corresponding to the model processing methods described in one or more embodiments of the specification may be the aforementioned service platform 100; the execution entity corresponding to the model processing methods described in one or more embodiments of the specification may also be the electronic device corresponding to the client, depending on the actual application environment. The implementation process of the model training system embodiments can be found in the following method embodiments and will not be repeated here.

[0040] based on Figure 1 The scene diagram shown is a detailed introduction to the model processing method provided by one or more embodiments of this specification.

[0041] See Figure 2 , provides a flow chart of a model processing method for one or more embodiments of this specification. This method can be implemented using a computer program and run on a background investigation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The model processing device can be a service platform.

[0042] Specifically, the model processing method includes:

[0043] S102, performing data replication processing on the first training data of the initial neural network model to obtain second training data;

[0044] The initial neural network model can be created by fitting one or more models such as a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN) model, an embedding model, a gradient boosting decision tree (GBDT) model, and a logistic regression (LR) model.

[0045] As can be understood, a neural network model is typically composed of a densely interconnected configuration of simple nonlinear analog processing elements within each of many nodes, and is a system model that mimics biological neurons. The neural network model is formed by connecting the input of at least one node (one node is equivalent to a neuron) to the output of each node, similar to the synaptic connections of real neurons. Each neuron expresses a specific output function, namely, an activation function, and the connection between each two neurons includes a connection strength, which is a weighted value acting on the signal passing through the connection.

[0046] In one or more embodiments of the present specification, a corresponding initial neural network model can be constructed based on an actual transaction scenario to perform model training to obtain a trained target neural network model. For example, the actual transaction scenario can be a content recommendation scenario, in which accurate content recommendations for different client users are achieved based on the target neural network model used for content recommendation, such as advertising content information recommendation and feed information flow display. In the aforementioned process, the target neural network model can be used to predict the user's click probability on a certain recommended content; for another example, the actual transaction scenario can be a user conversion scenario, in which the target neural network model is used to identify the corresponding user transaction data and output the conversion probability for the corresponding user, etc.

[0047] In one or more embodiments of this specification, the first training data can be understood as sample data for the initial neural network model. Some of the training data in the first training data can be pre-labeled with corresponding labels. The labeled training data can also be referred to as labeled training data. The labels are determined based on actual transaction scenarios. Unlabeled training data can be introduced into the first training data to enrich the training data and enhance the recognition effect after the model is launched.

[0048] In one feasible implementation, the first training data may include labeled training data and unlabeled training data. Labeled training data for a target transaction scenario may be obtained, followed by unlabeled online transaction data for the target transaction scenario. The first training data may then be generated based on the labeled training data and the online transaction data.

[0049] It can be understood that in actual online transaction scenarios, offline transaction data of users over a period of time is collected offline, and labels are annotated on these actual transaction data to obtain labeled training data. In order to improve the online application effect of the model, a large amount of online transaction data of online users in actual transaction scenarios is introduced to resist the data distribution differences between offline labeled training data and online transaction data. The model processing method of this specification focuses on the different representation information of the same training data in different model data processing stages in the model output, realizing a new model training processing method, which changes the need to annotate training data with a large number of labels in related technologies, and can get rid of the dependence on annotating training data while taking into account the accuracy of model recognition. A large amount of online transaction training data can be introduced in the model training stage for model training; the training data of online transactions is usually unlabeled training data, which also solves the problem of difficult model convergence caused by the introduction of a large amount of unlabeled online transaction data. It provides a method for introducing unlabeled online transaction data in online transaction scenarios to realize self-supervised learning of the model to train the model.

[0050] Furthermore, first training data is generated based on the labeled training data and the online transaction data. In some embodiments, the first training data generated by mixing a large amount of online transaction data can also be referred to as a pseudo model (data) set, so as to improve the model's expressive ability for online user representation, and improve the model's robustness in the face of changes in user distribution, improve the stability of online predictions, and realize self-supervised learning of the model by introducing unlabeled data in online transaction scenarios. In addition, the online transaction scenario can be an online information recommendation scenario.

[0051] Optionally, the amount of online transaction data may be greater than the amount of label training data.

[0052] In one or more embodiments of the present specification, in addition to obtaining labeled training data with clear labels during the data preparation stage, a large amount of all unlabeled user data, such as the aforementioned online transaction data, is also obtained, and then the first user training data for model training of the initial neural network model is obtained.

[0053] It is understood that the first user training data is usually directly input into the initial neural network model for training. However, in this application, the first user training data is copied to obtain the second user training data. Then, the first user training data and the second user training data are input into the initial neural network model together.

[0054] It should be noted that data replication for the training data of the input model is not performed because the amount of training data is insufficient.

[0055] Optionally, the first training data is copied to obtain the second training data, which may be done by a full copy method, that is, each data file in the first training data is copied to obtain the second training data that is completely identical to the first training data;

[0056] Optionally, the first training data is copied to obtain the second training data, and an incremental copy method can be used, that is, instead of copying each data file in the first training data, some data files in the first training data are copied to obtain the second training data; in this case, the second training data can be regarded as a subset of the first training data.

[0057] S104: Input the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data;

[0058] The first training data and the second training data can be understood as inputs of the initial neural network model in the model training phase;

[0059] The representation information can be understood as the output of the initial neural network model during the model training process, and the representation information can usually be in the form of a representation vector;

[0060] The first representation information is an actual output result obtained by performing model recognition on the first training data using the initial neural network model, that is, the representation information for the first training data;

[0061] The second representation information is an actual output result obtained by performing model recognition on the second training data using the initial neural network model, that is, the representation information for the second training data;

[0062] In one or more embodiments of the present specification, when the initial neural network model processes and identifies the first training data and the second training data, it controls at least one neuron in the initial neural network model (the neuron can be regarded as a computing processing node in the model) to be in a different node state in each round of model training stage. For example, the neuron can be controlled to be in a retreat state and not respond to data processing during this round of training; based on this, the initial neural network model will have data output differences for the same training data (the same training data in the first training data and the second training data) in different processing stages, that is, the output representation results (first representation information) of the initial neural network model for the first user training data and the output representation results (second representation information) for the second user training data have data differences; the model parameters are adjusted in the back propagation stage through the data differences to achieve the purpose of model training in the model training process: for the same training data in different processing stages, the representation information output by the model should be consistent, based on which a target neural network model that meets expectations can be trained.

[0063] S106: Perform model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model.

[0064] It can be understood that during the model training process, at least one data file in the first training data and / or the second training data is input into the initial neural network model to start the training of the initial neural network model. In the forward propagation training process, the model loss is calculated based on the loss function for each round of model output (that is, the first representation information and the second representation information of the corresponding data file), and the model loss is used to adjust the model parameters and even the model structure of the initial neural network model during the back propagation training process. The training goal of the model is to maximize the output similarity between the same training data at different processing stages, that is, the data (or information) similarity between the first representation information and the second representation information, so that the output difference between the first representation information and the second representation information is reduced (usually, the similarity between the representation vector of the first representation information and the representation vector of the second representation information is maximized). In other words, the similarity between the two features in the first representation information and the second representation information in the mapped vector space is maximized until the training end condition preset for the initial neural network model is met. The training end condition may include, for example, the value of the total loss function is less than or equal to the preset loss function threshold, the number of iterations reaches the preset number threshold, etc. The specific training end condition can be determined based on the actual business scenario and will not be repeated here.

[0065] In a specific implementation scenario, after obtaining the trained target neural network model, the target neural network model can be put online to obtain online user transaction data, input the user transaction data into the target neural network model, and output the transaction conversion probability for the user transaction data.

[0066] In one or more embodiments of the present specification, by performing data copy processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved. The target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched.

[0067] See Figure 3 , Figure 3 This is a flow chart of another embodiment of a model processing method proposed in one or more embodiments of this specification. Specifically:

[0068] S202: performing data replication processing on the first training data of the initial neural network model to obtain second training data;

[0069] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here.

[0070] S204: Inputting the first training data and the second training data into the initial neural network model;

[0071] In a feasible implementation, after inputting the first training data and the second training data into the initial neural network model, the initial neural network model is subjected to random retirement processing, and first representation information for the first training data and second representation information for the second training data are determined based on the initial neural network model after the random retirement processing;

[0072] It can be understood that by controlling the initial neural network model to perform random retirement processing in each round of model training, at least one neuron in the initial neural network model (neurons can be regarded as computing processing nodes in the model) is in a different node state. For example, the neuron can be controlled to be in a retired state and not respond to data processing during this round of training; based on this, the initial neural network model will have different data output representations for the same training data in different processing stages (the same training data in the first training data and the second training data).

[0073] S206: performing a first random falloff process on at least one first neuron in the initial neural network model, and outputting first representation information for first training data based on the initial neural network model after the first random falloff process;

[0074] The randomized backoff process involves "shielding" at least one neuron in the initial neural network model during the forward propagation of the training data in each round. This process effectively causes certain neurons to back off. These backed-off neurons do not participate in the model processing in this round. Instead, neurons in the initial neural network model that have not been randomly backed off participate in processing the training data in this round. The random backoff process is repeated in the next round.

[0075] The first random retreat processing can be understood as a random retreat processing process performed when the initial neural network model is used to perform model processing on at least one training data included in the first training data.

[0076] The first neuron is the subject of a first retirement process. During model processing of at least one training data item included in the first training data using the initial neural network model, the at least one first neuron is controlled to be "shielded" and not participate in the model processing of the first training data item in the current round. In the next round, the aforementioned random retirement process is repeated until model training is completed.

[0077] It can be understood that after performing the first random retreat processing on at least one first neuron in the initial neural network model, first representation information for the first training data will be output based on neurons other than all first neurons in the initial neural network model.

[0078] In one feasible embodiment, at least one random retreat layer is provided in the initial neural network model, and based on the at least one random retreat layer, the neurons surrounding the initial neural network model are controlled to perform random retreat processing. The random retreat layer includes a plurality of random retreat units (which can be understood as a type of neuron).

[0079] Optionally, a random retreat layer can be set in at least one subnetwork (such as an input layer network, a hidden layer network, a convolutional layer network, etc.) in a commonly used initial neural network model; the subnetwork is composed of several subnetwork layers and several random retreat layers. For example, taking the subnetwork as a hidden layer network as an example, the relevant hidden layer network usually only contains hidden layers. In this specification, several random retreat layers are set in the hidden layer network, and the random retreat layers are used to control the hidden units on the peripheral side to randomly retreat with a certain probability p in each round of training and not participate in the current round of model processing; for another example, taking the subnetwork as a convolutional network as an example, the relevant convolutional network usually only contains convolutional layers. In this specification, several random retreat layers are set in the convolutional network, and the random retreat layers are used to control the hidden units on the peripheral side to randomly retreat with a certain probability p in each round of training and not participate in the current round of model processing;

[0080] Indicative, such as Figure 4 As shown, Figure 4 It is a schematic diagram of a random retreat process, such as Figure 4 As shown, a random retreat layer can be set in at least one sub-network in a common initial neural network model, and the sub-network consists of several sub-network layers ( Figure 4 The network layer is shown in Figure 2) and several random retreat layers. In each training batch (equivalent to one round of training), during the forward propagation, a certain proportion of neurons (such as Figure 4 The neural network with random retreat processing shown in the figure will retain its current weights during the model training process, and set the outputs of the neurons that have undergone random retreat processing to 0. These selected neurons will restore the previously retained values ​​during the next training process, and randomly select some neurons again to repeat this process, that is, continue to repeat the aforementioned random retreat process until the model training is completed.

[0081] In a specific implementation scenario, several random retreat layers can be set in the hidden layer network in the initial neural network model, such as Figure 5 As shown, Figure 5 This is a schematic diagram of another random retreat processing scenario. In the initial neural network model after the settings are completed: the hidden layer network includes at least one random retreat layer and at least one hidden layer;

[0082] It can be understood that the random retreat layer is used to perform a first random retreat process on at least one first neuron in the hidden layer of the lower layer; Figure 4 The first neuron shown is Figure 5 The neurons in the current training batch shown have been subjected to random backoff processing; based on the initial neural network model after the first random backoff processing, the neurons other than the first neuron in the initial neural network model output first representation information for the first training data;

[0083] S208: Perform a second random falloff process on at least one second neuron in the initial neural network model, and output second representation information for the second training data based on the initial neural network model after the second random falloff process.

[0084] The difference between the second random retirement processing and the aforementioned first random retirement processing is that the first training data and the second training data are usually not processed in the same training batch, and the processing object neurons based on the random retirement processing in different training batches are usually different. The processing object of the first random retirement processing is the first neuron, and the processing object of the second random retirement processing is the second neuron.

[0085] The second neuron is the subject of a second retreat process. During model processing of at least one training data item included in the second training data using the initial neural network model, the at least one second neuron is controlled to be "shielded" and not participate in the current round of model processing of the second training data item. In the next round, the aforementioned random retreat process is repeated until model training is completed.

[0086] It can be understood that after performing a second random retreat process on at least one second neuron in the initial neural network model, second representation information for the second training data will be output based on neurons other than all second neurons in the initial neural network model.

[0087] In one feasible embodiment, at least one random retreat layer is provided in the initial neural network model. Based on the at least one random retreat layer, neurons surrounding the initial neural network model (e.g., neurons in the lower i layer, neurons in the upper j layer, etc.) are subjected to random retreat processing. The random retreat layer includes a plurality of random retreat units (which can be understood as a type of neuron).

[0088] It can be understood that the random retreat layer is used to perform a second random retreat process on at least one second neuron in the hidden layer of the lower layer; Figure 4 The second neuron shown is Figure 5 The neurons in the current training batch shown have been subjected to random backoff processing; based on the initial neural network model after the second random backoff processing, the neurons other than the second neuron in the initial neural network model output second representation information for the second training data;

[0089] It can be understood that the random retirement process will calculate the probability p of the relevant neural network (cell) and determine with probability p: in the current training batch, whether a certain neuron will be randomly retired during the processing of the current batch of training data in the training data. Therefore, the representation information generated by the network for the same batch of training data input in different training batches will be frozen by the corresponding neurons due to the random retirement process, making the representation information output by the model different each time for the same input batch of training data. The goal of model training is to maximize the similarity between two different representations output for the same training data, that is, to encourage the model to make full use of the feature relationship corresponding to the two different representations, and back-propagate to optimize the model parameters and model architecture, thereby enhancing the network's ability to extract user representations, and obtaining a neural network model with good recognition ability.

[0090] S210: Determine a model loss based on the first representation information and the second representation information, and perform model training on the initial neural network model based on the model loss to obtain a target neural network model.

[0091] It is understandable that in the training of a neural network model, label representation information will be set corresponding to the labeled training data. The label representation information can be understood as the expected result or expected output value of the model output.

[0092] In one or more embodiments of the present specification, the first training data may include unlabeled training data and labeled training data; and since the second training data is obtained by copying and processing the first training data, the second training data may include unlabeled training data and labeled training data.

[0093] It is understandable that the first training data / second training data usually contains a large amount of unlabeled data.

[0094] In a specific implementation scenario: there is at least the same unlabeled data in the first training data / the second training data; by executing the model processing method of this specification, self-supervised learning of the model of unlabeled data can be achieved, avoiding the dependence of the model self-supervised learning in the related art on the labeling of the training data. A quantification of the model loss of unlabeled data is achieved. The goal of model training in the specific implementation is to maximize the similarity between the two different representations output for the same training data, that is, to encourage the model to fully utilize the feature relationship corresponding to the two different representations, and back-propagate to optimize the model parameters and model architecture, thereby enhancing the network's ability to extract user representations, and a neural network model with good recognition ability can be obtained. The quantification of the model loss of unlabeled data is also based on the maximization of the similarity between the two different representations.

[0095] Optionally, third representation information for the first unlabeled data and fourth representation information for the second unlabeled data can be obtained from the first representation information, where the second unlabeled data is unlabeled data generated by copying the first labeled data. In a specific implementation, the source input training data corresponding to the first representation information is traversed to obtain the first unlabeled data whose source input training data is unlabeled data, thereby obtaining the third representation information corresponding to the first unlabeled data from the first representation information. Similarly, the source input training data corresponding to the second representation information is traversed to obtain the second unlabeled data whose source input training data is unlabeled data, thereby obtaining the fourth representation information corresponding to the second unlabeled data from the second representation information.

[0096] It can be understood that the third representation information and the fourth representation information are usually model outputs corresponding to the same training data in different training batches. For two different representations output for the same training data, the similarity between the two different representations is maximized. Specifically, the similarity is maximized by calculating the first loss. Generally, the smaller the loss, the higher the similarity between the two different representations.

[0097] In a specific implementation, based on the third representation information and the fourth representation information, a first loss for unlabeled data is determined;

[0098] In a feasible implementation, a first loss calculation formula for quantifying the model loss for unlabeled data is pre-set.

[0099] Determining a first loss for unlabeled data by inputting the third representation information and the fourth representation information into a first loss calculation formula;

[0100] The first loss calculation formula satisfies the following formula:

[0101]

[0102] Among them, L1 is the first loss, A i is the third representation information, B i is the fourth characterization information, kl() is the relative entropy function, and n is the number of data of the unlabeled data.

[0103] Understandable, A. i is the representation information corresponding to the i-th unlabeled data in the first training data. In some embodiments, the representation can be understood as a representation vector output by a model for the i-th unlabeled data. The representation vector is softmax processed to obtain the representation probability distribution, which is the A mentioned above. i” or B iBased on this, after obtaining different probability distributions of the i-th unlabeled data, the relative entropy is calculated based on the relative entropy function. The representation information can be understood as the representation vector and / or the representation probability distribution corresponding to the representation vector.

[0104] In one or more embodiments of the present specification, based on the hidden layer and the random retreat layer, as well as the first loss, the first loss can be used to encourage two different representations to be close, that is, to encourage the model to fully utilize the relationship between features and enhance the network's ability to extract user representations; at the same time, because the first loss is for unsupervised data, a method of introducing unlabeled training data into model training is implemented. In normal self-supervised learning, only labeled training data can be used, but here, through the introduction of the aforementioned method, unlabeled training data can be introduced, which expands the amount of data that can be used during model training, achieves rapid convergence of the model, and improves model processing efficiency. The entire model processing process can reduce dependence on labeled data.

[0105] In a feasible embodiment, the parameters of the initial neural network model can be adjusted directly based on the first loss during the backpropagation process, wherein the parameters include the weight values ​​and / or thresholds of each layer of neurons until the training end conditions preset for the initial neural network model are met. The training end conditions may include, for example, the value of the loss function is less than or equal to the preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The specific training end conditions can be determined based on the actual business scenario and will not be elaborated here.

[0106] In a specific implementation scenario, if the first training data / the second training data includes labeled training data, the labeled training data corresponds to the label representation information. The model loss calculation of the labeled data can be implemented based on the label representation information and the labeled training data.

[0107] For example, the initial neural network model includes an input network, a (neural) network, and a representation network. The (neural) network may include the above-mentioned random retreat layer.

[0108] During the forward propagation process, the first training data / second training data is input from the input network of the initial neural network model, passes through the transfer function (also known as activation function, conversion function) of the corresponding neurons (also known as nodes) (such as neurons in a (neural) network), and is then transmitted to the representation network. The state of each layer of neurons affects the state of the neurons in the next layer. The actual output value - the first representation information / the second representation information is calculated in the representation network. If there is labeled training data for the first training data / the second training data, the labeled training data corresponds to the labeled representation information. The actual output result of the corresponding labeled training data - the fifth representation information is obtained from the first representation information / the second representation information. The second loss for the labeled data is calculated based on the label representation information and the fifth representation information. In some embodiments, the second loss can also be understood as the expected loss between the actual output value of the model and the expected output value.

[0109] Further: obtaining fifth representation information for the label data from the first representation information and the second representation information to determine the label representation information corresponding to the label data; performing self-supervised learning based on the fifth representation information and the label representation information to determine a second loss for the label data; adjusting the parameters of the initial neural network model during the back propagation process based on the second loss, the parameters including the weight value and / or threshold of each layer of neurons

[0110] In a feasible implementation manner, the fifth representation information and the label representation information may be input into a second loss calculation formula to determine a second loss for the label data;

[0111] The second loss calculation formula satisfies the following formula:

[0112]

[0113] Among them, L2 is the first loss, C i is the third representation information, D i is the fourth representation information, ce() is the cross entropy function, and k is the number of data of the label data.

[0114] In a feasible embodiment, the parameters of the initial neural network model can be adjusted directly based on the second loss during the backpropagation process, wherein the parameters include the weight values ​​and / or thresholds of each layer of neurons until the training end conditions preset for the initial neural network model are met. The training end conditions may include, for example, the value of the loss function is less than or equal to the preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The specific training end conditions can be determined based on the actual business scenario and will not be elaborated here.

[0115] In a feasible implementation, the initial neural network model can be trained based on the first loss and the second loss until a preset training termination condition for the initial neural network model is met. The training termination condition may include, for example, the value of the total loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The specific training termination condition can be determined based on the actual transaction scenario and will not be further described here.

[0116] For example, Figure 6 As shown, Figure 6 This is a schematic diagram of a model training scenario involved in this specification. In the data input stage, the first training data "data i1, data i2...data in" is copied to obtain the second training data "data j1, data j2...data jn". The first training data and the second training data are input into the initial neural network model together, and in each batch of training, the first characterization information and the second characterization information output by the model are represented in the representation network. Then, the first loss and the second loss are calculated, and the total model loss for the initial neural network model is calculated based on the first loss and the second loss. The initial neural network model is trained based on the total model loss. That is, the parameters of the initial neural network model are adjusted during the backpropagation process based on the first loss and the second loss. The parameters include the weight values ​​and / or thresholds of each layer of neurons.

[0117] Indicatively, the first loss and the second loss can be weighted to obtain the total loss of the model;

[0118] For example, the first loss and the second loss can be input into the third loss calculation formula to calculate the total loss of the model as follows:

[0119] L=L1+L2*a

[0120] Among them, L is the total loss of the model, and a is the coefficient;

[0121] In some embodiments, the coefficient a may be determined based on a specific transaction scenario, for example, the coefficient a may be between 0.001 and 0.01.

[0122] In one or more embodiments of the present specification, by performing data copying processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved; the target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched; and based on random retreat processing, the disturbance of model noise is enhanced when the training data contains online transaction data, thereby realizing effective real-time data learning of the model, while expanding the model learning range and effectively enhancing the stability of the model after it is launched.

[0123] See Figure 7 , Figure 7 This is a flow chart of another embodiment of a model processing method proposed in one or more embodiments of this specification. Specifically:

[0124] S302: Obtain first training batch data of a current round in first training data for the initial neural network model, perform data copying on the first training batch data, and obtain second training batch data of a current round in second training data;

[0125] It can be understood that model training usually corresponds to multiple rounds of training batch data. The first training data can be understood as the training data set for the initial neural network model in model training. The first training data can be divided into several first training batch data according to the model training batch or the number of model training rounds; the second training data is the copy data of the first training data in model training. The second training data can be divided into several second training batch data according to the model training batch or the number of model training rounds;

[0126] Start model training from the first training batch of data in round 1;

[0127] In the first round of model training: obtain the first training batch data of the current first round, perform data replication processing on the first training batch data, and obtain the second training batch data of the first round in the second training data; then execute S304;

[0128] If the model end training condition for the initial neural network model is not met, the next round of model training is carried out;

[0129] In the second round of model training: obtain the first training batch data of the current second round, perform data replication processing on the first training batch data, and obtain the second training batch data of the second round in the second training data; then execute S304;

[0130] If the model end training condition for the initial neural network model is not met, the next round of model training is carried out; ....

[0132] In the i-th round of model training: obtaining the first training batch data of the current i-th round, performing data replication processing on the first training batch data to obtain the second training batch data of the i-th round in the second training data; then executing S304;

[0133] Until the model end training conditions for the initial neural network model are met, the model training is ended and the target neural network model is obtained.

[0134] S304: Input the first training batch data and the second training batch data into the initial neural network model to determine first representation information for the first training batch data and second representation information for the second training batch data.

[0135] It can be understood that the first training batch data is the sub-data of the first training data during the model training process, and the second training batch data is the sub-data of the second training data during the model training process. Specifically, "inputting the first training batch data and the second training batch data into the initial neural network model to determine the first representation information for the first training batch data and the second representation information for the second training batch data" can be referred to other method embodiments of this specification and will not be repeated here.

[0136] S306: Determine the first loss and / or the second loss for the current round of the initial neural network model based on the first representation information and the second representation information;

[0137] It can be understood that reference can be made to other method embodiments in this specification, which will not be repeated here.

[0138] S312: Determine whether model training for the initial neural network model is completed;

[0139] It is understandable that the training end condition is set in advance to determine whether to end the model training for the initial neural network model;

[0140] The training end conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. The specific training end conditions can be determined based on the actual transaction scenario and will not be detailed here.

[0141] S314: If yes, the initial neural network model is used as the target neural network model;

[0142] S316: If not, obtaining the first loss and / or the second loss of the current round for the initial neural network model;

[0143] It can be understood that reference can be made to other method embodiments in this specification, which will not be repeated here.

[0144] S318: Based on the first loss and / or the second loss, perform copy adjustment processing on the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data; and execute S304.

[0145] In one or more embodiments of the present specification, the data replication process is performed on the first training data to obtain the second training data. Instead of using a full replication method, the replication can be intelligently adjusted based on the model loss. That is, instead of replicating each data file in the first training data, the first training batch data of the next round in the first training data is adjusted by the data replication method to obtain the second training batch data of the next round in the second training data.

[0146] In one feasible embodiment, a data replication ratio is determined based on the first loss and / or the second loss of the current round. The data replication ratio is determined by setting a ratio parameter, and a corresponding ratio of data files in the first training batch data of the next round is replicated in the first training data to obtain the second training batch data. For example, if the data replication ratio of the current round is 80%, the data replication ratio is adjusted based on the first loss and / or the second loss, such that the data replication ratio of the next round is 70%. The first training batch data of the next round in the first training data is replicated based on the data replication ratio, to obtain the second training batch data of the next round in the second training data.

[0147] Optionally, a proportional mapping relationship between several reference loss ranges and their corresponding reference data ratios can be set. The proportional mapping relationship can be represented in the form of a mapping set, a mapping array, a mapping list, etc. In practice, the target reference loss range within which the model loss falls can be determined based on the aforementioned proportional mapping relationship, and the target data ratio corresponding to the target reference loss range can be obtained. The replication adjustment processing is performed based on the target data ratio, that is, data replication is performed at the target data ratio.

[0148] In a feasible implementation, multiple data replication modes can be set, with different data replication modes corresponding to different data replication amounts, such as data replication mode A for replicating data files every other time, data replication mode B for replicating data files every other time, and so on. The data replication mode for the next round can then be selected based on the model loss (e.g., the first loss, the second loss, or the total loss corresponding to the first and second losses).

[0149] It can be understood that by dynamically adjusting the data replication method of the next round based on the model loss of each round (the first loss and / or the second loss), it is possible to dynamically adjust the amount of data input to the model in each round based on the current model training effect, so as to achieve a better model training effect in the current round, save model training resource consumption compared to full replication, and improve the efficiency of model processing.

[0150] In one or more embodiments of the present specification, by performing data copying processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved; the target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched; and based on random retreat processing, the disturbance of model noise is enhanced when the training data contains online transaction data, thereby realizing effective real-time data learning of the model, while expanding the model learning range and effectively enhancing the stability of the model after it is launched.

[0151] The following will be combined Figure 8 , a detailed introduction to the model processing device provided in one or more embodiments of this specification is given. It should be noted that, Figure 4 The model processing device shown is used to execute the application Figures 2 and 3 For the convenience of explanation, only the parts related to this specification are shown in the method of the embodiment shown. For specific technical details not disclosed, please refer to this application. Figure 2 Figure 7 The embodiment shown.

[0152] See Figure 8, which shows a schematic diagram of the structure of a model processing device in one or more embodiments of this specification. The model processing device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the model processing device 1 includes a data replication module 11, a model processing module 12, and a model training module 13, which are specifically used to:

[0153] A data replication module 11 is used to perform data replication processing on the first training data for the initial neural network model to obtain second training data;

[0154] a model processing module 12, configured to input the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data;

[0155] The model training module 13 is used to perform model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model.

[0156] Optional, such as Figure 9 As shown, the model processing module 12 includes:

[0157] A data input unit 121 is configured to input the first training data and the second training data into the initial neural network model;

[0158] The information determination unit 122 is used to determine first representation information for the first training data and second representation information for the second training data based on the initial neural network model after random retirement processing by performing random retirement processing on the initial neural network model.

[0159] Optional, such as Figure 10 As shown, the information determination unit 122 includes:

[0160] A first determining subunit 1221 is configured to perform a first random backoff process on at least one first neuron in the initial neural network model, and output first representation information for first training data based on the initial neural network model after the first random backoff process;

[0161] The second determining subunit 1222 is configured to perform a second random backoff process on at least one second neuron in the initial neural network model, and output second representation information for the second training data based on the initial neural network model after the second random backoff process.

[0162] Optionally, the initial neural network model includes a hidden layer network, the hidden layer network includes at least one random retreat layer and at least one hidden layer, and the first determining subunit 1221 is specifically configured to:

[0163] Performing a first random retreat process on at least one first neuron in the hidden layer of the lower layer through the random retreat layer;

[0164] The second determining subunit 1222 is specifically configured to:

[0165] A second random retreat process is performed on at least one second neuron in the hidden layer of the lower layer through the random retreat layer.

[0166] Optional, such as Figure 11 As shown, the model training module 13 includes:

[0167] a loss determining unit 131, configured to determine a model loss based on the first representation information and the second representation information;

[0168] The model training unit 132 is used to perform model training on the initial neural network model based on the model loss to obtain a target neural network model.

[0169] Optionally, the loss determining unit 131 is specifically configured to:

[0170] Obtaining third representation information for first unlabeled data from the first representation information, and obtaining fourth representation information for second unlabeled data, where the second unlabeled data is unlabeled data generated by copying the first labeled data;

[0171] Determining a first loss for unlabeled data based on the third representation information and the fourth representation information;

[0172] The model training unit 132 is specifically used to perform model training on the initial neural network model based on the first loss.

[0173] Optionally, the loss determining unit 131 is specifically configured to:

[0174] Inputting the third representation information and the fourth representation information into a first loss calculation formula to determine a first loss for unlabeled data;

[0175] The first loss calculation formula satisfies the following formula:

[0176]

[0177] Among them, L1 is the first loss, A i is the third representation information, Bi is the fourth characterization information, kl() is the relative entropy function, and n is the number of data of the unlabeled data.

[0178] Optionally, the loss determining unit 131 is specifically configured to:

[0179] Acquire fifth representation information for the tag data from the first representation information and the second representation information, and determine tag representation information corresponding to the tag data;

[0180] Determining a second loss for the tag data based on the fifth representation information and the tag representation information;

[0181] The performing model training on the initial neural network model based on the model loss includes:

[0182] The initial neural network model is trained based on the first loss and the second loss.

[0183] Optionally, the loss determining unit 131 is specifically configured to:

[0184] Inputting the fifth representation information and the label representation information into a second loss calculation formula to determine a second loss for the label data;

[0185] The second loss calculation formula satisfies the following formula:

[0186]

[0187] Among them, L2 is the first loss, C i is the third representation information, D i is the fourth representation information, ce() is the cross entropy function, and k is the number of data of the label data.

[0188] Optionally, the device 1 is further used to: obtain the first loss and / or the second loss for the current round of the initial neural network model;

[0189] Based on the first loss and / or the second loss, copy and adjust the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data;

[0190] Optionally, the model training module 13 is further used to: input the first training batch data and the second training batch data into the initial neural network model to determine first representation information for the first training batch data and second representation information for the second training batch data.

[0191] Optionally, the model training module 13 is further configured to: determine a data replication ratio based on the first loss and / or the second loss;

[0192] Based on the data replication ratio, data replication processing is performed on the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data.

[0193] Optional, such as Figure 12 As shown, the device 1 further includes:

[0194] The data generation module 14 is used to obtain labeled training data for the target transaction scenario and obtain unlabeled online transaction data in the target transaction scenario;

[0195] The data generating module 14 is further configured to generate first training data based on the labeled training data and the online transaction data.

[0196] Optional, such as Figure 12 As shown, the device 1 further includes:

[0197] The model application module 15 inputs the user transaction data into the target neural network model and outputs the transaction conversion probability for the user transaction data.

[0198] It should be noted that the model processing device provided in the above embodiment, when executing the model processing method, only uses the division of the above-mentioned functional modules as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model processing device provided in the above embodiment and the model processing method embodiment are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0199] The above serial numbers in this specification are for description only and do not represent the advantages or disadvantages of the embodiments.

[0200] In one or more embodiments of the present specification, by performing data copy processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved; the target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched; and based on random retreat processing, the disturbance of model noise is enhanced when the training data contains online transaction data, thereby realizing effective real-time data learning of the model, while expanding the model learning range and effectively enhancing the stability of the model after it is launched.

[0201] An embodiment of this specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the model processing method as described in one or more of the above embodiments. The specific execution process can be found in the specific description of one or more of the above embodiments, and will not be repeated here.

[0202] An embodiment of this specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor and executes the model processing method as described in one or more of the above-mentioned embodiments. The specific execution process can be found in the specific description of one or more of the above-mentioned embodiments, and will not be repeated here.

[0203] See Figure 13 , provides a structural diagram of an electronic device for one or more embodiments of this specification. Figure 13 As shown, the electronic device 1000 may include: at least one processor 1001 , at least one network interface 1004 , a user interface 1003 , a memory 1005 , and at least one communication bus 1002 .

[0204] The communication bus 1002 is used to implement the connection and communication between these components.

[0205] The user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.

[0206] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0207] The processor 1001 may include one or more processing cores. The processor 1001 utilizes various interfaces and circuits to connect various components within the server 1000. It executes instructions, programs, code sets, or instruction sets stored in the memory 1005, and accesses data stored in the memory 1005 to perform various server 1000 functions and process data. Optionally, the processor 1001 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 1001 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 1001 and may be implemented as a separate chip.

[0208] Among them, the memory 1005 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Figure 13 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program.

[0209] exist Figure 13 In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 1001 can be used to call the application stored in the memory 1005, and specifically perform the following operations: copy the first training data for the initial neural network model to obtain second training data; input the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data; perform model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model.

[0210] In one embodiment, when the processor 1001 inputs the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data, it specifically performs the following steps: inputs the first training data and the second training data into the initial neural network model; performs random retirement processing on the initial neural network model, and determines the first representation information for the first training data and the second representation information for the second training data based on the initial neural network model after the random retirement processing.

[0211] In one embodiment, when the processor 1001 performs the random retirement processing on the initial neural network model and determines the first representation information for the first training data and the second representation information for the second training data based on the initial neural network model after the random retirement processing, it specifically performs the following steps: performing a first random retirement processing on at least one first neuron in the initial neural network model, and outputting the first representation information for the first training data based on the initial neural network model after the first random retirement processing; performing a second random retirement processing on at least one second neuron in the initial neural network model, and outputting the second representation information for the second training data based on the initial neural network model after the second random retirement processing.

[0212] In one embodiment, the processor 1001 executes the initial neural network model including a hidden layer network, the hidden layer network including at least one random retreat layer and at least one hidden layer, and the performing of the first random retreat process on at least one first neuron in the initial neural network model includes: performing the first random retreat process on at least one first neuron in the hidden layer of a lower layer through the random retreat layer;

[0213] In one embodiment, the processor 1001 performs the second random retreat processing on at least one second neuron in the initial neural network model, including: performing the second random retreat processing on at least one second neuron in the hidden layer of the lower layer through the random retreat layer.

[0214] In one embodiment, the processor 1001 performs model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model, including: determining a model loss based on the first representation information and the second representation information; and performing model training on the initial neural network model based on the model loss to obtain a target neural network model.

[0215] In one embodiment, when the processor 1001 determines the model loss based on the first representation information and the second representation information, it specifically performs the following steps: obtaining third representation information for first unlabeled data from the first representation information, and obtaining fourth representation information for second unlabeled data, where the second unlabeled data is unlabeled data generated by copying data from the first labeled data;

[0216] Determining a first loss for unlabeled data based on the third representation information and the fourth representation information;

[0217] The performing model training on the initial neural network model based on the model loss includes: performing model training on the initial neural network model based on the first loss.

[0218] In one embodiment, when the processor 1001 determines the first loss for the unlabeled data based on the third representation information and the fourth representation information, the processor 1001 specifically performs the following steps: inputting the third representation information and the fourth representation information into a first loss calculation formula to determine the first loss for the unlabeled data;

[0219] The first loss calculation formula satisfies the following formula:

[0220]

[0221] Among them, L1 is the first loss, A i is the third representation information, B i is the fourth characterization information, kl() is the relative entropy function, and n is the number of data of the unlabeled data.

[0222] In one embodiment, when executing the model processing method, the processor 1001 further performs the following steps:

[0223] Acquire fifth representation information for the tag data from the first representation information and the second representation information, and determine tag representation information corresponding to the tag data;

[0224] Determining a second loss for the tag data based on the fifth representation information and the tag representation information;

[0225] When the initial neural network model is trained based on the model loss, the following steps are specifically performed:

[0226] The initial neural network model is trained based on the first loss and the second loss.

[0227] In one embodiment, when the processor 1001 determines the second loss for the tag data based on the fifth representation information and the tag representation information, the processor 1001 specifically performs the following steps:

[0228] Inputting the fifth representation information and the label representation information into a second loss calculation formula to determine a second loss for the label data;

[0229] The second loss calculation formula satisfies the following formula:

[0230]

[0231] Among them, L2 is the first loss, C i is the third representation information, D i is the fourth representation information, ce() is the cross entropy function, and k is the number of data of the label data.

[0232] In one embodiment, when executing the model processing method, the processor 1001 further performs the following steps:

[0233] Obtaining the first loss and / or the second loss of a current round for the initial neural network model;

[0234] Based on the first loss and / or the second loss, copy and adjust the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data;

[0235] When the first training data and the second training data are input into the initial neural network model to determine the first representation information for the first training data and the second representation information for the second training data, the following steps are specifically performed:

[0236] The first training batch data and the second training batch data are input into the initial neural network model to determine first representation information for the first training batch data and second representation information for the second training batch data.

[0237] In one embodiment, the processor 1001 performs the copy adjustment processing on the first training batch data of the next round in the first training data based on the first loss and / or the second loss to obtain the second training batch data of the next round in the second training data: determines a data copy ratio based on the first loss and / or the second loss; and performs data copy processing on the first training batch data of the next round in the first training data based on the data copy ratio to obtain the second training batch data of the next round in the second training data.

[0238] In one embodiment, before executing the data copy processing for the first training data of the initial neural network model, the processor 1001 further performs the following steps: obtaining labeled training data for the target transaction scenario, obtaining unlabeled online transaction data in the target transaction scenario; and generating first training data based on the labeled training data and the online transaction data.

[0239] In one embodiment, after executing the step of obtaining the target neural network model, the processor 1001 further includes:

[0240] The user transaction data is input into the target neural network model, and the transaction conversion probability for the user transaction data is output.

[0241] In one or more embodiments of the present specification, by performing data copying processing on the first training data of the initial neural network model to obtain second training data, and inputting the first training data and the second training data into the initial neural network model for model training to determine first representation information for the first training data and second representation information for the second training data, and then performing model training on the initial neural network model based on the first representation information and the second representation information, a trained target neural network model can be obtained. By outputting different representation information based on the same training data in different model data processing stages, a new model training processing method is implemented, the model processing flow is optimized, and the model processing effect is improved; the target neural network model has good robustness against changes in online training data (such as user online transaction data), which can improve the prediction effect after the model is launched; and based on random retreat processing, the disturbance of model noise is enhanced when the training data contains online transaction data, thereby realizing effective real-time data learning of the model, while expanding the model learning range and effectively enhancing the stability of the model after it is launched.

[0242] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0243] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A model processing method, comprising: Obtaining labeled training data for a target transaction scenario, obtaining unlabeled online transaction data for the target transaction scenario, and generating first training data based on the labeled training data and the online transaction data, wherein the target transaction scenario is a transaction scenario for recommending content to a client user; Performing data replication processing on the first training data of the initial neural network model to obtain second training data; Inputting the first training data and the second training data into the initial neural network model, performing random retreat processing on the initial neural network model, and determining first representation information for the first training data and second representation information for the second training data based on the initial neural network model after the random retreat processing; Performing model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model, and performing content recommendation processing on the client based on the target neural network model; Among them, the random retirement processing is to shield at least one neuron that is retired in each round of the initial neural network model. The neurons that have undergone retirement processing do not participate in the model processing process of this round, and the neurons in the initial neural network model that have not undergone random retirement processing except the neurons that have undergone retirement processing participate in the processing of training data.

2. The method according to claim 1, wherein the determining, based on the initial neural network model after random retirement processing and first characterization information for the first training data and second characterization information for the second training data, comprises: performing a first random backoff process on at least one first neuron in the initial neural network model, and outputting first representation information for first training data based on the initial neural network model after the first random backoff process; A second random falloff process is performed on at least one second neuron in the initial neural network model, and second representation information for second training data is output based on the initial neural network model after the second random falloff process.

3. The method according to claim 2, wherein the initial neural network model comprises a hidden layer network, wherein the hidden layer network comprises at least one random retreat layer and at least one hidden layer, The performing a first random retreat process on at least one first neuron in the initial neural network model includes: Performing a first random retreat process on at least one first neuron in the hidden layer of the lower layer through the random retreat layer; The performing a second random retreat process on at least one second neuron in the initial neural network model comprises: A second random retreat process is performed on at least one second neuron in the hidden layer of the lower layer through the random retreat layer.

4. The method according to claim 1, wherein the training of the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model comprises: Determining a model loss based on the first representation information and the second representation information; The initial neural network model is trained based on the model loss to obtain a target neural network model.

5. The method according to claim 4, wherein determining the model loss based on the first representation information and the second representation information comprises: Obtaining third representation information for first unlabeled data from the first representation information, and obtaining fourth representation information for second unlabeled data, where the second unlabeled data is unlabeled data generated by copying the first labeled data; Determining a first loss for unlabeled data based on the third representation information and the fourth representation information; The performing model training on the initial neural network model based on the model loss includes: The initial neural network model is trained based on the first loss.

6. The method according to claim 5, wherein determining the first loss for unlabeled data based on the third representation information and the fourth representation information comprises: Inputting the third representation information and the fourth representation information into a first loss calculation formula to determine a first loss for unlabeled data; The first loss calculation formula satisfies the following formula: Among them, L1 is the first loss, A i is the third representation information, B i is the fourth characterization information, kl() is the relative entropy function, and n is the number of data of the unlabeled data.

7. The method according to claim 5, further comprising: Acquire fifth representation information for the tag data from the first representation information and the second representation information, and determine tag representation information corresponding to the tag data; Determining a second loss for the tag data based on the fifth representation information and the tag representation information; The performing model training on the initial neural network model based on the model loss includes: The initial neural network model is trained based on the first loss and the second loss.

8. The method according to claim 7, wherein determining the second loss for the tag data based on the fifth representation information and the tag representation information comprises: Inputting the fifth representation information and the label representation information into a second loss calculation formula to determine a second loss for the label data; The second loss calculation formula satisfies the following formula: Among them, L2 is the first loss, C i is the third representation information, D i is the fourth representation information, ce() is the cross entropy function, and k is the number of data of the label data.

9. The method according to claim 7, further comprising: Obtaining the first loss and / or the second loss of a current round for the initial neural network model; Based on the first loss and / or the second loss, copy and adjust the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data; Inputting the first training data and the second training data into the initial neural network model to determine first representation information for the first training data and second representation information for the second training data includes: The first training batch data and the second training batch data are input into the initial neural network model to determine first representation information for the first training batch data and second representation information for the second training batch data.

10. The method according to claim 9, wherein based on the first loss and / or the second loss, a first training batch data of the next round in the first training data is subjected to copy adjustment processing to obtain a second training batch data of the next round in the second training data: determining a data replication ratio based on the first loss and / or the second loss; Based on the data replication ratio, data replication processing is performed on the first training batch data of the next round in the first training data to obtain the second training batch data of the next round in the second training data.

11. The method according to any one of claims 1 to 10, further comprising: after obtaining the target neural network model; The user transaction data is input into the target neural network model, and the transaction conversion probability for the user transaction data is output.

12. A model processing device, comprising: The apparatus is configured to obtain labeled training data for a target transaction scenario, obtain unlabeled online transaction data for the target transaction scenario, and generate first training data based on the labeled training data and the online transaction data, wherein the target transaction scenario is a transaction scenario for recommending content to a client user. A data replication module is used to perform data replication processing on the first training data for the initial neural network model to obtain second training data; a model processing module, configured to input the first training data and the second training data into the initial neural network model, perform random retirement processing on the initial neural network model, and determine first representation information for the first training data and second representation information for the second training data based on the initial neural network model after the random retirement processing; a model training module, configured to perform model training on the initial neural network model based on the first representation information and the second representation information to obtain a target neural network model, and perform content recommendation processing to the client based on the target neural network model; Among them, the random retirement processing is to control the shielding processing of at least one neuron that has been retired in the initial neural network model in each round of model training. The neurons that have been retired do not participate in the current round of model processing process, and the neurons in the initial neural network model that have not been randomly retired except for the neurons that have been retired participate in the processing of training data.

13. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 11.

14. A computer program product, wherein the computer program product stores at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method steps according to any one of claims 1 to 11.

15. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Vaginal secretion wet sheet candida detection method based on Hough round detection and deep CNN (convolutional neural network)

    CN107099577A

  • Method for calculating shielding material accumulation factor based on BP neural network

    CN112733439A