Data processing method and apparatus
By establishing the mapping relationship between the identification of sample features and the gradient storage address during training the recommended model, the problem of low efficiency in searching for gradient storage address caused by hash conflicts is solved, and data processing and model training efficiency is improved.
Patent Information
- Application Number
- PCT/CN2024/118376
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-09-11
- Publication Date
- 2025-06-12
AI Technical Summary
During the training of the recommended model, the gradient storage address search efficiency caused by hash conflicts is low, resulting in low data processing efficiency.
By establishing the mapping relationship between the identification of sample features and the gradient storage address, the gradient storage address is quickly determined, and the gradient value is accumulated to the corresponding gradient storage unit, and the parameter value of the model parameters is adjusted.
The data processing efficiency in the training recommended model is improved, the need for gradient value accumulation calculation is reduced, and the model training efficiency is further improved.
Smart Images

Figure CN2024118376_12062025_PF_FP_ABST
Abstract
Description
Data processing method and device
[0001] This application claims priority to Chinese patent application filed on December 8, 2023, with application number 202311690834.5 and application name “Data Processing Method and Apparatus,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a data processing method and device. Background Art
[0003] The recommendation model is used to filter data content that users are interested in from massive data and recommend the filtered data content to users. The recommendation model can be obtained by training the server based on sample data. For example, for each sample data in the sample data set, the server extracts multiple sample features of the sample data, and the server calculates the gradient value corresponding to the model parameter corresponding to each sample feature in the multiple sample features and the hash value of each sample feature. The server establishes a mapping relationship between the hash value of each sample feature and the gradient storage address. Based on the mapping relationship, the server stores the gradient value corresponding to the model parameter corresponding to each sample feature to the gradient storage unit indicated by the corresponding gradient storage address. Based on the mapping relationship, the server obtains the gradient value stored at the gradient storage address corresponding to the hash value that identifies the same sample feature (that is, the gradient value stored in the gradient storage unit indicated by the gradient storage address), calculates the average value of the gradient value stored at the gradient storage address corresponding to the hash value that identifies the same sample feature, and adjusts the parameter value of the model parameter of the recommendation model based on the average value.
[0004] In the process of establishing the above-mentioned mapping relationship, in order to avoid hash conflicts, for sample features with the same hash value, the server obtains the hash value of the key code of these sample features, and the server establishes the mapping relationship based on the hash value of the key code of these sample features. When the server obtains the gradient value based on the mapping relationship, the server first determines the gradient storage address based on the hash value of the sample feature. For sample features with the same hash value, the server further determines the gradient storage address based on the hash value of the key code of these sample features, and then the server obtains the gradient value stored at the determined gradient storage address. Since for sample features with the same hash value, the server needs to further search the above-mentioned mapping relationship based on the hash value of the key code to determine the gradient storage address, this leads to low data processing efficiency during the training of the recommendation model.
[0005] Summary of the Invention
[0006] This application provides a data processing method and apparatus. The technical solution provided by this application can improve data processing efficiency during the training of recommendation models. The technical solution of this application is as follows.
[0007] In a first aspect, a data processing method is provided, the method comprising: determining a first gradient value corresponding to the first model parameter based on a loss value of first prediction data and an initial value of a first model parameter of a recommendation model, wherein the first prediction data is calculated by the recommendation model based on a first user feature and a first content feature, an identifier of a first sample feature in the first user feature and the first content feature is a first identifier, and the first model parameter corresponds to the first identifier; determining a first gradient storage address based on the first identifier (here, the identifier of the first sample feature) and a mapping relationship, the mapping relationship comprising a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address being used to indicate a first gradient storage unit, the first gradient storage unit being used to store an accumulated gradient value corresponding to the first model parameter; adding the first gradient value to the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; storing the first accumulated gradient value in the first gradient storage unit; and adjusting the parameter value of the first model parameter based on the accumulated gradient value stored in the first gradient storage unit (which may be the first accumulated gradient value, or may be an accumulated gradient value obtained by accumulating other gradient values corresponding to the first model parameter on the basis of the first accumulated gradient value).
[0008] The above mapping relationship may be a mapping relationship between sample feature identifiers and gradient storage addresses. For ease of description, the sample feature identifiers are referred to as feature identifiers or sample feature identifiers in some of the following descriptions. The identifier of the first sample feature is referred to as the first identifier, and the identifiers of other sample features may also be referred to as the first identifier. That is, the identifiers of different sample features may be the same or different.
[0009] Among them, the first user feature and the first content feature are both sample features of the first sample data. The loss value of the first prediction data refers to the loss of the first prediction data compared to the standard prediction data corresponding to the first sample data. The standard prediction data corresponding to the first sample data is the annotation value of the first sample data (or called annotation data). The first prediction data is the training prediction data corresponding to the first sample data, and the first prediction data is used to characterize the predicted correlation between the first user feature and the first content feature (that is, the predicted correlation), and the standard prediction data corresponding to the first sample data is used to characterize the standard correlation between the first user feature and the first content feature (that is, the labeled correlation). In the present application, the user features of different sample data may be the same or different, and the content features of different sample data may be the same or different.
[0010] The accumulated gradient value corresponding to the first model parameter is the accumulated value of the gradient values corresponding to the first model parameter. The first gradient storage unit is used to store the accumulated gradient value corresponding to the first model parameter, that is, the first gradient storage unit is used to store the accumulated gradient value corresponding to the model parameter corresponding to the sample feature identified as the first identifier. In this application, since the sample features identified as the first identifier all correspond to the first model parameter, this application is described as "the first gradient storage unit is used to store the accumulated gradient value corresponding to the first model parameter."
[0011] The data processing method of the present application can be executed by a single server or a server cluster consisting of multiple servers. The server can be a server for model training and for managing model parameters, for example, an artificial intelligence (AI) server, also known as a parameter server (PS).
[0012] The technical solution provided by this application is a mapping relationship between the identifier of the sample feature and the gradient storage address, and the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address. After the server determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the server can quickly determine the first gradient storage address according to the first identifier (here refers to the identifier of the first sample feature) and the mapping relationship, and then accumulate the first gradient value corresponding to the first model parameter to the accumulated gradient value stored at the first gradient storage address. When the server trains the recommendation model, the server can quickly determine the first gradient storage address according to the first identifier and the mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address, and adjust the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address, thereby improving the data processing efficiency during the model training process, thereby improving the model training efficiency. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform cumulative calculation of the gradient value when training the recommendation model, further improving the data processing efficiency and model training efficiency during the model training process. In this application, the first gradient storage address is used to indicate the first gradient storage unit. The accumulated gradient value stored at the first gradient storage address refers to the accumulated gradient value stored in the first gradient storage unit. For the sake of brevity, this application describes it as "the accumulated gradient value stored at the first gradient storage address." Other similar descriptions have the same meaning.
[0013] Optionally, determining the first gradient storage address based on the first identifier and the mapping relationship includes: determining the first index storage address based on the first identifier and the first mapping relationship, where the first mapping relationship is a mapping relationship between the identifier of the sample feature and the index storage address, the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, the first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; and determining the first gradient storage address based on the index information stored in the first index storage unit. Specifically, the server determines the first index storage unit based on the first index storage address, and then determines the first gradient storage address based on the index information stored in the first index storage unit.
[0014] The technical solution provided by the present application is that since the first mapping relationship is a mapping relationship between the identifier of the sample feature and the index storage address, and the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, after the server obtains the first gradient value corresponding to the first model parameter, the server can quickly determine the first index storage address based on the first identifier and the first mapping relationship, and then quickly determine the first gradient storage address based on the index information stored in the first index storage address, which can improve the efficiency of the server in determining the first gradient storage address, and thus improve the data processing efficiency during the model training process. In the present application, the first index storage address is used to indicate the first index storage unit, and the index information stored in the first index storage address refers to the accumulated gradient value stored in the first index storage unit. For the sake of simplicity, the present application describes it as "index information stored in the first index storage address". The meanings of other similar descriptions are similar.
[0015] Optionally, determining the first gradient storage address based on the index information stored in the first index storage unit includes: if the index information stored in the first index storage unit includes the gradient storage address, determining the gradient storage address stored in the first index storage unit as the first gradient storage address; if the index information stored in the first index storage unit does not include the gradient storage address, determining a first free storage unit in a pre-created first storage space as the first gradient storage unit, or creating a second storage space and determining a second free storage unit in the second storage space as the first gradient storage unit, and determining the address of the first gradient storage unit as the first gradient storage address. The first storage space and the second storage space each include at least one storage unit.
[0016] The technical solution provided by the present application is that, when the index information stored in the first index storage unit includes a gradient storage address, the server directly determines the gradient storage address stored in the first index storage unit as the first gradient storage address, so that the first gradient storage address can be quickly determined, which can improve the efficiency of the server in determining the first gradient storage address. In the case where the index information stored in the first index storage unit does not include a gradient storage address, the server can determine the first gradient storage address based on an idle storage unit. Therefore, regardless of whether the index information stored in the first index storage unit includes a gradient storage address, the server can determine the first gradient storage address for storing the accumulated gradient value corresponding to the model parameter corresponding to the sample feature of the first identifier, which can improve the flexibility of the server in determining the gradient storage address.
[0017] Optionally, the method further includes: the index information stored in the first index storage unit does not include the gradient storage address, and after the address of the first gradient storage unit is determined as the first gradient storage address, storing the first gradient storage address in the first index storage unit.
[0018] The technical solution provided by the present application is that, since the index information stored by the server in the first index storage unit does not have a gradient storage address, the server stores the first gradient storage address in the first index storage unit after determining the first gradient storage address. This makes it convenient for the server to quickly determine the first gradient storage address based on the index information stored in the first index storage unit.
[0019] Optionally, the index information stored in the first index storage unit includes status information, where the status information is used to indicate whether the storage state corresponding to the first identifier is an initial state or an accumulated state. The method further includes: when the status information indicates that the storage state corresponding to the first identifier is an initial state, determining that the index information stored in the first index storage unit does not include a gradient storage address; and when the status information indicates that the storage state corresponding to the first identifier is an accumulated state, determining that the index information stored in the first index storage unit includes a gradient storage address. For example, the first index storage unit includes a status field for recording status information and an address field for recording the gradient storage address, the address field being located after the status field, and the length of the status field being less than the length of the address field.
[0020] The technical solution provided by the present application is that, since the index information stored in the first index storage unit includes status information, the server can quickly determine whether the index information stored in the first index storage unit includes a gradient storage address based on the status information, thereby improving the efficiency of the server in determining whether the index information stored in the first index storage unit includes a gradient storage address.
[0021] Optionally, the method further includes: updating the status information after storing the first gradient storage address in the first index storage unit, when the status information indicates that the storage state corresponding to the first identifier is an initial state. The updated status information indicates that the storage state corresponding to the first identifier is an accumulation state. The server updating the status information can facilitate the server's subsequent rapid determination, based on the status information, that the index information stored in the first index storage unit includes the gradient storage address.
[0022] Optionally, adjusting the parameter value of the first model parameter based on the accumulated gradient value stored in the first gradient storage unit includes: determining an average gradient value corresponding to the first identifier based on the accumulated gradient values stored in the first gradient storage unit; and adjusting the parameter value of the first model parameter based on the average gradient value corresponding to the first identifier. The average gradient value corresponding to the first identifier is also the average value of the accumulated gradient values stored in the first gradient storage unit.
[0023] The technical solution provided by the present application is that since the first gradient storage unit stores the accumulated gradient value corresponding to the first model parameter, the server can quickly determine the average gradient value corresponding to the first identifier based on the accumulated gradient value stored in the first gradient storage unit, and adjust the parameter value of the first model parameter based on the average gradient value corresponding to the first identifier, which can improve the efficiency of training the recommendation model.
[0024] Optionally, both the first user feature and the first content feature are vectors, that is, both the user feature and the content feature are vectorized features.
[0025] The technical solution provided in this application can improve the efficiency of training the recommendation model because the server trains the recommendation model based on vectorized user features and content features, compared with training the recommendation model based on natural language user features and content features.
[0026] In a second aspect, a data processing device is provided, comprising modules for executing the data processing method provided in the first aspect or any alternative embodiment of the first aspect. The modules in the data processing device may be implemented using software, hardware, or a combination of software and hardware, and the modules may be arbitrarily combined or divided based on the specific implementation.
[0027] In a third aspect, a data processing device is provided, comprising a memory and a processor; the memory is configured to store a computer program; and the processor is configured to execute the computer program stored in the memory to cause the data processing device to perform the method provided in the first aspect or any optional embodiment of the first aspect. Optionally, the data processing device is a server.
[0028] In a fourth aspect, a computing device cluster is provided, comprising at least one computing device, each of the at least one computing device comprising a processor and a memory, the processor of the at least one computing device being used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the data processing method provided in the first aspect or any optional manner of the first aspect.
[0029] Optionally, the computing device is a server, and the computing device cluster is a server cluster.
[0030] That is, the present application provides a server cluster, comprising at least one server, each of the at least one server comprising a processor and a memory, the processor of the at least one server being used to execute instructions stored in the memory of the at least one server, so that the server cluster executes the data processing method provided in the first aspect or any optional method of the first aspect mentioned above.
[0031] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed, the method provided in the first aspect or any optional manner of the first aspect is implemented.
[0032] In a sixth aspect, a computer program product is provided, which includes a program or code, and when the program or code is executed, it implements the method provided in the first aspect or any optional manner of the first aspect.
[0033] The technical effects of the second to sixth aspects mentioned above can refer to the technical effects of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] FIG1 is a schematic diagram of a recommendation model according to an embodiment of the present application that obtains prediction data based on sample data;
[0035] FIG2 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0036] FIG3 is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0037] FIG4 is a schematic diagram of a training recommendation model provided in an embodiment of the present application;
[0038] FIG5 is a schematic diagram of another training recommendation model provided in an embodiment of the present application;
[0039] FIG6 is a flow chart of a data processing method provided in an embodiment of the present application;
[0040] FIG7 is a schematic diagram of a storage space provided in an embodiment of the present application;
[0041] FIG8 is a schematic diagram of another storage space provided in an embodiment of the present application;
[0042] FIG9 is a schematic diagram of a data processing device provided in an embodiment of the present application;
[0043] FIG10 is a schematic diagram of a server provided in an embodiment of the present application;
[0044] FIG11 is a schematic diagram of a server cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The embodiments of the present application will be described in further detail below with reference to the accompanying drawings.
[0046] The emergence and widespread use of the internet has brought users a vast amount of data, satisfying their data needs in the data age. However, with the rapid development of the internet, the amount of data has increased dramatically, making it difficult for users to find the content they are interested in from this massive data. Therefore, using recommendation models to filter out the content of interest to users from this massive data and recommend this content to them is of great significance. For example, recommendation models can filter out and recommend products, videos, news, and other content that users are interested in from massive data.
[0047] The recommendation model can recommend data content to users based on the correlation between user characteristics and content characteristics. For example, for any user, the recommendation model filters out data content whose content characteristics match the user characteristics from a large amount of data based on the user characteristics of the user, and recommends the filtered data content to the user. By way of example, the recommendation model filters out data content with specific content characteristics from a large amount of data based on the user characteristics of the user, and recommends the data content to the user, where the correlation between the specific content characteristics and the user characteristics is greater than a preset threshold. The recommendation model can be a variety of possible machine learning models, such as a deep neural network model and a convolutional neural network model. The recommendation model can be trained by a server based on sample data. By way of example, the server obtains a sample data set, and the server trains the recommendation model based on the sample data set. The sample data set includes multiple sample data, each sample data corresponding to a standard prediction data, and the standard prediction data corresponding to each sample data can be the labeled value (or labeled data) of the sample data, and the standard prediction data corresponding to each sample data can be obtained by manually or computer-annotating the sample data. Specifically, for each sample data in the sample data set, the server extracts multiple sample features of the sample data, each of the multiple sample features has an identifier, and the same sample features correspond to the same model parameters of the recommendation model. The multiple sample features include user features and content features. The server inputs the user features and the content features into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user features and the content features to obtain prediction data corresponding to the sample data. The prediction data is used to characterize the correlation between the user features and the content features. For ease of description, the prediction data calculated by the recommendation model based on the user features and the content features during the training process is referred to as training prediction data. For the training prediction data corresponding to each sample data, the server calculates the loss value of the training prediction data (the loss value of the training prediction data corresponding to each sample data is the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data). The server determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The server then determines the average value (i.e., the average gradient value) of the gradient values of the model parameters corresponding to the same sample features (i.e., the same model parameters corresponding to the same sample features). The server adjusts the parameter value of the model parameter based on the average gradient value to train the recommendation model. The server trains the recommendation model based on each sample data in the sample dataset until convergence conditions are reached, and then the server ends the training process of the recommendation model.The convergence condition may include at least one of the following: the number of training cycles reaches a preset number, the fluctuation of the loss value of the training prediction data obtained during multiple consecutive training cycles is small, and the loss value of the training prediction data obtained during multiple consecutive training cycles is less than a preset loss value. A training process is also called an iteration or a training step.
[0048] During the training of the recommendation model, after the server determines the gradient value corresponding to the model parameter corresponding to the sample feature, the server first stores the gradient value corresponding to the model parameter in the storage space of the server. The server determines the average value of the gradient value corresponding to the model parameter corresponding to the sample feature with the same identifier based on the gradient value corresponding to the model parameter stored in the storage space of the server. The storage space of the server includes multiple gradient storage units, which are used to store the gradient values corresponding to the model parameters. For example, for each sample data in the sample data set, after the server obtains the training prediction data corresponding to the sample data, the server calculates the loss value of the training recommendation data, calculates the hash value of each sample feature of the sample data, and determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The server establishes a mapping relationship between the hash value of each sample feature of the sample data in the sample data set and the gradient storage address. Based on the mapping relationship, the server stores the gradient value corresponding to the model parameter corresponding to the sample feature to the corresponding gradient storage address (that is, to the gradient storage unit indicated by the corresponding gradient storage address). Based on the mapping relationship, the server obtains the gradient value stored at the gradient storage address corresponding to the hash value identifying the same sample feature (i.e., the gradient value stored in the gradient storage unit indicated by the gradient storage address), calculates the average value of the gradient values stored at the gradient storage address corresponding to the hash value identifying the same sample feature, and adjusts the parameter value of the model parameter corresponding to the sample feature identifying the same sample feature based on the average value. In the process of establishing the above mapping relationship, in order to avoid hash conflicts, for sample features with the same hash value, the server further determines the key code of these sample features and calculates the hash value of the key code of these sample features. The server establishes the above mapping relationship based on the hash value of the key code of these sample features. When the server obtains the gradient value based on the mapping relationship, the server first determines the gradient storage address based on the hash value of the sample feature. For sample features with the same hash value, the server further determines the gradient storage address based on the hash value of the key code of these sample features, and then the server obtains the gradient value stored at the determined gradient storage address. However, since for sample features with the same hash value, the server needs to further search the above mapping relationship based on the hash value of the key code to determine the gradient storage address, this leads to low data processing efficiency during the training of the recommendation model. Among them, the process of establishing the above mapping relationship is also called a hash map process. The above mapping relationship may be a hash table, and the content stored in the hash table is a key-value pair. The key field in the key-value pair is used to record the hash value, and the value field is used to record the gradient storage address.
[0049] The sample data may be natural language data, and the sample features may be vectorized (embedding) features, that is, the sample features may be vectors. Natural language data is data represented in natural language, such as a sentence or a paragraph of text. Vectorized features are vectors obtained by extracting features from natural language data using embedding technology. Vectors not only reflect the differences between features (e.g., sample features) but also facilitate determining the distance between features (e.g., sample features). Vectors obtained by extracting features from natural language data using embedding technology can also be called dense vectors. For example, as shown in FIG1 , a server (not shown in FIG1 ) includes a vectorization module that extracts features from sample data 1 using embedding technology to obtain user features 11 and content features 12. Both user features 11 and content features 12 are vectors. The vectorization module inputs user features 11 and content features 12 into a recommendation model to be trained. The recommendation model to be trained performs feature cross-calculation on user features 11 and content features 12 to obtain predicted data 1 corresponding to the sample data 1. The predicted data 1 is used to represent the correlation between user features 11 and content features 12.
[0050] Since the types and quantities of sample data in e-commerce, video, news and other fields are huge, and the number of sample features can even reach over 10 billion, how to improve the training efficiency of recommendation models for such a large amount of sample data has always been an important issue of concern in the industry.
[0051] The technical solution of this application is introduced below, and first the implementation environment involved in the embodiments of this application is introduced.
[0052] As an example, Figure 2 is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a server 10 and a client 20. The client 20 is communicatively connected to the server 10. For example, the client 20 and the server 10 are communicatively connected via a network such as a local area network and the Internet. The client 20 can send a sample data set to the server 10, and the server 10 executes the data processing method provided by the embodiment of the present application based on the sample data set sent by the client 20 to train the recommendation model. In the implementation environment shown in Figure 2, the server 10 can implement the data processing method provided by the embodiment of the present application by running an executable program. For example, the executable program of the data processing method can be presented in the form of an application installation package. After the server 10 installs the application installation package, it can implement the data processing method by running the executable program.
[0053] As another example, FIG3 is a schematic diagram of another implementation environment provided by an embodiment of the present application. The implementation environment includes a server cluster 30 and a client 20. The client 20 is connected to the server cluster 30 in communication. For example, the client 20 is connected to the server cluster 30 via a network communication such as a local area network or the Internet. The client 20 can send a sample data set to the server cluster 30, and the server cluster 30 executes the data processing method provided by the embodiment of the present application based on the sample data set sent by the client 20 to train the recommendation model. As shown in FIG3, the server cluster 30 includes multiple servers (FIG3 is illustrated with 4 servers as an example), and the client 20 can send all or part of the sample data in the sample data set to each server in the server cluster 30, and each server in the server cluster 30 executes the data processing method provided by the embodiment of the present application based on the sample data sent by the client 20 to train the recommendation model. In the implementation environment shown in FIG3, the server cluster 30 can implement the data processing method provided by the embodiment of the present application by running an executable program. For example, each server in the server cluster 30 implements the data processing method provided by the embodiment of the present application by running an executable program. For example, the executable program of the data processing method is presented in the form of an application installation package, and after the server installs the application installation package, it can implement the data processing method by running the executable program. In some embodiments, each server in the server cluster 30 is also referred to as a node or a service node.
[0054] In one implementation, the client 20 can be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, or a smart vehicle-mounted device. The server cluster 30 can be a cloud computing service center. In the embodiment of the present application, the server is a server for model training, and the server can be used for managing model parameters. For example, the server is an artificial intelligence (AI) server, which can also be called a parameter server (PS).
[0055] For example, in the process of training a recommendation model based on sample data in a sample data set, the server obtains a recommendation model to be trained, extracts sample features of the sample data in the sample data set, and trains the recommendation model to be trained based on the sample features of the sample data in the sample data set. Since the sample features needed by the server in each training process are highly sparse, that is, the server usually only uses the sample features of part of the sample data in the sample data set, therefore, in each training process, the server initializes the model parameters that need to be adjusted in this training process (that is, the model parameters corresponding to the sample features needed in this training process) based on the initial values of the model parameters corresponding to the sample features needed in this training process. The initialized recommendation model is also the recommendation model to be trained in this training process. During each training process, the server inputs the user features and content features of each sample data required for this training process into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user features and the content features to obtain the prediction data corresponding to the sample data (i.e., the training prediction data). The server determines the gradient values (gradients) corresponding to the model parameters based on the training prediction data corresponding to the sample data and the initial values of the model parameters corresponding to each sample feature of the sample data, and the server determines the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier. The server adjusts the parameter value of the model parameter based on the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier.
[0056] According to the above description, the process of training the recommendation model can be executed by one server or by multiple servers. In the case where the process of training the recommendation model is executed by multiple servers, the multiple servers can adopt a distributed parallel training method.
[0057] As an example, Figure 4 is a schematic diagram of a training recommendation model provided by an embodiment of the present application. Figure 4 illustrates the process of training the recommendation model by a server as an example. The server can be the server 10 in the implementation environment shown in Figure 2, or it can be any server in the server cluster 30 provided by the implementation environment shown in Figure 3. As shown in Figure 4, the server includes a parameter management module and multiple execution modules. The execution module is also called a worker, and the execution module is responsible for loading sample data, forward calculation (calculating the loss value of the training prediction data) and reverse calculation (calculating the gradient value corresponding to the model parameters) during the model training process. For example, the execution module is responsible for interacting with the client to receive sample data sent by the client (i.e., loading the sample data). During each training process, for each sample data required for this training process received by the execution module, the execution module inputs the user features and content features of the sample data into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user features and the content features to obtain the training prediction data corresponding to the sample data. The execution module calculates the loss value of the training prediction data (i.e., the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data), and the execution module determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The execution module sends the gradient value corresponding to the model parameter corresponding to each sample feature of the sample data to the parameter management module. That is, the execution module is used to execute the process of calculating the training prediction data corresponding to the sample data and calculating the gradient value corresponding to the model parameter corresponding to each sample feature of the sample data based on the recommendation model to be trained. For example, as shown in FIG4 , sample data 1 includes user feature A and content feature B, sample data 2 includes user feature C and content feature B, sample data 3 includes user feature A and content feature D, and sample data 4 includes user feature C and content feature E. Assume that the model parameters of the recommendation model to be trained include model parameters A, B, D, and E, and that user feature A corresponds to model parameter A, content feature B corresponds to model parameter B, user feature C corresponds to model parameter C, content feature D corresponds to model parameter D, and content feature E corresponds to model parameter E. Execution module 1 calculates the training prediction data corresponding to sample data 1 based on user feature A of sample data 1 and content feature B of sample data 1 using the recommendation model to be trained. Execution module 1 calculates the loss value of the training prediction data corresponding to sample data 1. Execution module 1 determines a gradient value A1 corresponding to model parameter A based on the loss value of the training prediction data corresponding to sample data 1 and the initial value of model parameter A. Execution module 1 also determines a gradient value B1 corresponding to model parameter B based on the loss value of the training prediction data corresponding to sample data 1 and the initial value of model parameter B.Execution module 2 calculates the training prediction data corresponding to sample data 2 based on the user feature C of sample data 2 and the content feature B of sample data 2 using the recommendation model to be trained. Execution module 2 calculates the loss value of the training prediction data corresponding to sample data 2. Execution module 2 determines the gradient value C1 corresponding to model parameter C based on the loss value of the training prediction data corresponding to sample data 2 and the initial value corresponding to model parameter C. Furthermore, execution module 2 determines the gradient value B2 corresponding to model parameter B based on the loss value of the training prediction data corresponding to sample data 2 and the initial value corresponding to model parameter B. Execution module 3 calculates the training prediction data corresponding to sample data 3 based on the user feature A of sample data 3 and the content feature D of sample data 3 using the recommendation model to be trained. Execution module 3 calculates the loss value of the training prediction data corresponding to sample data 3. Execution module 3 determines the gradient value A2 corresponding to model parameter A based on the loss value of the training prediction data corresponding to sample data 3 and the initial value corresponding to model parameter A. Finally, execution module 3 determines the gradient value D1 corresponding to model parameter D based on the loss value of the training prediction data corresponding to sample data 3 and the initial value corresponding to model parameter D. The execution module 4 calculates the training prediction data corresponding to the sample data 4 based on the user feature C of the sample data 4 and the content feature E of the sample data 4 using the recommendation model to be trained. The execution module 4 calculates the loss value of the training prediction data corresponding to the sample data 4. The execution module 4 determines the gradient value C2 corresponding to the model parameter C based on the loss value of the training prediction data corresponding to the sample data 4 and the initial value of the model parameter C. Furthermore, the execution module 4 determines the gradient value E1 corresponding to the model parameter E based on the loss value of the training prediction data corresponding to the sample data 4 and the initial value of the model parameter E. In the process of sending the gradient value corresponding to the model parameter corresponding to the sample feature to the parameter management module, the execution module also sends the identifier of the sample feature to the parameter management module. For example, the execution module simultaneously sends the gradient value corresponding to the model parameter corresponding to the sample feature and the identifier of the sample feature to the parameter management module. The parameter management module is configured to receive and store the gradient values sent by the multiple execution modules, determine the average value of the gradient values of the model parameter corresponding to the sample features with the same identifier based on the stored gradient values, and adjust the parameter value of the model parameter based on the average value of the gradient values of the model parameter corresponding to the sample features with the same identifier.
[0058] As another example, Figure 5 is a schematic diagram of another training recommendation model provided by an embodiment of the present application. Figure 5 illustrates the training process of the recommendation model performed by two servers, Server 1 and Server 2. These two servers can be any two servers in the server cluster 30 provided in the implementation environment shown in Figure 2. As shown in Figure 5, Server 1 and Server 2 each include a parameter management module and multiple execution modules. The execution module is responsible for interacting with the client to receive sample data sent by the client. During each training process, for each sample data required for this training process received by the execution module, the execution module inputs the user characteristics and content characteristics of the sample data into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user characteristics and the content characteristics to obtain the training prediction data corresponding to the sample data. The execution module calculates the loss value of the training prediction data (that is, the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data), and the execution module determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The execution module sends the gradient value corresponding to the model parameter corresponding to each sample feature of the sample data to the parameter management module in the server where the execution module is located. That is, the execution module is used to execute the process of calculating the training prediction data corresponding to the sample data and calculating the gradient value corresponding to the model parameter corresponding to each sample feature of the sample data based on the recommendation model to be trained. For example, as shown in FIG5 , each execution module in server 1 sends the calculated gradient values corresponding to the model parameters corresponding to the sample features to parameter management module 1, and each execution module in server 2 sends the calculated gradient values corresponding to the model parameters corresponding to the sample features to parameter management module 2. When each execution module in server 1 sends the gradient values corresponding to the model parameters corresponding to the sample features to parameter management module 1, it also sends the identifier of the sample features to parameter management module 1. When each execution module in server 2 sends the gradient values corresponding to the model parameters corresponding to the sample features to parameter management module 2, it also sends the identifier of the sample features to parameter management module 2. For example, the execution modules simultaneously send the gradient values corresponding to the model parameters corresponding to the sample features and the identifier of the sample features to the corresponding parameter management module. Parameter management module 1 is configured to receive and store the gradient values sent by multiple execution modules in server 1, and parameter management module 2 is configured to receive and store the gradient values sent by multiple execution modules in server 2. As shown in FIG5 , parameter management modules 1 and 2 also synchronize their stored gradient values. In this way, each parameter management module in parameter management module 1 and parameter management module 2 can obtain the gradient values stored by the other parameter management module.After parameter management module 1 and parameter management module 2 synchronize the gradient values, at least one of parameter management module 1 and parameter management module 2 determines the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identification (i.e., the average gradient value), and adjusts the parameter value of the model parameter according to the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identification.
[0059] In an embodiment of the present application, the sample data sent by the client to the server may be natural language data, and the server may first perform feature extraction on the sample data through embedding technology to obtain multiple sample features of the sample data, and the multiple sample features include user features and content features, and the multiple sample features are all vectors. For example, the server also includes a vectorization module, which is used to interact with the client to receive the sample data sent by the client, and the vectorization module performs feature extraction on the sample data through embedding technology to obtain multiple sample features of the sample data. In an embodiment of the present application, the server trains a recommendation model based on vectorized sample features, which can improve the efficiency of training the recommendation model, so that in some embodiments, the vectorization module for extracting sample features can also be called an accelerator card.
[0060] Figures 2 and 3 are exemplary illustrations of the implementation environment of the present application and do not constitute a limitation on the implementation environment of the present application. As business needs change, the implementation environment of the present application can be adjusted according to business needs. In addition, Figures 4 and 5 are exemplary illustrations of training recommendation models and do not constitute a limitation on training recommendation models. For example, Figure 5 shows a schematic diagram of training a recommendation model in a data parallel manner, and the recommendation model can also be trained in a model parallel manner. Wherein, data parallelism means: deploying the same recommendation model to be trained on multiple servers, and distributing the sample data in the sample data set used for training to the multiple servers, and each server independently uses part of the sample data in the sample data set for model training. During each training process, each server synchronizes the calculated gradient value to other servers, so that each server can calculate the average gradient value of a training process. The average gradient value is the average value of the gradient values obtained by each server in the same training process. Each server can adjust the parameter values of the model parameters of the recommendation model according to the average gradient value. Model parallelism involves dividing the recommendation model into multiple sub-models and deploying them on different servers. During training, the sub-models on different servers perform calculations on the same sample data set, following the structural order of the recommendation model. During a single training session, gradient values are calculated from the calculations of the sub-models on different servers. Forward or backward propagation based on these gradient values updates the parameters of each sub-model. After multiple iterations of training to meet training requirements, the sub-models on each server are reorganized based on the recommendation model structure, resulting in a fully trained recommendation model.
[0061] The above is an introduction to the implementation environment of this application. The following introduces the method embodiment of this application.
[0062] Please refer to Figure 6, which shows a flow chart of a data processing method provided in an embodiment of the present application. The data processing method is executed by a server. The server can be server 10 in the implementation environment shown in Figure 2 or any server in the implementation environment shown in Figure 3. Referring to Figure 6, the method includes the following steps S601 to S605.
[0063] S601. Determine a first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter of the recommendation model, wherein the first prediction data is calculated by the recommendation model based on the first user feature and the first content feature, the identifier of the first sample feature in the first user feature and the first content feature is the first identifier, and the first model parameter corresponds to the first identifier.
[0064] Among them, the first user feature and the first content feature are both sample features of the first sample data. The first sample feature is any one of the first user feature and the first content feature. For example, the first sample feature is the first user feature, or the first sample feature is the first content feature. The first sample data is any one of the sample data sets used to train the recommendation model. The first sample data can be natural language data, and the first user feature and the first content feature can both be vectors. Optionally, the server performs feature extraction on the first sample data through embedding technology to obtain multiple sample features of the first sample data. The multiple sample features are all vectors. The multiple sample features include the first user feature and the first content feature. Each sample feature in the multiple sample features has an identifier, and the identifiers of different sample features in the multiple sample features may be the same or different. The identifiers of sample features of different sample data may be the same or different. Sample features with the same identifier correspond to the same model parameters of the recommendation model. For example, the multiple sample features of sample data 1 include user feature A and content feature B, and the multiple sample features of sample data 3 include user feature A and content feature D. The identifier of user feature A of sample data 1 and the identifier of user feature A of sample data 3 are both "01", and the identifier of user feature A of sample data 1 and the identifier of user feature A of sample data 3 both correspond to model parameter 1 of the recommendation model. The identifier of content feature B is "02", and content feature B corresponds to model parameter 2 of the recommendation model. The identifier of content feature D is "04", and content feature D corresponds to model parameter 4 of the recommendation model. Optionally, the server includes a vectorization module that extracts features from the first sample data using embedding technology.
[0065] In an embodiment of the present application, a recommendation model is used to recommend data content to a user based on the correlation between user features and content features. For example, the recommendation model is used to recommend data content with specific content features to a user with specific user features, and the specific user features match the specific content features. For example, the correlation between the specific user features and the specific content features is greater than a preset threshold. Among them, the first prediction data is calculated by the recommendation model based on the first user features and the first content features, and the first prediction data is used to characterize the correlation between the first user features and the first content features (that is, the predicted correlation). The loss value of the first prediction data is the loss value of the first prediction data compared to the standard prediction data corresponding to the first sample data. The standard prediction data corresponding to the first sample data can be the annotation value of the first sample data (or called annotation data). The standard prediction data corresponding to the first sample data can be obtained by manually or computer-annotating the first sample data. The standard prediction data corresponding to the first sample data can be used to characterize the standard correlation between the first user features and the first content features (that is, the annotated correlation). Optionally, after the server obtains the first user feature and the first content feature, the server inputs the first user feature and the first content feature into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the first user feature and the first content feature to obtain first prediction data (i.e., training prediction data). The server can use a loss function to calculate the loss value of the first prediction data. The server can use a gradient calculation formula to calculate the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter. The gradient calculation formula can be an expression of the derivative function of the loss function, the independent variable of the gradient calculation formula is the loss value of the prediction data and the initial value of the model parameter, and the dependent variable of the gradient calculation formula is the gradient value corresponding to the model parameter.
[0066] In an embodiment of the present application, the recommendation model to be trained may be a recommendation model obtained after initializing each model parameter of the recommendation model according to the initial values of the model parameters. Since the sample features required by the server in each training process are highly sparse, and the model parameters have a corresponding relationship with the identifiers of the sample features, the server may initialize the model parameters based on the initial values of the model parameters corresponding to the sample features required for this training process (that is, the model parameters corresponding to the identifiers of the sample features required for this training process), that is, update the parameter values of the model parameters to the initial values.
[0067] In an optional embodiment, the server includes a parameter storage space, which includes a plurality of parameter storage units, each parameter storage unit is used to store the parameter value of a model parameter (or each parameter storage unit corresponds to a model parameter), each parameter storage unit corresponds to a sample feature identifier (that is, the identifier of the sample feature), and each parameter storage unit is used to store the parameter value of the model parameter corresponding to the sample feature with the corresponding identifier. During each training process, for each model parameter that needs to be adjusted during this training process (that is, each model parameter corresponding to the sample feature required for this training process), the server obtains the parameter value of the model parameter from the parameter storage space, and the server determines the parameter value of the model parameter obtained from the parameter storage space as the initial value of the model parameter, and the server uses the initial value of the model parameter to initialize the model parameter. Among them, the parameter value stored in each parameter storage unit can be a parameter value set manually or by a computer, or it can be a parameter value updated to the parameter storage unit during the previous training process (that is, the server uses the parameter value of the model parameter calculated during the previous training process as the initial value of the model parameter during the next training process). For example, the first parameter storage unit in the parameter storage space is used to store the parameter value of the first model parameter. The first parameter storage unit corresponds to the first identifier (the identifier of the first sample data). The server determines the first parameter storage unit based on the identifier of the first sample data (that is, the first identifier). The server obtains the parameter value stored in the first parameter storage unit. The server determines the parameter value obtained from the first parameter storage unit as the initial value of the first model parameter. The server initializes the first model parameter using the initial value of the first model parameter.
[0068] In one example, please refer to Figures 7 and 8, which both show schematic diagrams of the storage space of the server provided in an embodiment of the present application. As shown in Figures 7 and 8, the storage space of the server includes a parameter storage space, and the parameter storage space includes a plurality of parameter storage units (Figures 7 and 8 take the plurality of parameter storage units as parameter storage units 1 to 5 as an example). Each parameter storage unit in the plurality of parameter storage units is used to store the parameter value of a model parameter, and each parameter storage unit corresponds to a sample feature identifier. For example, parameter storage unit 1 corresponds to sample feature identifier "01" and is used for the parameter value of model parameter 1 (model parameter 1 corresponds to sample feature identifier "01"), parameter storage unit 2 corresponds to sample feature identifier "02" and is used for the parameter value of model parameter 2 (model parameter 2 corresponds to sample feature identifier "02"), parameter storage unit 3 corresponds to sample feature identifier "03" and is used for the parameter value of model parameter 3 (model parameter 3 corresponds to sample feature identifier "03"), and so on. For example, the first model parameter is model parameter 1, and the first identifier is the sample feature identifier "01". The server determines that the first parameter storage unit is parameter storage unit 1 based on the first identifier (that is, the sample feature identifier "01"). The server obtains the parameter value stored in parameter storage unit 1. The server determines the parameter value obtained from parameter storage unit 1 as the initial value of model parameter 1. The server initializes model parameter 1 using the initial value of model parameter 1.
[0069] S602. Determine a first gradient storage address based on the first identifier and a mapping relationship, where the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter.
[0070] After the server determines the first gradient value corresponding to the first model parameter, since the first model parameter corresponds to the identifier of the first sample feature (ie, the first identifier), the server determines the first gradient storage address according to the first identifier and the mapping relationship.
[0071] Among them, the first gradient storage address is used to indicate the first gradient storage unit, and the first gradient storage unit stores the accumulated gradient value corresponding to the first model parameter, and the accumulated gradient value corresponding to the first model parameter is the accumulated value of the gradient value corresponding to the first model parameter. Optionally, the first gradient storage unit is also used to store the first identifier and the accumulated number of times corresponding to the first model parameter. The accumulated number of times corresponding to the first model parameter is the accumulated number of times the gradient value corresponding to the first model parameter, and can also be referred to as the accumulated number of times corresponding to the first identifier. This embodiment of the present application does not limit this. For example, the first gradient storage unit includes a sample feature identifier field, an accumulated number field, and an accumulated gradient value field. The first identifier is stored in the sample feature identifier field, the accumulated number of times is stored in the accumulated number field, and the accumulated gradient value is stored in the accumulated gradient value field.
[0072] The mapping relationship is a mapping relationship between the identifier of the sample feature (referred to as the sample feature identifier) and the gradient storage address. Optionally, the mapping relationship includes a first mapping relationship and a second mapping relationship. The first mapping relationship is a mapping relationship between the sample feature identifier and the index storage address, and the second mapping relationship is a mapping relationship between the index storage address and the index information. The index information may include the gradient storage address. Thus, the mapping relationship between the sample feature identifier and the gradient storage address is realized by the first mapping relationship and the second mapping relationship. For example, the first mapping relationship includes a one-to-one mapping relationship between multiple sample feature identifiers and multiple index storage addresses, and the second mapping relationship includes a one-to-one mapping relationship between the multiple index storage addresses and multiple index information. Each index storage address in the multiple index storage addresses is used to indicate an index storage unit, and each index storage unit is used to store the index information corresponding to the corresponding sample feature identifier. The first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, and the second mapping relationship includes a mapping relationship between the first index storage address and the first index information. The first index storage address is used to indicate the first index storage unit, and the first index storage unit is used to store the index information corresponding to the first identifier.
[0073] In an optional embodiment, the server determines the first index storage address based on the first identifier and the first mapping relationship, and the server determines the first gradient storage address based on the first index storage address and the second mapping relationship. Specifically, the first mapping relationship is a one-to-one correspondence between the sample feature identifier and the gradient storage address. The server searches for the first mapping relationship based on the first identifier, and the server determines the index storage address corresponding to the first identifier in the first mapping relationship as the first index storage address based on the search result. The server determines the index unit indicated by the first index storage address as the first index storage unit, obtains the index information stored in the first index storage unit (that is, the index information corresponding to the first identifier), and determines the first gradient storage address based on the index information stored in the first index storage unit.
[0074] In an example, the first mapping relationship is shown in Table 1 below.
[0075] Table 1 (first mapping relationship)
[0076] Referring to Table 1, the first mapping relationship includes the correspondence between the sample feature identifier "01" and the index storage address 1, the correspondence between the sample feature identifier "02" and the index storage address 2, the correspondence between the sample feature identifier "03" and the index storage address 3, the correspondence between the sample feature identifier "04" and the index storage address 4, and so on.
[0077] Continuing with Figures 7 and 8 , the server's storage space also includes an index storage space, which includes multiple index storage units (Figures 7 and 8 illustrate the multiple index storage units as index storage units 1-4). Each of these multiple index storage units is used to store index information corresponding to a corresponding sample feature identifier. The index information includes at least one of status information, a space identifier, and a gradient storage address. The status information indicates the storage status corresponding to the corresponding sample feature identifier. The gradient storage address indicates the gradient storage unit. The space identifier indicates the gradient storage space where the gradient storage unit indicated by the gradient storage address is located. For example, each of these multiple index storage units includes a status information field, a space identifier field, and a gradient storage address field. The status information field is used to store status information, the space identifier field is used to store the space identifier of the gradient storage space, and the gradient storage address field is used to store the gradient storage address. For example, index storage address 1 indicates index storage unit 1, which stores the index information corresponding to the sample feature identifier "01." Index storage address 2 is used to indicate index storage unit 2, and index storage unit 2 is used for the index information corresponding to the sample feature identifier "02". Index storage address 3 is used to indicate index storage unit 3, and index storage unit 3 is used to store the index information corresponding to the sample feature identifier "03". Index storage address 4 is used to indicate index storage unit 4, and index storage unit 4 is used to store the index information corresponding to the sample feature identifier "04". And so on. Assuming that the first identifier is the sample feature identifier "01", the server determines the index storage address 1 corresponding to the first identifier by searching the first mapping relationship shown in Table 1 according to the first identifier. The server determines index storage address 1 as the first index storage address, and then determines index storage unit 1 indicated by index storage address 1 as the first index storage unit.
[0078] In an optional embodiment, the server determines the first gradient storage address based on the index information stored in the first index storage unit, including: the server determines whether the index information stored in the first index storage unit includes the gradient storage address; when the index information stored in the first index storage unit includes the gradient storage address, the server determines the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include the gradient storage address, the server determines the first free storage unit in the pre-created first storage space as the first gradient storage unit, or the server creates a second storage space and determines the second free storage unit in the second storage space as the first gradient storage unit, and the server determines the address of the first gradient storage unit as the first gradient storage address, wherein the first storage space and the second storage space are both gradient storage spaces, and the first storage space and the second storage space respectively include at least one storage unit, and the at least one storage unit is both a gradient storage unit.
[0079] In an optional embodiment, when the index information stored by the server in the first index storage unit does not include a gradient storage address, the server determines whether there are free storage units in the pre-created first storage space. If there are free storage units in the first storage space, the server identifies the first free storage unit in the first storage space as the first gradient storage unit, where the first free storage unit is any free storage unit in the first storage space. If there are no free storage units in the first storage space, the server creates a second storage space, which includes at least one storage unit. The server identifies the second free storage unit in the second storage space as the first gradient storage unit, where the second free storage unit is any free storage unit in the second storage space. In one embodiment, the server sequentially traverses each gradient storage unit in the first storage space to determine whether there are free storage units in the first storage space. In another embodiment, the server records the identifiers of each free storage unit in the first storage space, and the server determines whether there are free storage units in the first storage space based on the identifiers of the free storage units recorded by the server. In another embodiment, the server records the number of busy storage units in the first storage space and the total number of storage units in the first storage space, and the server determines whether there are free storage units in the first storage space based on the number of busy storage units in the first storage space and the total number of storage units in the first storage space. For example, if the number of busy storage units in the first storage space is less than the total number of storage units in the first storage space, the server determines that there are free storage units in the first storage space. If the number of busy storage units in the first storage space is equal to the total number of storage units in the first storage space, the server determines that there are no free storage units in the first storage space. In an optional embodiment, if there are no free storage units in the first storage space, the server divides a storage area in the storage space of the server as a second storage space, and divides a plurality of gradient storage units in the second storage space, which is not limited in this embodiment of the present application.
[0080] In an optional embodiment, the index information stored in the first index storage unit includes state information, and the state information is used to indicate whether the storage state corresponding to the first identifier is an initial state or an accumulated state. For example, when the state information is initial state information, the state information is used to indicate that the storage state corresponding to the first identifier is an initial state. When the state information is accumulated state information, the state information is used to indicate that the storage state corresponding to the first identifier is an accumulated state. For example, the initial state information is 0, and the accumulated state information is any non-zero information, such as the accumulated state information is 1. The server can determine the storage state corresponding to the first identifier based on the state information stored in the first index storage unit. Specifically, when the state information stored in the first index storage unit indicates that the storage state corresponding to the first identifier is an initial state, the server determines that the index information stored in the first index storage unit does not include a gradient storage address. When the state information stored in the first index storage unit indicates that the storage state corresponding to the first identifier is an accumulated state, the server determines that the index information stored in the first index storage unit includes a gradient storage address. In an optional embodiment, the index information stored by the server in the first index storage unit does not include the gradient storage address, and after the server determines the address of the first gradient storage unit as the first gradient storage address, the server stores the first gradient storage address to the first index storage unit, and the server updates the status information stored in the first index storage unit, and the updated status information is used to indicate that the storage status corresponding to the first identifier is an accumulation status.
[0081] In one embodiment, please continue to refer to Figures 7 and 8. The first index storage unit is index storage unit 1. The index information stored in index storage unit 1 includes status information, and the status information is cumulative status information. The server determines that the storage status corresponding to the first identifier is a cumulative status based on the status information, and then determines that the index information stored in index storage unit 1 includes a gradient storage address (for example, the server determines that a gradient storage address is stored in the gradient storage address field in index storage unit 1 shown in Figures 7 and 8). The server determines the gradient storage address stored in index storage unit 1 as the first gradient storage address.
[0082] In another embodiment, referring again to Figures 7 and 8 , the first index storage unit is index storage unit 1. The index information stored in index storage unit 1 includes state information, and this state information is initial state information. The server determines, based on this state information, that the storage state corresponding to the first identifier is the initial state, and further determines that the index information stored in index storage unit 1 does not include a gradient storage address (for example, the server determines that the gradient storage address field in index storage unit 1 shown in Figures 7 and 8 does not store a gradient storage address). Assuming that the first storage space is gradient storage space 1, and the server determines that the index information stored in index storage unit 1 does not include a gradient storage address, the server determines whether there are any free storage units in gradient storage space 1. In one example, as shown in Figure 7 , the server determines that there are free storage units in gradient storage space 1 and determines the first free storage unit in gradient storage space 1 as the first gradient storage unit. For example, if the first free storage unit is gradient storage unit 11, the server determines gradient storage unit 11 as the first gradient storage unit. Afterwards, the server determines the address of gradient storage unit 11 as the first gradient storage address, stores the first gradient storage address in index storage unit 1 (specifically, in the gradient storage address field in index storage unit 1), updates the status information in index storage unit 1 to accumulated status information, and stores the space identifier of gradient storage space 1 in index storage unit 1 (specifically, stores the space identifier of gradient storage space 1 in the space identifier field in index storage unit 1). In another example, as shown in FIG8 , the server determines that there are no free storage units in gradient storage space 1, creates a second storage space, and determines a second free storage unit in the second storage space as the first gradient storage unit. Take the example that the second storage space is the gradient storage space 2, and the second free storage unit is the gradient storage unit 21, that is, the server determines the gradient storage unit 21 as the first gradient storage unit, and then the server determines the address of the gradient storage unit 21 as the first gradient storage address, the server stores the first gradient storage address to the index storage unit 1 (specifically, it is stored in the gradient storage address field in the index storage unit 1), and the server updates the status information in the index storage unit 1 to the cumulative status information, and the server stores the space identifier of the gradient storage space 2 in the index storage unit 1 (specifically, it stores the space identifier of the gradient storage space 2 in the space identifier field in the index storage unit 1). In an embodiment of the present application, the gradient storage space is located in a specific storage space of the server, and the specific storage space is used to create (or divide the gradient storage space), and the total capacity of the specific storage space can be determined based on the total number of sample features of the sample data in the sample data set. For example, the total capacity of the specific storage space is 2 n , n is the total number of sample features of the sample data in the sample dataset.
[0083] S603. Add the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value.
[0084] In an optional embodiment, the server determines the first gradient storage unit based on the first gradient storage address, obtains the accumulated gradient value stored in the first gradient storage unit, and adds the first gradient value to the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value. For example, the first gradient storage address is gradient storage address 1, which indicates gradient storage unit 11. That is, the first gradient storage unit is gradient storage unit 11. Referring to FIG. 7 , gradient storage unit 11 is found according to gradient storage address 1, and the gradient value in gradient storage unit 11 is added to the accumulated gradient value stored in gradient storage unit 11 to obtain the first accumulated gradient value.
[0085] S604. Store the first accumulated gradient value in a first gradient storage unit.
[0086] In an optional embodiment, the server uses the first accumulated gradient value to update the accumulated gradient value stored in the first gradient storage unit. Specifically, the server deletes the accumulated gradient value stored in the first gradient storage unit and then stores the first accumulated gradient value in the first gradient storage unit. In one example, referring to FIG7 , the first gradient storage unit is gradient storage unit 11, and the server deletes the accumulated gradient value stored in gradient storage unit 11 and then stores the first accumulated gradient value in gradient storage unit 11. In another example, referring to FIG8 , the first gradient storage unit is gradient storage unit 21, and the server deletes the accumulated gradient value stored in gradient storage unit 21 and then stores the first accumulated gradient value in gradient storage unit 21.
[0087] In an optional embodiment, the first gradient storage unit is further used to store the cumulative number of times corresponding to the first model parameter. The cumulative number of times corresponding to the first model parameter is the cumulative number of times the gradient value corresponding to the first model parameter is accumulated, or the cumulative number of times the gradient value corresponding to the first model parameter is accumulated in the first gradient storage unit. In the initial state, the cumulative number of times corresponding to the first model parameter is the initial number (for example, 0). After the server stores the first accumulated gradient value in the first gradient storage unit, the server also updates the cumulative number stored in the first gradient storage unit. For example, the server adds 1 to the cumulative number stored in the first gradient storage unit.
[0088] S605. Adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit.
[0089] In an optional embodiment, the server determines the average gradient value corresponding to the first identifier (that is, the average gradient value of the gradient values corresponding to the first model parameter) based on the accumulated gradient value stored in the first gradient storage unit, and the server adjusts the parameter value of the first model parameter based on the average gradient value corresponding to the first identifier. In one example, the server adjusts the parameter value of the first model parameter to the average gradient value corresponding to the first identifier. In another example, the server uses an optimizer to optimize the average gradient value corresponding to the first identifier to obtain an optimized gradient value, and the server adjusts the parameter value of the first model parameter to the optimized gradient value. The optimizer can be an adaptive moment estimation (Adam) optimizer or a stochastic gradient descent optimizer. The Adam algorithm used by the Adam optimizer is an algorithm that extends the stochastic gradient descent method.
[0090] In an optional embodiment, the first gradient storage unit further stores the number of accumulations corresponding to the first model parameter, that is, the number of accumulations corresponding to the first identifier. The server determines the average gradient value corresponding to the first identifier based on the accumulated gradient value stored in the first gradient storage unit and the number of accumulations stored in the first gradient storage unit. For example, the server determines the quotient of the accumulated gradient value stored in the first gradient storage unit and the number of accumulations stored in the first gradient storage unit as the average gradient value corresponding to the first identifier.
[0091] In an optional embodiment, after the server adjusts the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit, the server clears the first gradient storage unit to restore the first gradient storage unit to an idle state. For example, the server deletes the first identifier, the number of accumulated times, and the accumulated gradient value stored in the first gradient storage unit to clear the first gradient storage unit. Since the first gradient storage unit is associated with the first index storage unit (the first gradient storage address for indicating the first gradient storage unit is stored in the first index storage unit), after the server clears the first gradient storage unit, the server also resets the first index storage unit. For example, the server deletes the index storage address stored in the first index storage unit and resets the state information stored in the first index storage unit to the initial state information.
[0092] In an optional embodiment, after the server adjusts the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier, the server stores the adjusted parameter value of the first model parameter in the first parameter storage unit. For example, the server first determines the first parameter storage unit according to the first identifier, and then stores the adjusted parameter value of the first model parameter in the first parameter storage unit. Specifically, the server deletes the parameter value stored in the first parameter storage unit, and then stores the adjusted parameter value of the first model parameter in the first parameter storage unit.
[0093] The accumulated gradient value stored in the first gradient storage unit described in S605 can be the first accumulated gradient value, or it can be the accumulated gradient value obtained by accumulating other gradient values corresponding to the first model parameter on the basis of the first accumulated gradient value. The accumulated gradient value stored in the first gradient storage unit used by the server to adjust the parameter value of the first model parameter is the accumulated gradient value after the gradient values of the model parameters corresponding to the sample features with the same sample feature identifier are accumulated. The server can determine the number of sample features with the same sample feature identifier based on all the sample features obtained, and determine the total accumulated number of gradient values corresponding to the first model parameter based on the number of sample features, that is, the server determines that the number of accumulations in the first gradient storage unit reaches the total accumulated number, then the gradient values of the model parameters corresponding to the sample features with the same sample feature identifier have all been accumulated. The embodiment of the present application does not limit this.
[0094] The embodiment of the present application is illustrated by taking the first model parameter of the training recommendation model as an example. The recommendation model may include multiple model parameters, each of which corresponds to different sample feature identifiers. The server can train the model parameter according to the sample features indicated by the sample feature identifier corresponding to each model parameter. The training process of each model parameter in the multiple model parameters can refer to the training process of the first model parameter, and the embodiment of the present application will not be repeated here. In addition, the server may include a parameter management module and multiple execution modules. The above S601 can be executed by the execution module, and S602 to S605 can be executed by the parameter management module. For example, after the execution module determines the first gradient value corresponding to the first model parameter, the execution module first stores the first gradient value in the cache space of the execution module, and then the execution module sends the gradient values stored in the cache space to the parameter management module in sequence. While the execution module sends each gradient value to the parameter management module, it can also send the sample feature identifier corresponding to the gradient value (that is, the sample feature identifier corresponding to the model parameter corresponding to the gradient value) to the parameter management module. For example, the execution module sends the first identifier and the first gradient value to the parameter management module at the same time.
[0095] In summary, the technical solution provided by the embodiment of the present application is a mapping relationship between the identifier of the sample feature and the gradient storage address, and the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address. After the server determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the server can quickly determine the first gradient storage address according to the first identifier (here refers to the identifier of the first sample feature) and the mapping relationship, and then accumulate the first gradient value corresponding to the first model parameter to the accumulated gradient value stored at the first gradient storage address. When the server trains the recommendation model, the server can quickly determine the first gradient storage address according to the first identifier and the mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address, and adjust the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address, thereby improving the data processing efficiency during the model training process, and thus improving the model training efficiency. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform cumulative calculation of gradient values when training the recommendation model, further improving the data processing efficiency and model training efficiency during the model training process.
[0096] The above is an introduction to the method embodiments of the present application. The following describes the device embodiments of the present application, which are used to perform the method of the present application. For details not disclosed in the device embodiments, please refer to the method embodiments.
[0097] Please refer to Figure 9, which shows a schematic diagram of a data processing device 900 provided in an embodiment of the present application. The data processing device 900 can be a server or a functional component within the server, and can implement all or part of the steps of the data processing method provided in the embodiment shown in Figure 6. As shown in Figure 9, the data processing device 900 includes a first determination module 901, a second determination module 902, an accumulation module 903, a storage module 904, and an adjustment module 905.
[0098] A first determination module 901 is configured to determine a first gradient value corresponding to a first model parameter based on a loss value of first prediction data and an initial value of a first model parameter of a recommendation model. The first prediction data is calculated by the recommendation model based on first user features and first content features. The identifier of a first sample feature in the first user features and first content features is a first identifier, and the first model parameter corresponds to the first identifier. For the implementation of the first determination module 901, reference can be made to the description of step S601 above.
[0099] A second determining module 902 is configured to determine a first gradient storage address based on the first identifier and a mapping relationship. The mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address. The first gradient storage address indicates a first gradient storage unit, which is configured to store the accumulated gradient value corresponding to the first model parameter. The implementation process of the second determining module 902 can be found in the description of step S602 above.
[0100] The accumulation module 903 is configured to accumulate the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value. The implementation process of the accumulation module 903 may refer to the relevant description in the above step S603.
[0101] The storage module 904 is configured to store the first accumulated gradient value in the first gradient storage unit. The implementation process of the storage module 904 may refer to the relevant description in the above step S604.
[0102] The adjustment module 905 is configured to adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit. The implementation process of the adjustment module 905 may refer to the relevant description in the above step S605.
[0103] Optionally, the second determination module 902 is used to: determine the first index storage address based on the first identifier and the first mapping relationship, the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, the first index storage address is used to indicate the first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; determine the first gradient storage address based on the index information stored in the first index storage unit.
[0104] Optionally, the second determination module 902 is configured to: when the index information stored in the first index storage unit includes a gradient storage address, determine the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include a gradient storage address, determine a first free storage unit in a pre-created first storage space as the first gradient storage unit, or create a second storage space and determine a second free storage unit in the second storage space as the first gradient storage unit, and determine the address of the first gradient storage unit as the first gradient storage address, wherein the first storage space and the second storage space each include at least one storage unit.
[0105] Optionally, the storage module 904 is further configured to: after the second determination module 902 determines that the index information stored in the first index storage unit does not include the gradient storage address and determines the address of the first gradient storage unit as the first gradient storage address, store the first gradient storage address in the first index storage unit.
[0106] Optionally, the index information stored in the first index storage unit includes status information, and the status information is used to indicate whether the storage state corresponding to the first identifier is an initial state or an accumulated state. The second determination module 902 is further used to: when the status information is used to indicate that the storage state corresponding to the first identifier is an initial state, determine that the index information stored in the first index storage unit does not include a gradient storage address; when the status information is used to indicate that the storage state corresponding to the first identifier is an accumulated state, determine that the index information stored in the first index storage unit includes a gradient storage address.
[0107] Optionally, the data processing device 900 further includes: an updating module 906, configured to update the state information after storing the first gradient storage address in the first index storage unit when the state information indicates that the storage state corresponding to the first identifier is an initial state.
[0108] Optionally, the adjustment module 905 is configured to: determine an average gradient value corresponding to the first identifier based on the accumulated gradient value stored in the first gradient storage unit; and adjust a parameter value of the first model parameter based on the average gradient value corresponding to the first identifier.
[0109] Optionally, both the first user feature and the first content feature are vectors.
[0110] Optionally, the data processing device 900 can be a server or a functional component in the server, and the data processing device 900 includes a parameter management module and multiple execution modules. The above-mentioned first determination module 901 can be a sub-module in the execution module, and the second determination module 902, the accumulation module 903, the storage module 904, the adjustment module 905 and the update module 906 can all be sub-modules in the parameter management module.
[0111] In summary, the technical solution provided by the embodiment of the present application is that the mapping relationship is a mapping relationship between the identifier of the sample feature and the gradient storage address, and the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address. After the first determination module determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the second determination module can quickly determine the first gradient storage address according to the first identifier (here refers to the identifier of the first sample feature) and the mapping relationship, and then the storage module accumulates the first gradient value corresponding to the first model parameter to the accumulated gradient value stored at the first gradient storage address. When the server is training the recommendation model, the second determination module in the server can quickly determine the first gradient storage address according to the first identifier and the mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address. The adjustment module adjusts the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address, thereby improving the data processing efficiency during the model training process, thereby improving the model training efficiency. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform cumulative calculation of gradient values when training the recommendation model, further improving the data processing efficiency and model training efficiency during the model training process.
[0112] An embodiment of the present application provides a data processing device, comprising a memory and a processor. The memory is configured to store a computer program. The processor is configured to execute the computer program stored in the memory, causing the data processing device to perform some or all of the functions of the data processing method provided in the embodiment of the present application. The data processing device may be a server or a functional component within a server.
[0113] In the embodiments of the present application, the data processing device is described as a server. For example, FIG10 is a schematic diagram of the structure of a server 100 provided in an embodiment of the present application. As shown in FIG10 , the server 100 includes a processor 101, a memory 102, a communication interface 103, and a bus 104. The processor 101, the memory 102, and the communication interface 103 are connected to each other via the bus 104.
[0114] The processor 101 may include a general-purpose processor and / or a dedicated hardware chip. The general-purpose processor may include: a central processing unit (CPU), a microprocessor or a graphics processing unit (GPU). The CPU is, for example, a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The dedicated hardware chip is a hardware module for high-performance processing. The dedicated hardware chip includes at least one of a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or a network processor (NP). The processor 101 may also be an integrated circuit chip with signal processing capabilities. During implementation, some or all of the functions of the data processing method of the present application may be completed by the integrated logic circuit of the hardware in the processor 101 or instructions in the form of software.
[0115] Memory 102 is used to store computer programs, which include an operating system 102a and executable code (i.e., program instructions) 102b. Memory 102 may be, for example, a read-only memory or other type of static storage device capable of storing static information and instructions, a random access memory or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory, a read-only optical disc or other optical disc storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired executable code in the form of instructions or data structures and accessible by a computer, but not limited to these. For example, memory 102 is used to store an outbound port queue, etc. Memory 102 may be independent and connected to processor 101 via bus 104. Alternatively, memory 102 and processor 101 may be integrated. The memory 102 can store executable code. When the executable code stored in the memory 102 is executed by the processor 101, the processor 101 is used to perform some or all of the functions of the data processing method provided in the embodiments of the present application. For the implementation of the process executed by the processor 101, please refer to the relevant description of the aforementioned embodiments. The memory 102 may also include software modules and data required for other running processes, such as the operating system.
[0116] The communication interface 103 uses a transceiver module, such as, but not limited to, a transceiver, to communicate with other devices or communication networks. For example, the communication interface 103 can be any one or a combination of the following devices: a network interface (such as an Ethernet interface), a wireless network card, or other device with network access capabilities.
[0117] The bus 104 is any type of communication bus used to interconnect the internal devices of the server (e.g., the memory 102, the processor 101, and the communication interface 103). For example, a system bus is provided. The embodiments of the present application illustrate the interconnection of the aforementioned devices within the server via the bus 104. Optionally, the aforementioned devices within the server 100 may also be connected to each other using other connection methods besides the bus 104. For example, the aforementioned devices within the server 100 may be interconnected via an internal logical interface.
[0118] It should be noted that the above-mentioned multiple devices can be respectively arranged on independent chips, or at least partially or completely arranged on the same chip. Whether each device is independently arranged on different chips or integrated on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation form of the above-mentioned devices. The descriptions of the processes corresponding to the above-mentioned figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.
[0119] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product providing a program development platform includes one or more computer instructions. When these computer program instructions are loaded and executed on a server, all or part of the functions of the data processing method provided in the embodiments of the present application are implemented.
[0120] Furthermore, computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium stores computer program instructions that provide a program development platform.
[0121] An embodiment of the present application provides a computing device cluster, comprising at least one computing device. Each computing device in the at least one computing device includes a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs some or all of the functions of the data processing method provided in the embodiment of the present application.
[0122] Optionally, the computing device is a server, and the computing device cluster is a server cluster. That is, an embodiment of the present application provides a server cluster, the server cluster including at least one server, each of the at least one server including a processor and a memory, the processor of the at least one server being configured to execute instructions stored in the memory of the at least one server, so that the server cluster performs some or all of the functions of the data processing method provided in the embodiment of the present application.
[0123] Optionally, the structure of at least one server included in the server cluster may refer to the server 100 shown in Figure 10. The memory 102 of one or more servers 100 in the server cluster may store the same instructions for executing the data processing method.
[0124] In some possible implementations, the memory 102 of one or more servers 100 in the server cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more servers 100 can jointly execute the instructions for executing the data processing method.
[0125] It should be noted that the memory 102 in different servers 100 in the server cluster can store different instructions, each for executing a portion of the functions of the data processing method. That is, the instructions stored in the memory 102 in different servers 100 can implement the functions of one or more of the first determination module 901, the second determination module 902, the accumulation module 903, the storage module 904, the adjustment module 905, and the update module 906 shown in Figure 9.
[0126] In some possible implementations, one or more servers in a server cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG11 shows a possible implementation. As shown in FIG11 , multiple servers 110A, 110B, and 110C are connected via a network. Specifically, each server is connected to the network via a communication interface in the server. In this type of possible implementation, the servers 110A, 110B, and 110C include a bus 112, a processor 114, a memory 116, and a communication interface 118. The memory 116 in the server 110A stores instructions for the functions of the processor. At the same time, the memory 116 in the server 110B stores instructions for the functions of the processor. The memory 116 in the server 110C stores instructions for the functions of the processor.
[0127] It should be understood that the functions of server 110A shown in FIG11 may also be performed by multiple servers 110. Similarly, the functions of server 110B may also be performed by multiple servers 110. The functions of server 110C may also be performed by multiple servers 110. Furthermore, the deployment method of the modules used to implement the data processing method in the servers may also be adjusted according to application requirements.
[0128] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed (for example, by a server, a server cluster, a data processing device, one or more processors, etc.), it implements all or part of the steps of the data processing method provided in the above method embodiment.
[0129] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes a program or code. When the program or code is executed (for example, by a server, a server cluster, a data processing device, one or more processors, etc.), it implements all or part of the steps of the data processing method provided in the above method embodiment.
[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium (e.g., a solid-state hard disk).
[0131] It should be understood that the term "at least one" in this application refers to one or more, and "a plurality of" refers to two or more. In this application, unless otherwise specified, the symbol " / " generally means or, for example, A / B can mean A or B. The term "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, for the sake of clarity of description, this application uses words such as "first", "second", and "third" to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art will understand that words such as "first", "second", and "third" do not limit the quantity and execution order.
[0132] Different types of embodiments, such as method embodiments and device embodiments, provided in the embodiments of the present application can refer to each other. The order of operations in the method embodiments can be appropriately adjusted, and operations can be increased or decreased in response to circumstances. Any technician familiar with this technical field can easily think of changing methods within the technical scope disclosed in this application, and the methods should be covered within the scope of protection of this application, so they will not be repeated here.
[0133] In the corresponding embodiments provided in the present application, it should be understood that the disclosed devices and the like can be implemented through other structural methods. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. On the other hand, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical or other forms. The modules described as separate components may or may not be physically separated, and the components described as modules may or may not be physical modules, and may be located in one place or distributed on multiple network nodes. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0134] The above description is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: The method comprises: determining a first gradient value corresponding to a first model parameter of a recommendation model based on a loss value of first prediction data and an initial value of the first model parameter of the recommendation model, wherein the first prediction data is calculated by the recommendation model based on a first user feature and a first content feature, an identifier of a first sample feature in the first user feature and the first content feature is a first identifier, and the first model parameter corresponds to the first identifier; determining a first gradient storage address according to the first identifier and a mapping relationship, wherein the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter; Accumulating the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; storing the first accumulated gradient value in the first gradient storage unit; The parameter value of the first model parameter is adjusted according to the accumulated gradient value stored in the first gradient storage unit.
2. The method according to claim 1, characterized in that The determining the first gradient storage address according to the first identifier and the mapping relationship includes: Determine a first index storage address according to the first identifier and a first mapping relationship, where the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, where the first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; The first gradient storage address is determined according to the index information stored in the first index storage unit.
3. The method according to claim 2, characterized in that The determining the first gradient storage address according to the index information stored in the first index storage unit includes: In a case where the index information stored in the first index storage unit includes a gradient storage address, determining the gradient storage address stored in the first index storage unit as the first gradient storage address; In a case where the index information stored in the first index storage unit does not include a gradient storage address, a first free storage unit in a pre-created first storage space is determined as the first gradient storage unit, or a second storage space is created and a second free storage unit in the second storage space is determined as the first gradient storage unit, and an address of the first gradient storage unit is determined as the first gradient storage address, wherein the first storage space and the second storage space respectively include at least one storage unit.
4. The method according to claim 3, characterized in that The method further comprises: The index information stored in the first index storage unit does not include a gradient storage address, and after the address of the first gradient storage unit is determined as the first gradient storage address, the first gradient storage address is stored in the first index storage unit.
5. The method according to claim 4, characterized in that The index information stored in the first index storage unit includes state information, and the state information is used to indicate whether the storage state corresponding to the first identifier is an initial state or an accumulated state. The method further includes: When the state information indicates that the storage state corresponding to the first identifier is an initial state, determining that the index information stored in the first index storage unit does not include a gradient storage address; In a case where the state information is used to indicate that the storage state corresponding to the first identifier is an accumulation state, it is determined that the index information stored in the first index storage unit includes a gradient storage address.
6. The method according to claim 5, characterized in that The method further comprises: In a case where the state information is used to indicate that the storage state corresponding to the first identifier is an initial state, after the first gradient storage address is stored in the first index storage unit, the state information is updated.
7. The method according to any one of claims 1 to 6, characterized in that: The step of adjusting the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit includes: determining an average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; The parameter value of the first model parameter is adjusted according to the average gradient value corresponding to the first identifier.
8. The method according to any one of claims 1 to 7, characterized in that: The first user feature and the first content feature are both vectors.
9. A data processing device, characterized in that: The device comprises: a first determining module, configured to determine a first gradient value corresponding to a first model parameter of a recommendation model based on a loss value of first prediction data and an initial value of the first model parameter of the recommendation model, wherein the first prediction data is calculated by the recommendation model based on a first user feature and a first content feature, an identifier of a first sample feature in the first user feature and the first content feature is a first identifier, and the first model parameter corresponds to the first identifier; a second determining module, configured to determine a first gradient storage address according to the first identifier and a mapping relationship, wherein the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter; an accumulation module, configured to accumulate the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; A storage module, configured to store the first accumulated gradient value in the first gradient storage unit; An adjustment module is used to adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit.
10. The device according to claim 9, characterized in that The second determining module is used to: Determine a first index storage address according to the first identifier and a first mapping relationship, where the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, where the first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; The first gradient storage address is determined according to the index information stored in the first index storage unit.
11. The device according to claim 10, characterized in that The second determining module is used to: In a case where the index information stored in the first index storage unit includes a gradient storage address, determining the gradient storage address stored in the first index storage unit as the first gradient storage address; In a case where the index information stored in the first index storage unit does not include a gradient storage address, a first free storage unit in a pre-created first storage space is determined as the first gradient storage unit, or a second storage space is created and a second free storage unit in the second storage space is determined as the first gradient storage unit, and an address of the first gradient storage unit is determined as the first gradient storage address, wherein the first storage space and the second storage space respectively include at least one storage unit.
12. The device according to claim 11, characterized in that The storage module is further configured to: after the second determining module determines that the index information stored in the first index storage unit does not include a gradient storage address and determines the address of the first gradient storage unit as the first gradient storage address, store the first gradient storage address in the first index storage unit.
13. The device according to claim 12, characterized in that The index information stored in the first index storage unit includes state information, and the state information is used to indicate whether the storage state corresponding to the first identifier is an initial state or an accumulation state. The second determining module is further used for: When the state information indicates that the storage state corresponding to the first identifier is an initial state, determining that the index information stored in the first index storage unit does not include a gradient storage address; In a case where the state information is used to indicate that the storage state corresponding to the first identifier is an accumulation state, it is determined that the index information stored in the first index storage unit includes a gradient storage address.
14. The device according to claim 13, characterized in that The device also includes: An updating module is configured to update the state information after storing the first gradient storage address in the first index storage unit when the state information indicates that the storage state corresponding to the first identifier is an initial state.
15. The device according to any one of claims 9 to 14, characterized in that The adjustment module is used for: determining an average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; The parameter value of the first model parameter is adjusted according to the average gradient value corresponding to the first identifier.
16. The device according to any one of claims 9 to 15, characterized in that The first user feature and the first content feature are both vectors.
17. A data processing device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory so as to enable the data processing device to perform the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 8 is implemented.
19. A computer program product, characterized in that The computer program product comprises a program or a code, and when the program or the code is executed, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Data processing method and device
CN120123758A
Model parameter training method and device, server and storage medium
CN108491928A
Model training method, device and equipment, and storage medium
CN112541513A
Deep learning model training method and training single machine
CN114997416A
Model parameter updating method, apparatus and device, storage medium, and program product
WO2022193432A1