Data processing method and device
By establishing the mapping relationship between sample feature identification and gradient storage address during training recommended model, quickly determining the gradient storage address and accumulating gradient values, the problem of low data processing efficiency during training is solved, and more efficient model training is achieved.
Patent Information
- Application Number
- CN202311690834.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
During the training of the recommended model, data processing efficiency is low because it is necessary to further look up the mapping relationship based on the hash value of the key code to determine the gradient storage address.
By establishing a mapping relationship between the identification of sample features and the gradient storage address, the gradient storage address is quickly determined based on the identification, and the gradient value is accumulated into the corresponding storage unit to avoid the accumulation calculation of the gradient value.
The data processing efficiency and model training efficiency in the training recommended model are improved, and the computational complexity is reduced.
Smart Images

Figure CN120123758A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method and apparatus. Background Art
[0002] A recommendation model is used to screen data content that a user is interested in from a large amount of data and recommend the screened data content to the user. The recommendation model can be obtained by a server training according to sample data. For example, for each sample data in a sample data set, the server extracts multiple sample features of the sample data, the server calculates a gradient value corresponding to a model parameter corresponding to each sample feature among the multiple sample features and a hash value of each sample feature, the server establishes a mapping relationship between the hash value of each sample feature and a gradient storage address, and the server stores the gradient value corresponding to the model parameter corresponding to each sample feature into a gradient storage unit indicated by the corresponding gradient storage address based on the mapping relationship. The server obtains the gradient value stored at the gradient storage address corresponding to the hash value of the sample features with the same identifier (that is, the gradient value stored in the gradient storage unit indicated by the gradient storage address) based on the mapping relationship, the server calculates an average value of the gradient values stored at the gradient storage address corresponding to the hash value of the sample features with the same identifier, and adjusts a parameter value of the model parameter of the recommendation model based on the average value.
[0003] In the process of establishing the above mapping relationship, in order to avoid hash collisions, for sample features with the same hash value, the server obtains hash values of the key codes of these sample features, and the server establishes the mapping relationship based on the hash values of the key codes of these sample features. When the server obtains the gradient value based on the mapping relationship, the server first determines the gradient storage address according to the hash value of the sample feature. For sample features with the same hash value, the server further determines the gradient storage address according to the hash values of the key codes of these sample features, and then the server obtains the gradient value stored in the determined gradient storage address. Since for sample features with the same hash value, the server needs to further search the above mapping relationship according to the hash value of the key code to determine the gradient storage address, this results in low data processing efficiency in the process of training the recommendation model. Summary of the Invention
[0004] This application provides a data processing method and apparatus. The technical solution provided by this application can improve the data processing efficiency in the process of training a recommendation model. The technical solution of this application is as follows.
[0005] In a first aspect, a data processing method is provided. The method includes: determining a first gradient value corresponding to a first model parameter based on a loss value of first prediction data and an initial value of the first model parameter of a recommendation model, where the first prediction data is calculated by the recommendation model based on a first user feature and a first content feature, the identifier of the first sample feature in the first user feature and the first content feature is a first identifier, and the first model parameter corresponds to the first identifier; determining a first gradient storage address according to the first identifier (here referring to the identifier of the first sample feature) and a mapping relationship, where the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter; accumulating the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; storing the first accumulated gradient value in the first gradient storage unit; and adjusting the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit (which may be the first accumulated gradient value or an accumulated gradient value obtained by accumulating other gradient values corresponding to the first model parameter on the basis of the first accumulated gradient value).
[0006] Among them, the above mapping relationship may be a mapping relationship between the identifier of the sample feature and the gradient storage address. For the convenience of description, in some descriptions below, the identifier of the sample feature is simply referred to as the feature identifier or the sample feature identifier. The identifier of the first sample feature is the first identifier, and there may also be other sample feature identifiers that are the first identifier, that is, the identifiers of different sample features may be the same or different.
[0007] Among them, both the first user feature and the first content feature are sample features of the first sample data. The loss value of the first prediction data refers to the loss of the first prediction data compared to the standard prediction data corresponding to the first sample data. The standard prediction data corresponding to the first sample data is the annotation value (or called the annotation data) of the first sample data. The first prediction data is the training prediction data corresponding to the first sample data, and the first prediction data is used to represent the prediction correlation degree (that is, the predicted correlation degree) between the first user feature and the first content feature, and the standard prediction data corresponding to the first sample data is used to represent the standard correlation degree (that is, the annotated correlation degree) between the first user feature and the first content feature. In this application, the user features of different sample data may be the same or different, and the content features of different sample data may be the same or different.
[0008] Among them, the accumulated gradient value corresponding to the first model parameter is the accumulated value of the gradient value corresponding to the first model parameter. The first gradient storage unit is used to store the accumulated gradient value corresponding to the first model parameter, that is, the first gradient storage unit is used to store the accumulated gradient value corresponding to the model parameter corresponding to the sample feature marked with the first identifier. In this application, since the sample features marked with the first identifier all correspond to the first model parameter, this application is described as "the first gradient storage unit is used to store the accumulated gradient value corresponding to the first model parameter".
[0009] Among them, the data processing method of this application can be executed by a single server or by a server cluster composed of multiple servers. The server can be a server for model training, and the server is used for the management of model parameters. For example, the server is an artificial intelligence (AI) server, and the server is also called a parameter server (PS).
[0010] For the technical solution provided in this application, the mapping relationship is the mapping relationship between the identifier of the sample feature and the gradient storage address, and the mapping relationship includes the mapping relationship between the first identifier and the first gradient storage address. After the server determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the server can quickly determine the first gradient storage address according to the first identifier (here referring to the identifier of the first sample feature) and the mapping relationship, and then accumulate the first gradient value corresponding to the first model parameter into the accumulated gradient value stored at the first gradient storage address. When training the recommendation model, the server can quickly determine the first gradient storage address according to the first identifier and the mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address, and adjust the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address. Therefore, the data processing efficiency during the model training process can be improved, and then the model training efficiency can be improved. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform the accumulation calculation of the gradient values when training the recommendation model, further improving the data processing efficiency and model training efficiency during the model training process. In this application, the first gradient storage address is used to indicate the first gradient storage unit, and the accumulated gradient value stored at the first gradient storage address refers to the accumulated gradient value stored in the first gradient storage unit. For the sake of concise description in this application, it is described as "the accumulated gradient value stored at the first gradient storage address". The meaning of other similar descriptions is the same.
[0011] Optionally, determining the first gradient storage address according to the first identifier and the mapping relationship includes: determining the first index storage address according to the first identifier and the first mapping relationship, where the first mapping relationship is the mapping relationship between the identifier of the sample feature and the index storage address, the first mapping relationship includes the mapping relationship between the first identifier and the first index storage address, the first index storage address is used to indicate the first index storage unit, and the first index storage unit is used to store the index information corresponding to the first identifier; determining the first gradient storage address according to the index information stored in the first index storage unit. Specifically, the server determines the first index storage unit according to the first index storage address, and then determines the first gradient storage address according to the index information stored in the first index storage unit.
[0012] In the technical solution provided by this application, since the first mapping relationship is the mapping relationship between the identifier of the sample feature and the index storage address, and the first mapping relationship includes the mapping relationship between the first identifier and the first index storage address, after the server obtains the first gradient value corresponding to the first model parameter, the server can quickly determine the first index storage address according to the first identifier and the first mapping relationship, and then quickly determine the first gradient storage address according to the index information stored in the first index storage address, which can improve the efficiency of the server in determining the first gradient storage address, and further improve the data processing efficiency in the model training process. In this application, the first index storage address is used to indicate the first index storage unit, and the index information stored in the first index storage address refers to the accumulated gradient value stored in the first index storage unit. For the sake of concise description in this application, it is described as "the index information stored in the first index storage address". The meanings of other similar descriptions are the same.
[0013] Optionally, determining the first gradient storage address according to the index information stored in the first index storage unit includes: when the index information stored in the first index storage unit includes the gradient storage address, determining the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include the gradient storage address, determining the first free storage unit in the pre-created first storage space as the first gradient storage unit, or creating a second storage space and determining the second free storage unit in the second storage space as the first gradient storage unit, and determining the address of the first gradient storage unit as the first gradient storage address. Wherein, the first storage space and the second storage space each include at least one storage unit.
[0014] In the technical solution provided by this application, when the index information stored in the first index storage unit includes the gradient storage address, the server directly determines the gradient storage address stored in the first index storage unit as the first gradient storage address. Therefore, the first gradient storage address can be quickly determined, and the efficiency of the server in determining the first gradient storage address can be improved. When the index information stored in the first index storage unit does not include the gradient storage address, the server can determine the first gradient storage address based on the idle storage unit. Therefore, regardless of whether the index information stored in the first index storage unit includes the gradient storage address, the server can determine the first gradient storage address for storing the accumulated gradient value corresponding to the model parameter of the sample feature with the first identifier, which can improve the flexibility of the server in determining the gradient storage address.
[0015] Optionally, the method further includes: after determining that the index information stored in the first index storage unit does not include the gradient storage address and determining the address of the first gradient storage unit as the first gradient storage address, storing the first gradient storage address in the first index storage unit.
[0016] In the technical solution provided by this application, since the server stores the first gradient storage address in the first index storage unit after determining the first gradient storage address when the index information stored in the first index storage unit does not have the gradient storage address, it is convenient for the server to quickly determine the first gradient storage address according to the index information stored in the first index storage unit subsequently.
[0017] Optionally, the index information stored in the first index storage unit includes status information, and the status information is used to indicate that the storage status corresponding to the first identifier is the initial state or the accumulation state. The method further includes: when the status information is used to indicate that the storage status corresponding to the first identifier is the initial state, determining that the index information stored in the first index storage unit does not include the gradient storage address; when the status information is used to indicate that the storage status corresponding to the first identifier is the accumulation state, determining that the index information stored in the first index storage unit includes the gradient storage address. For example, the first index storage unit includes a status field for recording the status information and an address field for recording the gradient storage address. The address field is located after the status field, and the length of the status field is less than the length of the address field.
[0018] In the technical solution provided by this application, since the index information stored in the first index storage unit includes the status information, the server can quickly determine whether the index information stored in the first index storage unit includes the gradient storage address according to the status information, which can improve the efficiency of the server in determining whether the index information stored in the first index storage unit includes the gradient storage address.
[0019] Optionally, the method further includes: when the status information is used to indicate that the storage status corresponding to the first identifier is the initial status, after storing the first gradient storage address into the first index storage unit, updating the status information. The updated status information is used to indicate that the storage status corresponding to the first identifier is the accumulation status. The server updates the status information, which can facilitate the server to quickly determine according to the status information that the index information stored in the first index storage unit includes the gradient storage address.
[0020] Optionally, adjusting the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit includes: determining the average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; adjusting the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier. The average gradient value corresponding to the first identifier is also the average value of the accumulated gradient values stored in the first gradient storage unit.
[0021] In the technical solution provided by this application, since the first gradient storage unit stores the accumulated gradient value corresponding to the first model parameter, the server can quickly determine the average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit, and adjust the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier, which can improve the efficiency of training the recommendation model.
[0022] Optionally, both the first user feature and the first content feature are vectors. That is, both the user feature and the content feature are vectorized features.
[0023] In the technical solution provided by this application, since the server trains the recommendation model based on the vectorized user features and content features, compared with training the recommendation model based on the natural language user features and content features, it can improve the efficiency of training the recommendation model.
[0024] In a second aspect, a data processing device is provided, including modules for executing the data processing method provided in the first aspect or any optional manner of the first aspect. The modules in the data processing device can be implemented based on software, hardware, or a combination of software and hardware, and these modules can be arbitrarily combined or divided based on specific implementations.
[0025] In a third aspect, a data processing device is provided, including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the data processing device executes the method provided in the first aspect or any optional manner of the first aspect. Optionally, the data processing device is a server.
[0026] Fourthly, a computing device cluster is provided, including at least one computing device. Each computing device in the at least one computing device includes a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the data processing method provided in the first aspect or any optional implementation manner of the first aspect as described above.
[0027] Optionally, the computing device is a server, and the computing device cluster is a server cluster.
[0028] That is to say, the present application provides a server cluster, including at least one server. Each server in the at least one server includes a processor and a memory. The processor of the at least one server is configured to execute instructions stored in the memory of the at least one server, so that the server cluster executes the data processing method provided in the first aspect or any optional implementation manner of the first aspect as described above.
[0029] Fifthly, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium. When the computer program is executed, the method provided in the first aspect or any optional implementation manner of the first aspect as described above is implemented.
[0030] Sixthly, a computer program product is provided. The computer program product includes a program or code. When the program or code is executed, the method provided in the first aspect or any optional implementation manner of the first aspect as described above is implemented.
[0031] For the technical effects of the second aspect to the sixth aspect above, reference may be made to the technical effects of the first aspect, which will not be elaborated here. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of a recommendation model obtaining prediction data based on sample data provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of training a recommendation model provided by an embodiment of the present application;
[0036] Figure 5 It is a schematic diagram of another training of a recommendation model provided by an embodiment of the present application;
[0037] Figure 6It is a flowchart of a data processing method provided by an embodiment of the present application;
[0038] Figure 7 It is a schematic diagram of a storage space provided by an embodiment of the present application;
[0039] Figure 8 It is another schematic diagram of a storage space provided by an embodiment of the present application;
[0040] Figure 9 It is a schematic diagram of a data processing device provided by an embodiment of the present application;
[0041] Figure 10 It is a schematic diagram of a server provided by an embodiment of the present application;
[0042] Figure 11 It is a schematic diagram of a server cluster provided by an embodiment of the present application. Detailed implementation manners
[0043] The following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0044] The emergence and popularization of the Internet have brought a large amount of data to users, meeting the users' needs for data in the data era. However, with the substantial growth of the data volume brought about by the rapid development of the network, it is difficult for users to obtain the data content they are concerned about from the vast amount of data. Therefore, it is of great significance to screen out the data content concerned by users from the vast amount of data based on a recommendation model and recommend the screened data content to users. For example, based on the recommendation model, screen out the goods, videos, news, etc. concerned by users from the vast amount of data and recommend them to users.
[0045] Among them, the recommendation model can recommend data content to users based on the correlation between user features and content features. For example, for any user, the recommendation model filters out data content whose content features match the user features from the massive data based on the user features of the user, and recommends the filtered data content to the user. By way of example, the recommendation model filters out data content with specific content features from the massive data based on the user features of the user, and recommends the data content to the user, and the correlation between the specific content features and the user features is greater than a preset threshold. Among them, the recommendation model can be various possible machine learning models such as a deep neural network model, a convolutional neural network model, etc. The recommendation model can be trained by the server based on sample data. By way of example, the server obtains a sample data set, and the server trains the recommendation model based on the sample data set. Among them, the sample data set includes multiple sample data, each sample data corresponds to a standard prediction data, and the standard prediction data corresponding to each sample data can be the annotation value (or called annotation data) of the sample data, and the standard prediction data corresponding to each sample data can be obtained by manually or by a computer annotating the sample data. Specifically, for each sample data in the sample data set, the server extracts multiple sample features of the sample data, each sample feature in the multiple sample features has an identifier, and the sample features with the same identifier correspond to the same model parameter of the recommendation model. The multiple sample features include user features and content features. The server inputs the user features and the content features into the recommendation model to be trained, so that the recommendation model to be trained calculates based on the user features and the content features to obtain the prediction data corresponding to the sample data, and the prediction data is used to characterize the correlation between the user features and the content features. For the convenience of description, the prediction data calculated by the recommendation model based on the user features and the content features during the training process is called training prediction data. For the training prediction data corresponding to each sample data, the server calculates the loss value of the training prediction data (the loss value of the training prediction data corresponding to each sample data is the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data). The server determines the gradient value corresponding to the model parameter according to the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. Furthermore, the server determines the average value of the gradient values (i.e., the average gradient value) of the model parameters corresponding to the sample features with the same identifier (the sample features with the same identifier correspond to the same model parameter), and the server adjusts the parameter value of the model parameter based on the average gradient value to train the recommendation model. The server trains the recommendation model based on each sample data in the sample data set until the convergence condition is reached, and the server ends the training process of the recommendation model.Among them, the convergence condition may include at least one of the following: the number of training times reaches a preset number, the loss value of the training prediction data obtained during consecutive multiple training processes fluctuates less, and the loss value of the training prediction data obtained during consecutive multiple training processes is less than a preset loss value. Herein, one training process is also referred to as one iteration process or one training step (step).
[0046] During the process of training the recommendation model, after the server determines the gradient value corresponding to the model parameters corresponding to the sample features, the server first stores the gradient value corresponding to the model parameters in the storage space of the server. The server determines the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier based on the gradient value corresponding to the model parameters stored in the storage space of the server. Among them, the storage space of the server includes multiple gradient storage units, and these gradient storage units are used to store the gradient values corresponding to the model parameters. By way of example, for each sample data in the sample dataset, after the server obtains the training prediction data corresponding to the sample data, the server calculates the loss value of the training recommendation data, the server calculates the hash value of each sample feature of the sample data, and, the server determines the gradient value corresponding to the model parameters based on the loss value of the training prediction data and the initial value of the model parameters corresponding to each sample feature of the sample data. The server establishes a mapping relationship between the hash value of each sample feature of the sample data in the sample dataset and the gradient storage address, and the server stores the gradient value corresponding to the model parameters corresponding to the sample feature to the corresponding gradient storage address (that is, stores it to the gradient storage unit indicated by the corresponding gradient storage address) based on the mapping relationship. The server obtains the gradient value stored in the gradient storage address corresponding to the hash value of the sample features with the same identifier based on the mapping relationship (that is, the gradient value stored in the gradient storage unit indicated by the gradient storage address), and the server calculates the average value of the gradient values stored in the gradient storage address corresponding to the hash value of the sample features with the same identifier, and adjusts the parameter value of the model parameters corresponding to the sample features with the same identifier based on the average value. During the process of establishing the above mapping relationship, in order to avoid hash collisions, for sample features with the same hash value, the server further determines the key codes of these sample features and calculates the hash values of the key codes of these sample features, and the server establishes the above mapping relationship based on the hash values of the key codes of these sample features. When the server obtains the gradient value based on the mapping relationship, the server first determines the gradient storage address according to the hash value of the sample feature. For sample features with the same hash value, the server further determines the gradient storage address according to the hash value of the key codes of these sample features, and then the server obtains the gradient value stored in the determined gradient storage address. However, since for sample features with the same hash value, the server needs to further search the above mapping relationship according to the hash value of the key code to determine the gradient storage address, this results in low data processing efficiency during the process of training the recommendation model. Among them, the process of establishing the above mapping relationship is also called the hashmap process. The above mapping relationship can be a hash table, and the content stored in the hash table is key-value pairs. The key field in the key-value pair is used to record the hash value, and the value field is used to record the gradient storage address.
[0047] Among them, the sample data can be natural language data, and the sample features can be vectorized (embedding) features, that is, the sample features can be vectors. Natural language data is data represented in natural language. For example, a sentence or a paragraph of text is natural language data. Vectorized features are vectors obtained by extracting features from natural language data through the embedding technology. Vectors can not only reflect the differences between features (such as sample features), but also facilitate the determination of the distance between features (such as sample features). The vectors obtained by extracting features from natural language data through the embedding technology can also be called dense vectors. For example, as Figure 1 shown, the server ( Figure 1 not shown) includes a vectorization module. The vectorization module extracts features from sample data 1 through the embedding technology to obtain user feature 11 and content feature 12. Both user feature 11 and content feature 12 are vectors. The vectorization module inputs user feature 11 and content feature 12 into the recommendation model to be trained. The recommendation model to be trained performs feature cross calculation on user feature 11 and content feature 12 to obtain prediction data 1 corresponding to sample data 1. Prediction data 1 is used to characterize the correlation degree between user feature 11 and content feature 12.
[0048] Since in the fields of e-commerce, video, news, etc., the types and quantities of sample data are huge, and the number of sample features can even reach more than tens of billions, therefore, for a large amount of sample data, how to improve the training efficiency of the recommendation model has always been an important issue concerned by the industry.
[0049] Next, the technical solution of this application will be introduced. First, the implementation environment involved in the embodiments of this application will be introduced.
[0050] As an example, Figure 2 is a schematic diagram of an implementation environment provided by an embodiment of this application. The implementation environment includes server 10 and client 20. Client 20 is communicatively connected to server 10. For example, client 20 and server 10 are communicatively connected through a local area network, the Internet, or other networks. Client 20 can send a sample data set to server 10, and server 10 executes the data processing method provided by the embodiments of this application based on the sample data set sent by client 20 to train the recommendation model. In Figure 2 the shown implementation environment, server 10 can implement the data processing method provided by the embodiments of this application by running an executable program. For example, the executable program of this data processing method can be presented in the form of an application installation package. After server 10 installs this application installation package, it can implement this data processing method by running this executable program.
[0051] As another example, Figure 3It is a schematic diagram of another implementation environment provided by an embodiment of the present application. This implementation environment includes a server cluster 30 and a client 20. The client 20 is communicatively connected to the server cluster 30. For example, the client 20 and the server cluster 30 are communicatively connected through a network such as a local area network or the Internet. The client 20 can send a sample data set to the server cluster 30, and the server cluster 30 executes the data processing method provided by the embodiment of the present application based on the sample data set sent by the client 20 to train a recommendation model. As Figure 3 shown, the server cluster 30 includes multiple servers ( Figure 3 taking 4 servers as an example), the client 20 can send all or part of the sample data in the sample data set to each server in the server cluster 30, and each server in the server cluster 30 executes the data processing method provided by the embodiment of the present application based on the sample data sent by the client 20 to train a recommendation model. In Figure 3 the shown implementation environment, the server cluster 30 can implement the data processing method provided by the embodiment of the present application by running an executable program. For example, each server in the server cluster 30 implements the data processing method provided by the embodiment of the present application by running an executable program. By way of example, the executable program of this data processing method is presented in the form of an application installation package, and after the server installs this application installation package, it can implement this data processing method by running this executable program. In some embodiments, each server in the server cluster 30 is also referred to as a node or a service node.
[0052] In one implementation, the client 20 can be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device or a smart vehicle-mounted device, etc. The server cluster 30 can be a cloud computing service center. In the embodiment of the present application, the server is a server for model training, and the server can be used for the management of model parameters. For example, the server is an artificial intelligence (AI) server, and the server can also be referred to as a parameter server (PS).
[0053] For example, during the process of the server training the recommendation model based on the sample data in the sample dataset, the server obtains the recommendation model to be trained, extracts the sample features of the sample data in the sample dataset, and trains the recommendation model to be trained according to the sample features of the sample data in the sample dataset. Since the sample features required by the server are highly sparse during each training process, that is, the server usually only uses the sample features of some of the sample data in the sample dataset. Therefore, during each training process, the server initializes the model parameters to be adjusted during this training process (that is, the model parameters corresponding to the sample features required during this training process) based on the initial values of the model parameters corresponding to the sample features required during this training process. The initialized recommendation model is also the recommendation model to be trained during this training process. During each training process, the server inputs the user features and content features of each sample data required during this training process into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user features and the content features to obtain the predicted data corresponding to the sample data (that is, the training prediction data). The server determines the gradient value (gradients) corresponding to the model parameters based on the training prediction data corresponding to the sample data and the initial values of the model parameters corresponding to each sample feature of the sample data, and the server determines the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier. The server adjusts the parameter values of the model parameters based on the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier.
[0054] As can be seen from the above description, the process of training the recommendation model can be executed by one server or multiple servers. When the process of training the recommendation model is executed by multiple servers, the multiple servers can adopt a distributed parallel training method.
[0055] As an example, Figure 4 is a schematic diagram of a method for training a recommendation model provided by an embodiment of the present application. Figure 4 Taking the case where the process of training the recommendation model is executed by one server as an example. The server can be Figure 2 the server 10 in the shown implementation environment, or it can be Figure 3 any server in the server cluster 30 provided by the shown implementation environment. As Figure 4As shown, the server includes a parameter management module and multiple execution modules. The execution module is also called a worker. The execution module is responsible for processes such as loading sample data, forward calculation (calculating the loss value of the training prediction data), and backward calculation (calculating the gradient value corresponding to the model parameters) during the model training process. For example, the execution module is responsible for interacting with the client to receive the sample data sent by the client (i.e., loading the sample data). During each training process, for each sample data received by the execution module that is required for the current training process, the execution module inputs the user features of the sample data and the content features of the sample data into the recommendation model to be trained, so that the recommendation model to be trained calculates based on the user features and the content features to obtain the training prediction data corresponding to the sample data. The execution module calculates the loss value of the training prediction data (i.e., the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data), and, the execution module determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The execution module sends the gradient values corresponding to the model parameters corresponding to the respective sample features of the sample data to the parameter management module. That is, the execution module is used to execute the process of calculating the training prediction data corresponding to the sample data based on the recommendation model to be trained and calculating the gradient values corresponding to the model parameters corresponding to the respective sample features of the sample data. For example, as Figure 4As shown, sample data 1 includes user feature A and content feature B, sample data 2 includes user feature C and content feature B, sample data 3 includes user feature A and content feature D, and sample data 4 includes user feature C and content feature E. Assume that the model parameters of the recommendation model to be trained include model parameter A, model parameter B, model parameter D, and model parameter E, and user feature A corresponds to model parameter A, content feature B corresponds to model parameter B, user feature C corresponds to model parameter C, content feature D corresponds to model parameter D, and content feature E corresponds to model parameter E. Execution module 1 calculates the training prediction data corresponding to sample data 1 through the recommendation model to be trained, based on user feature A of sample data 1 and content feature B of sample data 1. Execution module 1 calculates the loss value of the training prediction data corresponding to sample data 1. Execution module 1 determines gradient value A1 corresponding to model parameter A based on the loss value of the training prediction data corresponding to sample data 1 and the initial value corresponding to model parameter A, and, execution module 1 determines gradient value B1 corresponding to model parameter B based on the loss value of the training prediction data corresponding to sample data 1 and the initial value corresponding to model parameter B. Execution module 2 calculates the training prediction data corresponding to sample data 2 through the recommendation model to be trained, based on user feature C of sample data 2 and content feature B of sample data 2. Execution module 2 calculates the loss value of the training prediction data corresponding to sample data 2. Execution module 2 determines gradient value C1 corresponding to model parameter C based on the loss value of the training prediction data corresponding to sample data 2 and the initial value corresponding to model parameter C, and, execution module 2 determines gradient value B2 corresponding to model parameter B based on the loss value of the training prediction data corresponding to sample data 2 and the initial value corresponding to model parameter B. Execution module 3 calculates the training prediction data corresponding to sample data 3 through the recommendation model to be trained, based on user feature A of sample data 3 and content feature D of sample data 3. Execution module 3 calculates the loss value of the training prediction data corresponding to sample data 3. Execution module 3 determines gradient value A2 corresponding to model parameter A based on the loss value of the training prediction data corresponding to sample data 3 and the initial value corresponding to model parameter A, and, execution module 3 determines gradient value D1 corresponding to model parameter D based on the loss value of the training prediction data corresponding to sample data 3 and the initial value corresponding to model parameter D. Execution module 4 calculates the training prediction data corresponding to sample data 4 through the recommendation model to be trained, based on user feature C of sample data 4 and content feature E of sample data 4. Execution module 4 calculates the loss value of the training prediction data corresponding to sample data 4. Execution module 4 determines gradient value C2 corresponding to model parameter C based on the loss value of the training prediction data corresponding to sample data 4 and the initial value corresponding to model parameter C, and, execution module 4 determines gradient value E1 corresponding to model parameter E based on the loss value of the training prediction data corresponding to sample data 4 and the initial value corresponding to model parameter E.Among them, when the execution module sends the gradient value corresponding to the model parameter corresponding to the sample feature to the parameter management module, it also sends the identifier of the sample feature to the parameter management module. For example, the execution module simultaneously sends the gradient value corresponding to the model parameter corresponding to the sample feature and the identifier of the sample feature to the parameter management module. The parameter management module is used to receive and store the gradient values sent by the multiple execution modules, determine the average value of the gradient values of the model parameters corresponding to the sample features with the same identifier according to the stored gradient values, and adjust the parameter values of the model parameters according to the average value of the gradient values of the model parameters corresponding to the sample features with the same identifier.
[0056] As another example, Figure 5 is a schematic diagram of another method for training a recommendation model provided by an embodiment of the present application. Figure 5 Taking the process of training the recommendation model as an example, which is executed by two servers, server 1 and server 2. These two servers can be Figure 2 any two servers in the server cluster 30 provided by the shown implementation environment. As Figure 5 shown, server 1 and server 2 respectively include a parameter management module and multiple execution modules. The execution module is responsible for interacting with the client to receive the sample data sent by the client. In each training process, for each sample data received by the execution module that is required for the current training process, the execution module inputs the user feature of the sample data and the content feature of the sample data into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the user feature and the content feature to obtain the training prediction data corresponding to the sample data. The execution module calculates the loss value of the training prediction data (that is, the loss value of the training prediction data compared to the standard prediction data corresponding to the sample data), and, the execution module determines the gradient value corresponding to the model parameter based on the loss value of the training prediction data and the initial value of the model parameter corresponding to each sample feature of the sample data. The execution module sends the gradient values corresponding to the model parameters corresponding to the respective sample features of the sample data to the parameter management module in the server where the execution module is located. That is, the execution module is used to execute the process of calculating the training prediction data corresponding to the sample data based on the recommendation model to be trained and calculating the gradient values corresponding to the model parameters corresponding to the respective sample features of the sample data. By way of example, as Figure 5As shown in the figure, each execution module in Server 1 sends the gradient value corresponding to the model parameter corresponding to the calculated sample feature to Parameter Management Module 1, and each execution module in Server 2 sends the gradient value corresponding to the model parameter corresponding to the calculated sample feature to Parameter Management Module 2. When each execution module in Server 1 sends the gradient value corresponding to the model parameter corresponding to the sample feature to Parameter Management Module 1, it also sends the identifier of the sample feature to Parameter Management Module 1. When each execution module in Server 2 sends the gradient value corresponding to the model parameter corresponding to the sample feature to Parameter Management Module 2, it also sends the identifier of the sample feature to Parameter Management Module 2. For example, the execution module sends the gradient value corresponding to the model parameter corresponding to the sample feature and the identifier of the sample feature to the corresponding parameter management module at the same time. Parameter Management Module 1 is used to receive and store the gradient values sent by multiple execution modules in Server 1, and Parameter Management Module 2 is used to receive and store the gradient values sent by multiple execution modules in Server 2. As Figure 5 shown, the gradient values stored in Parameter Management Module 1 and Parameter Management Module 2 are also synchronized with each other. In this way, each parameter management module in Parameter Management Module 1 and Parameter Management Module 2 can obtain the gradient values stored in the other parameter management module. After the gradient values are synchronized between Parameter Management Module 1 and Parameter Management Module 2, at least one of Parameter Management Module 1 and Parameter Management Module 2 determines the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier (i.e., the average gradient value), and adjusts the parameter value of the model parameter according to the average value of the gradient values corresponding to the model parameters corresponding to the sample features with the same identifier.
[0057] In the embodiment of the present application, the sample data sent by the client to the server may be natural language data. The server may first perform feature extraction on the sample data through the embedding technology to obtain multiple sample features of the sample data. The multiple sample features include user features and content features, and the multiple sample features are all vectors. By way of example, the server further includes a vectorization module. The vectorization module is used to interact with the client to receive the sample data sent by the client, and the vectorization module performs feature extraction on the sample data through the embedding technology to obtain multiple sample features of the sample data. In the embodiment of the present application, the server trains the recommendation model based on the vectorized sample features, which can improve the efficiency of training the recommendation model. Therefore, in some embodiments, the vectorization module for extracting sample features may also be referred to as an acceleration card.
[0058] Figure 2 and Figure 3 is an exemplary description of the implementation environment of the present application, and does not constitute a limitation on the implementation environment of the present application. With the change of business requirements, the implementation environment of the present application can be adjusted according to business requirements. In addition, Figure 4 andFigure 5 This is an exemplary description of training a recommendation model and does not constitute a limitation on training the recommendation model. For example Figure 5 The figure shows a schematic diagram of training a recommendation model using data parallelism. It is also possible to train the recommendation model using model parallelism. Among them, data parallelism means: deploying the same recommendation model to be trained on multiple servers, and dispersing the sample data in the sample dataset for training to these multiple servers. Each server independently uses a part of the sample data in the sample dataset for model training. During each training process, each server synchronizes the calculated gradient values to other servers, so that each server can calculate the average gradient value of one training process. The average gradient value is the average of the gradient values obtained by each server during the same training process. Each server can adjust the parameter values of the model parameters of the recommendation model according to the average gradient value. Model parallelism means: splitting the recommendation model into multiple sub-model parts, deploying different sub-model parts on different servers, and during training, according to the structure order of the recommendation model, the sub-model parts on different servers calculate the same sample data in the sample dataset. During one training process, gradient values can be obtained through the calculation of the sub-model parts on different servers, and the parameters of each sub-model part can be updated by forward propagation or backward propagation according to the gradient values. After multiple iterative trainings, after meeting the training requirements, the sub-model parts on each server are recombined according to the structure of the recommendation model, and a trained recommendation model can be obtained.
[0059] The above is the introduction to the implementation environment of this application. Next, the method embodiments of this application will be introduced.
[0060] Please refer to Figure 6 , which shows a flowchart of a data processing method provided by an embodiment of this application. This data processing method is executed by a server. The server can be Figure 2 the server 10 in the implementation environment shown or any server in the implementation environment shown Figure 3 . Referring to Figure 6 , this method includes the following steps S601 to S605.
[0061] S601. Determine the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter of the recommendation model. The first prediction data is calculated by the recommendation model based on the first user feature and the first content feature. The identifier of the first sample feature in the first user feature and the first content feature is the first identifier, and the first model parameter corresponds to the first identifier.
[0062] Among them, the first user feature and the first content feature are both sample features of the first sample data. The first sample feature is any one of the first user feature and the first content feature. For example, the first sample feature is the first user feature, or the first sample feature is the first content feature. The first sample data is any one of the sample data sets used to train the recommendation model. The first sample data may be natural language data, and the first user feature and the first content feature may both be vectors. Optionally, the server performs feature extraction on the first sample data through the embedding technology to obtain multiple sample features of the first sample data. The multiple sample features are all vectors. The multiple sample features include the first user feature and the first content feature. Each of the multiple sample features has an identifier, and the identifiers of different sample features in the multiple sample features may be the same or different. The identifiers of sample features of different sample data may be the same or different. Sample features with the same identifier correspond to the same model parameters of the recommendation model. For example, multiple sample features of sample data 1 include user feature A and content feature B, multiple sample features of sample data 3 include user feature A and content feature D, the identifier of user feature A of sample data 1 and the identifier of user feature A of sample data 3 are both "01", the identifier of user feature A of sample data 1 and the identifier of user feature A of sample data 3 both correspond to model parameter 1 of the recommendation model, the identifier of content feature B is "02", content feature B corresponds to model parameter 2 of the recommendation model, the identifier of content feature D is "04", content feature D corresponds to model parameter 4 of the recommendation model. Optionally, the server includes a vectorization module, which extracts features from the first sample data by using an embedding technology.
[0063] In an embodiment of the present application, a recommendation model is used to recommend data content to a user based on the correlation between user features and content features. For example, the recommendation model is used to recommend data content with specific content features to a user with specific user features, and the specific user features match the specific content features. For example, the correlation between the specific user features and the specific content features is greater than a preset threshold. Among them, the first prediction data is calculated by the recommendation model based on the first user features and the first content features, and the first prediction data is used to characterize the correlation between the first user features and the first content features (that is, the predicted correlation). The loss value of the first prediction data is the loss value of the first prediction data compared to the standard prediction data corresponding to the first sample data. The standard prediction data corresponding to the first sample data can be the annotation value of the first sample data (or called annotation data), and the standard prediction data corresponding to the first sample data can be obtained by manually or by computer annotating the first sample data. The standard prediction data corresponding to the first sample data can be used to characterize the standard correlation between the first user features and the first content features (that is, the annotated correlation). Optionally, after the server obtains the first user feature and the first content feature, the server inputs the first user feature and the first content feature into the recommendation model to be trained, so that the recommendation model to be trained performs calculations based on the first user feature and the first content feature to obtain first prediction data (i.e., training prediction data). The server may use a loss function to calculate the loss value of the first prediction data. The server may use a gradient calculation formula to calculate the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter. The gradient calculation formula may be an expression of a derivative function of the loss function, the independent variables of the gradient calculation formula are the loss value of the prediction data and the initial value of the model parameter, and the dependent variable of the gradient calculation formula is the gradient value corresponding to the model parameter.
[0064] In an embodiment of the present application, the recommendation model to be trained may be a recommendation model obtained after initializing each model parameter of the recommendation model according to the initial value of the model parameter. Since the sample features required by the server in each training process are highly sparse, and the model parameters have a corresponding relationship with the identifiers of the sample features, in each training process, the server may initialize the model parameters based on the initial values of the model parameters corresponding to the sample features required in this training process (that is, the model parameters corresponding to the identifiers of the sample features required in this training process), that is, update the parameter values of the model parameters to the initial values.
[0065] In an optional embodiment, the server includes a parameter storage space, which includes a plurality of parameter storage units, each parameter storage unit is used to store a parameter value of a model parameter (or each parameter storage unit corresponds to a model parameter), each parameter storage unit corresponds to a sample feature identifier (i.e., an identifier of a sample feature), and each parameter storage unit is used to store the parameter value of the model parameter corresponding to the sample feature with the corresponding identifier. In each training process, for each model parameter that needs to be adjusted in this training process (i.e., each model parameter corresponding to the sample feature required for this training process), the server obtains the parameter value of the model parameter from the parameter storage space, and the server determines the parameter value of the model parameter obtained from the parameter storage space as the initial value of the model parameter, and the server uses the initial value of the model parameter to initialize the model parameter. Among them, the parameter value stored in each parameter storage unit can be a parameter value set manually or by a computer, or it can be a parameter value updated to the parameter storage unit in the previous training process (i.e., the server uses the parameter value of the model parameter calculated in the previous training process as the initial value of the model parameter in the next training process). For example, the first parameter storage unit in the parameter storage space is used to store the parameter value of the first model parameter, the first parameter storage unit corresponds to the first identifier (the identifier of the first sample data), the server determines the first parameter storage unit according to the identifier of the first sample data (that is, the first identifier), the server obtains the parameter value stored in the first parameter storage unit, the server determines the parameter value obtained from the first parameter storage unit as the initial value of the first model parameter, and the server initializes the first model parameter with the initial value of the first model parameter.
[0066] For an example, see Figure 7 and Figure 8 , Figure 7 and Figure 8 All of them are schematic diagrams of the storage space of the server provided in the embodiments of the present application. Figure 7 and Figure 8 As shown, the storage space of the server includes a parameter storage space, and the parameter storage space includes a plurality of parameter storage units ( Figure 7 and Figure 8Take the multiple parameter storage units as parameter storage units 1 to 5 as an example. Each of the multiple parameter storage units is used to store a parameter value of a model parameter, and each parameter storage unit corresponds to a sample feature identifier. For example, parameter storage unit 1 corresponds to the sample feature identifier "01" and is used for the parameter value of model parameter 1 (model parameter 1 corresponds to the sample feature identifier "01"), parameter storage unit 2 corresponds to the sample feature identifier "02" and is used for the parameter value of model parameter 2 (model parameter 2 corresponds to the sample feature identifier "02"), parameter storage unit 3 corresponds to the sample feature identifier "03" and is used for the parameter value of model parameter 3 (model parameter 3 corresponds to the sample feature identifier "03"), and so on. For example, the first model parameter is model parameter 1, and the first identifier is the sample feature identifier "01". The server determines that the first parameter storage unit is parameter storage unit 1 based on the first identifier (that is, the sample feature identifier "01"), and the server obtains the parameter value stored in parameter storage unit 1. The server determines the parameter value obtained from parameter storage unit 1 as the initial value of model parameter 1, and the server initializes model parameter 1 using the initial value of model parameter 1.
[0067] S602. Determine a first gradient storage address according to a first identifier and a mapping relationship, wherein the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter.
[0068] After the server determines the first gradient value corresponding to the first model parameter, since the first model parameter corresponds to the identifier of the first sample feature (ie, the first identifier), the server determines the first gradient storage address according to the first identifier and the mapping relationship.
[0069] Among them, the first gradient storage address is used to indicate the first gradient storage unit, and the first gradient storage unit stores the accumulated gradient value corresponding to the first model parameter, and the accumulated gradient value corresponding to the first model parameter is the accumulated value of the gradient value corresponding to the first model parameter. Optionally, the first gradient storage unit is also used to store the first identifier and the accumulated number of times corresponding to the first model parameter. The accumulated number of times corresponding to the first model parameter is the accumulated number of times the gradient value corresponding to the first model parameter, and may also be referred to as the accumulated number of times corresponding to the first identifier, which is not limited in the embodiments of the present application. For example, the first gradient storage unit includes a sample feature identifier field, an accumulated number of times field, and an accumulated gradient value field, the first identifier is stored in the sample feature identifier field, the accumulated number of times is stored in the accumulated number of times field, and the accumulated gradient value is stored in the accumulated gradient value field.
[0070] Wherein, the mapping relationship is a mapping relationship between the identifier of the sample feature (referred to as the sample feature identifier) and the gradient storage address. Optionally, the mapping relationship includes a first mapping relationship and a second mapping relationship. The first mapping relationship is a mapping relationship between the sample feature identifier and the index storage address, and the second mapping relationship is a mapping relationship between the index storage address and the index information, and the index information may include the gradient storage address, thereby, the mapping relationship between the sample feature identifier and the gradient storage address is realized by the first mapping relationship and the second mapping relationship. For example, the first mapping relationship includes a one-to-one mapping relationship between multiple sample feature identifiers and multiple index storage addresses, and the second mapping relationship includes a one-to-one mapping relationship between the multiple index storage addresses and multiple index information, each index storage address in the multiple index storage addresses is used to indicate an index storage unit, and each index storage unit is used to store the index information corresponding to the corresponding sample feature identifier. Wherein, the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, and the second mapping relationship includes a mapping relationship between the first index storage address and the first index information, and the first index storage address is used to indicate the first index storage unit, and the first index storage unit is used to store the index information corresponding to the first identifier.
[0071] In an optional embodiment, the server determines the first index storage address based on the first identifier and the first mapping relationship, and the server determines the first gradient storage address based on the first index storage address and the second mapping relationship. Specifically, the first mapping relationship is a one-to-one correspondence between the sample feature identifier and the gradient storage address, and the server searches for the first mapping relationship based on the first identifier. The server determines the index storage address corresponding to the first identifier in the first mapping relationship as the first index storage address based on the search result. The server determines the index unit indicated by the first index storage address as the first index storage unit, and the server obtains the index information stored in the first index storage unit (that is, the index information corresponding to the first identifier), and the server determines the first gradient storage address based on the index information stored in the first index storage unit.
[0072] In an example, the first mapping relationship is shown in Table 1 below.
[0073] Table 1 (first mapping relationship)
[0074] Sample Feature Identification Index Storage Address 01 Index Storage Address 1 02 Index Storage Address 2 03 Index Storage Address 3 04 Index Storage Address 4 …… ……
[0075] Referring to Table 1, the first mapping relationship includes the correspondence between the sample feature identifier "01" and the index storage address 1, the correspondence between the sample feature identifier "02" and the index storage address 2, the correspondence between the sample feature identifier "03" and the index storage address 3, the correspondence between the sample feature identifier "04" and the index storage address 4, and so on.
[0076] Please continue to refer to Figure 7 andFigure 8 The storage space of the server also includes an index storage space, which includes a plurality of index storage units ( Figure 7 and Figure 8 Taking the multiple index storage units as index storage units 1 to 4 as an example, each of the multiple index storage units is used to store index information corresponding to the corresponding sample feature identifier. The index information includes at least one of state information, space identifier, and gradient storage address. The state information is used to indicate the storage state corresponding to the corresponding sample feature identifier. The gradient storage address is used to indicate the gradient storage unit. The space identifier is used to indicate the gradient storage space where the gradient storage unit indicated by the gradient storage address is located. For example, each of the multiple index storage units includes a state information field, a space identifier field, and a gradient storage address field, the state information field is used to store state information, the space identifier field is used to store the space identifier of the gradient storage space, and the gradient storage address field is used to store the gradient storage address. For example, index storage address 1 is used to indicate index storage unit 1, and index storage unit 1 is used to store index information corresponding to sample feature identifier "01". Index storage address 2 is used to indicate index storage unit 2, and index storage unit 2 is used for index information corresponding to sample feature identifier "02". Index storage address 3 is used to indicate index storage unit 3, and index storage unit 3 is used to store index information corresponding to sample feature identifier "03". Index storage address 4 is used to indicate index storage unit 4, and index storage unit 4 is used to store index information corresponding to sample feature identifier "04". And so on. Assuming that the first identifier is the sample feature identifier "01", the server searches for the first mapping relationship shown in Table 1 according to the first identifier to determine the index storage address 1 corresponding to the first identifier. The server determines index storage address 1 as the first index storage address, and then determines index storage unit 1 indicated by index storage address 1 as the first index storage unit.
[0077] In an optional embodiment, the server determines the first gradient storage address based on the index information stored in the first index storage unit, including: the server determines whether the index information stored in the first index storage unit includes the gradient storage address; when the index information stored in the first index storage unit includes the gradient storage address, the server determines the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include the gradient storage address, the server determines the first free storage unit in the pre-created first storage space as the first gradient storage unit, or the server creates a second storage space and determines the second free storage unit in the second storage space as the first gradient storage unit, and the server determines the address of the first gradient storage unit as the first gradient storage address, wherein the first storage space and the second storage space are both gradient storage spaces, and the first storage space and the second storage space respectively include at least one storage unit, and the at least one storage unit is both a gradient storage unit.
[0078] In an optional embodiment, when the index information stored in the first index storage unit by the server does not include the gradient storage address, the server determines whether there is a free storage unit in the pre-created first storage space. If there is a free storage unit in the first storage space, the server determines the first free storage unit in the first storage space as the first gradient storage unit, and the first free storage unit is any free storage unit in the first storage space. If there is no free storage unit in the first storage space, the server creates a second storage space, the second storage space includes at least one storage unit, and the server determines the second free storage unit in the second storage space as the first gradient storage unit, and the second free storage unit is any free storage unit in the second storage space. In one embodiment, the server sequentially traverses each gradient storage unit in the first storage space to determine whether there is a free storage unit in the first storage space. In another embodiment, the server records the identifiers of each free storage unit in the first storage space, and the server determines whether there is a free storage unit in the first storage space according to the identifiers of the free storage units recorded by the server. In another embodiment, the server records the number of busy storage units in the first storage space and the total number of storage units in the first storage space, and the server determines whether there are free storage units in the first storage space based on the number of busy storage units in the first storage space and the total number of storage units in the first storage space. For example, if the number of busy storage units in the first storage space is less than the total number of storage units in the first storage space, the server determines that there are free storage units in the first storage space. If the number of busy storage units in the first storage space is equal to the total number of storage units in the first storage space, the server determines that there are no free storage units in the first storage space. In an optional embodiment, if there are no free storage units in the first storage space, the server divides a storage area in the storage space of the server as a second storage space, and divides a plurality of gradient storage units in the second storage space, which is not limited in the embodiments of the present application.
[0079] In an optional embodiment, the index information stored in the first index storage unit includes state information, and the state information is used to indicate that the storage state corresponding to the first identifier is an initial state or an accumulated state. For example, when the state information is initial state information, the state information is used to indicate that the storage state corresponding to the first identifier is an initial state. When the state information is accumulated state information, the state information is used to indicate that the storage state corresponding to the first identifier is an accumulated state. For example, the initial state information is 0, and the accumulated state information is any non-0 information, for example, the accumulated state information is 1. The server can determine the storage state corresponding to the first identifier based on the state information stored in the first index storage unit. Specifically, when the state information stored in the first index storage unit by the server is used to indicate that the storage state corresponding to the first identifier is an initial state, it is determined that the index information stored in the first index storage unit does not include a gradient storage address. When the state information stored in the first index storage unit by the server is used to indicate that the storage state corresponding to the first identifier is an accumulated state, it is determined that the index information stored in the first index storage unit includes a gradient storage address. In an optional embodiment, the index information stored by the server in the first index storage unit does not include the gradient storage address, and after the server determines the address of the first gradient storage unit as the first gradient storage address, the server stores the first gradient storage address in the first index storage unit, and the server updates the status information stored in the first index storage unit, and the updated status information is used to indicate that the storage status corresponding to the first identifier is a cumulative state.
[0080] In one embodiment, please continue to refer to Figure 7 and Figure 8 , the first index storage unit is index storage unit 1, the index information stored in index storage unit 1 includes state information, and the state information is cumulative state information, the server determines that the storage state corresponding to the first identifier is a cumulative state according to the state information, and further determines that the index information stored in index storage unit 1 includes a gradient storage address (for example, the server determines Figure 7 and Figure 8 The gradient storage address field in the index storage unit 1 stores a gradient storage address), and the server determines the gradient storage address stored in the index storage unit 1 as the first gradient storage address.
[0081] In another embodiment, please continue to refer to Figure 7 and Figure 8 , the first index storage unit is index storage unit 1, the index information stored in index storage unit 1 includes state information, and the state information is initial state information, the server determines that the storage state corresponding to the first identifier is the initial state according to the state information, and further determines that the index information stored in index storage unit 1 does not include the gradient storage address (for example, the server determines Figure 7 andFigure 8 (the gradient storage address is not stored in the gradient storage address field in the index storage unit 1 shown). Assume that the first storage space is the gradient storage space 1. When the server determines that the index information stored in the index storage unit 1 does not include the gradient storage address, the server determines whether there is an idle storage unit in the gradient storage space 1. In one example, as Figure 7 shown, when the server determines that there is an idle storage unit in the gradient storage space 1, the server determines the first idle storage unit in the gradient storage space 1 as the first gradient storage unit. Taking the first idle storage unit being the gradient storage unit 11 as an example, that is, the server determines the gradient storage unit 11 as the first gradient storage unit. After that, the server determines the address of the gradient storage unit 11 as the first gradient storage address, stores the first gradient storage address into the index storage unit 1 (specifically, stores it into the gradient storage address field in the index storage unit 1), and the server updates the status information in the index storage unit 1 to the cumulative status information. Also, the server stores the space identifier of the gradient storage space 1 in the index storage unit 1 (specifically, stores the space identifier of the gradient storage space 1 in the space identifier field in the index storage unit 1). In another example, as Figure 8 shown, when the server determines that there is no idle storage unit in the gradient storage space 1, the server creates a second storage space and determines the second idle storage unit in the second storage space as the first gradient storage unit. Taking the second storage space being the gradient storage space 2 and the second idle storage unit being the gradient storage unit 21 as an example, that is, the server determines the gradient storage unit 21 as the first gradient storage unit. After that, the server determines the address of the gradient storage unit 21 as the first gradient storage address, stores the first gradient storage address into the index storage unit 1 (specifically, stores it into the gradient storage address field in the index storage unit 1), and the server updates the status information in the index storage unit 1 to the cumulative status information. Also, the server stores the space identifier of the gradient storage space 2 in the index storage unit 1 (specifically, stores the space identifier of the gradient storage space 2 in the space identifier field in the index storage unit 1). In the embodiments of the present application, the gradient storage space is located in a specific storage space of the server, and this specific storage space is used to create (or divide the gradient storage space). The total capacity of this specific storage space can be determined according to the total number of sample features of the sample data in the sample dataset. By way of example, the total capacity of this specific storage space is 2 n , where n is the total number of sample features of the sample data in the sample dataset.
[0082] S603. Add the first gradient value to the cumulative gradient value stored in the first gradient storage unit to obtain the first cumulative gradient value.
[0083] In an optional embodiment, the server determines the first gradient storage unit according to the first gradient storage address, the server obtains the accumulated gradient value stored in the first gradient storage unit, and the server adds the first gradient value to the accumulated gradient value stored in the first gradient storage unit to obtain the first accumulated gradient value. For example, the first gradient storage address is gradient storage address 1, and gradient storage address 1 is used to indicate gradient storage unit 11, that is, the first gradient storage unit is gradient storage unit 11. Figure 7 As shown, the gradient storage unit 11 is found according to the gradient storage address 1, and the gradient value in the gradient storage unit 11 and the accumulated gradient value stored in the gradient storage unit 11 are accumulated to obtain a first accumulated gradient value.
[0084] S604. Store the first accumulated gradient value in a first gradient storage unit.
[0085] In an optional embodiment, the server uses the first accumulated gradient value to update the accumulated gradient value stored in the first gradient storage unit. Specifically, the server deletes the accumulated gradient value stored in the first gradient storage unit, and then stores the first accumulated gradient value in the first gradient storage unit. Figure 7 , the first gradient storage unit is the gradient storage unit 11, the server deletes the accumulated gradient value stored in the gradient storage unit 11, and then stores the first accumulated gradient value in the gradient storage unit 11. In another example, referring to Figure 8 , the first gradient storage unit is the gradient storage unit 21 , the server deletes the accumulated gradient value stored in the gradient storage unit 21 , and then stores the first accumulated gradient value to the gradient storage unit 21 .
[0086] In an optional embodiment, the first gradient storage unit is also used to store the cumulative number of times corresponding to the first model parameter, the cumulative number of times corresponding to the first model parameter is the cumulative number of times the gradient value corresponding to the first model parameter is accumulated, or the cumulative number of times the gradient value corresponding to the first model parameter is accumulated in the first gradient storage unit. In the initial state, the cumulative number of times corresponding to the first model parameter is the initial number (for example, 0). After the server stores the first accumulated gradient value in the first gradient storage unit, the server also updates the cumulative number stored in the first gradient storage unit. For example, the server adds 1 to the cumulative number stored in the first gradient storage unit.
[0087] S605. Adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit.
[0088] In an optional embodiment, the server determines the average gradient value corresponding to the first identifier (that is, the average gradient value of the gradient values corresponding to the first model parameter) based on the accumulated gradient value stored in the first gradient storage unit, and the server adjusts the parameter value of the first model parameter based on the average gradient value corresponding to the first identifier. In one example, the server adjusts the parameter value of the first model parameter to the average gradient value corresponding to the first identifier. In another example, the server uses an optimizer to optimize the average gradient value corresponding to the first identifier to obtain an optimized gradient value, and the server adjusts the parameter value of the first model parameter to the optimized gradient value. Among them, the optimizer can be an adaptive moment estimation (Adam) optimizer or a stochastic gradient descent optimizer. The Adam algorithm used by the Adam optimizer is an algorithm that extends the stochastic gradient descent method.
[0089] In an optional embodiment, the first gradient storage unit also stores the cumulative number of times corresponding to the first model parameter, that is, the cumulative number of times corresponding to the first identifier, and the server determines the average gradient value corresponding to the first identifier based on the cumulative gradient value stored in the first gradient storage unit and the cumulative number of times stored in the first gradient storage unit. For example, the server determines the quotient of the cumulative gradient value stored in the first gradient storage unit and the cumulative number of times stored in the first gradient storage unit as the average gradient value corresponding to the first identifier.
[0090] In an optional embodiment, after the server adjusts the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit, the server clears the first gradient storage unit to restore the first gradient storage unit to an idle state. For example, the server deletes the first identifier, the accumulated number of times, and the accumulated gradient value stored in the first gradient storage unit to clear the first gradient storage unit. Since the first gradient storage unit is associated with the first index storage unit (the first gradient storage address for indicating the first gradient storage unit is stored in the first index storage unit), after the server clears the first gradient storage unit, the server also resets the first index storage unit. For example, the server deletes the index storage address stored in the first index storage unit and resets the state information stored in the first index storage unit to the initial state information.
[0091] In an optional embodiment, after the server adjusts the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier, the server stores the adjusted parameter value of the first model parameter in the first parameter storage unit. For example, the server first determines the first parameter storage unit according to the first identifier, and then stores the adjusted parameter value of the first model parameter in the first parameter storage unit. Specifically, the server deletes the parameter value stored in the first parameter storage unit, and then stores the adjusted parameter value of the first model parameter in the first parameter storage unit.
[0092] The cumulative gradient value stored in the first gradient storage unit described in S605 may be the first cumulative gradient value, or may be a cumulative gradient value obtained by adding other gradient values corresponding to the first model parameter on the basis of the first cumulative gradient value. Among them, the cumulative gradient value stored in the first gradient storage unit used by the server to adjust the parameter value of the first model parameter is the cumulative gradient value after adding the gradient values of the model parameters corresponding to the sample features with the same sample feature identifier. Among them, after the server obtains all the sample features, it can determine the number of sample features with the same sample feature identifier, and based on the number of this sample feature, determine the total number of cumulative times of the gradient value corresponding to the first model parameter, that is, when the server determines that the cumulative number of times in the first gradient storage unit reaches this total number of times, the gradient values of the model parameters corresponding to the sample features with the same sample feature identifier have all been cumulatively added. The embodiments of the present application do not limit this.
[0093] The embodiments of the present application illustrate by taking the first model parameter for training a recommendation model as an example. The recommendation model may include multiple model parameters, and these multiple model parameters correspond to different sample feature identifiers. The server can train the model parameters according to the sample features indicated by the sample feature identifiers corresponding to each model parameter. The training process of each model parameter among the multiple model parameters can refer to the training process of the first model parameter, and the embodiments of the present application will not elaborate here. In addition, the server may include a parameter management module and multiple execution modules. The above S601 may be executed by the execution module, and S602 to S605 may be executed by the parameter management module. By way of example, after the execution module determines the first gradient value corresponding to the first model parameter, the execution module first stores the first gradient value in the cache space of the execution module, and then the execution module sequentially sends the gradient values stored in the cache space to the parameter management module. When the execution module sends each gradient value to the parameter management module, it can also send the sample feature identifier corresponding to the gradient value (that is, the sample feature identifier corresponding to the model parameter corresponding to the gradient value) to the parameter management module. For example, the execution module sends the first identifier and the first gradient value to the parameter management module at the same time.
[0094] In summary, for the technical solution provided in the embodiment of the present application, since the mapping relationship is the mapping relationship between the identifier of the sample feature and the gradient storage address, and this mapping relationship includes the mapping relationship between the first identifier and the first gradient storage address. After the server determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the server can quickly determine the first gradient storage address according to the first identifier (here, it refers to the identifier of the first sample feature) and this mapping relationship, and then accumulate the first gradient value corresponding to the first model parameter into the accumulated gradient value stored at the first gradient storage address. When training the recommendation model, the server can quickly determine the first gradient storage address according to the first identifier and this mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address, and adjust the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address. Therefore, the data processing efficiency during the model training process can be improved, and thus the model training efficiency can be improved. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform the accumulation calculation of the gradient values when training the recommendation model, further improving the data processing efficiency and model training efficiency during the model training process.
[0095] The above is the introduction of the method embodiment of the present application. Next, the device embodiment of the present application will be introduced. The device of the present application is used to execute the method of the present application. For the details not disclosed in the device embodiment, please refer to the method embodiment.
[0096] Please refer to Figure 9 , which shows a schematic diagram of a data processing device 900 provided in an embodiment of the present application. The data processing device 900 may be a server or a functional component in the server. The data processing device 900 can implement all or part of the steps of the data processing method provided in the embodiment as shown in Figure 6 . As shown in Figure 9 , the data processing device 900 includes a first determination module 901, a second determination module 902, an accumulation module 903, a storage module 904, and an adjustment module 905.
[0097] The first determination module 901 is configured to determine the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter of the recommendation model. The first prediction data is calculated by the recommendation model based on the first user feature and the first content feature. The identifier of the first sample feature in the first user feature and the first content feature is the first identifier, and the first model parameter corresponds to the first identifier. The implementation process of the first determination module 901 may refer to the relevant description in step S601 above.
[0098] The second determination module 902 is configured to determine a first gradient storage address according to a first identifier and a mapping relationship. The mapping relationship includes the mapping relationship between the first identifier and the first gradient storage address. The first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store the accumulated gradient value corresponding to the first model parameter. The implementation process of the second determination module 902 may refer to the relevant description in step S602 above.
[0099] The accumulation module 903 is configured to accumulate the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value. The implementation process of the accumulation module 903 may refer to the relevant description in step S603 above.
[0100] The storage module 904 is configured to store the first accumulated gradient value into the first gradient storage unit. The implementation process of the storage module 904 may refer to the relevant description in step S604 above.
[0101] The adjustment module 905 is configured to adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit. The implementation process of the adjustment module 905 may refer to the relevant description in step S605 above.
[0102] Optionally, the second determination module 902 is configured to: determine a first index storage address according to the first identifier and a first mapping relationship. The first mapping relationship includes the mapping relationship between the first identifier and the first index storage address. The first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; determine the first gradient storage address according to the index information stored in the first index storage unit.
[0103] Optionally, the second determination module 902 is configured to: when the index information stored in the first index storage unit includes a gradient storage address, determine the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include a gradient storage address, determine the first free storage unit in the pre-created first storage space as the first gradient storage unit, or create a second storage space and determine the second free storage unit in the second storage space as the first gradient storage unit, and determine the address of the first gradient storage unit as the first gradient storage address, where the first storage space and the second storage space each include at least one storage unit.
[0104] Optionally, the storage module 904 is further configured to: after the second determination module 902 determines that the index information stored in the first index storage unit does not include a gradient storage address and determines the address of the first gradient storage unit as the first gradient storage address, store the first gradient storage address into the first index storage unit.
[0105] Optionally, the index information stored in the first index storage unit includes status information, which is used to indicate that the storage status corresponding to the first identifier is the initial state or the accumulation state. The second determination module 902 is further configured to: when the status information is used to indicate that the storage status corresponding to the first identifier is the initial state, determine that the index information stored in the first index storage unit does not include the gradient storage address; when the status information is used to indicate that the storage status corresponding to the first identifier is the accumulation state, determine that the index information stored in the first index storage unit includes the gradient storage address.
[0106] Optionally, the data processing device 900 further includes: an update module 906, configured to update the status information after storing the first gradient storage address into the first index storage unit when the status information is used to indicate that the storage status corresponding to the first identifier is the initial state.
[0107] Optionally, the adjustment module 905 is configured to: determine the average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; adjust the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier.
[0108] Optionally, both the first user feature and the first content feature are vectors.
[0109] Optionally, the data processing device 900 may be a server or a functional component in the server. The data processing device 900 includes a parameter management module and a plurality of execution modules. The above-mentioned first determination module 901 may be a sub-module in the execution module, and the second determination module 902, the accumulation module 903, the storage module 904, the adjustment module 905, and the update module 906 may all be sub-modules in the parameter management module.
[0110] In summary, for the technical solution provided in the embodiment of the present application, since the mapping relationship is the mapping relationship between the identifier of the sample feature and the gradient storage address, and this mapping relationship includes the mapping relationship between the first identifier and the first gradient storage address, after the first determination module determines the first gradient value corresponding to the first model parameter based on the loss value of the first prediction data and the initial value of the first model parameter, the second determination module can quickly determine the first gradient storage address according to the first identifier (here, it refers to the identifier of the first sample feature) and this mapping relationship. Then, the storage module accumulates the first gradient value corresponding to the first model parameter into the accumulated gradient value stored at the first gradient storage address. When the server trains this recommendation model, the second determination module in the server can quickly determine the first gradient storage address according to the first identifier and this mapping relationship, and then obtain the accumulated gradient value stored at the first gradient storage address. The adjustment module adjusts the parameter value of the first model parameter according to the accumulated gradient value stored at the first gradient storage address. Therefore, the data processing efficiency during the model training process can be improved, and thus the model training efficiency can be improved. In addition, since the first gradient storage address stores the accumulated gradient value corresponding to the first model parameter, rather than multiple gradient values corresponding to the first model parameter, the server does not need to perform the accumulation calculation of the gradient values when training this recommendation model, further improving the data processing efficiency and the model training efficiency during the model training process.
[0111] The embodiment of the present application provides a data processing device, including a memory and a processor. The memory is used to store a computer program. The processor is used to execute the computer program stored in the memory so that the data processing device executes some or all of the functions in the data processing method provided in the embodiment of the present application. The data processing device can be a server or a partial functional component in the server.
[0112] In the embodiment of the present application, the data processing device is taken as an example of a server for illustration. For example, Figure 10 is a schematic structural diagram of a server 100 provided in the embodiment of the present application. As Figure 10 shown, the server 100 includes a processor 101, a memory 102, a communication interface 103, and a bus 104. Among them, the processor 101, the memory 102, and the communication interface 103 are communicatively connected to each other through the bus 104.
[0113] The processor 101 may include a general-purpose processor and / or a dedicated hardware chip. The general-purpose processor may include: a central processing unit (CPU), a microprocessor, or a graphics processing unit (GPU). The CPU is, for example, a single-core processor (single-CPU), or a multi-core processor (multi-CPU). The dedicated hardware chip is a high-performance processing hardware module. The dedicated hardware chip includes at least one of a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a network processor (NP). The processor 101 may also be an integrated circuit chip with signal processing capabilities. During implementation, some or all of the functions of the data processing method of this application may be completed by the integrated logic circuit in the hardware of the processor 101 or instructions in software form.
[0114] The memory 102 is used to store computer programs, which include an operating system 102a and executable code (i.e., program instructions) 102b. The memory 102 is, for example, a read-only memory or other types of static storage devices that can store static information and instructions, or a random access memory or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory, a read-only optical disc, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. For example, the memory 102 is used to store the outbound port queue, etc. The memory 102 is, for example, independent and connected to the processor 101 through a bus 104. Or the memory 102 and the processor 101 are integrated together. The memory 102 can store executable code. When the executable code stored in the memory 102 is executed by the processor 101, the processor 101 is used to execute some or all of the functions of the data processing method provided in the embodiments of this application. The implementation manner of the processor 101 executing this process can be referred to the relevant descriptions in the foregoing embodiments. The memory 102 may also include software modules and data required for other running processes such as an operating system.
[0115] The communication interface 103 uses a transceiver module such as, but not limited to, a transceiver to achieve communication with other devices or communication networks. For example, the communication interface 103 can be any one or any combination of the following devices: a network interface (such as an Ethernet interface), a wireless network card, and other devices with network access functions.
[0116] The bus 104 is any type of communication bus for interconnecting the internal devices of the server (e.g., the memory 102, the processor 101, and the communication interface 103). For example, a system bus. The embodiment of the present application takes the interconnection of the above-mentioned devices inside the server through the bus 104 as an example. Optionally, the above-mentioned devices inside the server 100 can also be connected to each other in communication with each other using other connection methods other than the bus 104. For example, the above-mentioned devices inside the server 100 are interconnected through an internal logical interface.
[0117] It should be noted that the above-mentioned multiple devices can be respectively arranged on independent chips, or at least partially or completely arranged on the same chip. Whether to independently arrange each device on different chips or to integrate and arrange it on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation form of the above-mentioned devices. The descriptions of the processes corresponding to the above-mentioned figures have different focuses. For the parts not described in detail in a certain process, please refer to the relevant descriptions of other processes.
[0118] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product providing a program development platform includes one or more computer instructions, and when these computer program instructions are loaded and executed on a server, all or part of the functions of the data processing method provided in the embodiments of the present application are implemented.
[0119] Furthermore, computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium stores computer program instructions that provide a program development platform.
[0120] The embodiment of the present application provides a computing device cluster, including at least one computing device. Each computing device in the at least one computing device includes a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs part or all of the functions in the data processing method provided in the embodiment of the present application.
[0121] Optionally, the computing device is a server, and the computing device cluster is a server cluster. That is, an embodiment of the present application provides a server cluster, the server cluster includes at least one server, each of the at least one server includes a processor and a memory, and the processor of the at least one server is used to execute instructions stored in the memory of the at least one server, so that the server cluster performs part or all of the functions of the data processing method provided in the embodiment of the present application.
[0122] Optionally, the structure of at least one server included in the server cluster can be found in Figure 10 The server 100 is shown. The memory 102 in one or more servers 100 in the server cluster may store the same instructions for executing the data processing method.
[0123] In some possible implementations, the memory 102 of one or more servers 100 in the server cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more servers 100 may jointly execute instructions for executing the data processing method.
[0124] It should be noted that the memory 102 in different servers 100 in the server cluster can store different instructions, which are respectively used to execute part of the functions of the data processing method. That is, the instructions stored in the memory 102 in different servers 100 can be implemented as follows: Figure 9 The functions of one or more modules among the first determination module 901, the second determination module 902, the accumulation module 903, the storage module 904, the adjustment module 905 and the update module 906 are shown.
[0125] In some possible implementations, one or more servers in the server cluster may be connected via a network, which may be a wide area network or a local area network. Figure 11 A possible implementation is shown. Figure 11As shown, multiple servers 110A, 110B, and 110C are connected via a network. Specifically, the network is connected via a communication interface in each server. In this type of possible implementation, servers 110A, 110B, and 110C include a bus 112, a processor 114, a memory 116, and a communication interface 118. The memory 116 in server 110A stores instructions for the functions of the processor. At the same time, the memory 116 in server 110B stores instructions for the functions of the processor. The memory 116 in server 110C stores instructions for the functions of the processor.
[0126] It should be understood that Figure 11 The function of the server 110A shown in the figure may also be completed by multiple servers 110. Similarly, the function of the server 110B may also be completed by multiple servers 110. The function of the server 110C may also be completed by multiple servers 110. And the deployment mode of the modules for implementing the data processing method in the server may also be adjusted according to the application requirements.
[0127] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed (for example, executed by a server, a server cluster, a data processing device, one or more processors, etc.), it implements all or part of the steps of the data processing method provided in the above method embodiment.
[0128] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes a program or code, and when the program or code is executed (for example, executed by a server, a server cluster, a data processing device, one or more processors, etc.), it implements all or part of the steps of the data processing method provided in the above method embodiment.
[0129] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product, and the computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that contains one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium (e.g., a solid-state hard disk), etc.
[0130] It should be understood that the term "at least one" in this application refers to one or more, and "plurality" refers to two or more. In this application, unless otherwise specified, the symbol " / " generally means or, for example, A / B can represent A or B. The term "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, for the convenience of clear description, this application uses words such as "first", "second", and "third" to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first", "second", and "third" do not limit the quantity and execution order.
[0131] Different types of embodiments such as method embodiments and device embodiments provided in the embodiments of the present application can refer to each other, the order of operations of the method embodiments can be appropriately adjusted, and the operations can be increased or decreased in response to the circumstances. Any technician familiar with the technical field can easily think of changes within the technical scope disclosed in the present application, and the methods should be covered within the protection scope of the present application, so they will not be repeated here.
[0132] In the corresponding embodiments provided in the present application, it should be understood that the disclosed devices and the like can be implemented by other configuration methods. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical or other forms. The modules described as separate components may or may not be physically separated, and the components described as modules may or may not be physical modules, which may be located in one place or distributed on multiple network nodes. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0133] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A data processing method, characterized in that, the method includes: determining a first gradient value corresponding to the first model parameter based on a loss value of first prediction data and an initial value of a first model parameter of a recommendation model, where the first prediction data is calculated by the recommendation model based on first user features and first content features, an identifier of a first sample feature in the first user features and the first content features is a first identifier, and the first model parameter corresponds to the first identifier; determining a first gradient storage address according to the first identifier and a mapping relationship, where the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter; accumulating the first gradient value and an accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; storing the first accumulated gradient value into the first gradient storage unit; adjusting a parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit.
2. The method according to claim 1, characterized in that, the determining a first gradient storage address according to the first identifier and the mapping relationship includes: determining a first index storage address according to the first identifier and a first mapping relationship, where the first mapping relationship includes a mapping relationship between the first identifier and the first index storage address, the first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; determining the first gradient storage address according to the index information stored in the first index storage unit.
3. The method according to claim 2, characterized in that, the determining the first gradient storage address according to the index information stored in the first index storage unit includes: when the index information stored in the first index storage unit includes a gradient storage address, determining the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include a gradient storage address, determining a first free storage unit in a first storage space created in advance as the first gradient storage unit, or creating a second storage space and determining a second free storage unit in the second storage space as the first gradient storage unit, and determining an address of the first gradient storage unit as the first gradient storage address, where the first storage space and the second storage space each include at least one storage unit.
4. The method according to claim 3, characterized in that, the method further includes: after the index information stored in the first index storage unit does not include a gradient storage address and the address of the first gradient storage unit is determined as the first gradient storage address, storing the first gradient storage address into the first index storage unit.
5. The method according to claim 4, characterized in that, The index information stored in the first index storage unit includes status information, and the status information is used to indicate that the storage status corresponding to the first identifier is an initial state or an accumulation state. The method further includes: When the status information is used to indicate that the storage status corresponding to the first identifier is an initial state, determining that the index information stored in the first index storage unit does not include a gradient storage address; When the status information is used to indicate that the storage status corresponding to the first identifier is an accumulation state, determining that the index information stored in the first index storage unit includes a gradient storage address.
6. The method according to claim 5, wherein, the method further includes: When the status information is used to indicate that the storage status corresponding to the first identifier is an initial state, after storing the first gradient storage address into the first index storage unit, updating the status information.
7. The method according to any one of claims 1 to 6, wherein, adjusting the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit includes: determining an average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; adjusting the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier.
8. The method according to any one of claims 1 to 7, wherein, both the first user feature and the first content feature are vectors.
9. A data processing device, wherein, the device includes: a first determination module, configured to determine a first gradient value corresponding to the first model parameter based on a loss value of first prediction data and an initial value of a first model parameter of a recommendation model, where the first prediction data is calculated by the recommendation model based on a first user feature and a first content feature, the first sample feature in the first user feature and the first content feature has an identifier of a first identifier, and the first model parameter corresponds to the first identifier; a second determination module, configured to determine a first gradient storage address according to the first identifier and a mapping relationship, where the mapping relationship includes a mapping relationship between the first identifier and the first gradient storage address, the first gradient storage address is used to indicate a first gradient storage unit, and the first gradient storage unit is used to store an accumulated gradient value corresponding to the first model parameter; an accumulation module, configured to accumulate the first gradient value and the accumulated gradient value stored in the first gradient storage unit to obtain a first accumulated gradient value; a storage module, configured to store the first accumulated gradient value into the first gradient storage unit; an adjustment module, configured to adjust the parameter value of the first model parameter according to the accumulated gradient value stored in the first gradient storage unit.
10. The device according to claim 9, wherein, the second determination module is configured to: Determine a first index storage address according to the first identifier and the first mapping relationship, where the first mapping relationship includes the mapping relationship between the first identifier and the first index storage address, and the first index storage address is used to indicate a first index storage unit, and the first index storage unit is used to store index information corresponding to the first identifier; Determine the first gradient storage address according to the index information stored in the first index storage unit.
11. The apparatus according to claim 10, wherein, the second determination module is configured to: when the index information stored in the first index storage unit includes a gradient storage address, determine the gradient storage address stored in the first index storage unit as the first gradient storage address; when the index information stored in the first index storage unit does not include a gradient storage address, determine a first free storage unit in a pre-created first storage space as the first gradient storage unit, or create a second storage space and determine a second free storage unit in the second storage space as the first gradient storage unit, and determine the address of the first gradient storage unit as the first gradient storage address, where the first storage space and the second storage space each include at least one storage unit.
12. The apparatus according to claim 11, wherein, the storage module is further configured to: after the second determination module determines that the index information stored in the first index storage unit does not include a gradient storage address and determines the address of the first gradient storage unit as the first gradient storage address, store the first gradient storage address into the first index storage unit.
13. The apparatus according to claim 12, wherein, the index information stored in the first index storage unit includes status information, and the status information is used to indicate that the storage status corresponding to the first identifier is an initial state or an accumulation state, the second determination module is further configured to: when the status information is used to indicate that the storage status corresponding to the first identifier is an initial state, determine that the index information stored in the first index storage unit does not include a gradient storage address; when the status information is used to indicate that the storage status corresponding to the first identifier is an accumulation state, determine that the index information stored in the first index storage unit includes a gradient storage address.
14. The apparatus according to claim 13, wherein, the apparatus further includes: an update module, configured to update the status information after storing the first gradient storage address into the first index storage unit when the status information is used to indicate that the storage status corresponding to the first identifier is an initial state.
15. The apparatus according to any one of claims 9 to 14, wherein, the adjustment module is configured to: determine an average gradient value corresponding to the first identifier according to the accumulated gradient value stored in the first gradient storage unit; adjust the parameter value of the first model parameter according to the average gradient value corresponding to the first identifier.
16. The device according to any one of claims 9 to 15, wherein, both the first user feature and the first content feature are vectors.
17. A data processing device, wherein, it comprises a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program stored in the memory so that the data processing device executes the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, wherein, a computer program is stored in the computer-readable storage medium, and when the computer program is executed, the method according to any one of claims 1 to 8 is implemented.
19. A computer program product, wherein, the computer program product comprises a program or code, and when the program or code is executed, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Data processing method and apparatus
EP4811178A1
Data processing method and apparatus
WO2025118746A1