Data processing method and system, computer readable storage medium and computer program product

By merging the same data identifiers in the request at the processing node, the number of times the same data is obtained repeatedly during the training and inference of neural network model is reduced, the problem of excessive traffic is solved, and the system performance and computing efficiency are improved.

CN120371871APending Publication Date: 2025-07-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410115211.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

During the training and inference process of neural network models, general-purpose processors need to frequently obtain the same vector data from remote storage devices, resulting in excessive traffic and affecting computing efficiency and system performance.

Method used

By combining the same data identifiers in multiple requests at the processing node, a merged request is generated, and the corresponding data is obtained only once from the storage node, reducing the number of times the same data is obtained repeatedly, and deduplication operation is realized.

Benefits of technology

It effectively reduces traffic, improves system performance, reduces system jitter, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371871A_ABST
    Figure CN120371871A_ABST
Patent Text Reader

Abstract

Disclosed are a data processing method and system, a computer readable storage medium and a computer program product, relating to the field of artificial intelligence, the method comprising: a processing node obtaining a first request sent by each of a plurality of training nodes, each first request comprising a first data identifier of requested first data; the processing node combines the same data identifiers included in the plurality of first requests to generate at least one second request, and the at least one second request comprises a combined second data identifier; the processing node sends the at least one second request to the storage node, and obtains second data corresponding to the second data identifier from the storage node; and the processing node determines first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding training node. Therefore, the number of times of acquiring the same data is reduced, the data volume of the transmitted data is reduced, and the communication traffic is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a data processing method, system, computer-readable storage medium, and computer program product. Background Art

[0002] Currently, general-purpose processors and dedicated processors perform the calculations for training and inference of neural network models to improve the calculation efficiency. For neural network models such as recommendation systems and natural language processing, in one calculation, only some parameters (such as sparse parameters) in the neural network model need to participate in the calculation. For example, when vectorizing the parameters (embedding), the dedicated processor processes the vector data. Usually, the vector data is stored in a remote storage device, and the general-purpose processor needs to obtain the vector data from the remote storage device and feed it back to the dedicated processor, resulting in a large amount of communication. Summary of the Invention

[0003] This application provides a data processing method, system, computer-readable storage medium, and computer program product, thereby effectively reducing the communication volume.

[0004] In a first aspect, a data processing method is provided, which is applied to a data processing system. The data processing system includes a processing node, a plurality of training nodes, and a storage node. The training nodes are used to perform inference or training of a neural network model, and the storage node is used to store the data used by the training nodes to perform inference or training of the neural network model. The method includes: the processing node obtains a first request sent by each training node among the plurality of training nodes, and each first request includes a first data identifier of the requested first data; the processing node merges the same data identifiers included in the plurality of first requests to generate at least one second request, and at least one second request includes the merged second data identifier; the processing node sends at least one second request to the storage node and obtains the second data corresponding to the second data identifier from the storage node; the processing node determines the first data corresponding to the first data identifier from the second data and returns the first data to the corresponding training node.

[0005] Compared with the processing node processing each request among multiple requests, especially when multiple requests indicate obtaining the same data, the processing node needs to repeatedly obtain the same data, resulting in a large amount of communication for processing requests. The solution provided by this application performs a deduplication operation on multiple requests, that is, merges the same data identifiers in multiple requests, and obtains the data corresponding to the merged data identifier from the storage node. For requests containing the same data identifier, the same data corresponding to the data identifier is only obtained from the storage node once. Thus, by reducing the number of times of obtaining the same data, the amount of data transmitted is reduced, and the communication volume is effectively reduced.

[0006] In a possible implementation, the processing node and multiple training nodes are deployed on a training server. The processing node is executed by the general-purpose processor of the training server, and the training node is executed by the dedicated processor of the training server.

[0007] Thus, the communication volume and concurrency across different servers are reduced, the system performance is improved, and the system jitter is reduced.

[0008] In another possible implementation, the processing node merges the same data identifiers included in multiple first requests, including: the processing node retains one data identifier among the same data identifiers included in the multiple first requests to obtain the merged second data identifier.

[0009] The data identifier in the request is used to indicate data. If two requests contain the same data identifier, it means that the two requests indicate to obtain the same data. By merging the data identifiers included in the requests, duplicate removal operations are performed on multiple requests, reducing the number of times of obtaining the same data, reducing the amount of data to be transmitted, and effectively reducing the communication volume.

[0010] In another possible implementation, the processing node determines the first data corresponding to the first data identifier from the second data, including: the processing node determines the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.

[0011] In another possible implementation, returning the first data to the corresponding training node includes: the processing node returns the first data to the training node that issues the first data identifier.

[0012] In another possible implementation, the processing node further includes a cache for caching the second data obtained from the storage server. The method further includes: the processing node obtains a third request, determines whether the third data corresponding to the data identifier in the third request exists in the cache. If it exists, the third data is obtained from the cache and returned to the corresponding training node.

[0013] Thus, the processing node obtains data from the cache, reducing the possibility of the processing node obtaining data from the remote storage device and further reducing the communication volume across different devices.

[0014] In a second aspect, a data processing method is provided. The computing system to which the method is applied includes a general-purpose processor and a dedicated processor. The dedicated processor is used to train or infer an artificial intelligence model. The method includes: the general-purpose processor obtains a first request sent by each dedicated processor among a plurality of dedicated processors, and each first request includes a first data identifier of the first data requested; the general-purpose processor merges the same data identifiers included in the plurality of first requests to generate at least one second request, and the at least one second request includes a merged second data identifier; the general-purpose processor sends the at least one second request to a storage node and obtains second data corresponding to the second data identifier from the storage node; the general-purpose processor determines the first data corresponding to the first data identifier from the second data and returns the first data to the corresponding dedicated processor.

[0015] The dedicated processor obtains data and performs training calculations or inference calculations on the neural network model according to the data.

[0016] In a third aspect, a data processing apparatus is provided. The data processing apparatus includes various modules for performing the data processing method in the first aspect or any possible design of the first aspect. For example, the data processing apparatus is used to implement the functions of the general-purpose processor, and the data processing apparatus includes a communication module and a processing module.

[0017] The communication module is used to obtain a plurality of first requests of a plurality of training nodes, and each first request includes a first data identifier of the first data requested; the processing module is used to merge the same data identifiers included in the plurality of first requests; the processing module is further used to generate at least one second request, and the at least one second request includes a merged second data identifier; the communication module is further used to send the at least one second request to a storage node and obtain second data corresponding to the second data identifier from the storage node; the processing module is further used to determine the first data corresponding to the first data identifier from the second data; the communication module is further used to return the first data to the corresponding training node.

[0018] In a possible implementation manner, when the processing module merges the same data identifiers included in the plurality of first requests, it is specifically used to: retain one data identifier among the same data identifiers included in the plurality of first requests to obtain a merged second data identifier.

[0019] In another possible implementation manner, when the processing module determines the first data corresponding to the first data identifier from the second data, it is specifically used to: determine the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.

[0020] In another possible implementation manner, when the communication module is used to return the first data to the corresponding training node, it is specifically used to: return the first data to the training node that obtains the first data identifier corresponding thereto.

[0021] In another possible implementation, the processing node further includes a cache for caching the second data obtained from the storage server. The processing module is further configured to obtain a third request and determine whether the third data corresponding to the data identifier in the third request exists in the cache. If so, the third data is obtained from the cache. The communication module is further configured to return the corresponding training node.

[0022] In a fourth aspect, a computing system is provided. The computing system includes a general-purpose processor and a plurality of dedicated processors. The general-purpose processor executes the operation steps of the method in the first aspect or any one of the possible implementation manners of the first aspect. The plurality of dedicated processors perform training calculations or inference calculations on the neural network model according to the data.

[0023] In a fifth aspect, a data processing system is provided. The data processing system includes a training node, a processing node, and a storage node. The training node is configured to perform inference or training of the neural network model. The storage node is configured to store the data used by the training node to perform inference or training of the neural network model. The processing node executes the operation steps of the method in the first aspect or any one of the possible implementation manners of the first aspect.

[0024] In a sixth aspect, a computer device is provided. The computer device includes a memory and a plurality of processors. The memory is configured to store a set of computer instructions. When the processor executes the set of computer instructions, the processor executes the operation steps of the method in the first aspect or any one of the possible implementation manners of the first aspect.

[0025] In a seventh aspect, a computer-readable storage medium is provided, including: computer software instructions. When the computer software instructions run on the processor, the processor is caused to execute the operation steps of the method in the first aspect or any one of the possible implementation manners of the first aspect.

[0026] In an eighth aspect, a computer program product is provided. When the computer program product runs on a computer, the computer is caused to execute the operation steps of the method in the first aspect or any one of the possible implementation manners of the first aspect.

[0027] For the technical effects brought about by any one of the design manners in the second to eighth aspects, reference may be made to the technical effects brought about by the first aspect or different design manners in the first aspect, which will not be elaborated here.

[0028] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic structural diagram of a neural network provided by the present application;

[0030] Figure 2 Schematic diagram of a sparse parameter provided for this application;

[0031] Figure 3 Schematic diagram of a one-hot encoding and vectorization provided for this application;

[0032] Figure 4 Schematic diagram of a parameter storage provided for this application;

[0033] Figure 5 Schematic diagram of the architecture of a data processing system provided for this application;

[0034] Figure 6 Schematic diagram of the architecture of a computing system provided for this application;

[0035] Figure 7 Schematic diagram of the data storage of an ultra-large-scale model provided for this application;

[0036] Figure 8 Schematic diagram of a multi-level storage architecture provided for this application;

[0037] Figure 9 Schematic diagram of the flow of a data processing method provided for this application;

[0038] Figure 10 Schematic diagram of the structure of a data processing device provided for this application;

[0039] Figure 11 Schematic diagram of the structure of a computer device provided for this application. Detailed implementation manners

[0040] For ease of understanding, the main terms involved in this application are first explained.

[0041] Artificial Neural Network (ANN): Abbreviated as Neural Network (NN) or neural network-like. In the fields of machine learning and cognitive science, it is a mathematical model or computational model that mimics the structure and function of a biological neural network (the central nervous system of an animal, especially the brain). A neural network is a network formed by connecting multiple single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.

[0042] Each node represents a specific output function, called an activation function. The connection between every two nodes represents a weighted value for the signal passing through this connection, called the weight, which is equivalent to the memory of an artificial neural network. The output of the neural network varies according to the connection mode, weight values, and activation functions of the network. The neural network itself is usually an approximation of a certain algorithm or function in nature, or may be an expression of a logical strategy.

[0043] As Figure 1 shown, it is a schematic structural diagram of a neural network provided by this application. The neural network 100 includes N processing layers, where N is an integer greater than or equal to 3. The first layer of the neural network 100 is the input layer 110, which is responsible for receiving input signals, and the last layer of the neural network 100 is the output layer 130, which is responsible for outputting the processing results of the neural network. The other layers except the first layer and the last layer are intermediate layers 140, and these intermediate layers 140 together form the hidden layer 120. Each intermediate layer 140 in the hidden layer 120 can both receive input signals and output signals. The hidden layer 120 is responsible for the processing process of the input signals. Each layer represents a logical level of signal processing, and through multiple layers, the data signal can be processed through multiple levels of logic.

[0044] In some feasible embodiments, the input signals of the neural network can be various forms of signals such as video signals, voice signals, text signals, image signals, temperature signals, and engineering signals that can be processed by a computer.

[0045] Model training: It refers to training a neural network model using a training set so that the neural network model can predict or classify unknown data. During the training process of the neural network model, it learns according to the features and target values in the training set, and generates a neural network model after training. This neural network model can be used to predict or classify unknown data. Neural network model training is one of the most important links in machine learning, affecting the accuracy and reliability of the neural network model.

[0046] Parameters: It refers to the variables or weights that need to be learned or adjusted in the neural network model. Parameters can affect the prediction ability and performance of the neural network model. During the training process of the neural network model, the neural network model tries different parameter combinations to optimize the performance of the neural network model. Common parameters include weights, biases, learning rates, and regularization coefficients, etc. When using the neural network model for prediction, these parameters are used to calculate the output results.

[0047] Sparse parameters: Parameters that only partially participate in the calculation during one training process. For example, the parameters participating in forward calculation and backward update.

[0048] Exemplarily, Figure 2 is a schematic diagram of a sparse parameter provided for this application. As Figure 2 shown, in one training step, the weight part participating in the calculation only includes one-half of all the weights.

[0049] Generally, the scale of sparse parameters is very large. For example, taking a production-level recommendation system as an example, the data volume of sparse parameters reaches the level of 10TB - 30TB.

[0050] Recommendation system: An application that is built on the basis of massive data mining, provides personalized decision-making support and information services for users, and judges the items or services that the user currently needs or is interested in according to information such as the user's historical behavior, social relationships, points of interest, and the context environment.

[0051] In some neural network models such as recommendation systems and natural language processing, the input data usually contains discrete features. The embedding table (Embedding table) parameter is used to convert the input data into continuous vector data, and then the vector data is processed. In a single training process, only some of the parameters in the embedding table participate in the calculation and are trained and updated.

[0052] Embedding technology: Refers to a representation form of dense vectors. That is, converting sparse parameters into dense vector data. Embedding can represent the features of an object. The object can be, for example, height, gender, name, or item, etc.

[0053] Exemplarily, as Figure 3 shown, is a schematic diagram of one-hot encoding and Embedding provided for this application. Embedding is equivalent to performing smoothing on one-hot encoding, and one-hot encoding is equivalent to performing max pooling on Embedding.

[0054] During the training process of the neural network model, parameters (such as sparse parameters) are continuously adjusted to make the predicted value calculated from the input data and the parameters approximate the actual value. The parameters can be loaded into the storage medium of the dedicated processor, and the dedicated processor calculates and updates the parameters. For example, the dedicated processor includes, but is not limited to, a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), and an embedded neural network processor (NPU). The storage medium of the dedicated processor includes High Bandwidth Memory (HBM).

[0055] Exemplarily, as Figure 4 As shown in (a) of [], a schematic diagram of parameter storage provided by this application is shown. The Embedding table is entirely stored in the storage medium of the GPU (such as: HBM). The storage medium of the GPU is divided into two parts. One part is used to store data during the training process of the neural network model; the other part is used to store the Emb table. During the training process of the neural network model, the required Embedding vectors are pulled to the specified GPU in an All2All manner to achieve the purpose of sharing the Emb table. The Embedding table is simply referred to as the Emb table. The Emb table contains vector data of sparse parameters.

[0056] However, this solution has a large demand for GPUs. For example, it requires about 300 GPU cards to store a 10TB Emb table, resulting in extremely high training costs for the neural network model. In addition, due to the large number of GPUs required, the linearity of the system is poor, the performance of the entire system cannot be fully exerted, and the computing reliability is also poor. All GPUs use All2All to pull Embedding vectors to participate in training. If one GPU fails, it is easy to lose the data of the Emb table stored in the faulty GPU.

[0057] The storage capacity of the dedicated processor is only 16GB - 32GB, which is a drop in the bucket for dozens of TB of sparse parameters. Therefore, the sparse parameters can also be stored in storage media such as the volatile memory (such as: memory) or non-volatile memory (such as: solid state drive (SSD)) of the server.

[0058] In some embodiments, the sparse parameters are stored in a server including a general-purpose processor, and the general-purpose processor converts the sparse parameters into vector data. Then, the vector data is transmitted to the dedicated processor for calculation.

[0059] Exemplarily, asFigure 4 As shown in (b) of , it is a schematic diagram of parameter storage provided by this application. The entire Emb table is stored in the CPU cluster; for example, the entire Emb table is stored in a storage medium such as a volatile memory (e.g., memory) or a non-volatile memory (e.g., solid state drive (SSD)) in the CPU cluster. Among them, the CPU can provide functions for storing, querying, and updating the Emb table. The CPU cluster expands the storage capacity of the GPU. The GPU trains the neural network model. The CPU and the GPU cooperate to perform forward calculation and backward update of the neural network model; during forward calculation, the required Emb table is pulled from the CPU cluster through the pull interface; the GPU trains the neural network model based on the Emb table to obtain gradients; the gradients are pushed back to the CPU cluster through the push interface, and the Emb table is updated in the CPU cluster.

[0060] Since a large amount of Emb tables are transmitted through the network each time the neural network model is trained, that is, the pull operation and the push operation are executed, the calculation efficiency is relatively low.

[0061] In some other embodiments, if the server cannot store the complete Emb table, the Emb table is stored in the storage medium of the server and the remote storage device.

[0062] Exemplarily, as Figure 4 shown in (c) of , it is a schematic diagram of parameter storage provided by this application. When the storage capacity of the memory is insufficient, the SSD is also used to store the Emb table. The CPU can pull the required Emb table from the SSD and transmit the Emb table to the GPU.

[0063] For the Emb table corresponding to an ultra-large-scale model, the data volume is very large. For example, the data volume of the Emb table can reach 100TB. If the Emb table is stored in a dedicated processor, thousands of dedicated processors are required to store the Emb table and thousands of dedicated processors are used to train the ultra-large-scale model. The demand for dedicated processors is too large and it is almost impossible to achieve. If the Emb table is stored in a remote storage device, the general-purpose processor not only needs to execute queries and updates of the Emb table, but also executes a large number of pull operations and push operations, resulting in a large load on the general-purpose processor, consuming more network resources, generating a large amount of traffic, and the calculation efficiency of the neural network model is relatively low.

[0064] Moreover, when different dedicated processors perform model training or model inference calculations, they may require the same data. Then the general-purpose processor needs to perform pull operations and push operations on the same data multiple times, resulting in the consumption of more network resources and the generation of a large amount of traffic.

[0065] To solve the problem of a large amount of communication caused by the need to perform pull operations and push operations on the same data multiple times during the inference or training of a neural network model, the present application provides a data processing method. That is, a processing node obtains first requests sent by each of multiple training nodes, where each first request includes a first data identifier of the first data requested; the processing node merges the same data identifiers included in the multiple first requests to generate at least one second request, and the at least one second request includes the merged second data identifier; the processing node sends the at least one second request to a storage node and obtains the second data corresponding to the second data identifier from the storage node; the processing node determines the first data corresponding to the first data identifier from the second data and returns the first data to the corresponding training node.

[0066] Compared with the processing node processing each of multiple requests, especially when multiple requests indicate obtaining the same data, the processing node needs to repeatedly obtain the same data and perform pull operations and push operations on the same data multiple times, resulting in a large amount of communication for processing requests. The solution provided by the present application performs a deduplication operation on multiple requests, that is, merges the same data identifiers in multiple requests, obtains the data corresponding to the merged data identifier from the storage node, and for requests containing the same data identifier, only obtains the data corresponding to the same data identifier from the storage node once, that is, performs a pull operation and a push operation on the same data once. Thus, by reducing the number of times of obtaining the same data, the amount of data transmitted is reduced, effectively reducing the communication volume.

[0067] The following describes in detail the implementation manner of the data processing method provided by the present application with reference to the accompanying drawings.

[0068] Figure 5 It is a schematic diagram of the architecture of a data processing system provided by the present application. As Figure 5 shown, the data processing system 500 includes a client 510, a computing cluster 520, and a storage cluster 530.

[0069] The computing cluster 520 includes multiple computing nodes 521. The multiple computing nodes 521 can be connected based on high-speed interconnection technology through network devices (such as switches, network cards, etc.) to enable communication between the multiple computing nodes 521.

[0070] In some embodiments, the computing node 521 may include computing units with computing capabilities such as a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), and an embedded neural-network processing unit (NPU) to provide high-performance computing.

[0071] The computing cluster 520 further includes a control node 522. The control node 522 is used to manage and allocate tasks, and multiple tasks are executed in parallel by multiple computing nodes to improve the data processing rate.

[0072] In the present application, the control node 522 is further used to perform preprocessing operations on the data required for model training or model inference, and cooperate with multiple computing nodes 521 to manage the data. For example, the control node 522 vectorizes sparse parameters and converts them into dense vector data (such as: Emb table). Furthermore, when the control node 522 obtains a model training task or a model inference task, the vector data is stored in a specified storage medium according to the storage requirements and the storage resource characteristics of the system. The specified storage medium includes at least one of the storage media associated with the computing node, the storage media associated with the control node, or the storage media associated with the storage node. That is, for neural network models of different scales, due to the different amounts of data to be loaded, the control node 522 performs hierarchical storage on data with different storage requirements, that is, stores the data related to model training or model inference in the storage media of at least one of the computing node, the control node, or the storage node.

[0073] For example, for an ultra-large-scale model, when the control node 522 obtains a model training task or a model inference task, the vector data is stored in the storage media associated with the computing node, the storage media associated with the control node, or the storage media associated with the storage node according to the storage requirements and the storage resource characteristics of the system. That is, since the amount of data in the Emb table corresponding to the ultra-large-scale model is very large, the Emb table can be divided into multiple data blocks, and the multiple data blocks are hierarchically stored in the storage media associated with the computing node, the storage media associated with the control node, or the storage media associated with the storage node.

[0074] When the computing node 521 performs training calculation or inference calculation on the neural network model according to the data, if the data is stored in the storage medium of the computing node 521, the computing node 521 obtains the data related to model training or model inference from the storage medium of the computing node 521.

[0075] If the data is stored in the storage medium of the control node 522 or the storage medium of the storage node 531, the control node 522 obtains the data from the storage medium of the control node 522 or the storage medium of the storage node 531, loads the data into the computing node 521, and the computing node 521 performs training calculations or inference calculations on the neural network model according to the data.

[0076] Optionally, if the data is stored in the storage medium of the control node 522 or the storage medium of the storage node 531, the computing node 521 can also obtain the data related to model training or model inference from the storage medium of the control node 522 or the storage medium of the storage node 531.

[0077] In some embodiments, the control node 522 obtains the first requests sent by each computing node 521 among the multiple computing nodes 521, and each first request includes the first data identifier of the requested first data; the control node 522 merges the same data identifiers included in the multiple first requests to generate at least one second request, and at least one second request includes the merged second data identifier; the control node 522 sends at least one second request to the storage node, and obtains the second data corresponding to the second data identifier from the storage node; the control node 522 determines the first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding computing node 521.

[0078] The storage cluster 530 includes multiple storage nodes 531. One storage node 531 includes one or more controllers, network cards, and multiple hard disks. The hard disks are used to store data. The hard disks can be magnetic disks or other types of storage media, such as solid state drives or shingled magnetic recording hard disks, etc. The network cards are used to communicate with the computing nodes 521 included in the computing cluster 520. The controller is used to write data to the hard disk or read data from the hard disk according to the read / write data requests sent by the computing node 521. During the process of reading and writing data, the controller needs to convert the address carried in the read / write data request into an address that the hard disk can recognize.

[0079] The storage cluster 530 includes multiple storage nodes 531 and the computing cluster 520 includes multiple computing nodes 521 are connected based on high-speed interconnect technology through network devices (such as switches, network cards, etc.), enabling communication between the multiple computing nodes 521 and the multiple storage nodes 531.

[0080] In this application, the storage nodes 531 included in the storage cluster 530 can be used as the remote storage devices described in this application. For example, a storage node can be a server, and the storage node includes a CPU and a storage medium. In the case where the remaining storage capacity in the computing cluster is insufficient, the storage cluster 530 can provide a storage expansion function, that is, store the data related to model training or model inference, for example, the Emb table required for model training or model inference.

[0081] The client 510 communicates with the computing cluster 520 and the storage cluster 530 via the network 540. For example, the client 510 sends a request to the computing cluster 520 via the network 540, requesting the computing cluster 520 to perform model training or model inference. The network 540 can refer to an enterprise internal network (such as a Local Area Network (LAN)) or the Internet. The client 510 can refer to a computer connected to the network 540, and can also be called a workstation. Different clients can share resources on the network (such as computing resources, storage resources).

[0082] In some embodiments, the client 510 is installed with a client program 511. The client 510 runs the client program 511 to display a user interface (UI). The user 550 operates the user interface to submit a request. For example, the user 550 operates the user interface to submit a model training request or a model inference request. After the control node 522 obtains the request, it can obtain the data required for model training or model inference from the storage cluster 530, and store the data in a specified storage medium according to the storage requirements of the data and the storage resource characteristics of the system.

[0083] Optionally, the system administrator 560 can configure system information and the like through the client 510 by invoking the application platform interface (API) 512 or the command-line interface (CLI) interface 513. For example, the system information includes the storage policy of the data required for model training or model inference.

[0084] Figure 5 This is only a schematic diagram. The embodiments of the present application do not limit the device connection method, the number of devices, and the device form in the data processing system.

[0085] The present application does not limit the deployment form of the above computing node 521 and control node 522.

[0086] For example, the computing node 521 and the control node 522 can be deployed on the same server. The computing node can also be referred to as a dedicated processor. The control node can be referred to as a general-purpose processor. For example, the general-purpose processor can include a central processing unit (CPU). The dedicated processor includes computing nodes with computing capabilities such as GPUs, DPUs, and NPUs to provide high-performance computing. Alternatively, the computing node can be referred to as a device, a training card, or an acceleration card. The control node can be referred to as a host. Herein, the number of dedicated processors and general-purpose processors included in a server is not limited.

[0087] As another example, the computing node 521 and the control node 522 can each be an independent server, etc. The computing node 521 can be a dedicated server providing a heterogeneous computing architecture to provide high-performance computing. The control node 522 can be an ordinary general-purpose server. The dedicated server and the general-purpose server are interconnected via network devices based on high-speed interconnection technology. The dedicated server includes a storage medium, a general-purpose processor, and multiple dedicated processors. The dedicated server can be an artificial intelligence server. Multiple artificial intelligence servers are interconnected via a network, and multiple artificial intelligence servers form an AI cluster to implement model training calculations and model inference calculations. The general-purpose server includes a general-purpose processor.

[0088] The general-purpose server is used to manage and allocate tasks, instructing multiple computing nodes to execute multiple tasks in parallel to improve the data processing rate.

[0089] The general-purpose processor in the dedicated server is used to store vector data in a specified storage medium according to the storage requirements and the storage resource characteristics of the system. The specified storage medium includes at least one of the storage medium of the dedicated processor, the storage medium of the dedicated server, and the storage medium of the general-purpose server. The dedicated processor in the dedicated server is used to perform training calculations or inference calculations on the neural network model according to the data.

[0090] For ultra-large-scale models, the general-purpose processor in the dedicated server is also used to store vector data in the storage medium associated with the computing node, the storage medium associated with the control node, or the storage medium associated with the storage node.

[0091] It should be noted that the specific form of the storage medium described in this application is not limited. The storage medium includes a volatile memory pool or a non-volatile memory pool, or may include both volatile and non-volatile memories. For example, the storage medium of the dedicated processor includes at least one of HBM or SSD.

[0092] The following describes the hierarchical storage of data required for model training calculations or model inference calculations provided by this application.

[0093] Figure 6 The architecture diagram of a computing system provided for this application. The computing system provides a hierarchical storage architecture. As Figure 6 shown, the computing system 600 includes one or more dedicated servers 610 and one or more general servers 620. The dedicated server 610 contains a general-purpose processor and multiple dedicated processors. The general-purpose processor and the multiple dedicated processors can be co-deployed on the same server. The storage medium in the dedicated processor serves as the first-level storage layer, and the storage medium in the general-purpose processor serves as the second-level storage layer, that is, the storage medium in the general-purpose processor serves as local storage. The general server 620 serves as the third-level storage layer, that is, the storage server 620 serves as remote storage.

[0094] The storage medium in the dedicated processor is used to store data during model training or model inference, as well as all or part of the data of the Emb table.

[0095] The storage medium in the general-purpose processor is used to store all or part of the data of the Emb table, gradient accumulation, optimizer data, cache data, etc.

[0096] The general server is used to store all or part of the data of the Emb table, gradient accumulation, optimizer data, cache data, etc.

[0097] The general-purpose processor can serve as a processing node, and the dedicated processor can serve as a training node.

[0098] The general-purpose processor is used to perform preprocessing operations on the data required for model training or model inference. For example, vectorize the sparse parameters and convert them into dense vector data to obtain the Emb table. The general-purpose processor is also used to store the Emb table in the specified storage medium according to the storage requirements of the Emb table and the storage resource characteristics of the system. The specified storage medium includes the storage media of at least one of the above-mentioned first-level storage layer, second-level storage layer, and third-level storage layer.

[0099] The general-purpose processor can run multiple worker processes. One worker process can control one dedicated processor. For example, when the worker instructs the dedicated processor to perform model training calculation or model inference calculation, it can pull the Emb table and transfer it to the dedicated processor, and update the Emb table according to the gradient pushed back from the dedicated processor.

[0100] Optionally, the model described in this application may refer to a recommendation model.

[0101] Based on the above architecture, unified hierarchical storage solutions are proposed for medium-scale models, large-scale models, and extra-large-scale models respectively.

[0102] In a first possible implementation, the data volume of the Emb table corresponding to a medium-sized model is not large, and the required storage capacity is not large either. The Emb table can be stored in the primary storage layer, that is, stored in the storage medium of the dedicated processor.

[0103] In a second possible implementation, the data volume of the Emb table corresponding to a large-scale model is large, and the required storage capacity is also large. For example, the data volume of the Emb table can reach 10TB. If the Emb table is stored in the storage medium of the dedicated processor, taking the storage capacity of each dedicated processor as 32GB as an example, 300 dedicated processors are required to start training, and the demand for dedicated processors is too high. The Emb table can be stored in the secondary storage layer, that is, stored in the storage medium of the general-purpose processor.

[0104] In a third possible implementation, the data volume of the Emb table corresponding to an extra-large-scale model is extremely large, and the Emb table can be stored in the tertiary storage layer, that is, stored in the storage medium of the general-purpose server.

[0105] In some embodiments, the Emb table can be split into multiple data blocks according to the number of general-purpose servers, and the multiple data blocks of the Emb table are respectively stored in the storage media of multiple general-purpose servers. Among them, the splitting method of the Emb table is not limited in this application. The number of data blocks of the Emb table can also be less than the number of general-purpose servers, and the data blocks of the Emb table are stored in some of the multiple general-purpose servers. The number of data blocks of the Emb table can also be greater than or equal to the number of general-purpose servers, and the data blocks of the Emb table are stored in the storage media of multiple general-purpose servers. One general-purpose server stores one data block, or one general-purpose server can also store multiple data blocks.

[0106] Exemplarily, as Figure 7 shown, the general-purpose processor can store the Emb table of the extra-large-scale model in the general-purpose server.

[0107] After the general-purpose processor obtains a model training task request or a model inference task request, the worker in the general-purpose processor obtains the identifier of the data set from the request, and the identifier can indicate the Emb vector in the Emb table. The general-purpose processor obtains the Emb vector corresponding to the identifier and transfers it to the dedicated processor.

[0108] The general-purpose processor reads the Emb vector corresponding to the identifier from the general-purpose server and transfers it to the dedicated processor.

[0109] Optionally, the storage medium in the general-purpose processor is used to store a part of the Emb vectors in the full Emb table. The storage medium (such as HBM) in the dedicated processor is used to cache a part of the Emb vectors in the full Emb table.

[0110] When a dedicated processor trains or infers a neural network model, it can read the required Emb vectors from the HBM. After training, the Emb vectors in the HBM are updated according to the gradients. Alternatively, after training, the general-purpose processor can update the Emb table according to the gradients.

[0111] In some embodiments, triggered by a period / event, the dedicated processor periodically writes the Emb vectors in the HBM to the storage medium of the general-purpose processor.

[0112] Thus, the hierarchical storage architecture extends the storage method of the Emb table, supports the hierarchical storage of Emb tables for models of different scales, and various different hardware deployment forms, realizing a naturally scalable Emb table storage solution. One set of architecture supports three recommendation models: medium-scale Emb recommendation model, large-scale Emb recommendation model, and extra-large-scale Emb recommendation model, without the need to develop and deploy relevant software for each scenario.

[0113] For example, for a medium-scale model, by expanding the dedicated processor, the Emb table is stored in chunks in the storage media (Embedding Storage) of multiple dedicated processors, and multiple Embedding Storages are combined into a complete Emb table.

[0114] For another example, for a large-scale model, the Emb table is stored in the storage medium of the general-purpose processor, and the storage medium in the dedicated processor serves as a cache layer, reading the Emb table from the general-purpose processor as needed for model training or model inference.

[0115] For another example, for an extra-large-scale model, the Emb table is stored in a remote storage device, and the storage medium in the dedicated processor serves as a cache layer, reading the Emb table from the remote storage device as needed for model training or model inference.

[0116] It should be noted that the training node can run a worker process to pull the Emb vectors from the processing node, drive the dedicated processor to train or infer the neural network model, or push the gradients back to the processing node. The processing node can run a worker process to parse the received requests, pull the Emb vectors from the subsequent storage nodes or processing nodes according to the identifiers in the requests, and transfer them to the training node to instruct the training node to train or infer the neural network model.

[0117] In some embodiments, when there are many training nodes, the training nodes can be divided into multiple groups, and the requests sent by the training nodes within one group are processed by one processing node. The processing node can provide deduplication and caching services for the training nodes; finally, transfer the Emb vectors to the storage medium in the training nodes.

[0118] Exemplarily, as Figure 8 shown, it is a schematic diagram of a multi-level storage architecture provided by the present application. Multiple training nodes and processing nodes can be jointly deployed on the same device. For example, multiple training nodes and processing nodes are jointly deployed on a dedicated server. The storage nodes are deployed on general-purpose servers.

[0119] Among them, the Emb layer is used to convert sparse parameters into dense vector data. The DNN layer is used to train or infer a neural network model based on the dense vector data.

[0120] Optionally, one processing node can be used to process requests sent by multiple training nodes. The training nodes and the corresponding processing nodes within one group can be deployed on the same device, or the training nodes and the corresponding processing nodes within multiple groups can be deployed on the same device.

[0121] In practical applications, multiple general-purpose servers can be used to store the Emb table. The Emb table can be sliced into multiple data blocks according to the number of general-purpose servers, and the multiple data blocks of the Emb table are respectively stored in the storage media of multiple general-purpose servers.

[0122] The processing node can cache the Emb vectors. Then, the storage media in the training node, the processing node, and the storage node can jointly store the Emb table.

[0123] One processing node is used to receive requests sent by the training nodes within the corresponding one group, pull the Emb vectors from the processing node or the storage node, and transmit them to the training nodes.

[0124] In the present application, the requests sent by the training nodes within one group may contain the same data identifier. The processing node performs a deduplication operation on multiple requests, that is, combines the same data identifiers in multiple requests, obtains the data corresponding to the combined data identifier from the storage node, and for requests containing the same data identifier, only obtains the data corresponding to the same data identifier from the storage node once, that is, performs a pull operation and a push operation on the same data once. Thus, by reducing the number of times of obtaining the same data, the amount of data transmitted is reduced, and the communication volume is effectively reduced.

[0125] Next, the data processing process will be described in detail with reference to the accompanying drawings.

[0126] Figure 9 It is a schematic flowchart of a data processing method provided by the present application. Here, the transmission of sparse parameters related to the training calculation or inference calculation of the neural network model will be mainly described. As Figure 9 shown, the method includes the following steps 910 to step 960.

[0127] Step 910: The processing node obtains multiple first requests from multiple training nodes.

[0128] When model training or model inference is performed on multiple training nodes, the training nodes send requests to the processing node, and then the processing node receives multiple requests from the multiple training nodes. The multiple requests received from the multiple training nodes can be collectively referred to as first requests.

[0129] The data identifiers included in the multiple requests sent by the multiple training nodes can be collectively referred to as first data identifiers. The data indicated by the first data identifiers included in the multiple first requests can be collectively referred to as first data.

[0130] Each first request includes the first data identifier of the first data requested.

[0131] Optionally, the processing node obtains one first request sent by one training node, or the processing node obtains multiple first requests sent by one training node. This application does not limit the number of requests sent by the training nodes.

[0132] The first data identifiers included in the first requests sent by different training nodes may be the same or different. The first data identifiers included in the multiple first requests sent by the same training node may be the same or different. Understandably, a training node issues multiple requests, and the multiple requests indicate obtaining the same data, or the multiple requests indicate obtaining different data. Or, two different training nodes issue multiple requests, and the multiple requests indicate obtaining the same data, or the multiple requests indicate obtaining different data. That is, the first data indicated by the multiple first requests may be the same or different.

[0133] The processing node can perform a deduplication operation on the multiple requests, that is, merge and deduplicate the same data identifiers included in the multiple requests, which can reduce the communication volume of the requests sent to the storage node. The following steps 920 and 930 are described.

[0134] Step 920: The processing node merges the same data identifiers included in the multiple first requests.

[0135] If at least two requests among multiple first requests contain the same data identifier, it means that at least two requests are for obtaining the same data, and the processing node combines the data identifiers included in the at least two requests. Understandably, the processing node retains one of the same data identifiers included in the at least two requests. If the same data identifiers are divided into groups, the data identifiers included in the multiple first requests can be divided into one group or multiple groups. The data identifiers within each group are the same, and the data identifiers in different groups are different. The processing node can combine the same data identifiers within the group, that is, retain one of the same data identifiers included in the multiple first requests, to obtain the combined second data identifier. The combined second data identifier can include one data identifier or multiple data identifiers. When the combined second data identifier includes multiple data identifiers, these multiple data identifiers are all different. The data identifiers included in the multiple first requests include the combined second data identifier. The combined second data identifier is a partial identifier among the data identifiers included in the multiple first requests.

[0136] The data indicated by the data identifier includes sparse parameters related to the training calculation or inference calculation of the neural network model.

[0137] For example, among the multiple first requests, the first number of requests all contain the first identifier, and the first identifier indicates the first data, then the first number of requests are used to indicate obtaining the first data. Among the multiple first requests, the second number of requests all contain the second identifier, and the second identifier indicates the second data, then the second number of requests are used to indicate obtaining the second data. Each of the first number of requests indicates obtaining the same data, each of the second number of requests indicates obtaining the same data, and the data indicated by the first number of requests and the second number of requests is different. The processing node combines the first identifiers that the first number of requests all contain to obtain one first identifier, and the processing node combines the second identifiers that the second number of requests all contain to obtain one second identifier. The combined second data identifier includes the first identifier and the second identifier.

[0138] Step 930: The processing node generates at least one second request, and the at least one second request includes the combined second data identifier.

[0139] After the processing node combines the same data identifiers included in the multiple first requests, it generates at least one second request, and the at least one second request includes the combined second data identifier.

[0140] If there is a group of the same data identifiers among the multiple first requests, that is, the data identifiers included in the multiple first requests are all the same, that is, the multiple first requests are for requesting and obtaining the same data, then the processing node generates one second request.

[0141] If multiple first requests contain multiple sets of identical data identifiers, that is, the data identifiers included in multiple first requests are partially the same, that is, multiple first requests are used to request multiple different data. After the processing node merges the identical data identifiers included in the multiple first requests, multiple different data identifiers are obtained, and then the processing node generates multiple second requests, and one second request contains one of the multiple different data identifiers.

[0142] Understandably, the data identifiers included in the multiple first requests include at least one of the merged second data identifiers included in the at least one second request.

[0143] Step 940: The processing node sends at least one second request to the storage node and obtains the second data corresponding to the second data identifier from the storage node.

[0144] If the processing node does not cache the second data indicated to be obtained by at least one second request, it obtains the data indicated by the merged second data identifier included in at least one second request from the storage node. Understandably, for at least two requests indicating to obtain the same data, only one read operation is performed, and for requests indicating to obtain different data, at least two requests are processed separately. Since each second request is used to indicate obtaining different second data, the second data indicated to be obtained by each second request is obtained separately from the storage node.

[0145] Exemplarily, the processing node obtains 10 first requests sent by multiple training nodes. Among them, 5 first requests are used to indicate obtaining first data, and 5 first requests are used to indicate obtaining second data. The processing node does not cache the first data and the second data, so it sends two second requests to the storage node. One second request is used to obtain the first data, and one second request is used to obtain the second data. Thus, the processing node and the storage node only need to transmit two requests to obtain the data indicated by the 10 first requests. By reducing the number of times of obtaining the same data, the amount of data transmitted is reduced, and the communication volume is effectively reduced.

[0146] Optionally, after the processing node obtains the data, it can cache the data, so that it is convenient for the processing node to read the data from the storage medium of the processing node as soon as possible next time.

[0147] In some other embodiments, if the storage medium of the processing node stores data, the processing node obtains the data indicated by one of the at least two requests from the storage medium of the processing node. For example, the processing node obtains a third request, determines whether the third data corresponding to the data identifier in the third request exists in the cache. If it exists, the third data is obtained from the cache and returned to the corresponding training node.

[0148] Step 950: The processing node determines the first data corresponding to the first data identifier from the second data.

[0149] Since the data identifiers included in the multiple first requests include at least one merged second data identifier included in at least one second request, the second data obtained by the processing node according to the second data identifier includes the first data indicated by the data identifiers included in the multiple first requests. The processing node determines the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.

[0150] Step 960, the processing node returns the first data to the corresponding training node.

[0151] After receiving multiple first requests sent by multiple training nodes, the processing node can also store the identifiers of the multiple training nodes. After the processing node obtains the second data indicated by at least one second request, according to the identifiers of the training nodes that sent the multiple first requests, the processing node feeds back the first data corresponding to the first data identifier determined from the second data to the training nodes that sent the multiple first requests.

[0152] The request described in this application can indicate a pull request or a push request. A pull request can mean that after the processing node obtains data, it transmits the data to the training node. A push request can mean that the training node transmits training results such as gradients to the processing node, and the processing node updates the data in the storage node according to the training results.

[0153] In the data processing method provided by this application, the training node can send multiple pull requests or multiple push requests to the processing node. The processing node sends the multiple pull requests to the storage node to obtain data; or, after the processing node performs unified deduplication on the multiple push requests, it sends them to the storage node to update the data, reducing the communication volume by 50%. The processing node can cache the obtained Emb vectors, further reducing cross-network communication and improving system performance.

[0154] Bind multiple training nodes and the processing node. The requests of the training nodes are aggregated by the processing node, reducing the cross-physical server communication volume and concurrency, improving system performance, and reducing system jitter.

[0155] Deploy the training node and the processing node on the same device. The training node and the processing node communicate based on Linux SHM, further improving communication performance.

[0156] It can be understood that, in order to implement the functions in the above embodiments, the computer device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the manner of hardware or computer software driving hardware depends on the specific application scenarios and design constraints of the technical solution.

[0157] In the foregoing, in combination with Figures 1 to 9 , the data processing method provided according to the present application has been described in detail. Next, in combination with Figure 10 , the device provided according to the present application will be described. These devices can be used to implement the functions of the general-purpose processor in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, the device can be a general-purpose processor as shown in Figure 7 , and can also be a module (such as a chip) applied to a computer device.

[0158] As Figure 10 shown, the data processing device 1000 includes a communication module 1001, a processing module 1002, and a storage module 1003.

[0159] The data processing device 1000 is used to implement the functions of the processing nodes in the method embodiment shown in the above Figure 9 .

[0160] The communication module 1001 is used to obtain multiple first requests of multiple training nodes, and each first request includes a first data identifier of the requested first data. For example, the communication module 1001 is used to execute Figure 9 step 910 in

[0161] . Figure 9 The processing module 1002 is used to merge the same data identifiers included in the multiple first requests to generate at least one second request. For example, the processing module 1002 is used to execute

[0162] step 920 and step 930 in Figure 9 .

[0163] The communication module 1001 is further used to send at least one second request to the storage node. For example, the communication module 1001 is used to execute Figure 9 step 940 in

[0164] The communication module 1001 is further configured to return the first data to the corresponding training node. For example, the communication module 1001 is configured to execute Figure 9 step 960 in

[0165] The storage module 1003 is used to store the Emb table, identifiers, etc., for facilitating the training model or the inference model.

[0166] It should be understood that the data processing device 1000 according to the embodiments of the present application may be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It may also be implemented by software Figure 9 When implementing the method shown, and its respective modules may also be software modules. The data processing device 1000 and its respective modules may also be software modules.

[0167] The data processing device 1000 according to the embodiments of the present application may correspond to executing the methods described in the embodiments of the present application, and the above and other operations and / or functions of each unit in the data processing device 1000 are respectively for implementing Figure 9 the corresponding processes of each method in

[0168] Figure 11 This is a schematic structural diagram of a computer device 1100 provided by the present application. As Figure 11 shown, the computer device 1100 includes a processor 1110, a bus 1120, a memory 1130, a communication interface 1140, a memory 1150 (which may also be referred to as a main memory unit), and a processor 1160. The processor 1110, the processor 1160, the memory 1130, the memory 1150, and the communication interface 1140 are connected through the bus 1120.

[0169] It should be understood that in this embodiment, the processor 1110 may be a CPU, and the processor 1110 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0170] The computer device 1100 may further include a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the solution of the present application. For example, the processor 1160 may be a GPU or an NPU.

[0171] The communication interface 1140 is used to implement the communication between the computer device 1100 and external devices or components.

[0172] In the present application, when the computer device 1100 is used to implement Figure 9 the functions of the processing nodes shown, the processor 1110 is used to obtain multiple requests sent by the processor 1160, etc., so that the processor 1110 is used to merge the same data identifiers included in the multiple first requests, obtain the data indicated by the merged second data identifier, and feed back the data to the processor 1160 that issues the request.

[0173] The bus 1120 may include a path for transmitting information between the above components (such as the processor 1110, the memory 1150, and the storage 1130). In addition to the data bus, the bus 1120 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, all kinds of buses are labeled as the bus 1120 in the figure. The bus 1120 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 1120 may be divided into an address bus, a data bus, a control bus, etc.

[0174] As an example, the computer device 1100 may include multiple processors. The processor may be a multi-CPU processor. Here, the processor may refer to one or more devices, circuits, and / or computing units for processing data (such as computer program instructions).

[0175] It is worth noting that Figure 11 only takes the computer device 1100 including 1 processor 1110 and 1 memory 1130 as an example. Here, the processor 1110 and the memory 1130 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to service requirements. For example, the computer device 1100 includes multiple GPUs or multiple NPUs.

[0176] The memory 1150 may be a volatile memory pool or a non-volatile memory pool, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). The memory 1150 is used to store the Emb table, identifiers, etc.

[0177] The memory 1130 may correspond to the storage medium for storing information such as the Emb table and identifiers in the above method embodiments. For example, a disk, such as a mechanical hard disk or a solid-state drive.

[0178] The above computer device 1100 may be a general-purpose device or a special-purpose device. For example, the computer device 1100 may also be a server or other devices with computing capabilities.

[0179] It should be understood that the computer device 1100 according to this embodiment may correspond to the data processing device 1000 in this embodiment, and may correspond to the corresponding entity in any of the methods according to Figure 9 and the above and other operations and / or functions of each module in the data processing device 1000 respectively implement the corresponding processes of each method in Figure 9 . For the sake of brevity, they will not be elaborated here.

[0180] The method steps in this embodiment can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), register, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a computing device. Of course, the processor and the storage medium can also exist as discrete components in a computing device.

[0181] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid state drive (SSD). As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, Applied to a data processing system, the data processing system includes processing nodes, multiple training nodes, and storage nodes. The training nodes are used to perform inference or training of a neural network model, and the storage nodes are used to store data used by the training nodes to perform inference or training of the neural network model; The method includes: The processing node obtains multiple first requests of multiple training nodes, and each first request includes a first data identifier of the requested first data; The processing node merges the same data identifiers included in the multiple first requests; The processing node generates at least one second request, and the at least one second request includes the merged second data identifier; The processing node sends the at least one second request to the storage node and obtains the second data corresponding to the second data identifier from the storage node; The processing node determines the first data corresponding to the first data identifier from the second data and returns the first data to the corresponding training node.

2. The method according to claim 1, wherein The processing node and the multiple training nodes are deployed on a training server. The processing node is executed by a general-purpose processor of the training server, and the training node is executed by a dedicated processor of the training server.

3. The method according to claim 1 or 2, characterized in that, The processing node further includes a cache for caching the second data obtained from the storage server. The method further includes: The processing node obtains a third request, determines whether the third data corresponding to the data identifier in the third request exists in the cache. If it exists, the third data is obtained from the cache and returned to the corresponding training node.

4. The method according to any one of claims 1-3, characterized in that, The processing node merging the same data identifiers included in the multiple first requests includes: The processing node retains one data identifier among the same data identifiers included in the multiple first requests to obtain the merged second data identifier.

5. The method according to any one of claims 1-4, characterized in that, The processing node determining the first data corresponding to the first data identifier from the second data includes: The processing node determines the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.

6. The method according to any one of claims 1-5, characterized in that, Returning the first data to the corresponding training node includes: The processing node returns the first data to the training node that issued the first data identifier.

7. A data processing system, characterized in that, Includes: Training nodes, storage nodes, and processing nodes. The training nodes are used to perform inference or training of a neural network model, the storage nodes are used to store data used by the training nodes to perform inference or training of the neural network model, and the processing nodes are used to perform the operation steps of the method according to any one of claims 1-6 above.

8. A computer-readable storage medium, characterized in that, Includes: Computer software instructions; when the computer software instructions run in a processor, the processor is caused to perform the operation steps of the method according to any one of claims 1-6 above.

9. A computer program product, characterized in that, Includes: When a computer program product runs on a computer, the computer is caused to perform the operation steps of the method according to any one of claims 1-6 above.