Data processing method and system, computer readable storage medium, and computer program product
By processing the same data identification in the node merge request and adopting a hierarchical storage architecture, the problem of excessive traffic during neural network model training and inference is solved, and more efficient data transmission and computing efficiency is achieved.
Patent Information
- Application Number
- PCT/CN2025/070267
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-31
AI Technical Summary
During the training and inference process of neural network models, the existing technology requires frequent pulling and pushing sparse parameters from remote storage devices, resulting in excessive communication volume, low computing efficiency and poor system performance.
By merging the same data identifiers in multiple requests at the processing node, duplicate data acquisition for the storage node is reduced, and a hierarchical storage architecture and caching mechanism are adopted to optimize the data transmission process.
It effectively reduces traffic, improves system performance, reduces system jitter, and improves computing efficiency.
Smart Images

Figure CN2025070267_31072025_PF_FP_ABST
Abstract
Description
Data processing method, system, computer-readable storage medium, and computer program product
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 25, 2024, with application number 202410115211.3 and application name “Data processing method, system, computer-readable storage medium and computer program product”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method, system, computer-readable storage medium, and computer program product. Background Art
[0003] Currently, general-purpose processors and dedicated processors perform computations for training and inference of neural network models, improving computational efficiency. For neural network models such as recommendation systems and natural language processing, only a subset of the model's parameters (e.g., sparse parameters) are required for a single computation. For example, the parameters are vectorized (embedding), and dedicated processors process the vector data. Typically, vector data is stored in remote storage devices, and general-purpose processors need to retrieve the vector data from the remote storage devices and feed it back to the dedicated processors, resulting in a high volume of communication. Summary of the Invention
[0004] The present application provides a data processing method, system, computer-readable storage medium, and computer program product, thereby effectively reducing communication traffic.
[0005] In a first aspect, a data processing method is provided, which is applied to a data processing system. The data processing system includes a processing node, multiple training nodes, and a storage node. The training node is used to perform inference or training of a neural network model, and the storage node is used to store data used by the training node to perform inference or training of the neural network model. The method includes: the processing node obtains a first request sent by each of the multiple training nodes, each first request including a first data identifier of the requested first data; the processing node merges the same data identifiers included in the multiple first requests to generate at least one second request, the at least one second request including the merged second data identifier; the processing node sends the at least one second request to the storage node and obtains second data corresponding to the second data identifier from the storage node; the processing node determines the first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding training node.
[0006] Compared to a processing node processing each of multiple requests, especially when multiple requests indicate obtaining the same data, the processing node needs to repeatedly obtain the same data, resulting in a large amount of communication traffic for processing the requests. The solution provided by the present application performs a deduplication operation on multiple requests, that is, merges the same data identifiers in multiple requests, obtains the data corresponding to the merged data identifiers from the storage node, and obtains the data corresponding to the same data identifier only once from the storage node for requests containing the same data identifier. Thus, by reducing the number of times the same data is obtained, the amount of data transmitted is reduced, and the communication traffic is effectively reduced.
[0007] In a possible implementation, the processing node and multiple training nodes are deployed on a training server, the processing node is executed by a general-purpose processor of the training server, and the training node is executed by a dedicated processor of the training server.
[0008] This reduces the amount of communication and concurrency across different servers, improves system performance, and reduces system jitter.
[0009] In another possible implementation, the processing node merges the same data identifiers included in the multiple first requests, including: the processing node retains one data identifier from the same data identifiers included in the multiple first requests to obtain a merged second data identifier.
[0010] The data identifier in the request is used to identify the data. If two requests contain the same data identifier, it means that the two requests are requesting the same data. By combining the data identifiers contained in the requests, duplicate requests are deduplicated, reducing the number of times the same data is retrieved and the amount of data transmitted, effectively reducing communication traffic.
[0011] In another possible implementation, the processing node determines the first data corresponding to the first data identifier from the second data, including: the processing node determines the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.
[0012] In another possible implementation, returning the first data to the corresponding training node includes: the processing node returning the first data to the training node corresponding to the one that issued the first data identifier.
[0013] In another possible implementation, the processing node also includes a cache for caching the second data obtained from the storage server. The method also includes: the processing node obtains a third request, determines whether the third data corresponding to the data identifier in the third request exists in the cache, and if so, obtains the third data from the cache and returns it to the corresponding training node.
[0014] Thus, the processing node obtains data from the cache, reducing the possibility that the processing node needs to obtain data from the remote storage device, and further reducing the amount of communication across different devices.
[0015] In a second aspect, a data processing method is provided, wherein a computing system used by the method includes a general-purpose processor and a special-purpose processor. The special-purpose processor is used to train or infer an artificial intelligence model. The method includes: the general-purpose processor obtains a first request sent by each of a plurality of special-purpose processors, each first request including a first data identifier of the requested first data; the general-purpose processor merges the same data identifiers included in the plurality of first requests to generate at least one second request, the at least one second request including the merged second data identifier; the general-purpose processor sends the at least one second request to a storage node and obtains second data corresponding to the second data identifier from the storage node; the general-purpose processor determines the first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding special-purpose processor.
[0016] A dedicated processor acquires data and performs training or inference calculations on the neural network model based on the data.
[0017] In a third aspect, a data processing device is provided, the data processing device including modules for executing the data processing method of the first aspect or any possible design of the first aspect. For example, the data processing device is configured to implement the functions of a general-purpose processor, and the data processing device includes a communication module and a processing module.
[0018] A communication module is used to obtain multiple first requests from multiple training nodes, each first request including a first data identifier of the requested first data; a processing module is used to merge the same data identifiers included in the multiple first requests; the processing module is also used to generate at least one second request, at least one second request including a merged second data identifier; the communication module is also used to send at least one second request to a storage node and obtain second data corresponding to the second data identifier from the storage node; the processing module is also used to determine the first data corresponding to the first data identifier from the second data; the communication module is also used to return the first data to the corresponding training node.
[0019] In a possible implementation, when the processing module merges the same data identifiers included in multiple first requests, it is specifically configured to: retain one data identifier among the same data identifiers included in the multiple first requests to obtain a merged second data identifier.
[0020] In another possible implementation, when the processing module determines the first data corresponding to the first data identifier from the second data, it is specifically configured to: determine the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.
[0021] In another possible implementation, the communication module, when used to return the first data to the corresponding training node, is specifically used to: return the first data to the training node corresponding to the first data identifier.
[0022] In another possible implementation, the processing node also includes a cache for caching the second data obtained from the storage server. The processing module is also used to obtain a third request and determine whether the third data corresponding to the data identifier in the third request exists in the cache. If so, the third data is obtained from the cache. The communication module is also used to return to the corresponding training node.
[0023] In a fourth aspect, a computing system is provided, comprising a general-purpose processor and multiple special-purpose processors, wherein the general-purpose processor performs the steps of the method of the first aspect or any possible implementation of the first aspect. The multiple special-purpose processors perform training calculations or inference calculations on a neural network model based on data.
[0024] In a fifth aspect, a data processing system is provided, which includes a training node, a processing node and a storage node. The training node is used to perform reasoning or training of a neural network model, the storage node is used to store data used by the training node to perform reasoning or training of the neural network model, and the processing node executes the operating steps of the method in the first aspect or any possible implementation of the first aspect.
[0025] In a sixth aspect, a computer device is provided, comprising a memory and multiple processors, the memory being used to store a set of computer instructions; when the processor executes the set of computer instructions, the processor executes the operating steps of the method in the first aspect or any possible implementation of the first aspect.
[0026] In the seventh aspect, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a processor, the processor executes the operating steps of the method described in the first aspect or any possible implementation of the first aspect.
[0027] In an eighth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to execute the operating steps of the method described in the first aspect or any possible implementation of the first aspect.
[0028] The technical effects brought about by any design method in the second to eighth aspects can be referred to the technical effects brought about by the first aspect or different design methods in the first aspect, and will not be repeated here.
[0029] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a schematic diagram of the structure of a neural network provided by this application;
[0031] FIG2 is a schematic diagram of a sparse parameter provided by this application;
[0032] FIG3 is a schematic diagram of one-hot encoding and vectorization provided by this application;
[0033] FIG4 is a schematic diagram of parameter storage provided by this application;
[0034] FIG5 is a schematic diagram of the architecture of a data processing system provided by the present application;
[0035] FIG6 is a schematic diagram of the architecture of a computing system provided by the present application;
[0036] FIG7 is a schematic diagram of data storage of an ultra-large-scale model provided by the present application;
[0037] FIG8 is a schematic diagram of a multi-level storage architecture provided by the present application;
[0038] FIG9 is a flow chart of a data processing method provided by the present application;
[0039] FIG10 is a schematic structural diagram of a data processing device provided by the present application;
[0040] FIG11 is a schematic structural diagram of a computer device provided in this application. DETAILED DESCRIPTION
[0041] To facilitate understanding, the main terms involved in this application are first explained.
[0042] Artificial Neural Network (ANN): also known as Neural Network (NN) or Neural Network-like. In the fields of machine learning and cognitive science, it is a mathematical or computational model that mimics the structure and function of biological neural networks (the central nervous system of animals, particularly the brain). A neural network is a network formed by connecting multiple individual neurons together, meaning that the output of one neuron can be the input of another. The input of each neuron can be connected to the local receptive field of the previous layer to extract features from that local receptive field, which can be an area consisting of several neurons.
[0043] Each node represents a specific output function, called an activation function. Each connection between two nodes represents a weighted value for the signal passing through that connection, called a weight, which acts as the memory of the artificial neural network. The output of the neural network varies depending on the network's connection structure, weight values, and activation function. Neural networks themselves are often approximations of natural algorithms or functions, or they may express a logical strategy.
[0044] As shown in Figure 1, it is a schematic diagram of the structure of a neural network provided in the present application. The neural network 100 includes N processing layers, where N is an integer greater than or equal to 3. The first layer of the neural network 100 is the input layer 110, which is responsible for receiving input signals, and the last layer of the neural network 100 is the output layer 130, which is responsible for outputting the processing results of the neural network. The other layers excluding the first and last layers are intermediate layers 140, and these intermediate layers 140 together constitute the hidden layer 120. Each intermediate layer 140 in the hidden layer 120 can both receive input signals and output signals. The hidden layer 120 is responsible for the processing of the input signal. Each layer represents a logical level of signal processing. Through multiple layers, the data signal can be processed by multiple levels of logic.
[0045] In some feasible embodiments, the input signal of the neural network can be a video signal, a voice signal, a text signal, an image signal, a temperature signal, an engineering signal that can be processed by a computer, or other signals in various forms.
[0046] Model training involves using a training set to train a neural network model, enabling it to predict or classify unknown data. During the training process, neural network models learn based on the features and target values in the training set. Upon completion, a neural network model is generated, which can be used to predict or classify unknown data. Neural network model training is one of the most critical steps in machine learning, impacting the accuracy and reliability of the neural network model.
[0047] Parameters: These are the variables or weights that need to be learned or adjusted in a neural network model. Parameters can influence the predictive power and performance of a neural network model. During neural network training, different parameter combinations are tried to optimize the model's performance. Common parameters include weights, biases, learning rates, and regularization coefficients. These parameters are used to calculate the output when the neural network model is used for prediction.
[0048] Sparse parameters: Parameters that are only partially involved in the calculation during a training process. For example, parameters that participate in both forward calculations and backward updates.
[0049] For example, Figure 2 is a schematic diagram of a sparse parameter provided by this application. As shown in Figure 2, the weight part involved in the calculation in one training step only includes half of the total weight.
[0050] Typically, sparse parameters are very large. For example, in a production-level recommendation system, the amount of sparse parameter data can reach 10TB to 30TB.
[0051] Recommendation system: A type of application that provides users with personalized decision support and information services based on massive data mining, and determines the items or services that users currently need or are interested in based on information such as their historical behavior, social relationships, points of interest, and the context in which they are located.
[0052] In some neural network models, such as recommendation systems and natural language processing, input data often contains discrete features. Embedding table parameters are used to convert the input data into continuous vector data before processing the vector data. During a single training pass, only a subset of the embedding table parameters are involved in the calculation and training update.
[0053] Embedding technology: A dense vector representation. It converts sparse parameters into dense vector data. Embedding can represent object features, such as height, gender, name, or item.
[0054] For example, as shown in Figure 3, a schematic diagram of one-hot encoding and vectorization provided by this application is shown. Embedding is equivalent to smoothing one-hot encoding, and one-hot encoding is equivalent to performing max pooling on embedding.
[0055] During the training of a neural network model, parameters (e.g., sparse parameters) are continuously adjusted so that the predicted values after the input data and parameter calculations are close to the actual values. The parameters can be loaded into the storage medium of a dedicated processor, which then calculates and updates the parameters. For example, dedicated processors include, but are not limited to, graphics processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), and embedded neural-network processing units (NPUs). The storage medium of a dedicated processor includes high-bandwidth memory (HBM).
[0056] For example, as shown in (a) in Figure 4, this is a parameter storage diagram provided by the present application. The Embedding table is entirely stored in the storage medium of the GPU (such as HBM). The storage medium of the GPU is divided into two parts, one part is used to store data during the training process of the neural network model; the other part is used to store the Embedding table. During the training process of the neural network model, the required Embedding vectors are pulled to the specified GPU through the All2All method to achieve the purpose of sharing the Embedding table. The Embedding table is referred to as the Embedding table. The Embedding table contains vector data of sparse parameters.
[0057] However, this solution places a high demand on GPUs. For example, approximately 300 GPU cards are required to store a 10TB Embed table, making neural network model training extremely expensive. Furthermore, the large number of GPUs required leads to poor system linearity, insufficient overall system performance, and poor computational reliability. All GPUs use an All2All algorithm to pull Embedding vectors for training. If a GPU fails, the data stored in the Embedding table on the failed GPU is easily lost.
[0058] The storage capacity of a dedicated processor is only 16GB to 32GB, which is far too little for tens of TB of sparse parameters. Sparse parameters can also be stored in storage media such as the server's volatile memory (such as RAM) or non-volatile memory (such as a solid state drive (SSD)).
[0059] In some embodiments, the sparse parameters are stored in a server including a general-purpose processor, which converts the sparse parameters into vector data, which is then transferred to a dedicated processor for computation.
[0060] For example, as shown in (b) in FIG4 , a parameter storage diagram is provided in the present application. The Emb table is entirely stored in the CPU cluster; for example, the Emb table is entirely stored in a storage medium such as a volatile memory (such as memory) or a non-volatile memory (such as a solid state drive (SSD)) in the CPU cluster. Among them, the CPU can provide the functions of storing, querying, and updating the Emb table. The CPU cluster expands the storage capacity of the GPU. The GPU trains the neural network model. The CPU and GPU collaborate to perform forward calculations and reverse updates of the neural network model; during forward calculations, the required Emb table is pulled from the CPU cluster through the pull interface; the GPU trains the neural network model according to the Emb table to obtain the gradient; the gradient is pushed back to the CPU cluster through the push interface, and the Emb table is updated in the CPU cluster.
[0061] Since a large number of Emb tables are transmitted through the network each time a neural network model is trained, that is, pull and push operations are performed, the computational efficiency is low.
[0062] In other embodiments, if the server cannot store the complete Emb table, the Emb table is stored in the server's storage medium and a remote storage device.
[0063] For example, as shown in (c) of Figure 4, a parameter storage diagram provided by this application is provided. When the storage capacity of the memory is insufficient, the SSD is also used to store the Embed table. The CPU can pull the required Embed table from the SSD and transfer the Embed table to the GPU.
[0064] For ultra-large-scale models, the Embed table data volume is extremely large, reaching, for example, 100TB. If the Embed table were stored on dedicated processors, thousands of dedicated processors would be required to store the Embed table. Using thousands of dedicated processors to train an ultra-large-scale model would be prohibitively expensive. If the Embed table were stored on a remote storage device, general-purpose processors would not only need to query and update the Embed table but also perform numerous pull and push operations. This places a heavy load on the general-purpose processors, consumes a lot of network resources, generates a large amount of communication traffic, and reduces the computational efficiency of the neural network model.
[0065] Moreover, different dedicated processors may require the same data when performing model training or model inference calculations. In this case, the general-purpose processor needs to perform pull and push operations on the same data multiple times, resulting in the consumption of more network resources and a large amount of communication traffic.
[0066] In order to solve the problem that pull and push operations need to be performed multiple times on the same data during inference or training of a neural network model, resulting in a large amount of communication, the present application provides a data processing method, namely, a processing node obtains a first request sent by each of a plurality of training nodes, and each first request includes a first data identifier of the requested first data; the processing node merges the same data identifiers included in the plurality of first requests to generate at least one second request, and the at least one second request includes the merged second data identifier; the processing node sends the at least one second request to a storage node, and obtains the second data corresponding to the second data identifier from the storage node; the processing node determines the first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding training node.
[0067] Compared with a processing node processing each of multiple requests, especially when multiple requests indicate obtaining the same data, the processing node needs to repeatedly obtain the same data and perform pull and push operations on the same data multiple times, resulting in a large amount of communication traffic for processing requests. The solution provided by the present application performs a deduplication operation on multiple requests, that is, merges the same data identifiers in multiple requests, obtains the data corresponding to the merged data identifiers from the storage node, and for requests containing the same data identifier, obtains the data corresponding to the same data identifier only once from the storage node, that is, performs a pull and push operation on the same data once, thereby reducing the number of times the same data is obtained and the amount of data transmitted, effectively reducing the communication traffic.
[0068] The implementation of the data processing method provided in this application is described in detail below with reference to the accompanying drawings.
[0069] FIG5 is a schematic diagram of the architecture of a data processing system provided by the present application. As shown in FIG5 , the data processing system 500 includes a client 510 , a computing cluster 520 , and a storage cluster 530 .
[0070] The computing cluster 520 includes multiple computing nodes 521. The multiple computing nodes 521 can be connected through network devices (such as switches, network cards, etc.) based on high-speed interconnection technology, so that the multiple computing nodes 521 can communicate with each other.
[0071] In some embodiments, the computing node 521 may include computing units with computing capabilities such as a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), and an embedded neural-network processing unit (NPU) to provide high-performance computing.
[0072] The computing cluster 520 further includes a control node 522. The control node 522 is used to manage and allocate tasks, and multiple computing nodes execute multiple tasks in parallel to increase the data processing rate.
[0073] In the present application, the control node 522 is also used to perform preprocessing operations on the data required for model training or model reasoning, and cooperate with multiple computing nodes 521 to manage data. For example, the control node 522 vectorizes the sparse parameters and converts them into dense vector data (such as: Emb table). Furthermore, when the control node 522 obtains the model training task or the model reasoning task, the vector data is stored in the designated storage medium according to the storage requirements and the storage resource characteristics of the system. The designated storage medium includes at least one of the storage medium associated with the computing node, the storage medium associated with the control node, or the storage node. That is, for neural network models of different sizes, due to the different amounts of data that need to be loaded, the control node 522 performs hierarchical storage on data with different storage requirements, that is, the data related to model training or model reasoning is stored in the storage medium of at least one of the computing node, the control node, or the storage node.
[0074] For example, for an ultra-large-scale model, when the control node 522 receives a model training task or a model inference task, it stores the vector data in a storage medium associated with a computing node, a storage medium associated with a control node, or a storage medium associated with a storage node, based on the storage requirements and the system's storage resource characteristics. That is, because the Emb table corresponding to an ultra-large-scale model has a very large amount of data, the Emb table can be divided into multiple data blocks, and the multiple data blocks can be stored hierarchically in a storage medium associated with a computing node, a storage medium associated with a control node, or a storage medium associated with a storage node.
[0075] When computing node 521 performs training calculations or inference calculations on a neural network model based on data, if the data is stored in a storage medium of computing node 521 , computing node 521 obtains data related to model training or model inference from the storage medium of computing node 521 .
[0076] If the data is stored in the storage medium of the control node 522 or the storage medium of the storage node 531, the control node 522 obtains the data from the storage medium of the control node 522 or the storage medium of the storage node 531, loads the data to the computing node 521, and the computing node 521 performs training calculations or inference calculations on the neural network model based on the data.
[0077] Optionally, if the data is stored in the storage medium of the control node 522 or the storage medium of the storage node 531 , the computing node 521 may also obtain data related to model training or model inference from the storage medium of the control node 522 or the storage medium of the storage node 531 .
[0078] In some embodiments, the control node 522 obtains a first request sent by each computing node 521 in a plurality of computing nodes 521, and each first request includes a first data identifier of the requested first data; the control node 522 merges the same data identifier included in the plurality of first requests to generate at least one second request, and at least one second request includes the merged second data identifier; the control node 522 sends the at least one second request to the storage node, and obtains the second data corresponding to the second data identifier from the storage node; the control node 522 determines the first data corresponding to the first data identifier from the second data, and returns the first data to the corresponding computing node 521.
[0079] The storage cluster 530 includes multiple storage nodes 531. A storage node 531 includes one or more controllers, a network card, and multiple hard disks. The hard disks are used to store data. The hard disks can be magnetic disks or other types of storage media, such as solid-state drives or shingled magnetic recording hard disks. The network cards are used to communicate with the computing nodes 521 included in the computing cluster 520. The controllers are used to write data to or read data from the hard disks based on read / write data requests sent by the computing nodes 521. During the data reading and writing process, the controllers need to convert the addresses carried in the read / write data requests into addresses that the hard disks can recognize.
[0080] The storage cluster 530 includes multiple storage nodes 531 and the computing cluster 520 includes multiple computing nodes 521 which are connected through network devices (such as switches, network cards, etc.) based on high-speed interconnection technology, so that the multiple computing nodes 521 and the multiple storage nodes 531 can communicate with each other.
[0081] In this application, the storage nodes 531 included in the storage cluster 530 can serve as the remote storage devices described herein. For example, the storage node can be a server, which includes a CPU and storage media. If the remaining storage capacity in the computing cluster is insufficient, the storage cluster 530 can provide storage expansion capabilities, namely, storing data related to model training or model inference, such as the Embed table required for model training or model inference.
[0082] Client 510 communicates with computing cluster 520 and storage cluster 530 via network 540. For example, client 510 sends a request to computing cluster 520 via network 540, requesting that computing cluster 520 perform model training or model inference. Network 540 can be an internal enterprise network (e.g., a local area network (LAN)) or the Internet. Client 510 can be a computer connected to network 540, also known as a workstation. Different clients can share network resources (e.g., computing resources and storage resources).
[0083] In some embodiments, client 510 is installed with a client program 511. Client 510 runs client program 511 to display a user interface (UI). User 550 operates the UI to submit a request. For example, user 550 operates the UI to submit a model training request or a model inference request. After receiving the request, control node 522 can obtain the data required for model training or model inference from storage cluster 530 and store the data in a designated storage medium based on the data storage requirements and the system's storage resource characteristics.
[0084] Optionally, the system administrator 560 can call the application platform interface (API) 512 or the command-line interface (CLI) interface 513 through the client 510 to configure system information, for example, the system information includes the storage strategy for data required for model training or model inference.
[0085] FIG5 is merely a schematic diagram, and the embodiments of the present application do not limit the device connection method, device quantity, and device form in the data processing system.
[0086] This application does not limit the deployment form of the above-mentioned computing nodes 521 and control nodes 522.
[0087] For example, the computing node 521 and the control node 522 can be deployed on the same server. The computing node can also be called a dedicated processor. The control node can be called a general-purpose processor. For example, a general-purpose processor can include a central processing unit (CPU). Dedicated processors include computing nodes with computing capabilities such as GPUs, DPUs, and NPUs to provide high-performance computing. Alternatively, the computing node can be called a device, a training card, or an accelerator card. The control node can be called a host. There is no limit on the number of dedicated processors and general-purpose processors contained in a server.
[0088] For example, computing node 521 and control node 522 can each be an independent server. Computing node 521 can be a dedicated server that provides a heterogeneous computing architecture to provide high-performance computing. Control node 522 can be an ordinary general-purpose server. Dedicated servers and general-purpose servers are interconnected through network devices based on high-speed interconnection technology. Dedicated servers include storage media, general-purpose processors, and multiple dedicated processors. Dedicated servers can be artificial intelligence servers. Multiple artificial intelligence servers are interconnected via a network, and multiple artificial intelligence servers form an AI cluster to implement model training and model inference calculations. General-purpose servers include general-purpose processors.
[0089] A general-purpose server is used to manage and allocate tasks, instructing multiple computing nodes to execute multiple tasks in parallel to increase data processing speed.
[0090] The general-purpose processor in the dedicated server is configured to store the vector data on a designated storage medium based on storage requirements and the system's storage resource characteristics. The designated storage medium includes at least one of the storage medium of the dedicated processor, the storage medium of the dedicated server, and the storage medium of a general-purpose server. The dedicated processor in the dedicated server is configured to perform training or inference calculations on the neural network model based on the data.
[0091] For very large-scale models, the general-purpose processor in the dedicated server is further used to store the vector data in a storage medium associated with a computing node, a storage medium associated with a control node, or a storage medium associated with a storage node.
[0092] It should be noted that the specific form of the storage medium described in this application is not limited. The storage medium includes a volatile memory pool or a non-volatile memory pool, or may include both volatile and non-volatile memory. For example, the storage medium of a dedicated processor includes at least one of HBM or SSD.
[0093] The following describes the hierarchical storage of data required for model training calculations or model inference calculations provided by this application.
[0094] Figure 6 is a schematic diagram of the architecture of a computing system provided by the present application. The computing system provides a hierarchical storage architecture. As shown in Figure 6, the computing system 600 includes one or more dedicated servers 610 and one or more general-purpose servers 620. The dedicated server 610 includes a general-purpose processor and multiple dedicated processors. The general-purpose processor and multiple dedicated processors can be deployed together on the same server. The storage medium in the dedicated processor serves as the primary storage layer, and the storage medium in the general-purpose processor serves as the secondary storage layer, that is, the storage medium in the general-purpose processor serves as local storage. The general-purpose server 620 serves as the tertiary storage layer, that is, the storage server 620 serves as remote storage.
[0095] The storage medium in the dedicated processor is used to store data during model training or model inference, as well as all or part of the data in the Emb table.
[0096] The storage medium in the general-purpose processor is used to store all or part of the data in the Emb table, gradient accumulation, optimizer data, cache data, etc.
[0097] General servers are used to store all or part of the Emb table's data, gradient accumulation, optimizer data, and cached data.
[0098] General-purpose processors can be used as processing nodes, and dedicated processors can be used as training nodes.
[0099] The general-purpose processor is used to perform preprocessing operations on the data required for model training or model inference. For example, sparse parameters are vectorized and converted into dense vector data to obtain an Embed table. The general-purpose processor is also used to store the Embed table in a designated storage medium based on the Embed table's storage requirements and the system's storage resource characteristics. The designated storage medium includes storage media in at least one of the aforementioned primary, secondary, and tertiary storage layers.
[0100] A general-purpose processor can run multiple worker processes, and each worker process can control a dedicated processor. For example, a worker can instruct a dedicated processor to perform model training or model inference calculations, pull the Emb table and transmit it to the dedicated processor, and then update the Emb table based on the gradient pushed back from the dedicated processor.
[0101] Optionally, the model described in this application may refer to a recommendation model.
[0102] Based on the above architecture, a unified hierarchical storage solution is proposed for medium-scale models, large-scale models, and ultra-large-scale models.
[0103] In a first possible implementation, the Emb table corresponding to a medium-sized model has a small amount of data and requires a small amount of storage capacity. The Emb table can be stored in a primary storage layer, that is, the Emb table is stored in a storage medium in a dedicated processor.
[0104] In the second possible implementation, the Embed table corresponding to a large-scale model has a large amount of data and requires a large storage capacity. For example, the Embed table can have a data volume of up to 10TB. If the Embed table is stored on the storage medium of a dedicated processor, assuming each dedicated processor has a storage capacity of 32GB, 300 dedicated processors would be required to start training, which is too many dedicated processors. Alternatively, the Embed table can be stored in the secondary storage tier, that is, on the storage medium of a general-purpose processor.
[0105] In a third possible implementation, the data volume of the Emb table corresponding to the ultra-large-scale model is very large, and the Emb table can be stored in a tertiary storage layer, that is, the Emb table is stored in a storage medium in a general server.
[0106] In some embodiments, the Emb table can be divided into multiple data blocks according to the number of general servers, and the multiple data blocks of the Emb table can be stored in the storage media of multiple general servers respectively. The way of dividing the Emb table is not limited in this application. The number of data blocks of the Emb table can also be less than the number of general servers, and the data blocks of the Emb table are stored in some general servers among the multiple general servers. The number of data blocks of the Emb table can also be greater than or equal to the number of general servers, and the data blocks of the Emb table are stored in the storage media of multiple general servers. One general server stores one data block, or one general server can also store multiple data blocks.
[0107] For example, as shown in FIG7 , a general-purpose processor may store an ultra-large-scale model Emb table in a general-purpose server.
[0108] After the general-purpose processor receives a model training task request or a model inference task request, the worker in the general-purpose processor obtains the dataset identifier from the request. The identifier can indicate the Emb vector in the Emb table. The general-purpose processor obtains the Emb vector corresponding to the identifier and transmits it to the dedicated processor.
[0109] The general processor reads the Emb vector corresponding to the identifier from the general server and transmits it to the dedicated processor.
[0110] Optionally, the storage medium in the general-purpose processor is used to store a portion of the Emb vectors in the full Emb table, and the storage medium (eg, HBM) in the dedicated processor is used to cache a portion of the Emb vectors in the full Emb table.
[0111] When training or reasoning about a neural network model, a dedicated processor can read the required Emb vector from the HBM and update the Emb vector in the HBM based on the gradient after training. Alternatively, after training is complete, a general-purpose processor can update the Emb table based on the gradient.
[0112] In some embodiments, the dedicated processor periodically writes the Emb vector in the HBM to the storage medium of the general-purpose processor according to a cycle / event trigger.
[0113] The tiered storage architecture expands the storage options for Embed tables, supporting tiered storage for models of varying sizes and multiple hardware deployments, enabling a naturally scalable Embed table storage solution. A single architecture supports three recommendation models: medium-scale, large-scale, and extra-large, eliminating the need to develop and deploy specific software for each scenario.
[0114] For example, for a medium-sized model, dedicated processors are expanded to divide the Embedding table into blocks and store them in the storage media (Embedding Storage) of multiple dedicated processors. Multiple Embedding Storages are synthesized into a complete Embedding table.
[0115] For example, for large-scale models, the Emb table is stored in the storage medium of the general-purpose processor, and the storage medium in the dedicated processor serves as a cache layer. The Emb table is read from the general-purpose processor on demand for model training or model inference.
[0116] For example, for extremely large models, the Emb table is stored in a remote storage device, and the storage medium in the dedicated processor serves as a cache layer, reading the Emb table from the remote storage device on demand for model training or model inference.
[0117] It should be noted that the training node can run a worker process to pull the Emb vector from the processing node, drive the dedicated processor to train the neural network model or infer the neural network model, or push the gradient back to the processing node. The processing node can run a worker process to parse the received request and pull the Emb vector from the subsequent storage node or processing node based on the identifier in the request and transmit it to the training node, instructing the training node to train the neural network model or infer the neural network model.
[0118] In some embodiments, when there are many training nodes, the training nodes can be divided into multiple groups. Requests from training nodes in a group are processed by a single processing node. The processing node can provide deduplication and caching services for the training nodes and ultimately transfer the Emb vectors to the storage media in the training nodes.
[0119] For example, as shown in Figure 8, a schematic diagram of a multi-level storage architecture provided by this application is provided. Multiple training nodes and processing nodes can be deployed together on the same device. For example, multiple training nodes and processing nodes can be deployed together on a dedicated server. Storage nodes are deployed on a general-purpose server.
[0120] The Emb layer is used to convert sparse parameters into dense vector data, and the DNN layer is used to train or infer a neural network model based on dense vector data.
[0121] Optionally, a processing node can be used to process requests sent by multiple training nodes. Training nodes and corresponding processing nodes in a group can be deployed on the same device, or training nodes and corresponding processing nodes in multiple groups can be deployed on the same device.
[0122] In practical applications, the Emb table can be stored by multiple general servers. The Emb table can be divided into multiple data blocks according to the number of general servers, and the multiple data blocks of the Emb table are stored in the storage media of the multiple general servers respectively.
[0123] The processing node can cache the Emb vector. Then the storage medium in the training node, the processing node, and the storage node can collaboratively store the Emb table.
[0124] A processing node is used to receive a request sent by a training node in a corresponding group, pull the Emb vector from the processing node or storage node, and transmit it to the training node.
[0125] In the present application, the requests sent by the training nodes in a group may contain the same data identifier. The processing node performs deduplication operations on multiple requests, that is, merges the same data identifiers in multiple requests, obtains the data corresponding to the merged data identifiers from the storage node, and for requests containing the same data identifier, obtains the data corresponding to the same data identifier only once from the storage node, that is, performs a pull operation and a push operation on the same data. Thus, by reducing the number of times the same data is obtained and the amount of transmitted data is reduced, the communication volume is effectively reduced.
[0126] Next, the data processing process is described in detail with reference to the accompanying drawings.
[0127] FIG9 is a flow chart of a data processing method provided by the present application. Here, the transmission of sparse parameters related to training calculations or inference calculations of a neural network model is mainly described. As shown in FIG9 , the method includes the following steps 910 to 960.
[0128] Step 910: The processing node obtains multiple first requests from multiple training nodes.
[0129] When multiple training nodes perform model training or model inference, the training nodes send requests to the processing nodes, and the processing nodes receive multiple requests from the multiple training nodes. The multiple requests received from the multiple training nodes can be collectively referred to as first requests.
[0130] The data identifiers included in the multiple requests sent by the multiple training nodes and the data identifiers included in the multiple requests received from the multiple training nodes can be collectively referred to as first data identifiers. The data indicated by the first data identifiers included in the multiple first requests can be collectively referred to as first data.
[0131] Each first request includes a first data identifier of the requested first data.
[0132] Optionally, the processing node obtains a first request sent by a training node, or the processing node obtains multiple first requests sent by a training node. This application does not limit the number of requests sent by the training node.
[0133] The first data identifiers included in the first requests sent by different training nodes may be the same or different. The first data identifiers included in multiple first requests sent by the same training node may be the same or different. It is understood that when a training node issues multiple requests, the multiple requests may indicate the acquisition of the same data, or the multiple requests may indicate the acquisition of different data. Alternatively, when two different training nodes issue multiple requests, the multiple requests may indicate the acquisition of the same data, or the multiple requests may indicate the acquisition of different data. In other words, the first data indicated by the multiple first requests may be the same or different.
[0134] The processing node can perform deduplication on multiple requests, that is, merge and deduplication the same data identifiers contained in multiple requests, which can reduce the communication volume of requests sent to the storage node. The following is an explanation of steps 920 and 930.
[0135] Step 920: The processing node merges the same data identifiers included in multiple first requests.
[0136] If the data identifiers contained in at least two of the multiple first requests are the same, then at least two requests indicate obtaining the same data, and the processing node merges the data identifiers contained in the at least two requests. Understandably, the processing node retains one data identifier from the same data identifiers contained in the at least two requests. If the same data identifier is divided into a group, the data identifiers contained in the multiple first requests can be divided into one or more groups, the data identifiers in each group are the same, and the data identifiers in different groups are different. The processing node can merge the same data identifiers in the group, that is, retain one data identifier from the same data identifiers included in the multiple first requests to obtain a merged second data identifier. The merged second data identifier can include one data identifier or multiple data identifiers. When the merged second data identifier can include multiple data identifiers, the multiple data identifiers are all different. The data identifiers included in the multiple first requests include the merged second data identifier. The merged second data identifier is a partial identifier among the data identifiers included in the multiple first requests.
[0137] The data indicated by the data identifier includes sparse parameters related to training calculations or inference calculations of a neural network model.
[0138] For example, a first number of requests among multiple first requests all include a first identifier, and the first identifier indicates first data, then the first number of requests are used to indicate the acquisition of the first data. A second number of requests among multiple first requests all include a second identifier, and the second identifier indicates second data, then the second number of requests are used to indicate the acquisition of the second data. The data to be acquired indicated by each request in the first number of requests is the same, the data to be acquired indicated by each request in the second number of requests is the same, and the data to be acquired indicated by the first number of requests and the second number of requests are different. The processing node merges the first number of requests all including the first identifier to obtain a first identifier, and the processing node merges the second number of requests all including the second identifier to obtain a second identifier, and the merged second data identifier includes the first identifier and the second identifier.
[0139] Step 930: The processing node generates at least one second request, where the at least one second request includes the merged second data identifier.
[0140] The processing node merges the same data identifiers included in the multiple first requests to generate at least one second request, and the at least one second request includes the merged second data identifier.
[0141] If the multiple first requests contain a set of identical data identifiers, that is, the multiple first requests contain the same data identifiers, that is, the multiple first requests are used to request the same data, the processing node generates a second request.
[0142] If multiple first requests contain multiple sets of identical data identifiers, i.e., the data identifiers contained in the multiple first requests are partially identical, i.e., the multiple first requests are for requesting multiple different data. The processing node merges the identical data identifiers contained in the multiple first requests to obtain multiple different data identifiers, and then generates multiple second requests, each of which contains one of the multiple different data identifiers.
[0143] It is understandable that the data identifiers included in the multiple first requests include the merged second data identifier included in at least one second request.
[0144] Step 940: The processing node sends at least one second request to the storage node, and obtains second data corresponding to the second data identifier from the storage node.
[0145] If the processing node does not cache the second data indicated by the at least one second request, the processing node retrieves the data indicated by the merged second data identifier contained in the at least one second request from the storage node. Understandably, for at least two requests indicating the acquisition of the same data, only one read operation is performed, while for requests indicating the acquisition of different data, the at least two requests are processed separately. Because each second request indicates the acquisition of different second data, the second data indicated by each second request is retrieved separately from the storage node.
[0146] For example, a processing node receives 10 first requests from multiple training nodes, five of which are for obtaining first data and five for obtaining second data. If the processing node doesn't cache the first and second data, it sends two second requests to the storage node: one for obtaining the first data and one for obtaining the second data. Consequently, the processing node and the storage node only need to transmit two requests to obtain the data indicated by the 10 first requests. This reduces the number of times the same data is obtained and the amount of data transmitted, effectively reducing communication traffic.
[0147] Optionally, after obtaining the data, the processing node may cache the data, so that the processing node can read the data from the storage medium of the processing node as quickly as possible next time.
[0148] In other embodiments, if the storage medium of the processing node stores data, the processing node obtains data indicated by one of the at least two requests from the storage medium of the processing node. For example, the processing node obtains the third request, determines whether the third data corresponding to the data identifier in the third request is in the cache, and if so, obtains the third data from the cache and returns it to the corresponding training node.
[0149] Step 950: The processing node determines the first data corresponding to the first data identifier from the second data.
[0150] Because the data identifiers included in the multiple first requests include the merged second data identifier included in at least one second request, the second data obtained by the processing node based on the second data identifier includes the first data indicated by the data identifiers included in the multiple first requests. The processing node determines, from the second data, the first data corresponding to the first data identifier that is identical to the second data identifier.
[0151] Step 960: Process the node and return the first data to the corresponding training node.
[0152] After receiving multiple first requests from multiple training nodes, the processing node may also store the identifiers of the multiple training nodes. After obtaining the second data indicated by at least one second request, the processing node determines, based on the identifiers of the training nodes that issued the multiple first requests, the first data corresponding to the first data identifier from the second data and feeds it back to the training nodes that issued the multiple first requests.
[0153] The requests described herein may refer to pull requests or push requests. A pull request may refer to a processing node acquiring data and transferring it to a training node. A push request may refer to a training node transferring training results, such as gradients, to a processing node, which then updates the data in a storage node based on the training results.
[0154] In the data processing method provided by this application, a training node can send multiple pull requests or multiple push requests to a processing node. The processing node then sends the multiple pull requests to a storage node to retrieve data. Alternatively, the processing node can perform unified deduplication on multiple push requests and then send them to the storage node to update data, reducing communication traffic by 50%. The processing node can cache the acquired Emb vectors, further reducing cross-network communication and improving system performance.
[0155] Bind multiple training nodes and processing nodes, aggregate the training node requests through the processing nodes, reduce the communication volume and concurrency across physical servers, improve system performance, and reduce system jitter.
[0156] The training nodes and processing nodes are deployed on the same device, and the communication between the training nodes and processing nodes is based on Linux SHM, further improving communication performance.
[0157] It is understood that in order to implement the functions in the above embodiments, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.
[0158] The data processing method provided by the present application is described in detail above in conjunction with Figures 1 to 9. The device provided by the present application will be described below in conjunction with Figure 10. These devices can be used to implement the functions of the general-purpose processor in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment. In this embodiment, the device can be a general-purpose processor as shown in Figure 7, or it can be a module (such as a chip) applied to a computer device.
[0159] As shown in FIG. 10 , the data processing device 1000 includes a communication module 1001 , a processing module 1002 , and a storage module 1003 .
[0160] The data processing device 1000 is used to implement the functions of the processing node in the method embodiment shown in FIG. 9 .
[0161] The communication module 1001 is configured to obtain multiple first requests from multiple training nodes, each of which includes a first data identifier of the requested first data. For example, the communication module 1001 is configured to execute step 910 in FIG9 .
[0162] The processing module 1002 is configured to merge the same data identifiers included in the multiple first requests to generate at least one second request. For example, the processing module 1002 is configured to execute steps 920 and 930 in FIG9 .
[0163] The communication module 1001 is further configured to send at least one second request to the storage node. For example, the communication module 1001 is configured to execute step 940 in FIG9 .
[0164] The processing module 1002 is configured to determine the first data corresponding to the first data identifier from the second data. For example, the processing module 1002 is configured to execute step 950 in FIG9 .
[0165] The communication module 1001 is further configured to return the first data to the corresponding training node. For example, the communication module 1001 is configured to execute step 960 in FIG9 .
[0166] The storage module 1003 is used to store Emb tables, identifiers, etc. to facilitate model training or model inference.
[0167] It should be understood that the data processing device 1000 of the embodiment of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Alternatively, when the method shown in FIG. 9 is implemented by software, the data processing device 1000 and its modules can also be software modules.
[0168] According to the data processing device 1000 of the embodiment of the present application, it can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each unit in the data processing device 1000 are respectively for implementing the corresponding processes of each method in Figure 9. For the sake of brevity, they are not repeated here.
[0169] Figure 11 is a schematic diagram of the structure of a computer device 1100 provided in this application. As shown in Figure 11, computer device 1100 includes a processor 1110, a bus 1120, a memory 1130, a communication interface 1140, a memory 1150 (also referred to as a main memory unit), and a processor 1160. Processor 1110, processor 1160, memory 1130, memory 1150, and communication interface 1140 are connected via bus 1120.
[0170] It should be understood that in this embodiment, the processor 1110 may be a CPU, but may also be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0171] The computer device 1100 may also include a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application. For example, the processor 1160 may be a GPU or an NPU.
[0172] The communication interface 1140 is used to implement communication between the computer device 1100 and external devices or components.
[0173] In the present application, when the computer device 1100 is used to implement the functions of the processing node shown in FIG9 , the processor 1110 is used to obtain multiple requests sent by the processor 1160. The processor 1110 is used to merge the same data identifiers included in the multiple first requests, obtain the data indicated by the merged second data identifier, and feed the data back to the processor 1160 that issued the request.
[0174] The bus 1120 may include a path for transmitting information between the above-mentioned components (such as the processor 1110, the memory 1150, and the storage 1130). In addition to the data bus, the bus 1120 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 1120 in the figure. The bus 1120 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus 1120 can be divided into an address bus, a data bus, a control bus, etc.
[0175] As an example, computer device 1100 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions).
[0176] It is worth noting that FIG11 only uses a computer device 1100 including one processor 1110 and one memory 1130 as an example. Here, the processor 1110 and the memory 1130 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined based on business requirements. For example, the computer device 1100 may include multiple GPUs or multiple NPUs.
[0177] Memory 1150 may be a volatile memory pool or a nonvolatile memory pool, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 1150 is used to store Embed tables, flags, and the like.
[0178] The memory 1130 may correspond to a storage medium used to store information such as the Emb table and identifiers in the above method embodiment, for example, a disk such as a mechanical hard disk or a solid-state drive.
[0179] The computer device 1100 may be a general-purpose device or a dedicated device. For example, the computer device 1100 may be a server or other device with computing capabilities.
[0180] It should be understood that the computer device 1100 according to this embodiment may correspond to the data processing device 1000 in this embodiment, and may correspond to executing the corresponding subject according to any method in Figure 9, and the above-mentioned and other operations and / or functions of each module in the data processing device 1000 are respectively for implementing the corresponding processes of each method in Figure 9. For the sake of brevity, they will not be repeated here.
[0181] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist as discrete components in a computing device.
[0182] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that, Applied to a data processing system, the data processing system includes processing nodes, multiple training nodes, and storage nodes. The training nodes are used to perform inference or training of a neural network model, and the storage nodes are used to store data used by the training nodes to perform inference or training of the neural network model; The method includes: The processing node obtains multiple first requests of multiple training nodes. Each first request includes a first data identifier of the requested first data; The processing node merges the same data identifiers included in the multiple first requests; The processing node generates at least one second request, and the at least one second request includes the merged second data identifier; The processing node sends the at least one second request to the storage node and obtains the second data corresponding to the second data identifier from the storage node; The processing node determines the first data corresponding to the first data identifier from the second data and returns the first data to the corresponding training node.
2. The method according to claim 1, wherein The processing node and the multiple training nodes are deployed on a training server. The processing node is executed by a general-purpose processor of the training server, and the training node is executed by a dedicated processor of the training server.
3. The method according to claim 1 or 2, characterized in that, The processing node further includes a cache for caching the second data obtained from the storage server. The method further includes: The processing node obtains a third request and determines whether the third data corresponding to the data identifier in the third request exists in the cache. If it exists, the third data is obtained from the cache and returned to the corresponding training node.
4. The method according to any one of claims 1-3, characterized in that, The processing node merging the same data identifiers included in the multiple first requests includes: The processing node retains one of the same data identifiers included in the multiple first requests to obtain the merged second data identifier.
5. The method according to any one of claims 1-4, characterized in that, The processing node determining the first data corresponding to the first data identifier from the second data includes: The processing node determines the first data corresponding to the first data identifier that is the same as the second data identifier from the second data.
6. The method according to any one of claims 1-5, characterized in that, Returning the first data to the corresponding training node includes: The processing node returns the first data to the training node that issued the first data identifier.
7. A data processing system, characterized in that, Includes: Training nodes, storage nodes, and processing nodes. The training nodes are used to perform inference or training of a neural network model, the storage nodes are used to store data used by the training nodes to perform inference or training of the neural network model, and the processing nodes are used to perform the operation steps of the method according to any one of claims 1-6 above.
8. A computer-readable storage medium, characterized in that, Includes: Computer software instructions; when the computer software instructions run on a processor, the processor is caused to perform the operation steps of the method according to any one of claims 1-6 above.
9. A computer program product, characterized in that, Includes: When a computer program product runs on a computer, the computer is caused to perform the operation steps of the method according to any one of claims 1-6 above.
Citation Information
Patent Citations
Model training method and device, and domain name detection method and device
CN112926647A
Training method and device of graph neural network model, storage medium and electronic device
CN116910568A
Ticket embedding based on multi-dimensional it data
US20230186190A1
Method and device for training neural network
WO2020199914A1