System and method for vectorizing data in static data processing system
By designing a vectorized system in a static data processing system, the problem of high memory demand and incompatibility of static parallel strategies and dynamic processing in large-scale Embedding is solved, and efficient data processing and flexible storage solutions are achieved.
Patent Information
- Application Number
- CN202210642059.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2042-06-07
AI Technical Summary
In deep learning, existing large-scale Embedding implementation solutions face problems such as excessive memory demand, reduced communication efficiency and incompatible static parallel strategies with dynamic processing, resulting in waste of resources and increased difficulty in using them.
A vectorization system for static data processing systems is designed, and the vocabulary numbers are divided into multiple initial data vocabulary shards through the vocabulary division component, and the data vectorization forward and backward components are deployed. These components include multiple streamed executors for deduplication processing, data combing, prefetching and vectorized querying, enabling flexible hierarchical storage of data and multi-stage pipeline processing.
It realizes efficient processing of large-scale Embedding in static parallel data processing systems, reduces memory requirements, improves computing efficiency and flexibility in data exchange, supports dynamic encoding and built-in encoding, and does not require user manual encoding IDs.
Smart Images

Figure CN114996012B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a data processing technology. More specifically, the present disclosure relates to a system and method for performing vectorization on data during static data processing. Background Art
[0002] Nowadays, with the popularization of deep learning, more and more models and larger and larger amounts of data have made it impossible to implement the training of deep learning on a single computing device. For this reason, people have proposed distributed computing. With the popularization of distributed computing, large jobs or large tensors will be processed by splitting different parts of the data and deploying them to each computing device of different distributed data processing systems, and intermediate parameters need to be interacted during the calculation of each part.
[0003] In deep learning, the Embedding vector is an important module in large-scale deep recommendation systems. The implementation of large-scale Embedding is also one of the core problems that large-scale deep recommendation systems need to solve. There are mainly two technical methods: one is large-scale Embedding based on model parallelism, and the other is CPU / GPU hybrid Embedding. The following problems exist in these solutions: On the one hand, the model parallelism scheme needs to put all parameters into the video memory. As the scale expands, the more GPUs are required. For example, in the Criteo 1T dataset, there are approximately 1e9 feature IDs. If the embedding dimension (embedding_dims) is configured to 128, then a total of 512 GB of video memory is required to accommodate the Embedding parameters, which is equivalent to at least 13 A100 40G GPUs. If the adam optimizer is used, this value will become 3 times the previous one. According to the documents released by Baidu, Kuaishou, etc., the current adjusted scale is several orders of magnitude higher than the above. With the increase in the number of GPUs, it will bring higher network hardware requirements and cost requirements, and at the same time, it will also lead to a decrease in communication efficiency (frequent data interaction); compared with the large demand for video memory, the recommendation network often has a relatively low demand for computing. For the dataset of the above scale, in the latest results of MLPerf, using 64 80G A100s can complete training in 2 minutes. Another feature of the deep recommendation network is that it can deduplicate the IDs of the input vocabulary, which can significantly reduce the communication and computing requirements of Embedding. However, deduplication will produce dynamic results and cannot be well compatible with SBP (an existing static parallel strategy mechanism). Another problem caused by dynamics is memory amplification. Since the static parallel data processing system (e.g., OneFlow Lazy) uses static memory allocation, memory must be allocated according to the worst-case scenario. For example, each co-processor (e.g., a rank or GPU) has to send the deduplicated IDs to another rank to query the embedding. For example, if the number of IDs before deduplication is N, then on the destination rank, memory has to be allocated according to N * num_ranks, which will inevitably increase the memory requirement.However, the data distribution in the deep recommendation network is often not uniformly distributed, but follows a Power Law distribution. The CPU / GPU hybrid Embedding solution is designed based on this characteristic, deploying high-frequency IDs on the GPU and low-frequency IDs on the CPU. By expanding the Embedding scale through CPU memory, this solution also faces a similar compatibility issue between flexibility and staticity as before, and also requires users to first count the frequencies within the entire dataset, increasing the usage difficulty.
[0004] Therefore, there is a need for a solution that can utilize the advantages of large-scale parallel data processing of a static parallel data processing system (such as the OneFlow deep learning system) to solve the problem of large-scale Embedding, achieve flexible hierarchical storage, implement a multi-level pipeline for data loading, data exchange, table lookup, and calculation, and support built-in encoding without the need for the IDs input by the user to be continuously encoded within a certain range. Summary of the Invention
[0005] To this end, an object of the present invention is to solve at least the above problems. Specifically, the present disclosure provides a system for vectorizing data in a static data processing system, including: a vocabulary partitioning component for evenly partitioning the numbers of a batch of input vocabularies into multiple initial data vocabulary shards equal to the number of multiple coprocessors running in parallel and allocating and inputting them to the corresponding coprocessors, and a data vectorization forward component and a data vectorization backward component deployed corresponding to each coprocessor. The data vectorization forward component includes multiple data vectorization forward executors deployed in a streaming manner, and the data vectorization backward component includes multiple data vectorization backward executors deployed in a streaming manner. Wherein the multiple data vectorization forward executors perform streaming parallel processing of different stages on the input continuous batches of vocabularies, and include: a first deduplication forward executor for performing deduplication processing on the received initial vocabulary shard for the feature sequence numbers to obtain a first non-duplicate data vocabulary shard used by the local coprocessor and pre-storing a first data vocabulary shard for restoration; a vocabulary sorting forward executor for sorting out the data vocabulary that does not belong to the local coprocessor from the local first non-duplicate data vocabulary shard based on a predetermined segmentation rule according to the sequence number of the data and sending it to the coprocessor to which it belongs, and receiving the data vocabulary sent from the vocabulary sorting forward executors of other coprocessors, so as to form a first to-be-processed vocabulary shard belonging to the local coprocessor in this batch of vocabularies based on the predetermined segmentation rule; a second deduplication forward executor for performing deduplication processing on the first to-be-processed vocabulary when receiving the message that the vocabulary sorting forward executor has completed the exchange, to obtain a second non-duplicate data to-be-processed vocabulary shard used by the local coprocessor and pre-storing a second to-be-processed vocabulary shard for restoration; a vocabulary data prefetch forward executor for, when obtaining the message that it has completed the deduplication from the second deduplication forward executor, prefetching a predetermined number of vocabulary data according to the predetermined prefetch quantity based on the second non-duplicate data to-be-processed vocabulary shard used by the local coprocessor; a vocabulary vectorization query forward executor for, when obtaining the message that it has completed the prefetch from the vocabulary data prefetch forward executor and the message that the representation vectors of the previous batch of vocabularies have been updated and completed by the data vectorization backward component, performing vectorization query processing on the predetermined number of vocabulary data based on a predetermined vectorization rule, so as to obtain a predetermined number of representation vector shards corresponding to the predetermined number of vocabulary data fetched; a second restoration forward executor for, after obtaining the message that it has completed the vectorization query processing from the vocabulary vectorization query forward executor, performing corresponding vectorization restoration on the corresponding duplicate vocabularies in the pre-stored second to-be-processed vocabulary shard based on the pre-stored second to-be-processed vocabulary shard and the predetermined number of representation vector shards, to obtain a second duplicate representation vector shard corresponding to the pre-stored second to-be-processed vocabulary shard. Thus, the second duplicate representation vector shard and the predetermined number of representation vector shards form a second restored representation vector shard;The inverse-combing forward executor of the characterization vector, when obtaining the message of its completed restoration from the second restoration forward executor, returns and sends the characterization vectors corresponding to the data vocabulary obtained by the vocabulary-combing forward executor in the second restored characterization vector slices from other coprocessors to the other coprocessors and deletes them locally, and receives the characterization vectors corresponding to the data vocabulary sent out by the local vocabulary-combing forward executor returned from the inverse-combing forward executor of the characterization vector of other coprocessors, to form a first non-repetitive characterization vector slice corresponding to the first non-repetitive data vocabulary slice used by the local coprocessor; and the first restoration forward executor, after obtaining the message of its completed characterization vector exchange from the inverse-combing forward executor of the characterization vector, based on the pre-stored first data vocabulary slice for restoration and the first non-repetitive characterization vector slice, performs corresponding vectorized restoration on the corresponding repetitive vocabulary in the pre-stored first data vocabulary slice for restoration, to obtain a first repetitive characterization vector slice corresponding to the pre-stored first data vocabulary slice for restoration, whereby the first repetitive characterization vector slice and the first non-repetitive characterization vector slice form a first restored characterization vector slice.
[0006] A system for performing vectorization on data in a static data processing system according to the present disclosure, wherein the data vectorization backward component includes: a vector difference executor, which performs difference processing on the first restored characterization vector slice processed by the loss function executor; a first merge (Reduce) backward executor, which performs deduplication processing on the characterization vector difference result output by the vector difference executor based on the serial number of the pre-stored first data vocabulary slice for restoration, to obtain a first non-repetitive characterization vector difference result; a characterization vector difference combing backward executor, based on the serial number of the pre-stored second data vocabulary slice for restoration to be processed, splits the characterization vector difference result that does not belong to the local coprocessor from the first non-repetitive characterization vector difference result and sends it to the corresponding other coprocessors, and receives the characterization vector difference result that belongs to the local coprocessor sent by the characterization vector difference combing backward executor of other coprocessors, so as to form a second non-repetitive characterization vector difference result that belongs to the local coprocessor; a characterization vector value update backward executor, which updates the value of the second non-repetitive characterization vector difference result based on the value of the characterization vector obtained by the vocabulary vectorization query forward executor and a predetermined learning frequency; and a characterization vector update backward executor, which is used to update the initial data vocabulary slice by using the updated value of the second non-repetitive characterization vector difference result and the context of a predetermined number of pre-fetched vocabulary data.
[0007] A system for performing vectorization on data in a static data processing system according to the present disclosure, wherein the vocabulary partitioning component, the first deduplication forward execution body, the vocabulary sorting forward execution body, the vocabulary data prefetch forward execution body, the vocabulary vectorization query forward execution body, and the data vectorization backward component are deployed in four parallel task streams.
[0008] A system for performing vectorization on data in a static data processing system according to the present disclosure, wherein the vocabulary sorting forward execution body, the representation vector inverse sorting forward execution body, and the representation vector difference sorting backward execution body perform data exchange between different coprocessors through the NCCL communication method.
[0009] According to another aspect of the present disclosure, there is provided a method for performing vectorization on data in a static data processing system, including: dividing the numbers of a batch of input word lists evenly into a plurality of initial data word list shards equal to the number of multiple coprocessors running in parallel by a word list division component and allocating and inputting them to corresponding coprocessors, wherein data vectorization forward components and data vectorization backward components are deployed corresponding to each coprocessor, the data vectorization forward component includes a plurality of data vectorization forward execution bodies deployed in a streaming manner, and the data vectorization backward component includes a plurality of data vectorization backward execution bodies deployed in a streaming manner. The plurality of data vectorization forward execution bodies perform streaming parallel processing in different stages on consecutive batches of input word lists, and include: performing duplicate removal of feature serial numbers on the received initial word list shards by a first duplicate removal forward execution body to obtain a first non-duplicate data word list shard used by the local coprocessor and pre-storing a first data word list shard for restoration; dividing out the data word list that does not belong to the local coprocessor from the local first non-duplicate data word list shard based on a predetermined segmentation rule according to the serial number of the data by a word list sorting forward execution body and sending it to the coprocessor to which it belongs, and receiving the data word list sent from the word list sorting forward execution bodies of other coprocessors, so as to form a first to-be-processed word list shard belonging to the local coprocessor in this batch of word lists based on the predetermined segmentation rule; performing duplicate removal processing on the first to-be-processed word list when receiving the message that the word list sorting forward execution body has completed the exchange by a second duplicate removal forward execution body to obtain a second non-duplicate data to-be-processed word list shard used by the local coprocessor and pre-storing a second to-be-processed word list shard for restoration; when obtaining the message that it has completed duplicate removal from the second duplicate removal forward execution body by a word list data prefetch forward execution body, prefetching a predetermined number of word list data according to the predetermined prefetch quantity based on the second non-duplicate data to-be-processed word list shard used by the local coprocessor; performing vectorization query processing on the predetermined number of word list data based on a predetermined vectorization rule by a word list vectorization query forward execution body when obtaining the message that it has completed prefetching from the word list data prefetch forward execution body and the message that the representation vectors of the previous batch of word lists have been updated and completed by the data vectorization backward component, so as to obtain a predetermined number of representation vector shards corresponding to the predetermined number of word list data prefetched; after obtaining the message that it has completed vectorization query processing from the word list vectorization query forward execution body by a second restoration forward execution body, performing corresponding vectorization restoration on the corresponding duplicate word lists in the second to-be-processed word list shard for restoration pre-stored based on the second to-be-processed word list shard for restoration pre-stored and the predetermined number of representation vector shards to obtain a second duplicate representation vector shard corresponding to the second to-be-processed word list shard for restoration pre-stored, whereby the second duplicate representation vector shard and the predetermined number of representation vector shards form a second restored representation vector shard;The inverse-combing forward execution body of the representation vector, when obtaining the message of its completed restoration from the second restoration forward execution body, returns and sends the representation vectors corresponding to the data vocabulary obtained by the vocabulary-combing forward execution body in the second restored representation vector slices from other coprocessors to the other coprocessors and deletes them locally, and receives the representation vectors corresponding to the data vocabulary sent out by the local vocabulary-combing forward execution body returned from the inverse-combing forward execution body of the representation vector of other coprocessors, to form a first non-repetitive representation vector slice corresponding to the first non-repetitive data vocabulary slice used by the local coprocessor; and after the first restoration forward execution body obtains the message of its completed representation vector exchange from the inverse-combing forward execution body of the representation vector, based on the pre-stored first data vocabulary slice for restoration and the first non-repetitive representation vector slice, performs corresponding vectorized restoration on the corresponding repetitive vocabulary in the pre-stored first data vocabulary slice for restoration, to obtain a first repetitive representation vector slice corresponding to the pre-stored first data vocabulary slice for restoration, whereby the first repetitive representation vector slice and the first non-repetitive representation vector slice form a first restored representation vector slice.;
[0010] According to the method for vectorizing data in a static data processing system according to the present disclosure, wherein the data vectorization backward component: performs difference processing on the first restored representation vector slice processed by the loss function execution body through the vector difference execution body; performs deduplication processing on the representation vector difference result output by the vector difference execution body by the first merging backward execution body based on the serial number of the pre-stored first data vocabulary slice for restoration, to obtain a first non-repetitive representation vector difference result; based on the serial number of the pre-stored second data vocabulary slice to be processed for restoration, the representation vector difference combing backward execution body splits out the representation vector difference result that does not belong to the local coprocessor from the first non-repetitive representation vector difference result and sends it to the corresponding other coprocessors, and receives the representation vector difference result that belongs to the local coprocessor sent by the representation vector difference combing backward execution body of other coprocessors, so as to form a second non-repetitive representation vector difference result that belongs to the local coprocessor; updates the value of the second non-repetitive representation vector difference result through the representation vector value update backward execution body based on the value of the representation vector obtained by the vocabulary vectorization query forward execution body and a predetermined learning frequency; and updates the initial data vocabulary slice through the representation vector update backward execution body using the updated value of the second non-repetitive representation vector difference result and the context of the pre-fetched predetermined number of vocabulary data.
[0011] According to the method for vectorizing data in a static data processing system disclosed in the present invention, the word list partitioning component, the first deduplication forward execution body and the word list combing forward execution body, the word list data pre-fetching forward execution body and the word list vectorization query forward execution body, and the data vectorization backward component are deployed in four parallel task flows.
[0012] According to the method for vectorizing data in a static data processing system disclosed in the present invention, the vocabulary combing forward execution body, the representation vector back combing forward execution body, and the representation vector differential combing backward execution body exchange data between different coprocessors through the NCCL communication method.
[0013] By using the system and method for vectorizing data in a static data processing system according to the present disclosure, the problem of large-scale embedding can be better solved, flexible hierarchical storage can be achieved, and the embedding table can be placed on the GPU memory, host memory, or SSD. The current memory price (RMB) is about 1750000 / TB according to A100, the host memory price is about 64000 / TB, and the SSD price is about 1000 / TB. High-speed devices are used as caches for low-speed devices to achieve a balance between speed and capacity. By introducing a series of customized executors (such as Op) in the disclosed system, flexible data deduplication and exchange processing is achieved within the executor, thereby bypassing the static parallel strategy SBP and static limitations. In addition, the message mechanism of the task stream (Stream) and the executor (Actor) of OneFlow is fully utilized to realize data loading, data exchange, table lookup, and calculation of multi-level pipelines. And by introducing the prefetch mechanism, the data exchange between the cache and the underlying storage is scheduled to overlap with the forward and backward calculations. In addition, the present disclosure also supports built-in encoding, and does not require the ID of the vocabulary data entered by the user to be continuously encoded within a certain range, that is, it provides the function of dynamically inserting new feature IDs.
[0014] Other advantages, objectives and features of the present invention will be reflected in part through the following description, and in part will be understood by those skilled in the art through research and practice of the present invention. Brief Description of the Figures
[0015] Figure 1 Shown is a schematic diagram of a system for performing vectorization on data in a static data processing system according to the present disclosure.
[0016] Figure 2The figure shows a schematic process diagram of data exchange and duplicate removal between a vocabulary sorting forward executor and a second duplicate removal forward executor for a system that performs vectorization on data in a static data processing system, and forming a to-be-processed vocabulary shard with second non-duplicate data.
[0017] Figure 3 The figure shows a schematic process diagram of data prefetching by a vocabulary data prefetching forward executor 124 of a system that performs vectorization on data in a static data processing system.
[0018] Figure 4 The figure shows a schematic process diagram of data prefetching by a vocabulary vectorization query forward executor 125 of a system that performs vectorization on data in a static data processing system.
[0019] Figure 5 The figure shows a schematic diagram of the restoration process of a second restoration forward executor 126 and the reverse combing process of a characterization vector reverse combing forward executor 127 of a system that performs vectorization on data in a static data processing system.
[0020] Figure 6 The figure shows a schematic diagram of the combing process of a first merging backward executor 132 and a characterization vector differential combing backward executor 133 of a system that performs vectorization on data in a static data processing system.
[0021] Figure 7 The figure shows a timing schematic diagram of a parallel task flow of a system that performs vectorization on data in a static data processing system. Detailed implementation manners
[0022] The following further elaborates on the present invention in conjunction with embodiments and the accompanying drawings, so that those skilled in the art can implement it with reference to the description in the specification.
[0023] Here, exemplary embodiments will be described in detail, and their examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0024] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0025] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, one of two possible executors may be referred to as the first executor or the second executor hereinafter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0026] To enable those skilled in the art to better understand this disclosure, the following further describes this disclosure in detail in conjunction with the accompanying drawings and specific embodiments.
[0027] Figure 1 Shown is a schematic diagram of the principle of an example of a system for vectorizing data in a static data processing system according to this disclosure. As Figure 1 shown, a system 100 for vectorizing data in a static data processing system includes: a vocabulary partitioning component 110, a data vectorization forward component 120, and a data vectorization backward component 130. The data vectorization forward component 120 includes: a first deduplication forward executor 121, a vocabulary sorting forward executor 122, a second deduplication forward executor 123, a vocabulary data prefetch forward executor 124, a vocabulary vectorization query forward executor 125, a second restoration forward executor 126, a characterization vector inverse sorting forward executor 127, and a first restoration forward executor 128. The vocabulary partitioning component 110 evenly partitions the numbers of a batch of input vocabularies into multiple initial data vocabulary shards equal to the number of multiple coprocessors running in parallel and assigns them to the corresponding coprocessors. Corresponding to each coprocessor, data vectorization forward components 120-0, 120-1... 120-n, etc. and data vectorization backward components 130-0, 130-1... 130-n are deployed. The data vectorization forward component 120 includes multiple data vectorization forward executors deployed in a streaming manner, and the data vectorization backward component 130 includes multiple data vectorization backward executors deployed in a streaming manner. The multiple data vectorization forward executors perform different stages of streaming parallel processing on the input consecutive batches of vocabularies.
[0028] For the Embedding operation with parallel_num = n in parallel, a huge vocabulary is placed on different coprocessors (ranks or GPUs) according to certain partitioning rules. The number or sequence number (id) of the input vocabulary data for each batch is arranged in S0 (the 0th - dimensional split parallel distribution in the SBP parallel decision scheme). Therefore, different ids belong to different ranks according to the predetermined partitioning rules. However, after the data in the vocabulary is encoded after partitioning, features in one rank may be repeated and the ids of the features may not belong to their current rank according to the partitioning rules. During the process of representing the vocabulary as vectors, ids are the input ids, and the data types support int32 / int64 / int128 / uint32 / uint64 / uint128. The ids can be ordinal or nominal. Ordinal means ordinal number, for example, the total vocabulary is 100,000, and the id is in the range of 0 - 100,000. Nominal means non - ordinal. For example, there are 100,000 ids, but they are not encoded as 0 - 100,000, and encoding needs to be done according to the input id first. For ordinal numbers, direct indexing can be done like gather. Options are the configurations of Embedding, including name, partitioning, initialization method, update strategy, etc. Partitioning is the method of dividing the vocabulary ids across multiple cards, usually modulo or hash. For example, if there are 8 cards, id = 0 and id = 8 are assigned to card 0, id = 1 and id = 9 are assigned to card 1, and so on.
[0029] In each round of the training process of representing as vectors, first, the deduplication forward execution body 121 (id deduplication) performs deduplication processing on the feature sequence numbers of the received initial vocabulary shards (such as Figure 1 the vocabulary shard 0 in it), to obtain the first non - duplicate data vocabulary shard 121 - 1 (unique id subset) used by the local coprocessor and the first pre - stored data vocabulary shard 121 - 2 for restoration (reverse id subset for restoring deduplication). Here, the ids are all the ids of the features. The first non - duplicate data vocabulary shard 121 - 1 is the input id of the deduplicated batch.
[0030] Subsequently, the forward execution body 122 for vocabulary list sorting (id shuffle) splits the data vocabulary list in the local first non-duplicate data vocabulary list shard 121 that does not belong to the local coprocessor according to the sequence number of the data and sends it to the coprocessor to which it belongs, and receives the data vocabulary list sent by the forward execution body for vocabulary list sorting of other coprocessors, so as to form the first to-be-processed vocabulary list shard belonging to the local coprocessor in this batch of vocabulary lists according to the predetermined splitting rule. Generally speaking, for the input ids of the deduplicated batch, data partitioning and distribution are performed, and the ids belonging to different Ranks are distributed to each Rank, and the ids belonging to this Rank sent by each Rank are received. A subset of the collected ids is the first to-be-processed vocabulary list shard belonging to the local coprocessor. However, there are still duplicate ids in the first to-be-processed vocabulary list shard belonging to the local coprocessor collected from each rank according to the splitting rule. Therefore, in order to reduce the computational amount and save memory space, deduplication processing is still required. The sorting (shuffle) process described in this disclosure is to sort the ids scattered in each rank to the rank to which they should belong based on the splitting rule through NCCL data exchange. The following sorting of the representation vectors is also the same process.
[0031] For this reason, when the second deduplication forward execution body 123 receives the message that the forward execution body 122 for vocabulary list sorting has completed the exchange, it performs deduplication processing on the first to-be-processed vocabulary list shard, obtains the to-be-processed vocabulary list shard 123-1 (cur_rank unique id) of the second non-duplicate data for use by the local coprocessor, and pre-stores the second to-be-processed vocabulary list shard 123-2 (cur_rank reverse id) for restoration, for restoring deduplication. At this time, the number of executions of forward representation vectorization (embedding) in each rank is much less than the actual number, which is a huge saving for the data exchange volume and storage space in the subsequent data processing process. Although the second deduplication process is regarded as part of the id sorting here, it can actually also be a separate operation execution body.
[0032] Figure 2 Shown is a schematic diagram of the process of data exchange between the forward execution body 122 for vocabulary list sorting and the second deduplication forward execution body 123 in the system for performing vectorization on data in the static data processing system according to the present disclosure and forming the to-be-processed vocabulary list shard of the second non-duplicate data. As Figure 2As shown, for the input unique id, each Rank first performs a partition operation on the unique id, distributing the embeddings corresponding to the ids belonging to different ranks into different vocabulary shards (partitions). Eventually, parallel_num copies of vocabulary shards (partitions) are obtained. Partition 0 refers to the vocabulary of the embeddings corresponding to the ids on Rank0. Then, an AllGather operation is performed on the number of ids (num) in each partition to obtain the num_unique_id matrix. As Figure 2 shown, the rows of the matrix num_unique_id matrix are the num_unique_id of each partition obtained by each rank, and the columns of the matrix are the num_unique_id of each partition on each rank. The num_unique_id matrix can be used to obtain the data size during data distribution. For example, each Rank needs to send matrix[rank][0] ids to rank 0, send matrix[rank][0] ids to rank1, and receive matrix[0][rank] ids from rank0 and receive matrix[1][rank] ids from rank1. Next, the vocabulary sorting forward executor 122 performs data distribution. After sending / receiving (send / recv) via the NCCL communication method, each Rank obtains all the cur_rank ids belonging to the local Rank. That is, Rank0 obtains all the ids of partition 0, and Rank1 obtains all the data of partition1. Then, the second deduplication forward executor 123 performs deduplication one last time to obtain the cur_rank uniqueid and the cur_rank reverse_id used for restoring deduplication.
[0033] Next, when the vocabulary data prefetching forward executor 124 obtains the message that the second deduplication forward executor 123 has completed deduplication, it prefetches a predetermined amount of vocabulary data from the second deduplication-free data to-be-processed vocabulary shard 123-1 used by the local coprocessor according to the predetermined prefetch amount. Figure 3The following is a schematic diagram of the data prefetching process of the vocabulary data prefetching forward executor 124 of the system for performing vectorization on data in a static data processing system according to the present disclosure. Generally speaking, the main function of the vocabulary data prefetching forward executor 124 is to control the success rate of data acquisition during the streaming processing, that is, to be responsible for querying in advance the next predetermined number of batches of data to be processed. If the prefetching is successful, a success message is sent to the query executor 125, so that the query executor 125 can avoid repeated queries, improve the cache hit rate, reduce the time and waiting time for the embedding lookup, and make the data processing smoother. As Figure 3 shown, first, call the Test function of the Cache to check whether the embeddings corresponding to the keys are in the cache. The Test function returns the cache missing keys and the cache missing indices. Then, call the Get function of the kvStore (keyword value repository), passing in the cache missing keys, and return the kv Store values, the kv Store missing keys, and the kvStore missing indices, where the kv Store missing keys refer to the keys not found in the kv Store, and the kvStore missing indices refer to the indices in the kv Store values for the parts not found. Subsequently, initialize and fill the parts not found in the kv Store with initial values. Then, push (Put) the values into the Cache, and return the evicted keys and the evicted values. Finally, push (Put) the evicted keys and the evicted values into the kv Store.
[0034] When the vocabulary vectorization query forward executor 125 obtains the message indicating that its prefetching is completed from the vocabulary data prefetch forward executor and the message indicating that the representation vectors of the previous batch of vocabulary have been updated by the data vectorization backward component, it performs vectorization query processing on a predetermined number of vocabulary data based on a predetermined vectorization rule, so as to obtain a prefetch quantity of representation vector shards 125-1 corresponding to the predetermined number of vocabulary data fetched, that is, query the cur_rank unique embeddings corresponding to the cur_rank unique id. Since the vocabulary vectorization query forward executor 125 performs the query operation only after obtaining the message indicating that its prefetching is completed from the vocabulary data prefetch forward executor, the situation of query failure will not occur or the probability of query failure is relatively low. Figure 4 The following is a schematic diagram of the process of data prefetching by the vocabulary vectorization query forward executor 125 of the system for vectorizing data in the static data processing system according to the present disclosure. Specifically, the vocabulary vectorization query forward executor 125 is responsible for finding the corresponding values according to the keys during the forward calculation process. As Figure 4 shown, first, the vocabulary vectorization query forward executor 125 calls the Get function of the Cache to find the representation vectors (embeddings) corresponding to the keys. The Get function returns the values, cache missing keys, and cache missing indices. Cache missing keys refer to the keys not found in the cache, and cache missing indices refer to the indices of the keys not found in the values. Subsequently, the vocabulary vectorization query forward executor 125 calls the Get function of the kv Store to find the embeddings corresponding to the cache missing keys. The Get function returns the kv Store values, kv Store missing keys, and kv Store missing indices. Since the vocabulary data prefetch forward executor 124 has already performed prefetching, when the vocabulary vectorization query forward executor 125 queries, the corresponding keys must be found in the kv Store, so the kv Store missing keys should be empty.
[0035] Since the query operation performed by the vocabulary vectorized query forward executor 125 is not for all the IDs that should be processed by the local coprocessor in this batch, it is necessary to recover the pre-stored duplicate part. Since they are duplicate IDs, it is only necessary to perform duplicate restoration based on the existing query results. Therefore, after the second restoration forward executor 126 obtains the message that it has completed vectorized query processing from the vocabulary vectorized query forward executor 125, it performs corresponding vectorized restoration on the corresponding duplicate vocabulary in the pre-stored second to-be-processed vocabulary shard 123-2 for restoration based on the pre-stored second to-be-processed vocabulary shard 123-2 for restoration and the representation vector shard 125-1 of the prefetch quantity, to obtain a second duplicate representation vector shard corresponding to the pre-stored second to-be-processed vocabulary shard 123-2. Thus, the second duplicate representation vector shard and the representation vector shard 125-1 of the prefetch quantity form a second restored representation vector shard, which is the output result of the second restoration forward executor 126.
[0036] At this time, the second restored representation vector shard is not the corresponding representation vector for the initial vocabulary shard. Some of them are the representation vectors of the data exchanged from other coprocessors. Therefore, it is necessary to return the representation vectors of the data that do not belong to the local coprocessor to the coprocessor to which they originally belong, that is, to perform an inverse sorting of the representation vectors. For this purpose, when the representation vector inverse sorting forward executor 127 obtains the message that it has completed restoration from the second restoration forward executor 126, it returns and sends the representation vectors corresponding to the data vocabulary obtained by the vocabulary sorting forward executor 122 from other coprocessors in the second restored representation vector shard to the other coprocessors and deletes them locally, and receives the representation vectors corresponding to the data vocabulary sent out by the local vocabulary sorting forward executor 122 returned from the representation vector inverse sorting forward executor 127 of other coprocessors, to form a first non-duplicate representation vector shard 127-1 (unique embeddings) corresponding to the first non-duplicate data vocabulary shard 121-1 (unique id) used by the local coprocessor.
[0037] Figure 5 Shown is a schematic diagram of the restoration process of the second restoration forward executor 126 and the representation vector inverse sorting process of the representation vector inverse sorting forward executor 127 for vectorizing data in a static data processing system according to the present disclosure. As Figure 5As shown, after the second pre-restoration forward execution body 126 obtains the prefetch quantity of representation vector shards 125-1 (cur_rank uniqueembedding) from the vocabulary embedding lookup forward execution body 125, for the input prefetch quantity of representation vector shards 125-1 (cur_rank uniqueembedding), it first uses the pre-stored second shard of the vocabulary to be processed for restoration 123-2 (cur_rank reverseid) to restore the second restored representation vector shards (cur_rank embeddings) corresponding to the first shard of the vocabulary to be processed (i.e., the id before deduplication). Some of the corresponding feature ids in the second restored representation vector shards do not belong to the features of the data to be processed by the local coprocessor. Therefore, the embedding shuffle forward execution body 127 needs to perform an inverse sorting of the representation vectors on the basis of the second restored representation vector shards (cur_rank embeddings) corresponding to the set of ids before deduplication (i.e., 121-1) output by the second pre-restoration forward execution body 126, and finally obtain the first non-repeated representation vector shards 127-1 (unique embeddings) corresponding to the first non-repeated data vocabulary shard 121-1 (unique id) used by the local coprocessor.
[0038] Finally, after the first pre-restoration forward execution body 128 obtains the message indicating that it has completed the representation vector exchange from the embedding shuffle forward execution body 127, based on the pre-stored first data vocabulary shard 121-2 for restoration and the first non-repeated representation vector shards 127-1, it performs corresponding vectorized restoration on the corresponding repeated vocabulary in the pre-stored first data vocabulary shard 121-2 for restoration to obtain the first repeated representation vector shards corresponding to the pre-stored first data vocabulary shard for restoration. Thus, the first repeated representation vector shards and the first non-repeated representation vector shards form the first restored representation vector shards 128-1. Generally speaking, the first pre-restoration forward execution body 128 restores the embeddings arranged by partition to the order before the id partition, and finally obtains the first restored representation vector shards 128-1 (cur_batch_unique embeddings) of the current batch for subsequent calculations.
[0039] Return to Figure 1 , such as Figure 1As shown, after the first restored representation vector shard 128-1 (cur_batch_uniqueembeddings) passes through the loss function executor 140, it enters the backward data processing as the input of the backward data processing. The data vectorization backward component 130 includes: a vector difference executor 131, a first merge backward executor 132, a representation vector difference sorting backward executor 133, a representation vector value update backward executor 134, and a representation vector update backward executor 135. The vector difference executor 131 performs a difference operation on the first restored representation vector shard processed by the loss function executor to obtain the representation vector difference result 131-1 (embeddingdiff) corresponding to the current batch id. Subsequently, the first merge (Reduce) backward executor 132 performs a deduplication operation on the representation vector difference result 131-1 output by the vector difference executor 131 based on the sequence number of the pre-stored first data word list shard for restoration, to obtain the first non-repeated representation vector difference result. That is, for the embedding diff of this batch, the first merge backward executor 132 first performs a merge (reduce) operation to obtain the first non-repeated representation vector difference result 132-1 (unique embeddingdiff) corresponding to the unique id of this batch.
[0040] Subsequently, the representation vector difference sorting backward executor 133 (emb grad shuffle) based on the sequence number of the pre-stored second to-be-processed word list shard 121-2 for restoration, splits the representation vector difference result that does not belong to the local coprocessor from the first non-repeated representation vector difference result 132-1 and sends it to the corresponding other coprocessors, and receives the representation vector difference belonging to the local coprocessor sent from the representation vector difference sorting backward executor of other coprocessors, so as to form the second non-repeated representation vector difference result 133-1 (cur_rank uniqueembedding diff) belonging to the local coprocessor. Specifically, after obtaining the first non-repeated representation vector difference result 132-1 (uniqueembedding diff) of this batch, the representation vector difference sorting backward executor 133 sends the embedding diff corresponding to the ids of different Ranks to each Rank, and receives the embeddingdiff corresponding to the id of this Rank sent from each Rank. After collection, a reduce operation is performed on the cur_rank embedding diff to accumulate the embeddingdiff corresponding to the repeated ids, to obtain cur_rank unique embedding diff.
[0041] Figure 6 Shown is a schematic diagram of the first merged backward executor 132 of the system for performing vectorization on data in a static data processing system according to the present disclosure and the combing process characterizing the vector difference shuffle backward executor 133. As shown, the vector difference shuffle backward executor 133 (embedding gradient shuffle) distributes the embedding diff of the current batch obtained in the network to the ranks where the vocabulary tables of the embeddings corresponding to the ids are located, and finally obtains the embedding_diff corresponding to the vocabulary table on the current rank. The process of the vector difference shuffle backward executor 133 (embedding gradient shuffle) is very similar to the processes of the vocabulary combing forward executor 122 and the second deduplication forward executor 123. The vocabulary combing forward executor 122 distributes the ids of the current batch, while the vector difference shuffle backward executor 133 distributes the embedding diff of the current batch. Details are not elaborated here.
[0042] Return to Figure 1 , the vector value update backward executor 134 updates the value of the second non-duplicate vector difference result 133-1 (cur_rank unique embedding diff) based on the value 125-2 of the vector obtained by the vocabulary vectorization query forward executor 125 and a predetermined learning frequency (not shown, usually determined by the user according to needs). Then, the vector update backward executor 135 updates the initial data vocabulary shard using the updated value of the second non-duplicate vector difference result and the context of a predetermined number of vocabulary data prefetched.
[0043] Furthermore, in order to implement static streaming data processing, the vocabulary partitioning component 110, the first deduplication forward executor 121, the vocabulary combing forward executor 122, the vocabulary data prefetch forward executor 124, the vocabulary vectorization query forward executor 125, and the data vectorization backward component 130 can be deployed in four parallel task flows. Figure 7 Shown is a timing schematic diagram of the parallel task flows of the system for performing vectorization on data in a static data processing system according to the present disclosure. As Figure 7As shown, after the vocabulary partitioning component 110 partitions the vocabulary, it iteratively loads the data, loading a part of the data each time. The copy executors arranged in the task flow STREAM 1 copy the data from the host HOST to the coprocessor device. Subsequently, the first deduplication forward executor and the vocabulary sorting forward executor 122 perform sorting among each rank in the task flow STREAM2. The vocabulary data prefetch forward executor 124 and the vocabulary vectorized query forward executor 125 successively perform prefetching and query operations when deployed in the task flow STREAM 3. Finally, the backward processing component 130 deployed in the task flow STREAM 4 performs backward processing on the generated vectors. The executors in each task flow control the timing through messages, so as to perform streaming processing in the same task flow. During the execution of each batch, the query operation of the vocabulary vectorized query forward executor 125 is very time-consuming. If the id cannot be hit by the GPU Cache, it needs to be loaded from the memory or disk, which will result in very low throughput. The query operation of the vocabulary vectorized query forward executor 125 needs to depend on the end of the embedding update execution of the previous batch to execute, and cannot be executed in advance, and there is no chance to be masked by the forward and backward calculations. Therefore, the present application sets a vocabulary data prefetch forward executor 124 in the network to perform prefetching operations. For the input id, the corresponding embedding is put into the cache as much as possible to improve the cache hit rate of the subsequent query (Lookup) operation and reduce the latency. Through the timing control of the message control mode between the forward and backward executors, the embedding prefetching operation of the vocabulary data prefetch forward executor 124 is executed one iteration order (iter) in advance and overlapped with the forward and backward calculations, so that the time-consuming query operation is masked by the forward and backward calculations.
[0044] In addition, according to the system for vectorizing data in a static data processing system of the present disclosure, wherein the vocabulary sorting forward executor 122, the representation vector inverse sorting forward executor 127, and the representation vector differential sorting backward executor 133 perform data exchange between different coprocessors through the NCCL communication method.
[0045] Optionally, according to another aspect of the present disclosure, refer back to Figure 1 , in order to reduce the number of traversals of the above parallel scheme and the number of strategies for making parallel decisions and reduce data transmission, the initial logical node topology graph can be preprocessed. For example, through traversal, when a specific initial logical node has two or more downstream initial logical nodes, between the specific initial logical node and all of its
[0046] Through the parallel policy decision-making system for distributed data processing according to the present disclosure, it is possible to minimize the solution space faced by the parallel decision-making of distributed data processing from a global perspective, improve the feasibility of automatic parallel decision-making, reduce the difficulty of automatic parallel decision-making, and enable the parallel results obtained from the parallel decision-making to have lower computational costs and transmission costs. As a result, the computational efficiency of fixed computing resources for the same computing task is maximally improved, thereby accelerating the data processing speed. More importantly, it realizes the automation of parallel decision-making on the basis of approaching the lowest transmission cost as much as possible, greatly reducing the cost of manual debugging.
[0047] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that for those of ordinary skill in the art, all or any steps or components of the method and apparatus of the present disclosure can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present disclosure.
[0048] Therefore, the object of the present disclosure can also be achieved by running a program or a set of programs on any computing device. The computing device can be a well-known general-purpose device. Therefore, the object of the present disclosure can also be achieved only by providing a program product containing program code for implementing the method or apparatus. That is to say, such a program product also constitutes the present disclosure, and a storage medium storing such a program product also constitutes the present disclosure. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future.
[0049] It should also be noted that in the apparatus and method of the present disclosure, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure. And the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.
[0050] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A system for performing vectorization on data in a static data processing system, comprising: A vocabulary partitioning component for evenly dividing the numbers of a batch of vocabulary to be input into a plurality of initial data vocabulary fragments equal to the number of coprocessors running in parallel and distributing the input to the corresponding coprocessors, and a data vectorization forward component and a data vectorization backward component deployed corresponding to each coprocessor, wherein the data vectorization forward component includes a plurality of data vectorization forward execution bodies deployed in a streaming manner, and the data vectorization backward component includes a plurality of data vectorization backward execution bodies deployed in a streaming manner, wherein The plurality of data vectorization forward execution bodies execute stream parallel processing at different stages for input continuous batches of vocabulary tables, and include: A first deduplication forward execution body is used to perform feature sequence number deduplication processing on the received initial vocabulary fragments to obtain first non-repetitive data vocabulary fragments used by the local coprocessor and pre-store first data vocabulary fragments for restoration; The vocabulary combing forward execution body divides the local first non-repetitive data vocabulary fragments that do not belong to the local coprocessor based on the predetermined segmentation rule according to the sequence number of the data and sends them to the coprocessor to which they belong, and receives the data vocabulary sent from the vocabulary combing forward execution body of other coprocessors, thereby forming the first to-be-processed vocabulary fragments that belong to the local coprocessor based on the predetermined segmentation rule in the vocabulary of this batch; The second deduplication forward execution body is used to perform deduplication processing on the first to-be-processed vocabulary fragment upon receiving the message that the vocabulary combing forward execution body completes the exchange, obtain the second to-be-processed vocabulary fragment without duplicate data used by the local coprocessor, and pre-store the second to-be-processed vocabulary fragment for restoration; The vocabulary data pre-fetching forward execution body is used to pre-fetch a predetermined amount of vocabulary data according to the second non-repetitive data to-be-processed vocabulary fragments used by the local coprocessor according to a predetermined pre-fetch quantity when receiving a message that the de-duplication is completed from the second de-duplication forward execution body; The vocabulary vectorization query forward execution body, when obtaining a message that the pre-fetching is completed from the vocabulary data pre-fetching forward execution body and a message that the representation vectors of the previous batch of vocabulary are updated by the data vectorization backward component, performs vectorization query processing on a predetermined number of vocabulary data based on a predetermined vectorization rule, thereby obtaining a pre-fetched number of representation vector fragments corresponding to the pre-fetched predetermined number of vocabulary data; The second restoration forward execution body, after obtaining a message from the vocabulary vectorization query forward execution body that it has completed the vectorization query processing, performs corresponding vectorization restoration on the corresponding repeated vocabulary in the pre-stored second vocabulary fragments to be processed for restoration based on the pre-stored second vocabulary fragments to be processed for restoration and the pre-fetched number of representation vector fragments, and obtains a second repeated representation vector fragment corresponding to the pre-stored second vocabulary fragments to be processed for restoration, thereby forming a second restoration representation vector fragment with the second repeated representation vector fragment and the pre-fetched number of representation vector fragments; The representation vector back-combing forward executor, when receiving a message that the restoration is completed from the second restoration forward executor, returns the representation vector corresponding to the data vocabulary obtained by the vocabulary combing forward executor from other coprocessors in the second restoration representation vector segment to the other coprocessors and deletes it locally, and receives the representation vector corresponding to the data vocabulary sent by the local vocabulary combing forward executor returned by the representation vector back-combing forward executor of the other coprocessors, to form a first non-repetitive representation vector segment corresponding to the first non-repetitive data vocabulary segment used by the local coprocessor; and The first restoration forward executor, after obtaining the message that it has completed the exchange of representation vectors from the representation vector back-combing forward executor, performs corresponding vectorized restoration on the corresponding repeated vocabulary in the pre-stored first data vocabulary fragment for restoration based on the pre-stored first data vocabulary fragment for restoration and the first non-repetitive representation vector fragment, and obtains the first repeated representation vector fragment corresponding to the pre-stored first data vocabulary fragment for restoration, thereby forming a first restoration representation vector fragment with the first repeated representation vector fragment and the first non-repetitive representation vector fragment.
2. The system for performing vectorization on data in a static data processing system according to claim 1, wherein the data vectorization backward component comprises: A vector difference executor performs difference processing on the first restored representation vector slice after being processed by the loss function executor; A first merge backward execution body performs deduplication processing on the representation vector difference result output by the vector difference execution body based on the sequence number of the first data vocabulary fragment for restoration pre-stored, so as to obtain a first non-repetitive representation vector difference result; After the representation vector differential is combed, the execution body is directed to split the representation vector differential results that do not belong to the local coprocessor from the first non-repeating representation vector differential results based on the sequence number of the second to-be-processed word list fragment for restoration, and sends them to the corresponding other coprocessors, and receives the representation vector differential results belonging to the local coprocessor sent to the execution body after the representation vector differentials of other coprocessors are combed, thereby forming a second non-repeating representation vector differential result belonging to the local coprocessor; The characterization vector value updates the backward execution body, and based on the value of the characterization vector obtained by the vocabulary vectorization query forward execution body and the predetermined learning frequency, updates the value of the second non-repetitive characterization vector difference result; as well as The characterization vector update backward execution body is used to update the initial data vocabulary fragment using the updated value of the second non-repetitive characterization vector difference result and the context of the pre-fetched predetermined number of vocabulary data.
3. The system for performing vectorization on data in a static data processing system according to claim 1 or 2, wherein The vocabulary partitioning component, the first deduplication forward executor and the vocabulary combing forward executor, the vocabulary data pre-fetching forward executor and the vocabulary vectorization query forward executor, and the data vectorization backward component are deployed in four parallel task flows.
4. The system for performing vectorization on data in a static data processing system according to claim 2, wherein the vocabulary combing forward executor, the representation vector backward combing forward executor, and the representation vector differential combing backward executor exchange data between different coprocessors through NCCL communication.
5. A method for performing vectorization on data in a static data processing system, comprising: The word list partitioning component evenly divides the word list numbers of a batch to be input into multiple initial data word list fragments equal to the number of coprocessors running in parallel, and distributes the input to the corresponding coprocessors, wherein a data vectorization forward component and a data vectorization backward component are deployed corresponding to each coprocessor, the data vectorization forward component includes multiple data vectorization forward execution bodies deployed in a streaming manner, and the data vectorization backward component includes multiple data vectorization backward execution bodies deployed in a streaming manner, wherein The plurality of data vectorization forward execution bodies execute stream parallel processing at different stages for input continuous batches of vocabulary tables, and include: Performing feature sequence number deduplication processing on the received initial vocabulary fragments through the first deduplication forward execution body to obtain first non-repeated data vocabulary fragments used by the local coprocessor and pre-stored first data vocabulary fragments for restoration; The word list combing forward execution body splits the local first non-repetitive data word list fragment according to the sequence number of the data, and sends the data word list that does not belong to the local coprocessor based on the predetermined segmentation rule to the coprocessor to which it belongs, and receives the data word list sent from the word list combing forward execution body of other coprocessors, thereby forming the first to-be-processed word list fragment belonging to the local coprocessor based on the predetermined segmentation rule in the current batch of words; When receiving a message from the vocabulary combing forward execution body that the exchange is completed, the second deduplication forward execution body performs deduplication processing on the first vocabulary fragment to be processed, obtains a second vocabulary fragment to be processed without duplicate data used by the local coprocessor, and pre-stores the second vocabulary fragment to be processed for restoration; When the vocabulary data pre-fetch forward execution body obtains a message that the deduplication is completed from the second deduplication forward execution body, a predetermined amount of vocabulary data is pre-fetched according to the second non-repetitive data to-be-processed vocabulary slice used by the local coprocessor according to a predetermined pre-fetch quantity; When the forward executor of the vocabulary vectorization query obtains a message of completion of pre-fetching from the vocabulary data pre-fetching forward executor and a message of completion of updating of the representation vectors of the previous batch of vocabulary by the data vectorization backward component, the vectorization query processing is performed on the predetermined number of vocabulary data based on the predetermined vectorization rules, thereby obtaining the pre-fetched number of representation vector fragments corresponding to the pre-fetched predetermined number of vocabulary data; After the second restoration forward execution body obtains a message that the vectorized query processing is completed from the vocabulary vectorized query forward execution body, the corresponding repeated vocabulary in the pre-stored second vocabulary to be processed fragments for restoration is performed corresponding vectorized restoration based on the pre-stored second vocabulary to be processed fragments for restoration and the pre-fetched number of representation vector fragments, to obtain a second repeated representation vector fragment corresponding to the pre-stored second vocabulary to be processed fragment for restoration, thereby forming a second restored representation vector fragment with the second repeated representation vector fragment and the pre-fetched number of representation vector fragments; The representation vector back-combing forward executor, when receiving a message that the restoration is completed from the second restoration forward executor, returns the representation vector corresponding to the data vocabulary obtained by the vocabulary combing forward executor from other coprocessors in the second restoration representation vector segment to the other coprocessors and deletes it locally, and receives the representation vector corresponding to the data vocabulary sent by the local vocabulary combing forward executor returned by the representation vector back-combing forward executor of the other coprocessors, to form a first non-repetitive representation vector segment corresponding to the first non-repetitive data vocabulary segment used by the local coprocessor; and After the first restoration forward executor obtains the message that it has completed the exchange of representation vectors from the representation vector back-combing forward executor, based on the pre-stored first data vocabulary fragment for restoration and the first non-repetitive representation vector fragment, it performs corresponding vectorized restoration on the corresponding repeated vocabulary in the pre-stored first data vocabulary fragment for restoration, and obtains the first repeated representation vector fragment corresponding to the pre-stored first data vocabulary fragment for restoration, thereby forming a first restoration representation vector fragment with the first repeated representation vector fragment and the first non-repetitive representation vector fragment.
6. The method for performing vectorization on data in a static data processing system according to claim 5, wherein the data vectorization backward component: Performing differential processing on the first restored representation vector slice processed by the loss function executor through the vector differential executor; Performing deduplication processing on the representation vector difference result output by the vector difference execution body based on the sequence number of the first data vocabulary fragment for restoration pre-stored by the first merge backward execution body to obtain a first non-repetitive representation vector difference result; By combing the representation vector differentials, the execution body divides the representation vector differential results that do not belong to the local coprocessor from the first non-repeating representation vector differential results based on the sequence number of the second word list fragment to be processed for restoration, and sends them to the corresponding other coprocessors, and receives the representation vector differential results belonging to the local coprocessor sent to the execution body after combing the representation vector differentials of the other coprocessors, so as to form a second non-repeating representation vector differential result belonging to the local coprocessor; The backward execution body updates the value of the representation vector obtained by querying the forward execution body based on the vocabulary vectorization and the predetermined learning frequency through the representation vector value, and updates the value of the second non-repetitive representation vector difference result; as well as The backward execution body updates the initial data vocabulary fragments by using the updated value of the second non-repetitive representation vector difference result and the context of the pre-fetched predetermined number of vocabulary data through the representation vector update.
7. The method for performing vectorization on data in a static data processing system according to claim 5 or 6, wherein The vocabulary partitioning component, the first deduplication forward executor and the vocabulary combing forward executor, the vocabulary data pre-fetching forward executor and the vocabulary vectorization query forward executor, and the data vectorization backward component are deployed in four parallel task flows.
8. According to the method for vectorizing data in a static data processing system as described in claim 6, the vocabulary combing forward execution body, the representation vector back-combing forward execution body, and the representation vector differential combing backward execution body exchange data between different coprocessors through NCCL communication method.
Citation Information
Patent Citations
Query optimization method based on join index in data warehouse
CN104866608A
Text abstract generation method and device, equipment and storage medium
CN113268586A