Address mapping method and related apparatus
Patent Information
- Application Number
- CN202510287281.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2026-09-11
AI Technical Summary
[0056] The technical effects achieved by the second, third, fourth, fifth, and sixth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here.
Smart Images

Figure CN122733745A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an address mapping method and related apparatus. Background Technology
[0002] Recommendation systems are common business systems in internet applications, implemented using trained recommendation models. The training process of a recommendation model involves read and write operations on an address mapping table. The address mapping table records the mapping relationship between training data and the addresses of their corresponding embedding vectors. Training data might be, for example, "XX mobile phone" or data obtained by preprocessing "XX mobile phone" (e.g., hash calculation). During training, the embedding vectors corresponding to the training data are used for computation, and the addresses in the address mapping table represent the positions of these embedding vectors in the embedding table. After obtaining a piece of training data, a read operation is performed to query the address mapping table for the address corresponding to that training data. If the address is found (i.e., a match), the embedding vector for that training data is retrieved from the embedding table based on that address. If the address is not found (i.e., a match), the embedding vector corresponding to that training data is generated, stored in the embedding table, and then a write operation is performed to write the address of the training data and its corresponding embedding vector into the address mapping table, thus constructing an entry for the training data in the address mapping table.
[0003] Address mapping tables and embedding tables are typically built upon massive training data; current typical embedding tables have reached trillions of bytes (TB) or even tens of trillions of bytes. The sheer size of these embedding tables causes the address mapping process to consume over 40% of the end-to-end training time for recommendation models. Therefore, improving address mapping efficiency is a hot research topic in the industry. Summary of the Invention
[0004] This application provides an address mapping method and related apparatus, which can improve the address mapping efficiency during the training process of recommendation models, thereby improving the training efficiency of recommendation models. The technical solution is as follows:
[0005] Firstly, an address mapping method is provided, the method comprising:
[0006] In parallel, query operations are performed on N sets of training data for the recommendation model. The query operation is used to retrieve the address corresponding to the training data from an address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data. The N sets of training data are obtained by dividing multiple sets of training data for the recommendation model. N is an integer greater than 1. The multiple sets of training data are, for example, any batch of training data from multiple batches of training data for the recommendation model.
[0007] In related technologies, address mapping logic is executed serially on multiple training data sets. For each training data set, the address mapping logic includes read and write operations on the address mapping table, with the write operation executed immediately if the read operation fails. That is, the read and write operations in the address mapping logic of related technologies are tightly coupled. During the execution of address mapping logic on a single training data set, a mutex lock is required to lock the address mapping table, thus preventing the parallel execution of address mapping logic on the address mapping table.
[0008] However, in this application, multiple training data are divided into N training data, and the read and write operations on the address mapping table are decoupled. That is, during the query process, all operations on the address mapping table are read operations, without any write operations. This allows the query operation (i.e., the read operation on the address mapping table) to be performed in parallel on the N training data, improving the address mapping efficiency and thus increasing the training efficiency of the recommendation model.
[0009] The method further includes, before performing query operations on the N training data sets of the recommendation model in parallel, the following steps: acquiring multiple training data sets of the recommendation model; dividing the multiple training data sets into N training data sets; and, after all query operations on the N training data sets have been completed, writing first data and the address of the embedding vector corresponding to the first data set into the address mapping table, wherein the first data set consists of training data sets for which no corresponding address was found among the N training data sets. That is, after all read operations on the address mapping table are completed, if there is any training data that was not found, then a write operation is performed on the address mapping table.
[0010] Since reading operations on the address mapping table usually far outnumber writing operations during the training of recommendation models, this application decouples reading and writing operations on the address mapping table, enabling parallelization of reading operations, which greatly reduces the time consumption of reading operations and thus improves the training efficiency of recommendation models.
[0011] For example, the training process of the recommendation model includes multiple iterations, and the parallel query operation on the N training data is performed in all of these iterations. Thus, in the first iteration, there might be instances of a miss on the N training data. However, after the step of writing the first data and the address of the corresponding embedding vector to the address mapping table is completed in the first iteration, misses will not occur again in subsequent iterations unless there are special circumstances. Subsequent iterations only involve read operations on the address mapping table, but no write operations are performed. Therefore, it is evident that read operations far outnumber write operations during the training of the recommendation model.
[0012] In one possible implementation, the parallel query operation on the N training data sets includes: running N read threads in parallel across N processor cores, and performing query operations on the N training data sets in parallel across the N read threads. Each processor core runs one of the N read threads, and each read thread performs a query operation on one set of training data. That is, the technical solution of this application can fully utilize the multi-core capabilities of the processor and improve address mapping efficiency.
[0013] In one possible implementation, the first data is stored in a temporary variable space, which is used to store training data for which no corresponding address is found; before writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table, the method further includes: reading the first data from the temporary variable space.
[0014] Among these, training data for which no corresponding address was found may contain duplicates. Therefore, the temporary variable space may record identical training data. Alternatively, a deduplication operation can be performed to ensure that the training data recorded in the temporary variable space is not duplicated. The following sections will describe these cases separately.
[0015] In the first scenario, the temporary variable space may contain the same training data.
[0016] For example, the number of first data is at least one, and the method further includes: during the parallel execution of query operations on the N training data, whenever a first data is determined, the first data is written into the temporary variable space.
[0017] Accordingly, in one implementation, writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table includes: writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table based on the first data being read from the temporary variable space for the first time; and writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table based on the first data not being read from the temporary variable space for the first time, and the address corresponding to the first data still not being found in the address mapping table.
[0018] In the second scenario, the training data recorded in the temporary variable space has been deduplicated.
[0019] For example, the number of first data is at least one, and the method further includes: during the parallel execution of query operations on the N training data, whenever a first data is determined and no training data identical to the first data is stored in the temporary variable space, the first data is written into the temporary variable space.
[0020] Accordingly, in one implementation, the first data is read sequentially from the temporary variable space, and for each piece of the first data is read, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table.
[0021] In addition to accelerating address mapping by parallelizing read operations, this application can also divide the address mapping table into cold tables and hot tables according to the hotness of the data, thereby further improving the address mapping efficiency based on the cold tables and hot tables.
[0022] In one implementation, the address mapping table includes a cold table and a hot table, wherein the frequency of training data in the cold table is lower than the frequency of training data in the hot table; the parallel query operation on the N sets of training data includes: for a first training data in the N sets of training data, querying the address corresponding to the first training data from the hot table, wherein the first training data is any one of the N sets of training data; if the address corresponding to the first training data is not found in the hot table, then querying the address corresponding to the first training data from the cold table.
[0023] Extensive data analysis shows that hot data in the recommendation field is highly concentrated. Therefore, the number of entries in the hot table is much smaller than the number of entries in the cold table. However, since there is a large amount of hot data, the probability of each training data point being a hot data point is much greater than the probability of it being a cold data point. Therefore, querying the address corresponding to the training data from the hot table first has a higher probability of success (i.e., a hit), and a lower probability of failure. Thus, using hot and cold tables can further improve the speed of address mapping.
[0024] In one implementation, after acquiring multiple training data sets for the recommendation model, frequency statistics are determined, which record the number of times each acquired training data set appears. Based on these frequency statistics, cold and hot tables are identified.
[0025] The multiple training data can be any batch of training data from multiple batches of training data for the recommendation model. As an example, each time a batch of training data is acquired, a frequency statistics result is determined. That is, when acquiring non-first batch training data, the frequency statistics result needs to be updated. The more training data acquired, the more reliable the frequency statistics result.
[0026] In one possible implementation, the training data for the recommendation model includes multiple batches of training data, which are sequentially used to train the recommendation model. The hot list and the cold list are determined when the recommendation model is trained using the Mth batch of training data, where M is a positive integer. For example, after obtaining the Mth batch of training data, the cold list and the hot list are determined using the latest frequency statistics.
[0027] Secondly, an address mapping apparatus is provided, which has the function of implementing the address mapping method described in the first aspect. The address mapping apparatus includes one or more modules for implementing the address mapping method provided in the first aspect.
[0028] That is, an address mapping device is provided, the device comprising:
[0029] The acquisition module is used to acquire multiple training data sets for the recommendation model.
[0030] The data partitioning module is used to divide the multiple training data into N parts of training data, where N is an integer greater than 1;
[0031] A parallel query module is used to perform query operations on the N training data in parallel. The query operation is used to query the address corresponding to the training data from the address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data.
[0032] The write module is used to write first data and the address of the embedding vector corresponding to the first data into the address mapping table after all query operations on the N training data have been completed. The first data is training data among the N training data for which no corresponding address was found.
[0033] In one possible implementation, the parallel query module is specifically used for:
[0034] N read threads are run in parallel by N processor cores, and query operations are performed on the N training data in parallel by the N read threads. Each processor core runs one of the N read threads, and each read thread performs a query operation on one piece of training data.
[0035] In one possible implementation, the first data is stored in a temporary variable space, which is used to store training data for which no corresponding address was found.
[0036] The device further includes:
[0037] A reading module is used to read the first data from the temporary variable space.
[0038] In one possible implementation, the quantity of the first data is at least one, and the parallel query module is further configured to:
[0039] During the parallel query operation on the N sets of training data, whenever a first data point is determined, the first data point is written into the temporary variable space.
[0040] In one possible implementation, the write module is specifically used for:
[0041] Based on the first reading of the first data from the temporary variable space, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table; or,
[0042] Based on the fact that the first data is read from the temporary variable space for the first time, and the address corresponding to the first data cannot be found in the address mapping table, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table.
[0043] In another possible implementation, the quantity of the first data is at least one, and the parallel query module is further configured to:
[0044] During the parallel query operation on the N training data, whenever a first data is identified and no training data identical to the first data is stored in the temporary variable space, the first data is written into the temporary variable space.
[0045] In one possible implementation, the training process of the recommendation model includes multiple iterations, and the parallel query operation on the N training data is performed in each of the multiple iterations.
[0046] In one possible implementation, the address mapping table includes a cold table and a hot table, wherein the frequency of training data in the cold table is lower than the frequency of training data in the hot table;
[0047] The parallel query module is specifically used for:
[0048] For the first training data in the N training data sets, the address corresponding to the first training data is queried from the hot table. The first training data is any one of the N training data sets.
[0049] If the address corresponding to the first training data is not found in the hot table, then the address corresponding to the first training data is found in the cold table.
[0050] Thirdly, a computer device is provided, the computer device including a processor and a memory; the processor is configured to execute instructions stored in the memory to cause the computer device to perform the address mapping method provided in the first aspect.
[0051] In one possible implementation, the computer device may further include a communication bus for establishing a connection between the processor and the memory.
[0052] Fourthly, a computer device is provided, comprising a processor and a memory, the memory being used to store a program for executing the address mapping method provided in the first aspect, and to store data related to implementing the address mapping method provided in the first aspect. The processor is configured to execute the program stored in the memory.
[0053] In one possible implementation, the computer device may further include a communication bus for establishing a connection between the processor and the memory.
[0054] Fifthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the address mapping method described in the first aspect.
[0055] In a sixth aspect, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the address mapping method described in the first aspect above.
[0056] The technical effects achieved by the second, third, fourth, fifth, and sixth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description
[0057] Figure 1This is a schematic diagram of a data processing flow in a recommendation model provided in an embodiment of this application;
[0058] Figure 2 This is a schematic diagram illustrating an embedding vector calculation provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram of an address mapping method in related technologies;
[0060] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0061] Figure 5 This is a flowchart of an address mapping method provided in an embodiment of this application;
[0062] Figure 6 This is a probability distribution diagram corresponding to an embedded table provided in an embodiment of this application;
[0063] Figure 7 This is a schematic diagram of a method for determining the temperature of a hot or cold water meter, provided in an embodiment of this application.
[0064] Figure 8 This is a flowchart of another address mapping method provided in the embodiments of this application;
[0065] Figure 9 This is a schematic diagram of a multiple iteration process provided in an embodiment of this application;
[0066] Figure 10 This is a flowchart of yet another address mapping method provided in the embodiments of this application;
[0067] Figure 11 This is a schematic diagram of the structure of an address mapping device provided in an embodiment of this application. Detailed Implementation
[0068] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0069] First, the relevant technical background will be introduced.
[0070] Recommender systems are among the most common business systems in internet applications. As mentioned in the background, recommender systems involve querying and calculating embedding vectors, and the larger the embedding table, the higher the algorithm accuracy usually is. Based on this, leading internet companies improve algorithm accuracy and reap the benefits by expanding the size of the embedding table. Typical embedding tables (also known as vocabulary) have reached trillions (TB) to tens of trillions, and can process up to 1000 TB of user behavior per day. To improve the performance of recommender systems, related technologies widely use large-scale XPU clusters to accelerate recommender models.
[0071] The XPU can include, but is not limited to, various chips used for computing and acceleration, such as central processing units (CPUs), neural processing units (NPUs), graphics processing units (GPUs), and field-programmable gate arrays (FPGAs).
[0072] Currently, large-scale computing clusters (used to run recommendation models) in the industry already have a high number of GPUs, for example, they have expanded to 128-256 GPUs. This method of increasing the number of GPUs is known in the industry as horizontal scaling or scale-out. The opposite is vertical scaling, also known as scale-up, which expands the capacity of the computing cluster by adding more XPUs, memory, or hard drives to the nodes.
[0073] Since embedding vector computation involves massive (billions) feature queries and updates, it plays a crucial role in the overall performance of recommendation systems, especially sparse feature embedding vector computation. The industry typically employs high-density computing accelerators and high-speed device-to-device (D2D) heterogeneous computing systems to accelerate embedding vector computation. Among these, the address mapping module is the core module for embedding vector computation on the accelerator.
[0074] Figure 1 This is a schematic diagram of the data processing flow of a recommendation model provided in an embodiment of this application. For example... Figure 1As shown, training data can include numerical features and categorical features. Numerical features include data such as a user's age, height, and weight, while categorical features include data such as the user's browsing history (e.g., products, videos). The data processing flow of the recommendation model includes processing of numerical features and processing of categorical features. Numerical features can be fed into a multilayer perceptron (MLP) for processing, and then into a feature cross-layer. For categorical features, address mapping is first performed to retrieve the corresponding address. After retrieving the address, an embedding table is queried based on the address to obtain the embedding vector corresponding to that categorical feature. This embedding vector is then also fed into the feature cross-layer. The feature cross-layer can include various types of pooling modules, transformer modules, etc. After processing the data in the feature cross-layer, the resulting data can be further fed into deeper deep neural network (DNN) layers to complete subsequent training or inference processes.
[0075] Figure 2 This is a schematic diagram illustrating an embedding vector calculation provided in an embodiment of this application. The purpose of embedding vector calculation is to obtain, for example... Figure 1 The embedding vectors corresponding to the categorical features shown are calculated through address mapping and embedding table lookup. Address mapping maps the categorical features to physically accessible index addresses (referred to as addresses), while embedding table lookup retrieves the embedding vectors from the embedding table based on the mapped addresses. For example... Figure 2 As shown, the input data for address mapping is categorical features, with a data dimension of [S]. The data type can be integer values, where S represents the number of input categorical features, such as 1K, 8K, 32K, etc. The output data for address mapping is addresses, also with a data dimension of [S], representing the number of output addresses. The data dimension of the embedding table is [N, H]; N represents the total number of entries in the embedding table, such as 100 million to 10 billion; H represents the dimension of the embedding vector, also known as the embedding hidden layer dimension, such as 128 to 1024. The input data for embedding table queries is addresses, for example, S addresses can be input. The output data for embedding table queries is embedding vectors, for example, S embedding vectors corresponding one-to-one with the S addresses can be output.
[0076] It should be understood that in some embodiments, such as Figure 1As shown, a training dataset may include numerical features and categorical features. The numerical features and categorical features can be processed separately and then fused (e.g., both are fed into a feature cross-layer). In other embodiments, a training dataset may include categorical features but not numerical features. In this case, the training dataset can be directly mapped to an address and the embedding table can be queried without involving the processing of numerical features.
[0077] Since the key point of this application embodiment is address mapping acceleration and does not care whether the training process involves numerical features, for ease of description, this application embodiment fuzzifies the data that needs to be embedded vector calculation into training data and performs address mapping on the training data, that is, performs address mapping on the category features in the training data. The category features can be part of the training data or all of the data.
[0078] It should also be understood that, Figure 1 The recommendation model shown is intended to exemplify that embedding vector calculation is the foundation for the recommendation model's data processing via neural networks, and is not intended to limit the structure of the recommendation model in this application's embodiments. Those skilled in the art will understand that for any recommendation model structure involving embedding vector calculation, the technical solutions of this application's embodiments can be adopted.
[0079] It should also be noted that although the technical solution of this application's embodiments originates from solving related technical problems in recommendation systems, this technical solution can be adaptably applied to other domain systems, thereby accelerating address mapping. This technical solution is applicable to any recommendation training architecture, such as parameter server (PS) architecture, AllReduce architecture, etc.
[0080] Currently, there are two address mapping technologies in the industry: static mapping and dynamic mapping.
[0081] The core principle of static address mapping is to use a fixed-size (shape, typically [vocabulary_size, embedding_dimension]) embedding table to store embedding vectors. `vocabulary_size` represents the number of embedding vectors that can be written to the embedding table, and `embedding_dimension` represents the dimension of each embedding vector. Correspondingly, address mapping uses specific rules for static indexing. For example, relevant open-source frameworks use a sharding + offset method for static indexing. For instance, the embedding table can hold 2000 embedding vectors, corresponding to 2000 addresses. These 2000 addresses are divided into 10 groups, each group corresponding to one processor. Each CPU in the 10 processors is responsible for processing the mapping of one group of addresses. The address mapping process has two parameters: shard `i` and offset `j`. Shard `i` indicates that the i-th processor needs to process the current address mapping, and offset `j` indicates that the j-th address in the i-th group needs to be processed. The value of `i` ranges from 1 to 10, and the value of `j` ranges from 1 to 200.
[0082] However, static mapping suffers from problems such as difficulty in predicting the embedding table size, feature conflicts, and memory redundancy. Specifically, the embedding table size is generally determined by the spatial size of the categorical features (representing the types of features), but in online learning scenarios, the continuous addition of new categorical features makes it difficult to estimate the embedding table size. If the embedding table is set too small, different categorical features may find the same embedding vector, resulting in feature conflicts; if the embedding table is set too large, it will store embedding vectors that are unlikely to be found, resulting in memory redundancy.
[0083] Dynamic address mapping typically uses a hash table structure, where the address mapping table is a hash table used to store key-value pairs. The key in each key-value pair is a hash value, calculated by hashing the categorical feature; this hash value can be called the feature ID. The value in the key-value pair is the address corresponding to that categorical feature. The address mapping table and embedding table can be dynamically updated with and without entries based on training data. During dynamic address mapping, if F20 (a categorical feature) arrives but is not mapped to an address in the address mapping table (i.e., a miss), a new block of memory can be dynamically allocated to index and store the address corresponding to F20 (into the address mapping table) and the embedding vector (into the embedding table).
[0084] The advantage of dynamic mapping lies in its support for feature admission and rejection functions. This means it can filter out features with excessively low frequency, preventing overfitting and thus affecting training performance. Furthermore, the entries in the address mapping table and embedding table can be dynamically increased during model training. Therefore, dynamic mapping technology can avoid memory waste or insufficient mapping space, using memory resources in a more economical and flexible way, making it easier to deploy ultra-large-scale models. The embodiments of this application can be implemented using the concept of dynamic mapping.
[0085] While related technologies can reduce memory waste or insufficient mapping space through dynamic mapping, using mutexes for address mapping during the training of recommendation models leads to the serial execution of address mapping logic among multiple threads, making it difficult to efficiently utilize the multi-core capabilities of the processor. For example, see... Figure 3 Each processor core runs a thread that executes address mapping logic. For any given training data set, multiple threads running on multiple processor cores can preemptively acquire the right to operate on that training data. Once a thread acquires the right, it locks the address mapping table and executes the address mapping logic for that training data in kernel mode. During the address mapping process, the thread first performs a read operation on the address mapping table. If the address corresponding to the training data is found (i.e., a hit), the thread sends the address to the embedding table lookup thread so that the embedding table lookup thread can retrieve the embedding vector corresponding to the training data from the embedding table based on that address. If the address corresponding to the training data is not found (i.e., a miss), the thread writes the training data and its corresponding address to the address mapping table, and then sends the address to the embedding table lookup thread so that the embedding table lookup thread can retrieve the embedding vector corresponding to the training data from the embedding table based on that address. After the address mapping logic for that training data is completed, multiple processor cores continue to preempt the right to operate on the next training data set, and so on, until the address mapping logic for all training data sets is completed.
[0086] Therefore, it is evident that in the address mapping logic of related technologies, read and write operations on the address mapping table are tightly coupled. That is, for any training data, if a read operation on the address mapping table fails, a write operation must immediately follow. However, before the read operation, the thread is unaware whether a write operation should be performed on that training data. Based on this, related technologies require the use of mutex locks to lock the address mapping table, thus preventing the use of multiple processor cores for parallel lookup operations on the address mapping table, resulting in low address mapping efficiency.
[0087] This application provides a new address mapping method that can make fuller use of the processor's multi-core capabilities while accelerating address mapping.
[0088] The implementation environment of the embodiments of this application will be described next.
[0089] The address mapping method provided in this application can be applied to any node executing address mapping logic in a recommendation system, and this node can be any type of computer device. For example, the computer device can be any computer device in a computing cluster. The computer device may include one or more processors, such as a CPU, NPU, etc., and these one or more processors can cooperate to execute part or all of the program code of the recommendation model.
[0090] As described above, embedding vector computation includes address mapping and embedding table lookup. Therefore, in one implementation, the address mapping method provided in this embodiment is executed by the CPU, meaning the CPU performs the address mapping logic, while the embedding table lookup logic can be executed by other processors, such as an NPU or GPU. Of course, with the evolution of computing architectures, the address mapping logic and the embedding table lookup logic can be executed by the same processor, such as both by the CPU or both by the NPU.
[0091] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device includes one or more processors 401, a communication bus 402, a memory 403, and one or more communication interfaces 404.
[0092] Processor 401 may include, but is not limited to, one or more of CPU, NPU, GPU, or microprocessor, or may include one or more integrated circuits for implementing the solutions of this application, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. In some embodiments, the PLD is a complex programmable logic device (CPLD), FPGA, generic array logic (GAL), or any combination thereof.
[0093] The communication bus 402 is used to transmit information between the aforementioned components. In some embodiments, the communication bus 402 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0094] The memory 403 may be a read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory (CD-ROM), compressed optical disc, laser disc, digital versatile optical disc, Blu-ray disc, etc.), magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 403 exists independently and is connected to the processor 401 via a communication bus 402, or the memory 403 is integrated with the processor 401.
[0095] Communication interface 404 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 404 includes a wired communication interface and / or a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0096] In some embodiments, the computer device includes multiple processors, such as Figure 4 The processors 401 and 405 are shown. Processor 401 may include CPU0 and CPU1, etc., and processor 405 may include an NPU or GPU, etc. Each of these processors is a single-core processor or a multi-core processor. Here, "processor" refers to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0097] As one embodiment, the computer device also includes output devices and input devices. The output devices communicate with the processor 401 and are capable of displaying information in various ways. For example, the output devices may be liquid crystal displays (LCDs), light-emitting diode (LED) displays, cathode ray tube (CRT) displays, or projectors. The input devices communicate with the processor 401 and are capable of receiving user input in various ways. For example, the input devices may be mice, keyboards, touchscreen devices, or sensing devices.
[0098] In some embodiments, memory 403 stores program code 410 for executing the scheme of this application, and processor 401 is capable of executing the program code 410 stored in memory 403. The program code includes one or more software modules, and the computer device can implement the following by means of processor 401 and program code 410 in memory 403. Figure 5 The address mapping method provided in the embodiment.
[0099] It should be understood that the system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0100] Figure 5 This is a flowchart illustrating an address mapping method provided in an embodiment of this application. The method is applied to a computer device, which may be... Figure 4 The computer equipment shown. Please refer to... Figure 5 The method includes the following steps.
[0101] Step 501: Obtain multiple training data for the recommendation model.
[0102] In some embodiments, the computer device may include a CPU, and steps 501 to 504 can all be executed by the CPU in the computer device to obtain the address corresponding to each training data. Specifically, in step 501, the CPU in the computer device acquires the plurality of training data, for example, by acquiring the plurality of training data from the memory of the computer device.
[0103] The multiple training data can be determined based on a training sample set. The training sample set can be stored on the local computer device or on other computer devices, such as storage devices in a computing cluster. This application embodiment does not limit the storage location of the training sample set.
[0104] One way to determine multiple training data based on a training sample set is to obtain multiple original training samples from the training sample set and process these original training samples to obtain multiple training data. In other words, the training data can be processed data.
[0105] As an example, taking a raw training sample as an example, one or more preprocessing operations are performed on the raw training sample to obtain training data. These preprocessing operations may include feature extraction, normalization, and hash calculation. Taking hash calculation as an example, performing a hash calculation on the raw training sample yields training data, which is a hash value. In some embodiments, the hash value may represent the feature ID of the raw training sample; that is, the training data is the feature ID of the raw training sample.
[0106] Another way to determine multiple training data based on a training sample set is to obtain multiple original training samples from the training sample set and use these multiple original training samples as multiple training data. That is, the training data can be unprocessed data.
[0107] The operation of determining multiple training data based on the training sample set can be performed by the aforementioned computer device or by other computer devices; this application embodiment does not limit this. When this operation is performed by other computer devices, the aforementioned computer device can obtain multiple training data from those other computer devices. This operation can be performed by the CPU of the computer device, or by the NPU or GPU of the computer device, and this application embodiment does not limit this either.
[0108] As an example, computer device 1 preprocesses the training sample set to obtain multiple training data and sends the multiple training data to computer device 2. Computer device 2 stores the multiple training data in its memory, and the CPU of computer device 2 can read the multiple training data from the memory, such as reading it into the CPU's memory.
[0109] In one implementation, the training data for the recommendation model includes multiple batches of training data. These batches of training data are acquired sequentially and used to train the recommendation model. The multiple batches of training data in step 501 can be any one of these batches. That is, for any batch of training data, more efficient address mapping can be performed according to the address mapping method provided in the embodiments of this application.
[0110] In this embodiment of the application, the training process of the recommendation model includes multiple iterations. In each iteration, the recommendation model is trained once using multiple training data from step 501. That is, the multiple training data will participate in multiple iterations.
[0111] As an example, computer device 1 processes the training sample set to obtain multiple batches of training data and sends the multiple batches of training data to computer device 2. Computer device 2 stores the multiple batches of training data in its memory. The CPU of computer device 2 can read the multiple batches of training data from the memory in sequence. Each batch of training data read can obtain the multiple training data in step 501.
[0112] In one implementation, the computer device reads a batch of training data, uses this data to train the recommendation model for one round, then reads the next batch, and so on, until all the training data has been read once, completing one iteration of training for the recommendation model. After this iteration, the computer device again reads the multiple batches of training data sequentially, uses each batch to train the recommendation model for one round, then reads the next batch, and so on, until all the training data has been read once, completing another iteration of training for the recommendation model. In other words, the recommendation model is trained using the same multiple batches of training data in each iteration.
[0113] As an example, there are 10 batches of training data, each containing 1000 training data points. In each iteration, the computer sequentially reads these 10 batches of training data. After reading each batch, the recommendation model is trained once using the 1000 training data points in that batch. In this way, the recommendation model will go through one round of training with these 10 batches of training data in each iteration.
[0114] In another implementation, the computer device can read a batch of training data and use it to train the recommendation model multiple times. This multiple training rounds is the iterative process using the current batch of training data. After the multiple iterations using the current batch of training data are completed, the computer device reads the next batch of training data, and so on, until all the batches of training data have been read once, completing the multiple iterative training of the recommendation model.
[0115] As an example, there are 10 batches of training data, each batch containing 1000 training data points. The computer device reads these 10 batches of training data sequentially. After reading each batch of training data, the computer device uses the 1000 training data points in the current batch to perform multiple iterations of training on the recommendation model. After performing multiple iterations of training on the recommendation model using these 1000 training data points, the computer device reads the next batch of training data, and so on, to complete the training of the recommendation model.
[0116] In the two implementation methods described above, the steps of reading training data and training the recommendation model using the training data can be executed by different processors. For example, the step of reading training data can be executed by a CPU, and the step of training the recommendation model using the training data can be executed by an NPU or a GPU. Of course, with the evolution of computing systems and the development of chips, the steps of reading training data and training the recommendation model using the training data can also be executed by the same processor, such as both by a CPU, an NPU, or a GPU. This application embodiment does not limit this.
[0117] It should be understood that the above embodiments are described using the CPU executing step 501 as an example. In practice, the execution subject of step 501 can be flexibly set. That is, step 501 can be executed by the CPU in the above-mentioned computer device, or by other processors (such as NPU). This application embodiment does not limit this.
[0118] Step 502: Divide the multiple training data into N training data, where N is an integer greater than 1.
[0119] The embodiments of this application do not limit the specific implementation of step 502, such as whether the number of N training data sets is the same, the method of partitioning, or the number of sets. For example, the number of N training data sets can be the same or different; the partitioning method can be sequential or random.
[0120] As an example, the CPU can divide the multiple training data sets equally according to their order. Taking a total of 1000 training data sets with N=3 as an example, the CPU can determine the 1st to 333rd training data set as the first set, the 334th to 667th training data set as the second set, and the 668th to 1000th training data set as the third set.
[0121] In one possible implementation, in step 501, the CPU reads the multiple training data into memory, and in step 502, the N portions of training data after being divided are also stored in memory.
[0122] Steps 501 and 502 can both be executed by the CPU in the aforementioned computer device, or by other processors; this application embodiment does not limit this.
[0123] Step 503: Perform a query operation on N training data in parallel. This query operation is used to look up the address corresponding to the training data from the address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data.
[0124] In other words, the efficiency of address mapping is improved by performing query operations on N sets of training data in parallel.
[0125] In one implementation, N read threads can be run in parallel across N processor cores (or simply cores), and these N read threads can perform query operations on N sets of training data in parallel. Specifically, one processor core runs one of the N read threads, and each read thread performs a query operation on one set of training data. This fully utilizes the multi-core capabilities of the processor, thereby improving the utilization rate of the multiple processor cores.
[0126] It should be understood that the N processor cores can be multiple cores in the same processor or multiple cores in multiple processors, and the embodiments of this application do not limit this.
[0127] In the implementation of running N read threads in parallel using N processor cores, if the N processor cores are multiple cores from the same processor, the number of parts in step 502 can be greater than 1 and not exceed the total number of cores of the processor. Taking an 8-core processor as an example, the number of parts can be greater than 1 and not greater than 8, such as 3 or 5. The number of parts can be flexibly set according to the actual situation.
[0128] As an example, the N processor cores refer to multiple cores within the same CPU. That is, the CPU comprises multiple cores. After obtaining N sets of training data in step 501, the CPU can identify N cores from among the multiple cores and determine the correspondence between these N cores and the N sets of training data. It then passes an identifier for a corresponding set of training data to each of these N cores (this identifier can be the number / sequence range of the training data, or other possible identifiers, which are not limited). In this way, each core can read the corresponding set of training data from memory based on its identifier and then perform a query operation on the read training data.
[0129] The CPU can determine N cores from multiple cores in a manner that is based on a preset configuration, the load on the multiple cores, random selection, or other methods. This embodiment does not limit the specific method used. It should be understood that when N equals the total number of cores in the CPU, the CPU does not need to perform the step of determining N cores from multiple cores.
[0130] The process of each processor core performing a query operation on its corresponding set of training data can include: the processor core runs a read thread, which sequentially reads multiple training data sets included in the same set of training data from memory. For each training data set read, the address corresponding to that training data set is looked up in the address mapping table. In other words, multiple training data sets included in the same set of training data are processed serially by a single read thread.
[0131] In this embodiment, the query operation in step 503 is implemented through a read operation. That is, step 503 only involves reading the address mapping table and not writing to it. Therefore, during the implementation of step 502, there is no need to set a mutex lock as in related technologies. Instead, the query operation is performed on N sets of training data in parallel in a lock-free state, achieving parallelization of the read operation on the address mapping table. In other words, this embodiment decouples the read and write operations of the address mapping table, enabling the query operation on the address mapping table to be performed in parallel through step 503 without locking the address mapping table during the query process.
[0132] In addition to accelerating address mapping by parallelizing read operations, this application embodiment can also divide the address mapping table into cold tables and hot tables according to the data's hotness or coldness, thereby further improving address mapping efficiency based on cold tables and hot tables.
[0133] In other words, in one implementation, the address mapping table includes a cold table and a hot table, where the frequency of training data in the cold table is lower than that in the hot table. The frequency of a training data point refers to the number of times that training data appears repeatedly in the training dataset of the recommendation model, such as the number of times the training data appears repeatedly in multiple batches of training data. Based on this, the implementation process of performing query operations on N training data points in parallel can include: for the first training data point in the N training data points, querying the address corresponding to the first training data point from the hot table; if the address corresponding to the first training data point is not found in the hot table, then querying the address corresponding to the first training data point from the cold table. Here, the first training data point can be any one of the N training data points.
[0134] Extensive data analysis shows that hot topics in the recommendation field are highly concentrated. Figure 6 This is a probability distribution diagram corresponding to an embedded table provided in an embodiment of this application. Figure 6 The horizontal axis represents a large number of feature IDs (i.e., training data), and the vertical axis represents the probability of occurrence. The curve represents the probability of occurrence for each feature ID. A higher probability of occurrence indicates higher data popularity for that feature ID, and vice versa. Figure 6It can be seen intuitively that the feature IDs with a high probability of occurrence account for a very low percentage in the embedding table, with the statistical results showing that they only account for 0.82% of the embedding table, but the sample data corresponding to these feature IDs account for 75%.
[0135] Therefore, the number of entries in the hot table is much smaller than the number of entries in the cold table. However, there is a large amount of hot data, and the probability of each training data point being a hot data point is greater than the probability of it being a cold data point. Therefore, querying the address corresponding to the training data from the hot table first has a higher probability of success (i.e., a hit) and a lower probability of failure. In this way, using hot and cold tables can further improve address mapping efficiency.
[0136] In one implementation, the hot table can be stored in storage closer to the processor, such as in a cache or direct memory.
[0137] As an example, if there is enough space in the level 0 (L1) cache, or the level 1 (L1) cache, or the level 2 (L2) cache, the hot table can be stored in the L1 cache or the L2 cache. If there is not enough space in the L1 cache and the L2 cache, but there is enough space in the level 3 (L3) cache, the hot table can be stored in the L3 cache.
[0138] Accordingly, cold tables can be stored in storage spaces far from the processor, such as in memory or even on a hard disk. As an example, cold tables are stored in dynamic random access memory (DRAM).
[0139] Of course, hot tables and cold tables can also be stored in the same storage space, such as in memory. Even so, since hot tables have fewer entries, querying hot tables will take less time.
[0140] As described above, N training data sets can be processed in parallel by N processor cores. These N processor cores can be multiple cores within the same processor or multiple cores from multiple processors. Therefore, regardless of whether the address mapping table includes cold and hot tables, the storage location of the address mapping table can be very flexible. This application does not limit the storage location of the address mapping table; the storage location can be preset. Thus, the N processor cores can perform read operations on the address mapping table according to the preset storage location.
[0141] As an example, excluding cold and hot tables in the address mapping table, if the N processor cores are multiple cores within the same processor, the address mapping table can be stored in the memory of that processor, thereby reducing the time spent looking up the address mapping table. Alternatively, the address mapping table can be stored in the computer device's memory (such as a hard disk). If the N processor cores are multiple cores within multiple processors, the address mapping table can be stored in the memory of one of those processors, or it can be stored in the computer device's memory.
[0142] When the address mapping table includes a cold table and a hot table, if the N processor cores are multiple cores within the same processor, the hot table can be stored in the cache of that processor, and the cold table can be stored in the memory of that processor. If the N processors are multiple cores within multiple processors, then a copy of the hot table can be stored in the cache of each of the N processor cores, and the cold table can be stored in the memory of the computer device.
[0143] It should be understood that when the N processor cores are multiple cores in the same processor, the hot table is stored in the processor's cache, and the cold table is stored in the processor's memory, the time taken to complete the query operation of the address mapping table in parallel using the N cores of the processor is much lower than that of related technologies, and the address mapping acceleration effect is significant.
[0144] The process of determining cold and hot meters will be described next.
[0145] In one implementation, after acquiring multiple training data sets for the recommendation model, frequency statistics are determined, which record the number of times each acquired training data set appears. Based on these frequency statistics, cold and hot tables are identified.
[0146] As can be seen from the above, the training data for the recommendation model can include multiple batches of training data, which are used sequentially to train the recommendation model. Based on this, the hot and cold lists mentioned above can be determined when the recommendation model is trained using the Mth batch of training data from these multiple batches of training data, where M is a positive integer.
[0147] The aforementioned training data can be any batch of training data. In one implementation, the CPU reads a batch of training data from memory and performs a step to determine the frequency statistics result. If the multiple training data read this time is the Mth batch of training data, the cold table and hot table are determined based on the latest frequency statistics result. It should be understood that for each batch of training data read, the CPU also performs steps 501 to 504 to determine the address corresponding to this batch of training data. Subsequently, the embedding vector is queried from the embedding table based on the address, and the recommendation model is trained based on the embedding vector. That is, the training process of the recommendation model in this embodiment can include two parts: one part includes address mapping, embedding table query, and training based on embedding vectors; the other part includes frequency statistics of the training data and the determination of cold and hot tables.
[0148] In one implementation, M is an integer multiple of the cold and hot meter update cycle, and the cold and hot meter update cycle is K batches, where K is a positive integer. That is, the cold and hot meters are updated every K batches.
[0149] The embodiments of this application do not limit the value of K. The value of K can be set according to the total number of batches of training data. For example, K is 1 / 10 of the total number of batches. As an example, K is 1000, which means that the cold table and hot table are updated every 1000 batches.
[0150] In another implementation, M can be D + K * i, where D represents the lowest batch, K represents K batches (the cold and hot meter update cycle), and i is a variable with a positive integer range. That is, when the Dth batch is read, the cold and hot meters are determined for the first time. After the Dth batch, the cold and hot meters are updated every K batches.
[0151] This application does not limit the values of D and K; both D and K can be flexibly set according to the total number of batches. As an example, D is 1000 and K is 200, meaning that the cold and hot tables are determined for the first time when the 1000th batch is read. That is, a complete address mapping table is split into cold and hot tables. After the 1000th batch is read, the cold and hot tables are updated every 200 batches, i.e., when the 1200th, 1400th, 1600th, and so on batches are read. It should be understood that in this example, when the 1000th batch is read, the current frequency statistics are already relatively reliable, and the cold and hot tables can be determined based on these statistics. This ensures that the hot tables cover most of the hot data, resulting in a high hit rate.
[0152] It should be understood that in the two implementation methods described above, although the latest frequency statistics result is determined for each batch of training data read, the latest cold and hot tables are not determined based on the latest frequency statistics result every time.
[0153] In another implementation, M is j, and j is a variable whose value range is a positive integer. That is, for each batch of training data read, the latest frequency statistics result is determined, and the latest cold table and hot table are also determined based on the latest frequency statistics result.
[0154] In this application embodiment, one way to determine cold meters and hot meters based on frequency statistics results is as follows: according to the sorting results of the frequency in the frequency statistics results, determine the frequency inflection point in the frequency statistics results, and determine the cold meters and hot meters based on the frequency inflection point.
[0155] In one implementation, the size sorting result is the result of sorting the frequencies in the frequency statistics result in descending order. Based on this, the absolute value of the difference between the non-first frequency and the previous frequency in the size sorting result can be determined first to obtain the frequency difference value corresponding to the non-first frequency. Then, based on the frequency difference value corresponding to the non-first frequency in the size sorting result, the frequency inflection point can be determined.
[0156] One approach to determining the frequency inflection point based on the frequency difference corresponding to non-first frequencies in the sorting results is as follows: Determine the maximum value among the frequency differences corresponding to the first part of the frequencies. The first part of the frequencies includes the frequencies between the P-th and Q-th frequencies in the sorting results, where P and Q are integers not exceeding the total number of frequencies in the sorting results, and P is less than Q. The frequency corresponding to this maximum value is then determined as the frequency inflection point. In other words, by excluding very high-frequency and very low-frequency data, the frequency inflection point is determined from the remaining data to avoid unreasonable frequency inflection points that result in too many or too few hot data points.
[0157] The embodiments of this application do not limit the values of P and Q, and can be flexibly set according to the actual situation. As an example, P is 100 and Q is 150000, which means that the first 99 high-frequency data and the low-frequency data after 15000 are removed from the frequency statistics results, and the frequency inflection point is determined from the frequency between the 100th frequency and the 150000th frequency.
[0158] Of course, in some embodiments, no data may be removed from the frequency statistics results, i.e., P is 1 and Q is the total number of items in the frequency statistics results.
[0159] Figure 7 This is a schematic diagram illustrating a method for determining the temperature of a hot or cold water meter, as provided in an embodiment of this application. See also... Figure 7The training data for the recommendation model consists of multiple batches, each batch containing 5 IDs (which can be hash values). When the 6th batch is obtained, the frequency statistics are as follows: Figure 7 As shown, taking P=3 and Q=7 as an example, according to the technical solution of the above embodiment, the frequency difference can be determined by sorting the results according to the size of the frequency. The frequency difference corresponding to the 3rd ID to the 7th ID is 1, 0, 2, 0, 1 respectively, with a maximum value of 2. The frequency inflection point is 2, corresponding to ID0. Based on this, the number of hot IDs (referred to as hot IDs) is 5, the set of hot IDs is {ID1ID2ID3ID4ID0}, and the set of cold IDs is {ID8ID6ID7ID9}.
[0160] In another implementation, the size sorting result is the result of ascending the frequency statistics. Based on this, we can first determine the absolute value of the difference between the non-last frequency and the next frequency in the size sorting result to obtain the frequency difference value corresponding to the non-last frequency. Then, based on the frequency difference value corresponding to the non-last frequency in the size sorting result, we can determine the frequency inflection point. The specific implementation logic is similar to that of the descending sorting logic, and will not be described in detail here.
[0161] After determining the frequency inflection point, the entries in the embedding table containing training data whose frequency is less than the frequency inflection point can be designated as entries in the cold table, and the entries in the embedding table containing training data whose frequency is not less than the frequency inflection point can be designated as entries in the hot table.
[0162] The frequency statistics and the steps for determining cold and hot meters can be performed by the CPU or other processors, and this application embodiment does not limit this.
[0163] Step 504: Based on the fact that all query operations on the N training data have been completed, write the first data and the address of the embedding vector corresponding to the first data into the address mapping table. The first data is the training data in the N training data for which no corresponding address was found.
[0164] In other words, after all read operations on the N sets of training data are completed, write operations are then performed on the training data for which no corresponding address was found. In this way, read and write operations are executed in completely different time periods, and mutex locks are not required during the query operation on the N sets of training data.
[0165] It should be understood that "not found" here means the address mapping table was not hit, that is, the query operation failed or the read operation failed. A miss may or may not occur. Step 504, which writes the first data and the address of the embedding vector corresponding to the first data into the address mapping table, is performed in the event of a miss.
[0166] Step 504 can be executed by the CPU. For example, a processor core in the CPU can start a write thread to write the first data and the address of the embedding vector corresponding to the first data into the address mapping table. The processor core can be any processor core in the CPU. The write thread can sequentially perform write operations on the address mapping table for the portion of training data for which no corresponding address was found.
[0167] Alternatively, if the query operation for the N training data is performed by N processor cores, for each processor core, if there are training data in the training data for which no corresponding address is found, the processor core starts a write thread to write the training data in the training data for which no corresponding address is found and its corresponding address to the address mapping table.
[0168] In one implementation, the first data is stored in a temporary variable space. This temporary variable space stores training data for which no corresponding address was found. That is, the temporary variable space records N training data sets for which no corresponding address was found. Based on this, before writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table, the first data can be read from the temporary variable space first.
[0169] Among these, training data for which no corresponding address was found may contain duplicates. Therefore, the temporary variable space may record identical training data. Alternatively, a deduplication operation can be performed to ensure that the training data recorded in the temporary variable space is not duplicated. The following sections will describe these cases separately.
[0170] In the first scenario, the temporary variable space may contain the same training data.
[0171] In this embodiment, the number of first data points is at least one. During the parallel execution of query operations on N training data points, each time a first data point is identified, it is written into a temporary variable space. That is, the training data written to the temporary variable space is not deduplicated.
[0172] Based on this, the process of writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table can be as follows: if the first data is read from the temporary variable space for the first time, then write the first data and the address of the embedding vector corresponding to the first data into the address mapping table; or, if the first data is read from the temporary variable space for the first time but the address corresponding to the first data is still not found in the address mapping table, then write the first data and the address of the embedding vector corresponding to the first data into the address mapping table.
[0173] As an example, after step 503 is completed, a write thread is started. This write thread reads the first data sequentially from the temporary variable space. After reading the first piece of data, it directly writes the first piece of data and its corresponding address into the address mapping table. After that, for each piece of data read, it first performs a query operation on the address mapping table for that piece of data. If the address corresponding to the first piece of data is still not found in the address mapping table, it writes the first piece of data and its corresponding address into the address mapping table. This continues until all the first data in the temporary variable space has been processed, at which point the write thread can be closed.
[0174] It should be understood that multiple first data items in the temporary variable space are processed serially. Specifically, during the time interval between the completion of the query operation in step 503 and the reading of the first first data item from the temporary variable space, the address mapping table is not updated. Therefore, the address mapping table does not yet contain the address corresponding to the currently read first data item. Thus, the first data item and its corresponding address can be directly written to the address mapping table, during which the address mapping table is updated. Therefore, if a non-first data item is read from the temporary variable space, the address mapping needs to be performed again for the current first data item. That is, it needs to be determined whether the address corresponding to the current first data item can be found in the latest address mapping table. If the address corresponding to the current first data item is still not found in the address mapping table, then the current first data item and its corresponding address are written to the address mapping table. If the address corresponding to the current first data item can be found in the address mapping table, then there is no need to perform a write operation on the address mapping table; instead, the next first data item is read from the temporary variable space.
[0175] Of course, in some embodiments, for the first data read from the temporary variable space, it can be first determined whether the address corresponding to the first data can be found in the latest address mapping table. If it still cannot be found unless there are special circumstances, then the first data and its corresponding address are written into the address mapping table.
[0176] Figure 8 This is a flowchart of another address mapping method provided in an embodiment of this application. See also... Figure 8Multiple read threads run in parallel to perform lookup operations on the address mapping table on multiple training datasets. Each read thread sequentially reads multiple training data points from its corresponding dataset. Taking one training data point as a key, it first checks if the key exists in the address mapping table. If it does, it reads the corresponding address; otherwise, it records the key in a temporary variable. After the parallel read threads finish running, a mutually exclusive write thread is started to sequentially perform write operations on the address mapping table on the training data recorded in the temporary variable. This write thread first checks if the key read from the temporary variable exists in the address mapping table. If it does, it reads the corresponding address; otherwise, it writes a new entry (the currently read key and its corresponding address) to the address mapping table. It can then read the address corresponding to the key from the address mapping table again for subsequent embedding table lookups. Furthermore, if an eviction operation on the address mapping table and embedding table is triggered (described in detail later), a temporary write thread is started to delete the evictionable entries from the address mapping table through write operations.
[0177] The aforementioned temporary variable space can be L0 or L1 cache space, thereby reducing the time consumed in step 504. The temporary variable space can also be other cache or memory space, which is not limited in this embodiment.
[0178] As an example, if the N parallel processor cores are multiple cores of the same processor, the temporary variable space can be located in the cache of that processor; if the N processor cores are multiple cores of multiple processors, the temporary variable space can be located in the memory of one of the processors or in the memory of the computer device. This application does not limit this.
[0179] In the second scenario, the training data recorded in the temporary variable space has been deduplicated, meaning there is no duplicate training data in the temporary variable space.
[0180] In this embodiment, the number of first data points is at least one. During the parallel query operation on N training data points, whenever a first data point is identified and no identical training data point is stored in the temporary variable space, the first data point is written into the temporary variable space, thereby achieving the aforementioned deduplication. Conversely, if identical training data point is stored in the temporary variable space, the first data point is not written into the temporary variable space; instead, the query operation continues for the next training data point.
[0181] After step 503 is completed, a write thread is started. The write thread reads the first data from the temporary variable space in sequence. Each time the first data is read, the first data and its corresponding address are written into the address mapping table until all the first data in the temporary variable space has been processed, at which point the write thread can be closed.
[0182] Considering that the amount of training data in the temporary variable space after deduplication is reduced in the second case compared to the first case without deduplication, the reduction in training data may affect the training effect of the recommendation model. Therefore, in the second case, in order not to affect the training effect of the recommendation model due to the reduction in training data, if there are duplicates in the training data that were not found in step 503, multiple identifier sets are recorded. Each identifier set includes one or more data identifiers. Each data identifier is an identifier for a training data (such as number or sequence number). The training data identified by the data identifiers in an identifier set are the same. During step 504, for each training data in the temporary variable space, after performing a write operation on the address mapping table for that training data, the address corresponding to that training data in the address mapping table can be read for subsequent embedding table lookups. If there are other data identifiers in the identifier set where the identifier of the training data is located, the read address is determined to be the address corresponding to the training data identified by all data identifiers in that identifier set. Thus, after performing step 504, the addresses corresponding to all training data in step 501 can be obtained, and the embedding vectors corresponding to all training data in step 501 can be obtained in subsequent embedding table lookups. That is, the embedding vectors corresponding to the training data will not be reduced due to the deduplication operation in the temporary variable space, and the training effect of the recommendation model will not be affected.
[0183] It should be understood that the above-described method of using multiple identifier sets to record duplicate training data to ensure that the deduplication operation does not affect the training effect is not intended to limit the embodiments of this application. In related implementations, other methods can also be used to ensure that the deduplication operation does not affect the training effect, and the embodiments of this application will not provide examples of these methods.
[0184] As described above, the training process of the recommendation model involves multiple iterations, and querying N training data in parallel can be performed in all iterations. Thus, in the first iteration, there might be instances of missing data on the N training data. However, after step 504 is executed in the first iteration, missing data will not occur again in subsequent iterations unless there are special circumstances. Subsequent iterations only involve read operations on the address mapping table, but no write operations are performed. Therefore, read operations far outnumber write operations during the training of the recommendation model. This embodiment decouples read and write operations on the address mapping table, enabling parallel read operations and significantly reducing the time consumption of address mapping, thereby improving the training efficiency of the recommendation model.
[0185] Figure 9 This is a schematic diagram illustrating a multiple iteration process provided in an embodiment of this application. See also... Figure 9 In the first iteration, for the multiple training data after data splitting ( Figure 9 Taking three training data sets as an example, multiple processor cores run multiple read threads (thread 1, thread 2, thread 3, ...) to perform query operations (i.e., read the address mapping table, or simply read the table) in parallel on these multiple training data sets. Each processor core is responsible for one training data set. After all query operations on these multiple training data sets are completed, for each processor core's assigned training data set, if there are any missing training data sets, that processor core runs a write thread to perform the corresponding write operation on the missing training data sets to write to the address mapping table (or simply write the table). After the first iteration, as... Figure 9 As shown, table writes will no longer occur.
[0186] In related technologies, read and write operations on address mapping tables are coupled together, making it impossible to parallelize read operations. Assuming a total read operation on multiple training data takes 3 minutes, this application decouples read and write operations and parallelizes read operations, reducing the total read operation time on multiple training data to 1 minute.
[0187] In one possible implementation, the address mapping table and the embedded table also have an elimination mechanism, that is, entries in the address mapping table and the embedded table can be eliminated (i.e. deleted) according to certain conditions. This application does not limit the specific implementation of the elimination mechanism; it can be based on popularity, timed or periodic elimination, or elimination according to other conditions. As an example, an elimination operation is triggered every hour.
[0188] When an elimination mechanism is set for the address mapping table and the embedding table, a miss may occur during the non-first iteration in the above multiple iterations. However, the probability of a miss is very low. Therefore, in the event of a miss during the non-first iteration, a write thread can be temporarily started to perform a write operation on the address mapping table. Since the probability of this situation is very low, the performance impact on the embodiments of this application can be ignored.
[0189] Steps 501 to 504 above are performed during the training of the recommendation model. That is, during the training process, the read operation and the write operation are decoupled, and the read operation is parallelized to accelerate the address mapping. Furthermore, the hot and cold table can be determined and the address mapping can be further accelerated through the hot and cold table.
[0190] Typically, the training and inference processes of a model are consistent. Based on this, the execution flow of steps 501 to 503 above can be adaptively applied to the inference stage of the recommendation model, thereby accelerating address mapping during the inference process. This will be illustrated below.
[0191] During the inference process of the recommendation model, the computer device can acquire multiple inference data of the recommendation model, divide the multiple inference data into N inference data, and perform query operations on the N inference data in parallel. The query operation is used to look up the address corresponding to the inference data from the address mapping table.
[0192] The aforementioned inference data can be determined based on an inference sample set. This inference sample set can be stored on this computer device or on other computer devices, such as the storage device within a computing cluster (which can be a different cluster or the same cluster as the one used for training; this is not limited). This application embodiment does not limit the storage location of the inference sample set. Furthermore, the inference samples in the inference sample set can be determined based on user data; this application embodiment does not limit the method of obtaining the inference samples.
[0193] The implementation method for determining multiple inference data based on an inference sample set is similar to that for determining multiple training data based on a training sample set, and will not be elaborated here. As an example, a hash calculation is performed on an inference sample in the inference sample set to obtain an inference data, which is a hash value. In some embodiments, the hash value can represent the feature ID of the inference sample, that is, the inference data is the feature ID of the inference sample.
[0194] The operation of determining multiple inference data based on the inference sample set can be performed by the aforementioned computer device or by other computer devices; this application embodiment does not limit this.
[0195] A computer device can utilize N processor cores to run N read threads in parallel, and these N read threads can perform query operations on N sets of inference data in parallel. Specifically, one processor core runs one of the N read threads, and each read thread performs a query operation on one set of inference data. This fully utilizes the multi-core capabilities of the processor, thereby improving the utilization rate of multiple processor cores.
[0196] In one implementation, the address mapping table includes a cold table and a hot table, where the training data appears less frequently in the cold table than in the hot table. Based on this, the process of performing query operations on N sets of inference data in parallel can include: for the first set of inference data, querying the address corresponding to the first set of inference data from the hot table; if the address corresponding to the first set of inference data is not found in the hot table, then querying the address corresponding to the first set of inference data from the cold table. Here, the first set of inference data can be any one of the N sets of inference data.
[0197] If the address corresponding to the first inference data is not found in the address mapping table, the computer device can determine that the query has failed.
[0198] Figure 10 This is a flowchart of yet another address mapping method provided in an embodiment of this application. See also... Figure 10 The address mapping process can take multiple training data points as input (e.g., multiple category features). Address mapping is implemented using a read operation parallelization acceleration module and a hotspot caching acceleration module. Specifically, for the multiple input training data points, steps 502 and 503 are executed by the read operation parallelization acceleration module and the hotspot caching acceleration module to query the address corresponding to each training data point. The read operation parallelization acceleration module divides the multiple training data points into N parts and performs query operations on the N parts in parallel. For each training data point, the hotspot caching acceleration module queries the hot table for the address corresponding to that training data point. If the address corresponding to that training data point is not found in the hot table, the physical address index (referred to as the address) corresponding to that training data point is queried from the cold table. The output of the address mapping process is the address corresponding to the multiple training data points.
[0199] It should be understood that, Figure 10 The process of performing a write operation on the address mapping table when a lookup fails is not shown, and this does not imply that... Figure 10 The method shown does not involve writing to the address mapping table. Figure 10 This is only used as an example to explain two key concepts in the embodiments of this application. Figure 10 The implementation process of the method shown can be referred to Figures 5 to 9 Examples are not detailed here.
[0200] In this embodiment of the application, the computer device may include a CPU and an NPU (or GPU, etc.). When steps 501 to 504 are executed by the CPU, taking the example that the computer device also includes an NPU, the CPU may send multiple addresses corresponding to training data to the NPU. The NPU may perform a query operation on the embedding table based on the received addresses to obtain the embedding vectors corresponding to the multiple training data, and then train a recommendation model based on the embedding vectors corresponding to the multiple training data.
[0201] In one possible implementation, the computer device includes multiple NPUs for parallel training of the recommendation model. The embedding table includes multiple sub-tables, with one NPU storing one sub-table and different NPUs storing different sub-tables. The address range of the embedding vectors in the different sub-tables is different. The CPU can send the addresses corresponding to the multiple training data to the multiple NPUs based on the address range of the embedding vectors in each sub-table. The address received by each NPU is located within the address range of the embedding vectors in the sub-table stored by that NPU.
[0202] In another possible implementation, the embedding table is stored in any NPU or memory, allowing for lookup of the corresponding embedding vector based on its address. The lookup operation can be performed by any processor. After retrieving the embedding vector, the processor can send it to the NPU for training the recommendation model. Specifically, the processor can divide the retrieved embedding vector into multiple groups and send these groups to multiple NPUs. Each NPU receives one group of embedding vectors and uses it to train the recommendation model.
[0203] It should be understood that the above description of the embedded table query and subsequent training process is only used as an example to illustrate the entire training process of the recommendation model and is not intended to limit the embodiments of this application.
[0204] In summary, in the embodiments of this application, during the training process of the recommendation model, read and write operations on the address mapping table can be decoupled, thereby accelerating address mapping through parallelization of read operations (i.e., query operations), thus improving the training efficiency of the recommendation model. Furthermore, address mapping can also be accelerated using a hot / cold table.
[0205] Figure 11 This is a schematic diagram of an address mapping device provided in an embodiment of this application. The address mapping device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 4 The computer equipment shown. See also Figure 11The address mapping device includes: an acquisition module 1101, a data partitioning module 1102, a parallel query module 1103, and a write module 1104.
[0206] The acquisition module 1101 is used to acquire multiple training data of the recommendation model;
[0207] The data partitioning module 1102 is used to divide multiple training data into N parts, where N is an integer greater than 1;
[0208] The parallel query module 1103 is used to perform query operations on N training data in parallel. The query operation is used to query the address corresponding to the training data from the address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data.
[0209] The write module 1104 is used to write the first data and the address of the embedding vector corresponding to the first data into the address mapping table based on the completion of the query operations on N training data. The first data is the training data in the N training data for which no corresponding address was found.
[0210] In one possible implementation, the parallel query module 1103 is specifically used for:
[0211] N read threads are run in parallel using N processor cores, and query operations are performed on N sets of training data in parallel using N read threads. Each processor core runs one of the N read threads, and each read thread performs a query operation on one set of training data.
[0212] In one possible implementation, the first data is stored in a temporary variable space, which is used to store training data for which no corresponding address was found; the address mapping device further includes:
[0213] The read module is used to read the first data from the temporary variable space.
[0214] In one possible implementation, the number of first data items is at least one, and the parallel query module 1103 is further used for:
[0215] During the parallel execution of query operations on N sets of training data, each time a first set of data is determined, that first set of data is written into a temporary variable space.
[0216] In one possible implementation, module 1104 is written, specifically for:
[0217] Based on the first data read from the temporary variable space, the first data and the address of the embedding vector corresponding to the first data are written to the address mapping table; or,
[0218] If the first data is read from the temporary variable space for the first time, and the address corresponding to the first data cannot be found in the address mapping table, then the first data and the address of the embedding vector corresponding to the first data are written to the address mapping table.
[0219] In another possible implementation, the number of first data items is at least one, and the parallel query module 1103 is also used for:
[0220] During the parallel execution of query operations on N training data, whenever a first data point is identified and no identical training data point is stored in the temporary variable space, the first data point is written into the temporary variable space.
[0221] In one possible implementation, the training process of the recommendation model includes multiple iterations, with query operations performed in parallel on N sets of training data during each iteration.
[0222] In one possible implementation, the address mapping table includes a cold table and a hot table, where the frequency of training data in the cold table is lower than that in the hot table; the parallel query module 1103 is specifically used for:
[0223] For the first training data in N training data, look up the address corresponding to the first training data from the hot table. The first training data is any one of the N training data.
[0224] If the address corresponding to the first training data is not found in the hot table, then the address corresponding to the first training data is found in the cold table.
[0225] In one possible implementation, the training data for the recommendation model includes multiple batches of training data, which are used sequentially to train the recommendation model. The hot and cold tables are determined when the recommendation model is trained using the Mth batch of training data from the multiple batches, where M is a positive integer.
[0226] In this embodiment, during the training of the recommendation model, read and write operations on the address mapping table can be decoupled, and address mapping can be accelerated by parallelizing read operations (i.e., query operations), thereby improving the training efficiency of the recommendation model. Furthermore, address mapping can also be accelerated using a hot / cold table.
[0227] It should be noted that the address mapping device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the address mapping device and the address mapping method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0228] This application also provides a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the steps of the address mapping method shown in the above method embodiments.
[0229] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps of the address mapping method shown in the above method embodiments.
[0230] This application also provides a computer program that, when run on a computer, causes the computer to perform the steps of the address mapping method shown in the above method embodiments.
[0231] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0232] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.
[0233] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the training data involved in the embodiments of this application were all obtained under full authorization.
[0234] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An address mapping method, characterized in that, The method includes: Obtain multiple training data sets for the recommendation model; The multiple training data are divided into N training data, where N is an integer greater than 1; Query operations are performed in parallel on the N training data sets. The query operation is used to query the address corresponding to the training data from the address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data. Since all query operations on the N sets of training data have been completed, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table. The first data is the training data among the N sets of training data for which no corresponding address was found.
2. The method as described in claim 1, characterized in that, The parallel query operation on the N training data includes: N read threads are run in parallel by N processor cores, and query operations are performed on the N training data in parallel by the N read threads. Each processor core runs one of the N read threads, and each read thread performs a query operation on one piece of training data.
3. The method as described in claim 1 or 2, characterized in that, The first data is stored in a temporary variable space, which is used to store training data for which no corresponding address was found; Before writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table, the method further includes: Read the first data from the temporary variable space.
4. The method as described in claim 3, characterized in that, The number of the first data is at least one, and the method further includes: During the parallel query operation on the N sets of training data, whenever a first data point is determined, the first data point is written into the temporary variable space.
5. The method as described in claim 4, characterized in that, The step of writing the first data and the address of the embedding vector corresponding to the first data into the address mapping table includes: Based on the first reading of the first data from the temporary variable space, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table; or, Based on the fact that the first data is read from the temporary variable space for the first time, and the address corresponding to the first data cannot be found in the address mapping table, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table.
6. The method as described in claim 3, characterized in that, The number of the first data is at least one, and the method further includes: During the parallel query operation on the N training data, whenever a first data is identified and no training data identical to the first data is stored in the temporary variable space, the first data is written into the temporary variable space.
7. The method according to any one of claims 1-6, characterized in that, The training process of the recommendation model includes multiple iterations, and the parallel query operation on the N training data is performed in all of these iterations.
8. The method according to any one of claims 1-7, characterized in that, The address mapping table includes a cold table and a hot table, wherein the frequency of training data in the cold table is lower than the frequency of training data in the hot table; The parallel query operation on the N training data includes: For the first training data in the N training data sets, the address corresponding to the first training data is queried from the hot table. The first training data is any one of the N training data sets. If the address corresponding to the first training data is not found in the hot table, then the address corresponding to the first training data is found in the cold table.
9. The method as described in claim 8, characterized in that, The training data for the recommendation model includes multiple batches of training data, which are used sequentially to train the recommendation model. The hot list and the cold list are determined when the recommendation model is trained using the Mth batch of training data from the multiple batches of training data, where M is a positive integer.
10. An address mapping device, characterized in that, The device includes: The acquisition module is used to acquire multiple training data sets for the recommendation model. The data partitioning module is used to divide the multiple training data into N parts of training data, where N is an integer greater than 1; A parallel query module is used to perform query operations on the N training data in parallel. The query operation is used to query the address corresponding to the training data from the address mapping table. The address mapping table is used to record the mapping relationship between the training data and the address of the embedding vector corresponding to the training data. The write module is used to write first data and the address of the embedding vector corresponding to the first data into the address mapping table after all query operations on the N training data have been completed. The first data is training data among the N training data for which no corresponding address was found.
11. The apparatus as claimed in claim 10, characterized in that, The first data is stored in a temporary variable space, which is used to store training data for which no corresponding address was found; The device further includes: A reading module is used to read the first data from the temporary variable space.
12. The apparatus as claimed in claim 11, characterized in that, The quantity of the first data is at least one, and the parallel query module is further configured to: During the parallel query operation on the N sets of training data, whenever a first data point is determined, the first data point is written into the temporary variable space.
13. The apparatus as claimed in claim 12, characterized in that, The write module is specifically used for: Based on the first reading of the first data from the temporary variable space, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table; or, Based on the fact that the first data is read from the temporary variable space for the first time, and the address corresponding to the first data cannot be found in the address mapping table, the first data and the address of the embedding vector corresponding to the first data are written into the address mapping table.
14. The apparatus as claimed in claim 11, characterized in that, The quantity of the first data is at least one, and the parallel query module is further configured to: During the parallel query operation on the N training data, whenever a first data is identified and no training data identical to the first data is stored in the temporary variable space, the first data is written into the temporary variable space.
15. The apparatus according to any one of claims 10-14, characterized in that, The training process of the recommendation model includes multiple iterations, and the parallel query operation on the N training data is performed in all of these iterations.
16. The apparatus according to any one of claims 10-15, characterized in that, The address mapping table includes a cold table and a hot table, wherein the frequency of training data in the cold table is lower than the frequency of training data in the hot table; The parallel query module is specifically used for: For the first training data in the N training data sets, the address corresponding to the first training data is queried from the hot table. The first training data is any one of the N training data sets. If the address corresponding to the first training data is not found in the hot table, then the address corresponding to the first training data is found in the cold table.
17. A computer device, characterized in that, The computer device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computer device to perform the method as described in any one of claims 1-9.
18. A computer program product, characterized in that, The computer program product stores computer instructions, which, when executed by a processor, implement the method described in any one of claims 1-9.