Data redundancy perception recommendation system calculation method and system
By unifying the propagation stage of the embedding layer in the recommended model and identifying reusable data in real time, combined with the general NMP architecture, the problem of memory-intensive performance bottleneck in embedding layer training is solved, and highly efficient, low-latency, and low-energy-consuming recommended model training is achieved.
Patent Information
- Application Number
- CN202510488826.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The memory-intensive features of the embedded layer in the training of existing recommended models lead to performance bottlenecks, especially when large amounts of memory access and data handling during model training become performance bottlenecks.
By unifying the forward and backpropagation stages of the embedding layer during the data graph construction process of the recommended model, reusable data during the training process is identified in real time, and a general redundant-free near-storage processing (NMP) architecture is designed to provide redundant-free near-storage acceleration for the entire training process of the embedding layer.
It realizes the reduction of memory access and computing redundancy in recommended model training, significantly improves training efficiency and reduces energy consumption, adapts to the dynamic changes of embedded layer training, and supports backpropagation of embedded layer without additional storage overhead.
Smart Images

Figure CN120030239A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, specifically to the field of recommendation system technology, and in particular to a data redundancy-aware recommendation system calculation method and system. Background Art
[0002] The core goal of personalized recommendation systems is to use users' historical behavior data to predict the products or services that users may be interested in, thereby providing highly customized recommendations. In this process, recommendation models, especially deep learning recommendation models (DLRMs), play a vital role. These recommendation models can capture users' personalized needs and achieve accurate recommendations by learning users' behavior patterns.
[0003] The embedding layer in DLRMs is a key component of the model, which is responsible for converting the sparse features of users and products (such as the user's gender, age, browsing history, etc.) into dense vectors in a high-dimensional space. These dense vectors are then fed into the fully connected layer (FC layer) to generate the final recommendation results through further calculations, such as predicting the probability of a user clicking on an advertisement.
[0004] The embedding layer consists of multiple embedding tables, each of which corresponds to a categorical feature. For example, one embedding table may contain embedding vectors for all genders, and another embedding table may contain embedding vectors for all age groups. Each embedding vector in these tables obtains a numerical representation that can represent its corresponding feature through the learning process. In the recommendation model, this representation method of the embedding layer can effectively capture the personalized characteristics of users and use them for subsequent recommendation decisions.
[0005] Although the embedding layer plays an important role in the recommendation model, its memory-intensive nature also brings a series of challenges. The embedding layer needs to store a large number of embedding vectors, the total number of which can range from millions to billions, and each vector may require dozens or even hundreds of dimensions of storage space. This leads to a high demand for memory bandwidth in the embedding layer, especially during model training, where a large amount of memory access and data handling becomes a performance bottleneck.
[0006] To address this performance bottleneck, researchers discovered and exploited the locality characteristics of the embedding layer. Specifically, the embedding layer exhibits two main localities: one is clustered locality: in the embedding layer, only a few embedding vectors are frequently accessed, while most vectors are rarely touched. This phenomenon is very common in practical applications. For example, in a video recommendation system, the embedding vectors of some popular videos may be queried by a large number of users, while other unpopular videos are rarely accessed. The other is reduction locality: frequently accessed embedding vectors are more likely to be accessed multiple times at the same time, which means that when processing user requests, the system may operate on the same set of embedding vectors multiple times instead of processing each embedding vector independently.
[0007] To speed up the processing of the embedding layer, existing solutions try to exploit these two localities to reduce memory access and computation. These solutions identify and prepare reusable data in the preprocessing stage and reuse this data in the embedding query operation. However, these solutions have some limitations: ① Dynamic query unawareness: Existing solutions rely on predefined caches that are generated based on the analysis of historical user behavior. But in practical applications, the embedding access pattern changes dynamically over time, which makes the predefined hot vectors and their aggregation results may no longer meet the current needs. ② Storage overhead: The increased storage overhead for storing these aggregation results may be too expensive and undesirable in the main storage or GPU storage. ③ No support for embedding backpropagation: Existing solutions mainly focus on improving efficiency during online query processes, and they require offline preparation of high-frequency embedding vectors and their aggregation results. But during the training process, the dynamic changes in embedding access patterns and the update of embedded data make these pre-cached data no longer applicable in new training iterations. ④ Real-time gradient generation: In the backpropagation stage, gradients are generated in real time, which makes it impossible to analyze and prepare hot gradients and their aggregation results in advance.
[0008] To overcome these limitations, a novel non-redundant near-memory processing (NMP) solution is urgently needed, specifically targeting the embedding layer in recommendation model training, thereby achieving efficient, low-latency, and energy-efficient recommendation model training. Summary of the invention
[0009] The purpose of the present invention is to provide a data redundancy-aware recommendation system calculation method and system to address the deficiencies of the prior art.
[0010] The objective of the present invention is achieved through the following technical solutions: In a first aspect, an embodiment of the present invention provides a data redundancy-aware recommendation system calculation method, comprising the following steps: (1) Unification: In the process of building the data graph of the recommendation model, the forward propagation and backpropagation stages of the embedding layer are unified through the lookup and merge operations; (2) Identification: Real-time identification of reusable data in the training process during query and screening. Reusable data includes hot data and its partial aggregation results. (3) Reuse: Use the identified reusable data in the recommendation process to generate reference subgraphs and use them as candidate recommendation lists; (4) Acceleration: A general NMP architecture is designed to provide non-redundant near-storage acceleration for the entire training process of the embedding layer, from unification to reuse, and to process the reference subgraphs to accelerate the acquisition of the final candidate recommendation list. The NMP architecture integrates a dedicated processing unit in each DIMM device to perform lookup and merge operations and update the embedding vector.
[0011] Furthermore, in step (1), the process of constructing the data graph of the recommendation model specifically includes: First, the user's historical interaction data is collected as user preference data, and the feature information of the goods is collected as product feature data; then the user preference data and product feature data are fused to construct a multidimensional data graph as the data graph of the recommendation model; and corresponding indexes are generated for the nodes in the data graph.
[0012] Furthermore, in the step (1), in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through a search operation, and then merges the multiple embedding vectors into one vector through a merge operation; in the backward propagation stage, the repeated gradients are found through a search operation, and then the gradient merging process is implemented through a merge operation, and then the merged gradients are used to update the embedding vector.
[0013] Furthermore, in step (2), the query process specifically includes: dividing the user query into multiple sub-queries, and using the index to quickly search and locate candidate recommendation items related to the user query in the data graph; wherein the user query is the user's real-time behavior; In the step (2), the screening process specifically includes: sorting the multiple candidate recommendation items obtained during the query process according to the user's historical preferences, and then screening the most relevant N recommendations from the candidate recommendation items.
[0014] Furthermore, in step (3), the recommendation process specifically includes: generating a personalized recommendation list based on the screening results obtained in the screening process, that is, sorting the screened N candidate recommendation items according to their relevance scores, and generating a final personalized recommendation list; wherein the recommendation list can be customized according to user needs.
[0015] Furthermore, the step (3) specifically includes the following sub-steps: (3.1) Receiving: receiving an input query from training data; wherein the input query of the training data is composed of an input query of a user; (3.2) Preprocessing: Preprocess the input query to convert the user's natural language query into structured data that the system can understand; preprocessing includes: parsing, normalization and encoding; (3.3) Analysis: Based on the preprocessed query, the scoring mechanism is used to calculate the hot data and its partial aggregated results in the reusable data; the scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommendation item based on the user's historical behavior and user's historical preference data; each candidate recommendation item refers to each hot data and its partial aggregated results; (3.4) Storage: The calculated hot data and its partial aggregation results are stored in the reduction buffer and sorted according to their relevance scores; the reduction buffer is used to temporarily store candidate recommendations, hot data and its partial aggregation results, and request indexes; in the forward propagation phase, the reduction buffer is used to store the intermediate results of the embedding vector; in the backward propagation phase, the reduction buffer is used to store the intermediate results of the gradient; (3.5) Screening: Screen the hotspot data and their partial aggregation results in the specification buffer, and select the hotspot data and their partial aggregation results whose relevance scores are greater than the user-defined threshold; (3.6) Query: Query the filtered hot data and its partial aggregation results from the specification buffer; (3.7) Generation: Based on the hotspot data queried from the specification buffer and its partial aggregation results, a reference subgraph is generated and used as the generated candidate recommendation list.
[0016] Furthermore, the step (4) specifically includes the following sub-steps: (4.1) Instruction inflow: Executed in a pipelined manner, once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture; (4.2) Decoding: Take the alignment instruction from the instruction buffer and decode it; (4.3) Alignment operation: Decode the corresponding subarray to complete the alignment operation, that is, use the decoded alignment instructions to select recommended items from the candidate recommendation list. The selection process is based on user preferences and relevance scores. (4.4) Pattern bit mask generation: Based on the BitMap algorithm, four pattern bit masks are generated for the query read length to generate a relevance score for each candidate recommendation item based on user preferences; (4.5) Bitwise operations: Based on the relevance scores of the candidate recommendation items, iteratively select the N recommendation items with the highest relevance scores from the candidate recommendation list, and calculate the insertion, deletion, replacement, and matching bit vectors of each vertex through a series of AND, OR, and SHIFT bitwise operations to combine the N recommendation items into the final recommendation list.
[0017] A second aspect of an embodiment of the present invention provides a data redundancy-aware recommendation system, which is used to implement the above-mentioned data redundancy-aware recommendation system calculation method. The system includes: Data preprocessing module, which is used for data cleaning, standardization and feature extraction to convert raw data into a format for recommendation model training; The user preference-aware recommendation module is used to optimize the recommendation process by utilizing the user's historical behavior and user historical preference data, and adjust the recommendation strategy in real time based on user feedback; A receiver module, for receiving an input query from a user, the receiver module is responsible for interacting with the user interface, collecting the user's query information, and preprocessing the input query to convert the user's natural language query into structured data that the system can understand; A real-time data reuse identification module is used to identify reusable data in the forward propagation and back-propagation stages of the embedding layer in real time during the recommendation model training process; wherein, the real-time data reuse identification module predicts and marks hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access pattern in the training batch; An accelerator for providing near-storage acceleration without redundancy for the entire training process of the embedding layer of the recommendation model from unification to reuse based on a general NMP architecture designed in which the NMP architecture integrates a dedicated processing unit in each DIMM device to perform lookup and merge operations and update the embedding vector; A hardware acceleration instruction set for execution on an accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized for the characteristics of the recommendation model, including efficient calculation of the embedding layer and fast update of the gradient; and The recommendation generator module is used to utilize the identified reusable data in the recommendation process, generate reference subgraphs according to user preferences, process them, and call instructions in the hardware acceleration instruction set to execute on the accelerator to accelerate the acquisition of the final candidate recommendation list.
[0018] Furthermore, the system further comprises: An energy efficiency optimization module, used to monitor and adjust the use of hardware resources, the energy efficiency optimization module optimizes the energy efficiency ratio by dynamically adjusting the operating frequency and voltage of the processing unit; and / or Security module, used to protect the privacy and security of user data.
[0019] A third aspect of an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned data redundancy-aware recommendation system calculation method.
[0020] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention unifies the forward and reverse stages of the embedding layer through the search and merge (GnR) operation, and develops a general solution applicable to the entire embedding layer training process; the present invention can identify reusable data in the training process in real time through a data reusability awareness method, and reuse hot data and partial aggregation results to eliminate memory access and computational redundancy; the present invention implements this solution through a general NMP architecture, providing non-redundant near-storage acceleration for the entire embedding layer training process, and eliminating memory access and computational redundancy in the forward and reverse training stages of the embedding layer.
[0021] (2) The present invention can identify and utilize the reusability of data in real time, thereby reducing memory access and computational complexity; it is applicable not only to the forward propagation of the embedding layer but also to the backpropagation, enabling the present invention to provide acceleration in the entire training process of the recommendation model.
[0022] (3) The NMP solution of the present invention can adapt to the dynamic changes of embedding layer training without additional storage overhead. It does not need to pre-store hot vectors or their aggregation results. This real-time, non-redundant method enables the present invention to perform well in the training of recommendation models with different localities. Whether in low-locality models or high-locality models, the present invention can significantly improve performance and reduce energy consumption, and supports back propagation of the embedding layer, which has good practicality and promotion value.
[0023] (4) The present invention reduces the energy consumption of processor calculations, memory accesses, and data movement within the system during the recommendation model training process. By reducing the number of these operations, the memory access and calculation amount are reduced, and the energy consumption of the system is significantly reduced, making the training of the recommendation model more efficient.
[0024] (5) The present invention effectively reduces memory access and computational redundancy during recommendation model training, significantly improves training efficiency and reduces energy consumption. The present invention effectively solves the performance bottleneck problem of the embedding layer in recommendation model training by real-time recognition and reuse of data, which not only improves the efficiency of recommendation model training but also reduces energy consumption. It is of great significance to promote the development of personalized recommendation systems. It is not only suitable for online recommendation systems, but can also be extended to other fields that require large-scale data processing and machine learning algorithms, such as financial risk control, medical diagnosis, intelligent manufacturing, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is an example diagram of the forward propagation stage of the embedding layer in the recommendation model training process of the present invention; Figure 2 is an example diagram of the back propagation stage of the embedding layer in the recommendation model training process of the present invention; Figure 3 is an example diagram for eliminating memory access redundancy in the forward propagation stage of the present invention; Figure 4 is an example diagram for eliminating computational redundancy in the forward propagation stage of the present invention; Figure 5 It is a block diagram of the overall system architecture of the present invention. DETAILED DESCRIPTION
[0026] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0027] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0028] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0029] The present invention is described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the features of the following embodiments and implementations can be combined with each other.
[0030] The data redundancy-aware recommendation system computing method and system of the present invention is an innovative non-redundant near memory processing (NMP) solution, which is specially designed to accelerate the training process of the embedding layer in deep learning recommendation models (DLRMs). As the core component of DLRMs, the embedding layer is responsible for converting sparse user and item features into dense vector representations in high-dimensional space, which is crucial to the recommendation performance of the model.
[0031] The core idea of the present invention is to eliminate redundancy in memory access and computation by identifying and reusing reusable data in the embedding layer training process in real time. This process not only improves the efficiency of training, but also significantly reduces energy consumption.
[0032] The data redundancy-aware recommendation system calculation method of the present invention specifically includes the following steps: (1) Unification: In the process of constructing the data graph of the recommendation model, the forward propagation and backpropagation stages of the embedding layer are unified through the Gather-Reduce (GnR) operation to build a general framework applicable to the entire embedding layer training process. The general framework includes the unification of step (1), the identification of the following step (2), the reuse of the following step (3), and the acceleration of the following step (4).
[0033] Specifically, in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through a lookup operation, and then merges the multiple embedding vectors into one vector through a merge operation, such as Figure 1 Similarly, in the back-propagation phase, the repeated gradients are found through the search operation, and then the gradient merging process is realized through the merge operation, and then the merged gradients are used to update the embedded vector, as shown in Figure 2 Through this unified processing, the present invention can develop a general solution applicable to the entire embedding layer training process.
[0034] Furthermore, the construction process of the data graph of the recommendation model specifically includes: firstly, collecting the historical interaction data of users, including but not limited to the user's click, purchase, browsing and rating behaviors; at the same time, also collecting the characteristic information of the goods, such as the description, classification, price and user evaluation of the goods; by integrating the user preference data and the characteristic data of the goods, a multi-dimensional data graph is constructed as the data graph of the recommendation model, which can fully reflect the relationship between users and goods. And generate corresponding indexes for the nodes in the data graph so as to quickly retrieve them in the next step; the purpose of the index is to improve the query efficiency, and by constructing an efficient index structure, such as a hash table, inverted index or balanced tree, the search for nodes in the data graph can be completed quickly.
[0035] (2) Identification: Identify reusable data in the training process in real time during the query and screening process; the reusable data includes hot data and its partial aggregation results so that they can be reused in the next step.
[0036] Furthermore, the query process specifically includes: dividing the user query into multiple sub-queries, and using the index to quickly find and locate candidate recommendation items related to the user query in the data graph. The user query is the user's real-time behavior, such as search keywords, click events, or browsing history. These user queries are divided into multiple sub-queries, and the candidate recommendation items related to the user query are quickly found through the index structure.
[0037] It should be understood that user queries contain multi-end information and therefore need to be divided into different sub-queries according to different information combinations.
[0038] Furthermore, the screening process specifically includes: sorting multiple candidate recommendation items obtained during the query process according to the user's historical preferences, and then screening the most relevant N recommendations from the candidate recommendation items, which can effectively improve the accuracy and relevance of the recommendations.
[0039] (3) Reuse: In the recommendation process, the identified reusable data is used to generate reference subgraphs and use them as candidate recommendation lists, which can reduce memory access and computational redundancy and improve training efficiency.
[0040] Furthermore, the recommendation process specifically includes: generating a personalized recommendation list according to the screening results obtained in the screening process, that is, sorting the N candidate recommendation items after screening according to their relevance scores, and generating a final personalized recommendation list. Among them, the recommendation list can be customized according to the needs of the user, such as adjusting the number, type and order of recommendations.
[0041] In this embodiment, the elimination of memory access redundancy is achieved by analyzing the unique embedding vectors of each training batch and their reusability. In DLRMs training, usually only a few embedding vectors are accessed multiple times, while most embedding vectors are rarely accessed. The present invention utilizes this locality feature and significantly reduces the number of memory accesses by accessing only those unique and frequently reused embedding vectors. Specifically, the following steps are included: (3.1) Receiving: receiving input queries from training data, wherein the input queries of the training data are composed of user input queries.
[0042] (3.2) Preprocessing: Preprocess the input query to convert the user's natural language query into structured data that the system can understand. Preprocessing includes but is not limited to: parsing, normalization, and encoding.
[0043] (3.3) Analysis: Based on the preprocessed query, the scoring mechanism is used to calculate the hot data and its partial aggregation results in the reusable data. The scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommendation item based on the user's historical behavior and user's historical preference data; each candidate recommendation item refers to each hot data and its partial aggregation results.
[0044] (3.4) Storage: The calculated hot data and its partial aggregation results are stored in the reduction buffer and sorted according to their relevance scores so that subsequent steps can be processed efficiently. The reduction buffer is used to temporarily store candidate recommendations, to store hot data and its partial aggregation results, and to request indexes; in the forward propagation phase, the reduction buffer is used to store the intermediate results of the embedding vector; in the backward propagation phase, the reduction buffer is used to store the intermediate results of the gradient.
[0045] See also Figure 3 The present invention maintains a protocol buffer to record the request index and the corresponding embedding vector value of each query. When a new embedding vector is accessed, the system checks whether it has been cached in the protocol buffer. If so, the vector is directly reused, avoiding re-access to the memory.
[0046] (3.5) Screening: Screen the hot data and its partial aggregation results in the specification buffer, and select the hot data and its partial aggregation results with a relevance score greater than the user-defined threshold, which can reduce the number of candidate seed positions that need to be considered for calculation. This step is to ensure that only the most relevant candidate recommendation items can enter the next step of the recommendation generation process.
[0047] (3.6) Query: Query the filtered hot data and its partial aggregation results from the specification buffer for subsequent recommendation generation process.
[0048] (3.7) Generation: Based on the hotspot data queried from the specification buffer and its partial aggregation results, a reference subgraph is generated and used as the generated candidate recommendation list to provide basic data for the subsequent alignment step.
[0049] It should be noted that computational redundancy is eliminated by giving priority to accessing popular data and reusing some of their aggregated results. During the training of the embedding layer, some combinations of embedding vectors may appear multiple times, resulting in repeated computations. By identifying these frequently occurring combinations and storing their results after the first computation, subsequent repeated computations are avoided. Figure 4 As shown in Figure 3, by intelligently sorting data access requests, we prioritize those embedded vectors that are expected to be reused multiple times. This approach not only reduces the amount of computation but also improves the efficiency of data processing.
[0050] (4) Acceleration: Design a general NMP architecture to provide non-redundant near-storage acceleration for the entire training process of the embedding layer, from unification to reuse, and process the reference subgraphs to accelerate the acquisition of the final candidate recommendation list. The NMP architecture integrates dedicated processing units in each DIMM (Dual In-line Memory Module) device. These processing units, namely NMP cores, are responsible for performing lookup and merge operations and updating the embedding vectors.
[0051] It should be understood that the NMP core is a core component of the present invention, which is used to perform forward and reverse operations of the embedded layer. The design of the NMP core allows it to process data in a near-memory manner, thereby reducing the transmission of data between the memory and the processor.
[0052] (4.1) Instruction inflow: Once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture in a pipelined manner. The pipelined execution improves the processing efficiency of the system and allows multiple instructions to be processed in parallel.
[0053] (4.2) Decoding: Take the alignment instruction from the instruction buffer and decode it to convert the alignment instruction into specific operation steps.
[0054] (4.3) Alignment operation: The corresponding subarray is decoded to complete the alignment operation, that is, the decoded alignment instruction is used to select the recommended items from the candidate recommendation list. The selection process is based on user preferences and relevance scores.
[0055] (4.4) Pattern bit mask generation: Based on the BitMap algorithm, four pattern bit masks are generated for the query read length to generate a relevance score for each candidate recommendation item based on user preferences. The generation of relevance scores can be based on machine learning models such as collaborative filtering, deep learning, or reinforcement learning.
[0056] It should be understood that in the BitMap algorithm, BitMap is a very useful data structure, which uses a bit to mark the value corresponding to an element, and the key is the element. Since BitMap uses bits to store data, it can greatly save storage space.
[0057] (4.5) Bit operations: Based on the relevance scores of the candidate recommendation items, iteratively select the N recommended items with the highest relevance scores from the candidate recommendation list, and calculate the insertion, deletion, replacement, and matching bit vectors of each vertex through a series of AND, OR, and SHIFT bit operations to combine the N recommended items into the final recommendation list. This step ensures the personalization and accuracy of the recommendation list.
[0058] It is worth mentioning that the present invention also provides a data redundancy-aware recommendation system, which is used to implement the data redundancy-aware recommendation system calculation method in the above embodiment. Figure 5 As shown, the system includes a data preprocessing module, a user preference perception recommendation module, a receiver module, a real-time data reuse identification module, an accelerator, a hardware acceleration instruction set and a recommendation generator module, and the system is executed on a host to implement a data redundancy-aware recommendation system calculation method. The host includes multiple memory modules and memory channels, such as Figure 5 As shown in (a), the structure of each memory module is as follows Figure 5 As shown in (b) in .
[0059] In this embodiment, the data preprocessing module is used for data cleaning, standardization and feature extraction to convert the original data into a format suitable for recommendation model training, which can improve the training efficiency and accuracy of the recommendation model.
[0060] In this embodiment, the user preference-aware recommendation module is used to optimize the recommendation process by utilizing the user's historical behavior and user historical preference data, which can effectively improve the accuracy of the recommendation. The user preference-aware recommendation module includes an adaptive learning algorithm that can adjust the recommendation strategy in real time based on user feedback.
[0061] In this embodiment, the receiver module is used to receive input queries from users. The receiver module is responsible for interacting with the user interface, collecting user query information, and preprocessing the input queries to convert the user's natural language queries into structured data that the system can understand.
[0062] In this embodiment, the real-time data reuse identification module is used to identify reusable data in the forward propagation and back propagation stages of the embedding layer in real time during the recommendation model training process, wherein the real-time data reuse identification module predicts and marks hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access patterns in the training batch, thereby reducing unnecessary memory access and calculation.
[0063] In this embodiment, the accelerator is used to provide non-redundant near-storage acceleration for the entire training process from unification to reuse of the embedding layer of the recommendation model based on a designed general NMP architecture. The NMP architecture integrates a dedicated processing unit in each DIMM device to perform search and merge operations and update the embedding vector.
[0064] In this embodiment, the hardware acceleration instruction set is used to execute on the accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized according to the characteristics of the recommendation model, including efficient calculation of the embedding layer, fast update of the gradient, etc. Figure 5 As shown in (c) of Figure 5 The search and merging process shown in (i) in the figure performs efficient calculation of the embedding layer, and the fast update process of the gradient is as follows: Figure 5 As shown in ii).
[0065] In this embodiment, the recommendation generator module is used to use the identified reusable data in the recommendation process, generate reference subgraphs according to user preferences, process them, and call instructions in the hardware acceleration instruction set to execute on the accelerator to accelerate the acquisition of the final candidate recommendation list. The recommendation generator module can also adjust the recommendation strategy and output format according to different application scenarios and user needs.
[0066] Furthermore, the system also includes an energy efficiency optimization module, which is used to monitor and adjust the use of hardware resources to minimize energy consumption; the energy efficiency optimization module optimizes the energy efficiency ratio by dynamically adjusting the operating frequency and voltage of the processing unit.
[0067] Furthermore, the system also includes a security module, which is used to protect the privacy and security of user data; the security module implements data encryption and access control mechanisms to ensure that only authorized users and systems can access sensitive data.
[0068] In this embodiment, during the back propagation phase, an update mode is correspondingly provided, and the update mode is responsible for using the merged gradient to update the embedding vector, thereby ensuring that the recommendation model can learn from the training data and gradually optimize its recommendation performance.
[0069] Specifically, in the forward propagation stage of the embedding layer, as Figure 5 As shown in i), perform the following sub-steps: Step S10, search: Figure 5 As shown in operation ① in , the required embedding vector is found from the memory. This step is the starting point of data processing. Specifically, data is obtained through an efficient memory access mechanism, for example, access Figure 4 The state of the reduction buffer after vector ② in is as follows Figure 5 As shown in (d) in .
[0070] Step S11, comparison: Figure 5 As shown in operation ② in , the ID of the accessed embedded vector is compared with the request index in the specification buffer to determine whether the vector has been cached.
[0071] Step S12: Merge: Figure 5 As shown in operation ③ in , the accessed embedding vector and the corresponding partial aggregation result are sent to the computing unit (CU) for merging. This step involves element-level calculations and is the core of the forward propagation of the embedding layer.
[0072] Step S13, write back: Figure 5 As shown in operation ④ in , the result of the merge operation is written back to the specification buffer. This step ensures that the calculation result can be used by subsequent operations, thus avoiding repeated calculations.
[0073] Specifically, in the back-propagation phase of the embedding layer, as Figure 5 As shown in ii), perform the following sub-steps: Step S20, search for gradient: Figure 5 As shown in operation ⑤ in , the gradient information related to the embedded vector is accessed from the memory. This step is the starting point of back propagation, and it is necessary to obtain the update information of each embedded vector.
[0074] Step S21, compare gradients: Figure 5 As shown in operation ⑥ in , the vector ID is compared with the gradient index in the specification buffer to determine whether there is corresponding update information.
[0075] Step S22, Update: Update the embedding vector using the corresponding merged gradient. This step is the key to training the recommendation model, which ensures that the recommendation model can self-optimize based on the training data.
[0076] Step S23, write back update: Figure 5 As shown in operation ⑦ in , the updated embedding vector is written back to the memory. This step is the end point of backpropagation, which ensures that the update of the recommendation model parameters can be persisted.
[0077] In summary, the data redundancy-aware recommendation system calculation method and system described in the present invention include the following features: ① Real-time data reusability awareness: capable of identifying reusable hotspot data and its partial aggregation results in real time during the training process; ② Universal NMP architecture: capable of providing non-redundant near-storage acceleration for the entire training process of the embedding layer; ③ Reduced storage overhead: avoiding the additional storage overhead required to store pre-calculated hotspot vectors and their aggregation results; ④ Support for embedded back propagation: capable of adapting to the dynamic update of embedded data during the training process and real-time gradient generation in the back propagation stage. Through the above modules, the present invention achieves a significant acceleration of the recommendation system training process while maintaining the accuracy of recommendations and the security of user data, and is suitable for the real-time processing requirements of large-scale personalized recommendation services.
[0078] Corresponding to the above-mentioned embodiment of the data redundancy-aware recommendation system calculation method, the present invention also provides an embodiment of an electronic device.
[0079] An electronic device provided by an embodiment of the present invention includes one or more processors and a memory, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the data redundancy-aware recommendation system calculation method in the above embodiment.
[0080] The embodiments of the electronic device of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The electronic device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located to read the corresponding computer program instructions in the non-volatile memory into the memory and run it. From the hardware level, it is a hardware structure diagram of any device with data processing capabilities in which the electronic device of the present invention is located. In addition to the processor, memory, network interface, and non-volatile memory, any device with data processing capabilities in which the device in the embodiment is located can also include other hardware according to the actual function of the device with data processing capabilities, which will not be repeated here.
[0081] The implementation process of the functions and effects of each unit in the above electronic device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0082] For the electronic device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The electronic device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0083] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the data redundancy-aware recommendation system calculation method in the above embodiment is implemented.
[0084] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0085] The above embodiments are only used to illustrate the design ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.
Claims
1. A data redundancy-aware recommendation system calculation method, characterized in that: The following steps are involved: (1) Unification: In the process of building the data graph of the recommendation model, the forward propagation and backpropagation stages of the embedding layer are unified through the lookup and merge operations; (2) Identification: Real-time identification of reusable data in the training process during query and screening. Reusable data includes hot data and its partial aggregation results. (3) Reuse: Use the identified reusable data in the recommendation process to generate reference subgraphs and use them as candidate recommendation lists; (4) Acceleration: A general NMP architecture is designed to provide non-redundant near-storage acceleration for the entire training process of the embedding layer, from unification to reuse, and to process the reference subgraphs to accelerate the acquisition of the final candidate recommendation list. The NMP architecture integrates a dedicated processing unit in each DIMM device to perform lookup and merge operations and update the embedding vector.
2. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: In step (1), the process of constructing the data graph of the recommendation model specifically includes: First, the user's historical interaction data is collected as user preference data, and the feature information of the goods is collected as product feature data; then the user preference data and product feature data are fused to construct a multidimensional data graph as the data graph of the recommendation model; and corresponding indexes are generated for the nodes in the data graph.
3. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: In the step (1), in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through a search operation, and then merges the multiple embedding vectors into one vector through a merge operation; in the backward propagation stage, the repeated gradients are found through a search operation, and then the gradient merging process is implemented through a merge operation, and then the merged gradients are used to update the embedding vector.
4. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: In step (2), the query process specifically includes: dividing the user query into multiple sub-queries, and using the index to quickly search and locate candidate recommendation items related to the user query in the data graph; wherein the user query is the user's real-time behavior; In the step (2), the screening process specifically includes: sorting the multiple candidate recommendation items obtained during the query process according to the user's historical preferences, and then screening the most relevant N recommendations from the candidate recommendation items.
5. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: In the step (3), the recommendation process specifically includes: generating a personalized recommendation list based on the screening results obtained in the screening process, that is, sorting the screened N candidate recommendation items according to their relevance scores, and generating a final personalized recommendation list; wherein the recommendation list can be customized according to the needs of the user.
6. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: The step (3) specifically includes the following sub-steps: (3.1) Receiving: receiving an input query from training data; wherein the input query of the training data is composed of an input query of a user; (3.2) Preprocessing: Preprocess the input query to convert the user's natural language query into structured data that the system can understand; preprocessing includes: parsing, normalization and encoding; (3.3) Analysis: Based on the preprocessed query, the scoring mechanism is used to calculate the hot data and its partial aggregated results in the reusable data; the scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommendation item based on the user's historical behavior and user's historical preference data; each candidate recommendation item refers to each hot data and its partial aggregated results; (3.4) Storage: The calculated hot data and its partial aggregation results are stored in the reduction buffer and sorted according to their relevance scores; the reduction buffer is used to temporarily store candidate recommendations, hot data and its partial aggregation results, and request indexes; in the forward propagation phase, the reduction buffer is used to store the intermediate results of the embedding vector; in the backward propagation phase, the reduction buffer is used to store the intermediate results of the gradient; (3.5) Screening: Screen the hotspot data and their partial aggregation results in the specification buffer, and select the hotspot data and their partial aggregation results whose relevance scores are greater than the user-defined threshold; (3.6) Query: Query the filtered hot data and its partial aggregation results from the specification buffer; (3.7) Generation: Based on the hotspot data queried from the specification buffer and its partial aggregation results, a reference subgraph is generated and used as the generated candidate recommendation list.
7. The data redundancy-aware recommendation system calculation method according to claim 1, characterized in that: The step (4) specifically includes the following sub-steps: (4.1) Instruction inflow: Executed in a pipelined manner, once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture; (4.2) Decoding: Take the alignment instruction from the instruction buffer and decode it; (4.3) Alignment operation: Decode the corresponding subarray to complete the alignment operation, that is, use the decoded alignment instructions to select recommended items from the candidate recommendation list. The selection process is based on user preferences and relevance scores. (4.4) Pattern bit mask generation: Based on the BitMap algorithm, four pattern bit masks are generated for the query read length to generate a relevance score for each candidate recommendation item based on user preferences; (4.5) Bitwise operations: Based on the relevance scores of the candidate recommendation items, iteratively select the N recommendation items with the highest relevance scores from the candidate recommendation list, and calculate the insertion, deletion, replacement, and matching bit vectors of each vertex through a series of AND, OR, and SHIFT bitwise operations to combine the N recommendation items into the final recommendation list.
8. A data redundancy-aware recommendation system, used to implement the data redundancy-aware recommendation system calculation method according to any one of claims 1 to 7, characterized in that: The system comprises: Data preprocessing module, which is used for data cleaning, standardization and feature extraction to convert raw data into a format for recommendation model training; The user preference-aware recommendation module is used to optimize the recommendation process by utilizing the user's historical behavior and user historical preference data, and adjust the recommendation strategy in real time based on user feedback; A receiver module, for receiving an input query from a user, the receiver module is responsible for interacting with the user interface, collecting the user's query information, and preprocessing the input query to convert the user's natural language query into structured data that the system can understand; A real-time data reuse identification module is used to identify reusable data in the forward propagation and back-propagation stages of the embedding layer in real time during the recommendation model training process; wherein, the real-time data reuse identification module predicts and marks hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access pattern in the training batch; An accelerator for providing near-storage acceleration without redundancy for the entire training process of the embedding layer of the recommendation model from unification to reuse based on a general NMP architecture designed in which the NMP architecture integrates a dedicated processing unit in each DIMM device to perform lookup and merge operations and update the embedding vector; A hardware acceleration instruction set for execution on an accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized for the characteristics of the recommendation model, including efficient calculation of the embedding layer and fast update of the gradient; and The recommendation generator module is used to utilize the identified reusable data in the recommendation process, generate reference subgraphs according to user preferences, process them, and call instructions in the hardware acceleration instruction set to execute on the accelerator to accelerate the acquisition of the final candidate recommendation list.
9. The data redundancy-aware recommendation system according to claim 8, characterized in that: The system further comprises: An energy efficiency optimization module, used to monitor and adjust the use of hardware resources, the energy efficiency optimization module optimizes the energy efficiency ratio by dynamically adjusting the operating frequency and voltage of the processing unit; and / or Security module, used to protect the privacy and security of user data.
10. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the data redundancy-aware recommendation system calculation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for online parallel computing of recommended information, device for online parallel computing of recommended information, and server for online parallel computing of recommended information
CN104090894A
Tourism information processing and plan providing method
CN105468679A
Responsive recommendation method, system and equipment based on meta-learning
CN115409579A
Sequence recommendation method and system based on multilayer perceptron and self-attention mechanism
CN117708433A
Distributed recommendation method and system based on near data processing
CN119127727A