A Computational Method and System for a Data Redundancy-Aware Recommendation System
The data-aware recommendation system addresses memory bottlenecks in DLRMs by unifying forward and backward propagation stages and using NMP to identify and reuse hot data, improving training efficiency and reducing energy consumption.
Patent Information
- Application Number
- CN202510488826.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The embedding layer in the existing recommended model has memory-intensive features, resulting in high memory access and computing bottlenecks. The existing solutions cannot adapt to the dynamic embedding access mode, and there are problems with storage overhead and no support for embedding backpropagation.
Through the search and merging operations of the forward and backpropagation phases of the unified embedding layer, reusable data is identified in real time, reference sub-graphs are generated, and near-storage acceleration is used to reduce memory access and computing redundancy.
It realizes efficient and low-latency-free processing during the recommended model training process, adapts to the dynamic changes of the embedded layer, significantly improves performance and reduces energy consumption, and is suitable for personalized recommendation systems and other large-scale data processing fields.
Smart Images

Figure CN120030239B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, specifically to the field of recommendation system technology, and particularly relates to a calculation method and system for a recommendation system with data redundancy awareness. Background Art
[0002] The core goal of a personalized recommendation system is to use the historical behavior data of users to predict the products or services that users may be interested in, so as to provide highly customized recommendations. In this process, recommendation models, especially deep learning recommendation models (DLRMs), play a crucial role. These recommendation models can capture the personalized needs of users by learning the behavior patterns of users and achieve accurate recommendations.
[0003] The embedding layer in DLRMs is a key component of the model, which is responsible for converting the sparse features of users and products (such as the gender, age, browsing history, etc. of users) into dense vectors in a high-dimensional space. These dense vectors are then fed into the fully connected layer (FC layer), and the final recommendation results, such as the probability of predicting that a user clicks on an advertisement, are generated through further calculations.
[0004] The embedding layer consists of multiple embedding tables, and each embedding table corresponds to a categorical feature. For example, one embedding table may contain the embedding vectors of all genders, and another embedding table may contain the embedding vectors of all age groups. Each embedding vector in these tables has obtained a numerical representation that can represent its corresponding feature through the learning process. In the recommendation model, this representation method of the embedding layer can effectively capture the personalized features of users and use them for subsequent recommendation decisions.
[0005] Although the embedding layer plays an important role in the recommendation model, its memory-intensive nature also brings a series of challenges. The embedding layer needs to store a large number of embedding vectors, and the total number of these embedding vectors can range from millions to billions, and each vector may require storage space of dozens or even hundreds of dimensions. This leads to a high demand for memory bandwidth by the embedding layer, especially during the model training process, where a large number of memory accesses and data transfers become the bottleneck of performance.
[0006] To address this performance bottleneck, researchers have discovered and exploited the locality characteristics of the embedding layer. Specifically, the embedding layer exhibits two main types of locality: one is aggregation locality: in the embedding layer, only a few embedding vectors are frequently accessed, while most vectors are rarely touched. This phenomenon is very common in practical applications. For example, in a video recommendation system, the embedding vectors of some popular videos may be queried by a large number of users, while the embedding vectors of other unpopular videos are rarely accessed. The other is reduction locality: the embedding vectors that are frequently accessed are more likely to be accessed multiple times within the same time period, which means that when processing user requests, the system may operate on the same set of embedding vectors multiple times, rather than processing each embedding vector independently.
[0007] To accelerate the processing of the embedding layer, existing solutions attempt to utilize these two types of locality to reduce memory access and computational volume. These solutions identify and prepare reusable data during the preprocessing stage and reuse this data during the embedding query operation. However, these solutions have some limitations: ① Lack of awareness of dynamic queries: Existing solutions rely on predefined caches that are generated based on the analysis of historical user behavior. However, in practical applications, the embedding access patterns change dynamically over time, which makes the predefined popular vectors and their aggregated results may no longer meet the current requirements. ② Storage overhead: The additional storage overhead for storing these aggregated results may be too expensive and undesirable in the main storage or GPU storage. ③ Lack of support for embedding backpropagation: Existing solutions mainly focus on improving efficiency during the online query process. They require offline preparation of high-frequency embedding vectors and their aggregated results. However, during the training process, the dynamic changes in the embedding access patterns and the update of the embedding data make these pre-cached data no longer applicable in the new training iteration. ④ Real-time gradient generation: During the backpropagation stage, gradients are generated in real time, which makes it impossible to analyze and prepare popular gradients and their aggregated results in advance.
[0008] To overcome these limitations, there is an urgent need for a novel non-redundant near-memory processing (NMP) solution specifically for the embedding layer in recommendation model training, thereby achieving efficient, low-latency, and high energy efficiency ratio recommendation model training. Summary of the Invention
[0009] The objective of the present invention is to provide a data redundancy-aware recommendation system calculation method and system in view of the deficiencies of the prior art.
[0010] The objective of the present invention is achieved through the following technical solutions: In the first aspect of the embodiments of the present invention, a data redundancy-aware recommendation system calculation method is provided, including the following steps:
[0011] (1) Unification: During the construction of the data graph of the recommendation model, the forward and backward propagation stages of the embedding layer are uniformly processed through lookup and merging operations;
[0012] (2) Identification: Reusable data during the training process is identified in real time during the querying and filtering process; among them, the reusable data includes hot data and its partial aggregation results;
[0013] (3) Reuse: The identified reusable data is utilized during the recommendation process to generate a reference subgraph, which is used as the candidate recommendation list;
[0014] (4) Acceleration: A general NMP architecture is designed to provide redundant-free near-memory acceleration for the entire training process of the embedding layer, that is, from unification to reuse, and the reference subgraph is processed to accelerate the acquisition of the final candidate recommendation list; among them, the NMP architecture integrates dedicated processing units within each DIMM device for performing lookup and merging operations and updating embedding vectors.
[0015] Further, in the step (1), the construction process of the data graph of the recommendation model specifically includes:
[0016] First, collect the historical interaction data of users as user preference data, and at the same time collect the feature information of commodities as commodity feature data; then fuse the user preference data and the commodity feature data to construct a multi-dimensional data graph as the data graph of the recommendation model; and generate corresponding indexes for the nodes in the data graph.
[0017] Further, in the step (1), in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through a lookup operation, and then merges the multiple embedding vectors into one vector through a merging operation; in the backward propagation stage, duplicate gradients are found through a lookup operation, and then the gradient merging process is realized through a merging operation, and the embedding vectors are updated using the merged gradients.
[0018] Further, in the step (2), the query process specifically includes: splitting the user query into multiple sub-queries, and using the index to quickly find and locate the candidate recommendation items related to the user query in the data graph; among them, the user query is the real-time behavior of the user;
[0019] In the step (2), the filtering process specifically includes: sorting the multiple candidate recommendation items obtained in the query process according to the user's historical preferences, and then screening the most relevant N recommendations from the candidate recommendation items.
[0020] Further, in the step (3), the recommendation process specifically includes: generating a personalized recommendation list according to the screening results obtained in the screening process, that is, sorting the N candidate recommended items after screening according to their relevance scores, and generating a final personalized recommendation list; wherein, the recommendation list can be customized according to the user's needs.
[0021] Further, the step (3) specifically includes the following sub-steps:
[0022] (3.1) Receiving: Receiving an input query from the training data; wherein, the input query of the training data is composed of the user's input query;
[0023] (3.2) Preprocessing: Preprocessing the input query to convert the user's natural language query into structured data that the system can understand; wherein, the preprocessing includes: parsing, normalization, and encoding;
[0024] (3.3) Analyzing: According to the preprocessed query, using a scoring mechanism to calculate the hot data and its partial aggregation results in the reusable data; wherein, the scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommended item according to the user's historical behavior and user historical preference data; each candidate recommended item refers to each hot data and its partial aggregation results;
[0025] (3.4) Storing: Storing the calculated hot data and its partial aggregation results in a specification buffer, and sorting them according to their relevance scores; wherein, the specification buffer is used to temporarily store candidate recommended items, and is used to store hot data and its partial aggregation results as well as request indexes; in the forward propagation stage, the specification buffer is used to store the intermediate results of the embedding vectors; in the backward propagation stage, the specification buffer is used to store the intermediate results of the gradients;
[0026] (3.5) Screening: Screening the hot data and its partial aggregation results in the specification buffer, and screening out the hot data and its partial aggregation results whose relevance scores are greater than the user-defined threshold;
[0027] (3.6) Querying: Querying the screened hot data and its partial aggregation results from the specification buffer;
[0028] (3.7) Generating: Based on the hot data and its partial aggregation results queried from the specification buffer, generating a reference subgraph and using it as the generated candidate recommendation list.
[0029] Further, the step (4) specifically includes the following sub-steps:
[0030] (4.1)Instruction inflow: Executed in a pipelined manner. Once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture;
[0031] (4.2)Decoding: The alignment instructions are fetched from the instruction buffer and decoded;
[0032] (4.3)Alignment operation: The corresponding sub-arrays are obtained through decoding to complete the alignment operation, that is, the recommended items are selected from the candidate recommendation list using the decoded alignment instructions, and the selection process is based on user preferences and relevance scores;
[0033] (4.4)Pattern bitmask generation: Based on the BitMap algorithm, four pattern bitmasks are generated for the query read length to generate a relevance score for each candidate recommendation item according to user preferences;
[0034] (4.5)Bitwise operation: Based on the relevance scores of the candidate recommendation items, the N candidate recommendation items with the highest relevance scores are iteratively selected from the candidate recommendation list, and the insertion, deletion, replacement, and matching bit vectors of each vertex are calculated through a series of AND, OR, and SHIFT bitwise operations to combine the N candidate recommendation items into the final recommendation list.
[0035] In the second aspect of the embodiments of the present invention, a data redundancy-aware recommendation system is provided for implementing the above-mentioned data redundancy-aware recommendation system calculation method. The system includes:
[0036] A data preprocessing module for data cleaning, standardization, and feature extraction to convert the original data into a format for recommendation model training;
[0037] A user preference-aware recommendation module for optimizing the recommendation process using the user's historical behavior and user historical preference data, and adjusting the recommendation strategy in real time according to user feedback;
[0038] A receiver module for receiving input queries from users. This receiver module is responsible for interacting with the user interface, collecting the user's query information, and preprocessing the input queries to convert the user's natural language queries into structured data that the system can understand;
[0039] A real-time data reuse identification module for identifying reusable data in the forward and backward propagation stages of the embedding layer during the recommendation model training process; wherein, this real-time data reuse identification module predicts and marks the hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access patterns in the training batches;
[0040] An accelerator, which is used for a general NMP architecture based on design, and provides redundant-free near-memory acceleration for the entire training process of the embedding layer of the recommendation model from unification to reuse; wherein, the NMP architecture integrates dedicated processing units in each DIMM device for performing lookup and merge operations and updating embedding vectors;
[0041] A hardware acceleration instruction set, which is used to execute on the accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized according to the characteristics of the recommendation model, including efficient computing of the embedding layer and rapid update of gradients; and
[0042] A recommendation generator module, which is used to utilize the identified reusable data during the recommendation process, generate a reference subgraph according to user preferences, process it, and call the instructions in the hardware acceleration instruction set to execute on the accelerator to accelerate the acquisition of the final candidate recommendation list.
[0043] Furthermore, the system further includes:
[0044] An energy efficiency optimization module, which is used to monitor and regulate the use of hardware resources, and the energy efficiency optimization module realizes the optimization of the energy efficiency ratio by dynamically adjusting the working frequency and voltage of the processing unit; and / or
[0045] A security module, which is used to protect the privacy and security of user data.
[0046] A third aspect of the embodiments of the present invention provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned data redundancy-aware recommendation system calculation method.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] (1) The present invention unifies the forward and backward stages of the embedding layer through the lookup and merge (GnR) operation, and develops a general solution applicable to the entire training process of the embedding layer; the present invention can real-time identify the reusable data during the training process through the data reusable awareness method, and reuse the hot data and partial aggregation results to eliminate memory access and computing redundancy; the present invention implements this solution through a general NMP architecture, provides redundant-free near-memory acceleration for the entire training process of the embedding layer, and eliminates memory access and computing redundancy in the forward and backward training stages of the embedding layer.
[0049] (2) The present invention can real-time identify and utilize the reusability of data, thereby reducing memory access and computing amount; it is not only applicable to the forward propagation of the embedding layer, but also applicable to the backward propagation, so that the present invention can provide acceleration during the entire training process of the recommendation model.
[0050] (3) The NMP solution of the present invention can adapt to the dynamic changes in the embedding layer training without additional storage overhead. It does not require pre-storing popular vectors or their aggregated results. This real-time and redundancy-free method enables the present invention to perform excellently in the training of recommendation models with different localities. Whether in low-locality models or high-locality models, the present invention can significantly improve performance and reduce energy consumption, and supports backpropagation of the embedding layer, having good practicality and promotion value.
[0051] (4) The present invention reduces the energy consumption in the training process of the recommendation model from the calculations of the processor, memory access, and data transfer within the system. By reducing the number of these operations, it reduces memory access and computational volume, significantly reducing the energy consumption of the system and making the training of the recommendation model more efficient.
[0052] (5) The present invention effectively reduces memory access and computational redundancy in the training process of the recommendation model, significantly improving the training efficiency and reducing energy consumption. By real-time identifying and reusing data, the present invention effectively solves the performance bottleneck problem of the embedding layer in the training of the recommendation model, not only improving the training efficiency of the recommendation model but also reducing energy consumption, which is of great significance for promoting the development of personalized recommendation systems. It is not only applicable to online recommendation systems but also can be extended to other fields that require large-scale data processing and machine learning algorithms, such as financial risk control, medical diagnosis, intelligent manufacturing, etc. Brief Description of the Drawings
[0053] Figure 1 is an example diagram of the forward propagation stage of the embedding layer in the training process of the recommendation model of the present invention;
[0054] Figure 2 is an example diagram of the backpropagation stage of the embedding layer in the training process of the recommendation model of the present invention;
[0055] Figure 3 is an example diagram of the present invention for eliminating memory access redundancy in the forward propagation stage;
[0056] Figure 4 is an example diagram of the present invention for eliminating computational redundancy in the forward propagation stage;
[0057] Figure 5 is a block diagram of the overall architecture of the system of the present invention. Detailed Embodiments
[0058] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0059] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0060] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0061] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners may be combined with each other.
[0062] The data redundancy-aware recommendation system calculation method and system of the present invention is an innovative redundancy-free near-memory processing (NMP) solution specifically designed to accelerate the training process of the embedding layer in deep learning recommendation models (DLRMs). As a core component of DLRMs, the embedding layer is responsible for converting sparse user and item features into dense vector representations in a high-dimensional space, and this process is crucial for the recommendation performance of the model.
[0063] The core idea of the present invention is to eliminate redundancy in memory access and calculation by real-time identifying and reusing reusable data during the training process of the embedding layer. This process not only improves the training efficiency but also significantly reduces energy consumption.
[0064] The data redundancy-aware recommendation system calculation method of the present invention specifically includes the following steps:
[0065] (1) Unification: In the process of constructing the data graph of the recommendation model, the forward and backward propagation stages of the embedding layer are uniformly processed through the Gather-Reduce (GnR) operation to construct a general framework applicable to the entire training process of the embedding layer. This general framework includes the unification in step (1), the identification in step (2) below, the reuse in step (3) below, and the acceleration in step (4) below.
[0066] Specifically, in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through the gather operation and then merges the multiple embedding vectors into one vector through the reduce operation, as Figure 1 shown; similarly, in the backward propagation stage, duplicate gradients are found through the gather operation, and then the gradient merging process is achieved through the reduce operation, and the embedding vectors are updated using the merged gradients, as Figure 2 shown. Through this unified processing, the present invention can develop a general solution applicable to the entire training process of the embedding layer.
[0067] Further, the process of constructing the data graph of the recommendation model specifically includes: First, collect the historical interaction data of users, including but not limited to users' click, purchase, browsing, and rating behaviors; at the same time, also collect the feature information of products, such as product descriptions, classifications, prices, and user evaluations, etc.; by fusing user preference data and product feature data, construct a multi-dimensional data graph as the data graph of the recommendation model, which can comprehensively reflect the relationship between users and products. And generate corresponding indexes for the nodes in the data graph for quick retrieval in the following steps; the purpose of the index is to improve the query efficiency. By constructing an efficient index structure, such as a hash table, an inverted index, or a balanced tree, etc., the search for nodes in the data graph can be quickly completed.
[0068] (2) Identification: In the query and filtering process, identify the reusable data in the training process in real time; among them, the reusable data includes hot data and its partial aggregation results for reuse in the following steps.
[0069] Further, the query process specifically includes: splitting the user query into multiple sub-queries and quickly finding and locating the candidate recommendation items related to the user query in the data graph using the index. Among them, the user query is the real-time behavior of the user, such as search keywords, click events, or browsing history, etc. Split these user queries into multiple sub-queries and quickly find the candidate recommendation items related to the user query through the index structure.
[0070] It should be understood that the user query contains multi-terminal information, so it needs to be divided into different sub-queries according to different information combinations.
[0071] Further, the screening process specifically includes: sorting multiple candidate recommendation items obtained during the query process according to the user's historical preferences, and then screening out the most relevant N recommendations from the candidate recommendation items, so as to effectively improve the accuracy and relevance of the recommendations.
[0072] (3) Reuse: Utilize the identified reusable data during the recommendation process to generate a reference subgraph and use it as a candidate recommendation list, which can reduce memory access and computational redundancy and improve training efficiency.
[0073] Further, the recommendation process specifically includes: generating a personalized recommendation list based on the screening results obtained during the screening process, that is, sorting the N screened candidate recommendation items according to their relevance scores and generating a final personalized recommendation list. Among them, the recommendation list can be customized according to the user's needs, such as adjusting the number, type, and order of recommendations, etc.
[0074] In this embodiment, by analyzing the unique embedding vectors of each training batch and their reusability, the elimination of memory access redundancy is achieved. In DLRMs training, usually only a few embedding vectors are accessed multiple times, while most embedding vectors are accessed less frequently. The present invention utilizes this locality characteristic and significantly reduces the number of memory accesses by only accessing those unique and frequently reused embedding vectors. Specifically, it includes the following steps:
[0075] (3.1) Receive: Receive the input query from the training data. Among them, the input query of the training data is composed of the user's input query.
[0076] (3.2) Preprocess: Preprocess the input query to convert the user's natural language query into structured data that the system can understand. Among them, the preprocessing includes but is not limited to: parsing, normalization, and encoding, etc.
[0077] (3.3) Analyze: According to the preprocessed query, use the scoring mechanism to calculate the hot data and its partial aggregation results in the reusable data. Among them, the scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommendation item according to the user's historical behavior and user historical preference data; each candidate recommendation item refers to each hot data and its partial aggregation results.
[0078] (3.4) Store: Store the calculated hot data and its partial aggregation results in the reduction buffer and sort them according to their relevance scores so that subsequent steps can be processed efficiently. Among them, the reduction buffer is used to temporarily store candidate recommendation items, store hot data and its partial aggregation results as well as request indexes; in the forward propagation stage, the reduction buffer is used to store the intermediate results of the embedding vectors; in the backward propagation stage, the reduction buffer is used to store the intermediate results of the gradients.
[0079] See Figure 3 , in the present invention, a specification buffer is maintained to record the request index of each query and the corresponding embedded vector value. When a new embedded vector is accessed, the system checks whether it has been cached in the specification buffer. If so, the vector is directly reused, avoiding repeated access to memory.
[0080] (3.5) Screening: Screen the hot data in the specification buffer and its partial aggregation results, and screen out the hot data and its partial aggregation results whose relevance scores are greater than the user-defined threshold, which can reduce the number of candidate seed positions that need to be considered for calculation. This step is to ensure that only the most relevant candidate recommendations can enter the next recommendation generation process.
[0081] (3.6) Query: Query the screened hot data and its partial aggregation results from the specification buffer for the subsequent recommendation generation process.
[0082] (3.7) Generation: Based on the hot data and its partial aggregation results queried from the specification buffer, generate a reference subgraph and use it as the generated candidate recommendation list to provide basic data for the subsequent alignment step.
[0083] It should be noted that computational redundancy is eliminated by preferentially accessing popular data and reusing their partial aggregation results. During the training of the embedding layer, certain combinations of embedded vectors may appear multiple times, resulting in repeated calculations. By identifying these frequently occurring combinations and storing their results after the first calculation, subsequent repeated calculations are avoided. As Figure 4 shown, by intelligently sorting data access requests, those embedded vectors that are expected to be reused multiple times are preferentially processed. This method not only reduces the amount of computation but also improves the efficiency of data processing.
[0084] (4) Acceleration: Design a general NMP architecture to provide redundancy-free near-memory acceleration for the entire training process of the embedding layer, that is, from unification to reuse, and process the reference subgraph to accelerate the acquisition of the final candidate recommendation list. Among them, the NMP architecture integrates dedicated processing units within each DIMM (Dual In-line Memory Module) device, and these processing units, namely the NMP cores, are responsible for performing lookup and merge operations as well as updating of embedded vectors.
[0085] It should be understood that the NMP core is the core component of the present invention for performing forward and backward operations of the embedding layer. The design of the NMP core allows it to process data in a near-memory manner, thereby reducing the transfer of data between memory and the processor.
[0086] (4.1)Instruction inflow: Executed in a pipelined manner. Once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture. The pipelined execution improves the processing efficiency of the system and allows multiple instructions to be processed in parallel.
[0087] (4.2)Decoding: The alignment instructions are fetched from the instruction buffer and decoded to convert the alignment instructions into specific operation steps.
[0088] (4.3)Alignment operation: The corresponding sub-array is decoded to complete the alignment operation, that is, the recommended items are selected from the candidate recommendation list using the decoded alignment instructions, and the selection process is based on user preferences and relevance scores.
[0089] (4.4)Pattern bitmask generation: Based on the BitMap algorithm, four pattern bitmasks are generated for the query read length to generate a relevance score for each candidate recommendation item according to user preferences. The generation of the relevance score can be based on machine learning models such as collaborative filtering, deep learning, or reinforcement learning, etc.
[0090] It should be understood that in the BitMap algorithm, BitMap is a very useful data structure that uses a single bit to mark the value corresponding to a certain element (value), and the key is the element. Since BitMap uses bits to store data, it can greatly save storage space.
[0091] (4.5)Bitwise operations: Based on the relevance scores of the candidate recommendation items, the N candidate recommendation items with the highest relevance scores are iteratively selected from the candidate recommendation list, and the insertion, deletion, replacement, and matching bit vectors of each vertex are calculated through a series of AND, OR, and SHIFT bitwise operations to combine the N candidate recommendation items into the final recommendation list. This step ensures the personalization and accuracy of the recommendation list.
[0092] It is worth mentioning that the present invention also provides a data redundancy-aware recommendation system for implementing the data redundancy-aware recommendation system calculation method in the above embodiments. As Figure 5 shown, the system includes a data preprocessing module, a user preference-aware recommendation module, a receiver module, a real-time data reuse identification module, an accelerator, a hardware acceleration instruction set, and a recommendation generator module. The system is executed on the host to implement the data redundancy-aware recommendation system calculation method. The host includes multiple memory modules and memory channels, as Figure 5 shown in (a) therein, and the structural composition of each memory module is as Figure 5 shown in (b) therein.
[0093] In this embodiment, the data preprocessing module is used for data cleaning, standardization, and feature extraction to convert the original data into a format suitable for training the recommendation model, which can improve the training efficiency and accuracy of the recommendation model.
[0094] In this embodiment, the user preference-aware recommendation module is used to optimize the recommendation process by leveraging the user's historical behavior and user historical preference data, which can effectively improve the accuracy of recommendations. The user preference-aware recommendation module includes an adaptive learning algorithm that can adjust the recommendation strategy in real time according to user feedback.
[0095] In this embodiment, the receiver module is used to receive input queries from users. The receiver module is responsible for interacting with the user interface, collecting the user's query information, and preprocessing the input query to convert the user's natural language query into structured data that the system can understand.
[0096] In this embodiment, the real-time data reuse identification module is used to identify reusable data in the forward and backward propagation stages of the embedding layer during the training process of the recommendation model. Among them, the real-time data reuse identification module predicts and marks hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access patterns in the training batches, thereby reducing unnecessary memory access and calculations.
[0097] In this embodiment, the accelerator is based on a general NMP architecture design and provides redundant-free near-memory acceleration for the entire training process of the recommendation model's embedding layer from unification to reuse. The NMP architecture integrates dedicated processing units within each DIMM device to perform lookup and merge operations and update the embedding vectors.
[0098] In this embodiment, the hardware acceleration instruction set is used to execute on the accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized according to the characteristics of the recommendation model, including efficient calculations of the embedding layer, rapid updates of gradients, etc., as shown in (c) of Figure 5 ; among them, efficient calculations of the embedding layer are performed through the lookup and merge process shown in i) of Figure 5 , and the rapid update process of gradients is shown in ii) of Figure 5 .
[0099] In this embodiment, the recommendation generator module is used to utilize the identified reusable data during the recommendation process, generate a reference subgraph according to user preferences, process it, and call the instructions in the hardware acceleration instruction set to execute on the accelerator to accelerate the acquisition of the final candidate recommendation list. The recommendation generator module can also adjust the recommendation strategy and output format according to different application scenarios and user requirements.
[0100] Furthermore, the system also includes an energy efficiency optimization module, which is used to monitor and regulate the use of hardware resources to minimize energy consumption; this energy efficiency optimization module realizes the optimization of the energy efficiency ratio by dynamically adjusting the working frequency and voltage of the processing unit.
[0101] Furthermore, the system also includes a security module, which is used to protect the privacy and security of user data; this security module implements data encryption and access control mechanisms to ensure that only authorized users and systems can access sensitive data.
[0102] In this embodiment, during the backpropagation stage, there is a corresponding update mode, which is responsible for using the combined gradients to update the embedding vectors, ensuring that the recommendation model can learn from the training data and gradually optimize its recommendation performance.
[0103] Specifically, during the forward propagation stage of the embedding layer, as Figure 5 shown in i) of
[0104] Step S10, Lookup: As shown in operation ① of Figure 5 , look up the required embedding vectors from the memory. This step is the starting point of data processing, and specifically obtains data through an efficient memory access mechanism. For example, the state of the reduction buffer after accessing vector ② in Figure 4 is as shown in (d) of Figure 5 .
[0105] Step S11, Comparison: As shown in operation ② of Figure 5 , compare the ID of the accessed embedding vector with the request index in the reduction buffer to determine whether the vector has been cached.
[0106] Step S12, Merge: As shown in operation ③ of Figure 5 , send the accessed embedding vector and the corresponding partial aggregation result to the computing unit (CU) for merging operation. This step involves element-level calculations and is the core of the forward propagation of the embedding layer.
[0107] Step S13, Write-back: As shown in operation ④ of Figure 5 , write the result of the merge operation back to the reduction buffer. This step ensures that the calculation results can be utilized by subsequent operations, thus avoiding repeated calculations.
[0108] Specifically, during the backpropagation stage of the embedding layer, as Figure 5 shown in ii) of
[0109] Step S20, Gradient Lookup: As shown in Figure 5As shown in operation ⑤ in , look up and access the gradient information related to the embedding vector in memory. This step is the starting point of backpropagation and requires obtaining the update information for each embedding vector.
[0110] Step S21, Compare gradients: As Figure 5 shown in operation ⑥ in , compare the vector ID with the gradient index in the reduction buffer to determine whether there is already corresponding update information.
[0111] Step S22, Update: Update the embedding vector using the corresponding combined gradient. This step is crucial for the training of the recommendation model, ensuring that the recommendation model can optimize itself according to the training data.
[0112] Step S23, Write back the update: As Figure 5 shown in operation ⑦ in , write the updated embedding vector back to memory. This step is the end point of backpropagation, ensuring that the update of the recommendation model parameters can be persisted.
[0113] In summary, the data redundancy-aware recommendation system calculation method and system of the present invention have the following characteristics: ① Real-time data reusability awareness: It can identify reusable hot data and its partial aggregation results in real time during the training process; ② General NMP architecture: It can provide non-redundant near-memory acceleration for the entire training process of the embedding layer; ③ Reduce storage overhead: Avoid the additional storage overhead required to store pre-computed hot vectors and their aggregation results; ④ Support embedding backpropagation: It can adapt to the dynamic update of embedding data during the training process and real-time gradient generation in the backpropagation stage. Through the above modules, the present invention achieves a significant acceleration of the training process of the recommendation system, while maintaining the accuracy of recommendations and the security of user data, and is applicable to the real-time processing requirements of large-scale personalized recommendation services.
[0114] Corresponding to the foregoing embodiments of the data redundancy-aware recommendation system calculation method, the present invention also provides an embodiment of an electronic device.
[0115] An electronic device provided by an embodiment of the present invention includes one or more processors and a memory, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the data redundancy-aware recommendation system calculation method in the above embodiments.
[0116] Embodiments of the electronic device of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. Embodiments of the electronic device can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From a hardware perspective, it is a hardware structure diagram of any device with data processing capabilities where the electronic device of the present invention is located. In addition to the processor, memory, network interface, and non-volatile memory, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0117] For the implementation processes of the functions and roles of each unit in the above-mentioned electronic device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0118] For embodiments of the electronic device, since they basically correspond to method embodiments, relevant parts can be referred to the partial descriptions of the method embodiments. The embodiments of the electronic device described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0119] Embodiments of the present invention also provide a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the data redundancy-aware recommendation system calculation method in the above embodiments.
[0120] The computer-readable storage medium may be an internal storage unit of any data processing-capable device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any data processing-capable device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any data processing-capable device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any data processing-capable device, and may also be used to temporarily store data that has been output or is to be output.
[0121] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A calculation method for a data redundancy-aware recommendation system, characterized in that The following steps are involved: (1) Unification: In the process of building the data graph of the recommendation model, the forward propagation and backpropagation stages of the embedding layer are unified through the search and merge operations; (2) Identification: Real-time identification of reusable data in the training process during query and screening; reusable data includes hot data and its partial aggregation results; (3) Reuse: Use the identified reusable data in the recommendation process to generate reference subgraphs and use them as candidate recommendation lists; The step (3) specifically includes the following sub-steps: (3.1) Receiving: receiving an input query from training data; wherein the input query of the training data is composed of an input query of a user; (3.2) Preprocessing: Preprocess the input query to convert the user's natural language query into structured data that the system can understand; preprocessing includes: parsing, normalization and encoding; (3.3) Analysis: Based on the preprocessed query, the scoring mechanism is used to calculate the hot data and its partial aggregation results in the reusable data; the scoring mechanism is a machine learning model that can assign a relevance score to each candidate recommendation item based on the user's historical behavior and user's historical preference data; each candidate recommendation item refers to each hot data and its partial aggregation results; (3.4) Storage: The calculated hotspot data and its partial aggregation results are stored in the reduction buffer and sorted according to their relevance scores; the reduction buffer is used to temporarily store candidate recommendation items, to store hotspot data and its partial aggregation results, and to request indexes; in the forward propagation phase, the reduction buffer is used to store the intermediate results of the embedding vector; in the backward propagation phase, the reduction buffer is used to store the intermediate results of the gradient; (3.5) Screening: Screen the hotspot data and their partial aggregation results in the specification buffer, and select the hotspot data and their partial aggregation results whose correlation scores are greater than a user-defined threshold; (3.6) Query: Query the filtered hot data and its partial aggregation results from the specification buffer; (3.7) Generation: Based on the hotspot data queried from the specification buffer and its partial aggregation results, a reference subgraph is generated and used as the candidate recommendation list; (4) Acceleration: A general NMP architecture is designed to provide non-redundant near-storage acceleration for the entire training process of the embedding layer, from unification to reuse, and to process the reference subgraphs to accelerate the acquisition of the final candidate recommendation list. The NMP architecture integrates a dedicated processing unit in each DIMM device to perform search and merge operations and update the embedding vector.
2. The method for calculating a data redundancy-aware recommendation system according to claim 1, wherein In the step (1), the process of constructing the data graph of the recommendation model specifically includes: First, the user's historical interaction data is collected as user preference data, and the feature information of the goods is collected as product feature data; then the user preference data and product feature data are fused to construct a multidimensional data graph as the data graph of the recommendation model; and corresponding indexes are generated for the nodes in the data graph.
3. The method for calculating a data redundancy-aware recommendation system according to claim 1, wherein In step (1), in the forward propagation stage, the embedding layer retrieves multiple embedding vectors through a lookup operation, and then merges the multiple embedding vectors into one vector through a merging operation; In the backward propagation stage, duplicate gradients are found through a lookup operation, and then the gradient merging process is achieved through a merging operation, and the embedding vectors are updated using the merged gradients.
4. The method for calculating a data redundancy-aware recommendation system according to claim 1, wherein In step (2), the query process specifically includes: splitting the user query into multiple sub-queries, and quickly finding and locating candidate recommendations related to the user query in the data graph using an index; where the user query is the user's real-time behavior; In step (2), the filtering process specifically includes: sorting the multiple candidate recommendations obtained in the query process according to the user's historical preferences, and then filtering out the most relevant N recommendations from the candidate recommendations.
5. The method for calculating a data redundancy-aware recommendation system according to claim 1, wherein In step (3), the recommendation process specifically includes: generating a personalized recommendation list based on the filtering results obtained in the filtering process, that is, sorting the filtered N candidate recommendations according to their relevance scores, and generating a final personalized recommendation list; where the recommendation list can be customized according to the user's needs.
6. The method for calculating a data redundancy-aware recommendation system according to claim 1, wherein Step (4) specifically includes the following sub-steps: (4.1) Instruction inflow: Executed in a pipeline manner. Once the reference subgraph is generated, the alignment instructions flow into the instruction buffer of the processing unit in the NMP architecture; (4.2) Decoding: Taking out the alignment instructions from the instruction buffer and decoding them; (4.3) Alignment operation: Decoding to obtain the corresponding sub-array to complete the alignment operation, that is, using the decoded alignment instructions to select recommendations from the candidate recommendation list, and the selection process is based on user preferences and relevance scores; (4.4) Pattern bitmask generation: Based on the BitMap algorithm, generating four pattern bitmasks for the query read length to generate a relevance score for each candidate recommendation according to user preferences; (4.5) Bitwise operation: Based on the relevance scores of the candidate recommendations, iteratively selecting the N candidate recommendations with the highest relevance scores from the candidate recommendation list, and calculating the insertion, deletion, replacement, and matching bit vectors of each vertex through a series of AND, OR, and SHIFT bitwise operations to combine the N recommendations into a final recommendation list.
7. A data redundancy-aware recommendation system for implementing the data redundancy-aware recommendation system calculation method according to any one of claims 1-6, characterized in that, The system includes: A data preprocessing module for data cleaning, standardization, and feature extraction to convert the raw data into a format for recommendation model training; A user preference-aware recommendation module for optimizing the recommendation process using the user's historical behavior and user historical preference data, and adjusting the recommendation strategy in real time according to user feedback; A receiver module for receiving input queries from the user. This receiver module is responsible for interacting with the user interface, collecting the user's query information, and preprocessing the input query to convert the user's natural language query into structured data that the system can understand; A real-time data reuse recognition module, which is used to recognize reusable data in the forward propagation and backward propagation stages of the embedding layer during the training process of the recommendation model; wherein, the real-time data reuse recognition module predicts and marks hot data and its partial aggregation results that will be reused in subsequent operations by analyzing the data access patterns in the training batches. An accelerator, which is used to provide redundant-free near-memory acceleration for the entire training process of the recommendation model embedding layer from unification to reuse based on a general NMP architecture designed; wherein, the NMP architecture integrates dedicated processing units within each DIMM device to perform lookup and merge operations and update of embedding vectors. A hardware acceleration instruction set, which is used to be executed on the accelerator to accelerate the training process of the recommendation model; the hardware acceleration instruction set is optimized according to the characteristics of the recommendation model, including efficient computation of the embedding layer and rapid update of gradients; and A recommendation generator module, which is used to utilize the recognized reusable data during the recommendation process, generate a reference subgraph according to user preferences, process it, and call the instructions in the hardware acceleration instruction set to be executed on the accelerator to accelerate the acquisition of the final candidate recommendation list.
8. The data redundancy-aware recommendation system according to claim 7, wherein The system further includes: An energy efficiency optimization module, which is used to monitor and regulate the use of hardware resources, and the energy efficiency optimization module realizes the optimization of the energy efficiency ratio by dynamically adjusting the working frequency and voltage of the processing unit; and / or A security module, which is used to protect the privacy and security of user data.
9. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the data redundancy-aware recommendation system calculation method according to any one of claims 1-6.
Citation Information
Patent Citations
Tourism information processing and plan providing method
CN105468679A
Sequence recommendation method and system based on multilayer perceptron and self-attention mechanism
CN117708433A