Inference method and device, system, equipment, medium and product of recommendation system
Patent Information
- Application Number
- CN202510288537.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2026-09-11
AI Technical Summary
大模型的推理时延通常较高,难以满足推荐系统的时延要求
[0043] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
Smart Images

Figure CN122736712A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of recommendation technology, and in particular to a reasoning method, reasoning device, recommendation system, computing device cluster, computer-readable storage medium, and computer program product for a recommendation system. Background Technology
[0002] With the continuous development of artificial intelligence (AI), especially the rapid development of large language models (LLM) represented by generative pre-trained transformers (GPT), more and more operators are choosing to apply large models such as LLM to recommendation systems for generative recommendations.
[0003] For example, video platforms can use large models like LLM (Lifecycle Management) to infer user behavior based on browsing, playback, and commenting, thereby recommending personalized videos. Similarly, e-commerce platforms can use large models like LLM to infer the match between candidate products and users based on user behavior such as adding items to favorites, shopping carts, and placing orders, thus recommending products accordingly.
[0004] However, recommender systems have strict latency requirements, typically within 15 milliseconds (ms). Large models often have high inference latency, making it difficult to meet the latency requirements of recommender systems. Therefore, providing a low-latency inference method has become a key focus in the industry. Summary of the Invention
[0005] This application provides an inference method for a recommender system. Based on the business characteristic of fixed historical behavior sequences of users in recommender systems, this method modifies the autoregressive decoding of large models into parallel decoding. By using multiple inference nodes for parallel decoding, the inference latency is significantly reduced, meeting the latency requirements of recommender systems. This application also provides an inference device, recommender system, computing device cluster, computer-readable storage medium, and computer program product corresponding to the above method.
[0006] Firstly, this application provides a reasoning method for a recommender system. This method can be applied to recommender systems. The recommender system can be a software system, such as a software system in an e-commerce platform for implementing product recommendations, or a software system in a video platform for implementing video recommendations. The aforementioned software system can be a standalone software system, or it can be integrated into other systems as a plugin, service, etc. In some possible implementations, the recommender system may also include a hardware system; for example, the recommender system may include a cluster of computing devices, which executes the reasoning method of this application when running.
[0007] Specifically, the recommender system can receive recommendation inference requests, which request the inference of a target item from the candidate set of the recommender system. The target item is at least one candidate item in the candidate set. The recommender system then schedules multiple candidate items to multiple inference nodes for parallel decoding. Each inference node decodes at least one candidate item according to the recommender model, and each inference node decodes different candidate items. The recommender model is obtained by training a large model. The recommender system aggregates the decoding results from the multiple inference nodes and determines the target item based on the aggregated decoding results.
[0008] Based on the business characteristic of fixed historical behavior sequences of users in recommendation systems, this method modifies the autoregressive decoding of large models to parallel decoding. Specifically, multiple candidate options in the candidate set are scheduled to multiple inference nodes. Each inference node decodes at least one of the multiple candidate options, and each inference node decodes different candidate options. This enables parallel inference by multiple inference nodes, significantly reducing inference latency and meeting the latency requirements of recommendation systems.
[0009] In some possible implementations, the candidate set includes multiple batches of candidates. The recommender system can determine the required memory for each batch of candidates within the multiple batches based on at least one of the following: batch size, sequence length of the input sequence, vector dimension of the candidate candidates, hidden layer dimension, or number of transformer layers in the recommender model. The recommender system can then determine multiple inference nodes based on the required memory.
[0010] This method estimates the memory requirements of each batch of candidates, thereby scheduling each batch of candidates to appropriate inference nodes for parallel decoding, improving decoding performance, shortening decoding time, and thus reducing inference latency.
[0011] In some possible implementations, for the first batch of candidates in multiple batches, the recommendation system determines the first inference node from the inference resource pool. The remaining GPU memory of the first inference node is greater than or equal to the required GPU memory of the first batch of candidates.
[0012] This method pools the resources of inference nodes to form an inference resource pool, and performs unified scheduling and management of the inference resource pool to improve resource utilization while meeting demand.
[0013] In some possible implementations, the remaining video memory of the first inference node is greater than or equal to the required video memory of the first batch of candidates, and the remaining video memory of the first inference node is the smallest. In other words, the first inference node can include the inference node whose remaining video memory is closest to the required video memory among the nodes whose remaining video memory is greater than or equal to the required video memory. This can reduce resource waste of the inference node and maximize the resource utilization of the inference node.
[0014] In some possible implementations, the recommender system can also update the remaining GPU memory of the first inference node or remove the first inference node from the inference resource pool. Accordingly, for candidates in a second batch out of multiple batches, the recommender system can determine the second inference node from the updated inference resource pool.
[0015] Updating the remaining GPU memory of the first inference node allows it to continue to be used in subsequent scheduling processes, improving its resource utilization. Removing the first inference node from the inference resource pool enables different batches of candidate nodes to be scheduled for parallel decoding, improving parallel performance.
[0016] In some possible implementations, before scheduling multiple candidate options from the candidate set to multiple inference nodes for parallel decoding, the recommender system can also prefill the user's historical behavior sequence and store the prefilled result in a key-value cache. The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options. This method leverages the business characteristic of fixed historical behavior sequences in recommender systems, completing prefilling in advance to further reduce latency.
[0017] In some possible implementations, when user behavior data is detected, the recommendation system can update the user's historical behavior sequence based on the behavior data, and then update the pre-filled results based on the updated historical behavior sequence. This allows for timely updates to the pre-filled results, ensuring the accuracy and efficiency of subsequent decoding.
[0018] In some possible implementations, the key-value cache resides in a distributed memory pool across multiple inference nodes. This distributed memory pool can be shared by multiple inference nodes, which can access it as if it were local memory, thereby improving query performance and consequently increasing the efficiency of decoding based on the key-value cache.
[0019] Secondly, this application provides a reasoning device. The device includes:
[0020] A communication module is used to receive a recommendation reasoning request, the recommendation reasoning request being used to request the reasoning of a target item from the candidate set of the recommendation system, the target item being at least one candidate item in the candidate set;
[0021] The scheduling module is used to schedule multiple candidate options in the candidate set to multiple inference nodes for parallel decoding. Each of the multiple inference nodes is used to decode at least one candidate option among the multiple candidate options according to the recommendation model, and each inference node is used to decode different candidate options among the multiple candidate options. The recommendation model is obtained by training a large model.
[0022] The aggregation module is used to aggregate the decoding results of the multiple inference nodes and determine the target item based on the aggregated decoding results.
[0023] In some possible implementations, the candidate set includes multiple batches of candidates, and the scheduling module is further configured to:
[0024] The required video memory for each batch of candidates in the plurality of batches is determined based on at least one of the batch size, the sequence length of the input sequence, the vector dimension of the candidate, the hidden layer dimension, or the number of layers of the converter in the recommendation model.
[0025] The multiple inference nodes are determined based on the stated requirements.
[0026] In some possible implementations, the scheduling module is specifically used for:
[0027] For the first batch of candidates in the multiple batches, a first inference node is determined from the inference resource pool. The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch.
[0028] In some possible implementations, the remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch, and the remaining video memory of the first inference node is minimized.
[0029] In some possible implementations, the device further includes:
[0030] The resource management module is used to update the remaining video memory of the first inference node or delete the first inference node from the inference resource pool.
[0031] The scheduling module is also used for:
[0032] For the candidates in the second batch among the multiple batches, a second inference node is determined from the updated inference resource pool.
[0033] In some possible implementations, the device further includes:
[0034] The prefill module is used to prefill the user's historical behavior sequence before scheduling multiple candidate options in the candidate set to multiple inference nodes for parallel decoding, and to store the prefill result in a key-value cache. The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options.
[0035] In some possible implementations, the pre-filling module is also used for:
[0036] When the user's behavioral data is detected, the user's historical behavioral sequence is updated based on the behavioral data;
[0037] The pre-filled results are updated based on the updated historical behavior sequence.
[0038] In some possible implementations, the key-value cache resides in a distributed memory pool of the plurality of inference nodes.
[0039] Thirdly, this application provides a recommendation system. The recommendation system includes a management node and an inference node. The management node is used to execute the inference method of the recommendation system as described in the first aspect or any implementation thereof, thereby achieving collaborative recommendation with the inference node.
[0040] Fourthly, this application provides a computing device cluster. The computing device cluster includes at least one computing device, which includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory to cause the computing device or the computing device cluster to perform the reasoning method of the recommendation system as described in the first aspect or any implementation thereof.
[0041] Fifthly, this application provides a computer-readable storage medium storing instructions that instruct a computing device or a cluster of computing devices to execute the reasoning method of the recommendation system described in the first aspect or any implementation thereof.
[0042] In a sixth aspect, this application provides a computer program product containing instructions that, when run on a computing device or a cluster of computing devices, causes the computing device or cluster of computing devices to execute the reasoning method of the recommendation system described in the first aspect or any implementation thereof.
[0043] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0044] To more clearly illustrate the technical methods of this application, the accompanying drawings used will be briefly described below.
[0045] Figure 1 An architecture diagram of a large-model-based recommender system provided for this application;
[0046] Figure 2 A hardware architecture diagram of a recommendation system provided in this application;
[0047] Figure 3 A framework diagram of software deployed on a server provided for this application;
[0048] Figure 4 A flowchart of a reasoning method for a recommender system provided in this application;
[0049] Figure 5 A schematic diagram illustrating the reasoning process of a generative recommender system and the reasoning process of a large model provided in this application;
[0050] Figure 6 This application provides a schematic diagram of constructing a distributed memory pool;
[0051] Figure 7 A schematic diagram illustrating a scenario for a reasoning method in a recommendation system provided in this application;
[0052] Figure 8 A schematic diagram illustrating parallel decoding of multiple batches of candidates provided in this application;
[0053] Figure 9 A schematic diagram of the structure of a reasoning device provided in this application;
[0054] Figure 10 A schematic diagram of the structure of a computing device provided in this application;
[0055] Figure 11 This is a schematic diagram of the structure of a computing device cluster provided in this application. Detailed Implementation
[0056] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0057] First, some technical terms involved in the embodiments of this application will be introduced.
[0058] A recommender system (RS) is used to filter information from massive amounts of data based on a user's ratings or preferences, and then recommend such information to the user. It's worth noting that recommender systems can also provide search functionality; users can enter keywords into a search interface, and the recommender system will recommend information corresponding to those keywords.
[0059] Recommendation systems can recommend content including, but not limited to, audio, video, text, images, or multimodal information. Multimodal information refers to information that includes two or more modalities. For clarity, consider the following example: news is a type of multimodal information combining text and images. Recommendation systems can include music playback applications, video platforms, etc. For instance, a music playback application can infer potentially interesting audio from its music library based on the audio the user has played (songs, instrumental music, etc.) and their saved audio tracks, and then recommend such audio to the user. Similarly, a video platform can infer other videos the user might be interested in based on their video ratings and recommend those videos to the user.
[0060] Generative recommender systems utilize generative artificial intelligence (GAI) to implement recommendation functions. Generative AI can learn and simulate the inherent patterns of things, generating new content with logic and coherence based on user input data. Generative recommender systems can include large-model-based recommender systems. Large-model-based recommender systems are reasoning systems, and the large model can be a large language model (LLM) used for reasoning, including but not limited to various generative models based on transformer structures, such as generative pre-trained transformers (GPT). In some cases, the terms "large model" and "generative model" can be simply referred to as "model."
[0061] Figure 1 An architecture diagram of a large model-based recommender system is shown. The recommender system 10 includes an inference engine 100, an inference service 200, and a model inference component 300. The above modules of the recommender system 10 are described in detail below.
[0062] The inference engine 100 serves as an entry point for recommendation inference requests, forwarding user recommendation inference requests to the inference service 200. In some possible implementations, the inference engine 100 may also integrate a search engine; therefore, the inference engine 100 may include a search inference engine.
[0063] The Inference Service 200 provides deployment and operation capabilities for inference services. It typically includes the following components: inference engine service tools (IE service tools), an inference service client (IE client), service policy management (IE management service, IE MS), an inference server (IE server), and inference backends (IE backends). The inference engine service tools provide performance testing, accuracy testing, and visualization capabilities for large-scale model inference, and support throughput improvements through configuration. The inference service client and its accompanying inference server provide complete inference service capabilities, including communication protocols, request and response interfaces for interfacing with the inference server; these interfaces can be provided to user applications. The service policy management provides operation and maintenance capabilities, including model Pod-level and instance-level management within Pods, service quality monitoring, model updates, fault rescheduling, and load balancing. The inference server provides model inference service capabilities and supports command-line deployment of RESTful services.
[0064] like Figure 1 As shown, the inference server includes an endpoint, a General Model Inference Scheduler (GMIS), and a Backend Manager. The endpoint provides RESTful interfaces, inference service protocols, and interface encapsulations for inference service developers, supporting request interfaces from mainstream inference frameworks. The GMIS provides multi-instance scheduling capabilities, enabling a scalable architecture from inference task scheduling to task execution, adapting to various inference methods. The Backend Manager provides a unified abstract interface for different inference engines and models, facilitating extensibility and reducing modifications required due to changes in inference engines and models.
[0065] Model inference component 300 can be a large model inference component under the inference engine, providing large model inference capabilities. For example, model inference component 300 can include IE LLM. The overall architecture of model inference component 300 can be divided into three layers, including modeling and text generator, and large model manager.
[0066] Modeling offers deeply customized and optimized built-in modules and models. Built-in modules include attention, embedding, column linear, row linear, and multilayer perceptron (MLP) modules, supporting online weight splitting and loading. Built-in models are networked using these modules, supporting tensor splitting and various quantization methods. Users can also customize their models by referring to examples and using built-in modules. After compilation and optimization, the networked model generates an executable graph that accelerates inference on computing cards.
[0067] Text Generator is responsible for model configuration, initialization, loading, autoregressive inference process, post-processing, etc., and provides a unified autoregressive inference interface to LLMManager, supporting parallel decoding and plug-in operation.
[0068] LLM Manager is responsible for state management and task scheduling. It implements user request batching based on scheduling strategies, manages key-value (KV) cache based on a unified memory pool, returns inference results, and provides a state monitoring interface.
[0069] Recommender systems based on large models can perform concurrent inference on a large number of recommendation inference requests. However, large models use autoregressive decoding to complete the inference, resulting in significant inference latency. For example, when the candidate set of a recommender system is 2K (meaning the candidate set includes 2000+ candidates), the inference latency ranges from 3s to 5s, which is significantly different from the required latency of approximately 15ms to 50ms for recommender systems, making it difficult to meet the latency requirements.
[0070] In view of this, this application provides a reasoning method for a recommender system. This method can be applied to recommender systems. The recommender system can be a software system, such as a software system in an e-commerce platform for implementing product recommendations, or a software system in a video platform for implementing video recommendations. The aforementioned software system can be a standalone software system, or it can be integrated into other systems as a plugin, service, etc. In some possible implementations, the recommender system may also include a hardware system; for example, the recommender system may include a cluster of computing devices, which executes the reasoning method of this application when running.
[0071] Specifically, the recommender system receives a recommendation inference request, which requests the inference of a target item to be recommended to the user from the candidate set of the recommender system. The target item is at least one candidate item in the candidate set. The recommender system then schedules multiple candidate items to multiple inference nodes for parallel decoding. Each inference node decodes at least one candidate item according to the recommendation model, and each inference node decodes different candidate items. The recommendation model is obtained by training a large model. The recommender system aggregates the decoding results from the multiple inference nodes and determines the target item based on the aggregated decoding results.
[0072] Based on the business characteristic of fixed historical behavior sequences of users in recommendation systems, this method modifies the autoregressive decoding of large models to parallel decoding. Specifically, multiple candidate options in the candidate set are scheduled to multiple inference nodes. Each inference node decodes at least one of the multiple candidate options, and each inference node decodes different candidate options. This enables parallel inference by multiple inference nodes, significantly reducing inference latency and meeting the latency requirements of recommendation systems.
[0073] To make the technical solution of this application clearer and easier to understand, the architecture of the recommendation system of this application is described below with reference to the accompanying drawings.
[0074] First, see Figure 2 The diagram illustrates a system architecture for a recommendation system 10, which includes a recommendation server 202 and multiple inference servers 204. The recommendation server 202 serves as the entry point for recommendation inference requests, forwarding user recommendation inference requests to the inference servers 204. It should be noted that the recommendation server 202 can use an independent management node to forward user recommendation inference requests to multiple inference servers 204 for parallel decoding to accelerate inference. In some implementations, the recommendation server 202 itself can also act as a management node, forwarding user recommendation inference requests to multiple inference servers 204 for parallel decoding to accelerate inference. Alternatively, the recommendation server 202 can directly forward user recommendation inference requests to the inference servers 204; in this case, the inference server 204 receiving the recommendation inference request acts as a "management node." The "management node" can decide which inference nodes participate in parallel decoding. Inference nodes are nodes used to implement inference functions. For example, when the inference server 204 is a high-performance AI server, the inference nodes can be computing cards such as Neural-network Processing Units (NPUs) or Graphics Processing Units (GPUs) within the inference server 204.
[0075] Figure 2The example illustrates how recommendation server 202 directly forwards user recommendation inference requests to inference server 204. Inference server 204, acting as a management node, schedules multiple inference nodes for parallel decoding to accelerate inference.
[0076] In specific implementation, the management node receives recommendation inference requests, such as those forwarded by recommendation server 202. The recommendation inference request requests the inference of target items to be recommended to the user from the candidate set of recommendation system 10. The target item is at least one candidate item in the candidate set. Candidate items are objects that recommendation system 10 can recommend; for example, candidate items can be products, books, audio, video, news, etc. The management node also schedules multiple candidate items from the candidate set to multiple inference nodes for parallel decoding. Multiple inference nodes can include multiple computing cards in one inference server 204, or multiple computing cards in multiple inference servers 204. The computing card can be an AI accelerator such as an NPU or GPU. Each accelerator can be considered an inference node. Each inference node decodes at least one candidate item from the multiple candidate items based on the recommendation model trained on a large model. Each inference node decodes different candidate items from the multiple candidate items. The management node also summarizes the decoding results from the multiple inference nodes and determines the target item based on the summarized decoding results.
[0077] Then, see Figure 3 The diagram illustrates a framework of software deployed on a server. Firmware 302 and driver 304 can be installed on the inference server 204. Firmware 302 is typically a program written to read-only memory that can directly control and interact with the hardware, and check for any hardware errors. Driver 304 is specifically a small piece of code added to the operating system, containing information about the hardware. When a computer program requests to interact with certain hardware, driver 304 can act as a translator of instructions between the hardware and the program using it. For example, firmware 302 can control and interact with device 24, and check for any errors in device 24; driver 304 can act as a translator of instructions between device 24 and the program using it.
[0078] Furthermore, when the hardware architecture of the inference server 204 adopts a heterogeneous computing architecture (including computing architectures using computing units with different types of instruction sets), a heterogeneous computing framework 306 can also be installed on the inference server 204. In a distributed computing scenario, the heterogeneous computing framework 306 can be a heterogeneous computing framework for neural networks (Compute Architecture for Neural Net, CANN). CANN can support users to quickly build AI applications by providing multi-level programming interfaces. Here, AI applications refer to applications built based on AI models (such as LLM). It should be noted that the heterogeneous computing framework 306 is an optional framework. The inference method of the recommendation system in this application embodiment can be executed even without installing the above framework on the inference server 204. The role of the above framework is to improve the inference efficiency of the recommendation system.
[0079] It should be noted that firmware 302, driver 304 or heterogeneous computing framework 306 can be pre-installed when the inference server 204 leaves the factory, or users can install them themselves according to their business needs.
[0080] Users can install a deep learning framework 308 on the inference server 204. The deep learning framework 308 is used to construct a large-scale computational graph by compiling the methods for implementing the model, and to automatically perform gradient calculations within the computational graph. This computational graph is also called the graph compilation result. Thus, the recommendation system can perform distributed computation based on the graph compilation result, thereby achieving parallel inference. Depending on the compilation method, deep learning frameworks can be divided into frameworks that support static compilation and frameworks that support dynamic compilation. Users can choose to install one or more deep learning frameworks 308 on the inference server 204 according to their business needs. In some embodiments, the deep learning framework 308 may not be installed on the inference server 204; in this case, the inference server 204 can implement the model from scratch using a programming language such as Python.
[0081] Users can also install a software development kit (SDK) 310 for AI accelerators (such as NPUs) on the inference server 204. The SDK 310 provides an extremely easy-to-use application programming interface (API) to accelerate the development of high-performance AI applications.
[0082] For example, software development kit 310 may include a recommendation SDK (denoted as Rec SDK) for search recommendation. The Rec SDK provides a recommendation framework based on AI accelerators (such as NPUs) to support large-scale recommendation scenarios and facilitate efficient training of recommendation models. In this application, developers can efficiently train large models based on the recommendation SDK to obtain recommendation models, and then build inference devices based on the trained recommendation models. Inference nodes can run the inference devices to execute the recommendation inference methods of this application. During runtime, the inference devices can collaborate with the recommendation devices in recommendation server 202 to complete recommendations.
[0083] Furthermore, users can deploy model library 312 on inference server 204. Model library 312 includes AI models implemented using a unified framework, which have standardized parameters and APIs. The AI models include reusable configuration items defined within the unified framework. This reduces the configuration work required for the AI models.
[0084] Users can train a recommendation model based on AI models (such as LLM) in model library 312 and the APIs provided by the Rec SDK in software development kit 310. After completing the above preparations, the recommendation device in recommendation system 10 is used to receive recommendation inference requests. The inference device is used to receive recommendation inference requests, such as receiving recommendation inference requests forwarded by the recommendation device, and to schedule multiple candidate items in the candidate set to multiple inference nodes for parallel decoding. Each inference node is used to decode at least one candidate item from the multiple candidate items according to the recommendation model, and each inference node is used to decode different candidate items from the multiple candidate items. The recommendation model is obtained by training a large model. This large model can be, for example, a model in model library 312. Developers can use the recommendation SDK to train the above large model to obtain the recommendation model. The inference device is also used to summarize the decoding results of multiple inference nodes and determine the target item based on the summarized decoding results.
[0085] based on Figure 2 or Figure 3 The present application also provides a reasoning method for the recommender system 10 shown in the accompanying drawings. The reasoning method for the recommender system provided in this application will be described in detail below with reference to the accompanying drawings.
[0086] See Figure 4 The flowchart shown represents a reasoning method for a recommender system, which includes the following steps:
[0087] S402, Recommendation system 10 receives recommendation reasoning requests.
[0088] A recommendation inference request is used to request the inference of target items to be recommended to the user from the candidate set of the recommendation system. The target item is at least one candidate item in the candidate set. Recommendation inference requests may be generated by a business system in response to a user's page access or search action. In some examples, recommendation inference requests may also be generated by a business system in response to a system task, which may include, but is not limited to, a scheduled recommendation task. The business system may vary depending on the business type. For example, a business system may include a video business system, a news business system, or an e-commerce business system. The candidate set refers to the collection of objects that the recommendation system can recommend; each recommendable object in the candidate set can be considered a candidate item. The candidate items may vary depending on the business type. For example, candidate items may include at least one of candidate products, candidate text, candidate audio, candidate video, or candidate image.
[0089] Specifically, when a user triggers access to the homepage or home page of the business system through a client, or when a user triggers a search operation through a client, the business system can respond to the user's page access or search operation by generating a recommendation inference request. The recommendation system 10 can receive the aforementioned recommendation inference request sent by the business system.
[0090] It should be noted that when a user triggers a search, their needs are clear; therefore, search-triggered recommendations can be considered a form of proactive recommendation. Conversely, when a user triggers a page visit, their needs may be ambiguous; therefore, page visit-triggered recommendations can be considered a form of passive recommendation.
[0091] S404, Recommender system 10 schedules multiple candidate options in the candidate set to multiple inference nodes for parallel decoding.
[0092] Each of the multiple inference nodes is used to decode at least one candidate option from a plurality of candidate options based on the recommendation model, and each inference node is used to decode different candidate options from the plurality of candidate options. The recommendation model is obtained by training a large model. In this application, the large model can include large-scale models built based on a transformer architecture, including but not limited to LLMs such as GPT. The parameter scale of the large model can reach tens of billions, hundreds of billions, or even trillions.
[0093] Before scheduling multiple candidate options from the candidate set to multiple inference nodes for parallel decoding, the recommender system 10 (e.g., the inference server 204 in the recommender system 10) can also prefill the user's historical behavior sequence and store the prefilled result in a key-value cache (KV cache). The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options.
[0094] The pre-filling stage (also known as the initialization stage) and the decoding stage in the large model inference process are explained in detail below.
[0095] The pre-filling stage refers to the process by which the large model processes the tokens in the input prompt. A token, also called a word or sub-word, is the smallest unit of text data in natural language processing. A token can be a word, a portion of a word, a character, or a sub-word. Tokenization is the process of breaking down text into its smallest units for easier processing and analysis by computing devices. Token processing can involve feeding tokens into the large model for forward propagation until the first output token is generated. During forward propagation, the transformer layer of the large model uses a self-attention mechanism to compute the query vector Q, key vector K, and value vector V of the input sequence X through linear transformations. Typically, the query vector Q, key vector K, and value vector V can be obtained by multiplying the input sequence X by the corresponding weight matrix.
[0096] In this system, the query vector helps the large model formulate a question, representing the information the model wants to focus on or acquire. In the self-attention mechanism, the query vector is used to calculate similarity with the key vector to determine the degree of association between different words in the input sequence. The key vector helps the large model find the key content of the question. The key vector is typically multiplied by the query vector to obtain an attention score, which reflects the association between different words in the input sequence. The value vector provides the actual answer or information. Specifically, the large model generates the final output representation by multiplying the attention score by the value vector and performing a weighted sum.
[0097] Decoding: This process generates tokens sequentially using an autoregressive method. Each generated token is added to the input sequence and fed back into the model to generate the next token. The token generation process stops when a specific stopping token (also called an end marker) is generated or when a user-defined termination condition is met (e.g., reaching the maximum sequence length).
[0098] Autoregression is a commonly used generation method, especially in language models such as GPT. Autoregression is primarily used for sequence generation tasks, generating each element in the sequence step-by-step while using previously generated elements as input to predict the next element. Large-scale autoregressive decoding is based on the probabilistic chain rule to generate the next word. For example, a large-scale model can iteratively output the second generated word based on the input sequence and the first generated word, then iteratively output the third generated word based on the input sequence, the first generated word, and the second generated word, and so on. A large-scale model can iteratively output the (i+1)th generated word based on the input sequence, the first generated word, the second generated word, and so on.
[0099] Considering that the user's historical behavior sequence in a recommendation system is fixed, i.e., the input sequence is fixed, and there is no strong dependency between the objects recommended to the user based on the user's historical behavior sequence, the recommendation system can use a parallel decoding method to directly return the decoding results of all candidate options at once. The historical behavior sequence can include a sequence of actions performed by the user on candidate options. Taking an e-commerce business example, for user u, suppose that in a historical time period, the user performed action1 on item1 and action2 on item2. Actions can include, but are not limited to, clicking, adding to favorites, adding to cart, placing an order, and rating, etc. In this example, action1 can be clicking, and action2 can be adding to favorites. The historical behavior sequence can be a sequence formed by products and actions. The action performed on a certain product (such as item1) (such as action1) is used to form a token in the historical behavior sequence.
[0100] See Figure 5 The diagram shows a generative recommendation system reasoning process and a large model reasoning process. After the input sequence is pre-filled, the large model uses an autoregressive decoding method to decode it during reasoning, thereby generating word units one by one. The generative recommendation system schedules multiple candidate options (e.g., m candidate options) to multiple reasoning nodes for parallel decoding.
[0101] Furthermore, the pre-filling and decoding stages involve significant repetitive computation when calculating query vectors, key vectors, and value vectors. Larger models can cache key and value vectors during the pre-filling stage so that the decoding stage can retrieve the key and value vectors of new input terms by querying the cache. The key-value cache can include both key and value caches. In some examples, key and value caches can be merged, for instance, storing key and value vectors in a unified cache.
[0102] Specifically, the recommender system can pre-fill the user's historical behavior sequence and store the pre-filled results in a key-value cache. For example, the recommender system can calculate a key vector and a value vector for each word in the historical behavior sequence, storing the key vector in the key cache and the value vector in the value cache. The key-value cache is used to accelerate the decoding of at least one candidate option from multiple candidate options. Specifically, when the inference node performs decoding, it can query the key-value cache to decode the candidate options. Decoding by querying instead of calculating can significantly improve decoding efficiency.
[0103] User behavior triggered within a historical timeframe remains unchanged; therefore, the sequence of user behavior is fixed, allowing the recommendation system to pre-populate the database. Specifically, the recommendation system can construct a distributed memory pool, for example, a distributed memory pool built around the memory media of a computing device cluster, supporting multiple inference nodes. This distributed memory pool can be shared by multiple inference nodes, which can access it as if it were their local memory. See also... Figure 6 The diagram illustrates a method for constructing a distributed memory pool. The computing device cluster includes multiple AI servers 600. The host side of each AI server 600 includes a CPU 602, dynamic random-access memory (DRAM) 604, and a solid-state disk (SSD) 606. The device side of each AI server 600 includes an NPU 608, which comprises a computing unit 6082 and high-bandwidth memory (HBM) 6084. Inference nodes can be the NPU 608 within the AI server 600. The DRAM 604 and SSD 606 on the host side of the AI server 600 can provide some memory for constructing the distributed memory pool.
[0104] In recommendation scenarios, distributed memory pools are used to store sparse tables or key-value caches. A sparse table is a dynamically programmed data structure that reduces storage space requirements by dividing a large continuous interval into several smaller intervals and storing the maximum and minimum values within each smaller interval. A sparse table can use a two-dimensional array to represent the maximum and minimum values of each smaller interval, where the first dimension represents the starting position of the interval and the second dimension represents the length of the interval. For example, candidate embeddings are often sparse; a sparse table can store these embeddings (also called vectors), accelerating the computation process of the recommendation system. For any given query interval, the recommendation system can quickly retrieve the maximum and minimum values of that interval by calculating its index in the sparse table. A key-value cache is a storage structure that reduces the latency of data access. In large model inference processes, the key vectors and value vectors of the same terms need to be accessed multiple times; therefore, a key-value cache can be used to store these vectors.
[0105] Recommendation systems can initialize a key-value cache in a distributed memory pool. The initial key-value cache can be empty. The system can then calculate key and value vectors based on the user's historical behavior sequence and store them in the cache, without waiting for the user to trigger a page visit or search. Furthermore, when user behavior data is detected, the system can update the user's historical behavior sequence and update the pre-filled results accordingly. For example, the system can calculate key and value vectors for words in new behavior data and update the cache with these new vectors.
[0106] S406. The recommendation system 10 summarizes the decoding results of multiple inference nodes and determines the target item based on the summarized decoding results.
[0107] Specifically, the decoding result can include the probability of the next token. In some examples, this probability can be a logarithmic probability (logits). The recommender system can aggregate the probabilities decoded by multiple inference nodes, and then determine the next token (output token) based on the probabilities of the next token decoded by each inference node. The recommender system 10 can perform inverse tokenization based on the next token to obtain the target item. Specifically, the recommender system 10 can determine multiple possible next tokens based on the probability of the next token, and then determine multiple target items based on the multiple possible next tokens.
[0108] In some possible implementations, the recommender system 10 can determine the next word based on its probability using a set decoding strategy. The decoding strategy indicates how to select words based on the probabilities output by the recommender model. For example, the decoding strategy can include, but is not limited to, greedy decoding, sampling decoding, or beam search. Greedy decoding typically selects the word with the largest logit, sampling decoding treats the logits as a multinomial distribution and randomly samples words from it, and beam search maintains a set of candidate words and expands upon it at each time step by selecting the most likely candidate words.
[0109] Based on the above description, this application provides an inference method for a recommender system. This method, taking into account the fixed historical behavior sequences of users in recommender systems, modifies the autoregressive decoding of a large model into parallel decoding. Specifically, multiple candidate options in the candidate set are scheduled to multiple inference nodes. Each inference node decodes at least one candidate option from the multiple candidate options, and each inference node decodes different candidate options from the multiple candidate options. This achieves parallel inference across multiple inference nodes, significantly reducing inference latency and meeting the latency requirements of recommender systems.
[0110] This application designs an m-Falcon parallel inference method based on the characteristic that users' historical behavior sequences are fixed in recommender systems. Here, m represents the number of candidate options; it should be noted that the number of candidate options in the candidate set can also be referred to as the size of the candidate set. In this method, the candidate set includes multiple batches of candidate options. The recommender system can determine the required GPU memory for each batch of candidate options based on at least one of the following: batch size, input sequence length, candidate option vector dimension, hidden layer dimension, or the number of transformer layers in the recommender model. For example, the recommender system can estimate the required GPU memory for each batch of candidate options based on the modeling method. Then, the recommender system determines multiple inference nodes based on the required GPU memory. The remaining GPU memory (or available GPU memory) of each inference node is greater than or equal to the required GPU memory of the candidate options scheduled to that inference node.
[0111] See Figure 7 The diagram illustrates a scenario of a reasoning method in a recommendation system. Recommendation system 10 includes a recommendation server 202 and multiple inference servers 204. Recommendation server 202 provides the entry point for recommendation inference requests and is responsible for forwarding these requests to the inference servers 204. Specifically, recommendation server 202 can determine the target server from among the multiple inference servers 204 to handle the recommendation inference request based on a load balancing strategy, and then forward the recommendation inference request to the target server.
[0112] Each inference server 204 includes a pre-filling module, a decoding module, and a parallel inference service. Taking a target server among multiple inference servers as an example, the target server can first perform pre-computation based on the user's historical behavior sequence, storing the key vector and value vector in a key-value cache. When it receives a recommendation inference request forwarded by the recommendation server 202, a parallel inference service (such as m-falcon) is launched to process the recommendation inference request.
[0113] The parallel inference service processes recommendation inference requests as follows:
[0114] First, the parallel inference service divides the m candidate options into... One batch;
[0115] Among them, b m Specify the batch size. Parallel inference services can use sequential partitioning or random sampling to divide the m candidate options into... Batch size. For example, the parallel inference service can process batches 1 to b in the candidate set. m Each candidate is assigned to batch 1, and the b-th candidate in the candidate set is... m +1 candidate up to 2b m Each candidate is assigned to batch 2, and so on, until the mb-th batch is assigned to batch 2. m Candidates +1 to m are assigned to batches.
[0116] Then, the parallel inference service models the memory requirements of each batch of candidates to schedule candidates from different batches to different inference nodes. The memory requirements during inference can include the memory needed for model weights, the memory needed for inputs and outputs, and the memory needed for intermediate results. The cache used for intermediate results can include the memory used by the key-value cache (KV Cache).
[0117] Based on this, the parallel inference service can model the memory requirements of each batch of candidates using the following formula:
[0118] ((5H*H+4H+S)*SD+10BSH)*L(1)
[0119] Where B represents the batch size, S represents the sequence length of the input sequence, D represents the embedding size of the candidate vectors, H represents the hidden size, and L represents the number of transformer layers in the recommendation model (e.g., the number of times the transformer is repeated in the recommendation model).
[0120] For each batch of candidates, the parallel inference service can select and schedule inference nodes that meet the memory requirements of the candidates in that batch. The scheduling process is described in detail below.
[0121] In some possible implementations, the parallel inference service can determine inference nodes batch by batch. Specifically, for the candidates in the first batch (e.g., batch i) of multiple batches, the parallel inference service determines the first inference node from the inference resource pool. The first inference node includes one or more inference nodes, and the remaining GPU memory of the first inference node is greater than or equal to the required GPU memory of the candidates in the first batch. To fully utilize the resources of the inference nodes and improve resource utilization, the remaining GPU memory of the first inference node is greater than or equal to the required GPU memory of the candidates in the first batch, and the remaining GPU memory of the first inference node is minimized. In other words, the first inference node can include the inference node whose remaining GPU memory is closest to the required GPU memory among the nodes whose remaining GPU memory is greater than or equal to the required GPU memory.
[0122] Among them, the parallel inference service can determine candidate nodes from the inference resource pool for the first batch of candidates in multiple batches, and then determine the first inference node based on the remaining memory of the candidate nodes.
[0123] For example, the parallel inference service can compare the remaining GPU memory of inference nodes in the inference resource pool with the required GPU memory of the first batch of candidates, and determine candidate nodes from the inference resource pool whose remaining GPU memory is greater than or equal to the required GPU memory of the first batch of candidates. The parallel inference service can then select one inference node from the candidate nodes as the first inference node. To fully utilize the resources of the inference nodes and improve resource utilization, the parallel inference service can select the inference node whose remaining GPU memory is closest to the required GPU memory from the candidate nodes as the first inference node. In other words, the first inference node can be the inference node with the smallest remaining GPU memory among the candidate nodes.
[0124] For example, the parallel inference service can sort the inference nodes in the inference resource pool according to their remaining GPU memory size. Then, based on the GPU memory requirements of the first batch of candidate nodes, it can find the inference node with the closest remaining GPU memory to the required GPU memory from the sorted inference nodes, and determine the first inference node based on this node. Specifically, if the remaining GPU memory of the inference node closest to the required GPU memory is greater than or equal to the required GPU memory, the first inference node can be either the inference node with the closest remaining GPU memory to the required GPU memory, or the inference node whose remaining GPU memory ranks before the closest inference node when sorted from largest to smallest. If the remaining GPU memory of the inference node closest to the required GPU memory is less than the required GPU memory, the first inference node can be the inference node whose remaining GPU memory ranks before the closest inference node when sorted from largest to smallest.
[0125] It should be noted that when inference nodes are sorted in ascending order of remaining video memory, if the remaining video memory of the inference node whose remaining video memory is closest to the required video memory is greater than or equal to the required video memory, the first inference node can be either the inference node whose remaining video memory is closest to the required video memory, or the inference node ranked after the closest inference node; if the remaining video memory of the inference node whose remaining video memory is closest to the required video memory is less than the required video memory, the first inference node can be the inference node ranked after the closest inference node.
[0126] If the remaining GPU memory of all inference nodes in the inference resource pool is insufficient to meet the GPU memory requirements of the first batch of candidates, the parallel inference service can designate multiple inference nodes as the first inference nodes for the first batch of candidates. The sum of the remaining GPU memory of these multiple inference nodes should be greater than or equal to the required GPU memory of the first batch of candidates. Similarly, to fully utilize the resources of the inference nodes, the sum of the remaining GPU memory of the multiple inference nodes should be as close as possible to the required GPU memory to avoid resource waste.
[0127] Furthermore, the parallel inference service can update the remaining GPU memory of the first inference node. For example, the parallel inference service can determine the difference between the remaining GPU memory of the first inference node and the required GPU memory of the first batch of candidates, and update the remaining GPU memory of the first inference node based on this difference. Alternatively, the parallel inference service can remove the first inference node from the inference resource pool. Then, for the candidates in the second batch among multiple batches, the parallel inference service determines the second inference node from the updated inference resource pool. The specific implementation of the parallel inference service determining the second inference node can be found in the description of the parallel inference service determining the first inference node, and will not be repeated here.
[0128] In other possible implementations, the parallel inference service can determine the inference nodes corresponding to multiple batches of candidate options in batches. Specifically, the parallel inference service can construct a planning model based on the memory requirements of multiple batches of candidate options and the remaining memory of inference nodes in the inference resource pool, and configure constraints on the planning model. These constraints may include ensuring that the remaining memory of each inference node is greater than or equal to the memory requirements of the candidate options scheduled to that inference node. The parallel inference service can determine the inference nodes corresponding to multiple batches of candidate options by solving the planning model.
[0129] Next, the decoding modules at different inference nodes read the key vectors and value vectors from the key-value cache layer by layer for each batch of candidates, and then perform decoding based on the key vectors and value vectors. Specifically, the recommendation model may include an L-layer transformer. For the first layer computation, the decoding module can read the key vectors and value vectors required for the first layer computation from the KV Cache. Then, during the first layer computation, it reads the key vectors and value vectors required for the second layer computation in parallel, and so on. The decoding module can read the key vectors and value vectors required for the (i+1)th layer computation in parallel during the i-th layer computation, thereby achieving mutual masking of reading time and computation time, further shortening the inference latency.
[0130] like Figure 8 As shown, the two inference nodes read the embedded representations of multiple candidate options allocated to them from the distributed memory pool, and read the corresponding key vectors and value vectors from the key-value cache for decoding to obtain the decoding result.
[0131] It should be noted that during inference, the recommendation model can be converted into a whole graph using TorchAir. Accordingly, the decoding module can perform inference based on the whole graph, reducing operator scheduling overhead.
[0132] Finally, the target server aggregates the decoding results from multiple inference nodes and determines the target item based on the aggregated decoding results.
[0133] As described above, the inference method of the recommender system in this application uses the m-falcon concept to perform distributed modeling and execution of the candidate set of the recommender system, achieving lightweight and scalable inference. It can quickly adapt to changes in the number of candidates or changes in the users in the recommendation scenario, thus meeting the inference requirements. Furthermore, this method, combined with the business characteristics of the recommendation scenario, achieves low-latency inference through caching.
[0134] Based on the aforementioned reasoning method for recommendation systems, this application also provides a reasoning apparatus. For example... Figure 9 As shown, the inference device 900 includes:
[0135] The communication module 902 is used to receive a recommendation reasoning request, the recommendation reasoning request being used to request the reasoning of a target item from the candidate set of the recommendation system, the target item being at least one candidate item in the candidate set;
[0136] The scheduling module 904 is used to schedule multiple candidate options in the candidate set to multiple inference nodes for parallel decoding. Each of the multiple inference nodes is used to decode at least one candidate option among the multiple candidate options according to the recommendation model, and each inference node is used to decode different candidate options among the multiple candidate options. The recommendation model is obtained by training a large model.
[0137] The aggregation module 906 is used to aggregate the decoding results of the multiple inference nodes and determine the target item based on the aggregated decoding results.
[0138] For example, the communication module 902, scheduling module 904, and aggregation module 906 described above can be implemented in hardware or in software. Specifically, the scheduling module 904 and aggregation module 906 can be... Figure 7 The modules shown are in the parallel inference service.
[0139] When implemented in software, the communication module 902, scheduling module 904, and aggregation module 906 can be applications running on computing devices (e.g., inference server 204), such as computing engines. These applications can also be virtualized and provided to users as virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, or container services. VM services can be services that use virtualization technology to create virtual machine (VM) resource pools on multiple physical hosts, providing VMs to users on demand. BMS services are services that use virtualization technology to create BMS resource pools on multiple physical hosts, providing BMS services to users on demand. Container services are services that use virtualization technology to create container resource pools on multiple physical hosts, providing containers to users on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines, featuring secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0140] When implemented in hardware, the communication module 902 can be a transceiver module such as a transceiver, and the scheduling module 904 and the aggregation module 906 can include at least one computing device (such as a server) or at least one computing core. Alternatively, the scheduling module 904 and the aggregation module 906 can also be devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0141] In some possible implementations, the candidate set includes multiple batches of candidates, and the scheduling module 904 is further configured to:
[0142] The required video memory for each batch of candidates in the plurality of batches is determined based on at least one of the batch size, the sequence length of the input sequence, the vector dimension of the candidate, the hidden layer dimension, or the number of layers of the converter in the recommendation model.
[0143] The multiple inference nodes are determined based on the stated requirements.
[0144] In some possible implementations, the scheduling module 904 is specifically used for:
[0145] For the first batch of candidates in the multiple batches, a first inference node is determined from the inference resource pool. The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch.
[0146] In some possible implementations, the remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch, and the remaining video memory of the first inference node is minimized.
[0147] In some possible implementations, the device 900 further includes:
[0148] Resource management module 908 is used to update the remaining video memory of the first inference node or delete the first inference node from the inference resource pool.
[0149] The scheduling module 904 is also used for:
[0150] For the candidates in the second batch among the multiple batches, a second inference node is determined from the updated inference resource pool.
[0151] Similar to the scheduling module 904, the resource management module 908 can be implemented in hardware or software. When implemented in software, the resource management module 908 can be an application running on a computing device (e.g., inference server 204). This application can also be virtualized and provided to users as a virtualization service such as a BMS, VM, or container. When implemented in hardware, the resource management module 908 can include at least one computing device (such as a server) or at least one computing core. Alternatively, the resource management module 908 can also be a device implemented using an ASIC or a PLD.
[0152] In some possible implementations, the device 900 further includes:
[0153] The prefill module 909 is used to prefill the user's historical behavior sequence before scheduling multiple candidate options in the candidate set to multiple inference nodes for parallel decoding, and to store the prefill result in a key-value cache. The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options.
[0154] Similar to the resource management module 908, the pre-filling module 909 can be implemented in hardware or software. When implemented in software, the pre-filling module 909 can be an application running on a computing device (e.g., inference server 204). This application can also be virtualized and provided to users as a virtualization service such as a BMS, VM, or container. When implemented in hardware, the pre-filling module 909 can include at least one computing device (e.g., a server) or at least one computing core. Alternatively, the pre-filling module 909 can also be a device implemented using an ASIC or a PLD.
[0155] In some possible implementations, the pre-filling module 909 is further configured to:
[0156] When the user's behavioral data is detected, the user's historical behavioral sequence is updated based on the behavioral data;
[0157] The pre-filled results are updated based on the updated historical behavior sequence.
[0158] In some possible implementations, the key-value cache resides in a distributed memory pool of the plurality of inference nodes.
[0159] This application also provides a computing device. The computing device may include the aforementioned inference server 204, used for implementing... Figure 9 The inference device 900 shown here functions as, or executes the inference method of the aforementioned recommendation system. The hardware structure of the computing device is described below with reference to the accompanying drawings.
[0160] Computing devices can be homogeneous servers or heterogeneous servers. Homogeneous servers include servers that use the same type of processor to provide computing power. Heterogeneous servers include servers that use multiple different types of processors to provide computing power, and are typically used in specialized fields. These multiple different types of processors can include processors using different instruction sets. For example, a heterogeneous server can include at least one of the following: Neural-network Processing Unit (NPU), Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), Application-specific Integrated Circuit (ASIC), Complex Programmable Logic Device (CPLD), Field-programmable Gate Array (FPGA), and Central Processing Unit (CPU).
[0161] To accelerate computation and reduce inference latency, computing devices can be heterogeneous servers. For example, a computing device can be an AI server built on CPUs and NPUs, or an AI server built on CPUs and GPUs. This AI server can leverage large models for recommendations.
[0162] The following example, using a heterogeneous server, illustrates the hardware architecture of a computing device.
[0163] The computing device 1000 includes a host 22 and at least one device 24. The host 22 and the at least one device 24 are connected. The device 24 is used to accelerate the computing performance of the host 22.
[0164] The host 22 includes a processor and memory. The processor can be a central processing unit (CPU), and the memory can be a dual in-line memory module (DIMM). Specifically, the DIMM can be of double data rate (DDR) type, such as DDR4 DIMM memory. Figure 2In the example, host 22 includes 4 CPUs and 4 DDR4 DIMM groups, with each CPU connected to one DDR4 DIMM group, and each DDR4 DIMM group including 8 DDR4 DIMMs. Multiple CPUs of host 22 can be connected via hydra interfaces to form a hydra mesh.
[0165] Optionally, host 22 may also include interfaces, such as one or more of the following: Serial Advanced Technology Attachment (SATA) interface, Next Generation Non-Volatile Memoryexpress (NVMe) interface, and Gigabit Ethernet (GE) interface. Host 22 may also include memory. The memory may include SATA-enabled memory or NVMe-enabled memory, such as a SATA-enabled hard disk drive (HDD) or an NVMe-enabled solid-state drive (SSD).
[0166] Device 24 includes a computing card, also known as an accelerator, for accelerating the CPU. Figure 2 In the example, the accelerator included in device 24 may be a neural network processing unit (NPU). In other possible implementations of the embodiments of this application, device 24 may also include more types or more numbers of accelerators, such as GPUs or TPUs.
[0167] This application also provides a computing device cluster. For example... Figure 11 As shown, the computing device cluster includes at least one computing device 1000. The memory of one or more computing devices 1000 in the computing device cluster may store instructions for the same inference unit 900 to execute the inference method of the recommendation system.
[0168] In some possible implementations, one or more computing devices 1000 in the computing device cluster may also store partial instructions for the inference device 900 to execute the inference method of the recommendation system. In other words, a combination of one or more computing devices 1000 can jointly execute the instructions for the inference device 900 to execute the inference method of the recommendation system.
[0169] It should be noted that the memory in different computing devices 1000 in the computing device cluster can store different instructions for executing some functions of the inference device 900.
[0170] This application also provides a recommendation system. The recommendation system includes management nodes and inference nodes. The management node can be an independent node, a recommendation server, or an inference server that receives recommendation inference requests forwarded by the recommendation server. The inference node is used to implement the inference function, including the computing card in the inference server used to implement the inference function, such as an AI accelerator like an NPU or GPU. The management node executes the aforementioned inference method of the recommendation system, thereby collaborating with the inference nodes to achieve the recommendation function.
[0171] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned reasoning method of the recommendation system.
[0172] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute the reasoning method of the recommendation system described above.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. An inference method of a recommendation system, characterized by, The method includes: Receive a recommendation reasoning request, the recommendation reasoning request being used to request the reasoning of a target item from the candidate set of the recommendation system, the target item being at least one candidate item in the candidate set; Multiple candidate options in the candidate set are scheduled to multiple inference nodes for parallel decoding. Each inference node is used to decode at least one candidate option according to the recommendation model, and each inference node is used to decode different candidate options. The recommendation model is obtained by training a large model. The decoding results of the multiple inference nodes are summarized, and the target item is determined based on the summarized decoding results.
2. The method according to claim 1, characterized in that, The candidate set includes multiple batches of candidates, and the method further includes: The required video memory for each batch of candidates in the plurality of batches is determined based on at least one of the batch size, the sequence length of the input sequence, the vector dimension of the candidate, the hidden layer dimension, or the number of layers of the converter in the recommendation model. The multiple inference nodes are determined based on the stated requirements.
3. The method according to claim 2, characterized in that, The step of determining the plurality of inference nodes based on the required memory includes: For the first batch of candidates in the multiple batches, a first inference node is determined from the inference resource pool. The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch.
4. The method according to claim 3, characterized in that, The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch, and the remaining video memory of the first inference node is the smallest.
5. The method according to claim 3 or 4, characterized in that, The method further includes: Update the remaining video memory of the first inference node or delete the first inference node from the inference resource pool; For the candidates in the second batch among the multiple batches, a second inference node is determined from the updated inference resource pool.
6. The method according to any one of claims 1 to 5, characterized in that, Before scheduling multiple candidates in the candidate set to multiple inference nodes for parallel decoding, the method further includes: The user's historical behavior sequence is prefilled, and the prefilled result is stored in a key-value cache. The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options.
7. The method according to claim 6, characterized in that, The method further includes: When the user's behavioral data is detected, the user's historical behavioral sequence is updated based on the behavioral data; The pre-filled results are updated based on the updated historical behavior sequence.
8. The method according to claim 6 or 7, characterized in that, The key-value cache is located in a distributed memory pool of the multiple inference nodes.
9. A reasoning device, characterized in that, The device includes: A communication module is used to receive a recommendation reasoning request, the recommendation reasoning request being used to request the reasoning of a target item from the candidate set of the recommendation system, the target item being at least one candidate item in the candidate set; The scheduling module is used to schedule multiple candidate options in the candidate set to multiple inference nodes for parallel decoding. Each of the multiple inference nodes is used to decode at least one candidate option among the multiple candidate options according to the recommendation model, and each inference node is used to decode different candidate options among the multiple candidate options. The recommendation model is obtained by training a large model. The aggregation module is used to aggregate the decoding results of the multiple inference nodes and determine the target item based on the aggregated decoding results.
10. The apparatus according to claim 9, characterized in that, The candidate set includes multiple batches of candidates, and the scheduling module is further used for: The required video memory for each batch of candidates in the plurality of batches is determined based on at least one of the batch size, the sequence length of the input sequence, the vector dimension of the candidate, the hidden layer dimension, or the number of layers of the converter in the recommendation model. The multiple inference nodes are determined based on the stated requirements.
11. The apparatus according to claim 10, characterized in that, The scheduling module is specifically used for: For the first batch of candidates in the multiple batches, a first inference node is determined from the inference resource pool. The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch.
12. The apparatus according to claim 11, characterized in that, The remaining video memory of the first inference node is greater than or equal to the video memory required by the candidates in the first batch, and the remaining video memory of the first inference node is the smallest.
13. The apparatus according to claim 11 or 12, characterized in that, The device further includes: The resource management module is used to update the remaining video memory of the first inference node or delete the first inference node from the inference resource pool. The scheduling module is also used for: For the candidates in the second batch among the multiple batches, a second inference node is determined from the updated inference resource pool.
14. The apparatus according to any one of claims 9 to 13, characterized in that, The device further includes: The prefill module is used to prefill the user's historical behavior sequence before scheduling multiple candidate options in the candidate set to multiple inference nodes for parallel decoding, and to store the prefill result in a key-value cache. The key-value cache is used to accelerate the decoding of at least one of the multiple candidate options.
15. The apparatus according to claim 14, characterized in that, The pre-filling module is also used for: When the user's behavioral data is detected, the user's historical behavioral sequence is updated based on the behavioral data; The pre-filled results are updated based on the updated historical behavior sequence.
16. The apparatus according to claim 14 or 15, characterized in that, The key-value cache is located in a distributed memory pool of the multiple inference nodes.
17. A recommendation system, characterized in that, The recommendation system includes management nodes and inference nodes, wherein the management nodes are used to execute the inference method of the recommendation system as described in any one of claims 1 to 8.
18. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, the at least one computing device including at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor executes the computer-readable instructions to cause the computing device cluster to perform the reasoning method of the recommendation system as described in any one of claims 1 to 8.
19. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the reasoning method of the recommendation system according to any one of claims 1 to 8.
20. A computer program product, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the reasoning method of the recommendation system according to any one of claims 1 to 8.