Sequential recommendation with diffusion mechanism based on information fusion

CN122734166APending Publication Date: 2026-09-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510287205.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2026-09-11

Smart Images

  • Figure CN122734166A_ABST
    Figure CN122734166A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, apparatuses and computer program products for sequence recommendation with diffusion mechanism based on information fusion. A historical interaction sequence can be received. The historical interaction sequence can include a set of interaction items arranged in an interaction time. A set of interaction item embedding vectors can be obtained from a candidate item embedding vector set. The candidate item embedding vector set can correspond to a candidate item set. The set of interaction item embedding vectors can correspond to the set of interaction items. The set of interaction items can be included in the candidate item set. A first predicted embedding vector can be generated by performing relevance encoding on the set of interaction item embedding vectors. A second predicted embedding vector can be generated by a diffusion mechanism based on fusion of the candidate item embedding vector set and a recommendation item distribution. The first predicted embedding vector can serve as a conditioning signal for the diffusion mechanism. A set of recommendation item embedding vectors matching the second predicted embedding vector can be retrieved from the candidate item embedding vector set. A set of recommendation items corresponding to the set of recommendation item embedding vectors can be obtained from the candidate item set.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] With the development of internet technology and the growth of online information, recommendation services are playing an increasingly important role in many online platforms. Recommendation services aim to provide users with a personalized experience by analyzing data associated with them to recommend information or content that they may be interested in. Recommendation services can be applied in various scenarios, such as e-commerce platforms, video platforms, music platforms, social media platforms, and news platforms. Summary of the Invention

[0002] This invention is provided to introduce a set of concepts, which will be further described in the following detailed description. This invention is not intended to identify key or essential features of the protected subject matter, nor is it intended to limit the scope of the protected subject matter.

[0003] Embodiments of this disclosure provide a method, apparatus, and computer program product for sequence recommendation utilizing a diffusion mechanism based on information fusion. A historical interaction sequence can be received. The historical interaction sequence may include a set of interaction items arranged chronologically. A set of interaction item embedding vectors can be obtained from a set of candidate item embedding vectors. The set of candidate item embedding vectors may correspond to a candidate item set. The set of interaction item embedding vectors may correspond to the set of interaction items. The set of interaction items may be included in the candidate item set. A first predicted embedding vector can be generated by performing relevance encoding on the set of interaction item embedding vectors. A second predicted embedding vector can be generated by fusing the candidate item embedding vector set and the distribution of recommended items through a diffusion mechanism. The first predicted embedding vector can serve as a conditional signal for the diffusion mechanism. A set of recommended item embedding vectors matching the second predicted embedding vector can be retrieved from the set of candidate item embedding vectors. A set of recommended items corresponding to the set of recommended item embedding vectors can be obtained from the candidate item set.

[0004] It should be noted that one or more of the above aspects include the features specifically pointed out in the following detailed description and claims. Certain illustrative features of the one or more aspects are set forth in detail in the following specification and drawings. These features merely indicate various ways in which the principles of each aspect can be implemented, and this disclosure is intended to include all such aspects and their equivalents. Attached Figure Description

[0005] The following description will take into account several aspects disclosed, which are provided to illustrate rather than limit the aspects disclosed.

[0006] Figure 1 An exemplary process for sequence recommendation utilizing an information fusion-based diffusion mechanism according to embodiments of this disclosure is illustrated.

[0007] Figure 2 An exemplary process for generating a second predictive embedding vector using a diffusion mechanism, according to an embodiment of this disclosure, is shown.

[0008] Figure 3A An exemplary process for fusion processing according to an embodiment of this disclosure is shown.

[0009] Figure 3B An exemplary process for fusion processing according to an embodiment of this disclosure is shown.

[0010] Figure 4 An exemplary process for training an encoder and decoder according to embodiments of this disclosure is shown.

[0011] Figure 5 An exemplary process for obtaining a set of recommendations is illustrated according to an embodiment of this disclosure.

[0012] Figure 6 A flowchart is shown of an exemplary method for sequence recommendation utilizing an information fusion-based diffusion mechanism according to an embodiment of this disclosure.

[0013] Figure 7 An exemplary apparatus for sequence recommendation utilizing an information fusion-based diffusion mechanism is shown according to an embodiment of this disclosure.

[0014] Figure 8 An exemplary apparatus for sequence recommendation utilizing an information fusion-based diffusion mechanism is shown according to an embodiment of this disclosure. Detailed Implementation

[0015] This disclosure will now be discussed with reference to various exemplary embodiments. It should be understood that this discussion of embodiments is merely intended to enable those skilled in the art to better understand and thus implement the embodiments of this disclosure, and is not intended to teach any limitation on the scope of this disclosure.

[0016] Sequential recommendation is a widely used recommendation technique in recommendation services. It aims to predict a user's future interactions based on their historical interaction sequences. A historical interaction sequence is an ordered sequence of objects interacted with by the user over a past time period, arranged chronologically. This time period can be of any length, such as hours, days, or months, and can be determined based on various factors such as the application scenario, business needs, and processing resources. In addition to the historical interaction sequence, sequential recommendation typically involves a candidate set, which is a collection of objects that might be recommended to the user. For example, the predicted objects a user might interact with in the future can be selected from this candidate set. In this paper, the objects involved in sequential recommendation are referred to as items. For clarity, objects previously interacted with by the user can be called interaction items, predicted objects a user might interact with in the future can be called recommended items, and objects that might be recommended to the user can be called candidate items. Accordingly, the aforementioned candidate set is referred to as the candidate item set below. Items can have different representations in different application scenarios. For example, in an e-commerce platform, an item can refer to a product; in a video platform, an item can refer to a video; in a music platform, an item can refer to music; in a social media platform, an item can refer to content posted on that platform; in a news platform, an item can refer to news; and so on. In this article, interactions can include various user behaviors, such as browsing, clicking, purchasing, commenting, and saving. In some cases, a single historical interaction sequence can be formed based on a single type of interaction (e.g., a single type of user behavior). For example, in an e-commerce platform, a historical interaction sequence can be an ordered sequence of products a user has purchased; that is, the historical interaction sequence is formed based on purchasing behavior. In some cases, a single historical interaction sequence can be formed based on multiple types of interactions. For example, in a social media platform, a historical interaction sequence can be an ordered sequence of content a user has liked and saved; that is, the historical interaction sequence is formed based on liking and saving behaviors.

[0017] The basic idea behind sequence recommendation is to extract user behavior patterns from historical interaction sequences and infer user interests based on these patterns, thereby predicting items the user might interact with in the future. However, in real life, user behavior patterns can be complex. On one hand, user behavior patterns may contain relatively stable or deterministic components, which could be caused, for example, by user interests or preferences. On the other hand, user behavior patterns may also contain relatively random components, which could be caused, for example, by the influence of various external factors, changes in user interests or preferences, and so on. Therefore, achieving high-quality sequence recommendation based on such complex behavior patterns is very challenging.

[0018] This disclosure proposes a sequential recommendation mechanism utilizing a diffusion mechanism based on information fusion. In this disclosure, sequential recommendation is achieved by considering both relatively deterministic and relatively random components of user behavior patterns. For example, a set of candidate embedding vectors can be used to capture the relatively deterministic components, and a distribution of recommendation items can be used to capture the relatively random components. Then, the candidate embedding vector set and the recommendation item distribution can be fused in a diffusion mechanism to predict recommendations. The term "embedding vector" refers to a vector representation of the original data that carries the feature information and / or inherent relationships of the original data and is suitable for processing and analysis by machine learning models. As previously mentioned, sequential recommendation involves a set of candidate items. In this disclosure, each candidate item in the candidate set can be represented as an embedding vector, which is referred to herein as a candidate embedding vector. Accordingly, the set including candidate embedding vectors can be referred to as a candidate embedding vector set. The candidate set can correspond to the candidate embedding vector set. The candidate embedding vector set can characterize the correlation between different candidate items. For example, the similarity between different candidate embedding vectors can indicate the correlation between corresponding candidate items. The higher the similarity between different candidate embedding vectors, the higher the correlation between the corresponding candidate options; conversely, the lower the similarity between different candidate embedding vectors, the lower the correlation between the corresponding candidate options. The correlation between different candidate options is usually relatively stable or deterministic and can reflect the relatively deterministic part of user behavior patterns. Therefore, the set of candidate embedding vectors can be understood as a representation of deterministic information. The term "distribution" is usually used to describe the probability of a random variable taking different values. In this paper, the recommendation item distribution can characterize the likelihood or probability of different candidate options being recommended. The recommendation item distribution can reflect the relatively random part of user behavior patterns and can therefore be understood as a representation of random information. In this paper, information fusion refers to the fusion of deterministic and random information. A diffusion mechanism is a mechanism for generating prediction samples by simulating the gradual evolution of a noise distribution into a target distribution. Diffusion mechanisms typically include forward diffusion and backward diffusion processes. Forward diffusion is the process of gradually adding noise to the original data to generate a noise distribution, while backward diffusion is the process of gradually removing noise from the noise distribution to recover the original data. The training phase of a diffusion mechanism typically involves forward and backward diffusion processes, while the inference phase usually involves backward diffusion. In this paper, the information fusion-based diffusion mechanism refers to a diffusion mechanism that performs predictions based on the fusion of deterministic and stochastic information.In the embodiments of this disclosure, the diffusion mechanism integrates the candidate embedding vector set with the recommendation distribution, that is, integrates deterministic information with random information, to obtain richer information. Based on this rich information, it is possible to more effectively mine user behavior patterns and infer user interests, thereby providing more reliable prediction results and greatly improving the performance of sequence recommendation.

[0019] In this paper, for a specific recommendation service, the candidate set can include all items involved in the recommendation service, or it can include a subset of items involved in the recommendation service, for example, based on the user's historical interactions. As an example, if the candidate set includes a subset of items, this subset can include items that the user has interacted with (referred to as interaction items), other items that the user has not interacted with but are associated with interaction items, etc. In either case, interaction items are typically included as candidates in the candidate set. Accordingly, a set of embedding vectors corresponding to a set of interaction items included in the historical interaction sequence can be obtained from the candidate embedding vector set. In this paper, the embedding vectors corresponding to interaction items are referred to as interaction item embedding vectors. Relevance encoding can be performed on a set of interaction item embedding vectors to generate a first predicted embedding vector. The first predicted embedding vector can be understood as the items that the user may interact with in the future, predicted based on the relevance of a set of interaction item embedding vectors. Then, a diffusion mechanism can be used to generate a second predicted embedding vector by fusing the candidate embedding vector set with the recommendation item distribution and using the first predicted embedding vector as a conditional signal. After generating the second predicted embedding vector, a set of embedding vectors matching the second predicted embedding vector can be retrieved from the candidate embedding vector set. This set of embedding vectors is referred to herein as a set of recommended item embedding vectors. Then, a set of candidate items corresponding to the set of recommended item embedding vectors can be obtained from the candidate item set as the final set of recommended items. Based on the above process, the technical effects of the embodiments of this disclosure can include: on the one hand, by fusing the candidate embedding vector set with the recommended item distribution, that is, by simultaneously considering deterministic and random information associated with user behavior patterns, richer and more comprehensive information can be provided for the prediction process of the diffusion mechanism; on the other hand, by using the first predicted embedding vector as a conditional signal, a clearer directional guidance can be provided for the prediction process. Therefore, through such a diffusion mechanism, more reliable prediction results can be generated, thereby providing more accurate recommended items, which can significantly improve the performance of sequence recommendation.

[0020] In one implementation, the process from receiving a historical interaction sequence to generating a second predicted embedding vector can be implemented using an encoder-decoder architecture. The encoder receives a historical interaction sequence including a set of interaction items arranged chronologically, obtains a set of interaction item embedding vectors corresponding to the set of interaction items from a set of candidate embedding vectors, and generates a first predicted embedding vector by performing relevance encoding on the set of interaction item embedding vectors. The set of candidate embedding vectors can be generated during the encoder's training phase. Since the encoder performs predictions based on a set of interaction item embedding vectors obtained from the set of candidate embedding vectors, and the set of candidate embedding vectors represents deterministic information, the encoder can also be understood as performing predictions based on deterministic information. The decoder generates a second predicted embedding vector based on the fusion of the set of candidate embedding vectors and the distribution of recommended items through a diffusion mechanism. The first predicted embedding vector can serve as a conditional signal for the diffusion mechanism. Since the distribution of recommended items represents random information, the decoder can be understood as performing predictions based on the fusion of deterministic and random information from the encoder through a diffusion mechanism; that is, the decoder performs predictions using a diffusion mechanism based on information fusion. Accordingly, the decoder can also be viewed as a diffusion model. This architecture, which includes an encoder and a decoder, enables the efficient implementation of embodiments of this disclosure.

[0021] The embodiments disclosed herein can be applied to various application scenarios, such as e-commerce platforms, video platforms, music platforms, social platforms, news platforms, etc.

[0022] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0023] Figure 1 An exemplary process 100 for sequence recommendation utilizing an information fusion-based diffusion mechanism according to embodiments of the present disclosure is illustrated. Process 100 may involve an encoder 110, a decoder 140, and a search model 150. The encoder 110 may perform predictions based on deterministic information. The decoder 140 may perform predictions based on a fusion of deterministic and stochastic information from the encoder via a diffusion mechanism. The search model 150 may obtain recommendations based on the prediction results of the decoder 140.

[0024] like Figure 1As shown, encoder 110 may include an embedding layer 120. Embedding layer 120 can refer to a functional layer used to convert input data into embedding vectors. At embedding layer 120, a historical interaction sequence 101 can be received. Historical interaction sequence 101 may include a set of interaction items arranged according to interaction time. This set of interaction items may be a set of items that the user interacted with within a past time period. The past time period can have any length, which can be set according to various factors such as actual application scenarios, business needs, and processing resources. As an example, the past time period can be several hours in the past, several days in the past, etc. Considering that the number of interaction items within a past time period may not be fixed in different scenarios, in order to maintain the consistency of encoder 110's processing of any historical interaction sequence, the sequence length of the historical interaction sequence can be preset according to various factors (e.g., application scenarios, business needs, encoder processing resources, etc.). The sequence length can be equal to the total number of items in the historical interaction sequence. Here, it is assumed that the sequence length is m, where m is a positive integer greater than 1. Accordingly, historical interaction sequence 101 can have a preset sequence length m. If the number of interaction items in the past time period is equal to m, then these m interaction items can be selected to form the historical interaction sequence 101. If the number of interaction items in the past time period is greater than m, then the last m interaction items can be selected to form the historical interaction sequence 101. If the number of interaction items in the past time period is less than m, then these interaction items can be included in the historical interaction sequence 101, and additional padding items can be added to the historical interaction sequence 101 to make the length of the historical interaction sequence 101 reach m. Padding items can be located, for example, at the beginning or end of the historical interaction sequence 101. Padding items can have preset values, such as 0. Figure 1 In this context, assuming the number of interaction items in the past time period reaches m, the historical interaction sequence 101 may include interaction items 101-1, ..., 101-m. In one implementation, the historical interaction sequence 110 may include the identification (ID) of interaction item 101-1, ..., the identification of interaction item 101-m.

[0025] At the embedding layer 120, a set of interaction item embedding vectors 122 can be obtained from the candidate embedding vector set 103, which includes interaction item embedding vectors corresponding to the interaction items in the historical interaction sequence 101. The candidate embedding vector set 103 can correspond to the candidate set 102, that is, each candidate embedding vector in the candidate embedding vector set 103 can represent a candidate item in the candidate set 102. In one implementation, each candidate embedding vector can represent the ID of a candidate item. As mentioned above, interaction items are included as candidates in the candidate set, therefore, the candidate set 102 can include at least interaction items 101-1, ..., interaction items 101-m, etc. In this way, the interaction item embedding vectors corresponding to interaction items 101-1, ..., interaction items 101-m can be found from the candidate embedding vector set 103, that is, interaction item embedding vectors 122-1, ..., interaction item embedding vectors 122-m. The candidate embedding vector set 103 can also be called a candidate embedding vector lookup table. In this paper, we assume that the total number of candidate options in candidate option set 102 is N, and that the dimension of the candidate option embedding vector is D. Then, candidate option embedding vector set 103 can have a dimension of N×D. D can be a positive integer greater than 1, such as 64, 128, etc. N can be a positive integer greater than 1. Typically, N is much larger than D. For example, in large-scale recommendation services, the value of N may be in the range of thousands, tens of thousands, or even larger. Candidate option embedding vector set 103 can be pre-generated during the training phase of encoder 110 and stored at embedding layer 120.

[0026] Encoder 110 may also include a relevance coding layer 130. The relevance coding layer 130 may refer to a functional layer for performing relevance coding. The relevance coding layer 130 may be implemented using various suitable mechanisms. For example, the relevance coding layer 130 may be implemented using a self-attention mechanism. As an example, the relevance coding layer 130 may be implemented as a Transformer. At the relevance coding layer 130, relevance coding may be performed on a set of interaction item embedding vectors 122 to generate a first predicted embedding vector 132. In one implementation, the relevance between a set of interaction items (i.e., interaction items 101-1, ..., interaction items 101-m) may be extracted based on the set of interaction item embedding vectors 122, and the first predicted embedding vector 132 may be generated based on the relevance between the set of interaction items. The relevance between a set of interaction items may, for example, include the co-occurrence relationship of the set of interaction items in the historical interaction sequence 101, such as co-occurrence, adjacent occurrence, etc., in the historical interaction sequence 101. The first predicted embedding vector 132 may characterize the items that the user may interact with in the future as predicted by the relevance coding layer 130. The first predicted embedding vector 132 can have the same dimension, D, as the candidate embedding vector and the interaction embedding vector.

[0027] Since the encoder 110 obtains a set of interaction term embedding vectors 122 from the candidate term embedding vector set 103 and generates a first prediction embedding vector 132 based on the set of interaction term embedding vectors 122, and the candidate term embedding vector set 103 is a representation of deterministic information, the encoder 110 can be understood as performing prediction based on deterministic information, that is, generating the first prediction embedding vector 132 based on deterministic information.

[0028] The above processing procedure of encoder 110 can be expressed by equation (1):

[0029]

[0030] Where c represents the first predicted embedding vector 132, i1,i2,…,i m These represent interaction items 101-1 to 101-m, It can represent encoder 110, and It can represent the parameters in encoder 110.

[0031] If the historical interaction sequence 101 also includes padding items, the padding items can be converted into padding item embedding vectors at the embedding layer 120. These padding item embedding vectors, along with a set of interaction item embedding vectors, can then be input into the relevance coding layer 130. However, since the padding items do not have any practical meaning, after performing relevance coding on the set of interaction item embedding vectors and padding item embedding vectors at the relevance coding layer 130, the last embedding vector in the relevance coding result corresponding to a non-padding item can be used as the first predicted embedding vector 132.

[0032] In addition to encoder 110, process 100 also involves decoder 140. Decoder 140 can implement a diffusion mechanism based on information fusion. The input of decoder 140 may include candidate embedding vector set 103 and first predicted embedding vector 132. Furthermore, the input of decoder 140 may also include recommendation distribution 104. At decoder 140, a second predicted embedding vector 142 can be generated based on the fusion of candidate embedding vector set 103 and recommendation distribution 104 using the diffusion mechanism. The first predicted embedding vector 132 can serve as a conditional signal for the diffusion mechanism; therefore, the first predicted embedding vector 132 can also be referred to as a conditional embedding vector. The following will refer to... Figure 2 An exemplary process of decoder 140 is described in more detail below.

[0033] Furthermore, process 100 also involves search model 150. Search model 150 may refer to a model capable of performing embedding vector retrieval and outputting recommendations based on the retrieved embedding vectors.

[0034] Search model 150 may include a recommendation embedding vector retrieval layer 160. The recommendation embedding vector retrieval layer 160 may refer to a functional layer for performing recommendation embedding vector retrieval. At the recommendation embedding vector retrieval layer 160, a set of recommendation embedding vectors 162 matching the second predicted embedding vector 142 can be retrieved from the candidate embedding vector set 103. In one implementation, an Approximate Nearest Neighbor (ANN) mechanism can be used to retrieve a set of recommendation embedding vectors 162. The AAN mechanism is a mechanism that can quickly find data points most similar to the query given a query. Its basic idea is to first find several data points similar to the query, and then return the final retrieval results based on the similarity between the query and these data points. In this paper, the ANN mechanism may include various applicable mechanisms, such as the inverted file (IVF) mechanism, the hierarchical navigable small world (HNSW) mechanism, and the disk-based approximate nearest neighbor (DiskANN) mechanism, etc. As an example, the second predicted embedding vector 142 can be used as a query to select a subset of candidate embedding vectors from the candidate embedding vector set 103. The similarity between the second predicted embedding vector 142 and each candidate embedding vector in the subset is calculated. Then, based on the calculated similarity, a set of recommended embedding vectors 162 is selected. The selection of the candidate embedding vector subset can be achieved through various appropriate strategies, depending on the specific ANN mechanism. For example, in the HNSW mechanism, the candidate embedding vector subset is selected based on the graph structure. The similarity between embedding vectors can be determined through various appropriate metrics, such as Euclidean distance, cosine similarity, etc. Since the ANN mechanism does not require calculating similarity over the entire candidate embedding vector set 103, the retrieval of recommended embedding vectors can be achieved efficiently.

[0035] Search model 150 may also include a recommendation mapping layer 170. Recommendation mapping layer 170 may refer to a functional layer used to map retrieved recommendation embedding vectors to recommendations. At recommendation mapping layer 170, a set of candidate items corresponding to a set of recommendation embedding vectors 162 can be obtained from candidate set 102 to form a set of recommendations 172.

[0036] It should be understood that the above is combined Figure 1The process 100 described is merely exemplary. Depending on the specific application requirements, modules or steps in process 100 can be replaced or modified in any way, and process 100 may include more or fewer modules or steps. For example, encoder 110 may include fewer or more functional layers, and search model 150 may also include more or fewer functional layers.

[0037] Figure 2 An exemplary process 200 for generating a second predicted embedding vector using a diffusion mechanism according to an embodiment of the present disclosure is shown. In process 200, a second predicted embedding vector can be generated using a diffusion mechanism based on the fusion of a set of candidate embedding vectors and a distribution of recommended items, wherein a first predicted embedding vector can serve as a conditional signal for the diffusion mechanism.

[0038] like Figure 2 As shown, process 200 may involve decoder 210. Decoder 210 may correspond to Figure 1 The decoder 140 is used in this context. Since the diffusion mechanism is an iterative process based on time steps, the decoder 210 can iteratively perform operations over multiple time steps until a second predicted embedding vector 205 is generated. Each iteration can be represented by a time step. Correspondingly, the number of iterations can correspond to the number of time steps. The number of time steps (i.e., the number of iterations) can be set according to various factors, such as the actual application scenario, business requirements, the processing latency of the decoder 210, and processing resources. As an example, the number of time steps can be preset to a positive integer between 6 and 20. In the following description, an iteration corresponding to an exemplary time step 203 will be used as an example.

[0039] Decoder 210 may include fusion layer 220. Fusion layer 220 may refer to a functional layer for performing fusion processing. At fusion layer 220, fusion processing can be performed on candidate embedding vector set 201 and current recommendation distribution 202 to generate fusion vector 222. Candidate embedding vector set 201 may correspond to... Figure 1 The candidate embedding vector set 103. The current recommendation item distribution 202 can correspond to Figure 1 The recommended item distribution 202 is 104. For the first iteration, the current recommended item distribution 202 can be a preset initial recommended item distribution. The initial recommended item distribution can be various appropriate noise distributions, such as Gaussian noise distribution. For each iteration after the first iteration, the current recommended item distribution 202 can be the recommended item distribution generated by the previous iteration. As mentioned earlier, the candidate embedding vector set represents deterministic information, while the recommended item distribution represents random information. Therefore, through fusion processing, the generated fusion vector contains both deterministic and random information, which helps to improve the accuracy of subsequent predictions.

[0040] In one implementation, the dimension of the current recommendation distribution 202 can be equal to the total number of candidate items in the candidate item set, and the fusion processing at the fusion layer 220 can make the dimension of the fusion vector 222 smaller than the dimension of the current recommendation distribution 202. This allows the subsequent splicing layer 240 and denoising layer 250 to be processed based on a smaller-dimensional vector, effectively reducing processing costs and accelerating processing speed, thereby significantly improving overall recommendation efficiency. For example, the dimension of the current recommendation distribution 202 can be N. By performing fusion processing with the N×D candidate item embedding vector set 201 and the N-dimensional current recommendation distribution 202, the dimension of the generated fusion vector 222 can be D, which is the same as the dimension of the candidate item embedding vector. As mentioned earlier, N is usually a relatively large value, especially in large-scale recommendation scenarios. Since D is much smaller than N, performing subsequent processing based on a D-dimensional fusion vector can greatly save processing resources and improve processing efficiency.

[0041] In one implementation, the fusion process may include a weighted aggregation operation. For example, a weighted aggregation operation may be performed on the candidate embedding vector set 201 based on the current recommendation distribution 202 to generate a fused vector 222. Through weighted aggregation, the candidate embedding vector set and the recommendation distribution can be efficiently fused. The weighted aggregation operation can be implemented in various suitable ways. As an example, the weighted aggregation operation may be implemented based on an attention mechanism. For example, a weighted sum may be performed on the candidate embedding vector set 201 using an attention mechanism to generate a fused vector 222. The current recommendation distribution 202 may correspond to the attention score in the attention mechanism, and the candidate embedding vector set 201 may correspond to the value in the attention mechanism. As another example, a weighted operation may be performed on the candidate embedding vector set 201 based on the current recommendation distribution 202 to generate a weighted vector set. Then, an aggregation operation may be performed on the weighted vector set using a pooling mechanism to generate the fused vector 222.

[0042] Decoder 210 may also include a time-step projector 230. The time-step projector 230 can refer to a functional module for converting time steps into vectors. At the time-step projector 230, a time-step vector 232 corresponding to time step 203 can be generated. For example, the time-step projector 230 can be a sinusoidal projector. Through the time-step vector 232, decoder 210 can clearly know which time step the current recommendation distribution is in, thereby better performing subsequent denoising processing. Furthermore, the time-step vector 232 can also provide a clear rhythmic indication for the decoder's prediction process. In one implementation, the dimension of the time-step vector 232 can be D, which is the same as the dimension of the fusion vector 222 and the candidate embedding vector. This simplifies the complexity of subsequent processing and reduces processing costs.

[0043] The decoder 210 may also include a splicing layer 240. The splicing layer 240 may refer to a functional layer for performing vector splicing processing. At the splicing layer 240, the first predicted embedding vector 204, the time step vector 232, and the fusion vector 222 can be spliced ​​together to generate a spliced ​​vector 242. The first predicted embedding vector 204 may correspond to... Figure 1 The first predicted embedding vector 204, time step vector 232, and fusion vector 222 can be concatenated together according to a preset concatenation order. For example, the concatenation can be performed with the first predicted embedding vector 204 first, the time step vector 232 in the middle, and the fusion vector 222 last. Through this concatenation process, the concatenated vector 242 will contain richer information, thereby promoting the generation of a more accurate second predicted embedding vector. In one implementation, the dimensions of the first predicted embedding vector 204, time step vector 232, and fusion vector 222 are all D, so the dimension of the concatenated vector 242 can be 3×D.

[0044] The decoder 210 may also include a denoising layer 250. The denoising layer 250 can refer to a functional layer for performing denoising processing. At the denoising layer 250, denoising processing can be performed on the concatenated vector 242 to generate an intermediate prediction embedding vector 252. For example, noise in the concatenated vector 242 can be predicted, and the predicted noise can be removed from the concatenated vector 242 to generate the intermediate prediction embedding vector 252. Therefore, the intermediate prediction embedding vector 252 at each time step will contain less noise and will be closer to the desired prediction embedding vector than at the previous time step. The denoising layer 250 can be implemented using various suitable mechanisms, such as Transformer, Convolutional Neural Network (CNN), Multilayer Perceptron (MLP), etc.

[0045] If the preset number of iterations has been reached, the intermediate prediction embedding vector 252 can be output as the second prediction embedding vector 205, and the iteration process can be terminated.

[0046] Decoder 210 may also include a distribution generation layer 260. Distribution generation layer 260 may refer to a functional layer for generating a distribution of recommended items. If a preset number of iterations has not been reached, at distribution generation layer 260, a next recommended item distribution 262 can be generated based on the intermediate predicted embedding vector 252. For example, a multiplication operation can be performed on the intermediate predicted embedding vector 252 and the candidate embedding vector set 201 to generate the next recommended item distribution 262. As an example, the dot product of the intermediate predicted embedding vector 252 and each candidate embedding vector in the candidate embedding vector set 201 can be calculated to obtain a set of dot product results, which can then be converted into a set of probabilities using a probability distribution function to form the next recommended item distribution 262. The next recommended item distribution 262 can continue to be provided to fusion layer 220 to continue performing the next iteration, such as continuing to perform the fusion processing, splicing processing, denoising processing described above, and possibly performing recommended item distribution generation (e.g., if the preset number of iterations has not yet been reached).

[0047] Since decoder 210 generates a fusion vector 222 by performing a fusion process on the candidate embedding vector set 201 and the current recommendation distribution 202, and performs prediction based on the concatenated vector 242 containing the fusion vector 222, and since the candidate embedding vector set 201 represents deterministic information and the current recommendation distribution 202 represents random information, decoder 210 can be understood as performing prediction based on the fusion of deterministic and random information through a diffusion mechanism. In other words, decoder 210 performs prediction using a diffusion mechanism based on information fusion.

[0048] The process by which decoder 210 generates the next recommendation distribution can be expressed by equation (2):

[0049] x t-1 =f θ (x t ,t,c,E) (2)

[0050] Where t represents the time step, x t-1 Let c represent the distribution of the next recommendation for the next time step t-1, c represent the first predicted embedding vector 204, E represent the set of candidate embedding vectors 201, and f represent the distribution of the next recommendation for the next time step t-1. θ Let θ represent decoder 210, and let θ represent the parameters in decoder 210.

[0051] It should be understood that the above is combined Figure 2 The described process 200 is merely exemplary. Depending on the specific application requirements, process 200 can be modified in any way. For example, in different implementations, decoder 210 may include fewer or more functional layers. As an example, the stitching layer 240 and denoising layer 250 can be merged into a single functional layer, allowing stitching and denoising to be implemented by a single layer. As another example, the fusion layer 220 can be split into multiple functional layers, allowing these multiple functional layers to jointly implement the fusion process.

[0052] Figure 3A An exemplary process 300A for fusion processing according to an embodiment of the present disclosure is illustrated. In process 300A, a fusion processing can be performed on a set of candidate embedding vectors and a current recommendation distribution using an attention mechanism to generate a fusion vector. Such fusion processing may include performing a weighted summation operation on the set of candidate embedding vectors using an attention mechanism, wherein the current recommendation distribution may correspond to an attention score in the attention mechanism, and the set of candidate embedding vectors may correspond to a value in the attention mechanism.

[0053] like Figure 3A As shown, process 300A involves fusion layer 310. Fusion layer 310 can correspond to... Figure 2 The fusion layer 220 is included. The fusion layer 310 may include a linear layer 320. The linear layer 320 may refer to a functional layer for performing linear transformations. At the linear layer 320, a linear transformation can be performed on the candidate embedding vector set 301 to generate a linearly transformed vector set 322. The candidate embedding vector set 301 may correspond to... Figure 1 The candidate embedding vector set 103 or Figure 2 The candidate embedding vector set 201 is described. In one implementation, the candidate embedding vector set 103 can be viewed as a matrix, and a linear transformation can be performed on it using a weight matrix. For example, matrix multiplication can be performed between the weight matrix and the candidate embedding vector set 103 to generate a linearly transformed vector set 322. The weight matrix can be set based on various factors, such as actual business needs and application scenarios. For example, the information contained in the candidate embedding vector set 301 can be recombinated or adjusted using the weight matrix.

[0054] The fusion layer 310 may also include a Softmax layer 330. The Softmax layer 330 may refer to a functional layer used to perform normalization processing. At the softmax layer 330, normalization processing can be performed on the current recommendation distribution 302, thereby generating a normalized distribution 332. The current recommendation distribution 302 may correspond to... Figure 2 The current recommended items are distributed as follows: 202.

[0055] The fusion layer 310 may also include a multiplication layer 340. The multiplication layer 340 may refer to a functional layer for performing multiplication operations. At the multiplication layer 340, matrix multiplication can be performed on the normalized distribution 332 and the linearly transformed vector set 322 to generate the fused vector 342.

[0056] The process in fusion layer 310 can be understood as the multiplication operation between attention scores and values ​​in the attention mechanism. The current recommendation distribution 202 can be understood as the attention scores in the attention mechanism, the normalized distribution 332 as the attention weights, and the linearly transformed vector set 322 as the values. The multiplication operation between the normalized distribution 332 and the linearly transformed vector set 322 can be understood as the multiplication operation between attention weights and values.

[0057] For reference Figure 2 As described, the fusion process can be performed iteratively step by step. The process in fusion layer 310 can be expressed by equation (3):

[0058] O t =softmax(x t )×(W v E) (3)

[0059] Where t represents the time step, O t Represents the fusion vector 342, x t This indicates that the current recommended item distribution is 302, W. v Let E represent the weight matrix at linear layer 320, and let E represent the set of candidate embedding vectors 301.

[0060] As mentioned earlier, x t It can be an N-dimensional distribution, and E can be an N×D dimensional matrix. Then, O... t It will be a D-dimensional vector. Since D is much smaller than N, the O-dimensional vector produced by the fusion process... t The dimension is much smaller than x t This allows for a significant reduction in the computational cost of subsequent processing (e.g., stitching, denoising) and an improvement in computational efficiency.

[0061] In one implementation, a subset of elements from the current recommendation distribution 302 can be selected for fusion processing. For example, at the multiplication layer 340, the top K elements from the normalized distribution 332 can be selected to perform matrix multiplication with the K corresponding vectors in the linearly transformed vector set 322. This can more effectively reduce the computational cost of fusion processing and further improve computational efficiency. The top K elements from the normalized distribution 332 could, for example, be the top K elements with the largest values ​​in the normalized distribution 332.

[0062] It should be understood that the above is combined Figure 3A The described process 300A is merely exemplary. Depending on the specific application requirements, process 300A can be modified in any way. For example, in different application scenarios, the fusion layer 310 may include fewer or more functional layers. In one implementation, the fusion layer 310 may not include the linear layer 320; in this case, matrix multiplication can be performed on the normalized distribution and candidate embedding vector set at the multiplication layer 340 to generate the fused vector.

[0063] Figure 3B An exemplary process 300B for fusion processing according to an embodiment of the present disclosure is illustrated. In process 300B, a pooling mechanism can be used to perform fusion processing on a set of candidate embedding vectors and a current recommendation distribution to generate a fused vector. Such fusion processing may include first performing a weighted operation on the set of candidate embedding vectors based on the current recommendation distribution, and then performing an aggregation operation using a pooling mechanism.

[0064] like Figure 3B As shown, process 300B may involve fusion layer 360. Fusion layer 360 may correspond to Figure 2 The fusion layer 220 is included. The fusion layer 360 may include a weighting layer 370. The weighting layer 370 may refer to a functional layer for performing weighting operations. At the weighting layer 370, a weighting operation can be performed on the candidate embedding vector set 351 based on the current recommendation item distribution 352, thereby generating a weighted vector set 372. The candidate embedding vector set 351 may correspond to... Figure 1 The candidate embedding vector set 103 or Figure 2 The set of candidate embedding vectors in the dataset is 201. The current recommendation distribution 352 can correspond to... Figure 2The current recommendation distribution 352 is used as a weight to perform a weighted operation on the candidate embedding vector set 351, thereby generating a weighted vector set 372. For example, each value in the current recommendation distribution 352 can correspond to a candidate embedding vector in the candidate embedding vector set 351, and the value can represent the probability that the candidate corresponding to the candidate embedding vector is a recommendation. Each value in the current recommendation distribution 352 can be multiplied by the corresponding candidate embedding vector to obtain a weighted vector. The weighted vector set 372 can be a set including the obtained weighted vectors. In another implementation, the candidate embedding vector set 351 and the current recommendation distribution 352 can be processed separately using an MLP mechanism to generate an MLP-processed vector set and an MLP-processed distribution. Then, the MLP-processed distribution can be used as a weight to perform a weighted operation on the MLP-processed vector set, thereby generating a weighted vector set 372. Because the MLP mechanism facilitates feature combination and / or further abstraction, the MLP-processed distribution and the MLP-processed vector set can contain richer information. Weighting operations can be implemented using various suitable methods. For example, weighting operations can be implemented using a Lookup function. In this paper, a Lookup function can refer to a function used to establish a correspondence between a set of weights and a set of vectors and multiply the corresponding weights by the vectors. Accordingly, the parameters of the Lookup function can include a set of weights and a set of vectors. In such an implementation, the candidate embedding vector set 351 and the current recommendation distribution 352 can both be directly used as parameters of the Lookup function, or the MLP-processed vector set and the MLP-processed distribution can be used as parameters of the Lookup function.

[0065] The fusion layer 360 may also include a pooling layer 380. The pooling layer 380 may refer to a functional layer used to perform aggregation operations. At the pooling layer 380, a pooling mechanism can be used to perform aggregation operations on the weighted vector set 372, thereby generating the fusion vector 382. For example, the pooling mechanism may include average pooling, weighted pooling, etc. For weighted pooling, the pooling layer 380 may also include a linear layer (in... Figure 3B (Not shown in the image), this linear layer can generate weights for weighted pooling. In one implementation, the pooling mechanism can be implemented as a summation operation.

[0066] For reference Figure 2 As described, the fusion process can be performed iteratively step by step. The process in fusion layer 360 can be expressed by equation (4):

[0067] O t =POOLER(LOOKUP(E T ,xt (4)

[0068] Where t represents the time step, O t Represents the fusion vector 382, ​​x t Let E represent the current distribution of recommended items (352), E represent the set of candidate embedding vectors (351), and POOLER() represent the pooler corresponding to pooling layer 380.

[0069] When the candidate embedding vector set 351 and the current recommendation distribution 352 are processed separately using the MLP mechanism, the process in the fusion layer 360 can be expressed by equation (5):

[0070] O t =POOLER(LOOKUP(f(E)) T ,g(x t (5)

[0071] Where f(E) represents the set of vectors processed by MLP, and g(x) t ) represents the distribution processed by MLP.

[0072] As mentioned earlier, x t It can be an N-dimensional distribution, and E can be an N×D dimensional matrix. Then, O... t It will be a D-dimensional vector. Since D is much smaller than N, the O-dimensional vector produced by the fusion process... t The dimension is much smaller than x t This allows for a significant reduction in the computational cost of subsequent processing (e.g., stitching, denoising) and an improvement in computational efficiency.

[0073] It should be understood that the above is combined Figure 3B The process 300B described is merely exemplary. It can be modified in any way according to actual application requirements. For example, the pooling layer 380 can be replaced with a functional layer based on a self-attention mechanism, thus enabling aggregation operations to be performed on the weighted vector set 372 using the self-attention mechanism.

[0074] Figure 4 An exemplary process 400 for training an encoder and decoder according to an embodiment of this disclosure is illustrated. In process 400, the encoder and decoder may be trained iteratively until they meet their performance requirements. For example, Figure 1 The encoder 110 in the code can correspond to the encoder trained through process 400, while Figure 1 Decoder 140 or Figure 2 The decoder 210 in the process can correspond to the decoder trained by process 400.

[0075] Process 400 may include encoder training at 410. At 410, encoder training may be performed based on training sequence 401 to generate encoder predicted embedding vectors 412. Training sequence 401 may be a sequence of real user interactions. For example, training sequence 401 may include a set of training items arranged according to interaction time. In one implementation, encoder training may include performing the following operations in each iteration: obtaining a set of training embedding vectors corresponding to a set of training items in training sequence 401 from the current candidate embedding vector set 402, and then generating encoder predicted embedding vectors 412 by performing relevance encoding on the set of training embedding vectors. In the first iteration, the current candidate embedding vector set 402 may be a randomly initialized set. In implementations where the encoder includes an embedding layer and a relevance encoding layer (e.g., see [reference]...), the encoder may be trained on a randomized basis. Figure 1 In this process, obtaining a set of training embedding vectors can be performed at the embedding layer, and generating the encoder's predicted embedding vectors can be performed at the relevance coding layer. Encoder training can be performed based on various appropriate mechanisms, such as contrastive learning. Contrastive learning aims to learn representations of data samples such that the representations of similar data samples are close and the representations of dissimilar data samples are far apart. Through contrastive learning, the encoder's ability to represent terms can be enhanced.

[0076] Process 400 may further include encoder loss calculation at 420. At 420, encoder loss 422 may be calculated based on encoder predicted embedding vector 412. In one implementation, encoder loss 422 may be calculated using encoder predicted embedding vector 412 and true recommendation embedding vector. The true recommendation embedding vector may correspond to a true recommendation 403 that was recommended and interacted with by the user after training sequence 401. Encoder loss 422 may be calculated using various appropriate loss functions, such as contrastive learning loss functions.

[0077] Process 400 may also include decoder training at 430. At 430, decoder training can be performed based on the true recommendations 403, the current candidate embedding vector set 402, and the encoder-predicted embedding vector 412. Since the decoder implements a diffusion mechanism, decoder training can include both forward and backward diffusion processes. For ease of description, assume that the number of training time steps for decoder training is Y.

[0078] In the forward diffusion process, a real distribution corresponding to the real recommendation item 403 may be first generated, and then noise adding processing is iteratively performed on the real distribution according to an increasing direction from training time step 1 to training time step Y, so as to generate a target noise distribution. The real distribution can characterize the distribution of the real recommendation item 403 in a candidate item set. For any training time step yforward (1≤yforward≤Y) in the forward diffusion process, a corresponding current noise distribution can be calculated. When yforward=1, that is, at the first iteration of the forward diffusion process, the current noise distribution is obtained after performing noise adding processing on the real distribution; when 1<yforward≤Y, the current noise distribution is obtained after performing noise adding processing on the noise distribution obtained at the training time step yforward-1. The current noise distribution calculated at the training time step Y can be used as the target noise distribution.

[0079] In the reverse diffusion process, denoising processing is iteratively performed based on the target noise distribution according to a decreasing direction from training time step Y to training time step 1, so as to generate a decoder prediction distribution 432. In one implementation, at any training time step y in the reverse diffusion process reverse (1≤y reverse ≤Y), the following operations are performed: generating a training fusion vector by performing fusion processing on a current prediction distribution and a current candidate item embedding vector set 402; generating a training time step vector corresponding to the training time step y reverse ; generating a training concatenation vector by concatenating the training fusion vector, the training time step vector and an encoder prediction embedding vector 412; performing denoising processing on the training concatenation vector to generate a training intermediate prediction embedding vector; generating an intermediate prediction distribution based on the training intermediate prediction embedding vector. When y reverse =Y, that is, at the first iteration of the reverse diffusion process, the current prediction distribution may be the target noise distribution generated through the forward diffusion process; when 1≤y reverse <Y, the current prediction distribution may be the intermediate prediction distribution generated at y reverse +1. The decoder prediction distribution 432 may be the intermediate prediction distribution generated at y reverse . As an example, when y reverse =1, it indicates that the reverse diffusion process has iterated to the training time step 1, and the decoder prediction distribution 432 at this time may be the intermediate prediction distribution generated at the training time step 1. As another example, the reverse diffusion process does not have to iterate all the way to the training time step 1, but can stop at any y reverse (1<y reverse ≤Y), and the decoder prediction distribution 432 at this time may be the intermediate prediction distribution at the training time step y reverseThe intermediate prediction distribution generated at time step Y effectively saves training resources. As another example, the backdiffusion process can consist of only one training time step operation; that is, the decoder prediction distribution 432 can be an intermediate prediction distribution generated at training time step Y.

[0080] In one implementation, as referenced Figure 2 As described, the decoder may include a fusion layer, a time-step projector, a stitching layer, a denoising layer, and a distribution generation layer. Accordingly, the generation of the training fusion vector may be performed at the fusion layer, the generation of the training time-step vector may be performed at the time-step projector, the generation of the training stitching vector may be performed at the stitching layer, the generation of the training intermediate prediction embedding vector may be performed at the denoising layer, and the generation of the intermediate prediction distribution may be performed at the distribution generation layer. Furthermore, the decoder may also include a noisy layer for performing noisy processing during the forward diffusion process.

[0081] Process 400 may also include the calculation of the decoder loss at 440. At 440, the decoder loss 442 can be calculated based on the decoder prediction distribution 432. For example, the decoder loss 442 can be calculated using the decoder prediction distribution 432 and a reference distribution. The reference distribution may vary depending on which training time step the decoder prediction distribution 432 corresponds to. For example, the decoder prediction distribution 432 may be at y reverse The intermediate prediction distribution generated at y reverse When = 1, the reference distribution can be the true distribution described above; when 1 <y reverse When ≤Y, the reference distribution can be the training time step y in the forward diffusion process. forward =y reverse The current noise distribution is calculated at -1. The decoder loss 442 can be calculated using any suitable loss function, such as the KL divergence loss function.

[0082] Process 400 may also include determining at 450 whether the encoder and decoder performance meets the performance requirements. At 450, it can be determined whether the encoder and decoder performance meets the requirements based on encoder loss 422 and decoder loss 442. In one implementation, a joint loss can be calculated based on encoder loss 422 and decoder loss 442. If the joint loss meets the joint loss requirement, it can be determined that the encoder and decoder meet the performance requirements. If the joint loss does not meet the joint loss requirement, it can be determined that the encoder and decoder do not meet the performance requirements. If the encoder and decoder meet the performance requirements, training can end at 470. If the encoder and decoder do not meet the performance requirements, the parameters in the encoder and decoder can be updated at 460, and then training can continue for the next iteration. Updating the parameters in the encoder may include, for example, updating the parameters in equation (1) above. And update the current candidate embedding vector set 402. Updating the parameters in the decoder may include, for example, updating the parameter θ in equation (2) above. The next iteration of training can be performed based on the updated encoder, the updated decoder, and the updated candidate embedding vector set. After training is completed, the encoder that can be used in the inference phase is finally obtained (e.g., Figure 1 The encoder 110 and decoder (e.g., in the code) are shown. Figure 1 Decoder 140 or Figure 2 The decoder 210 in the middle) and the candidate embedding vector set (e.g., corresponding to Figure 1 The candidate embedding vector set 103 Figure 2 The candidate embedding vector set 201 Figure 3A The candidate embedding vector set 301 or Figure 3B The candidate embedding vector set in 351).

[0083] It should be understood that the above text, in combination with... Figure 4 The process 400 described is merely exemplary. Depending on the actual application requirements, the steps in process 400 can be replaced or modified in any way, and process 400 may include more or fewer steps.

[0084] Figure 5 An exemplary process 500 for obtaining a set of recommendations according to an embodiment of this disclosure is illustrated. In process 500, a set of recommendation embedding vectors that match a second predicted embedding vector can be retrieved from a set of candidate embedding vectors using an ANN mechanism, and then a set of recommendations corresponding to the set of recommendation embedding vectors is obtained from the candidate set. The ANN mechanism may include, for example, an IVF mechanism, an HNSW mechanism, a DiskANN mechanism, etc.

[0085] Process 500 may include the selection of a subset of candidate embedding vectors at 510. At 510, a subset 512 of candidate embedding vectors may be selected from the set 502 of candidate embedding vectors based on a second predicted embedding vector 501. The second predicted embedding vector 501 may correspond to... Figure 1 The second predicted embedding vector 142 or Figure 2 The second predicted embedding vector 205 in the dataset. The candidate embedding vector set 502 can correspond to... Figure 1 The candidate embedding vector set 103 Figure 2 The candidate embedding vector set 201 Figure 3A The candidate embedding vector set 301 or Figure 3B The candidate embedding vector set is 351. Selection of a subset of candidate embedding vectors can be achieved using various appropriate strategies, depending on the specific ANN mechanism.

[0086] Process 500 may also include a similarity calculation at 520. At 520, the similarity between the second predicted embedding vector 501 and each candidate embedding vector in the subset of candidate embedding vectors 512 can be calculated to obtain a set of similarities 522. The similarity can be determined by various appropriate metrics, such as Euclidean distance, cosine similarity, etc.

[0087] Process 500 may also include the selection of recommended item embedding vectors at 530. At 530, a set of candidate item embedding vectors can be selected from the subset 512 of candidate item embedding vectors based on a set of similarities 522, as a set of recommended item embedding vectors 532. For example, the subset 512 of candidate item embedding vectors can be sorted in descending order of a set of similarities 522, and a predetermined number of candidate item embedding vectors with the highest ranking can be selected as a set of recommended item embedding vectors 532.

[0088] Process 500 may also include the acquisition of recommendations at 540. At 540, a set of candidate items corresponding to a set of recommendation item embedding vectors 532 can be acquired from the candidate set 503, as a set of recommendations 542. The candidate set 503 may correspond to... Figure 1 The candidate set is 102.

[0089] In steps 510 to 530 above, since similarity is calculated using a subset of candidate embedding vectors instead of the entire set of candidate embedding vectors, computational costs are significantly reduced, retrieval efficiency is greatly improved, and thus recommendation efficiency is effectively enhanced. Therefore, the ANN mechanism is particularly suitable for large-scale recommendation scenarios.

[0090] Process 500, for example, can be achieved through Figure 1 This is implemented using search model 150. Steps 510 to 530 can be performed at the recommendation embedding vector retrieval layer 160 in search model 150, and step 540 can be performed at the recommendation mapping layer 170 in search model 150.

[0091] It should be understood that the above text, in combination with... Figure 5 The described process 500 is merely exemplary. Depending on the specific application requirements, the steps in process 500 can be replaced or modified in any way, and process 500 may include more or fewer steps. For example, in different implementations, step 510 can be omitted. In this case, the similarity between the second predicted embedding vector and each candidate embedding vector in the candidate embedding vector set can be directly calculated, and then a set of recommended embedding vectors can be selected based on the calculated similarity.

[0092] Figure 6A flowchart is shown of an exemplary method 600 for sequence recommendation utilizing an information fusion-based diffusion mechanism according to an embodiment of this disclosure.

[0093] At position 610, a historical interaction sequence can be received. The historical interaction sequence can include a set of interaction items arranged in order of interaction time.

[0094] At position 620, a set of interaction item embedding vectors can be obtained from the candidate embedding vector set. The candidate embedding vector set can correspond to a candidate set. A set of interaction item embedding vectors can correspond to a set of interactions. A set of interactions can be included in the candidate set.

[0095] At position 630, a first predicted embedding vector can be generated by performing relevance encoding on a set of interaction term embedding vectors.

[0096] At position 640, a second predicted embedding vector can be generated through a diffusion mechanism, based on the fusion of the candidate embedding vector set and the recommendation distribution. The first predicted embedding vector can serve as a conditional signal for the diffusion mechanism.

[0097] At position 650, a set of recommended item embedding vectors that match the second predicted embedding vector can be retrieved from the candidate embedding vector set.

[0098] At position 660, a set of recommendations corresponding to a set of recommendation embedding vectors can be obtained from the candidate set.

[0099] In one implementation, generating the first predicted embedding vector may include: generating the first predicted embedding vector by performing relevance encoding on a set of interaction item embedding vectors using a self-attention mechanism.

[0100] In one implementation, each candidate embedding vector in the candidate embedding vector set can represent the identifier of a candidate in the candidate set.

[0101] In one implementation, generating the second predicted embedding vector may include iteratively performing the following operations at each time step of the diffusion mechanism: generating a fused vector by performing a fusion process on the candidate embedding vector set and the current recommendation distribution; and generating an intermediate predicted embedding vector by performing a denoising process based at least on the fused vector. The current recommendation distribution at the first time step may be a pre-defined initial recommendation distribution. The current recommendation distribution at each time step other than the first time step may be generated based on the intermediate predicted embedding vector generated at the previous time step. The intermediate predicted embedding vector generated at the last time step may serve as the second predicted embedding vector.

[0102] The initial recommendation distribution can be a Gaussian noise distribution.

[0103] Generating a fusion vector can include performing a weighted aggregation operation on the set of candidate embedding vectors based on the current distribution of recommended items.

[0104] Weighted aggregation operations can include performing a weighted summation on the candidate embedding vector set using an attention mechanism to generate a fused vector. The current recommendation distribution can correspond to the attention score in the attention mechanism. The candidate embedding vector set can correspond to the value in the attention mechanism.

[0105] The weighted aggregation operation may include: performing a weighted operation on the candidate embedding vector set based on the current recommendation item distribution to generate a weighted vector set; and performing an aggregation operation on the weighted vector set using a pooling mechanism to generate a fused vector.

[0106] Generating an intermediate prediction embedding vector may include: generating a time step vector corresponding to the time step; generating a concatenated vector by concatenating the fusion vector, the time step vector, and the first prediction embedding vector; and generating an intermediate prediction embedding vector by performing denoising processing on the concatenated vector.

[0107] The fusion process can make the dimension of the fusion vector smaller than the dimension of the current recommendation distribution. The dimension of the current recommendation distribution can be equal to the total number of candidate items in the candidate item set.

[0108] In one implementation, obtaining a set of interaction item embedding vectors and generating a first predicted embedding vector can be performed by an encoder, which can be trained using a contrastive learning mechanism.

[0109] The set of candidate embedding vectors can be pre-generated by the encoder during the training phase.

[0110] The generation of the second predicted embedding vector can be performed by a decoder, which can be trained by using KL divergence as a loss function.

[0111] In one implementation, retrieving a set of recommendation embedding vectors that match the second predicted embedding vector may include: using an ANN mechanism to retrieve a set of recommendation embedding vectors that match the second predicted embedding vector from a set of candidate embedding vectors.

[0112] It should be understood that method 600 may also include any steps / processes recommended by the sequence of information fusion-based diffusion mechanisms according to embodiments of the present disclosure as described above.

[0113] Figure 7 An exemplary apparatus 700 for sequence recommendation utilizing an information fusion-based diffusion mechanism is shown according to an embodiment of this disclosure.

[0114] The device 700 may include a historical interaction sequence receiving module 710, an interaction item embedding vector acquisition module 720, a first predicted embedding vector generation module 730, a second predicted embedding vector generation module 740, a recommendation item embedding vector retrieval module 750, and a recommendation item acquisition module 760. The historical interaction sequence receiving module 710 can receive historical interaction sequences. The historical interaction sequence may include a set of interaction items arranged according to interaction time. The interaction item embedding vector acquisition module 720 can acquire a set of interaction item embedding vectors from a set of candidate embedding vectors. The set of candidate embedding vectors may correspond to a set of candidate items. A set of interaction item embedding vectors may correspond to a set of interaction items. A set of interaction items may be included in the set of candidate items. The first predicted embedding vector generation module 730 can generate a first predicted embedding vector by performing relevance encoding on a set of interaction item embedding vectors. The second predicted embedding vector generation module 740 can generate a second predicted embedding vector through a diffusion mechanism, based on the fusion of the set of candidate embedding vectors and the distribution of recommendation items. The first predicted embedding vector can serve as a conditional signal for the diffusion mechanism. The recommendation item embedding vector retrieval module 750 can retrieve a set of recommendation item embedding vectors that match the second predicted embedding vector from the candidate item embedding vector set. The recommendation item acquisition module 760 can acquire a set of recommendations corresponding to a set of recommendation item embedding vectors from the candidate item set.

[0115] Furthermore, the apparatus 700 may also include any other modules configured to perform any operation of a sequence recommendation utilizing an information fusion-based diffusion mechanism according to embodiments of the present disclosure as described above.

[0116] Figure 8 An exemplary apparatus 800 for sequence recommendation utilizing an information fusion-based diffusion mechanism is shown according to an embodiment of the present disclosure.

[0117] The apparatus 800 may include at least one processor 810. The apparatus 800 may also include a memory 820 coupled to the at least one processor 810. The memory 820 may store computer-executable instructions. When executed, the computer-executable instructions cause the at least one processor 810 to perform the following operations: receive a historical interaction sequence, the historical interaction sequence including a set of interaction items arranged chronologically; obtain a set of interaction item embedding vectors from a set of candidate embedding vectors, the set of candidate embedding vectors corresponding to a set of candidate items, the set of interaction item embedding vectors corresponding to a set of interaction items, the set of interaction items being included in the candidate set; generate a first predicted embedding vector by performing relevance encoding on the set of interaction item embedding vectors; generate a second predicted embedding vector by a diffusion mechanism based on the fusion of the set of candidate embedding vectors and the distribution of recommended items, the first predicted embedding vector serving as a conditional signal for the diffusion mechanism; retrieve a set of recommended item embedding vectors from the set of candidate embedding vectors that match the second predicted embedding vector; and obtain a set of recommended items from the set of candidate items that correspond to the set of recommended item embedding vectors.

[0118] Furthermore, at least one processor 810 may also be configured to perform any other operations of the sequence recommendation method utilizing an information fusion-based diffusion mechanism according to embodiments of the present disclosure as described above.

[0119] Embodiments of this disclosure propose a computer program product for sequence recommendation utilizing an information fusion-based diffusion mechanism. The computer program product may include a computer program. The computer program is executed by at least one processor to: receive a historical interaction sequence, the historical interaction sequence including a set of interaction items arranged chronologically; obtain a set of interaction item embedding vectors from a candidate item embedding vector set, the candidate item embedding vector set corresponding to a candidate item set, the set of interaction item embedding vectors corresponding to a set of interaction items, the set of interaction items being included in the candidate item set; generate a first predicted embedding vector by performing relevance encoding on the set of interaction item embedding vectors; generate a second predicted embedding vector by fusing the candidate item embedding vector set and the distribution of recommended items through a diffusion mechanism, the first predicted embedding vector serving as a conditional signal for the diffusion mechanism; retrieve a set of recommended item embedding vectors from the candidate item embedding vector set that match the second predicted embedding vector; and obtain a set of recommended items from the candidate item set that correspond to the set of recommended item embedding vectors. Furthermore, the computer program may also be executed by at least one processor to perform any other operations of the method for sequence recommendation utilizing an information fusion-based diffusion mechanism according to embodiments of this disclosure as described above.

[0120] Embodiments of this disclosure can be implemented on a non-transitory computer-readable medium. The non-transitory computer-readable medium may include instructions. When executed, the instructions may cause one or more processors to perform any step / process of the sequence recommendation method utilizing an information fusion-based diffusion mechanism according to embodiments of this disclosure as described above.

[0121] It should be understood that all operations in the methods described above are merely exemplary, and this disclosure is not limited to any operation in the methods or the order of such operations, but should cover all other equivalent transformations under the same or similar concept.

[0122] Furthermore, unless otherwise specified or clearly indicated from the context that the application is to the singular form, the articles “a” and “an” as used in this specification and the appended claims should generally be interpreted as meaning “one” or “one or more”.

[0123] It should also be understood that all modules in the apparatus described above can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. Furthermore, any of these modules can be further functionally divided into sub-modules or combined together.

[0124] Processors have been described in conjunction with various devices and methods. These processors can be implemented using electronic hardware, computer software, or any combination thereof. Whether these processors are implemented as hardware or software will depend on the specific application and the overall design constraints imposed on the system. As an example, the processors, any portions of processors, or any combinations of processors given in this disclosure can be implemented as microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable processing units configured to perform the various functions described in this disclosure. The functionality of the processors, any portions of processors, or any combinations of processors given in this disclosure can be implemented as software executed by a microprocessor, microcontroller, DSP, or other suitable platform.

[0125] Software should be broadly considered as representing instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, procedures, functions, etc. Software may reside on a computer-readable medium. Computer-readable media may include, for example, memory, which may be, for example, magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical disks, smart cards, flash memory devices, random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, or removable disks. Although memory is shown as separate from the processor in several aspects set forth in this disclosure, memory may also reside within the processor (e.g., in caches or registers).

[0126] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in this disclosure that are known or about to be known to those skilled in the art shall be covered by the claims.

Claims

1. A method for sequence recommendation utilizing a diffusion mechanism based on information fusion, comprising: Receive a historical interaction sequence, the historical interaction sequence comprising a set of interaction items arranged according to interaction time; Obtain a set of interaction item embedding vectors from the candidate embedding vector set, wherein the candidate embedding vector set corresponds to the candidate set, the set of interaction item embedding vectors corresponds to the set of interaction items, and the set of interaction items is included in the candidate set; A first predicted embedding vector is generated by performing relevance encoding on the set of interaction item embedding vectors; A second predicted embedding vector is generated by fusing the candidate embedding vector set and the recommendation distribution through a diffusion mechanism, and the first predicted embedding vector serves as a conditional signal for the diffusion mechanism. Retrieve a set of recommendation embedding vectors from the candidate embedding vector set that match the second predicted embedding vector; as well as Obtain a set of recommendations from the candidate set that corresponds to the set of recommendation embedding vectors.

2. The method of claim 1, wherein, The generation of the first predicted embedding vector includes: The first predicted embedding vector is generated by performing relevance encoding on the set of interaction item embedding vectors using a self-attention mechanism.

3. The method according to claim 1, wherein, Each candidate embedding vector in the candidate embedding vector set represents the identifier of a candidate in the candidate set.

4. The method of claim 1, wherein, The generation of the second predicted embedding vector includes iteratively performing the following operations at each time step of the diffusion mechanism: A fusion vector is generated by performing a fusion process on the candidate embedding vector set and the current recommendation item distribution; and An intermediate prediction embedding vector is generated by performing denoising processing based at least on the fused vector. in, The current distribution of recommended items at the first time step is the preset initial distribution of recommended items. The current recommendation distribution at each time step other than the first time step is generated based on the intermediate prediction embedding vector generated at the previous time step, and The intermediate prediction embedding vector generated at the last time step is used as the second prediction embedding vector.

5. The method of claim 4, wherein, The generated fusion vector includes: The fusion vector is generated by performing a weighted aggregation operation on the candidate embedding vector set based on the current recommendation item distribution.

6. The method of claim 5, wherein, The weighted aggregation operation includes: The candidate embedding vector set is weighted and summed using an attention mechanism to generate the fused vector. Wherein, the current recommendation item distribution corresponds to the attention score in the attention mechanism, and the candidate item embedding vector set corresponds to the value in the attention mechanism.

7. The method of claim 5, wherein, The weighted aggregation operation includes: Based on the current recommendation item distribution, a weighted operation is performed on the candidate item embedding vector set to generate a weighted vector set; and The weighted vector set is aggregated using a pooling mechanism to generate the fused vector.

8. The method of claim 4, wherein, The generation of intermediate prediction embedding vectors includes: Generate the time step vector corresponding to this time step; A concatenated vector is generated by concatenating the fusion vector, the time step vector, and the first prediction embedding vector; and The intermediate prediction embedding vector is generated by performing the denoising process on the concatenated vector.

9. The method according to claim 4, wherein, The fusion process makes the dimension of the fusion vector smaller than the dimension of the current recommendation item distribution, and The dimension of the current recommendation distribution is equal to the total number of candidate items in the candidate item set.

10. The method according to claim 1, wherein, The acquisition of a set of interaction item embedding vectors and the generation of the first predicted embedding vector are performed by an encoder, which is trained using a contrastive learning mechanism.

11. The method according to claim 10, wherein, The set of candidate embedding vectors is pre-generated by the encoder during the training phase.

12. The method according to claim 1, wherein, The generation of the second predicted embedding vector is performed by a decoder, which is trained by using KL divergence as a loss function.

13. The method according to claim 4, wherein, The initial recommendation distribution is a Gaussian noise distribution.

14. The method of claim 1, wherein, The set of recommendation embedding vectors that match the second predicted embedding vector includes: Using the approximate nearest neighbor mechanism, a set of recommended item embedding vectors that match the second predicted embedding vector are retrieved from the set of candidate embedding vectors.

15. An apparatus for sequence recommendation utilizing a diffusion mechanism based on information fusion, comprising: At least one processor; as well as A memory storing computer-executable instructions, which, when executed, cause the at least one processor to: Receive a historical interaction sequence, the historical interaction sequence comprising a set of interaction items arranged according to interaction time; Obtain a set of interaction item embedding vectors from the candidate embedding vector set, wherein the candidate embedding vector set corresponds to the candidate set, the set of interaction item embedding vectors corresponds to the set of interaction items, and the set of interaction items is included in the candidate set; A first predicted embedding vector is generated by performing relevance encoding on the set of interaction item embedding vectors; A second predicted embedding vector is generated by fusing the candidate embedding vector set and the recommendation distribution through a diffusion mechanism, and the first predicted embedding vector serves as a conditional signal for the diffusion mechanism. Retrieve a set of recommendation embedding vectors from the candidate embedding vector set that match the second predicted embedding vector; as well as Obtain a set of recommendations from the candidate set that corresponds to the set of recommendation embedding vectors.

16. The apparatus of claim 15, wherein, The generation of the first predicted embedding vector includes: The first predicted embedding vector is generated by performing relevance encoding on the set of interaction item embedding vectors using a self-attention mechanism.

17. The apparatus of claim 15, wherein, The generation of the second predicted embedding vector includes iteratively performing the following operations at each time step of the diffusion mechanism: A fusion vector is generated by performing a fusion process on the candidate embedding vector set and the current recommendation item distribution; and An intermediate prediction embedding vector is generated by performing denoising processing based at least on the fused vector. in, The current distribution of recommended items at the first time step is the preset initial distribution of recommended items. The current recommendation distribution at each time step other than the first time step is generated based on the intermediate prediction embedding vector generated at the previous time step, and The intermediate prediction embedding vector generated at the last time step is used as the second prediction embedding vector.

18. The apparatus of claim 17, wherein, The generated fusion vector includes: The fusion vector is generated by performing a weighted aggregation operation on the candidate embedding vector set based on the current recommendation item distribution.

19. The apparatus according to claim 17, wherein, The fusion process makes the dimension of the fusion vector smaller than the dimension of the current recommendation item distribution, and The dimension of the current recommendation distribution is equal to the total number of candidate items in the candidate item set.

20. A computer program product utilizing a diffusion mechanism based on information fusion for sequence recommendation, comprising a computer program that is executed by at least one processor for: Receive a historical interaction sequence, the historical interaction sequence comprising a set of interaction items arranged according to interaction time; Obtain a set of interaction item embedding vectors from the candidate embedding vector set, wherein the candidate embedding vector set corresponds to the candidate set, the set of interaction item embedding vectors corresponds to the set of interaction items, and the set of interaction items is included in the candidate set; A first predicted embedding vector is generated by performing relevance encoding on the set of interaction item embedding vectors; A second predicted embedding vector is generated by fusing the candidate embedding vector set and the recommendation distribution through a diffusion mechanism, and the first predicted embedding vector serves as a conditional signal for the diffusion mechanism. Retrieve a set of recommendation embedding vectors from the candidate embedding vector set that match the second predicted embedding vector; as well as Obtain a set of recommendations from the candidate set that corresponds to the embedding vector of the set of recommendations.