Conditional guidance diffusion-based generative recommendation method and device and storage medium

Through a generative recommendation method based on condition-guided diffusion, users' noise data, short-term and long-term interactive project sequences are used to generate targeted user interest vectors, solving the problem that traditional methods cannot achieve guided and targeted user interest generation, and improving the accuracy and diversity of recommendations.

CN119939030APending Publication Date: 2025-05-06TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510045640.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional project recommendation methods cannot achieve guided and targeted user interest vector generation, and it is difficult to capture the complexity and diversity of user interests.

Method used

A generative recommendation method based on condition-guided diffusion is adopted. By obtaining the noise data, short-term interactive project sequences and long-term interactive project sequences of the target user, the diffusion guidance information is determined using a pre-trained coding model, and the candidate project vector corresponding to the multiple candidate recommendation project classifications is generated, and the items to be recommended are finally determined based on similarity.

Benefits of technology

It realizes guiding the diffusion model to the subdistribution space in a specific distribution space to generate targeted user interest vectors, solving the problem that traditional methods cannot achieve guided and targeted user interest generation, and improving the accuracy and diversity of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939030A_ABST
    Figure CN119939030A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a generative recommendation method and device based on conditional guidance diffusion and a storage medium. Noise data, a short-term interaction item sequence, a long-term interaction item sequence and a specific embedding vector of a target user are acquired; using a coding model to determine diffusion guidance information of the target user based on the specific embedded vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data; determining a plurality of candidate item vectors based on the diffusion guidance information; based on the similarity between each candidate item vector and the actual item vector, determining a to-be-recommended item of the target user; on the premise that the specific distribution space of the target user is generated based on the specific embedded vector, the diffusion model is further guided to the specific sub-distribution space to generate the category of the targeted item of the target user, so that the guidance and targeted user interest vector generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a generative recommendation method, device and storage medium based on conditional guided diffusion. Background Art

[0002] With the popularization of the Internet and mobile devices, digital information has penetrated into every aspect of people's daily lives. Faced with the problem of information overload caused by massive data, recommendation systems have been widely used in scenarios such as location services, e-commerce, and online videos due to their advantages in meeting users' personalized needs for data. Among them, recommendation systems are used to recommend items that users may be interested in from a large number of options. The items can be commodities, videos, etc.

[0003] The traditional recommendation system determines recommended items by representing users and items as vectors in a high-dimensional space, and determining the similarity between users and items by calculating the distance or angle between the user vector and the item vector.

[0004] However, in complex recommendation systems, users usually have diverse and changing interests, which may exist in different fields or topics at the same time. The user vector generated by the traditional single vector embedding method can only express one point in the interest space, which is a compromise of multiple interests of the user. This expression method is difficult to make recommendations given a specific interest point, and it is difficult to characterize the complexity of the user's interest distribution. Based on this, some studies have tried to capture the multiple interests of users through single distribution embedding technology. Single distribution embedding technology can model user interests as a potential distribution rather than a single vector. This distribution can capture the user's preference for items in different categories. Multiple user representations can be obtained by sampling from this distribution multiple times, each representing an interest of the user.

[0005] However, although the single distribution embedding technology can express a certain degree of user uncertainty, there is a lack of clear semantic distinction between different sampling results, and it is difficult to correspond to a specific interest of the user, and it is impossible to generate targeted recommendation information. Summary of the invention

[0006] In view of this, the present disclosure proposes a generative recommendation method, device and storage medium based on conditional guided diffusion, which can solve the problem that traditional item recommendation methods cannot achieve guided and targeted user interest vector generation.

[0007] According to one aspect of the present disclosure, a generative recommendation method based on conditional guided diffusion is provided, the method comprising:

[0008] Acquire noise data, a short-term interaction item sequence, and a long-term interaction item sequence of a target user; wherein the short-term interaction item sequence is used to indicate item information interacted by the target user in a first time period; and the long-term interaction item sequence is used to indicate item information interacted by the target user in a second time period; the duration of the first time period is shorter than the duration of the second time period, and the first time period and the second time period are determined based on a current recommendation time;

[0009] Obtaining a pre-trained specific embedding vector of a target user, where the specific embedding vector is used to distinguish interest distributions of different target users;

[0010] Using a pre-trained encoding model, based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data, determining the diffusion guidance information of the target user;

[0011] Determining candidate item vectors corresponding to a plurality of candidate recommendation item categories based on the diffusion guidance information;

[0012] Based on the similarity between each candidate item vector and the actual item vector, the to-be-recommended item for the target user is determined.

[0013] In a possible implementation, the training process of the specific embedding vector and the encoding model includes:

[0014] generating a user ID of the target user, wherein the user IDs of different target users are different;

[0015] Acquire a short-term interaction item sequence sample, a long-term interaction item sequence sample, and a target item vector of a next interaction item of the target user;

[0016] Using the encoding model to be trained, determining an initialized specific embedding vector corresponding to the user identifier;

[0017] Based on the diffusion model to be trained, the target item vector is forward diffused for t steps to obtain a noise data sample at time step t;

[0018] Using the encoding model to be trained, determining an encoding result based on the initialized specific embedding vector, the short-term interaction item sequence sample, the long-term interaction item sequence sample and the noise data sample;

[0019] Based on the similarity between the encoding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, the model parameters of the encoding model to be trained and the model parameters of the diffusion model to be trained are iteratively updated to obtain a trained encoding model, a trained diffusion model and a trained specific embedding vector, the model parameters of the encoding model include the initialized specific embedding vector, the trained specific embedding vector is obtained by iteratively updating the initialized specific embedding vector, and the trained diffusion model is used to perform inverse denoising on the noise data based on the diffusion guidance information to generate the candidate item vector.

[0020] In a possible implementation, the iteratively updating the model parameters of the encoding model to be trained and the model parameters of the diffusion model to be trained based on the similarity between the encoding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, to obtain the trained encoding model, the trained diffusion model and the trained specific embedding vector, includes:

[0021] Obtaining a negative sample item vector of the next interactive item of the target user;

[0022] Determining a user loss based on a first similarity between the initialized specific embedding vector and the target item vector, and a second similarity between the initialized specific embedding vector and the negative sample item vector;

[0023] Determining a recommendation loss based on a third similarity between the encoding result and the target item vector and a fourth similarity between the encoding result and the negative sample item vector;

[0024] Determining a diffusion loss based on a diffusion parameter of the diffusion model to be trained, the encoding result and the target item vector;

[0025] Combining the loss values ​​of the user loss, the recommendation loss, and the diffusion loss, iteratively update the model parameters of the encoding model to be trained and the diffusion parameters of the diffusion model to be trained to maximize the first similarity and the third similarity, minimize the second similarity and the fourth similarity, and minimize the diffusion loss, to obtain a trained encoding model, a trained diffusion model, and a trained specific embedding vector.

[0026] In a possible implementation, using the encoding model to be trained to determine the encoding result based on the initialized specific embedding vector, the short-term interaction item sequence sample, the long-term interaction item sequence sample, and the noise data sample includes:

[0027] Based on the first encoding model to be trained, encoding the initialized specific embedding vector to obtain a long-term interest encoding result;

[0028] Based on the second encoding model to be trained, encoding the short-term interactive item sequence samples to obtain a short-term interest encoding result;

[0029] Determining category distribution data of historical interaction items of the target user based on the long-term interaction item sequence samples;

[0030] Based on the third encoding model to be trained, encoding the category distribution data to obtain the category preference encoding result;

[0031] Unconditionally masking the short-term interest encoding result and the category preference encoding result based on a preset random variable to obtain a converted short-term interest encoding result and a converted category preference encoding result;

[0032] Based on the fusion model to be trained, information fusion is performed on the long-term interest encoding result, the converted short-term interest encoding result, the converted category preference encoding result, the noise data sample, and the embedded representation of time step t to obtain the encoding result.

[0033] In a possible implementation, determining the category distribution data of the historical interaction items of the target user based on the long-term interaction item sequence samples includes:

[0034] The frequency of occurrence of each category in the historical interaction items in the historical interaction is determined to obtain the category distribution data.

[0035] In a possible implementation, using a pre-trained coding model to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence, and the noise data includes:

[0036] Based on a pre-trained first encoding model, encoding the specific embedding vector to obtain a long-term interest feature;

[0037] Based on a pre-trained second encoding model, encoding the short-term interactive item sequence to obtain a short-term interest feature;

[0038] Based on the long-term interaction item sequence, determining the category distribution characteristics of the historical interaction items and the category vectors of the L historical interaction items with the highest interaction frequency of the target user;

[0039] Based on the pre-trained third encoding model, each category vector is encoded to obtain a conditional encoding of each category; and the category distribution feature is encoded to obtain a conditional encoding of the category distribution;

[0040] generating first guidance information based on the long-term interest feature, the short-term interest feature, the conditional encoding of the category distribution, the noise data, and the embedded representation of the time step t;

[0041] Generate second guidance information of the lth category based on the long-term interest feature, the short-term interest feature, the conditional encoding of the lth category, the noise data and the embedded representation of the time step t; wherein l is a positive integer less than or equal to L;

[0042] The diffusion guidance information includes the first guidance information and the second guidance information.

[0043] In a possible implementation, determining the category vectors of the L historical interaction items with the highest interaction frequency of the target user includes:

[0044] One-hot vectors are constructed for the L historical interaction items with the highest interaction frequency to obtain the category vector for each category.

[0045] In a possible implementation, the determining of candidate item vectors corresponding to multiple candidate recommendation item categories based on the diffusion guidance information is expressed by the following formula:

[0046]

[0047] Among them, β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); ∈ represents sampling noise; α t =1-β t ; represents the diffusion guidance information, x t represents the noise data at time step t, x t-1 Represents the noise data at time step t-1; when x is obtained t-1 Then, let t=t-1, and after multiple iterations until t=1, x0 is obtained, where x0 represents the candidate item vector.

[0048] According to another aspect of the present disclosure, a generative recommendation device based on conditional guided diffusion is provided, the device comprising:

[0049] A first acquisition module is used to acquire noise data, a short-term interaction item sequence and a long-term interaction item sequence of a target user; wherein the short-term interaction item sequence is used to indicate item information interacted by the target user in a first time period; and the long-term interaction item sequence is used to indicate item information interacted by the target user in a second time period; the length of the first time period is shorter than the length of the second time period, and the first time period and the second time period are determined based on a current recommendation time;

[0050] A second acquisition module is used to acquire a pre-trained specific embedding vector of a target user, where the specific embedding vector is used to distinguish interest distributions of different target users;

[0051] A data encoding module, configured to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data using a pre-trained encoding model;

[0052] A diffusion guidance module, configured to determine candidate item vectors corresponding to a plurality of candidate recommendation item categories based on the diffusion guidance information;

[0053] The item recommendation module is used to determine the item to be recommended for the target user based on the similarity between each candidate item vector and the actual item vector.

[0054] In a possible implementation manner, the device further includes:

[0055] An identification generating module, used to generate a user identification of the target user, where different target users have different user identifications;

[0056] A sample acquisition module, used to acquire a short-term interaction item sequence sample, a long-term interaction item sequence sample, and a target item vector of a next interaction item of the target user;

[0057] A vector determination module, used to determine the initialized specific embedding vector corresponding to the user identifier using the encoding model to be trained;

[0058] A noise generation module, used for performing t-step forward diffusion on the target item vector based on the diffusion model to be trained to obtain a noise data sample at time step t;

[0059] A sample encoding module, used to use the encoding model to be trained to determine an encoding result based on the initialized specific embedding vector, the short-term interaction item sequence samples, the long-term interaction item sequence samples and the noise data samples;

[0060] A model training module is used to iteratively update the model parameters of the coding model to be trained and the model parameters of the diffusion model to be trained based on the similarity between the coding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, so as to obtain a trained coding model, a trained diffusion model and a trained specific embedding vector, wherein the model parameters of the coding model include the initialized specific embedding vector, the trained specific embedding vector is obtained by iteratively updating the initialized specific embedding vector, and the trained diffusion model is used to perform inverse denoising on the noise data based on the diffusion guidance information to generate the candidate item vector.

[0061] In a possible implementation, the model training module includes:

[0062] A negative sample acquisition unit, used to acquire a negative sample item vector of a next interactive item of the target user;

[0063] A first loss determination unit, configured to determine a user loss based on a first similarity between the initialized specific embedding vector and the target item vector, and a second similarity between the initialized specific embedding vector and the negative sample item vector;

[0064] a second loss determination unit, configured to determine a recommendation loss based on a third similarity between the encoding result and the target item vector, and a fourth similarity between the encoding result and the negative sample item vector;

[0065] a third loss determination unit, configured to determine a diffusion loss based on a diffusion parameter of the diffusion model to be trained, the encoding result and the target item vector;

[0066] A model parameter updating unit is used to iteratively update the model parameters of the encoding model to be trained and the diffusion parameters of the diffusion model to be trained based on the loss values ​​of the user loss, the recommendation loss and the diffusion loss, so as to maximize the first similarity and the third similarity, minimize the second similarity and the fourth similarity, and minimize the diffusion loss, so as to obtain a trained encoding model, a trained diffusion model and a trained specific embedding vector.

[0067] In a possible implementation, the sample encoding module includes:

[0068] A first encoding unit, configured to encode the initialized specific embedding vector based on a first encoding model to be trained to obtain a long-term interest encoding result;

[0069] A second encoding unit is used to encode the short-term interactive item sequence samples based on a second encoding model to be trained to obtain a short-term interest encoding result;

[0070] A distribution determination unit, configured to determine category distribution data of historical interaction items of the target user based on the long-term interaction item sequence samples;

[0071] a third encoding unit, configured to encode the category distribution data based on a third encoding model to be trained to obtain the category preference encoding result;

[0072] An unconditional masking unit, used for unconditionally masking the short-term interest encoding result and the category preference encoding result respectively based on a preset random variable to obtain a converted short-term interest encoding result and a converted category preference encoding result;

[0073] The coding fusion unit is used to perform information fusion on the long-term interest coding result, the converted short-term interest coding result, the converted category preference coding result, the noise data sample, and the embedded representation of the time step t based on the fusion model to be trained to obtain the coding result.

[0074] In a possible implementation manner, the distribution determination unit is specifically configured to:

[0075] The frequency of occurrence of each category in the historical interaction items in the historical interaction is determined to obtain the category distribution data.

[0076] In a possible implementation, the data encoding module includes:

[0077] a fourth encoding unit, configured to encode the specific embedding vector based on a pre-trained first encoding model to obtain a long-term interest feature;

[0078] a fifth encoding unit, configured to encode the short-term interaction item sequence based on a pre-trained second encoding model to obtain a short-term interest feature;

[0079] A category determination unit, configured to determine, based on the long-term interaction item sequence, category distribution characteristics of historical interaction items and category vectors of L historical interaction items with the highest interaction frequency of the target user;

[0080] a sixth encoding unit, configured to encode each category vector based on a pre-trained third encoding model to obtain a conditional encoding of each category; and to encode the category distribution feature to obtain a conditional encoding of the category distribution;

[0081] A first generating unit, configured to generate first guidance information based on the long-term interest feature, the short-term interest feature, the conditional encoding of the category distribution, the noise data, and the embedded representation of the time step t;

[0082] A second generating unit, configured to generate second guidance information of the lth category based on the long-term interest feature, the short-term interest feature, the conditional encoding of the lth category, the noise data and the embedded representation of the time step t; wherein l is a positive integer less than or equal to L;

[0083] The diffusion guidance information includes the first guidance information and the second guidance information.

[0084] In a possible implementation manner, the category determination unit is specifically configured to:

[0085] One-hot vectors are constructed for the L historical interaction items with the highest interaction frequency to obtain the category vector for each category.

[0086] In a possible implementation, the diffusion guidance process of the diffusion guidance module is expressed by the following formula:

[0087]

[0088] Among them, β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); ∈ represents sampling noise; α t =1-β t ; represents the diffusion guidance information, x t represents the noise data at time step t, x t-1 Represents the noise data at time step t-1; when x is obtained t-1 Then, let t=t-1, and after multiple iterations until t=1, x0 is obtained, where x0 represents the candidate item vector.

[0089] According to another aspect of the present disclosure, a generative recommendation device based on conditional guided diffusion is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0090] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0091] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0092] By obtaining noise data, short-term interaction item sequences and long-term interaction item sequences of target users; obtaining pre-trained specific embedding vectors of target users; using the pre-trained encoding model, based on the specific embedding vectors, short-term interaction item sequences, long-term interaction item sequences and noise data, determining the diffusion guidance information of the target user; based on the diffusion guidance information, determining the candidate item vectors corresponding to multiple candidate recommendation item categories; based on the similarity between each candidate item vector and the actual item vector, determining the target user's to-be-recommended items; on the premise of the specific distribution space of the target user generated based on the specific embedding vector, the diffusion model can be further guided to a specific sub-distribution space to generate categories of targeted items for the target user, thereby solving the problem that traditional item recommendation methods cannot achieve the generation of guided and targeted user interest vectors, and achieving the generation of guided and targeted user interest vectors.

[0093] At the same time, since the user's long-term interests, short-term interests and category preferences are comprehensively considered to achieve personalized encoding, and the user distribution is learned based on the conditional guided diffusion model, it is possible to achieve guideable user interest generation for more accurate and diverse recommendations.

[0094] In addition, by encoding the category vectors of the L historical interaction items with the highest interaction frequency of the target user, the conditional coding of each category is obtained, and then the candidate item vector corresponding to each category is generated by reverse diffusion. The target user can be provided with recommended items corresponding to different categories, thereby realizing selective user interest vector generation.

[0095] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0097] Figure 1 A flowchart showing a generative recommendation method based on conditional guided diffusion according to an embodiment of the present disclosure is shown;

[0098] Figure 2 A flowchart showing a process of determining diffusion guidance information of a target user according to an embodiment of the present disclosure;

[0099] Figure 3 A flowchart showing a model training method according to an embodiment of the present disclosure is shown;

[0100] Figure 4 A flowchart showing a process of determining an encoding result according to an embodiment of the present disclosure is shown;

[0101] Figure 5 A flowchart showing a model updating process according to an embodiment of the present disclosure is shown;

[0102] Figure 6 A schematic diagram showing a generative recommendation process based on conditional guided diffusion according to an embodiment of the present disclosure;

[0103] Figure 7 A block diagram of a generative recommendation device based on conditional guided diffusion according to an embodiment of the present disclosure is shown;

[0104] Figure 8 FIG. 4 is a block diagram of a generative recommendation device based on conditional guided diffusion according to another embodiment of the present disclosure. DETAILED DESCRIPTION

[0105] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0106] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0107] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0108] First, several terms involved in this application are introduced.

[0109] Diffusion Models: A type of deep learning method based on probabilistic generative models. Diffusion models gradually transform data into noise by simulating the physical diffusion process, and then learn the reverse process to gradually restore the original data from the noise to achieve high-quality generation effects.

[0110] The core idea of ​​the diffusion model includes two main processes:

[0111] 1. Forward Diffusion Process: gradually add noise to the data to transform it into pure noise. This process is usually a Markov chain, which is implemented by adding a small amount of Gaussian noise at each time step.

[0112] 2. Reverse Generation Process: Learn to gradually remove noise from the noise and restore the original data. This process is also a Markov chain, but in the opposite direction. Starting from standard Gaussian noise, through the learned reverse diffusion process, the noise is gradually removed to generate new data that is basically consistent with the real data distribution.

[0113] Since the traditional modeling method based on single distribution embedding technology can only express an area in the interest space, different vectors can be sampled from this area according to the distribution function learned by the model for recommendation. However, it is still impossible to achieve guided, selective and targeted user interest vector generation.

[0114] Based on this, in order to better solve the problem of multi-interest modeling, this application proposes an embedding technology based on multiple distributions. By learning a specific distribution for each interest of each user, multiple interests can be expressed simultaneously. This not only maintains the advantages of distribution-based embedding methods in uncertainty modeling, but also ensures that different interests are clearly distinguished. Through multi-distribution embedding technology, the recommendation system can capture the diverse interests of users more finely, while enriching the recommended content, improving the recommendation quality and user satisfaction.

[0115] The following is a detailed introduction to the generative recommendation method based on conditional guided diffusion provided by this application.

[0116] Figure 1 The flowchart of the generative recommendation method based on conditional guided diffusion according to an embodiment of the present disclosure is shown. This embodiment takes the method used in an electronic device with processing capabilities such as a user terminal or a server as an example for explanation. The user terminal can be a computer, a mobile phone, a tablet computer, etc. This embodiment does not limit the type of electronic device. Figure 1 As shown, the method includes:

[0117] Step 101: Obtain noise data, short-term interaction item sequences, and long-term interaction item sequences of a target user.

[0118] Since the electronic device does not determine the next item that the target user needs to interact with during initialization, the next item is initialized as a randomly distributed noise data. The noise data is used to reflect the randomness of the next item that the target user needs to interact with.

[0119] Exemplarily, obtaining noise data of a target user includes: The noise data is obtained by sampling in . Where I represents the unit matrix.

[0120] In this application, the project includes resources obtained by the target user through the network, which can be product information, video, audio, text content, etc. This embodiment does not limit the implementation method of the project. Generally, the project has one or at least two categories, and the target user will interact with the project of his interest. The interaction method includes but is not limited to: clicking, or browsing time longer than the preset time, etc. This embodiment does not limit the way the target user interacts with the project. Each project is divided into one or more categories, and each category includes multiple project contents, for example: the project includes: movie category, life diary category, novel recommendation category, etc.; for the movie category, it includes the video content of movie 1 and the video content of movie 2; for the life diary category, it includes at least one video diary posted by different users, and the video diary posted by a user may also be divided into the travel category; for the novel recommendation category, it includes at least one video introduction of the novel, etc.

[0121] The short-term interaction item sequence is used to indicate the item information interacted by the target user within the first time period. The first time period is determined based on the current recommendation moment. The current recommendation moment refers to the moment when the next item needs to be recommended to the target user. Exemplarily, the first time period is the time period of the first duration closest to the current recommendation moment, for example: the first time period is the time period within the first duration before the current recommendation moment, and the first duration can be 10 minutes, 2 hours, etc.; for another example: the first time period is the time period of the M items interacted by the target user before the current recommendation moment. This embodiment does not limit the method for determining the first time period.

[0122] Exemplarily, the project information of each project indicated by the short-term interactive project sequence includes: the interactive time of the project, and the category of the project. In other implementations, the project information may also include other content, such as: the release time of the project, the name of the project, the release end information, etc. This embodiment does not limit the content of the project information.

[0123] Exemplarily, the short-term interactive project sequence is obtained by transforming the project information of each project interacted by the target user in the first time period using an embedding model. An embedding model refers to a mathematical model that maps data (such as text, images, etc.) into a continuous vector space. The converted vector captures the characteristics and semantic information of the data, so that similar data points are close to each other in the vector space, while dissimilar data points are far apart. Optionally, the embedding model can be a word embedding (Word2Ve) model, a global vector (GloVe) model, a bidirectional encoder representation (Bidirectional Encoder Representations from Transformers, BERT) model, etc. This embodiment does not limit the implementation method of the embedding model.

[0124] Optionally, the short-term interaction item sequence may be sent by other devices, or generated by the electronic device based on the interaction operation of the target user. This embodiment does not limit the generation method of the short-term interaction item sequence.

[0125] Taking the example of an electronic device generating a short-term interaction item sequence based on the interaction operation of the target user, for the target user u, each item indicated by the interaction operation of the target user u is arranged according to the time sequence information, and a short-term interaction item sequence is generated based on the item information of the M items that have been interacted before the current recommendation moment (i.e., the M items that have been interacted most recently). in, is the embedding representation of the mth item in the sequence, which is obtained by transforming the item information of the mth item through the embedding model. Where m is a positive integer less than or equal to M, and M is a positive integer. d represents the dimension of each embedding representation in the short-term interaction item sequence.

[0126] The long-term interaction item sequence is used to indicate the item information that the target user interacts with in the second time period, wherein the duration of the first time period is less than the duration of the second time period, and the second time period is determined based on the current recommendation time. Exemplarily, the second time period is the time period of the second time period closest to the current recommendation time, for example: the second time period is the time period within the second time period before the current recommendation time, and the second time period is greater than the first time period, such as 1 month, half a year, etc.; for another example: the second time period is the multiple items that the target user interacted with before the current recommendation time. The present embodiment does not limit the method for determining the second time period. in, These are all items that the target user can interact with.

[0127] The project information of each project indicated by the long-term interactive project sequence includes: the interactive time of the project and the category of the project. In other implementations, the project information may also include other content, such as: the release time of the project, the name of the project, the release end information, etc. This embodiment does not limit the content of the project information.

[0128] In one example, the short-term interaction item sequence is an interaction item sequence in the first time period closest to the current recommendation time in the long-term interaction item sequence.

[0129] Optionally, the long-term interaction item sequence may be sent by other devices, or generated by the electronic device based on the interaction operation of the target user. This embodiment does not limit the generation method of the long-term interaction item sequence.

[0130] Step 102, obtaining a pre-trained specific embedding vector of the target user.

[0131] The specific embedding vector is used to distinguish the interest distribution of different target users. In this embodiment, by obtaining the specific embedding vector of the target user, specific information of each target user, or conditional information for the target user, can be generated for the diffusion process to achieve personalized distribution characterization across users.

[0132] In this embodiment, different target users correspond to different user identifiers, each user identifier corresponds to a specific embedding vector, and the specific embedding vector corresponding to the target user can be determined based on the user identifier of the target user.

[0133] The user identification may be a randomly generated string, or may be the identification card number of the target user, etc. This embodiment does not limit the implementation method of the user identification.

[0134] Exemplarily, different target users correspond to different specific embedding vectors, thereby distinguishing the long-term interest distribution of different target users. The training process of the specific embedding vector is detailed below, and this embodiment will not be described in detail here.

[0135] Optionally, step 102 may be executed after step 101, or may be executed before step 101, or may be executed synchronously with step 101. This embodiment does not limit the execution order between steps 102 and 101.

[0136] Step 103, using a pre-trained encoding model, based on a specific embedding vector, a short-term interaction item sequence, a long-term interaction item sequence and noise data, determines diffusion guidance information of the target user.

[0137] In one example, a specific embedding vector, a short-term interaction item sequence, and a long-term interaction item sequence correspond to different encoding models, respectively, to obtain diffusion guidance information.

[0138] Accordingly, refer to Figure 2 ,Using the pre-trained encoding model, based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data, the diffusion guidance information of the target user is determined, including the following steps:

[0139] Step 1031, based on the pre-trained first encoding model, encode the specific embedding vector to obtain the long-term interest feature.

[0140] Exemplarily, the first coding model is established based on a multilayer perceptron (MLP) model. In other implementations, the first coding model can also be established based on neural network models such as a convolutional neural network (CNN) or a recurrent neural network (RNN). This embodiment does not limit the implementation method of the first coding model.

[0141] Based on a pre-trained first encoding model, encoding a specific embedding vector includes: inputting the specific embedding vector into the pre-trained first encoding model to obtain a long-term interest feature.

[0142] Assume that a particular embedding vector is represented as The first encoding model is MLP, then the long-term interest feature c l It can be expressed by the following formula:

[0143] c l =MLP(h u )

[0144] in, d represents the dimension of a specific embedding vector and long-term feature of interest.

[0145] The training process of the first coding model is detailed below and will not be elaborated in detail in this embodiment.

[0146] Step 1032: Encode the short-term interaction item sequence based on the pre-trained second encoding model to obtain short-term interest features.

[0147] In this embodiment, by encoding the short-term interaction project sequence and the long-term interaction project sequence, a personalized distribution characterization for the same target user can be achieved, that is, given the specific distribution space of the target user (i.e., the long-term interest characteristics), the diffusion model is further guided to a specific sub-distribution space, thereby determining the items to be recommended. Exemplarily, the sub-distribution space includes a sub-distribution space indicating short-term interests, and a sub-distribution space indicating the category preference of the target user. Among them, the sub-distribution space indicating short-term interests is modeled by the second encoding model. The sub-distribution space of category preference is modeled by the third encoding model described below. Among them, interest and category preference are different concepts. Category preference can clearly correspond to the category of the project, while interest may not correspond to any category of the project. Interest is an abstract concept. The interest may cover the categories of multiple projects, or only involve a part of the projects in a certain category of the project.

[0148] Exemplarily, the second coding model is established based on a Transformer encoder. In other implementations, the second coding model may also be established based on a neural network model such as Bidirectional Encoder Representations from Transformers (BERT) or a sparse transformer (Sparse Transformer). This embodiment does not limit the implementation method of the second coding model.

[0149] Based on the pre-trained second encoding model, the short-term interaction item sequence is encoded, including: inputting the short-term interaction item sequence into the pre-trained second encoding model to obtain a short-term interest feature.

[0150] Referring to the above, assume that the short-term interaction item sequence is The second encoding model is the Transformer encoder, then the short-term interest feature g s It can be expressed by the following formula:

[0151]

[0152] The training process of the second coding model is detailed below and will not be elaborated in detail in this embodiment.

[0153] Step 1033: Determine the category distribution characteristics of the historical interaction items and the category vectors of the L historical interaction items with the highest interaction frequencies of the target user based on the long-term interaction item sequence.

[0154] The category distribution characteristics of the historical interaction items are used to indicate the interaction frequency of the target user interacting with items of each category.

[0155] Exemplarily, determining the category distribution characteristics of the target user's historical interaction items includes: for the embedded representation of each item in the long-term interaction item sequence, determining the category to which the embedded representation belongs; counting the number of embedded representations corresponding to each category, and obtaining the frequency of the target user's interaction with the category based on the ratio of the number to the total number of embedded representations.

[0156] Assume that the set of items in the long-term interaction item sequence is Category distribution characteristics It can be expressed by the following formula:

[0157]

[0158] in, is an indicative function, C i represents the set of categories to which each item i in the long-term interaction item sequence belongs. The category to which item i belongs may be one category or at least two categories. k It represents the kth category among all the categories to which each item in the long-term interaction item sequence belongs, where k is less than or equal to A positive integer, A collection of categories representing all items that the user can interact with. For each category c k , if category c k and the category set C to which item i belongs i If one of the categories c′ is the same, then category c k The count of is increased by 1, if the category c k and the category set C to which item i belongs i If one of the categories c′ in is different, then category c k The count of remains unchanged; in the category set C i After the comparison is completed, the value of i is increased by 1, and the category c is changed again. k The set of categories C to which item i belongs i Compare; loop through and get That is, category c k The number of occurrences in a long-term interaction item sequence. It represents the total number of times each category to which each item in the long-term interaction item sequence belongs appears in the long-term interaction item sequence; express Indicates category c k The frequency of occurrence in a long sequence of interactive items.

[0159] Exemplarily, determining the category vectors of the L historical interaction items with the highest interaction frequency of the target user includes: constructing a one-hot vector for the L historical interaction items with the highest interaction frequency to obtain a category vector for each category. L is a positive integer, and L is pre-stored in the electronic device.

[0160] Distribution of features by category The frequency with which the target user interacts with each category of items can be determined. Therefore, based on the category distribution characteristics The L historical interaction items with the highest interaction frequency of the target user can be determined.

[0161] Among them, one-hot vector is a method in machine learning and data representation for converting categorical data into numerical form. In one-hot vector encoding, each category is represented as a vector whose length is equal to the total number of categories, and each element in the vector corresponds to a category. For a given category, only one element in the one-hot vector corresponding to the category has a value of 1, and the values ​​of the remaining elements are 0. The position of the element with a value of 1 corresponds to the index of the category. In this embodiment, the dimension of the one-hot vector is similar to the category distribution feature. The dimension of one-hot vector is

[0162] For example, the one-hot vector of category 1 is [1,0,0,…,0]; the one-hot vector of category 2 is [0,1,0,…,0]; the one-hot vector of category 3 is [0,0,1,…,0], and so on. The number of 0s in the one-hot vector of each category is The difference is the position of "1".

[0163] Based on the above principles, the category vector of each category can be expressed as:

[0164]

[0165] in, Represents the category vector of the l-th category, where l is a positive integer less than or equal to L.

[0166] In this embodiment, by obtaining the category vector of each category, specific guidance information for each category can be obtained, so that the finally generated item vector can correspond to the category clearly defined by the item, thereby achieving selective interest vector generation and recommendation.

[0167] Step 1034, based on the pre-trained third coding model, encode each category vector to obtain the conditional coding of each category; and encode the category distribution feature to obtain the conditional coding of the category distribution.

[0168] Exemplarily, the third coding model is established based on the MLP model. In other implementations, the third coding model may also be established based on a neural network model such as CNN or RNN. This embodiment does not limit the implementation method of the third coding model.

[0169] Based on a pre-trained third encoding model, the category distribution features are encoded to obtain conditional encoding of the category distribution, including: inputting the category distribution features into the pre-trained third encoding model to obtain conditional encoding of the category distribution.

[0170] Assume that the category distribution feature is expressed as The third encoding model is MLP, then the conditional encoding of the category distribution g c It can be expressed by the following formula:

[0171]

[0172] Based on the pre-trained third encoding model, each category vector is encoded to obtain the conditional encoding of each category, including: inputting the category vector into the pre-trained third encoding model to obtain the conditional encoding of each category.

[0173] Assume that each category vector is represented as The third encoding model is MLP, then the conditional encoding of each category It can be expressed by the following formula:

[0174]

[0175] Conditional coding for each category Composition Collection

[0176] Step 1031, step 1032, and steps 1033-1034 may be executed simultaneously, or executed sequentially. This embodiment does not limit the execution order of step 1031, step 1032, and steps 1033-1034.

[0177] Step 1035, generating first guidance information based on the long-term interest feature, the short-term interest feature, the conditional encoding of the category distribution, the noise data and the embedded representation of the time step t.

[0178] Specifically, information fusion is performed on long-term interest features, short-term interest features, conditional coding of category distribution, noise data, and embedded representation of time step t to obtain the first guidance information.

[0179] In one example, information fusion is performed on long-term interest features, short-term interest features, conditional coding of category distribution, noise data, and embedded representation of time step t through a pre-trained fusion model. The fusion model can be established based on MLP, or it can also be established based on a neural network model such as CNN or RNN. This embodiment does not limit the implementation method of the fusion model. The training process of the fusion model is described in detail below, and this embodiment will not be described in detail here.

[0180] At this time, the encoding form of the first guidance information can be expressed as:

[0181]

[0182] Among them, h u represents a specific embedding vector, represents a sequence of short-term interactive items, represents the category distribution characteristics, x t Represents noise data, t represents the diffusion time step. The value of t is initialized to T, T is a preset value, such as T = 32, and t is an integer that decreases from T to 0 as the reverse diffusion proceeds, and the reverse diffusion ends.

[0183] In other embodiments, the information fusion method may also be implemented through a vector connection operation, and this embodiment does not limit the information fusion method.

[0184] Taking the fusion model based on MLP as an example, the first guidance information It can be expressed by the following formula:

[0185]

[0186] Among them, c l represents the long-term interest feature, g s represents short-term interest features, g c represents the conditional encoding of the category distribution, δ represents the control x t The hyperparameter that affects δ is that a small δ is used to enhance the personalization of the target user, while a large δ is used to enhance randomness. δ is preset in the electronic device. t represents the noise data, e t represents the embedding representation at time step t. Represents vector concatenation.

[0187] For example, the embedding representation e at time step t is t It is calculated based on cosine and sinine functions. When the vector dimension d is an even number, the calculation method is to concatenate the cosine code of d / 2 and the similarity code of d / 2, which is specifically expressed as follows:

[0188]

[0189] When the vector dimension d is an odd number, it is represented as follows:

[0190]

[0191] Wherein, i is an integer ranging from 1 to d / 2; D is a preset number of steps, and the value of D is generally large, such as 10000. In other embodiments, the value of D may also be other values. This embodiment does not limit the value of D.

[0192] Step 1036, based on the long-term interest features, the short-term interest features, the conditional coding of the l-th category, the noise data and the embedded representation of the time step t, generate the second guidance information of the l-th category; wherein the diffusion guidance information includes the first guidance information and the second guidance information.

[0193] Specifically, information fusion is performed based on the long-term interest feature, the short-term interest feature, the conditional coding of the lth category, the noise data, and the embedded representation of the time step t to obtain the second guidance information. The information fusion method is the same as the information fusion method in step 1035, and this embodiment will not be repeated here.

[0194] Second guidance information It can be expressed by the following formula:

[0195]

[0196] Among them, c l represents the long-term interest feature, g s Indicates short-term interest characteristics, represents the conditional encoding of the lth category, δ represents the control x t The hyperparameters that affect x t represents the noise data at time step t, e t represents the embedding representation at time step t.

[0197] In other implementations, the same coding model can also be used to encode based on specific embedding vectors, short-term interaction item sequences, and long-term interaction item sequences. This embodiment does not limit the generation method of long-term interest features, short-term interest features, conditional coding of the lth category, and conditional coding of category distribution.

[0198] Optionally, in other implementations, the category vector may not be generated, and each category vector may not be encoded; or, the category distribution feature may not be generated, and the category distribution feature may not be encoded.

[0199] Optionally, step 1035 may be executed before step 1036, or may be executed during step 1036, or may be executed synchronously with step 1036. This embodiment does not limit the execution order between steps 1035 and 1036.

[0200] Step 104: determining candidate item vectors corresponding to multiple candidate recommendation item categories based on the diffusion guidance information.

[0201] In one example, determining candidate item vectors corresponding to multiple candidate recommendation item categories based on diffusion guidance information includes: performing reverse denoising on noise data based on a diffusion model and diffusion guidance information to obtain candidate item vectors corresponding to multiple candidate recommendation item categories. This process can be expressed by the following formula:

[0202]

[0203] Among them, β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); ∈ represents sampling noise; α t =1-β t ; Represents the diffusion guidance information. t represents the noise data at time step t, x t-1 Represents the noise data at time step t-1; when x is obtained t-1 Then, let t = t-1, and obtain x0 through multiple iterations of the above formula until t = 1, where x0 represents the candidate item vector

[0204] Referring to the above, if the diffusion guidance information includes the first guidance information And the second guidance information Then in the above formula for Accordingly, through the above-mentioned reverse denoising process, the candidate item vectors of the candidate items corresponding to the multiple candidate recommendation item categories can be obtained.

[0205] The training process of the diffusion model is described below, and this embodiment will not be described in detail here.

[0206] Step 105 : determining the items to be recommended for the target user based on the similarity between each candidate item vector and the actual item vector.

[0207] Since the candidate item vectors inversely reconstructed by the diffusion model may not correspond to the actual item vectors of the actual items, in this embodiment, the actual items can be determined by determining the actual item vectors that are most similar to the candidate item vectors.

[0208] Exemplarily, the actual item vector is an embedded representation of a candidate item in a candidate item set. The candidate item set may include all items, or may be items selected from all items and corresponding to the candidate recommendation item classification, or may be items selected according to other rules. This embodiment does not limit the method for determining the candidate item set.

[0209] If the similarity between the candidate item vector and the actual item vector is achieved through vector dot product, the similarity between each candidate item vector and the actual item vector can be expressed as follows:

[0210]

[0211] Where u represents the target user, i represents the i-th candidate item in the candidate item set, and s(u,i) represents the candidate item vector of target user u and the actual item vector h of candidate item i. i The maximum similarity between .

[0212] Assuming that the number of items to be recommended is a preset value K, where K is a positive integer less than or equal to the size of the candidate item set, the first K candidate items with the highest s(u,i) are selected as the items to be recommended.

[0213] In summary, the generative recommendation method based on conditional guided diffusion provided in this embodiment obtains noise data, short-term interaction item sequences and long-term interaction item sequences of the target user; obtains a pre-trained specific embedding vector of the target user; uses a pre-trained encoding model to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data; determines the candidate item vectors corresponding to multiple candidate recommendation item classifications based on the diffusion model; determines the target user's to-be-recommended items based on the similarity between each candidate item vector and the actual item vector; and can further guide the diffusion model to a specific sub-distribution space to generate categories of targeted items for the target user, on the premise of a specific distribution space of the target user generated based on the specific embedding vector, thereby solving the problem that the traditional item recommendation method cannot achieve the generation of guided and targeted user interest vectors, and achieving the generation of guided and targeted user interest vectors.

[0214] At the same time, since the user's long-term interests, short-term interests and category preferences are comprehensively considered to achieve personalized encoding, and the user distribution is learned based on the conditional guided diffusion model, it is possible to achieve guideable user interest generation for more accurate and diverse recommendations.

[0215] In addition, by encoding the category vectors of the L historical interaction items with the highest interaction frequency of the target user, the conditional coding of each category is obtained, and then the candidate item vector corresponding to each category is generated by reverse diffusion. The target user can be provided with recommended items corresponding to different categories, thereby realizing selective user interest vector generation.

[0216] Next, the training process of specific embedding vectors, encoding models, and diffusion models is introduced.

[0217] Figure 3 A flow chart of a model training method according to an embodiment of the present disclosure is shown. This embodiment takes the method used in an electronic device with processing capabilities such as a user terminal or a server as an example for explanation. The electronic device and the electronic device in the generative recommendation method embodiment are the same device or different devices. This embodiment does not limit the device type of the electronic device. Figure 3 As shown, the method includes:

[0218] Step 301: Generate a user ID of a target user.

[0219] Wherein, different target users have different user identifiers. Optionally, the method of generating the target user identifier includes but is not limited to: randomly generating a character string of a preset length to obtain the user identifier; or obtaining the identity information of the target user to obtain the user identifier. This embodiment does not limit the method of generating the user identifier.

[0220] Step 302: Obtain a short-term interaction item sequence sample, a long-term interaction item sequence sample, and a target item vector of the next interaction item of the target user.

[0221] The next interactive item refers to the item that the target user interacts with next, and the target item vector is the embedding representation of the next interactive item. That is, x0 = h i , where h i represents the target item vector, and x0 represents the data to be diffused.

[0222] For details on how to obtain long-term interaction item sequence samples, please refer to the relevant description of long-term interaction item sequence. The difference is that the long-term interaction item sequence samples are determined based on the collection time of the noise data samples. This embodiment does not elaborate on the method for obtaining long-term interaction item sequence samples.

[0223] The method for obtaining short-term interaction item sequence samples is detailed in the relevant description of short-term interaction item sequence. The difference is that the short-term interaction item sequence samples are determined based on the collection time of the noise data samples. This embodiment does not elaborate on the method for obtaining short-term interaction item sequence samples.

[0224] Optionally, step 302 may be executed after step 301, or may be executed before step 301, or may be executed synchronously with step 301. This embodiment does not limit the execution order between steps 301 and 302.

[0225] Step 303: Use the encoding model to be trained to determine the initialized specific embedding vector corresponding to the user identifier.

[0226] In this embodiment, the front end of the encoding model to be trained has a matrix representing a specific embedding vector, each element in the matrix is ​​initialized as an initialization value as a model parameter, and each row of elements in the matrix constitutes a specific embedding vector; different user identifiers are mapped to different rows of the matrix, thereby obtaining the initialized specific embedding vector corresponding to the user identifier.

[0227] Step 304 , based on the diffusion model to be trained, forward diffusion is performed on the target item vector for t steps to obtain a noise data sample at time step t.

[0228] The t-step forward diffusion of the target item vector can be expressed as follows:

[0229]

[0230] Among them, q(x t |x t-1 ) represents the noise data sample x at the given previous time step t-1 t-1 Generate a noise data sample x at time step t t The conditional probability of , t is an integer that increases from 0 to each time step. Among them, t is a positive integer less than or equal to T, T is the total time step of forward diffusion, for example: T is 32. β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); I represents the identity matrix. Represents the noise data sample x t For is the mean, and β t I is the covariance matrix of the Gaussian distribution.

[0231] Through the re-parameterization technique and the additivity principle of Gaussian distribution, the above forward diffusion formula can calculate the noise data sample x at time step t from x0 in one step.t ,Right now:

[0232]

[0233] Among them, x0 represents the target item vector; ∈ represents sampling noise, which follows Gaussian distribution β t are trainable model parameters of the diffusion model.

[0234] Step 305 , using the encoding model to be trained, based on the initialized specific embedding vector, the short-term interaction item sequence samples, the long-term interaction item sequence samples and the noise data samples, determines the encoding result.

[0235] If a specific embedding vector, a short-term interaction item sequence, and a long-term interaction item sequence correspond to different coding models respectively, then based on the above, it can be seen that the coding models to be trained include a first coding model to be trained, a second coding model to be trained, a third coding model to be trained, and a fusion model to be trained.

[0236] Accordingly, refer to Figure 4 , using the encoding model to be trained, based on the initialized specific embedding vector, the short-term interaction item sequence sample, the long-term interaction item sequence sample and the noise data sample, the encoding result is determined, including the following steps:

[0237] Step 3051, based on the first coding model to be trained, encode the initialized specific embedding vector to obtain a long-term interest coding result.

[0238] Based on the first coding model to be trained, encoding the initialized specific embedding vector includes: inputting the initialized specific embedding vector into the first coding model to be trained to obtain a long-term interest coding result.

[0239] According to the principle of the first coding model above, assume that the specific embedding vector initialized is represented as h u , The first encoding model to be trained is MLP, then the long-term interest encoding result c l It can be expressed by the following formula:

[0240] c l =MLP(h u )

[0241] in, d represents the dimension of the initialized specific embedding vector and the long-term feature of interest.

[0242] Step 3052: Encode the short-term interaction item sequence samples based on the second encoding model to be trained to obtain a short-term interest encoding result.

[0243] Based on the second encoding model to be trained, encoding the short-term interaction item sequence samples includes: inputting the short-term interaction item sequence samples into the second encoding model to be trained to obtain a short-term interest encoding result.

[0244] According to the principle of the second coding model above, assuming that the short-term interaction item sequence sample is The second encoding model to be trained is the Transformer encoder, then the short-term interest encoding result g s It can be expressed by the following formula:

[0245]

[0246] Step 3053: Determine the category distribution data of the target user's historical interaction items based on the long-term interaction item sequence samples.

[0247] The category distribution data is used to indicate the interaction frequency of the target user interacting with items of each category in the long-term interaction item sequence sample.

[0248] Exemplarily, determining the category distribution data of the target user's historical interaction items based on the long-term interaction item sequence samples includes: determining the frequency of occurrence of each category in the historical interaction items in the historical interactions to obtain the category distribution data.

[0249] Specifically, for the embedded representation of each item in the long-term interaction item sequence sample, determine the category to which the embedded representation belongs; count the number of embedded representations corresponding to each category, and based on the ratio of this number to the total number of embedded representations, obtain the frequency of occurrence of the category in historical interactions. For specific determination methods, see the category distribution characteristics above. The method for determining the value of is not described in detail in this embodiment.

[0250] Step 3054: Encode the category distribution data based on the third encoding model to be trained to obtain a category preference encoding result.

[0251] Based on the third encoding model to be trained, encoding the category distribution data includes: inputting the category distribution data into the third encoding model to be trained to obtain a category preference encoding result.

[0252] According to the principle of the third coding model above, it is assumed that the category distribution data is expressed as The third encoding model to be trained is MLP, then the category preference encoding result g c It can be expressed by the following formula:

[0253]

[0254] Optionally, in order to prevent over-reliance on categories with higher frequency during diffusion generation and to capture all category information when characterizing user distribution, a Dropout layer is introduced during training to classify category distribution data. Randomly setting some elements in to 0 can reduce the interdependence between neurons in the model and improve the generalization ability of the model.

[0255] Correspondingly, at this time, the category preference encoding result g c It can be expressed by the following formula:

[0256]

[0257] Optionally, step 3051, step 3052, and steps 3053-3054 may be executed synchronously, or executed sequentially. This embodiment does not limit the execution order of step 3051, step 3052, and steps 3053-3054.

[0258] Step 3055, unconditionally mask the short-term interest encoding result and the category preference encoding result based on the preset random variables to obtain the converted short-term interest encoding result and the converted category preference encoding result.

[0259] In this embodiment, in order to improve the generalization ability of the model, the forward diffusion process is assisted by unconditional masking. s and category preference encoding result g c With probability u s ,u c Randomly replace with dummy tokens for unconditional diffusion For example, two random variables m are set s ~Bernoulli s ) and m c ~Bernoulli c ), where m s ~Bernoulli s ) represents the random variable m s The obedience parameter is u s The Bernoulli distribution is a discrete probability distribution that describes random experiments with only two possible outcomes, usually called "success" and "failure". s is the probability of success, that is, m s The probability of taking the value is 1, and the probability of failure is 1-u s Similarly, m c ~Bernoulli c ) represents the random variable m c The obedience parameter is uc Bernoulli distribution, at this time, m c The probability of taking the value 1 is u c , m c The probability of taking the value as 0 is 1-u c . Accordingly, the short-term interest encoding results and the category preference encoding results are unconditionally masked based on random variables, which can be expressed as follows:

[0260] g′ s =(1-m s )·g s +m s ·ψ s ,

[0261] g′ c =(1-m c )·g c +m c ·ψ c .

[0262] Among them, m s Encoding result g for short-term interest s The corresponding random variable, m c The result of encoding category preference g c The corresponding random variable, ψ s Encoding result g for short-term interest s The corresponding virtual token, ψ c The result of encoding category preference g c The corresponding virtual token.

[0263] In other implementations, the generalization ability of the model may be improved by only using one of the Dropout layer and the unconditional mask. This embodiment does not limit the method of enhancing the generalization ability of the model used in the model training process.

[0264] Step 3056, based on the fusion model to be trained, information fusion is performed on the long-term interest encoding result, the converted short-term interest encoding result, the converted category preference encoding result, the noise data sample, and the embedded representation of the time step t to obtain the encoding result.

[0265] The method for obtaining the embedded representation of time step t and the type of fusion model to be trained are described above and will not be described in detail in this embodiment.

[0266] Taking the fusion model to be trained based on MLP as an example, the encoding result It can be expressed by the following formula:

[0267]

[0268] Among them, c lrepresents the long-term interest encoding result, g′ s represents the short-term interest encoding result after conversion, g′ c represents the converted category preference encoding result, δ represents the control x t The hyperparameters that affect x t represents the noise data sample, e t represents the embedding representation at time step t.

[0269] Step 306, based on the similarity between the encoding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, iteratively update the model parameters of the encoding model to be trained and the model parameters of the diffusion model to be trained to obtain the trained encoding model, the trained diffusion model and the trained specific embedding vector, the model parameters of the encoding model include the initialized specific embedding vector, the trained specific embedding vector is obtained by iteratively updating the initialized specific embedding vector, and the trained diffusion model is used to perform inverse denoising on the noise data based on the diffusion guidance information to generate a candidate item vector.

[0270] Exemplarily, the encoding model to be trained and the diffusion model to be trained are trained based on user loss, recommendation loss and diffusion loss. The user loss is used to calculate the similarity between a specific embedding vector in the model parameters and the target item vector. The recommendation loss is used to calculate the similarity between the encoding result and the target item vector. The diffusion loss is used to train the diffusion parameters of the diffusion model, such as the noise ratio β at time step t in the above text. t .

[0271] Accordingly, refer to Figure 5 , based on the similarity between the encoding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, iteratively update the model parameters of the encoding model to be trained and the model parameters of the diffusion model to be trained to obtain the trained encoding model, the trained diffusion model and the trained specific embedding vector, including the following steps:

[0272] Step 3061, obtaining the negative sample item vector of the next interactive item of the target user.

[0273] The negative sample item vector is an embedded representation of an item that the target user will not interact with. Exemplarily, the negative sample item vector is obtained by converting N items obtained by random sampling into corresponding embedded representations. The value of N is a positive integer.

[0274] In other embodiments, the negative sample item vector may also be determined from item vectors other than the long-term interaction item sequence samples. This embodiment does not limit the method for obtaining the negative sample item vector.

[0275] Step 3062, determining the user loss based on a first similarity between the initialized specific embedding vector and the target item vector, and a second similarity between the initialized specific embedding vector and the negative sample item vector.

[0276] User loss It can be expressed by the following formula:

[0277]

[0278] Wherein, σ(·) represents a similarity calculation function. In this embodiment, σ(·) is a sigmoid function, which can be expressed as h u represents the specific embedding vector for initialization, h i represents the target item vector, Represents the negative sample item vector.

[0279] In this example, the similarity calculation is obtained by calculating the cosine similarity metric and finally simplifying the dot product between two vectors. In actual implementation, the similarity calculation can also be other methods. This embodiment does not limit the calculation method of the similarity.

[0280] Step 3063: Determine the recommendation loss based on the third similarity between the encoding result and the target item vector, and the fourth similarity between the encoding result and the negative sample item vector.

[0281] Recommendation loss It can be expressed by the following formula:

[0282]

[0283] Among them, σ(·) represents the similarity calculation function, represents the encoding result, x0 represents the target item vector, Represents the negative sample item vector.

[0284] Since t is large, x t Approaching Gaussian noise will make it difficult for personalized coding to predict the encoding result close to the target item vector x0 Based on this, in this embodiment, only when t≤τ, The recommended loss is calculated when τ is τ, and T is the total time step of forward diffusion. In actual implementation, the value range of τ may also be other values, and this embodiment does not limit the value range of τ.

[0285] Step 3064, determining the diffusion loss based on the diffusion parameters of the diffusion model, the encoding result and the target item vector.

[0286] Diffusion loss It can be expressed by the following formula:

[0287]

[0288] Among them, α t =1-β t , β t is the trainable diffusion parameter of the diffusion model. Although, both the recommendation loss and the diffusion loss are based on x0 and The two have different focuses. In order to make the encoding model and the diffusion model more suitable for the recommendation task, the diffusion loss is used to train the diffusion model.

[0289] Optionally, steps 3062-3064 may be executed sequentially or synchronously, and this embodiment does not limit the execution order of steps 3062-3064.

[0290] Step 3065, combining the loss values ​​of user loss, recommendation loss and diffusion loss, iteratively update the model parameters of the encoding model to be trained and the diffusion parameters of the diffusion model to be trained to maximize the first similarity and the third similarity, minimize the second similarity and the fourth similarity, and minimize the diffusion loss, to obtain the trained encoding model, the trained diffusion model and the trained specific embedding vector.

[0291] In one example, a weighted result of user loss, recommendation loss and diffusion loss is determined to obtain a loss value, and based on the loss value, model parameters of the encoding model to be trained and diffusion parameters of the diffusion model to be trained are iteratively updated through a back propagation algorithm to obtain a trained encoding model and a trained diffusion model.

[0292] Among them, the weighted result of determining user loss, recommendation loss and diffusion loss can be expressed by the following formula:

[0293]

[0294] in, represents the loss value, represents the recommendation loss, Indicates user loss, represents diffusion loss, λ user Represents the weight of user loss, used to control Impact on loss value, λ diff Represents the weight of diffusion loss, which is used to control Impact on loss value.

[0295] In other embodiments, a weight may also be set for the recommendation loss. This embodiment does not limit the combination of user loss, recommendation loss and diffusion loss.

[0296] Since the model parameters of the trained encoding model include the trained specific embedding vector, the specific embedding vector corresponding to the target user can be obtained for use in recommending items to the target user. Exemplarily, since the trained specific embedding vector is a part of the trained encoding model, in step 102 of the above embodiment, the user ID of the target user can be input into the trained encoding model to obtain the specific embedding vector corresponding to the user ID.

[0297] To summarize, the model training method provided in this embodiment can refer to the similarity between the encoding result and the target item vector, as well as the similarity between the initialized specific embedding vector and the target item vector, to train the specific embedding vector corresponding to each target user, and obtain the encoding model and diffusion model suitable for different target users, which can improve the recommendation accuracy.

[0298] In addition, by introducing unconditional masks during the training process, the uncertainty in the model training process can be increased, thereby enhancing the generalization ability of the model.

[0299] In one possible implementation, when there is a new target user, in order to obtain the specific embedding vector corresponding to the target user, the above-mentioned model training method can be executed again; and / or, in order to ensure the accuracy of item recommendation, the above-mentioned model training method is executed once every preset time period to adapt to the possible changing interests and category preferences of each target user.

[0300] Optionally, the preset time period may be one month, or half a year, and this embodiment does not limit the value of the preset time period.

[0301] In order to more clearly understand the generative recommendation method provided by this application, a complete example of the method from model training to model application is given below. Figure 6 ,For the model training stage and the model inference stage, the encoding model is first used to perform conditional encoding and guided encoding respectively (refer to Figure 6 The difference is that the encoding model in the model training phase is the encoding model to be trained, and the model parameters are not accurate enough at this time, while the encoding model in the model inference phase is the trained encoding model, and the model parameters can be used for item recommendation.

[0302] The conditional encoding includes: embedding the specific embedding vector h corresponding to the target user u Input the first encoding model (MLP) to get the long-term interest c l During the training phase, h u Initialized to a randomly set value; during the application phase, h u is the vector representation after training.

[0303] The guided encoding includes two aspects: on the one hand, the short-term interaction sequence corresponding to the target user (i.e., the short-term interaction item sequence sample in the training phase or the short-term interaction item sequence in the application phase) Input the second encoding model (Transformer) to get the short-term interest g s .

[0304] On the other hand, the category distribution is determined based on the long-term interaction sequence corresponding to the target user (i.e., the long-term interaction item sequence sample in the training phase or the long-term interaction item sequence in the application phase), and the category preference g is obtained based on the category distribution and the third coding model (MLP). c .

[0305] In the process of guiding encoding, the process of guiding encoding in the training phase is slightly different from that in the application phase. For example, in the training phase, short-term interest g s and category preference g c Perform unconditional masking, and / or, require class preference g c In the application stage, it is also necessary to determine the category codes of the L categories with the highest target user interaction frequency, and use the third coding model to encode each category code. The specific description is as above, and this embodiment will not be repeated here.

[0306] Afterwards, the noise data at time step t, the embedded representation at time step t, and the long-term interest c l , short-term interests s , category preference g c Perform information fusion to obtain guidance information (i.e., the encoding results in the training phase or the diffusion guidance information in the application phase).

[0307] refer to Figure 6 In the training phase shown in part (b) of the figure, the target item vector x0=h based on the target user's next interactive item i and the guidance information of each time step, through the forward diffusion of time step t, the encoding result of time step t is obtained, combined with the specific embedding vector h u With the target item vector h i Determined user losses Encoding result based on time step t The recommendation loss determined by the target item vector x0 And the encoding result at time step t The recommendation loss determined by the target item vector x0 The encoding model and the diffusion model are trained to obtain a trained encoding model, a trained diffusion model, and a trained specific embedding vector.

[0308] refer to Figure 6 In the application phase shown in part (c), for each target user, The noise data of time step T is obtained by sampling, and the noise data is reversely denoised based on the diffusion guidance information starting from time step T to obtain the candidate item vector x0. The candidate item vector is compared with the actual item vector h i After the similarity comparison, the actual item vector with the highest similarity is output as the item to be recommended.

[0309] Figure 7 A block diagram of a generative recommendation device based on conditional guided diffusion according to an embodiment of the present disclosure is shown. The device at least includes: a first acquisition module 710, a second acquisition module 720, a data encoding module 730, a diffusion guidance module 740 and an item recommendation module 750.

[0310] A first acquisition module 710 is used to acquire noise data, a short-term interaction item sequence and a long-term interaction item sequence of a target user; wherein the short-term interaction item sequence is used to indicate item information interacted by the target user in a first time period; and the long-term interaction item sequence is used to indicate item information interacted by the target user in a second time period; the length of the first time period is shorter than the length of the second time period, and the first time period and the second time period are determined based on the current recommendation time;

[0311] A second acquisition module 720 is used to acquire a pre-trained specific embedding vector of a target user, where the specific embedding vector is used to distinguish interest distributions of different target users;

[0312] A data encoding module 730, configured to use a pre-trained encoding model to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data;

[0313] A diffusion guidance module 740 is used to determine candidate item vectors corresponding to multiple candidate recommendation item categories based on the diffusion guidance information;

[0314] The item recommendation module 750 is used to determine the item to be recommended for the target user based on the similarity between each candidate item vector and the actual item vector.

[0315] Optionally, the device further comprises:

[0316] An identification generating module, used to generate a user identification of the target user, where different target users have different user identifications;

[0317] A sample acquisition module, used to acquire a short-term interaction item sequence sample, a long-term interaction item sequence sample, and a target item vector of a next interaction item of the target user;

[0318] A vector determination module, used to determine the initialized specific embedding vector corresponding to the user identifier using the encoding model to be trained;

[0319] A noise generation module, used for performing t-step forward diffusion on the target item vector based on the diffusion model to be trained to obtain a noise data sample at time step t;

[0320] A sample encoding module, used to use the encoding model to be trained to determine an encoding result based on the initialized specific embedding vector, the short-term interaction item sequence samples, the long-term interaction item sequence samples and the noise data samples;

[0321] A model training module is used to iteratively update the model parameters of the coding model to be trained and the model parameters of the diffusion model to be trained based on the similarity between the coding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, so as to obtain a trained coding model, a trained diffusion model and a trained specific embedding vector, wherein the model parameters of the coding model include the initialized specific embedding vector, the trained specific embedding vector is obtained by iteratively updating the initialized specific embedding vector, and the trained diffusion model is used to perform inverse denoising on the noise data based on the diffusion guidance information to generate the candidate item vector.

[0322] Optionally, the model training module includes:

[0323] A negative sample acquisition unit, used to acquire a negative sample item vector of a next interactive item of the target user;

[0324] A first loss determination unit, configured to determine a user loss based on a first similarity between the initialized specific embedding vector and the target item vector, and a second similarity between the initialized specific embedding vector and the negative sample item vector;

[0325] a second loss determination unit, configured to determine a recommendation loss based on a third similarity between the encoding result and the target item vector, and a fourth similarity between the encoding result and the negative sample item vector;

[0326] a third loss determination unit, configured to determine a diffusion loss based on a diffusion parameter of the diffusion model to be trained, the encoding result and the target item vector;

[0327] A model parameter updating unit is used to iteratively update the model parameters of the encoding model to be trained and the diffusion parameters of the diffusion model to be trained based on the loss values ​​of the user loss, the recommendation loss and the diffusion loss, so as to maximize the first similarity and the third similarity, minimize the second similarity and the fourth similarity, and minimize the diffusion loss, so as to obtain a trained encoding model, a trained diffusion model and a trained specific embedding vector.

[0328] Optionally, the sample encoding module includes:

[0329] A first encoding unit, configured to encode the initialized specific embedding vector based on a first encoding model to be trained to obtain a long-term interest encoding result;

[0330] A second encoding unit is used to encode the short-term interactive item sequence samples based on a second encoding model to be trained to obtain a short-term interest encoding result;

[0331] A distribution determination unit, configured to determine category distribution data of historical interaction items of the target user based on the long-term interaction item sequence samples;

[0332] a third encoding unit, configured to encode the category distribution data based on a third encoding model to be trained to obtain the category preference encoding result;

[0333] An unconditional masking unit, used for unconditionally masking the short-term interest encoding result and the category preference encoding result respectively based on a preset random variable to obtain a converted short-term interest encoding result and a converted category preference encoding result;

[0334] The coding fusion unit is used to perform information fusion on the long-term interest coding result, the converted short-term interest coding result, the converted category preference coding result, the noise data sample, and the embedded representation of the time step t based on the fusion model to be trained to obtain the coding result.

[0335] Optionally, the distribution determination unit is specifically configured to:

[0336] The frequency of occurrence of each category in the historical interaction items in the historical interaction is determined to obtain the category distribution data.

[0337] Optionally, the data encoding module includes:

[0338] a fourth encoding unit, configured to encode the specific embedding vector based on a pre-trained first encoding model to obtain a long-term interest feature;

[0339] a fifth encoding unit, configured to encode the short-term interaction item sequence based on a pre-trained second encoding model to obtain a short-term interest feature;

[0340] A category determination unit, configured to determine, based on the long-term interaction item sequence, category distribution characteristics of historical interaction items and category vectors of L historical interaction items with the highest interaction frequency of the target user;

[0341] a sixth encoding unit, configured to encode each category vector based on a pre-trained third encoding model to obtain a conditional encoding of each category; and to encode the category distribution feature to obtain a conditional encoding of the category distribution;

[0342] A first generating unit, configured to generate first guidance information based on the long-term interest feature, the short-term interest feature, the conditional encoding of the category distribution, the noise data, and the embedded representation of the time step t;

[0343] A second generating unit, configured to generate second guidance information of the lth category based on the long-term interest feature, the short-term interest feature, the conditional encoding of the lth category, the noise data and the embedded representation of the time step t; wherein l is a positive integer less than or equal to L;

[0344] The diffusion guidance information includes the first guidance information and the second guidance information.

[0345] Optionally, the category determination unit is specifically configured to:

[0346] One-hot vectors are constructed for the L historical interaction items with the highest interaction frequency to obtain the category vector for each category.

[0347] Optionally, the diffusion guidance process of the diffusion guidance module is represented by the following formula:

[0348]

[0349] Among them, β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); ∈ represents sampling noise; α t =1-β t ; represents the diffusion guidance information, x t represents the noise data at time step t, x t-1 Represents the noise data at time step t-1; when x is obtained t-1Then, let t=t-1, and after multiple iterations until t=1, x0 is obtained, where x0 represents the candidate item vector.

[0350] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0351] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0352] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0353] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0354] Figure 8 1 is a block diagram of a generative recommendation apparatus 1900 based on conditional guided diffusion according to an exemplary embodiment. For example, the apparatus 1900 may be provided as a server or a terminal device. Figure 8 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0355] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , MacOS Noise Data M , Uni noise data M , Linu noise data M , FreeBSDTM or similar.

[0356] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to perform the above method.

[0357] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A generative recommendation method based on conditional guided diffusion, characterized in that: The method comprises: Acquire noise data, a short-term interaction item sequence, and a long-term interaction item sequence of a target user; wherein the short-term interaction item sequence is used to indicate item information interacted by the target user in a first time period; and the long-term interaction item sequence is used to indicate item information interacted by the target user in a second time period; the duration of the first time period is shorter than the duration of the second time period, and the first time period and the second time period are determined based on a current recommendation time; Obtaining a pre-trained specific embedding vector of a target user, where the specific embedding vector is used to distinguish interest distributions of different target users; Using a pre-trained encoding model, based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data, determining the diffusion guidance information of the target user; Determining candidate item vectors corresponding to a plurality of candidate recommendation item categories based on the diffusion guidance information; Based on the similarity between each candidate item vector and the actual item vector, the to-be-recommended item for the target user is determined.

2. The method according to claim 1, characterized in that The training process of the specific embedding vector and the encoding model includes: generating a user ID of the target user, wherein the user IDs of different target users are different; Acquire a short-term interaction item sequence sample, a long-term interaction item sequence sample, and a target item vector of a next interaction item of the target user; Using the encoding model to be trained, determining an initialized specific embedding vector corresponding to the user identifier; Based on the diffusion model to be trained, the target item vector is forward diffused for t steps to obtain a noise data sample at time step t; Using the encoding model to be trained, determining an encoding result based on the initialized specific embedding vector, the short-term interaction item sequence sample, the long-term interaction item sequence sample and the noise data sample; Based on the similarity between the encoding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, the model parameters of the encoding model to be trained and the model parameters of the diffusion model to be trained are iteratively updated to obtain a trained encoding model, a trained diffusion model and a trained specific embedding vector, the model parameters of the encoding model include the initialized specific embedding vector, the trained specific embedding vector is obtained by iteratively updating the initialized specific embedding vector, and the trained diffusion model is used to perform inverse denoising on the noise data based on the diffusion guidance information to generate the candidate item vector.

3. The method according to claim 2, characterized in that The iterative updating of the model parameters of the coding model to be trained and the model parameters of the diffusion model to be trained based on the similarity between the coding result and the target item vector, and the similarity between the initialized specific embedding vector and the target item vector, to obtain the trained coding model, the trained diffusion model and the trained specific embedding vector, comprises: Obtaining a negative sample item vector of the next interactive item of the target user; Determining a user loss based on a first similarity between the initialized specific embedding vector and the target item vector, and a second similarity between the initialized specific embedding vector and the negative sample item vector; Determining a recommendation loss based on a third similarity between the encoding result and the target item vector and a fourth similarity between the encoding result and the negative sample item vector; Determining a diffusion loss based on a diffusion parameter of the diffusion model to be trained, the encoding result and the target item vector; Combining the loss values ​​of the user loss, the recommendation loss, and the diffusion loss, iteratively update the model parameters of the encoding model to be trained and the diffusion parameters of the diffusion model to be trained to maximize the first similarity and the third similarity, minimize the second similarity and the fourth similarity, and minimize the diffusion loss, to obtain a trained encoding model, a trained diffusion model, and a trained specific embedding vector.

4. The method according to claim 2, characterized in that: The using the encoding model to be trained to determine the encoding result based on the initialized specific embedding vector, the short-term interaction item sequence sample, the long-term interaction item sequence sample and the noise data sample comprises: Based on the first encoding model to be trained, encoding the initialized specific embedding vector to obtain a long-term interest encoding result; Based on the second encoding model to be trained, encoding the short-term interactive item sequence samples to obtain a short-term interest encoding result; Determining category distribution data of historical interaction items of the target user based on the long-term interaction item sequence samples; Based on the third encoding model to be trained, encoding the category distribution data to obtain the category preference encoding result; Unconditionally masking the short-term interest encoding result and the category preference encoding result based on a preset random variable to obtain a converted short-term interest encoding result and a converted category preference encoding result; Based on the fusion model to be trained, information fusion is performed on the long-term interest encoding result, the converted short-term interest encoding result, the converted category preference encoding result, the noise data sample, and the embedded representation of time step t to obtain the encoding result.

5. The method according to claim 4, characterized in that The determining, based on the long-term interaction item sequence samples, the category distribution data of the historical interaction items of the target user comprises: The frequency of occurrence of each category in the historical interaction items in the historical interaction is determined to obtain the category distribution data.

6. The method according to claim 4, characterized in that The using of the pre-trained coding model to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data includes: Based on a pre-trained first encoding model, encoding the specific embedding vector to obtain a long-term interest feature; Based on a pre-trained second encoding model, encoding the short-term interaction item sequence to obtain a short-term interest feature; Based on the long-term interaction item sequence, determining the category distribution characteristics of the historical interaction items and the category vectors of the L historical interaction items with the highest interaction frequency of the target user; Based on the pre-trained third encoding model, each category vector is encoded to obtain a conditional encoding of each category; and the category distribution feature is encoded to obtain a conditional encoding of the category distribution; generating first guidance information based on the long-term interest feature, the short-term interest feature, the conditional encoding of the category distribution, the noise data, and the embedded representation of the time step t; Generate second guidance information of the lth category based on the long-term interest feature, the short-term interest feature, the conditional encoding of the lth category, the noise data and the embedded representation of the time step t; wherein l is a positive integer less than or equal to L; The diffusion guidance information includes the first guidance information and the second guidance information.

7. The method according to claim 6, characterized in that Determining the category vectors of the L historical interaction items with the highest interaction frequency of the target user includes: One-hot vectors are constructed for the L historical interaction items with the highest interaction frequency to obtain the category vector for each category.

8. The method according to any one of claims 1 to 7, characterized in that: The candidate item vectors corresponding to the plurality of candidate recommendation item categories are determined based on the diffusion guidance information, which is expressed by the following formula: Among them, β t is the noise ratio at time step t, which is used to control the amount of noise added at each time step; as t increases, β t Gradually increasing, β t ∈(0,1); ∈ represents sampling noise; α t =1-β t ; represents the diffusion guidance information, x t represents the noise data at time step t, x t-1 Represents the noise data at time step t-1; when x t-1 Then, let t=t-1, and after multiple iterations until t=1, x0 is obtained, where x0 represents the candidate item vector.

9. A generative recommendation device based on conditional guided diffusion, characterized in that: The device comprises: A first acquisition module is used to acquire noise data, a short-term interaction item sequence and a long-term interaction item sequence of a target user; wherein the short-term interaction item sequence is used to indicate item information interacted by the target user in a first time period; and the long-term interaction item sequence is used to indicate item information interacted by the target user in a second time period; the length of the first time period is shorter than the length of the second time period, and the first time period and the second time period are determined based on a current recommendation time; A second acquisition module is used to acquire a pre-trained specific embedding vector of a target user, where the specific embedding vector is used to distinguish interest distributions of different target users; A data encoding module, configured to determine the diffusion guidance information of the target user based on the specific embedding vector, the short-term interaction item sequence, the long-term interaction item sequence and the noise data using a pre-trained encoding model; A diffusion guidance module, configured to determine candidate item vectors corresponding to a plurality of candidate recommendation item categories based on the diffusion guidance information; The item recommendation module is used to determine the item to be recommended for the target user based on the similarity between each candidate item vector and the actual item vector.

10. A generative recommendation device based on conditional guided diffusion, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method described in any one of claims 1 to 8 when executing the instructions stored in the memory.

11. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.