Data processing method, model training method, content recommendation method and related products
By generating multiple interest distributions and concatenating vectors, the problem of insufficient recommendation accuracy and diversity in recommendation systems is solved. This approach improves recommendation diversity while maintaining accuracy, thereby enhancing the ability to identify and recommend products based on users' diverse interests.
Patent Information
- Application Number
- CN202511185726.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-05
AI Technical Summary
Existing recommendation systems are insufficient in terms of recommendation accuracy and diversity, making it difficult to simultaneously cover users' diverse interests.
By acquiring the object features and content interaction sequences of the target object, K multi-interest distributions are predicted. Each multi-interest distribution includes C interest indices, which belong to different interest sub-dictionaries. Multi-interest vectors are generated by training the interest model and concatenating the vectors for content recommendation.
While ensuring the accuracy of recommendations, it improves the diversity of recommendations, enhances the ability to identify and recommend diverse and potentially novel interests, and strengthens the diversity and novelty of the recommendation system.
Smart Images

Figure CN121071199A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular relates to a data processing method, a model training method, a content recommendation method and related products. BACKGROUND
[0002] In the retrieval stage of the recommendation system, large-scale recall often relies on vector retrieval technology, for example, ANN (Approximate Nearest Neighbor) retrieval. The traditional method can map the user into a unified vector, and recommend the content to the user through a single vector. However, the user often has diversified interests, and a single vector is difficult to cover multiple types of content that the user may like at the same time, thereby causing the recommendation system to have deficiencies in the accuracy and diversity of recommendation. SUMMARY
[0003] The present application provides a data processing method, a model training method, a content recommendation method and related products, which can improve the diversity of recommendation while ensuring the accuracy of recommendation.
[0004] In a first aspect, a data processing method is provided, and the method comprises:
[0005] obtaining an object feature of a target object and a content interaction sequence corresponding to the target object; the content interaction sequence comprises content that the target object has historically interacted with;
[0006] based on the object feature and the content interaction sequence, predicting K multi-interest distributions associated with the target object; K is an integer greater than or equal to 1;
[0007] Each of the K multi-interest distributions comprises C interest indexes, and the C interest indexes in each of the multi-interest distributions belong to C interest sub-dictionaries respectively, the C interest sub-dictionaries correspond to C different interest dimensions respectively, and the interest sub-dictionary of each interest dimension comprises a plurality of discrete interest vectors; C is an integer greater than 1.
[0008] In a second aspect, a data processing apparatus is provided, and the apparatus comprises:
[0009] an obtaining module configured to obtain an object feature of a target object and a content interaction sequence corresponding to the target object; the content interaction sequence comprises content that the target object has historically interacted with;
[0010] a predicting module configured to predict, based on the object feature and the content interaction sequence, K multi-interest distributions associated with the target object; K is an integer greater than or equal to 1;
[0011] Each of the K multi-interest distributions includes C interest indexes, the C interest indexes in each of the multi-interest distributions respectively belong to C interest sub-dictionaries, the C interest sub-dictionaries respectively correspond to different C interest dimensions, and the interest sub-dictionary of each interest dimension includes a plurality of discrete interest vectors; C is an integer greater than 1.
[0012] In a third aspect, a model training method is provided, and the method includes:
[0013] obtaining an initial interest model, sample object information of a sample object, a sample content interaction sequence corresponding to the sample object, and sample interaction content corresponding to the sample object; the sample content interaction sequence includes content that has been interacted with by the sample object, and the sample interaction content is content that has been interacted with by the sample object after the sample content interaction sequence;
[0014] obtaining a sample interest probability distribution of the sample object according to the sample object information and the sample content interaction sequence, and determining a total model loss of the initial interest model according to the sample interest probability distribution, the sample interaction content, the sample object information, and the sample content interaction sequence;
[0015] training at least part of the initial interest model according to the total model loss to obtain an interest model.
[0016] In a fourth aspect, a model training device is provided, and the device includes:
[0017] a data obtaining module configured to obtain an initial interest model, sample object information of a sample object, a sample content interaction sequence corresponding to the sample object, and sample interaction content corresponding to the sample object; the sample content interaction sequence includes content that has been interacted with by the sample object, and the sample interaction content is content that has been interacted with by the sample object after the sample content interaction sequence;
[0018] a prediction module configured to obtain a sample interest probability distribution of the sample object according to the sample object information and the sample content interaction sequence, and determine a total model loss of the initial interest model according to the sample interest probability distribution, the sample interaction content, the sample object information, and the sample content interaction sequence;
[0019] a training module configured to train at least part of the initial interest model according to the total model loss to obtain an interest model.
[0020] In a fifth aspect, a content recommendation method is provided, and the method includes:
[0021] obtaining a representation vector of a target object and K multi-interest vectors associated with the target object; each of the multi-interest vectors includes C discrete interest vectors, and each of the discrete interest vectors corresponds to an interest dimension or an interest category; K is an integer greater than or equal to 1, and C is an integer greater than 1;
[0022] based on the representation vector and the K multi-interest vectors, K retrieval vectors are obtained;
[0023] a plurality of candidate contents matching each retrieval vector are obtained;
[0024] target content is determined from the plurality of candidate contents, so as to recommend the target content to the target object.
[0025] In a sixth aspect, a content recommendation apparatus is provided, and the apparatus comprises:
[0026] a vector obtaining module configured to obtain a representation vector of a target object and K multi-interest vectors associated with the target object; each multi-interest vector comprises C discrete interest vectors, each discrete interest vector corresponds to an interest dimension or an interest category; K is an integer greater than or equal to 1, and C is an integer greater than 1;
[0027] the vector obtaining module is configured to obtain K retrieval vectors based on the representation vector and the K multi-interest vectors;
[0028] a content determining module configured to obtain a plurality of candidate contents matching each retrieval vector;
[0029] the content determining module is configured to determine target content from the plurality of candidate contents, so as to recommend the target content to the target object.
[0030] In a seventh aspect, an electronic device is provided, which comprises a processor and a memory, the memory is configured to store computer program code, the computer program code comprises computer instructions, and in the case that the processor executes the computer instructions, the electronic device executes the first aspect and any one of the embodiments thereof, or executes the third aspect and any one of the embodiments thereof, or executes the fifth aspect and any one of the embodiments thereof.
[0031] In an eighth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, the computer program comprises program instructions, and in the case that the program instructions are executed by a processor, the processor executes the first aspect and any one of the embodiments thereof, or executes the third aspect and any one of the embodiments thereof, or executes the fifth aspect and any one of the embodiments thereof.
[0032] In a ninth aspect, a computer program product is provided, and the computer program product comprises computer program or instructions, and in the case that the computer program or instructions are run on a computer, the computer executes the first aspect and any one of the embodiments thereof, or executes the third aspect and any one of the embodiments thereof, or executes the fifth aspect and any one of the embodiments thereof.
[0033] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the application.
[0034] The embodiment of the present application can obtain K multi-interest distributions associated with the target object, each of the K multi-interest distributions including C interest indexes, and the C interest indexes in each multi-interest distribution belong to C interest sub-dictionaries respectively, the C interest sub-dictionaries correspond to different C interest dimensions respectively, and the interest sub-dictionary of each interest dimension includes a plurality of discrete interest vectors. Therefore, the embodiment of the present application can cover multi-interests through multi-interest quantification (i.e., C discrete interest vectors) and multi-interest generation (i.e., K multi-interest distributions) at the same time to model the preferences of the target object, and then determine the target content for recommending to the target object in the content set through multi-interest recommendation, thereby improving the diversity of recommendation while ensuring the accuracy of recommendation. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the drawings needed to be used in the embodiments of the present application or the background art will be described below.
[0036] Figure 1 A scene schematic diagram for data interaction provided by the embodiment of the present application;
[0037] Figure 2 A flow schematic diagram of a content recommendation method provided by the embodiment of the present application;
[0038] Figure 3 A flow schematic diagram of a data processing method provided by the embodiment of the present application;
[0039] Figure 4 A structure schematic diagram of an interest model provided by the embodiment of the present application;
[0040] Figure 5 A flow schematic diagram of a model training method provided by the embodiment of the present application;
[0041] Figure 6 A structure schematic diagram of a content recommendation device provided by the embodiment of the present application;
[0042] Figure 7 A structure schematic diagram of a data processing device provided by the embodiment of the present application;
[0043] Figure 8 A structure schematic diagram of a model training device provided by the embodiment of the present application;
[0044] Figure 9 A hardware structure schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0046] The terms “first”, “second”, and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device.
[0047] Reference herein to “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily independent or alternative embodiments to each other. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0048] The content recommendation method provided by the embodiments of the present application is executed by a content recommendation device, the data processing method is executed by a data processing device, and the model training method is executed by a model training device. The content recommendation device, the data processing device and the model training device can be any electronic device that can execute the technical solutions disclosed in the embodiments of the present application. For example, the electronic device can be a terminal device or a server. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart home appliance (for example, a smart television), a wearable device, a vehicle-mounted terminal, an aircraft, and the like.
[0049] The embodiments of the present application can be applicable to a recommendation system (for example, applicable to a recall layer and a coarse ranking layer of a recommendation system), which can recommend personalized content to a user (for example, an object, which can be referred to as an account of the user for the convenience of understanding) so as to make the user consume the recommended content, for example, recommending content to the object when the object searches through a terminal device, recommending content to the object when the object browses an application client in the terminal device. The content can include video (for example, a movie and a TV series), audio (for example, music), text (for example, news), image (for example, a photo album), an electronic commodity (for example, a TV and a washing machine), and the like. Specific new media platforms of the recommendation system can include an e-commerce platform, a content community, a short video platform, and the like. Specific businesses of the recommendation system can include a TV recommendation business, a music recommendation business, a news recommendation business, a photo album recommendation business, a commodity recommendation business, and the like. Here, the specific businesses and the specific new media platforms of the recommendation system will not be enumerated one by one. For the convenience of understanding, the embodiments of the present application can refer to the content recommended to the object as target content, refer to the object to which the target content is recommended as a target object, and the number of target content can be one or more.
[0050] It should be understood that the related data (for example, object information, sample object information, a content interaction sequence, a sample content interaction sequence, and sample interaction content) in the embodiments of the present application must comply with the requirements of relevant laws and regulations when being collected and used, obtain the informed consent or separate consent of a personal information subject, and carry out subsequent data use and processing behavior within the scope of authorization of laws and regulations and the personal information subject.
[0051] The embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Please refer to Figure 1 , Figure 1 A scene schematic diagram for performing data interaction is provided in the embodiments of the present application. For the convenience of understanding, it is taken as an example that the method is executed by a server (for example, a server 10a) here, and it is taken as an example that content is recommended to an object (for example, an object 10c, which can also be referred to as a target object here) when the object browses an application client in a terminal device (for example, a terminal device 10b) here. The object 10c is an object who logs in the application client in the terminal device 10b.
[0052] As shown in Figure 1 , the object 10c can perform an opening operation on the application client in the terminal device 10b. In this way, the terminal device 10b can send a recommendation request to the server 10a in response to the opening operation performed by the object 10c on the application client. In this way, the server can receive the recommendation request sent by the terminal device 10b and determine target content for recommending to the object 10c in a content set.
[0053] The server 10a can obtain object information of the object 10c and a content interaction sequence corresponding to the object 10c, the object information is a data set for describing a target object (i.e., the object 10c) (for example, the object information can include basic information, personal preferences, social attributes, feedback evaluations, etc.), and the content interaction sequence includes content that the target object (i.e., the object 10c) has historically interacted with (for example, the interaction can include likes, comments, forwards, views, clicks, collections, etc.), and the content interaction sequence is used to represent the object behavior of the target object (i.e., the object 10c), and here it is illustrated by taking an example in which the content in the content interaction sequence includes content 11a, content 11b, content 11c, …, and content 11d. The server 10a can perform multi-interest prediction on the object 10c according to the object information of the object 10c and the content interaction sequence corresponding to the object 10c, and obtain K multi-interest distributions associated with the object 10c, where K can be an integer greater than or equal to 1, and here it is illustrated by taking an example in which the K multi-interest distributions include multi-interest distribution 12a, …, and multi-interest distribution 12b.
[0054] Each of the K multi-interest distributions includes C interest information, and the C interest information in each multi-interest distribution respectively belongs to C interest sub-dictionaries in the interest dictionary, and the C interest sub-dictionaries respectively correspond to interest information of different dimensions (i.e., each interest sub-dictionary corresponds to a different dimension of interest), where C can be an integer greater than 1, and here, C is taken as 3 for illustration, and the 3 interest sub-dictionaries can include interest sub-dictionary 13a, interest sub-dictionary 13b, and interest sub-dictionary 13c, for example, interest sub-dictionary 13a corresponds to interest information of a category dimension (i.e., the dimension of the interest information in interest sub-dictionary 13a is category, for example, the category can be a music category, an art category, a movie category, a data category), interest sub-dictionary 13b corresponds to interest information of a style dimension (i.e., the dimension of the interest information in interest sub-dictionary 13b is style, for example, the style can be a music style, an art style, an architectural style, a literature style), and interest sub-dictionary 13c corresponds to interest information of a popularity dimension (i.e., the dimension of the interest information in interest sub-dictionary 13c is popularity, for example, the popularity can be high popularity, medium popularity, and low popularity). Each of the C interest sub-dictionaries includes at least two interest information of the same dimension and a discrete interest vector corresponding to each of the at least two interest information, interest sub-dictionary 13a includes interest information 14a, …, interest information 14b, and interest sub-dictionary 13a further includes discrete interest vector 15a corresponding to interest information 14a, …, discrete interest vector 15b corresponding to interest information 14b, interest sub-dictionary 13b includes interest information 16a, …, interest information 16b, and interest sub-dictionary 13b further includes discrete interest vector 17a corresponding to interest information 16a, …, discrete interest vector 17b corresponding to interest information 16b, and interest sub-dictionary 13c includes interest information 18a, …, interest information 18b, and interest sub-dictionary 13c further includes discrete interest vector 19a corresponding to interest information 18a, …, discrete interest vector 19b corresponding to interest information 18b, for example, interest information 14a can be a music category, interest information 14b can be a movie category, interest information 16a can be an art style, interest information 16b can be an architectural style, interest information 18a can be high popularity, and interest information 18b can be medium popularity.
[0055] The server 10a can perform vector lookup on the C interest sub-dictionaries according to the C interest information in each multi-interest distribution, obtain C discrete interest vectors respectively corresponding to each multi-interest distribution, perform vector splicing on the C discrete interest vectors respectively corresponding to each multi-interest distribution, and obtain a multi-interest vector respectively corresponding to each multi-interest distribution (i.e., the multi-interest vector 20a corresponding to the multi-interest distribution 12a, …, the multi-interest vector 20b corresponding to the multi-interest distribution 12b). For ease of understanding, the multi-interest distribution 12a is taken as an example here. The multi-interest distribution 12a includes the interest information 14a in the interest sub-dictionary 13a, the interest information 16a in the interest sub-dictionary 13b, and the interest information 18b in the interest sub-dictionary 13c. The interest sub-dictionary 13a is subjected to vector lookup according to the interest information 14a in the multi-interest distribution 12a, to obtain the discrete interest vector 15a. The interest sub-dictionary 13b is subjected to vector lookup according to the interest information 16a in the multi-interest distribution 12a, to obtain the discrete interest vector 17a. The interest sub-dictionary 13c is subjected to vector lookup according to the interest information 18b in the multi-interest distribution 12a, to obtain the discrete interest vector 19b. The discrete interest vector 15a, the discrete interest vector 17a, and the discrete interest vector 19b are determined as the C discrete interest vectors corresponding to the multi-interest distribution 12a. The C discrete interest vectors (i.e., the discrete interest vector 15a, the discrete interest vector 17a, and the discrete interest vector 19b) corresponding to the multi-interest distribution 12a are subjected to vector splicing, to obtain the multi-interest vector 20a corresponding to the multi-interest distribution 12a.
[0056] The server 10a can determine target content for recommendation to the object 10c in the content set according to the K multi-interest vectors (i.e., the multi-interest vector 20a, …, the multi-interest vector 20b), the object information of the object 10c, and the content in the content set (the content in the content set is content to be recommended), and return the target content to the terminal device 10b, so that the object 10c performs a trigger operation in the terminal device 10b for the target content. For example, the trigger operation here can be browsing news or watching a movie.
[0057] It can be seen that the embodiment of the present application can generate K multi-interest vectors for the target object, each of which includes C discrete interest vectors of different dimensions, and the C discrete interest vectors are used to represent interest information of different dimensions, thereby representing the diversified interests of the target object through multi-interest quantization (i.e., C discrete interest vectors) and multi-interest generation (i.e., K multi-interest vectors), finely modeling the multi-interest, improving the ability to capture potential interest changes of the target object, covering multiple types of content that the target object may like, improving the diversity of the recommendation while ensuring the accuracy of the recommendation, i.e., improving the identification and recommendation ability of the recommendation system for diversified interests and potential novel interests, predicting and capturing new demands that the object may produce, thereby achieving significant improvement in business indicators (e.g., retention rate, click rate, distribution efficiency, etc.). In addition, the embodiment of the present application is particularly effective for new objects whose potential interests have not yet appeared in historical behaviors, or objects with extremely high interest diversification, or new objects with less historical behavior, improving the recall ability of unpopular, long-tail or fresh content.
[0058] Please refer to Figure 2 , Figure 2 A flowchart of a content recommendation method provided by the embodiment of the present application is shown.
[0059] In step S101, a representation vector of a target object and K multi-interest vectors associated with the target object are obtained.
[0060] Specifically, the target object is subjected to multi-interest prediction according to the object information of the target object and the content interaction sequence corresponding to the target object, to obtain N multi-interest distribution probabilities corresponding to N multi-interest distributions respectively, and the N multi-interest distributions are filtered according to the N multi-interest distribution probabilities, to obtain K multi-interest distributions associated with the target object in the N multi-interest distributions. The content interaction sequence includes content that the target object has interacted with historically, and the content interaction sequence includes content at T time points (i.e., time point 1, …, time point T, which can be continuous or discontinuous, and the T time points are time points before the current time point) or T time stamps. Each of the N multi-interest distributions includes C interest information (or each of the K multi-interest distributions includes C interest information), the C interest information in each multi-interest distribution belongs to C interest sub-dictionaries in an interest dictionary respectively, the interest dictionary belongs to a multi-interest maintenance sub-model in an interest model, the C interest sub-dictionaries correspond to interest information of different dimensions respectively, each of the C interest sub-dictionaries includes at least two interest information of the same dimension and discrete interest vectors corresponding to the at least two interest information respectively, and the C interest sub-dictionaries can discretize an interest space through vector quantization, so as to effectively solve the interest collapse (or interest collapse, which refers to that when a plurality of interest vectors are extracted or generated, the plurality of interest vectors are too strong in mutual “overlap” or “homogenization”, resulting in that the plurality of interest vectors have low distinguishability and cannot effectively represent different interest aspects. Here, the interest vector can be a discrete interest vector. In other words, the interest collapse refers to that most content is quantized into very few interest information or discrete interest vectors); here, K can be an integer greater than or equal to 1, here, C can be an integer greater than 1, here, T can be an integer greater than 1, and here, N can be an integer greater than K. The N multi-interest distributions are obtained by combining the interest information in each interest sub-dictionary, the N multi-interest distributions are different from each other, N is determined by the number of interest information in each interest sub-dictionary, or in other words, N is obtained by multiplying the number of interest information in the C interest sub-dictionaries (for example, if C is equal to 3 and the number of interest information in the three interest sub-dictionaries is 4, 5 and 6 respectively, then N is equal to 4*5*6=120). Each multi-interest vector includes C discrete interest vectors, and each discrete interest vector corresponds to an interest dimension or an interest category.
[0061] The interest dictionary can be regarded as a combination of the C interest sub-dictionaries, and the C interest sub-dictionaries form a Cartesian product to construct the interest dictionary. For ease of understanding, the relationship between the interest dictionary and the C interest sub-dictionaries can be seen from the following formula (1):
[0062]
[0063] where C denotes the number of interest sub-dictionaries, denotes the c-th interest sub-dictionary among the C interest sub-dictionaries, M c denotes the number of interest information (or discrete interest vectors) in the c-th interest sub-dictionary, d denotes the size of the discrete interest vectors, the size of the discrete interest vectors in the C interest sub-dictionaries are the same, E * denotes all possible combinations of interest information (i.e., N multi-interest profiles).
[0064] It can be understood that the object feature is obtained from the object information of the target object, and the object feature and the content interaction sequence corresponding to the target object are input into the multi-interest generation sub-model in the interest model. The object feature is a subset of the object information. The multi-interest generation sub-model includes a generative network and a feature conversion network. The generative network learns the sequential behavior (i.e., the content interaction sequence) of the target object conditioned on the object feature. The specific type of the generative network and the feature conversion network is not limited in the present application. For example, the generative network can be a GPT model (Generative Pre-trained Transformers), and the feature conversion network can be a BERT model (Bidirectional Encoder Representations from Transformers), a VGG model (Visual Geometry Group), etc. The object feature and the content in the content interaction sequence are respectively extracted by the multi-interest generation sub-model to obtain a key object vector corresponding to the object feature and a historical content vector corresponding to the content in the content interaction sequence. The object feature is extracted by the feature conversion network (which can also be referred to as a first feature conversion network) to obtain a key object vector corresponding to the object feature. The content in the content interaction sequence is extracted by the feature conversion network (which can also be referred to as a second feature conversion network) to obtain a historical content vector corresponding to the content in the content interaction sequence. One content in the content interaction sequence corresponds to one historical content vector, and T contents in the content interaction sequence correspond to T historical content vectors. The key object vector and the historical content vector corresponding to the content in the content interaction sequence are concatenated by the multi-interest generation sub-model (i.e., the key object vector and the T historical content vectors are concatenated). An object content vector is obtained. The distribution of the object content vector is predicted by the multi-interest generation sub-model to obtain a distribution prediction vector corresponding to N multi-interest distributions. The object content vector is predicted by the generative network to obtain a candidate prediction vector. The candidate prediction vector is fully connected (which can be realized by a fully connected layer) to obtain a distribution prediction vector corresponding to N multi-interest distributions. The distribution prediction vector is normalized (which can be realized by a normalization exponential function) to obtain a normalized distribution prediction vector (i.e., a posterior distribution) corresponding to N multi-interest distributions. The multi-interest distribution probability corresponding to each of the N multi-interest distributions is obtained from the normalized distribution prediction vector.The N multi-interest distribution probabilities are probabilities of the target object possibly appearing in the corresponding multi-interest distribution at the next moment (i.e., moment T+1) predicted by the multi-interest generation sub-model, that is, the multi-interest generation sub-model can learn the multi-interest distribution that the target object is likely to appear in at the next moment.
[0065] The N multi-interest distribution probabilities are sorted to obtain sorted N multi-interest distribution probabilities, the multi-interest distribution probabilities are iteratively obtained from the sorted N multi-interest distribution probabilities in descending order, the sum of the iteratively obtained multi-interest distribution probabilities is determined, and if the sum of the iteratively obtained multi-interest distribution probabilities is greater than or equal to a probability threshold (the probability threshold is greater than 0 and less than 1, and the specific value of the probability threshold is not limited in the embodiments of the present application, for example, the probability threshold can be equal to 0.9, 0.95, etc.), the iteratively obtained multi-interest distribution probability is determined as a candidate multi-interest distribution probability. The number of the candidate multi-interest distribution probabilities is S, which can be an integer greater than or equal to K. Here, the number of the iteratively obtained multi-interest distribution probabilities is greater than or equal to K when the sum of the iteratively obtained multi-interest distribution probabilities is greater than or equal to the probability threshold. The S candidate multi-interest distribution probabilities are normalized (the normalization can be realized by a normalization exponential function) to obtain normalized candidate multi-interest distribution probabilities corresponding to the S candidate multi-interest distribution probabilities, respectively, S multi-interest distributions corresponding to the S normalized candidate multi-interest distribution probabilities are obtained from the N multi-interest distributions (i.e., the S candidate multi-interest distribution probabilities are obtained from the N multi-interest distributions), and the K multi-interest distributions associated with the target object are obtained from the S multi-interest distributions according to the S normalized candidate multi-interest distribution probabilities. The S normalized candidate multi-interest distribution probabilities refer to the probabilities that the corresponding multi-interest distributions are the multi-interest distributions associated with the target object, in other words, the probability that the multi-interest distribution corresponding to a larger normalized candidate multi-interest distribution probability is the multi-interest distribution associated with the target object is greater, and the probability that the multi-interest distribution corresponding to a smaller normalized candidate multi-interest distribution probability is the multi-interest distribution associated with the target object is smaller. Optionally, if the sum of the iteratively obtained multi-interest distribution probabilities is less than the probability threshold, the multi-interest distribution probabilities are iteratively obtained from the un-iterated multi-interest distribution probabilities (i.e., the multi-interest distribution probabilities in the N multi-interest distribution probabilities that have not been iterated) in descending order until the sum of the iteratively obtained multi-interest distribution probabilities is greater than or equal to the probability threshold.
[0066] The specific process of obtaining K multi-interest distributions associated with the target object from N multi-interest distributions according to the probability threshold and N multi-interest distribution probabilities can be referred to as a Nucleus sampling strategy. The Nucleus sampling strategy can randomly sample from a candidate distribution set (the candidate distribution set includes S multi-interest distributions) in which the probability accumulation (i.e., the sum of the obtained multi-interest distribution probabilities) is greater than τ (i.e., the probability threshold). The S multi-interest distributions are multi-interest distributions with high likelihood, balance diversity and accuracy, and capture the potential preferences of the target object. In other words, the Nucleus sampling strategy can ensure that the selection of K multi-interest distributions is concentrated on the most likely set (i.e., the set composed of S multi-interest distributions), while allowing some randomness to improve novelty, discarding low-probability multi-interest distributions (i.e., multi-interest distributions other than S multi-interest distributions in N multi-interest distributions), and reducing the fatigue of homogeneous content.
[0067] Optionally, if the sum of the obtained multi-interest distribution probabilities is greater than or equal to the probability threshold, and the number of the obtained multi-interest distribution probabilities is less than K, then the multi-interest distribution probabilities in the untraversed multi-interest distribution probabilities are traversed in descending order until the number of the obtained multi-interest distribution probabilities is equal to K. The K obtained multi-interest distribution probabilities are determined as candidate multi-interest distribution probabilities, and K multi-interest distributions corresponding to the K candidate multi-interest distribution probabilities are obtained from N multi-interest distributions. The K multi-interest distributions are determined as the multi-interest distributions associated with the target object.
[0068] It can be understood that K-1 multi-interest distributions associated with the target object are obtained from the S multi-interest distributions according to the S normalized candidate multi-interest distribution probabilities. The specific process of obtaining K-1 multi-interest distributions associated with the target object from the S multi-interest distributions according to the S normalized candidate multi-interest distribution probabilities can be referred to the description of obtaining K multi-interest distributions associated with the target object from the S multi-interest distributions according to the S normalized candidate multi-interest distribution probabilities, and will not be described here. According to the interest information in the K-1 multi-interest distributions, the selected times corresponding to at least two interest information in each interest sub-dictionary are determined. The selected times refer to the number of multi-interest distributions in which the corresponding interest information belongs to the K-1 multi-interest distributions. For example, 3 multi-interest distributions in the K-1 multi-interest distributions include a certain interest information, and the selected times corresponding to the interest information is equal to 3. 0 multi-interest distributions in the K-1 multi-interest distributions include a certain interest information, and the selected times corresponding to the interest information is equal to 0 (i.e., the initial selected times). A candidate multi-interest distribution is obtained from S-K+1 multi-interest distributions according to the S normalized candidate multi-interest distribution probabilities, and the selected times corresponding to the interest information in each interest sub-dictionary is updated (or the selected times corresponding to the interest information in the candidate multi-interest distribution is continued to be updated, and the update here can be self-increment) according to the interest information in the candidate multi-interest distribution, to obtain the updated selected times corresponding to at least two interest information in each interest sub-dictionary (or the updated selected times corresponding to the interest information in the candidate multi-interest distribution). The S-K+1 multi-interest distributions are multi-interest distributions in the S multi-interest distributions except the K-1 multi-interest distributions. If the updated selected times corresponding to at least two interest information in each interest sub-dictionary (or the updated selected times corresponding to the interest information in the candidate multi-interest distribution) are all less than the selected times threshold (the specific value of the selected times threshold is not limited in the embodiments of the present application, for example, the selected times threshold can be equal to 3), the candidate multi-interest distribution is determined as the multi-interest distribution associated with the target object, and K multi-interest distributions associated with the target object are obtained.
[0069] Optionally, if there is an updated selected time greater than or equal to the selected time threshold in the updated selected times corresponding to the at least two interest information in each interest sub-dictionary (or the updated selected times corresponding to the interest information in the candidate multi-interest distribution), it is determined that the candidate multi-interest distribution is not the multi-interest distribution associated with the target object, the updated selected times corresponding to the interest information in the candidate multi-interest distribution are restored (i.e., the updated selected times corresponding to the interest information in the candidate multi-interest distribution are restored to the selected times before the update), the S-K multi-interest distributions (i.e., the multi-interest distributions other than the candidate multi-interest distribution in the S-K+1 multi-interest distributions) are obtained according to the S normalized candidate multi-interest distribution probabilities, and the updated multi-interest distribution is determined as the multi-interest distribution associated with the target object according to the interest information in the updated multi-interest distribution and the selected times corresponding to the at least two interest information in each interest sub-dictionary, thereby obtaining the K multi-interest distributions associated with the target object. The specific process of determining the updated multi-interest distribution as the multi-interest distribution associated with the target object according to the interest information in the updated multi-interest distribution and the selected times corresponding to the at least two interest information in each interest sub-dictionary can be referred to the description of determining the candidate multi-interest distribution as the multi-interest distribution associated with the target object according to the candidate multi-interest distribution and the selected times corresponding to the at least two interest information in each interest sub-dictionary, which will not be repeated here.
[0070] The specific process of obtaining the K multi-interest distributions associated with the target object from the S multi-interest distributions according to the selected time threshold and the S normalized candidate multi-interest distribution probabilities can be referred to as a controllable aggregation strategy. The controllable aggregation strategy can prevent the discrete vectors (i.e., discrete interest vectors) or interest information in the interest sub-dictionary from being repeated too many times in the K multi-interest distributions, further improve the coverage of the multi-interest, and thus improve the diversity of the K multi-interest distributions.
[0071] Optionally, the N multi-interest distribution probabilities are sorted to obtain sorted N multi-interest distribution probabilities, the K multi-interest distribution probabilities are obtained from the sorted N multi-interest distribution probabilities in descending order, the K multi-interest distributions corresponding to the K multi-interest distribution probabilities are obtained from the N multi-interest distributions, and the K multi-interest distributions are determined as the multi-interest distributions associated with the target object.
[0072] The K multi-interest distributions include a multi-interest distribution H j , and the multi-interest distribution H j may be any one of the K multi-interest distributions. Here, j can be an integer less than or equal to K. The multi-interest distribution H jvector search (i.e., according to the multi-interest distribution H j vector search), obtaining the multi-interest distribution H j corresponding to the C discrete interest vectors, vector splicing is performed on the multi-interest distribution H j corresponding to the C discrete interest vectors, vector splicing is performed on the multi-interest distribution H j corresponding to the C discrete interest vectors.
[0073] Optionally, vector fusion is performed on the C discrete interest vectors corresponding to each multi-interest distribution respectively (the specific manner of vector fusion is not limited in the embodiments of the present application, for example, the vector fusion can be weighted summation or addition), obtaining the multi-interest vector corresponding to each multi-interest distribution respectively, wherein, the multi-interest distribution H j vector fusion is performed on the C discrete interest vectors corresponding to each multi-interest distribution respectively, obtaining the multi-interest vector corresponding to each multi-interest distribution respectively, wherein, the multi-interest distribution H j corresponding to the C discrete interest vectors.
[0074] It should be understood that the object information of the target object and the content interaction sequence corresponding to the target object are polled to obtain K multi-interest distributions associated with the target object by performing multi-interest prediction on the target object according to the object information and the content interaction sequence, and the K multi-interest distributions are stored in the interest cache (i.e., Top-K interest cache, User Top-K Interest Cache). Wherein, polling the object information and the content interaction sequence to obtain the K multi-interest distributions can update the K multi-interest distributions in the interest cache at a lower frequency in a timely manner (i.e., periodically), or update the K multi-interest distributions in the interest cache when the object behavior updates significantly to obtain more accurate K multi-interest distributions. In this way, when receiving the recommendation request (or search service request) sent by the target object, the K multi-interest distributions are obtained from the interest cache, and the step of performing vector lookup on the C interest sub-dictionaries according to the C interest information in each multi-interest distribution is performed to obtain the C discrete interest vectors corresponding to each multi-interest distribution. Wherein, the interest cache can reduce the real-time generation overhead in the online inference stage, avoid real-time calculation of the K multi-interest distributions, reduce the online calculation pressure caused by real-time multi-interest generation, improve the efficiency of obtaining the K multi-interest distributions, and further improve the efficiency of obtaining the K multi-interest vectors, improve the efficiency of similarity search according to the K multi-interest vectors, and greatly improve the QPS (Queries Per Second, Queries Per Second) performance of the recommendation system. In addition, the embodiments of the present application can realize online service through the interest cache, and the cache mechanism of the interest cache can allow the sub-models (here, the sub-models include the multi-interest generation sub-model and the multi-interest maintenance sub-model) in the interest model to use models with arbitrary complexity without increasing the online inference delay.
[0075] Optionally, the index of each of the K multi-interest distributions is stored into the interest cache, and when a recommendation request (or a search service request) sent by a target object is received, the index of each of the K multi-interest distributions is obtained from the interest cache, and the C interest sub-dictionaries are vector searched according to the index of each of the K multi-interest distributions, to obtain C discrete interest vectors corresponding to each of the multi-interest distributions. The index of the multi-interest distribution can represent C interest information in each of the multi-interest distributions, and the index of the multi-interest distribution can be an index for the interest dictionary or an index for the interest sub-dictionary. The index for the interest dictionary and the index for the interest sub-dictionary can be converted to each other. The index for the interest dictionary refers to representing the multi-interest distribution by one index value, and the index for the interest sub-dictionary refers to representing the multi-interest distribution by C index values. Here, taking C equal to 3 as an example, for example, when the index of the multi-interest distribution is the index for the interest dictionary, the index of the multi-interest distribution can be 0, which can represent the first interest information in the first interest sub-dictionary, the first interest information in the second interest sub-dictionary, and the first interest information in the third interest sub-dictionary; when the index of the multi-interest distribution is the index for the interest dictionary, the index of the multi-interest distribution can be 1, which can represent the first interest information in the first interest sub-dictionary, the first interest information in the second interest sub-dictionary, and the second interest information in the third interest sub-dictionary; when the index of the multi-interest distribution is the index for the interest sub-dictionary, the index of the multi-interest distribution can be 123, 1 can represent the second interest information in the first interest sub-dictionary, 2 can represent the third interest information in the second interest sub-dictionary, and 3 can represent the fourth interest information in the third interest sub-dictionary; when the index of the multi-interest distribution is the index for the interest sub-dictionary, the index of the multi-interest distribution can be 032, 0 can represent the first interest information in the first interest sub-dictionary, 3 can represent the fourth interest information in the second interest sub-dictionary, and 2 can represent the third interest information in the third interest sub-dictionary.
[0076] When the index of the multi-interest distribution is the index for the interest dictionary, the index for the interest dictionary is converted into the index for the interest sub-dictionary, C index values are obtained from the index for the interest sub-dictionary, the corresponding interest sub-dictionaries are vector searched according to the C index values, to obtain discrete interest vectors corresponding to the C index values, and the C discrete interest vectors are determined as the C discrete interest vectors corresponding to the multi-interest distribution. When the index of the multi-interest distribution is the index for the interest sub-dictionary, C index values are obtained from the index for the interest sub-dictionary, the corresponding interest sub-dictionaries are vector searched according to the C index values, to obtain discrete interest vectors corresponding to the C index values, and the C discrete interest vectors are determined as the C discrete interest vectors corresponding to the multi-interest distribution.
[0077] Therefore, K multi-interest distributions associated with the target object are obtained. Each multi-interest distribution includes C interest indexes respectively corresponding to C interest sub-dictionaries, and the C interest sub-dictionaries respectively correspond to different C interest dimensions. The interest sub-dictionary of each interest dimension includes a plurality of discrete interest vectors respectively corresponding to a plurality of interest categories in the interest dimension. Based on the discrete interest vectors of the C interest indexes in each multi-interest distribution in the corresponding interest sub-dictionary, a multi-interest vector corresponding to each multi-interest distribution is obtained.
[0078] In step S102, K retrieval vectors are obtained based on the representation vector and the K multi-interest vectors.
[0079] In step S103, a plurality of candidate contents matching each retrieval vector are obtained.
[0080] Specifically, the content vector of the content in the content set is obtained, and the similarity between the target object and the content in the content set is determined according to the retrieval vector and the content vector. According to the similarity between the target object and the content in the content set, a plurality of candidate contents matching the retrieval vector in the content set are determined.
[0081] In step S104, the target content is determined from the plurality of candidate contents, so as to recommend the target content to the target object.
[0082] Specifically, the K multi-interest vectors, the object information, and the content in the content set are input into a multi-interest retrieval sub-model in the interest model. The multi-interest retrieval sub-model is used to retrieve related content (i.e., the target content), and the multi-interest retrieval sub-model includes an object tower network, a content tower network, and an object interest fusion network. The object tower network and the content tower network can constitute a dual-tower model (Dual-Tower Model), and the specific types of the object tower network, the content tower network, and the object interest fusion network are not limited in the embodiments of the present application. The object information is feature-extracted by the object tower network to obtain a representation vector corresponding to the object information, and the content in the content set is feature-extracted by the content tower network to obtain a content feature corresponding to the content in the content set. The representation vector and the K multi-interest vectors are respectively feature-fused by the object interest fusion network to obtain K retrieval vectors respectively corresponding to the K multi-interest distributions. The representation vector and the multi-interest distribution H j corresponding to the multi-interest vector are feature-fused by the object interest fusion network to obtain the multi-interest distribution H j corresponding to the retrieval vector; the representation vector and the multi-interest distribution H j corresponding to the multi-interest vector are vector-spliced by the object interest fusion network to obtain the multi-interest distribution H j corresponding to the object interest splicing vector, and the multi-interest distribution H ja representation vector in the corresponding object interest splicing vector and the multi-interest distribution H j perform feature fusion on the corresponding multi-interest vectors to obtain the multi-interest distribution H j the corresponding retrieval vectors. Obtain the feature dot product (i.e., dot product) between the content features and the K retrieval vectors respectively, and determine the maximum feature dot product in the K feature dot products as the similarity between the target object and the content in the content set. Wherein, the specific way of determining the similarity is not limited in the embodiments of the present application, here taking the feature dot product as an example to illustrate that the cosine similarity, Euclidean distance, etc. can also be determined as the similarity. According to the similarity between the target object and the content in the content set, determine the target content in the content set for recommendation to the target object. For example, when the embodiments of the present application are applied to the recall layer of the recommendation system, according to the similarity between the target object and the content in the content set, recall the content in the content set, and obtain the target content for recommendation to the target object from the recalled content; when the embodiments of the present application are applied to the coarse ranking layer of the recommendation system, add the content recalled by the recall layer of the recommendation system to the content set, and according to the similarity between the target object and the content in the content set, perform coarse ranking on the content in the content set to obtain the content after coarse ranking and screening, and obtain the target content for recommendation to the target object from the content after coarse ranking and screening.
[0083] For ease of understanding, the specific process of determining the K retrieval vectors corresponding to the K multi-interest distributions respectively according to the K multi-interest vectors, object information and content in the content set can be referred to formula (2) as follows:
[0084]
[0085] Wherein, here taking one of the K multi-interest distributions (for example, the multi-interest distribution H j ) as an example to illustrate, denotes the multi-interest vector, denotes the K multi-interest distributions (or the index of each multi-interest distribution in the K multi-interest distributions, the index of the multi-interest distribution can also be referred to as the interest index), u denotes the representation vector, concat denotes vector splicing, and z fusion (·) denotes the object interest fusion network, u k denotes the retrieval vector.
[0086] An embodiment of the present application proposes a generative multi-interest recommendation framework (GemiRec), which is used to output multiple differentiated interest vectors (here, the interest vector can be a multi-interest vector) to match different contents in the retrieval or recommendation link, so as to improve the coverage and accuracy of the recommendation, and aims to predict the content that the object is most likely to be interested in at the next time step (i.e., the next moment). For ease of understanding, the specific process of obtaining the similarity between the target object and the content in the content set can refer to the following formula (3):
[0087]
[0088] wherein u∈U represents any object (for example, the target object) in the object set (i.e., U), f k (·) represents a feature point product (or similarity) calculated according to the kth multi-interest distribution of the target object, represents the similarity between the target object and the content in the content set, represents the content in the content set (the content set includes the content that the target object is likely to interact at time t+1), represents the content interaction sequence, represents the content that the target object interacts at time t.
[0089] As can be seen, the embodiment of the present application can map the target object into K multi-interest vectors, and each of the K multi-interest vectors includes C discrete interest vectors, so as to cover the multi-interest through multi-interest quantization (i.e., C discrete interest vectors) and multi-interest generation (i.e., K multi-interest vectors) at the same time, to model the preference of the target object, and then determine the target content for recommending to the target object through multi-interest recommendation, while ensuring the accuracy of the recommendation, to improve the diversity of the recommendation.
[0090] Please refer to Figure 3 , Figure 3 A flowchart of a data processing method provided by the embodiment of the present application.
[0091] Step S201, obtaining the object feature of the target object and the content interaction sequence corresponding to the target object;
[0092] The content interaction sequence includes the content that the target object has historically interacted with;
[0093] Step S202, based on the object feature and the content interaction sequence, predicting K multi-interest distributions associated with the target object;
[0094] Specifically, according to the object feature and the content interaction sequence, an interest probability distribution of the target object is obtained. The interest probability distribution is used to represent probabilities of different multi-interest distributions (i.e., multi-interest distribution probabilities corresponding to the N multi-interest distributions respectively). Sampling is performed on the interest probability distribution to obtain K multi-interest distributions associated with the target object, in other words, according to the N multi-interest distribution probabilities, the N multi-interest distributions are screened to obtain K multi-interest distributions associated with the target object from the N multi-interest distributions.
[0095] The interest probability distribution of the target object is obtained according to the object feature and the content interaction sequence by using a multi-interest generation sub-model, and the multi-interest generation sub-model includes a generative model.
[0096] The K multi-interest distributions each include C interest indexes, the C interest indexes in each multi-interest distribution belong to C interest sub-dictionaries respectively, the C interest sub-dictionaries correspond to different C interest dimensions respectively, and the interest sub-dictionary of each interest dimension includes a plurality of discrete interest vectors. C is an integer greater than 1.
[0097] The specific process of steps S201-S202 can be referred to the description of steps S101-S104 in the above Figure 2 The description of steps S101-S104 in the above
[0098] It can be seen that the embodiments of the present application can obtain K multi-interest distributions associated with the target object, each of the K multi-interest distributions includes C interest indexes, the C interest indexes in each multi-interest distribution belong to C interest sub-dictionaries respectively, the C interest sub-dictionaries correspond to different C interest dimensions respectively, and the interest sub-dictionary of each interest dimension includes a plurality of discrete interest vectors. Therefore, the embodiments of the present application can simultaneously cover multi-interest through multi-interest quantization (i.e., C discrete interest vectors) and multi-interest generation (i.e., K multi-interest distributions) to model the preference of the target object, and then determine the target content for recommending to the target object in the content set through multi-interest recommendation, thereby ensuring the accuracy of recommendation and improving the diversity of recommendation.
[0099] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of an interest model provided by the embodiments of the present application. As Figure 4As shown, in the reasoning phase (i.e., the prediction phase), the interest model can include an IDMM module (Interest Dictionary Maintenance Module), a MIPDM module (Multi-Interest Posterior Distribution Module), and a MIRM module (Multi-Interest Retrieval Module), the IDMM module corresponds to a multi-interest maintenance sub-model (i.e., an IDMM model), the MIPDM module corresponds to a multi-interest generation sub-model (i.e., a MIPDM model), and the MIRM module corresponds to a multi-interest retrieval sub-model (i.e., a MIRM model). Among them, the IDMM module maintains a vector quantization dictionary (i.e., an interest dictionary) containing multiple groups of discrete discrete interest vectors, the MIPDM module uses a generative model to predict the object interest distribution that may appear at the next moment, to capture the potential and not yet explicitly appearing in the historical behavior of the interest of the object, and the MIRM module uses multiple interest representations (i.e., multiple interest vectors) to calculate the similarity to retrieve the content.
[0100] The interest dictionary in the multi-interest maintenance sub-model includes C interest sub-dictionaries, which can specifically include interest sub-dictionary 1, interest sub-dictionary 2, …, and interest sub-dictionary C. The number of interest information (or discrete interest vectors) in the interest sub-dictionary 1 is M1, the number of interest information (or discrete interest vectors) in the interest sub-dictionary 2 is M2, …, and the number of interest information (or discrete interest vectors) in the interest sub-dictionary C is M C C. M1 interest information (or discrete interest vectors) in the interest sub-dictionary 1, M2 interest information (or discrete interest vectors) in the interest sub-dictionary 2, …, and M C C interest information (or discrete interest vectors) in the interest sub-dictionary C can be combined to obtain M1*M2*…*M C C multi-interest vectors.
[0101] As Figure 4 shown, taking a target object as an example, the object features of the target object and the content interaction sequence corresponding to the target object are obtained, and the object features and the content interaction sequence are input into the multi-interest generation sub-model. Among them, the content interaction sequence includes the content that the target object has interacted with in the past. Here, taking the content interaction sequence including content i1 (i.e., the above ), content i2 (i.e., the above ), …, and content i t (i.e., the above ) as an example; the multi-interest generation sub-model includes a feature conversion network (i.e., Z feat ), a generative network (i.e., Z GPT) and a fully connected layer (i.e., Z FC The object feature is extracted by the feature conversion network to obtain a key object vector corresponding to the object feature. The content in the content interaction sequence is extracted by the multi-interest generation sub-model to obtain a historical content vector corresponding to the content in the content interaction sequence. The key object vector and the historical content vector corresponding to the content in the content interaction sequence are concatenated by the multi-interest generation sub-model to obtain an object content vector. The object content vector is distributed predicted by the generative network to obtain a candidate prediction vector. The candidate prediction vector is fully connected by the fully connected layer to obtain a distribution prediction vector corresponding to N multi-interest distributions. The distribution prediction vector is normalized by a normalization exponential function to obtain a normalized distribution prediction vector corresponding to the N multi-interest distributions. The multi-interest distribution probabilities corresponding to the N multi-interest distributions are obtained from the normalized distribution prediction vector. The K multi-interest distributions associated with the target object are obtained by screening the N multi-interest distributions according to the N multi-interest distribution probabilities. The index (i.e., interest index) of each multi-interest distribution in the K multi-interest distributions is stored in the interest cache. Here, the index of the multi-interest distribution is taken as an example of the index of the interest sub-dictionary. The K indexes include the index of the first multi-interest distribution i1, the index of the second multi-interest distribution i2, …, and the index of the Kth multi-interest distribution i K The interest cache includes the index of each multi-interest distribution in the K multi-interest distributions associated with the object u1 (for example, the index of the first multi-interest distribution i1 is 123, the index of the second multi-interest distribution i2 is 345, …, and the index of the Kth multi-interest distribution i K The interest cache includes the index of each multi-interest distribution in the K multi-interest distributions associated with the object u2 (for example, the index of the first multi-interest distribution i1 is 041, the index of the second multi-interest distribution i2 is 574, …, and the index of the Kth multi-interest distribution i K The interest cache includes the index of each multi-interest distribution in the K multi-interest distributions associated with the object u3 (for example, the index of the first multi-interest distribution i1 is 212, the index of the second multi-interest distribution i2 is 340, …, and the index of the Kth multi-interest distribution i K The interest cache includes the index of each multi-interest distribution in the K multi-interest distributions associated with the object u3 (for example, the index of the first multi-interest distribution i1 is 212, the index of the second multi-interest distribution i2 is 340, …, and the index of the Kth multi-interest distribution i
[0102] As Figure 4As shown, when receiving the recommendation request sent by the target object, the index of each multi-interest distribution in the K multi-interest distributions is obtained from the interest cache, the C interest sub-dictionaries are vector searched according to the index of each multi-interest distribution, C discrete interest vectors respectively corresponding to each multi-interest distribution are obtained, the C discrete interest vectors respectively corresponding to each multi-interest distribution are vector spliced, a multi-interest vector respectively corresponding to each multi-interest distribution is obtained, and the K multi-interest vectors belong to M1*M2*…*M C interest vectors. Wherein, the multi-interest vector can also be called a multi-interest retrieval vector, and here the target object is taken as an example of object u1, the C interest sub-dictionaries are vector searched according to the index 123 of the 1st multi-interest distribution i1, the C discrete interest vectors corresponding to the 1st multi-interest distribution i1 are obtained, the C discrete interest vectors corresponding to the 1st multi-interest distribution i1 are vector spliced, and the multi-interest vector corresponding to the 1st multi-interest distribution i1 is obtained; the C interest sub-dictionaries are vector searched according to the index 345 of the 2nd multi-interest distribution i2, the C discrete interest vectors corresponding to the 2nd multi-interest distribution i2 are obtained, the C discrete interest vectors corresponding to the 2nd multi-interest distribution i2 are vector spliced, and the multi-interest vector corresponding to the 2nd multi-interest distribution i2 is obtained; …; the C interest sub-dictionaries are vector searched according to the index 678 of the Kth multi-interest distribution i K , the C discrete interest vectors corresponding to the Kth multi-interest distribution i K are obtained, the C discrete interest vectors corresponding to the Kth multi-interest distribution i K are vector spliced, and the multi-interest vector corresponding to the Kth multi-interest distribution i K is obtained.
[0103] As shown in Figure 4 , the object information of the target object is obtained, the K multi-interest vectors of the target object, the object information and the content in the content set are input into the multi-interest retrieval sub-model. Wherein, the multi-interest retrieval sub-model includes an object tower network, a content tower network and an object interest fusion network, and here a candidate content (i.e. a first candidate content) in the content set is taken as an example, and the candidate content can be any one content in the content set. The object information is feature extracted through the object tower network to obtain a representation vector corresponding to the object information; the candidate content is feature extracted through the content tower network to obtain a content feature (i.e. v) corresponding to the candidate content. The representation vector and the K multi-interest vectors are respectively feature fused through the object interest fusion network to obtain a retrieval vector (i.e. u k , or u1, u2, …, u K). The feature dot product between the content feature and the K search vectors is obtained, the maximum feature dot product in the K feature dot products is determined as the similarity between the target object and the candidate content. According to the similarity between the target object and the content in the content set, the target content (for example, the target content can be the candidate content) in the content set is determined for recommendation to the target object.
[0104] Please refer to Figure 5 , Figure 5 A flowchart of a model training method provided by an embodiment of the present application is shown.
[0105] In step S301, an initial interest model, sample object information of a sample object, a sample content interaction sequence corresponding to the sample object, and a sample interaction content corresponding to the sample object are obtained.
[0106] The initial interest model includes an initial multi-interest generation sub-model, an initial multi-interest search sub-model, and an initial multi-interest maintenance sub-model. The sample object information is a data set for describing the sample object (for example, the sample object information can include basic information, personal preferences, social attributes, feedback evaluations, etc.). The sample content interaction sequence includes content that the sample object has interacted with (for example, the interaction can include likes, comments, forwards, views, clicks, collections, etc.). The sample content interaction sequence is used to represent the object behavior of the sample object. The sample interaction content is content that the sample object has interacted with after the sample content interaction sequence. The sample content interaction sequence includes content at T time points (i.e., time point 1, …, time point T, the T time points can be continuous or discontinuous) or T time stamps. The sample interaction content is content that the sample object has interacted with at the next time point (i.e., time point T+1) of the T time points.
[0107] In step S302, a sample interest probability distribution of the sample object is obtained according to the sample object information and the sample content interaction sequence. A model total loss of the initial interest model is determined according to the sample interest probability distribution, the sample interaction content, the sample object information, and the sample content interaction sequence.
[0108] Specifically, the first model total loss is determined according to the sample object information, the sample content interaction sequence, the sample interaction content, the initial multi-interest generation sub-model, the initial multi-interest retrieval sub-model, and the initial multi-interest maintenance sub-model, the initial multi-interest retrieval sub-model is adjusted in parameters according to the first model total loss, and a candidate multi-interest retrieval sub-model is obtained. The second model total loss is determined according to the sample object information, the sample content interaction sequence, the sample interaction content, the initial multi-interest generation sub-model, the candidate multi-interest retrieval sub-model, and the initial multi-interest maintenance sub-model, the initial multi-interest maintenance sub-model is adjusted in parameters according to the second model total loss, and a candidate multi-interest maintenance sub-model is obtained; the first dictionary loss is determined according to the sample object information, the sample content interaction sequence, the sample interaction content, the candidate multi-interest retrieval sub-model, and the initial multi-interest maintenance sub-model, the initial interest dictionary in the initial multi-interest maintenance sub-model is adjusted in dictionary according to the first dictionary loss, and a candidate interest dictionary is obtained. The third model total loss is determined according to the sample object information, the sample content interaction sequence, the sample interaction content, the initial multi-interest generation sub-model, the candidate multi-interest retrieval sub-model, and the candidate multi-interest maintenance sub-model, the candidate multi-interest retrieval sub-model, the candidate multi-interest maintenance sub-model, and the initial multi-interest generation sub-model are adjusted in parameters according to the third model total loss, and a multi-interest retrieval sub-model, a multi-interest maintenance sub-model, and a multi-interest generation sub-model are obtained; the second dictionary loss is determined according to the sample object information, the sample content interaction sequence, the sample interaction content, the candidate multi-interest retrieval sub-model, and the candidate multi-interest maintenance sub-model, the candidate interest dictionary is adjusted in dictionary according to the second dictionary loss, and an interest dictionary in the multi-interest maintenance sub-model is obtained.
[0109] The initial multi-interest maintenance sub-model, the candidate multi-interest maintenance sub-model, and the multi-interest maintenance sub-model can be collectively referred to as a first generalization network model, and the initial multi-interest maintenance sub-model, the candidate multi-interest maintenance sub-model, and the multi-interest maintenance sub-model belong to the names of the first generalization network model at different times. In the training stage, the first generalization network model can be referred to as the initial multi-interest maintenance sub-model and the candidate multi-interest maintenance sub-model, and in the inference stage, the first generalization network model can be referred to as the multi-interest maintenance sub-model. The initial multi-interest retrieval sub-model, the candidate multi-interest retrieval sub-model, and the multi-interest retrieval sub-model can be collectively referred to as a second generalization network model, and the initial multi-interest retrieval sub-model, the candidate multi-interest retrieval sub-model, and the multi-interest retrieval sub-model belong to the names of the second generalization network model at different times. In the training stage, the second generalization network model can be referred to as the initial multi-interest retrieval sub-model and the candidate multi-interest retrieval sub-model, and in the inference stage, the second generalization network model can be referred to as the multi-interest retrieval sub-model. The initial multi-interest generation sub-model and the multi-interest generation sub-model can be collectively referred to as a third generalization network model, and the initial multi-interest generation sub-model and the multi-interest generation sub-model belong to the names of the third generalization network model at different times. In the training stage, the third generalization network model can be referred to as the initial multi-interest generation sub-model, and in the inference stage, the third generalization network model can be referred to as the multi-interest generation sub-model. Similarly, the initial interest dictionary, the candidate interest dictionary, and the interest dictionary are names of different stages, the initial interest dictionary, the candidate interest dictionary, and the interest dictionary can be collectively referred to as a dictionary, the initial interest dictionary and the candidate interest dictionary are names of the training stage, and the interest dictionary is a name of the inference stage. The initial interest sub-dictionary, the candidate interest sub-dictionary, and the interest sub-dictionary are names of different stages, the initial interest sub-dictionary, the candidate interest sub-dictionary, and the interest sub-dictionary can be collectively referred to as a sub-dictionary, the initial interest sub-dictionary and the candidate interest sub-dictionary are names of the training stage, and the interest sub-dictionary is a name of the inference stage.
[0110] It should be understood that, based on the sample interaction content, the first model loss of the initial multi-interest maintenance sub-model is determined; based on the sample object information and the sample content interaction sequence, multi-interest prediction is performed on the sample object to obtain N normalized sample distribution prediction vectors corresponding to the multi-interest distributions; based on the normalized sample distribution prediction vectors and the sample interaction content, the second model loss of the initial multi-interest generation sub-model is determined. In other words, based on the sample object information and the sample content interaction sequence, the sample interest probability distribution of the sample object is obtained; based on the sample interest probability distribution and the sample interaction content, the second model loss of the initial multi-interest generation sub-model is determined; based on the sample object information, the sample content interaction sequence, and the sample interaction content, the positive sample parameters and negative sample parameters are determined; based on the positive sample parameters and negative sample parameters, the third model loss of the initial multi-interest retrieval sub-model is determined. Based on the first model loss, the second model loss, and the third model loss, the total first model loss (or the total model loss of the initial interest model) is determined.
[0111] The process involves inputting the sample interaction content features corresponding to the sample interaction content into an initial multi-interest maintenance sub-model. This sub-model comprises an initial interest encoding network, an initial interest decoding network, and C initial interest sub-dictionaries. The initial interest encoding network encodes the sample interaction content features to obtain the encoded content features. Based on these encoded content features, vector mapping is performed on the C initial interest sub-dictionaries to obtain the residual vector and the discrete interest vector corresponding to each initial interest sub-dictionary. These C discrete interest vectors are then concatenated to obtain the sample multi-interest vector corresponding to the sample interaction content. The initial interest decoding network decodes the sample multi-interest vector to obtain the decoded content features. Based on the C initial interest sub-dictionaries, the sample interaction content features, the decoded content features, the C discrete interest vectors, and the C residual vectors, the first model loss of the initial multi-interest maintenance sub-model is determined.
[0112] The process involves inputting sample object features and sample content interaction sequences from the sample object information into the initial multi-interest generation sub-model. The sample key object vector corresponding to the sample object features and the sample historical content vector corresponding to the content in the sample content interaction sequence are concatenated to obtain the sample object content vector. Distribution prediction is performed on the sample object content vector to obtain normalized sample distribution prediction vectors corresponding to N multi-interest distributions. These normalized vectors include the sample multi-interest distribution probabilities corresponding to the N multi-interest distributions. In other words, distribution prediction is performed on the sample object content vector to obtain the sample interest probability distribution, which represents the probability of different multi-interest distributions (i.e., the sample multi-interest distribution probabilities corresponding to the N multi-interest distributions). The sample multi-interest distributions of the sample interaction content for the C initial interest sub-dictionaries are obtained. Based on the sample multi-interest distributions and the normalized sample distribution prediction vectors, the second model loss of the initial multi-interest generation sub-model is determined. In other words, the sample multi-interest distributions of the sample interaction content are obtained, and the second model loss of the initial multi-interest generation sub-model is determined based on the sample multi-interest distributions and the sample interest probability distributions. The sample multi-interest distribution is determined by the interest information of the C sample discrete interest vectors in the corresponding initial interest sub-dictionary.
[0113] The initial multi-interest retrieval sub-model inputs the sample multi-interest vectors corresponding to the sample interaction content, the sample object information, and the content from the sample content interaction sequence. Feature fusion is performed on the sample representation vector corresponding to the sample object information and the sample multi-interest vector to obtain the sample retrieval vector. The feature dot product between the sample content features corresponding to the content in the sample content interaction sequence and the sample retrieval vector is determined as the positive sample parameter. Negative content samples and negative interest samples for the sample object are obtained, and negative sample parameters are determined based on the negative content samples, negative interest samples, sample object information, sample content interaction sequence, and sample interaction content. The third model loss of the initial multi-interest retrieval sub-model is determined based on the positive and negative sample parameters.
[0114] Specifically, the first loss weight corresponding to the first model loss, the second loss weight corresponding to the second model loss, and the third loss weight corresponding to the third model loss are obtained. The first model loss, the second model loss, and the third model loss are then weighted and summed based on these three weights to obtain the total loss of the first model. The first loss weight, the second loss weight, and the third loss weight are hyperparameters that control the relative importance of the first model loss, the second model loss, and the third model loss. This embodiment does not limit the specific values of the first loss weight, the second loss weight, and the third loss weight. For example, the first loss weight can be equal to 0.3, the second loss weight can be equal to 0.3, and the third loss weight can be equal to 0.4.
[0115] For ease of understanding, the specific process of determining the total loss of the first model based on the first model loss, the second model loss, and the third model loss can be found in the following formula (4):
[0116] L total =λ1L IDMM +λ2L MIPDM +λ3L MIRM (4)
[0117] Among them, L IDMM L represents the first model loss of the initial multi-interest maintenance sub-model. MIPDM L represents the second model loss of the initial multi-interest generating sub-model. MIRM λ1 represents the first loss weight, λ2 represents the second loss weight, and λ3 represents the third loss weight. L represents the third loss weight. total This represents the total loss of the first model.
[0118] It should be understood that, based on the sample interaction content, the first model loss of the initial multi-interest maintenance sub-model is determined; positive and negative sample parameters are determined based on the sample object information, sample content interaction sequence, and sample interaction content; and the fourth model loss of the candidate multi-interest retrieval sub-model is determined based on the positive and negative sample parameters. The specific process of determining the fourth model loss of the candidate multi-interest retrieval sub-model based on the positive and negative sample parameters can be found in the description of determining the third model loss of the initial multi-interest retrieval sub-model based on the positive and negative sample parameters, and will not be repeated here. The initial interest dictionary in the initial multi-interest maintenance sub-model is obtained. Based on the C initial interest sub-dictionaries in the initial interest dictionary, the first model loss, and the fourth model loss, the first dictionary loss corresponding to each of the C initial interest sub-dictionaries is determined. Subtraction operations are performed on the C initial interest sub-dictionaries and their corresponding first dictionary losses to obtain the candidate interest sub-dictionaries corresponding to the C initial interest sub-dictionaries. These C candidate interest sub-dictionaries are then determined as the candidate interest dictionary.
[0119] Specifically, the partial derivatives of the first model loss with respect to the C initial interest sub-dictionaries in the initial interest dictionary (i.e., the first partial derivatives, which represent the gradients from the initial multi-interest maintenance sub-model and are used to determine how the initial interest sub-dictionaries affect the first model loss) are obtained. The partial derivatives of the fourth model loss with respect to the C initial interest sub-dictionaries in the initial interest dictionary (i.e., the second partial derivatives, which represent the gradients from the candidate multi-interest retrieval sub-model and are used to determine how the initial interest sub-dictionaries affect the fourth model loss) are also obtained. The first loss weight corresponding to the first model loss and the fourth loss weight corresponding to the fourth model loss are also obtained. The first loss weight and the fourth loss weight are hyperparameters that control the relative importance of the first and second partial derivatives. This embodiment does not limit the specific values of the first loss weight and the fourth loss weight; for example, the first loss weight can be equal to 0.3, and the fourth loss weight can be equal to 0.4. The first and second partial derivatives are weighted and summed according to the first and fourth loss weights to obtain the candidate dictionary losses corresponding to the C initial interest sub-dictionaries. The C candidate dictionary losses are then multiplied by the learning rate (this embodiment does not limit the specific value of the learning rate; for example, the learning rate can be equal to 0.01) to obtain the first dictionary losses corresponding to the C initial interest sub-dictionaries.
[0120] For ease of understanding, based on the sample object information, sample content interaction sequence, sample interaction content, candidate multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model, the initial interest dictionary in the initial multi-interest maintenance sub-model is adjusted to obtain the candidate interest dictionary. The specific process can be found in the following formula (5):
[0121]
[0122] Here, we will take one of the C initial interest sub-dictionaries as an example for explanation, L IDMM L represents the first model loss of the initial multi-interest maintenance sub-model. MIRM This represents the fourth model loss of the candidate multi-interest retrieval sub-model. Let E represent the partial derivative, λ1 represent the first loss weight, λ3 represent the fourth loss weight, and E represent the partial derivative. c and Let η represent the initial sub-dictionary of interests, and η represent the learning rate. Represents a candidate interest sub-dictionary.
[0123] The specific process for determining the total loss of the second model based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest generation sub-model, candidate multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model, and the specific process for determining the total loss of the third model based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest generation sub-model, candidate multi-interest retrieval sub-model, and candidate multi-interest maintenance sub-model, can be found in the description of determining the total loss of the first model based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest generation sub-model, initial multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model. These details will not be repeated here. The total loss of the first model, the total loss of the second model, and the total loss of the third model... The total loss of the three models can be collectively referred to as the total model loss. The specific process of determining the second dictionary loss based on sample object information, sample content interaction sequence, sample interaction content, candidate multi-interest retrieval sub-model, and candidate multi-interest maintenance sub-model can be found in the description of determining the first dictionary loss based on sample object information, sample content interaction sequence, sample interaction content, candidate multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model; this will not be repeated here. The specific process of adjusting the candidate interest dictionaries based on the second dictionary loss to obtain the interest dictionary in the multi-interest maintenance sub-model can be found in the description of adjusting the initial interest dictionary in the initial multi-interest maintenance sub-model based on the first dictionary loss to obtain the candidate interest dictionary; this will not be repeated here.
[0124] Therefore, to mitigate the potential training instability caused by the simultaneous updates of three modules (i.e., three sub-models), this application embodiment can employ a three-stage training strategy. First, in the first stage, the MIRM module (i.e., the initial multi-interest retrieval sub-model) can be trained independently, while other modules remain frozen and the interest dictionary (or initial interest dictionary) remains unchanged. Then, in the second stage, the IDMM module (i.e., the initial multi-interest maintenance sub-model) can be trained independently, while other modules remain frozen until the interest dictionary (or initial interest dictionary) converges. Finally, in the third stage, the MIRM module, IDMM module, and MIPDM module can be jointly trained, allowing simultaneous updates of the interest dictionary (i.e., candidate interest dictionary). This three-stage training strategy ensures that each module is optimally adjusted in isolation, thereby achieving better coordination and stability.
[0125] To prevent interest collapse (i.e., overlapping or high similarity among multiple interest information or discrete interest vectors), this embodiment can use clustering-based initialization for each interest sub-dictionary in the first training batch to obtain an initial interest dictionary. For example, the centroids obtained by K-means clustering can be used as the initial discrete interest vectors in the initial interest sub-dictionary. Optionally, this embodiment can also initialize any one of the C initial interest sub-dictionaries using existing interest sub-dictionaries (i.e., interest sub-dictionaries trained before the current time step) to further accelerate convergence.
[0126] Optionally, based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest generation sub-model, initial multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model, the total loss of the first model is determined. Based on the total loss of the first model, the parameters of the initial multi-interest generation sub-model, initial multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model are adjusted to obtain the multi-interest retrieval sub-model, multi-interest maintenance sub-model, and multi-interest generation sub-model. Based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model, the third dictionary loss is determined. Based on the third dictionary loss, the initial interest dictionary in the initial multi-interest maintenance sub-model is adjusted to obtain the interest dictionary in the multi-interest maintenance sub-model. The specific process of determining the third dictionary loss based on sample object information, sample content interaction sequence, sample interaction content, initial multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model can be found in the description of determining the first dictionary loss based on sample object information, sample content interaction sequence, sample interaction content, candidate multi-interest retrieval sub-model, and initial multi-interest maintenance sub-model, and will not be repeated here. The specific process of adjusting the initial interest dictionary in the initial multi-interest maintenance sub-model based on the third dictionary loss to obtain the interest dictionary in the multi-interest maintenance sub-model can be found in the description of adjusting the initial interest dictionary in the initial multi-interest maintenance sub-model based on the first dictionary loss to obtain the candidate interest dictionary, and will not be repeated here.
[0127] The first dictionary loss, second dictionary loss, and third dictionary loss can be collectively referred to as dictionary loss. Optionally, based on the first model loss and the third model loss, the dictionary loss (i.e., the third dictionary loss) corresponding to each of the C initial interest sub-dictionaries is determined. The C initial interest sub-dictionaries are then updated according to the dictionary losses. The specific process of updating the C initial interest sub-dictionaries according to the dictionary losses can be found in the description of updating the C initial interest sub-dictionaries according to the first dictionary loss above, and will not be repeated here.
[0128] The following steps, S3021-S3023, describe how to determine the first model loss of the initial multi-interest maintenance sub-model, the second model loss of the initial multi-interest generation sub-model, and the third model loss of the initial multi-interest retrieval sub-model.
[0129] Step S3021: Determine the first model loss of the initial multi-interest maintenance sub-model based on the sample interaction content;
[0130] Specifically, the sample interaction content features corresponding to the sample interaction content are obtained, and these features are input into the initial multi-interest maintenance sub-model. The sample interaction content features are obtained by feature extraction of the sample interaction content through the initial content pyramid network in the initial multi-interest retrieval sub-model. The initial multi-interest maintenance sub-model includes an initial interest encoding network, an initial interest decoding network, and an initial interest dictionary. This embodiment does not limit the specific types of the initial interest encoding network and the initial interest decoding network; they are names of different stages, and the initial interest decoding network is also a name for the training stage, while the interest encoding network and interest decoding network are names for the inference stage. The initial interest dictionary includes C initial interest sub-dictionaries, and the C initial interest sub-dictionaries include M initial interest sub-dictionaries. i Initial interest sub-dictionary M i The initial interest sub-dictionary can be any one of C initial interest sub-dictionaries, where i can be a positive integer less than or equal to C. The sample interaction content features are encoded using an initial interest encoding network to obtain the encoded content features corresponding to the sample interaction content features. Vector mapping is then performed on the C initial interest sub-dictionaries based on the encoded content features to obtain the residual vector and the sample discrete interest vector corresponding to each of the C initial interest sub-dictionaries. The C sample discrete interest vectors are then concatenated (vector concatenation can better capture the multidimensionality of object interests) to obtain the sample multi-interest vector corresponding to the sample interaction content. Optionally, the C sample discrete interest vectors are fused to obtain the sample multi-interest vector corresponding to the sample interaction content. The sample multi-interest vector is then decoded using an initial interest decoding network to obtain the decoded content features corresponding to the sample multi-interest vector. The initial interest decoding network is used to reconstruct the sample multi-interest vector into sample interaction content features. Based on the sample interaction content features, decoded content features, C sample discrete interest vectors, C residual vectors, and C initial interest sub-dictionaries, the first model loss of the initial multi-interest maintenance sub-model is determined.
[0131] Specifically, based on the sample interaction content features, decoded content features, C discrete interest vectors, and C residual vectors, the encoding and decoding losses of the initial multi-interest maintenance sub-model are determined. Based on the C initial interest sub-dictionaries, the sub-regularization losses corresponding to each of the C initial interest sub-dictionaries are determined. The C sub-regularization losses are then summed to obtain the regularization loss of the initial multi-interest maintenance sub-model. Based on the encoding and decoding losses and the regularization loss, the first model loss of the initial multi-interest maintenance sub-model is determined. Optionally, the encoding and decoding losses and the regularization loss are both used as the first model loss of the initial multi-interest maintenance sub-model.
[0132] Specifically, the encoding / decoding weights corresponding to the encoding / decoding loss and the regularization weights corresponding to the regularization loss are obtained. The encoding / decoding loss and the regularization loss are then weighted and summed based on the encoding / decoding weights and the regularization weights to obtain the first model loss of the initial multi-interest maintenance sub-model. The encoding / decoding weights and regularization weights are hyperparameters that control the relative importance of the encoding / decoding loss and the regularization loss. This embodiment does not limit the specific values of the encoding / decoding weights and the regularization weights; for example, the encoding / decoding weight can be equal to 1, and the regularization weight can be equal to 0.5.
[0133] For ease of understanding, the specific process of determining the first model loss of the initial multi-interest maintenance sub-model based on the encoding / decoding loss and regularization loss can be found in the following formula (6):
[0134]
[0135] in, Indicates encoding / decoding loss. Let λ represent the regularization loss. reg This represents the regularization weight. Here, we'll use an example where the encoding / decoding weight is equal to 1. L IDMM This represents the first model loss of the initial multi-interest maintenance sub-model.
[0136] It is understandable that if the initial interest sub-dictionary M i If the first of the C initial interest sub-dictionaries is the encoded content feature, then the initial interest sub-dictionary M is determined. i The corresponding residual vector is used to obtain the encoded content features and the initial interest sub-dictionary M. i The vector distance between at least two initial discrete interest vectors in the initial interest sub-dictionary M i Obtain the initial discrete interest vector corresponding to the minimum vector distance between at least two vector distances, and determine the initial discrete interest vector corresponding to the minimum vector distance as the initial interest sub-dictionary M. i The corresponding discrete interest vector of the sample; if the initial interest sub-dictionary M iIf it is not the first of the C initial interest sub-dictionaries, then for the initial interest sub-dictionary M... i-1 The corresponding residual vector and the initial interest sub-dictionary M i-1 The initial interest sub-dictionary M is obtained by subtracting the corresponding discrete interest vectors of the samples. i The corresponding residual vector is used to obtain the initial interest sub-dictionary M. i The corresponding residual vectors are respectively related to the initial interest sub-dictionary M. i The vector distance between at least two initial discrete interest vectors in the initial interest sub-dictionary M i Obtain the initial discrete interest vector corresponding to the minimum vector distance between at least two vector distances, and determine the initial discrete interest vector corresponding to the minimum vector distance as the initial interest sub-dictionary M. i The corresponding discrete interest vector of the sample. Wherein, the initial interest sub-dictionary M i-1 For the initial interest sub-dictionary M i The previous initial interest sub-dictionary, C initial interest sub-dictionaries, allow the vector mapping to be repeated C times. For example, initial interest sub-dictionary M i-1 This can be a second initial interest sub-dictionary, initial interest sub-dictionary M i This can be a third initial interest sub-dictionary; this application does not limit the specific method of determining vector distance, for example, determining vector distance through cosine similarity, Euclidean distance, etc. Therefore, the embodiments of this application can use RQ-VAE (Residual Quantized Variational AutoEncoder) to map sample interaction content to a set of discrete interest vectors (i.e., C sample discrete interest vectors), that is, map the sample interaction content features to C initial sub-dictionaries, thereby realizing vector quantization (that is, mapping continuous vectors to a set of discrete vectors, a set of discrete vectors being a multi-dimensional discrete interest representation), effectively avoiding the unconstrained drift of continuous vectors in high-dimensional space, and to a certain extent preventing interest collapse, avoiding providing only a limited recommendation range for objects during retrieval.
[0137] For ease of understanding, the specific process of determining the sample multi-interest vector based on the sample interaction content can be found in the following formula (7):
[0138]
[0139] in, This represents the discrete interest vector of the sample corresponding to the first initial interest sub-dictionary. This represents the discrete interest vector of the samples corresponding to the second initial interest sub-dictionary, ... Let m represent the discrete interest vector of the sample corresponding to the Cth initial interest sub-dictionary. 1The m-th element of the first initial interest sub-dictionary 1 line, m 2 The m-th element of the second initial interest sub-dictionary 2 line, ..., m C The m-th element of the C-th initial interest sub-dictionary C Okay, z quan (v) and denoted as sample multi-interest vector (i.e., multi-dimensional interest embedding), v represents the sample interaction content feature corresponding to the sample interaction content, and concat represents vector concatenation.
[0140] Furthermore, the residual vector corresponding to the c-th initial interest sub-dictionary is r c-1 This represents the residual vector corresponding to the (c-1)th initial interest sub-dictionary. This represents the discrete interest vector of the sample corresponding to the (c-1)th initial interest sub-dictionary. Specifically, r1 = z enc (sg[v]) represents the residual vector corresponding to the first initial interest sub-dictionary, v represents the sample interaction content feature corresponding to the sample interaction content, and z enc Let r represent the initial interest encoding network. Here, sg[·] represents the gradient stopping operator, which is used to ensure that the gradient of r only updates the initial interest encoding network and the initial interest decoding network, and does not update the sample interaction content feature v or the sample discrete interest vector e.
[0141] For ease of understanding, the specific process of determining the encoding and decoding loss of the initial multi-interest maintenance sub-model based on the sample interaction content features, decoding content features, C sample discrete interest vectors, and C residual vectors can be found in the following formula (8):
[0142]
[0143] Where v represents the sample interaction content features, z quan (v) represents the sample multi-interest vector, z dec (z quan (v)) represents the decoded content features, r c This represents the residual vector corresponding to the c-th initial interest sub-dictionary. This represents the discrete interest vector of the sample corresponding to the c-th initial interest sub-dictionary, and sg[·] represents the gradient stopping operator. β represents the encoding / decoding loss, and β represents the hyperparameter. This application does not limit the specific value of β. For example, β can be equal to 0.1.
[0144] It is understandable that the initial interest sub-dictionary M... i The mean vector of interest embedding is obtained by performing a mean operation on at least two initial discrete interest vectors in the initial interest sub-dictionary M.i At least two initial discrete interest vectors are subtracted from the interest embedding mean vector to obtain at least two difference embedding vectors. These at least two difference embedding vectors are then used to construct a centering matrix. The transpose of the centering matrix is obtained, and the matrix is then used based on the transpose, the centering matrix, and the initial interest sub-dictionary M. i The number of initial discrete interest vectors in the dictionary determines the initial interest sub-dictionary M. i The corresponding covariance matrix. Based on the initial interest sub-dictionary M. i The corresponding covariance matrix determines the initial interest sub-dictionary M. i The corresponding sub-regularization losses are calculated by summing the C sub-regularization losses to obtain the regularization loss of the initial multi-interest maintenance sub-model. This regularization loss ensures the diversity of the initial discrete interest vectors (or discrete interest vectors within the interest sub-dictionary) in the initial interest sub-dictionary.
[0145] For ease of understanding, based on the initial interest sub-dictionary M i Given at least two initial discrete interest vectors, determine the initial interest sub-dictionary M. i The specific process for obtaining the corresponding covariance matrix can be found in the following formula (9):
[0146]
[0147] Here, we will take the c-th initial interest sub-dictionary out of the C initial interest sub-dictionaries as an example for explanation. E represents the centering matrix corresponding to the c-th initial interest sub-dictionary. c Let represent the matrix constructed from at least two initial discrete interest vectors in the c-th initial interest sub-dictionary. Let represent the mean vector of interest embeddings corresponding to the c-th initial interest sub-dictionary. M represents the transpose of the centering matrix corresponding to the c-th initial interest sub-dictionary. c Cov(E) represents the number of initial discrete interest vectors in the c-th initial interest sub-dictionary. c Let ) represent the covariance matrix corresponding to the c-th initial interest sub-dictionary.
[0148] For ease of understanding, the specific process of determining the regularization loss of the initial multi-interest maintenance sub-model based on the covariance matrices corresponding to the C initial interest sub-dictionaries can be found in the following formula (10):
[0149]
[0150] Among them, Cov(E) c Let represent the covariance matrix corresponding to the c-th initial interest sub-dictionary, ‖·‖ F Describing the Frobenius norm, This represents the regularization loss.
[0151] Therefore, since the first model loss of the initial multi-interest maintenance sub-model includes regularization loss, and the first model loss of the initial multi-interest maintenance sub-model can be used to determine the dictionary loss (e.g., the first dictionary loss), embodiments of this application can introduce a regularization term (i.e., regularization loss) into the dictionary loss to enhance the distinctiveness within each interest sub-dictionary, thereby minimizing redundancy and improving uniqueness and comprehensiveness.
[0152] Step S3022: Based on the sample object information and the sample content interaction sequence, obtain the sample interest probability distribution of the sample object; based on the sample interest probability distribution and the sample interaction content, determine the second model loss of the initial multi-interest generation sub-model.
[0153] Specifically, sample object features are obtained from sample object information, and these features, along with the sample content interaction sequence, are input into the initial multi-interest generation sub-model. Here, sample object features are a subset of sample object information; the initial multi-interest generation sub-model includes an initial generative network and an initial feature transformation network. The initial generative network learns the sequential behavior of sample objects (i.e., the sample content interaction sequence) conditioned on the sample object features. The initial generative network and generative network are names of different stages, as are the initial feature transformation network and feature transformation network. The initial generative network and initial feature transformation network are names of the training stage, and the generative network and feature transformation network are names of the inference stage. The initial multi-interest generation sub-model extracts features from the sample object features and the content in the sample content interaction sequence, respectively, to obtain the sample key object vector corresponding to the sample object features and the sample historical content vector corresponding to the content in the sample content interaction sequence. Specifically, the process involves: extracting features from the sample objects using an initial feature transformation network (also known as the first initial feature transformation network), resulting in a sample key object vector; extracting features from the content in the sample content interaction sequence using a second initial feature transformation network, resulting in a sample history content vector corresponding to the content in the sample content interaction sequence; one content item in the sample content interaction sequence corresponds to one sample history content vector, and T content items in the sample content interaction sequence correspond to T sample history content vectors. The initial multi-interest generation sub-model concatenates the sample key object vector and the sample history content vectors corresponding to the content in the sample content interaction sequence (i.e., concatenates the sample key object vector and the T sample history content vectors) to obtain the sample object content vector. Finally, the initial multi-interest generation sub-model predicts the distribution of the sample object content vector, resulting in N sample distribution prediction vectors corresponding to the multi-interest distributions. Specifically, the initial generative network predicts the distribution of the sample object content vector to obtain candidate prediction vectors. These candidate prediction vectors are then processed by a fully connected layer (the fully connected layer can be implemented using an initial fully connected layer; the initial fully connected layer is the name of the training phase, and the fully connected layer is the name of the inference phase), resulting in N sample distribution prediction vectors corresponding to multiple interest distributions. Finally, the initial multi-interest generative sub-model normalizes these sample distribution prediction vectors (normalization can be achieved using a normalized exponential function), resulting in N normalized sample distribution prediction vectors corresponding to the multiple interest distributions.The normalized sample distribution prediction vector includes the sample multi-interest distribution probabilities corresponding to N multi-interest distributions, where N is an integer greater than 1. These N sample multi-interest distribution probabilities represent the probabilities predicted by the initial multi-interest generation sub-model for the sample object to exhibit a corresponding multi-interest distribution at the next time step (i.e., time T+1) after T time steps. In other words, the initial multi-interest generation sub-model can learn the multi-interest distribution of the sample object at the next time step. The sample interaction content is obtained by considering the sample multi-interest distributions of C initial interest sub-dictionaries. The label vector corresponding to the sample interaction content is determined based on these distributions. The sample multi-interest distribution is determined by the interest information of C discrete sample interest vectors in their corresponding initial interest sub-dictionaries. That is, the sample multi-interest distribution includes the interest information of C discrete sample interest vectors in their corresponding initial interest sub-dictionaries. The label vector includes label parameters corresponding to N multi-interest distributions. The label parameter corresponding to the sample multi-interest distribution is the first label parameter (equal to 1), and the label parameters corresponding to the N-1 multi-interest distributions (i.e., the multi-interest distributions other than the sample multi-interest distribution among the N multi-interest distributions) are all second label parameters (equal to 0). The second model loss of the initial multi-interest generation sub-model is determined based on the normalized sample distribution prediction vector and the label vector. Specifically, the second model loss of the initial multi-interest generation sub-model is determined based on the loss function, the normalized sample distribution prediction vector, and the label vector; here, the cross-entropy loss function is used as an example for illustration.
[0154] For ease of understanding, the specific process of determining the second model loss of the initial multi-interest generation sub-model based on the cross-entropy loss function, the normalized sample distribution prediction vector, and the label vector can be found in the following formula (11):
[0155]
[0156] Where, p u Represents the normalized sample distribution prediction vector. Let L represent the label vector, log represent the natural logarithm, and L... MIPDM This represents the second model loss of the initial multi-interest generating sub-model.
[0157] For ease of understanding, the specific process of performing multi-interest prediction on sample objects based on sample object features and sample content interaction sequences to obtain N normalized sample distribution prediction vectors corresponding to multi-interest distributions can be found in the following formula (12):
[0158] p u =softmax(z) FC (z GPT (S′ u (12)
[0159] Here, we take the initial generative network as the GPT model as an example for illustration, S′ u =concat(z feat (u feat ),S u ) represents the sample object content vector, u feat z represents the features of the sample object. feat (u feat ) represents the key object vector of the sample. In This represents a vector of sample history content; concat indicates vector concatenation; z GPT Let z represent the initial generative network. FC This represents the initial fully connected layer, softmax represents the normalized exponential function, and p u This represents the normalized sample distribution prediction vector.
[0160] Step S3023: Determine the positive sample parameters and negative sample parameters based on the sample object information, sample content interaction sequence, and sample interaction content; determine the third model loss of the initial multi-interest retrieval sub-model based on the positive sample parameters and negative sample parameters.
[0161] Specifically, the sample multi-interest vectors corresponding to the sample interaction content, sample object information, and content from the sample content interaction sequence are input into the initial multi-interest retrieval sub-model. This initial multi-interest retrieval sub-model includes an initial object tower network, an initial content tower network, and an initial object-interest fusion network. The initial object tower network and initial content tower network can form a dual-tower model. The initial object tower network and object tower network are names of different stages; the initial content tower network and content tower network are names of different stages; the initial object-interest fusion network and object-interest fusion network are names of different stages; the initial object tower network, initial content tower network, and initial object-interest fusion network are names of the training stage; and the object tower network, content tower network, and object-interest fusion network are names of the inference stage. The initial object tower network extracts features from the sample object information to obtain the sample representation vector corresponding to the sample object information. The initial content tower network extracts features from the content in the sample content interaction sequence to obtain the sample content features corresponding to the content in the sample content interaction sequence. The initial object-interest fusion network fuses the sample representation vector and the sample multi-interest vector to obtain the sample retrieval vector. Specifically, the initial object interest fusion network concatenates the sample representation vector and the sample multi-interest vector to obtain the sample object interest concatenated vector. Then, the initial object interest fusion network fuses the sample representation vector and the sample multi-interest vector within this concatenated vector to obtain the sample retrieval vector. The feature dot product (i.e., dot product) between the sample content features and the sample retrieval vector is obtained and defined as the positive sample parameter. Content negative samples for the sample object are obtained, and the first negative sample parameter is determined based on the content negative samples, sample object information, and sample interaction content. Interest negative samples for the sample object are obtained, and the second negative sample parameter is determined based on the interest negative samples, sample object information, and sample content interaction sequence. The first and second negative sample parameters are defined as the negative sample parameters; in other words, the negative sample parameters include both the first and second negative sample parameters. The third model loss of the initial multi-interest retrieval sub-model is determined based on the positive and negative sample parameters. Specifically, the third model loss of the initial multi-interest retrieval sub-model is determined based on the loss function, positive sample parameters, and negative sample parameters. In other words, the third model loss of the initial multi-interest retrieval sub-model is determined based on the loss function, positive sample parameters, first negative sample parameters, and second negative sample parameters. Here, we take the softmax loss (i.e., log-likelihood loss) function as an example to illustrate the concept. The softmax loss function can distinguish positive samples (i.e., sample object information, sample content interaction sequence, and sample interaction content) from negative interest samples (i.e., negative interests) and negative content samples (i.e., negative items). That is, (sample object information, sample content interaction sequence, and sample interaction content) can be called positive samples.Optionally, the third model loss of the initial multi-interest retrieval sub-model can be determined based on the sample content features and the sample retrieval vector.
[0162] For ease of understanding, the specific process of determining the positive sample parameters based on the sample object information, the sample content interaction sequence, and the sample interaction content can be found in the following formula (13):
[0163]
[0164] in, Represents the sample multi-interest vector, m * Represents the sample multi-interest distribution (or the index of the sample multi-interest distribution), u represents the sample representation vector, concat represents vector concatenation, and z fusion (·) represents the initial object interest fusion network, v i This represents the sample content features corresponding to the content in the sample content interaction sequence. This represents the parameters of the positive sample.
[0165] For ease of understanding, the specific process of determining the third model loss of the initial multi-interest retrieval sub-model based on the softmax loss function, positive sample parameters, and negative sample parameters can be found in the following formula (14):
[0166]
[0167] in, Represents the parameters of positive samples. Indicates the parameters of the first negative sample. This represents the parameter of the second negative sample. Represents the sample multi-interest vector, m * Let v represent the sample multi-interest distribution (or the index of the sample multi-interest distribution), u represent the sample representation vector, and v represent the multi-interest distribution. i This represents the sample content features corresponding to the content in the sample content interaction sequence. This represents a negative sample of the content. Indicates negative samples of interest. m represents a content in a negative sample. - Let L represent a multi-interest distribution in the negative samples of interest, e denote the base, log(·) denote the logarithmic function, and L MIRM This represents the third model loss of the initial multi-interest retrieval sub-model.
[0168] The process involves: obtaining a sample content set; identifying the content in the sample content set excluding the content of the sample content interaction sequence (i.e., the content within the sample content interaction sequence) as negative samples for the sample object; or, optionally, obtaining a sample content set excluding the content historically interacted with by the sample object as negative samples for the sample object. The sample content set includes historically interacted content from objects other than the sample object. Based on the negative samples, sample object information, and sample interaction content, a first negative sample parameter is determined. (Sample object information, negative samples, and sample interaction content) can be referred to as negative samples. The specific process for determining the first negative sample parameter based on the negative samples, sample object information, and sample interaction content can be found in the description of determining positive sample parameters based on sample object information, sample content interaction sequence, and sample interaction content, and will not be repeated here.
[0169] For ease of understanding, the specific process of obtaining negative samples of the sample object can be found in the following formula (15):
[0170]
[0171] in, Let i represent the set of sample content, and let i represent the content in the sample content interaction sequence. This represents a negative sample of the content.
[0172] Specifically, a first-class set of multi-interest distributions and a second-class set of multi-interest distributions are obtained from N multi-interest distributions. The first-class set includes multi-interest distributions among the N multi-interest distributions whose prediction counts exceed a prediction count threshold. The prediction count for each of the N multi-interest distributions refers to the number of times the probability of the corresponding multi-interest distribution in the normalized sample distribution prediction vector is the probability of the maximum sample multi-interest distribution. In other words, the multi-interest distributions in the first-class set are popular multi-interest distributions. The second-class set includes multi-interest distributions randomly sampled from the N multi-interest distributions; that is, the multi-interest distributions in the second-class set are randomly sampled multi-interest distributions. To obtain the sample interaction content, based on the sample multi-interest distributions of C initial interest sub-dictionaries, the first and second types of multi-interest distribution sets are merged into a merged multi-interest distribution set. The multi-interest distributions in the merged multi-interest distribution set, excluding the sample multi-interest distributions, are identified as negative interest samples for the sample object. Optionally, the multi-interest distributions in the first type of multi-interest distribution set, excluding the sample multi-interest distributions, are identified as negative interest samples for the sample object, and the multi-interest distributions in the second type of multi-interest distribution set, excluding the sample multi-interest distributions, are also identified as negative interest samples for the sample object. Here, the sample multi-interest distributions are determined by the interest information of the C sample discrete interest vectors in their corresponding initial interest sub-dictionaries, and the negative interest samples are unrelated multi-interest distributions. The parameters of the second negative sample are determined based on the negative interest samples, sample object information, and sample content interaction sequence. Among them, (sample object information, sample content interaction sequence and interest negative sample) can be called negative sample. The specific process of determining the second negative sample parameter based on the interest negative sample, sample object information and sample content interaction sequence can be found in the description of determining the positive sample parameter based on sample object information, sample content interaction sequence and sample interaction content, which will not be repeated here.
[0173] For ease of understanding, the specific process of obtaining interest negative samples for a given sample object can be found in the following formula (16):
[0174]
[0175] in, This represents the first type of multi-interest distribution set. Let m represent the set of second-class multi-interest distributions. * This indicates the multi-interest distribution of the sample (or the index of the multi-interest distribution of the sample). This represents negative samples of interest.
[0176] Therefore, embodiments of this application can incorporate negative content (i.e., negative content samples) and negative interest (i.e., negative interest samples) into the training softmax loss function through negative sampling, thereby distinguishing between correct pairings (i.e., positive samples) and incorrect pairings (i.e., negative samples).
[0177] Step S303: Train at least a portion of the initial interest model based on the total model loss to obtain the interest model.
[0178] Interest models and interest dictionaries can be used to perform the above. Figure 2 The corresponding embodiments of steps S101-S104 and the above Figure 3 In the corresponding embodiments, steps S201-S202, and steps S101-S104 and S201-S202, enable the target object to consume the recommended target content, and then continuously fine-tune or incrementally learn the interest model and interest dictionary based on the consumed target content to ensure the accuracy of the interest model and interest dictionary in the time dimension.
[0179] Therefore, the initial interest model, sample object information, sample content interaction sequence corresponding to the sample object, and sample interaction content corresponding to the sample object are obtained. Based on the sample object information, sample content interaction sequence, and sample interaction content, at least a portion of the initial multi-interest retrieval sub-model, initial multi-interest maintenance sub-model, and initial multi-interest generation sub-model are adjusted to obtain the multi-interest retrieval sub-model, multi-interest maintenance sub-model, and multi-interest generation sub-model. Additionally, the initial interest dictionary in the initial multi-interest maintenance sub-model is adjusted to obtain the interest dictionary in the multi-interest maintenance sub-model. The multi-interest maintenance sub-model, multi-interest generation sub-model, and multi-interest retrieval sub-model are then defined as the interest model.
[0180] Therefore, this embodiment of the application can adjust the parameters of the initial multi-interest retrieval sub-model, the initial multi-interest maintenance sub-model, and the initial multi-interest generation sub-model in the initial interest model based on the sample object information, the sample content interaction sequence corresponding to the sample object, and the sample interaction content corresponding to the sample object, to obtain the multi-interest retrieval sub-model, the multi-interest maintenance sub-model, and the multi-interest generation sub-model. Furthermore, it can adjust the initial interest dictionary in the initial multi-interest maintenance sub-model to obtain the interest dictionary in the multi-interest maintenance sub-model, and determine the multi-interest retrieval sub-model, the multi-interest maintenance sub-model, and the multi-interest generation sub-model as the interest model. Further, this embodiment of the application can perform multi-interest prediction on the target object in the interest model to obtain K multi-interest distributions associated with the target object. Each of the K multi-interest distributions includes C interest information, and the C interest information in each multi-interest distribution belongs to C interest sub-dictionaries in the interest dictionary. Since the C interest sub-dictionaries correspond to interest information of different dimensions, each multi-interest distribution includes interest information of C dimensions. Thus, when performing vector lookups on the C interest sub-dictionaries based on the C interest information in each multi-interest distribution, discrete interest vectors for each multi-interest distribution across C dimensions can be obtained. The resulting multi-interest vector, obtained by concatenating the C discrete interest vectors, reflects the multiple interests of each multi-interest distribution (i.e., each multi-interest distribution reflects a combination of multiple interests, rather than relying on a single dominant interest). Therefore, this embodiment can map the target object to K multi-interest vectors, where each of the K multi-interest vectors includes C discrete interest vectors. This simultaneously covers multi-interest through multi-interest quantization (i.e., C discrete interest vectors) and multi-interest generation (i.e., K multi-interest vectors), modeling the target object's preferences. Furthermore, multi-interest recommendation determines the target content to be recommended to the target object within the content set, ensuring recommendation accuracy while improving recommendation diversity.
[0181] Please refer to the following: Figure 4 , Figure 4 This is a schematic diagram of the structure of an interest model provided in an embodiment of this application. For example... Figure 4 As shown, during the training phase, the initial interest model may include the IDMM module, the MIPDM module, and the MIRM module. The IDMM module corresponds to the initial multi-interest maintenance sub-model (i.e., the IDMM model), the MIPDM module corresponds to the initial multi-interest generation sub-model (i.e., the MIPDM model), and the MIRM module corresponds to the initial multi-interest retrieval sub-model (i.e., the MIRM model).
[0182] The initial interest dictionary in the initial multi-interest maintenance sub-model includes C initial interest sub-dictionaries. Specifically, these C initial interest sub-dictionaries can be initial interest sub-dictionary 1, initial interest sub-dictionary 2, ..., initial interest sub-dictionary C. Initial interest sub-dictionary 1 contains M1 pieces of interest information (or initial discrete interest vectors), initial interest sub-dictionary 2 contains M2 pieces of interest information (or initial discrete interest vectors), ..., initial interest sub-dictionary C contains M... C There are M1 interest information items (or initial discrete interest vectors) in the initial interest sub-dictionary 1, M2 interest information items (or initial discrete interest vectors) in the initial interest sub-dictionary 2, ..., M1 interest information items (or initial discrete interest vectors) in the initial interest sub-dictionary C. C Each piece of interest information (or initial discrete interest vector) can be combined to obtain M1*M2*…*M C An initial multi-interest vector.
[0183] like Figure 4 The example shown illustrates how to obtain the sample interaction content (i.e., content i) corresponding to a sample object. t+1 The sample interaction content refers to the content that the sample object interacted with at time T+1 (the sample interaction content is the real value). Features of the sample interaction content are extracted through the initial content pyramid network in the initial multi-interest retrieval sub-model to obtain the sample interaction content features corresponding to the sample interaction content. These sample interaction content features are then input into the initial multi-interest maintenance sub-model. The initial multi-interest maintenance sub-model includes the initial interest encoding network (i.e., Z...). enc ), Initial Interest Decoding Network (i.e., Z) dec The initial interest dictionary is used to encode the sample interaction content features through an initial interest encoding network, obtaining the encoded content features corresponding to the sample interaction content features. Based on the encoded content features, vector mapping is performed on C initial interest sub-dictionaries to obtain the residual vector and the sample discrete interest vector corresponding to each of the C initial interest sub-dictionaries. The C sample discrete interest vectors are concatenated to obtain the sample multi-interest vector corresponding to the sample interaction content. Since the sample interaction content is the true value, the sample multi-interest vector determined based on the sample interaction content is also the true value. The sample multi-interest vector is decoded through an initial interest decoding network to obtain the decoded content features corresponding to the sample multi-interest vector. Based on the sample interaction content features, decoded content features, C sample discrete interest vectors, C residual vectors, and C initial interest sub-dictionaries, the first model loss (i.e., L0) of the initial multi-interest maintenance sub-model is determined. IDMM ).
[0184] like Figure 4As shown, the sample object features and the corresponding sample content interaction sequence of the sample object are obtained, and then input into the initial multi-interest generation sub-model. The sample content interaction sequence includes the content that the sample object has historically interacted with; here, the sample content interaction sequence includes content i1 (i.e., the aforementioned...). ), content i2 (i.e., the above) ),..., content i t (i.e., the above) Taking this as an example, the initial multi-interest generation sub-model includes the initial feature transformation network (i.e., Z). feat ), initial generative network (i.e., Z) GPT ) and the initial fully connected layer (i.e., Z) FC The initial feature transformation network extracts features from the sample objects to obtain the sample key object vectors corresponding to the sample object features. The initial multi-interest generation sub-model extracts features from the content in the sample content interaction sequence to obtain the sample history content vectors corresponding to the content in the sample content interaction sequence. The initial multi-interest generation sub-model concatenates the sample key object vectors and the sample history content vectors corresponding to the content in the sample content interaction sequence to obtain the sample object content vector. The initial multi-interest generation sub-model predicts the distribution of the sample object content vectors to obtain the sample candidate prediction vectors. The initial fully connected layer processes the sample candidate prediction vectors to obtain the sample distribution prediction vectors corresponding to N multi-interest distributions. The normalized exponential function normalizes the sample distribution prediction vectors to obtain the normalized sample distribution prediction vectors corresponding to N multi-interest distributions. The sample multi-interest distributions of the sample interaction content are obtained for C initial interest sub-dictionaries. The label vectors corresponding to the sample interaction content are determined based on the sample multi-interest distributions. The second model loss (i.e., L) of the initial multi-interest generation sub-model is determined based on the normalized sample distribution prediction vectors and the label vectors. MIPDM ).
[0185] like Figure 4As shown, the sample object information is obtained, and the sample multi-interest vector corresponding to the sample interaction content, the sample object information, and the content in the sample content interaction sequence are input into the initial multi-interest retrieval sub-model. The initial multi-interest retrieval sub-model includes an initial object tower network, an initial content tower network, and an initial object interest fusion network. Here, we take the candidate content (i.e., the second candidate content) in the sample content interaction sequence as an example. The candidate content can be any content in the sample content interaction sequence. The initial object tower network extracts features from the sample object information to obtain the sample representation vector corresponding to the sample object information; the initial content tower network extracts features from the candidate content to obtain the sample content features (i.e., v) corresponding to the candidate content. The initial object interest fusion network fuses the sample representation vector and the sample multi-interest vector to obtain the sample retrieval vector (i.e., u). k The first step is to obtain the feature dot product between the sample content features and the sample retrieval vector, and use this feature dot product as the positive sample parameter. Next, content negative samples for the sample object are obtained, and the first negative sample parameter is determined based on the content negative samples, sample object information, and sample interaction content. Then, interest negative samples for the sample object are obtained, and the second negative sample parameter is determined based on the interest negative samples, sample object information, and sample content interaction sequence. Finally, the third model loss (i.e., L0) of the initial multi-interest retrieval sub-model is determined based on the positive sample parameter, the first negative sample parameter, and the second negative sample parameter. MIRM ).
[0186] Furthermore, based on the first model loss, the second model loss, and the third model loss, the total loss of the first model is determined. The first stage of the three-stage training strategy is then implemented based on this total loss. Similarly, the second and third stages of the three-stage training strategy are implemented through the steps described above, resulting in a multi-interest retrieval sub-model, a multi-interest maintenance sub-model, and a multi-interest generation sub-model. The specific processes of the second and third stages will not be elaborated here. Optionally, the parameters of the initial multi-interest generation sub-model, the initial multi-interest retrieval sub-model, and the initial multi-interest maintenance sub-model are adjusted based on the total loss of the first model to obtain the multi-interest retrieval sub-model, the multi-interest maintenance sub-model, and the multi-interest generation sub-model, thus implementing the first-stage training strategy.
[0187] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0188] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.
[0189] Please see Figure 6 ,Figure 6 This is a schematic diagram of a content recommendation device provided in an embodiment of the present application. The content recommendation device 1 includes: a vector acquisition module 11 and a content determination module 12.
[0190] The vector acquisition module 11 is used to acquire the representation vector of the target object and K multi-interest vectors associated with the target object; each multi-interest vector includes C discrete interest vectors, and each discrete interest vector corresponds to an interest dimension or interest category; K is an integer greater than or equal to 1, and C is an integer greater than 1;
[0191] Vector acquisition module 11 is used to obtain K retrieval vectors based on the representation vector and K multi-interest vectors;
[0192] Content determination module 12 is used to obtain multiple candidate contents that match each retrieval vector;
[0193] The content determination module 12 is used to determine the target content from multiple candidate contents in order to recommend the target content to the target object.
[0194] Among them, the content determination module 12 is specifically used to obtain the content vector of the content in the content set;
[0195] The content determination module 12 is specifically used to determine the similarity between the target object and the content in the content set based on the retrieval vector and the content vector;
[0196] The content determination module 12 is specifically used to determine multiple candidate contents in the content set that match the retrieval vector based on the similarity between the target object and the content in the content set.
[0197] Among them, the vector acquisition module 11 is specifically used to acquire K multi-interest distributions associated with the target object. Each multi-interest distribution includes C interest indices corresponding to C interest sub-dictionaries. The C interest sub-dictionaries correspond to different C interest dimensions. The interest sub-dictionary of each interest dimension includes multiple discrete interest vectors corresponding to multiple interest categories in that interest dimension.
[0198] The vector acquisition module 11 obtains the multi-interest vector corresponding to each multi-interest distribution based on the discrete interest vectors of the C interest indices in the corresponding interest sub-dictionary of each multi-interest distribution.
[0199] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The data processing device 2 includes: an acquisition module 21 and a prediction module 22.
[0200] The acquisition module 21 is used to acquire the object characteristics of the target object and the content interaction sequence corresponding to the target object; the content interaction sequence includes the content that the target object has interacted with in the past.
[0201] Prediction module 22 is used to predict K multi-interest distributions associated with the target object based on object features and content interaction sequences; K is an integer greater than or equal to 1.
[0202] Each of the K multi-interest distributions includes C interest indices. Each of the C interest indices in the multi-interest distribution belongs to a C interest sub-dictionary. Each of the C interest sub-dictionaries corresponds to a different C interest dimension. Each interest dimension's interest sub-dictionary includes multiple discrete interest vectors; C is an integer greater than 1.
[0203] Among them, the prediction module 22 is specifically used to obtain the interest probability distribution of the target object based on the object features and content interaction sequence. The interest probability distribution is used to characterize the probability of different multi-interest distributions.
[0204] The prediction module 22 is specifically used to sample the interest probability distribution to obtain K multiple interest distributions associated with the target object.
[0205] Among them, based on object features and content interaction sequences, the interest probability distribution of the target object is obtained using a multi-interest generation sub-model, which includes a generative model.
[0206] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application. The model training device 4 includes: a data acquisition module 41, a prediction module 42, and a training module 43.
[0207] The data acquisition module 41 is used to acquire the initial interest model, sample object information of the sample object, sample content interaction sequence corresponding to the sample object, and sample interaction content corresponding to the sample object; the sample content interaction sequence includes the content that the sample object has interacted with in the past, and the sample interaction content is the content that the sample object has interacted with after the sample content interaction sequence.
[0208] Prediction module 42 is used to obtain the sample interest probability distribution of the sample object based on the sample object information and the sample content interaction sequence, and to determine the total model loss of the initial interest model based on the sample interest probability distribution, sample interaction content, sample object information and sample content interaction sequence.
[0209] Training module 43 is used to train at least a portion of the initial interest model based on the total model loss to obtain the interest model.
[0210] The initial interest model includes an initial multi-interest generation sub-model, an initial multi-interest retrieval sub-model, and an initial multi-interest maintenance sub-model.
[0211] Prediction module 42 is specifically used to determine the first model loss of the initial multi-interest maintenance sub-model based on the sample interaction content;
[0212] Prediction module 42 is specifically used to obtain the sample interest probability distribution of the sample object based on the sample object information and the sample content interaction sequence, and to determine the second model loss of the initial multi-interest generation sub-model based on the sample interest probability distribution and the sample interaction content.
[0213] Prediction module 42 is specifically used to determine positive sample parameters and negative sample parameters based on sample object information, sample content interaction sequence and sample interaction content, and to determine the third model loss of the initial multi-interest retrieval sub-model based on the positive sample parameters and negative sample parameters.
[0214] The prediction module 42 is specifically used to determine the total model loss of the initial interest model based on the first model loss, the second model loss, and the third model loss.
[0215] Among them, the prediction module 42 is specifically used to input the sample interaction content features corresponding to the sample interaction content into the initial multi-interest maintenance sub-model; the initial multi-interest maintenance sub-model includes an initial interest encoding network, an initial interest decoding network and C initial interest sub-dictionaries;
[0216] Prediction module 42 is specifically used to encode the sample interaction content features through the initial interest coding network to obtain the encoded content features corresponding to the sample interaction content features;
[0217] Prediction module 42 is specifically used to perform vector mapping on C initial interest sub-dictionaries according to the features of the encoded content, to obtain the residual vector and the sample discrete interest vector corresponding to each of the C initial interest sub-dictionaries, and to concatenate the C sample discrete interest vectors to obtain the sample multi-interest vector corresponding to the sample interaction content.
[0218] Prediction module 42 is specifically used to decode the sample multi-interest vector through the initial interest decoding network to obtain the decoded content features corresponding to the sample multi-interest vector;
[0219] The prediction module 42 is specifically used to determine the first model loss of the initial multi-interest maintenance sub-model based on C initial interest sub-dictionaries, sample interaction content features, decoded content features, C sample discrete interest vectors, and C residual vectors.
[0220] Among them, the prediction module 42 is specifically used to input the sample object features and sample content interaction sequence in the sample object information into the initial multi-interest generation sub-model;
[0221] Prediction module 42 is specifically used to concatenate the sample key object vector corresponding to the sample object features and the sample historical content vector corresponding to the content in the sample content interaction sequence to obtain the sample object content vector.
[0222] The prediction module 42 is specifically used to predict the distribution of the sample object content vector to obtain the sample interest probability distribution; the sample interest probability distribution is used to characterize the probability of different multi-interest distributions.
[0223] The prediction module 42 is specifically used to obtain the sample multi-interest distribution of the sample interaction content, and to determine the second model loss of the initial multi-interest generation sub-model based on the sample multi-interest distribution and the sample interest probability distribution.
[0224] Among them, the prediction module 42 is specifically used to input the sample multi-interest vector, sample object information and sample content interaction sequence corresponding to the sample interaction content into the initial multi-interest retrieval sub-model;
[0225] Prediction module 42 is specifically used to perform feature fusion on the sample representation vector and sample multi-interest vector corresponding to the sample object information to obtain the sample retrieval vector;
[0226] Prediction module 42 is specifically used to determine the positive sample parameters by the feature dot product between the sample content features corresponding to the content in the sample content interaction sequence and the sample retrieval vector.
[0227] The prediction module 42 is specifically used to obtain content negative samples and interest negative samples for the sample object, and to determine the negative sample parameters based on the content negative samples, interest negative samples, sample object information, sample content interaction sequence and sample interaction content.
[0228] Among them, the training module 43 is also used to determine the dictionary loss corresponding to each of the C initial interest sub-dictionaries based on the first model loss and the third model loss;
[0229] Training module 43 is also used to update the C initial interest sub-dictionaries respectively based on dictionary loss.
[0230] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0231] Please see Figure 9 , Figure 9This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device 3 includes a processor 31 and a memory 32. Optionally, the electronic device 3 also includes an input device 33 and an output device 34. The processor 31, memory 32, input device 33, and output device 34 are coupled together via connectors, which include various interfaces, transmission lines, or buses, etc., and are not limited in this embodiment. It should be understood that in the various embodiments of this application, coupling refers to mutual connection in a specific way, including direct connection or indirect connection through other devices, such as through various interfaces, transmission lines, buses, etc.
[0232] Processor 31 may include one or more processors, such as one or more central processing units (CPUs). If the processor is a CPU, it may be a single-core CPU or a multi-core CPU. Optionally, processor 31 may be a processor group consisting of multiple CPUs, with the multiple processors coupled to each other via one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in the embodiments of this application.
[0233] The memory 32 can be used to store computer program instructions, as well as various types of computer program code, including program code for executing the scheme of this application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.
[0234] Input device 33 is used to input data and / or signals, and output device 34 is used to output data and / or signals. Input device 33 and output device 34 can be independent devices or an integrated device.
[0235] It is understood that in this embodiment of the application, the memory 32 can be used not only to store related instructions, but also to store related data. This embodiment of the application does not limit the specific data stored in the memory.
[0236] Understandable Figure 9This is merely a simplified design of an electronic device. In practical applications, the electronic device may also include other necessary components, including, but not limited to, any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of this application are within the protection scope of this application.
[0237] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0238] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will also readily understand that the various embodiments of this application have different focuses, and for the sake of convenience and brevity, the same or similar parts may not be repeated in different embodiments. Therefore, parts not described or not described in detail in one embodiment can be referred to the descriptions in other embodiments.
[0239] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0240] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0241] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0242] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0243] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data processing method, characterized in that, The method includes: Obtain the object characteristics of the target object and the content interaction sequence corresponding to the target object; the content interaction sequence includes the content that the target object has historically interacted with. Based on the object features and the content interaction sequence, K multi-interest distributions associated with the target object are predicted; where K is an integer greater than or equal to 1. Each of the K multi-interest distributions includes C interest indices. The C interest indices in each multi-interest distribution belong to C interest sub-dictionaries. The C interest sub-dictionaries correspond to different C interest dimensions. Each interest dimension's interest sub-dictionary includes multiple discrete interest vectors. C is an integer greater than 1.
2. The method according to claim 1, characterized in that, The step of predicting K multi-interest distributions associated with the target object based on the object features and the content interaction sequence includes: Based on the object features and the content interaction sequence, the interest probability distribution of the target object is obtained, and the interest probability distribution is used to characterize the probability of different multi-interest distributions; The interest probability distribution is sampled to obtain K multiple interest distributions associated with the target object.
3. The method according to claim 2, characterized in that, The interest probability distribution of the target object is obtained by using a multi-interest generation sub-model based on the object features and the content interaction sequence. The multi-interest generation sub-model includes a generative model.
4. A model training method, characterized in that, The method includes: Obtain the initial interest model, sample object information of the sample object, sample content interaction sequence corresponding to the sample object, and sample interaction content corresponding to the sample object; the sample content interaction sequence includes the content that the sample object has interacted with in the past, and the sample interaction content is the content that the sample object has interacted with after the sample content interaction sequence; Based on the sample object information and the sample content interaction sequence, the sample interest probability distribution of the sample object is obtained. Based on the sample interest probability distribution, the sample interaction content, the sample object information, and the sample content interaction sequence, the total model loss of the initial interest model is determined. At least a portion of the initial interest model is trained based on the total loss of the model to obtain an interest model.
5. The method according to claim 4, characterized in that, The initial interest model includes an initial multi-interest generation sub-model, an initial multi-interest retrieval sub-model, and an initial multi-interest maintenance sub-model; The step of obtaining the sample interest probability distribution of the sample object based on the sample object information and the sample content interaction sequence, and determining the total model loss of the initial interest model based on the sample interest probability distribution, the sample interaction content, the sample object information, and the sample content interaction sequence, includes: Based on the sample interaction content, determine the first model loss of the initial multi-interest maintenance sub-model; Based on the sample object information and the sample content interaction sequence, the sample interest probability distribution of the sample object is obtained, and based on the sample interest probability distribution and the sample interaction content, the second model loss of the initial multi-interest generation sub-model is determined. Based on the sample object information, the sample content interaction sequence, and the sample interaction content, positive sample parameters and negative sample parameters are determined, and based on the positive sample parameters and negative sample parameters, the third model loss of the initial multi-interest retrieval sub-model is determined; The total model loss of the initial interest model is determined based on the first model loss, the second model loss, and the third model loss.
6. The method according to claim 5, characterized in that, The step of determining the first model loss of the initial multi-interest maintenance sub-model based on the sample interaction content includes: The sample interaction content features corresponding to the sample interaction content are input into the initial multi-interest maintenance sub-model; the initial multi-interest maintenance sub-model includes an initial interest encoding network, an initial interest decoding network, and C initial interest sub-dictionaries; The sample interaction content features are encoded by the initial interest coding network to obtain the encoded content features corresponding to the sample interaction content features. Based on the encoded content features, the C initial interest sub-dictionaries are vector-mapped to obtain the residual vector and the sample discrete interest vector corresponding to each of the C initial interest sub-dictionaries. The C sample discrete interest vectors are then concatenated to obtain the sample multi-interest vector corresponding to the sample interaction content. The sample multi-interest vector is decoded by the initial interest decoding network to obtain the decoded content features corresponding to the sample multi-interest vector. The first model loss of the initial multi-interest maintenance sub-model is determined based on the C initial interest sub-dictionaries, the sample interaction content features, the decoded content features, the C sample discrete interest vectors, and the C residual vectors.
7. The method according to claim 5, characterized in that, The step of obtaining the sample interest probability distribution of the sample object based on the sample object information and the sample content interaction sequence, and determining the second model loss of the initial multi-interest generation sub-model based on the sample interest probability distribution and the sample interaction content, includes: The sample object features and the sample content interaction sequence in the sample object information are input into the initial multi-interest generation sub-model; The sample object content vector is obtained by concatenating the sample key object vector corresponding to the sample object feature and the sample history content vector corresponding to the content in the sample content interaction sequence. The sample object content vector is subjected to distribution prediction to obtain the sample interest probability distribution; the sample interest probability distribution is used to characterize the probability of different multi-interest distributions. Obtain the sample multi-interest distribution of the sample interaction content, and determine the second model loss of the initial multi-interest generation sub-model based on the sample multi-interest distribution and the sample interest probability distribution.
8. The method according to claim 5, characterized in that, The step of determining positive and negative sample parameters based on the sample object information, the sample content interaction sequence, and the sample interaction content includes: The sample multi-interest vector corresponding to the sample interaction content, the sample object information, and the content in the sample content interaction sequence are input into the initial multi-interest retrieval sub-model; The sample representation vector corresponding to the sample object information and the sample multi-interest vector are fused to obtain the sample retrieval vector; The feature dot product between the sample content features corresponding to the content in the sample content interaction sequence and the sample retrieval vector is determined as the positive sample parameter; Obtain content negative samples and interest negative samples for the sample object, and determine the negative sample parameters based on the content negative samples, the interest negative samples, the sample object information, the sample content interaction sequence, and the sample interaction content.
9. The method according to claim 6, characterized in that, The method further includes: Based on the first model loss and the third model loss, determine the dictionary loss corresponding to each of the C initial interest sub-dictionaries; The C initial interest sub-dictionaries are updated according to the dictionary loss.
10. A content recommendation method, characterized in that, The method includes: Obtain the representation vector of the target object and K multi-interest vectors associated with the target object; each multi-interest vector includes C discrete interest vectors, each discrete interest vector corresponding to an interest dimension or interest category; K is an integer greater than or equal to 1, and C is an integer greater than 1; Based on the representation vector and the K multiple interest vectors, K retrieval vectors are obtained; Retrieve multiple candidate contents that match each retrieval vector; The target content is determined from the plurality of candidate contents in order to recommend the target content to the target object.
11. The method according to claim 10, characterized in that, The step of obtaining multiple candidate contents that match each retrieval vector includes: Retrieve the content vector of the content in the content collection; Based on the retrieval vector and the content vector, the similarity between the target object and the content in the content set is determined; Based on the similarity between the target object and the content in the content set, multiple candidate contents that match the retrieval vector are determined in the content set.
12. The method according to claim 10, characterized in that, The process of obtaining the representation vector of the target object and the K multiple interest vectors associated with the target object includes: Obtain K multi-interest distributions associated with the target object. Each multi-interest distribution includes C interest indices corresponding to C interest sub-dictionaries. The C interest sub-dictionaries correspond to different C interest dimensions. Each interest dimension's interest sub-dictionary includes multiple discrete interest vectors corresponding to multiple interest categories in that interest dimension. Based on the discrete interest vectors of the C interest indices in each multi-interest distribution in the corresponding interest sub-dictionary, the multi-interest vector corresponding to each multi-interest distribution is obtained.
13. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the object characteristics of the target object and the content interaction sequence corresponding to the target object; the content interaction sequence includes the content that the target object has interacted with in the past. The prediction module is used to predict K multi-interest distributions associated with the target object based on the object features and the content interaction sequence; where K is an integer greater than or equal to 1. Each of the K multi-interest distributions includes C interest indices. The C interest indices in each multi-interest distribution belong to C interest sub-dictionaries. The C interest sub-dictionaries correspond to different C interest dimensions. Each interest dimension's interest sub-dictionary includes multiple discrete interest vectors. C is an integer greater than 1.
14. A model training device, characterized in that, The device includes: The data acquisition module is used to acquire an initial interest model, sample object information of sample objects, sample content interaction sequence corresponding to the sample objects, and sample interaction content corresponding to the sample objects; the sample content interaction sequence includes the content that the sample objects have interacted with in the past, and the sample interaction content is the content that the sample objects have interacted with after the sample content interaction sequence; The prediction module is used to obtain the sample interest probability distribution of the sample object based on the sample object information and the sample content interaction sequence, and to determine the total model loss of the initial interest model based on the sample interest probability distribution, the sample interaction content, the sample object information and the sample content interaction sequence. The training module is used to train at least a portion of the initial interest model based on the total loss of the model, so as to obtain an interest model.
15. A content recommendation device, characterized in that, The device includes: The vector acquisition module is used to acquire the representation vector of the target object and K multi-interest vectors associated with the target object; each multi-interest vector includes C discrete interest vectors, each discrete interest vector corresponding to an interest dimension or interest category; K is an integer greater than or equal to 1, and C is an integer greater than 1. The vector acquisition module is used to obtain K retrieval vectors based on the representation vector and the K multi-interest vectors; The content determination module is used to obtain multiple candidate contents that match each retrieval vector; The content determination module is used to determine target content from the plurality of candidate contents, so as to recommend the target content to the target object.
16. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1 to 12.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 12.
18. A computer program product, characterized in that, The computer program product includes a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Video recommendation method and device and electronic equipment
CN116366923A
Recommendation method, device and equipment
CN116561431A
Recall model training method, object recall method, content recommendation method and content recommendation system
CN120256958A
Method and system for recommending content
US20170193106A1