Model training method, embedding generation method, recommendation method and related products

By using the state space model to train the language model, using the embedding sequence of objects and objects to generate predictive embeddings, the problem of insufficient ability of the language model to predict object embeddings in the recommended scenario is solved, and more efficient and accurate prediction is achieved.

CN120541520APending Publication Date: 2025-08-26XINGIN INFORMATION TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604057.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

How to train language models to make them have the ability to predict the embedding of items that interact with objects to improve the accuracy and efficiency of recommendations.

Method used

The language model is trained using the state space model, using the embedding sequences of objects and objects to generate predictive embeddings, and updating model parameters through positive samples and loss functions to improve the model's prediction ability.

Benefits of technology

This reduces the computational complexity of the model, improves the speed and accuracy of generating prediction embeddings, and enhances the prediction ability of the model in recommended scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541520A_ABST
    Figure CN120541520A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an embedding generation method, a recommendation method and related products. The model training method comprises the steps that a first language model is acquired, and the first language model comprises a state space model; a first sequence of first item embedding is obtained, the first item embedding including embedding of first interaction items that have interacted with a first object. And inputting the first sequence into a first language model, so that the first language model generates a first prediction embedding, and the item represented by the first prediction embedding is an item interacting with the first object after the first interaction item. A first loss of the first language model is determined based on the first predicted embedding and a first positive sample, the first positive sample including an item actually interacting with the first object after the first interacting item. And based on the first loss, updating parameters of the first language model to obtain a target language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, an embedding generation method, a recommendation method, and related products. Background Art

[0002] In the recommendation scenario, by predicting the items that the object will interact with and recommending the items to the object, the probability of the object interacting with the recommended items can be increased, and the items recommended to the object can be determined, thereby improving the object's experience.

[0003] Thanks to their powerful performance, language models are increasingly being used in recommendation scenarios, including predicting items that interact with an object based on language models. Specifically, language models can be used to predict the embeddings of items that interact with an object, and then the items that interact with the object can be determined based on these embeddings. Therefore, how to train language models to predict the embeddings of items that interact with an object is an urgent problem that needs to be solved. Summary of the Invention

[0004] The present application provides a model training method, an embedding generation method, a recommendation method and related products, wherein the related products include a model training device, an embedding generation device, a recommendation device, an electronic device, a computer-readable storage medium and a computer program product.

[0005] In a first aspect, a model training method is provided, the method comprising:

[0006] Obtaining a first language model, wherein the first language model includes a state space model (SSM);

[0007] Obtaining a first sequence of first item embeddings, wherein the first item embeddings include embeddings of first interactive items that have interacted with a first object;

[0008] Inputting the first sequence into the first language model so that the first language model generates a first predicted embedding, wherein the first predicted embedding represents an item that interacts with the first object after the first interactive item;

[0009] determining a first loss for the first language model based on the first predicted embedding and a first positive sample, the first positive sample comprising an item that actually interacted with the first object after the first interactive item;

[0010] Based on the first loss, parameters of the first language model are updated to obtain a target language model.

[0011] In combination with any embodiment of the present application, obtaining the first language model includes:

[0012] inputting the first item embeddings in the first sequence into a second language model in sequence, so that the second language model generates a second predicted embedding based on the input first item embeddings, where the item represented by the second predicted embedding is an item that interacts with the first object after the first interactive item input into the second language model;

[0013] determining a second loss for the second language model based on the second predicted embedding and a second positive sample, the second positive sample including an item that actually interacted with the first object after the first interactive item was input to the second language model;

[0014] Based on the second loss, parameters of the second language model are updated to obtain the first language model.

[0015] In conjunction with any embodiment of the present application, the time when the first interactive object corresponding to the embedding of the first object interacts with the first object is the first interaction time;

[0016] The step of sequentially inputting the first item embeddings in the first sequence into a second language model so that the second language model generates a second predicted embedding based on the input first item embeddings includes:

[0017] The first item embeddings in the first sequence are sequentially input into a second language model in order of the first interaction time from earliest to latest, so that the second language model generates a second predicted embedding based on the input first item embeddings.

[0018] In conjunction with any embodiment of the present application, inputting the first item embedding in the first sequence into the first language model so that the first language model generates a first predicted embedding includes:

[0019] The first item embeddings in the first sequence are sequentially input into the first language model in order of the first interaction time from earliest to latest, so that the first language model generates the first predicted embedding.

[0020] In combination with any embodiment of the present application, the first positive sample includes an object that interacts with the first object when exposed to the first object;

[0021] The determining, based on the first predicted embedding and the first positive sample, a first loss of the first language model includes:

[0022] The first loss is determined based on the first predicted embedding, the first positive sample, and a negative sample, where the negative sample includes an item that does not interact with the first object when exposed to the first object.

[0023] In combination with any embodiment of the present application, the first language model also includes an attention mechanism.

[0024] In a second aspect, a method for generating an embedding is provided, the method comprising:

[0025] Obtaining a second sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object;

[0026] Inputting the second sequence into a target language model so that the target language model generates a target prediction embedding based on the second item embedding, wherein the target language model is trained based on the first aspect and any embodiment thereof, and the item represented by the target prediction embedding is an item that interacts with the second object after the second interactive item;

[0027] Based on the target prediction embedding, an object embedding of the second object is determined, and the object embedding is used to determine an item that interacts with the second object after the second interactive item.

[0028] A third aspect provides a recommendation method, the method comprising:

[0029] Obtaining an object embedding of a second object, wherein the object embedding is obtained based on the method provided by the second aspect;

[0030] determining a target embedding that matches the object embedding from candidate embeddings, the candidate embeddings being embeddings of candidate items;

[0031] The candidate item corresponding to the target embedding is determined to be an item to be recommended to the second object.

[0032] In a fourth aspect, a model training device is provided, comprising:

[0033] an acquiring unit, configured to acquire a first language model, wherein the first language model includes a state space model;

[0034] The acquisition unit is further configured to acquire a first sequence of first item embeddings, where the first item embeddings include embeddings of first interactive items that have interacted with the first object;

[0035] a generating unit, configured to input the first sequence into the first language model, so that the first language model generates a first predicted embedding, wherein the first predicted embedding represents an item that interacts with the first object after the first interactive item;

[0036] a determining unit, configured to determine a first loss of the first language model based on the first predicted embedding and a first positive sample, the first positive sample comprising an item that actually interacted with the first object after the first interactive item;

[0037] An updating unit is configured to update parameters of the first language model based on the first loss to obtain a target language model.

[0038] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0039] inputting the first item embeddings in the first sequence into a second language model in sequence, so that the second language model generates a second predicted embedding based on the input first item embeddings, where the item represented by the second predicted embedding is an item that interacts with the first object after the first interactive item input into the second language model;

[0040] determining a second loss for the second language model based on the second predicted embedding and a second positive sample, the second positive sample including an item that actually interacted with the first object after the first interactive item was input to the second language model;

[0041] Based on the second loss, parameters of the second language model are updated to obtain the first language model.

[0042] In conjunction with any embodiment of the present application, the time when the first interactive object corresponding to the embedding of the first object interacts with the first object is the first interaction time;

[0043] The acquisition unit is configured to:

[0044] The first item embeddings in the first sequence are sequentially input into a second language model in order of the first interaction time from earliest to latest, so that the second language model generates a second predicted embedding based on the input first item embeddings.

[0045] In combination with any embodiment of the present application, the generating unit is configured to:

[0046] The first item embeddings in the first sequence are sequentially input into the first language model in order of the first interaction time from earliest to latest, so that the first language model generates the first predicted embedding.

[0047] In combination with any embodiment of the present application, the first positive sample includes an object that interacts with the first object when exposed to the first object;

[0048] The determining unit is configured to:

[0049] The first loss is determined based on the first predicted embedding, the first positive sample, and a negative sample, where the negative sample includes an item that does not interact with the first object when exposed to the first object.

[0050] In combination with any embodiment of the present application, the first language model also includes an attention mechanism.

[0051] In a fifth aspect, an embedding generation device is provided, the embedding generation device comprising:

[0052] an acquiring unit, configured to acquire a second sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object;

[0053] a generating unit, configured to input the second sequence into a target language model, so that the target language model generates a target prediction embedding based on the second item embedding, wherein the target language model is trained based on the first aspect and any embodiment thereof, and the item represented by the target prediction embedding is an item that interacts with the second object after the second interactive item;

[0054] A determining unit is configured to determine an object embedding of the second object based on the target prediction embedding, wherein the object embedding is used to determine an item that interacts with the second object after the second interactive item.

[0055] In a sixth aspect, a recommendation device is provided, the recommendation device comprising:

[0056] an acquiring unit, configured to acquire an object embedding of a second object, wherein the object embedding is obtained based on the method provided by the second aspect;

[0057] a determining unit for determining a target embedding that matches the object embedding from candidate embeddings, the candidate embeddings being embeddings of candidate items;

[0058] The determining unit is further configured to determine that the candidate item corresponding to the target embedding is the item to be recommended to the second object.

[0059] In the seventh aspect, an electronic device is provided, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code comprising computer instructions, and when the processor executes the computer instructions, the electronic device executes the first aspect and any embodiment thereof, or the electronic device executes the method of the second aspect, or the electronic device executes the method of the third aspect.

[0060] In an eighth aspect, another electronic device is provided, comprising: a processor, a sending device, an input device, an output device and a memory, wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any embodiment thereof, or the electronic device executes the method of the second aspect, or the electronic device executes the method of the third aspect.

[0061] In the ninth aspect, a computer-readable storage medium is provided, in which a computer program is stored. The computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the first aspect and any embodiment thereof, or the processor is caused to execute the method of the second aspect, or the processor is caused to execute the method of the third aspect.

[0062] In the tenth aspect, a computer program product is provided, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the above-mentioned first aspect and any embodiment thereof, or the computer is caused to execute the method of the above-mentioned second aspect, or the computer is caused to execute the method of the above-mentioned third aspect.

[0063] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0064] In an embodiment of the present application, the first sequence is a sequence of first item embeddings, wherein the first item embeddings include the embeddings of the first interactive item that interacted with the first object. Therefore, the first item embeddings in the first sequence include information about the first interactive item. The first sequence is input into a first language model, so that the first language model processes the first sequence and uses the information from the first item embeddings to generate a first predicted embedding. The first predicted embedding represents an item that interacted with the first object after the first interactive item. Because the first language model includes a state-space model, processing the first sequence by the first language model can reduce computational complexity, thereby increasing the speed of generating the first predicted embedding. A first loss for the first language model is then determined based on the first predicted embeddings and the first positive sample. Based on the first loss, the parameters of the first language model are updated to obtain a target language model. This can increase the speed of obtaining the target language model through training and enable the target language model to predict the embedding of the next item that will interact with the object based on the sequence of embeddings of the items that the object has interacted with. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0066] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0067] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;

[0068] Figure 2 A schematic diagram of the model structure of an attention mechanism model provided in an embodiment of the present application;

[0069] Figure 3 A schematic diagram of the model structure of another attention mechanism provided in an embodiment of the present application;

[0070] Figure 4 A schematic diagram of the structure of a state space model provided in an embodiment of the present application;

[0071] Figure 5 A schematic diagram of the structure of a first language model provided in an embodiment of the present application;

[0072] Figure 6 A schematic diagram of the structure of a Mamba hybrid expert layer provided in an embodiment of the present application;

[0073] Figure 7A schematic diagram of an embodiment of the present application providing a method for compressing a first interactive item using a third language model to obtain a first item embedding;

[0074] Figure 8 A schematic diagram of a first language model provided in an embodiment of the present application processing a first sequence to generate a first predictive embedding;

[0075] Figure 9 A schematic diagram of a second language model provided in an embodiment of the present application processing a first sequence to generate a second predictive embedding;

[0076] Figure 10 A schematic diagram of a flow chart of an embedding generation method provided in an embodiment of the present application;

[0077] Figure 11 A flowchart of a recommended method provided in an embodiment of the present application;

[0078] Figure 12 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0079] Figure 13 A schematic structural diagram of an embedding generation device provided in an embodiment of the present application;

[0080] Figure 14 A schematic diagram of the structure of a recommended device provided in an embodiment of the present application;

[0081] Figure 15 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0082] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0083] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0084] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0085] The embodiments of the present application provide a model training method, based on which a target language model can be trained, wherein the target language model is used to predict the embedding of an item that will interact with an object next, based on items that have interacted with the object. The model training method is performed by a model training device, wherein the model training device can be any electronic device that can execute the technical solution disclosed in the embodiments of the method of the present application. Optionally, the model training device can be one of the following: a computer or a server.

[0086] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The following describes the embodiment of the present application in conjunction with the drawings in the embodiment of the present application. Figure 1 , Figure 1 A flowchart of a model training method provided in an embodiment of the present application.

[0087] 101. Obtain a first language model, where the first language model includes a state-space model.

[0088] In the embodiments of the present application, the language model (including the first language model, the target language model, and the second language model to be mentioned below) is a natural language model, that is, the language model has natural language processing (NLP) capabilities. Optionally, the language model is a large language model (LLM).

[0089] In the embodiments of the present application, a state-space model is a mathematical model used to describe dynamic systems. The inclusion of the state-space model in the first language model means that the model structure of the first language model includes the state-space model. When the first language model includes the state-space model, after data is input into the first language model, the state-space model can be used to process the input data. When the input data includes a sequence, processing the input data using the state-space model has low computational complexity.

[0090] Optionally, for processing input data including sequences, the computational complexity of the state-space model is lower than that of the attention mechanism-based model (hereinafter referred to as the attention mechanism model for ease of expression).

[0091] Optional, Figure 2 A schematic diagram of the model structure of an attention mechanism model provided in an embodiment of the present application. Figure 2 The attention mechanism model shown is a transformer model. Figure 2 As shown, the attention mechanism model includes normalization, attention mechanism, and multilayer perceptron (MLP). Specifically, after the data is input into the attention mechanism model, the input data is first normalized, and then the normalized data is processed using the attention mechanism, and then the processed data is spliced ​​with the input data to obtain first spliced ​​data. The first spliced ​​data is then normalized, and then the normalized data is processed using MLP to obtain first perceptron output data, and then the first spliced ​​data and the first perceptron output data are spliced ​​to obtain the output of the attention mechanism model. Optionally, normalization includes root mean square layer normalization (RMSNorm).

[0092] Optional, Figure 3 This is a schematic diagram of the model structure of another attention mechanism provided in the embodiment of this application. Figure 3 As shown, the attention mechanism model includes normalization, attention mechanism, and mixture of experts (MoE). Specifically, after the data is input into the attention mechanism model, the input data is first normalized, and then the normalized data is processed using the attention mechanism. The processed data is then concatenated with the input data to obtain first concatenated data. The first concatenated data is then normalized, and then the normalized data is processed using MoE to obtain expert output data. The first concatenated data and the expert output data are then concatenated to obtain the output of the attention mechanism model. Optionally, normalization includes RMS Norm.

[0093] Specifically, for an input data containing a sequence of n characters (tokens), the embedding matrix Where d represents the dimension of the embedding. The attention mechanism model processes the input data as follows:

[0094] Q=XW q ,K=XW k ,V=XW v …Formula (1)

[0095]

[0096] Among them, X represents the embedding matrix of the input data, and Y represents the embedding matrix obtained by processing the embedding of the input data by the attention mechanism model. That is W v The dimension of W is d×d. q Represents the transformation matrix corresponding to the query in the attention mechanism model, W k Represents the transformation matrix corresponding to the key in the attention mechanism model, W v W represents the transformation matrix corresponding to the value in the attention mechanism model. q 、W k 、W v are all learnable parameters. For the sake of simplicity, formulas (1) and (2) represent the processing of the attention mechanism in the attention mechanism model, and do not represent the processing of the attention mechanism model other than the attention mechanism, for example, the processing of the MLP in the attention mechanism model is not represented.

[0097] Take the last row of the embedding matrix Y as the n+1th character predicted by the attention mechanism model through the following formula:

[0098] y n+1 =Y [n,:] …Formula (3)

[0099] Among them, y n+1 Indicates the n+1th character, Y [n,:] Represents the last row of the embedding matrix Y.

[0100] From formula (1), formula (2) and formula (3), we can see that the main computational complexity of the attention mechanism model is located in QK T Part, specifically, the computational complexity here is O(n 2 ). It can be seen that when the sequence length n of the input data becomes larger, the computational complexity of the attention mechanism model will increase significantly.

[0101] The state space model can be expressed as follows:

[0102]

[0103] y n+1 =Ch n+1 …Formula (5)

[0104] in, C are all learnable parameters. n is the nth row of the embedding matrix X of the input data, h n is the hidden state, y n+1 Indicates the n+1th character. In formula (4), h nIt can be obtained by recursively calculating the formula of the first row, and then in formula (5) according to the second row of the embedding matrix X and h n You can get y n+1 .

[0105] Since in formula (4) and formula (5), h n and x n They are all one-dimensional row vectors, so the computational complexity of the state-space model is O(n), that is, the computational complexity of the state-space model is linearly related to the sequence length n of the input data.

[0106] By comparison, for input data with a sequence length of n, the computational complexity of the state space model is O(n), while the computational complexity of the attention mechanism model is O(n). 2 ). Therefore, the computational complexity of the state-space model is lower than that of the attention mechanism model.

[0107] Optional, Figure 4 This is a schematic diagram of the structure of a state space model provided in an embodiment of the present application. Figure 4 As shown, the state space model includes normalization, Mamba model (Mamba), MLP (i.e. Figure 4 Specifically, after inputting data into the state-space model, the input data is first normalized, then processed using the Mamba model, and then the processed data is concatenated with the input data to obtain second concatenated data. The second concatenated data is then normalized, then processed using the MLP to obtain second perceptron output data, and the second concatenated data and the second perceptron output data are concatenated to obtain the output of the state-space model. Optionally, the normalization includes RMS Norm.

[0108] Optional, Figure 5 This is a schematic diagram of the structure of a first language model provided in an embodiment of the present application. Figure 5 As shown in the figure, the first language model includes: 3 state-space models, 4 Mamba mixture of experts layers, and 1 attention mechanism model. Specifically, after the data is input into the first language model, it is processed in sequence by 1 state-space model, 1 Mamba mixture of experts layer, 1 state-space model, 1 Mamba mixture of experts layer, 1 attention mechanism model, 1 Mamba mixture of experts layer, 1 state-space model, and 1 Mamba mixture of experts layer to obtain the output of the first language model.

[0109] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a Mamba hybrid expert layer provided in an embodiment of the present application. Figure 6As shown, the Mamba hybrid expert layer includes normalization, a Mamba model, and a hybrid expert model. Specifically, after the data is input into the Mamba hybrid expert layer, the input data is first normalized, then the normalized data is processed using the Mamba model, and then the processed data is concatenated with the input data to obtain third concatenated data. The third concatenated data is then normalized, then the normalized data is processed using the hybrid expert model to obtain third perceptron output data, and then the third concatenated data and the third perceptron output data are concatenated to obtain the output of the Mamba hybrid expert layer. Optionally, normalization includes RMS Norm.

[0110] 102. Obtain a first sequence of first item embeddings, wherein the first item embeddings include embeddings of first interactive items that have interacted with a first object.

[0111] In the embodiments of the present application, objects (including the first object and the second object mentioned below) can be any object. Articles (including) can be any item. Optionally, the items include: multimedia content, wherein the multimedia content includes one or more of the following: images, audio, text, and video. For example, multimedia content includes text. For another example, multimedia content can include text and images. For another example, multimedia content includes audio, text, and video.

[0112] A first interactive item is an item that has interacted with the first object, meaning that an interaction has occurred between the first object and the first interactive item. Optionally, the interaction between the object and the item indicates the object's interest in the item. Optionally, interactive behaviors include: clicking, liking, commenting, forwarding, and collecting. Accordingly, items clicked by the first object, liked by the first object, commented on by the first object, forwarded by the first object, and collected by the first object are all first interactive items. The embedding of a first interactive item is a first item embedding. The first item embedding includes information about the first interactive item.

[0113] Optionally, the number of first interactive items is greater than 1, and accordingly, the number of first item embeddings is also greater than 1. The first sequence is the sequence of all first item embeddings. For example, the first interactive items include first interactive item a and first interactive item b, wherein the embedding of first interactive item a is first item embedding c, and the embedding of first interactive item b is first item embedding d. The first sequence is the sequence of first item embedding c and first item embedding d. In the first sequence, the order of first item embedding c is 1, and the order of first item embedding d is 2. Alternatively, in the first sequence, the order of first item embedding c is 2, and the order of first item embedding d is 1.

[0114] Optionally, the time at which the first interactive item corresponding to the first item embedding interacts with the first object is the first interaction time. The earlier the first interaction time corresponding to the first item embedding, the higher the order of the first item embedding in the first sequence. For example, the first interactive items include first interactive item a and first interactive item b, where the embedding of first interactive item a is first item embedding c, and the embedding of first interactive item b is first item embedding d. If the time at which the first object interacts with first interactive item a is t1, and the time at which the first object interacts with first interactive item b is t2, then the time at which the first interactive item interacts with the first object corresponding to first item embedding c is t1. The time at which the first interactive item interacts with the first object corresponding to first item embedding d is t2. If t1 is earlier than t2, then the order of first item embedding c in the first sequence is earlier than the order of first item embedding d in the first sequence. If t2 is earlier than t1, then the order of first item embedding d in the first sequence is earlier than the order of first item embedding c in the first sequence.

[0115] In one implementation for obtaining a first item embedding, a third language model is used to compress the first interactive item based on a first prompt to obtain the first item embedding. For example, the first prompt includes "Please compress the first interactive item into a single word." Optionally, the third language model is an LLM. The third language model may be the same as or different from the first language model. The first prompt is used to guide the third language model to compress the first interactive item. Optionally, the information content of the result obtained by the third language model from compressing the first interactive item is less than that of the first interactive item.

[0116] Optionally, after inputting the first prompt word into the third language model, the third language model may encode the first prompt word to obtain a hidden state sequence of the first prompt word. Encoding the hidden state sequence yields a processing result, which is the result of writing the compressed training search content into the first prompt word. Because the hidden state sequence is used to determine the processing result, the hidden state sequence includes the hidden states of the processing result, i.e., the hidden state sequence includes the hidden states of the result obtained by compressing the first interactive item. The hidden state sequence corresponds to the first prompt word. Specifically, if the word at the reference position in the first prompt word is referred to as the reference word, then the hidden state at the reference position in the hidden state sequence is the embedding of the reference word encoded by the third language model. Therefore, the third language model can obtain the first item embedding based on the hidden state in the hidden state sequence that corresponds to the compressed result of the first interactive item.

[0117] Optionally, the model training device uses the third language model to compress the first interactive item based on the first prompt word (prompt), and the implementation process of obtaining the first item embedding can be expressed by the following formula:

[0118]

[0119] Among them, LLM(·) means calling the third language model for operation, <prompt>Indicates the first prompt word, Represents the first item embedding. .emd[-1] represents taking out the embedding generated by the third language model. Represents the first item embedding obtained by compressing the first interactive item through the third language model. Right now The dimension size is 1×d, where d is the dimension of the embedding generated by the third language model.

[0120] The first interactive item is compressed by the third language model to obtain a first item embedding, and information about the first interactive item can be represented by the first item embedding. The first item embedding includes information representing the interaction behavior between the first object and the first interactive item.

[0121] Optional, Figure 7 This is a schematic diagram of an embodiment of the present application providing a method for compressing a first interactive item through a third language model to obtain a first item embedding. Figure 7 As shown, the third language model is an LLM based on a pure decoder (decoder-only) architecture, that is, the third language model processes the sequence in the order of the sequence. Figure 7 In the example, the first prompt word compresses the first interactive item. The first interactive item includes {"title":xxx,"content":yyy,"release date":zzz}. After the first prompt word and the first interactive item are input into the third language model, the first item embedding can be output.

[0122] Optionally, when the number of first interactive items is greater than 1, the first item embedding of each first interactive item can be obtained by inputting all first interactive items into the third language model. For example, X = {x1, x2, ..., x n }, where X represents the set of first interactive items, x1, x2,…, x n Represents n first interactive items. After X is input into the third language model, it can generate Where E represents the set of first item embeddings, n first-item embeddings representing the n first-interactive items.

[0123] In an implementation method of obtaining the first sequence, the model training device receives the first sequence input by the user through an input component, wherein the input component includes: a keyboard, a mouse, a touch screen, a touchpad, and an audio input device.

[0124] In another implementation of obtaining the first sequence, the model training device receives the first sequence sent by the terminal. Optionally, the terminal can be any of the following: a mobile phone, a computer, a tablet computer, or a server.

[0125] 103. Input the first sequence into a first language model so that the first language model generates a first prediction embedding, wherein the item represented by the first prediction embedding is an item that interacts with the first object after the first interactive item.

[0126] In an embodiment of the present application, the first item embedding is processable by a language model. Optionally, the dimension of the first item embedding is the same as the dimension of the characters that the first language model can process. Optionally, if the dimension of the first item embedding in the first sequence is greater than the dimension of the characters that the first language model can process, the first item embedding is reduced in dimension so that the dimension of the first item embedding is the same as the dimension of the characters that the first language model can process, and the reduced dimension embedding is then input into the first language model.

[0127] After the model training device inputs the first sequence into the first language model, the first language model can predict the embedding of the object that interacts with the first object next after interacting with the first interactive object based on the information of the first object embedding in the first sequence, and obtain a first predicted embedding.

[0128] Optionally, when the number of first item embeddings in the first sequence is greater than one, the model training device sequentially inputs the first item embeddings into the first language model in order of first interaction time, from earliest to latest, so that the first language model generates a first predicted embedding based on the first item embeddings in the first sequence. For example, the time at which the first interactive item corresponding to the first item embedding c interacts with the first object is t1. The time at which the first interactive item corresponding to the first item embedding d interacts with the first object is t2. If t1 is earlier than t2, the model training device first inputs the first item embedding c into the first language model, and then inputs the first item embedding d into the first language model.

[0129] Since the order of first interaction times represents the order in which the first object interacted with the first interactive item corresponding to the first item embedding, the order in which the first object interacted with the first interactive item can better reflect the trend of the first object's interaction with the item, which is helpful for predicting the embedding of the item with which the first object will interact next. Therefore, when the first item embeddings in the first sequence are sequentially input into the first language model in order of first interaction times, the first language model can use information about the first item embeddings and the first interaction times to generate the first predicted embedding, thereby improving the accuracy of the first predicted embedding.

[0130] Optionally, the earlier the first interaction time corresponding to the first item embedding, the earlier the first item embedding is in the first sequence. The model training device inputs the first item embeddings in the first sequence into the first language model in the order in which the first items are embedded in the first sequence. The first item embeddings may be input into the first language model sequentially in order of first interaction time from earliest to latest.

[0131] Optionally, the item corresponding to the first predicted embedding is the item that first interacted with the first object after the first interactive item. For example, the first interactive items include first interactive item a and first interactive item b. The first object interacted with first interactive item a at time t1, and the first object interacted with first interactive item b at time t2. If t1 is earlier than t2, then the item corresponding to the first predicted embedding is the item that first interacted with the first object after t2.

[0132] Optionally, the model training device inputs the embedding of the first sequence and the first object into the first language model so that the first language model generates a first predicted embedding, wherein the embedding of the first object includes attribute information of the first object, for example, the embedding of the first object includes the age and gender of the first object.

[0133] Optionally, inputting the embeddings of the first sequence and the first object into the first language model so that the first language model generates a first predicted embedding can be expressed as follows:

[0134]

[0135] Among them, LLM(·) means calling the first language model for operation,<user profile> Represents the embedding of the first object. Represents the first sequence of n first item embeddings. .emd[-1] means taking out the embedding generated by the first language model.

[0136] Optional, Figure 8 A schematic diagram of a first language model provided in an embodiment of the present application processing a first sequence to generate a first prediction embedding. Figure 8 As shown, first, each first item in the first sequence is embedded into the MLP (i.e. Figure 8 The multi-layer perceptron in ( ) is used to reduce the dimension of each first item embedding to the dimension of the embedding that the first language model can handle. Then the MLP (i.e. Figure 8 The embedding after dimensionality reduction by the multi-layer perceptron in the

[15] is input to the first language model, and the first language model processes and outputs the first predicted embedding.

[0137] 104. Determine a first loss of the first language model based on the first predicted embedding and the first positive sample.

[0138] In this embodiment of the present application, the first positive sample includes items that actually interacted with the first object after the first interactive item. For example, the first interactive item includes first interactive item a. The first object interacted with first interactive item a at time t1, and the first object interacted with item b at time t2. If t1 is earlier than t2, then item b is the item that actually interacted with the first object after the first interactive item.

[0139] Optionally, the first positive sample includes an item that interacts with the first subject when exposed to the first subject. For example, the first positive sample includes document a, where, when document a is exposed to the first subject, the first subject interacts with document a. When the item is exposed to the first subject, the first subject interacts with the item, indicating that the first subject has a high degree of interest in the item. Therefore, the first positive sample includes items that the first subject has a high degree of interest in.

[0140] Since the item corresponding to the first predicted embedding is the item that interacted with the first object after the first interactive item, and the first positive sample includes the item that actually interacted with the first object after the first interactive item, the first positive sample can be used as a basis for evaluating the accuracy of the first predicted embedding, and further determine the loss of the first language model, that is, the first loss.

[0141] In one possible implementation, the model training device determines a first loss based on a first similarity between the first predicted embedding and the embedding of the first positive sample, wherein the first similarity is negatively correlated with the first loss.

[0142] In another possible implementation, the model training device determines a first loss based on the first predicted embedding, the first positive sample, and the negative sample, where the negative sample includes items that did not interact with the first object. Because the negative sample includes items that did not interact with the first object, the negative sample can also serve as a basis for evaluating the accuracy of the first predicted embedding. Therefore, the model training device can determine the first loss based on the first predicted embedding, the first positive sample, and the negative sample.

[0143] Optionally, the model training device determines a first similarity between the first predicted embedding and the embedding of the first positive sample, determines a second similarity between the first predicted embedding and the embedding of the negative sample, and determines a first loss based on the first similarity and the second similarity, wherein the similarity is negatively correlated with the first loss and the second similarity is positively correlated with the first loss.

[0144] Optionally, negative samples include items that, when exposed to a first subject, did not interact with the first subject. For example, a negative sample includes document b. When document b is exposed to the first subject, the first subject does not interact with document b. If the first subject does not interact with the item when exposed to the item, this indicates that the first subject is relatively uninterested in the item. Therefore, negative samples include items that the first subject is relatively uninterested in.

[0145] In another possible implementation, the model training device determines a contrastive loss function value as the first loss based on the first predicted embedding, the first positive sample, and the negative sample. Optionally, the contrastive loss function is an information noise contrastive estimation (InfoNCE) loss.

[0146] Optionally, the model training device determines the first loss by the following formula:

[0147]

[0148]

[0149] in, is the first loss. b is the number of first objects. is the first predicted embedding. is the embedding of the first positive example. is the embedding of the kth negative sample among the m negative samples. exp(·) represents the exponential function with the natural number e as the base. s(a, b) indicates the similarity between a and b. ∑ represents the summation. τ is the temperature coefficient.

[0150] 105. Based on the first loss, update the parameters of the first language model to obtain a target language model.

[0151] The model training device updates the parameters of the first language model based on the first loss until the first loss converges, stops updating the parameters of the first language model, and obtains the target language model.

[0152] In an embodiment of the present application, the first sequence is a sequence of first item embeddings, wherein the first item embeddings include the embeddings of the first interactive item that interacted with the first object. Therefore, the first item embeddings in the first sequence include information about the first interactive item. The first sequence is input into a first language model, so that the first language model processes the first sequence and uses the information from the first item embeddings to generate a first predicted embedding. The first predicted embedding represents an item that interacted with the first object after the first interactive item. Because the first language model includes a state-space model, processing the first sequence by the first language model can reduce computational complexity, thereby increasing the speed of generating the first predicted embedding. A first loss for the first language model is then determined based on the first predicted embeddings and the first positive sample. Based on the first loss, the parameters of the first language model are updated to obtain a target language model. This can increase the speed of obtaining the target language model through training and enable the target language model to predict the embedding of the next item that will interact with the object based on the sequence of embeddings of the items that the object has interacted with.

[0153] As an optional implementation, the model training device performs the following steps during the execution of step 101:

[0154] 201. Input the first item embeddings in the first sequence into the second language model in sequence, so that the second language model generates a second predicted embedding based on the input first item embeddings.

[0155] In this embodiment of the present application, the items represented by the second predicted embedding are items that interact with the first object after the first interactive item input into the second language model. For example, the first sequence includes the first item embedding c of the first interactive item a and the first item embedding d of the first interactive item b. When the first item embeddings in the first sequence are sequentially input into the second language model, the first item embedding c is first input into the second language model, followed by the first item embedding d. After the first item embedding c is input into the second language model, and before the first item embedding d is input into the second language model, the input first item embeddings include the first item embedding c. Based on the first item embedding c, the second predicted embedding generated by the second language model is the embedding of the item that interacts with the first object after the first interactive item a. After both the first item embedding c and the first item embedding d are input into the second language model, the input first item embeddings include the first item embedding c and the first item embedding d. Based on the first item embedding c and the first item embedding d, the second predicted embedding generated by the second language model is the embedding of the item that interacts with the first object after the first interactive item a and the first interactive item b.

[0156] Optionally, the model training device sequentially inputs the first item embeddings in the first sequence into the second language model in order of first interaction time, so that the second language model generates a second predicted embedding based on the input first item embeddings. Since the order of first interaction time, from earliest to latest, represents the order in which the first object interacts with the first interactive item corresponding to the first item embedding, the order in which the first object interacts with the first interactive item can better reflect the trend of interaction between the first object and the item, which is conducive to predicting the embedding of the item with which the first object will interact next. Therefore, when the first item embeddings in the first sequence are sequentially input into the second language model in order of first interaction time, the second language model can generate a second predicted embedding using information about the input first item embeddings and information about the first interaction time corresponding to the input first item embeddings, thereby improving the accuracy of the second predicted embedding.

[0157] Optional, Figure 9 A schematic diagram of a second language model provided in an embodiment of the present application processing a first sequence to generate a second prediction embedding. Figure 9 As shown, the first sequence includes the following n-1 first item embeddings: First, the embedding of each first item in the first sequence is input into the MLP (i.e. Figure 9 The multi-layer perceptron in the is used to reduce the dimension of each first item embedding to the dimension of the embedding that the second language model can handle. Then the MLP (i.e. Figure 9 The embedding after dimensionality reduction by the multi-layer perceptron in is input to the second language model, which then processes and outputs the following n-1 second predicted embeddings: in, is based on generated, is based on and generated, is based on generated.

[0158] 202. Determine a second loss of the second language model based on the second predicted embedding and the second positive sample.

[0159] In an embodiment of the present application, the second positive sample includes an item that actually interacts with the first object after the first interactive item has been input into the second language model. For example, the first sequence includes the first item embedding c of the first interactive item a and the first item embedding d of the first interactive item b. When the first item embeddings in the first sequence are sequentially input into the second language model, the first item embedding c is first input into the second language model, and then the first item embedding d is input into the second language model. After the first item embedding c is input into the second language model, before the first item embedding d is input into the second language model, the input first item embedding includes the first item embedding c. The second predicted embedding generated by the second language model based on the first item embedding c is the embedding of the item that interacts with the first object after the first interactive item a. Accordingly, the second positive sample includes the embedding of the item that actually interacts with the first object after the first interactive item a. After both the first item embedding c and the first item embedding d are input into the second language model, the input first item embedding includes the first item embedding c and the first item embedding d. The second predicted embedding generated by the second language model based on the first item embedding c and the first item embedding d is the embedding of the item that interacts with the first object after the first interactive item a and the first interactive item b. Accordingly, the second positive sample includes the embedding of the item that actually interacts with the first object after the first interactive item a and the first interactive item b.

[0160] Since the item corresponding to the second predicted embedding is the item that interacted with the first object after the first interactive item was input, and the second positive sample includes the item that actually interacted with the first object after the first interactive item was input, the second positive sample can be used as a basis for evaluating the accuracy of the second predicted embedding, and further determine the loss of the second language model, that is, the second loss.

[0161] In one possible implementation, the model training device determines the second loss based on a third similarity between the second predicted embedding and the embedding of the second positive sample, wherein the third similarity is negatively correlated with the second loss.

[0162] In another possible implementation, the model training device determines the second loss based on the second predicted embedding, the second positive sample, and the negative sample. Because the negative sample includes items that did not interact with the first object, the negative sample can also serve as a basis for evaluating the accuracy of the second predicted embedding. Therefore, the model training device can determine the first loss based on the second predicted embedding, the second positive sample, and the negative sample.

[0163] Optionally, the model training device determines a first similarity between the first predicted embedding and the embedding of the first positive sample. Determine a second similarity between the first predicted embedding and the embedding of the negative sample. Determine a first loss based on the first similarity and the second similarity, wherein the similarity is negatively correlated with the first loss and the second similarity is positively correlated with the first loss. Optionally, the negative sample includes an item that has interacted with a third object, wherein the third object is different from the first object.

[0164] In another possible implementation, the model training device determines a contrastive loss function value based on the second predicted embedding, the second positive sample, and the negative sample as the second loss. Optionally, the contrastive loss function is InfoNCE loss.

[0165] Optionally, the model training device determines the second loss by the following formula:

[0166]

[0167] in, is the second loss. b is the number of first objects. is the j-th second predicted embedding generated by the second language model. is the embedding of the second positive sample corresponding to the jth second predicted embedding. is the embedding of the kth negative sample among the m negative samples. exp(·) represents the exponential function with the natural number e as the base. s(a, b) indicates the similarity between a and b. ∑ represents the summation. τ is the temperature coefficient.

[0168] Optionally, in formula (10), the j-th second predicted embedding is generated based on the (j-1)-th first item embedding that is first input to the second language model in the first sequence, and the embedding of the second positive sample corresponding to the j-th second predicted embedding is the j-th first item embedding that is input to the second language model in the first sequence.

[0169] 203. Based on the second loss, update the parameters of the second language model to obtain the first language model.

[0170] The model training device updates the parameters of the second language model based on the second loss until the second loss converges, stops updating the parameters of the second language model, and obtains the first language model.

[0171] The model training device trains the second language model using this implementation to generate a first language model. This enables the first language model to predict the embedding of an item that will interact with the object next, based on the input embeddings of items that have interacted with the object. Thus, during the training process from steps 101 to 105, the first language model can generate a first predicted embedding based on the first sequence.

[0172] As an optional implementation, the first language model further includes an attention mechanism, thereby improving the processing effect of the first language model on the first sequence, and further improving the accuracy of the first predicted embedding obtained by processing the first sequence.

[0173] In this embodiment, the first language model includes a state-space model and an attention mechanism. Because the state-space model of the first language model can reduce computational complexity, the first language model has lower computational complexity than an attention mechanism model that does not include a state-space model. Furthermore, because the attention mechanism processes sequences more effectively than the state-space model, the first language model processes the first sequence more effectively than the state-space model, thereby improving the accuracy of the first predictive embedding generated by processing the first sequence.

[0174] The present application also provides an embedding generation method. After training a target language model using the model training method described above, the method generates an embedding for an object based on the target language model. The embedding generation method is performed by an embedding generation device, which can be any electronic device capable of executing the technical solutions disclosed in the present application. Optionally, the embedding generation device can be any of the following: a computer or a server.

[0175] See also Figure 10 , Figure 10 A flowchart of an embedding generation method provided in an embodiment of the present application.

[0176] 1001. Obtain a second sequence of second item embeddings, where the second item embeddings include embeddings of second interactive items that have interacted with a second object.

[0177] In this embodiment of the present application, a second interactive item is an item that has interacted with a second object, i.e., an interaction has occurred between the second object and the second interactive item. Optionally, the interaction between the object and the item indicates that the object is interested in the item. Optionally, items clicked by the second object, liked by the second object, commented on by the second object, forwarded by the second object, or collected by the second object are all second interactive items. The embedding of the second interactive item is a second item embedding. The second item embedding includes information about the second interactive item.

[0178] Optionally, the number of second interactive items is greater than 1, and accordingly, the number of second item embeddings is also greater than 1. The second sequence is the sequence of all second item embeddings. For example, the second interactive items include a second interactive item a and a second interactive item b, where the embedding of second interactive item a is a second item embedding c, and the embedding of second interactive item b is a second item embedding d. The second sequence is the sequence of second item embedding c and second item embedding d. In the second sequence, the order of second item embedding c is 1, and the order of second item embedding d is 2. Alternatively, in the second sequence, the order of second item embedding c is 2, and the order of second item embedding d is 1.

[0179] Optionally, the time at which the second interactive item corresponding to the second item embedding interacts with the second object is the second interaction time. The earlier the second interaction time corresponding to the second item embedding, the higher the order of the second item embedding in the second sequence. For example, the second interactive items include a second interactive item a and a second interactive item b, where the embedding of second interactive item a is a second item embedding c, and the embedding of second interactive item b is a second item embedding d. If the time at which the second object interacts with second interactive item a is t1, and the time at which the second object interacts with second interactive item b is t2, then the time at which the second interactive item interacts with the second object corresponding to the second item embedding c is t1. The time at which the second interactive item interacts with the second object corresponding to the second item embedding d is t2. If t1 is earlier than t2, then the order of the second item embedding c in the second sequence is earlier than the order of the second item embedding d in the second sequence. If t2 is earlier than t1, then the order of the second item embedding d in the second sequence is earlier than the order of the second item embedding c in the second sequence.

[0180] In one implementation of obtaining the second item embedding, the third language model is used to compress the second interactive item based on a second prompt word to obtain the second item embedding. For example, the second prompt word includes "please compress the second interactive item into one word."

[0181] 1002. Input the second sequence into a target language model so that the target language model generates a target prediction embedding based on the second item embedding.

[0182] The target language model in the embodiments of the present application is trained using the model training method described above. Because the training method described above enables the target language model to predict the embedding of the next item that the object will interact with based on the sequence of embeddings of items that the object has interacted with, the embedding generation device inputs the second sequence into the target language model, enabling the target language model to predict the embedding of the next item that will interact with the second object based on the sequence of embeddings of the second item, i.e., the second predicted embedding.

[0183] 1003. Determine an object embedding of a second object based on the target predicted embedding.

[0184] In this embodiment of the present application, object embeddings are used to determine the items that interact with the second object after the second interactive item. Because the target prediction embedding is the embedding of the items that interact with the second object after the second sequence, the target prediction embedding can be used to determine the items that interact with the second object after the second interactive item.

[0185] In one possible implementation, the embedding generation device uses the target prediction embedding as the object embedding of the second object. In another possible implementation, the embedding generation device obtains an initial embedding of the second object, wherein the initial embedding is the embedding of the second object extracted using the twin-tower model, and the initial embedding includes information for determining the items that interact with the second object after the second interactive item. Based on the initial embedding and the target prediction embedding, the object embedding of the second object is determined. The information for determining the items that interact with the second object after the second interactive item extracted using the twin-tower model is different from the information included in the target prediction embedding for determining the items that interact with the second object after the second interactive item. Therefore, the embedding generation device determines the object embedding of the second object based on the initial embedding and the target prediction embedding, which can enrich the information of the object embedding of the second object and thereby improve the accuracy of the object embedding.

[0186] Optionally, the embedding generation device obtains an object embedding for the second object by concatenating the initial embedding and the target predicted embedding. Optionally, when the dimension of the initial embedding is different from the dimension of the target predicted embedding, the embedding generation device uses an MLP to make the dimension of the initial embedding the same as the dimension of the target predicted embedding, and then concatenates the initial embedding and the target predicted embedding of the same dimension to obtain the object embedding for the second object.

[0187] In an embodiment of the present application, the second sequence is a sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object. The target language model is trained based on the model training method described above, and therefore includes a state-space model. The target language model is capable of predicting the embeddings of items that will interact with the object next based on the sequence of embeddings of items with which the object has interacted. Therefore, after obtaining the second sequence, the embedding generation device inputs the second sequence into the target language model, enabling the target language model to generate a target predicted embedding based on the second item embeddings, wherein the target predicted embedding represents items that interacted with the second object after the second interactive item. Based on the target predicted embeddings, the object embedding of the second object can be determined, wherein the object embedding is used to determine items that interacted with the second object after the second interactive item. Because the target language model includes a state-space model, generating the target predicted embedding using the target language model can reduce the computational complexity of generating the target predicted embedding, increase the speed of generating the target predicted embedding, and thereby increase the speed of determining the object embedding of the second object.

[0188] The present application also provides a recommendation method. After obtaining an object embedding for the second object using the embedding generation method described above, the method recommends items for the second object based on the object embedding. The recommendation method is performed by a recommendation device, which can be any electronic device capable of executing the technical solutions disclosed in the present application. Optionally, the recommendation device can be any of the following: a computer or a server.

[0189] See also Figure 11 , Figure 11 A flowchart of a recommended method provided in an embodiment of the present application.

[0190] 1101. Obtain an object embedding of a second object.

[0191] The object embedding in this step is obtained based on the embedding generation method described above.

[0192] 1102. Determine a target embedding that matches the object embedding from the candidate embeddings.

[0193] In this embodiment of the present application, a candidate embedding is an embedding of a candidate item, where a candidate item is an item to be confirmed as being recommended to a second object. In one possible implementation, the recommendation device calculates a fourth similarity between the object embedding and the candidate embedding, and determines that the candidate embedding is the target embedding if the fourth similarity is greater than or equal to a similarity threshold. In another possible implementation, the recommendation device calculates a fourth similarity between the object embedding and the candidate embedding, and determines that the candidate embedding corresponding to the maximum fourth similarity is the target embedding.

[0194] 1103. Determine the candidate item corresponding to the target embedding as the item to be recommended to the second object.

[0195] Because the object embedding includes information for determining the item that interacted with the second object after the second interactive item, the target embedding matches the object embedding, indicating a high probability that the candidate item corresponding to the target embedding is the item that interacted with the second object after the second interactive item. Therefore, the recommendation device determines that the candidate item corresponding to the target embedding is the item to be recommended to the second object.

[0196] In this embodiment of the present application, the object embedding is obtained based on the embedding generation method described above. Therefore, the object embedding includes information used to determine items that interact with the second object after the second interactive item. After obtaining the object embedding of the second object, the recommendation device determines a target embedding from the candidate embeddings that matches the object embedding. It then determines the candidate item corresponding to the target embedding as the item to be recommended to the second object, thereby increasing the probability that the item to be recommended to the second object is the item that interacted with the second object after the second interactive item.

[0197] In one possible implementation scenario, candidate items include candidate documents, where the candidate documents include text and images. A material library includes candidate documents and candidate embeddings for the candidate documents. After obtaining the object embedding of a second object, the recommendation device determines a target embedding from the material library that matches the object embedding. The candidate document corresponding to the target embedding is then determined as an item to be recommended to the second object, and the candidate document corresponding to the target embedding is recommended to the second object.

[0198] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0199] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0200] See also Figure 12 , Figure 12 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present application. The model training device 1 includes: an acquisition unit 11, a generation unit 12, a determination unit 13, and an update unit 14, wherein:

[0201] An acquiring unit 11 is configured to acquire a first language model, where the first language model includes a state space model;

[0202] The acquisition unit 11 is further configured to acquire a first sequence of first item embeddings, where the first item embeddings include embeddings of first interactive items that have interacted with the first object;

[0203] a generating unit 12 configured to input the first sequence into the first language model so that the first language model generates a first predicted embedding, wherein the first predicted embedding represents an item that interacts with the first object after the first interactive item;

[0204] a determining unit 13, configured to determine a first loss of the first language model based on the first predicted embedding and a first positive sample, where the first positive sample includes an item that actually interacted with the first object after the first interactive item;

[0205] The updating unit 14 is configured to update the parameters of the first language model based on the first loss to obtain a target language model.

[0206] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0207] inputting the first item embeddings in the first sequence into a second language model in sequence, so that the second language model generates a second predicted embedding based on the input first item embeddings, where the item represented by the second predicted embedding is an item that interacts with the first object after the first interactive item input into the second language model;

[0208] determining a second loss for the second language model based on the second predicted embedding and a second positive sample, the second positive sample including an item that actually interacted with the first object after the first interactive item was input to the second language model;

[0209] Based on the second loss, parameters of the second language model are updated to obtain the first language model.

[0210] In conjunction with any embodiment of the present application, the time when the first interactive object corresponding to the embedding of the first object interacts with the first object is the first interaction time;

[0211] The acquisition unit 11 is configured to:

[0212] The first item embeddings in the first sequence are sequentially input into a second language model in order of the first interaction time from earliest to latest, so that the second language model generates a second predicted embedding based on the input first item embeddings.

[0213] In combination with any embodiment of the present application, the generating unit 12 is configured to:

[0214] The first item embeddings in the first sequence are sequentially input into the first language model in order of the first interaction time from earliest to latest, so that the first language model generates the first predicted embedding.

[0215] In combination with any embodiment of the present application, the first positive sample includes an object that interacts with the first object when exposed to the first object;

[0216] The determining unit 13 is configured to:

[0217] The first loss is determined based on the first predicted embedding, the first positive sample, and a negative sample, where the negative sample includes an item that does not interact with the first object when exposed to the first object.

[0218] In combination with any embodiment of the present application, the first language model also includes an attention mechanism.

[0219] In an embodiment of the present application, the first sequence is a sequence of first item embeddings, wherein the first item embeddings include the embeddings of the first interactive item that interacted with the first object. Therefore, the first item embeddings in the first sequence include information about the first interactive item. The first sequence is input into a first language model, so that the first language model processes the first sequence and uses the information from the first item embeddings to generate a first predicted embedding. The first predicted embedding represents an item that interacted with the first object after the first interactive item. Because the first language model includes a state-space model, processing the first sequence by the first language model can reduce computational complexity, thereby increasing the speed of generating the first predicted embedding. A first loss for the first language model is then determined based on the first predicted embeddings and the first positive sample. Based on the first loss, the parameters of the first language model are updated to obtain a target language model. This can increase the speed of obtaining the target language model through training and enable the target language model to predict the embedding of the next item that will interact with the object based on the sequence of embeddings of the items that the object has interacted with.

[0220] See also Figure 13 , Figure 13 This is a schematic diagram of the structure of an embedding generation device provided in an embodiment of the present application. The embedding generation device 2 includes: an acquisition unit 21, a generation unit 22, and a determination unit 23, wherein:

[0221] An acquiring unit 21 is configured to acquire a second sequence of second item embeddings, where the second item embeddings include embeddings of second interactive items that have interacted with a second object;

[0222] a generating unit 22 configured to input the second sequence into a target language model, so that the target language model generates a target prediction embedding based on the second item embedding, wherein the target language model is trained based on the first aspect and any embodiment thereof, and the item represented by the target prediction embedding is an item that interacts with the second object after the second interactive item;

[0223] A determination unit 23 is configured to determine an object embedding of the second object based on the target prediction embedding, wherein the object embedding is used to determine an item that interacts with the second object after the second interactive item.

[0224] In an embodiment of the present application, the second sequence is a sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object. The target language model is trained based on the model training method described above, and therefore includes a state-space model. The target language model is capable of predicting the embeddings of items that will interact with the object next based on the sequence of embeddings of items with which the object has interacted. Therefore, after obtaining the second sequence, the embedding generation device inputs the second sequence into the target language model, enabling the target language model to generate a target predicted embedding based on the second item embeddings, wherein the target predicted embedding represents items that interacted with the second object after the second interactive item. Based on the target predicted embeddings, the object embedding of the second object can be determined, wherein the object embedding is used to determine items that interacted with the second object after the second interactive item. Because the target language model includes a state-space model, generating the target predicted embedding using the target language model can reduce the computational complexity of generating the target predicted embedding, increase the speed of generating the target predicted embedding, and thereby increase the speed of determining the object embedding of the second object.

[0225] See also Figure 14 , Figure 14 This is a schematic diagram of the structure of a recommendation device provided in an embodiment of the present application. The recommendation device 3 includes: an acquisition unit 31 and a determination unit 32, wherein:

[0226] An acquiring unit 31 is configured to acquire an object embedding of a second object, where the object embedding is obtained based on the method provided in the second aspect;

[0227] a determining unit 32 for determining a target embedding that matches the object embedding from candidate embeddings, the candidate embeddings being embeddings of candidate items;

[0228] The determining unit 32 is further configured to determine that the candidate item corresponding to the target embedding is the item to be recommended to the second object.

[0229] In this embodiment of the present application, the object embedding is obtained based on the embedding generation method described above. Therefore, the object embedding includes information used to determine items that interact with the second object after the second interactive item. After obtaining the object embedding of the second object, the recommendation device determines a target embedding from the candidate embeddings that matches the object embedding. It then determines the candidate item corresponding to the target embedding as the item to be recommended to the second object, thereby increasing the probability that the item to be recommended to the second object is the item that interacted with the second object after the second interactive item.

[0230] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0231] Figure 15 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device 4 includes a processor 41 and a memory 42. Optionally, the electronic device 4 also includes an input device 43 and an output device 44. The processor 41, the memory 42, the input device 43 and the output device 44 are coupled via a connector, and the connector includes various interfaces, transmission lines or buses, etc., which are not limited in the embodiments of the present application. It should be understood that in each embodiment of the present application, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, for example, connection through various interfaces, transmission lines, buses, etc.

[0232] The processor 41 may include one or more processors, for example, one or more central processing units (CPUs). In the case where the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU. Alternatively, the processor 41 may be a processor group consisting of multiple CPUs, wherein the multiple processors are coupled to each other via one or more buses. Alternatively, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.

[0233] The memory 42 can be used to store computer program instructions and various computer program codes, including the program code for executing the solution of the present application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or portable compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0234] The input device 43 is used to input data and / or signals, and the output device 44 is used to output data and / or signals. The input device 43 and the output device 44 can be independent devices or an integrated device.

[0235] It can be understood that in the embodiment of the present application, the memory 42 can be used not only to store relevant instructions, but also to store relevant data. The embodiment of the present application does not limit the specific data stored in the memory.

[0236] It is understandable that Figure 15 Only a simplified design of an electronic device is shown. In actual applications, the electronic device may further include other necessary components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the scope of protection of the present application.

[0237] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0238] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here. Those skilled in the art will also clearly understand that the descriptions of the various embodiments of this application have different focuses. For the convenience and brevity of description, the same or similar parts may not be repeated in different embodiments. Therefore, for parts not described or not described in detail in a certain embodiment, reference can be made to the descriptions of other embodiments.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0240] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0241] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0242] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0243] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing related hardware to perform the processes. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.< / prompt>

Claims

1. A model training method, characterized in that: The method comprises: Acquire a first language model, wherein the first language model includes a state-space model; Obtaining a first sequence of first item embeddings, wherein the first item embeddings include embeddings of first interactive items that have interacted with a first object; Inputting the first sequence into the first language model so that the first language model generates a first predicted embedding, wherein the first predicted embedding represents an item that interacts with the first object after the first interactive item; determining a first loss for the first language model based on the first predicted embedding and a first positive sample, the first positive sample comprising an item that actually interacted with the first object after the first interactive item; Based on the first loss, parameters of the first language model are updated to obtain a target language model.

2. The method according to claim 1, characterized in that The obtaining of the first language model includes: inputting the first item embeddings in the first sequence into a second language model in sequence, so that the second language model generates a second predicted embedding based on the input first item embeddings, where the item represented by the second predicted embedding is an item that interacts with the first object after the first interactive item input into the second language model; determining a second loss for the second language model based on the second predicted embedding and a second positive sample, the second positive sample including an item that actually interacted with the first object after the first interactive item was input to the second language model; Based on the second loss, parameters of the second language model are updated to obtain the first language model.

3. The method according to claim 2, characterized in that The time when the first interactive object corresponding to the embedding of the first object interacts with the first object is the first interaction time; The step of sequentially inputting the first item embeddings in the first sequence into a second language model so that the second language model generates a second predicted embedding based on the input first item embeddings includes: The first item embeddings in the first sequence are sequentially input into a second language model in order of the first interaction time from earliest to latest, so that the second language model generates a second predicted embedding based on the input first item embeddings.

4. The method according to claim 3, characterized in that Inputting the first item embedding in the first sequence into the first language model so that the first language model generates a first predicted embedding includes: The first item embeddings in the first sequence are sequentially input into the first language model in order of the first interaction time from earliest to latest, so that the first language model generates the first predicted embedding.

5. The method according to any one of claims 1 to 4, characterized in that The first positive sample includes an item that interacts with the first object when exposed to the first object; The determining, based on the first predicted embedding and the first positive sample, a first loss of the first language model includes: The first loss is determined based on the first predicted embedding, the first positive sample, and a negative sample, where the negative sample includes an item that does not interact with the first object when exposed to the first object.

6. The method according to any one of claims 1 to 4, characterized in that The first language model also includes an attention mechanism.

7. A method for generating an embedding, characterized in that: The method comprises: Obtaining a second sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object; inputting the second sequence into a target language model so that the target language model generates a target prediction embedding based on the second item embedding, wherein the target language model is trained based on the method of any one of claims 1 to 6, and the item represented by the target prediction embedding is an item that interacts with the second object after the second interactive item; Based on the target prediction embedding, an object embedding of the second object is determined, and the object embedding is used to determine an item that interacts with the second object after the second interactive item.

8. A recommendation method, characterized in that: The method comprises: Obtaining an object embedding of a second object, wherein the object embedding is obtained based on the method of claim 7; determining a target embedding that matches the object embedding from candidate embeddings, the candidate embeddings being embeddings of candidate items; The candidate item corresponding to the target embedding is determined to be an item to be recommended to the second object.

9. A model training device, characterized in that: The model training device comprises: an acquiring unit, configured to acquire a first language model, wherein the first language model includes a state space model; The acquisition unit is further configured to acquire a first sequence of first item embeddings, where the first item embeddings include embeddings of first interactive items that have interacted with the first object; a generating unit, configured to input the first sequence into the first language model, so that the first language model generates a first predicted embedding, wherein the first predicted embedding represents an item that interacts with the first object after the first interactive item; a determining unit, configured to determine a first loss of the first language model based on the first predicted embedding and a first positive sample, the first positive sample comprising an item that actually interacted with the first object after the first interactive item; An updating unit is configured to update parameters of the first language model based on the first loss to obtain a target language model.

10. An embedding generation device, characterized in that: The embedding generating device comprises: an acquiring unit, configured to acquire a second sequence of second item embeddings, wherein the second item embeddings include embeddings of second interactive items that have interacted with the second object; a generating unit, configured to input the second sequence into a target language model, so that the target language model generates a target prediction embedding based on the second item embedding, wherein the target language model is trained based on the method of any one of claims 1 to 6, and the item represented by the target prediction embedding is an item that interacts with the second object after the second interactive item; A determining unit is configured to determine an object embedding of the second object based on the target prediction embedding, wherein the object embedding is used to determine an item that interacts with the second object after the second interactive item.

11. A recommendation device, characterized in that: The recommended device includes: an acquiring unit, configured to acquire an object embedding of a second object, wherein the object embedding is obtained based on the method according to claim 7; a determining unit for determining a target embedding that matches the object embedding from candidate embeddings, the candidate embeddings being embeddings of candidate items; The determining unit is further configured to determine that the candidate item corresponding to the target embedding is the item to be recommended to the second object.

12. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 6, or the electronic device executes the method according to claim 7, or the electronic device executes the method according to claim 8.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 6, or the method according to claim 7, or the method according to claim 8.

14. A computer program product, characterized in that The computer program product includes a computer program or instructions; when the computer program or instructions are run on a computer, the computer is enabled to execute the method according to any one of claims 1 to 6, or the computer is enabled to execute the method according to claim 7, or the computer is enabled to execute the method according to claim 8.