Intention learning method and device, electronic equipment, storage medium and product
By using the end-to-end learningable clustering model in intention learning, the local optimization problem caused by the separation of E-step and M-steps in the EM framework is solved, and more efficient user behavior clustering and recommendation system performance improvement is achieved.
Patent Information
- Application Number
- CN202510126391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
In the existing intention learning scheme, the separation of E and M steps based on the EM framework results in the fact that each step can only be optimized within the local scope and cannot be fully coordinated, which leads to the problem of local optimal solutions and slow convergence speed.
By obtaining the user behavior sequence, embedding it into the latent space, obtaining the user behavior vector, and iteratively training the equity prediction model based on the overall loss minimization goal of the equity prediction model to obtain an end-to-end learningable clustering model.
The overall end-to-end optimization is achieved, and user behavior data is effectively clustered, providing more accurate user interest and intention information for the recommendation system, and improving the accuracy and performance of the recommendation system.
Smart Images

Figure CN120067717A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an intention learning method, apparatus, electronic device, storage medium, and product. Background Art
[0002] In related intention learning solutions, intention learning is usually performed based on the Expectation-Maximization (EM) framework. The solution alternately performs the expectation (E) step and the maximization (M) step during the learning process to gradually improve the model performance. However, due to the separation of the E step and the M step in the EM framework, each step can only be optimized within its respective local range and cannot be comprehensively coordinated, resulting in problems of getting stuck in local optimal solutions and slow convergence speed in related intention learning solutions. Summary of the Invention
[0003] The present disclosure provides an intention learning method, apparatus, electronic device, storage medium, and product to solve the problem in related technologies that due to the separation of the E step and the M step in the EM framework, each step can only be optimized within its respective local range and cannot be comprehensively coordinated.
[0004] A first aspect embodiment of the present disclosure proposes an intention learning method, which includes:
[0005] Obtain a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of a first user at corresponding times;
[0006] Embed the first user behavior sequence into a latent space to obtain a first user behavior vector;
[0007] Based on the first user behavior vector, taking the minimum overall loss of the rights and interests prediction model as the goal, iteratively train the rights and interests prediction model to obtain a target model, where the target model is an end-to-end learnable clustering model;
[0008] Obtain a second user behavior sequence, and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
[0009] In one embodiment, embedding the first user behavior sequence into a latent space to obtain a first user behavior vector includes:
[0010] Use an encoder to embed the first user behavior sequence into a latent space to obtain the behavior sequence embeddings of the user at different times;
[0011] Perform an aggregation process on the behavior sequence embeddings of the user at different times to obtain a first user behavior vector.
[0012] In one embodiment, based on the first user behavior vector, with the goal of minimizing the overall loss of the rights and interests prediction model, the rights and interests prediction model is iteratively trained to obtain a target model, including:
[0013] Initialize the cluster center as learnable neural network parameters;
[0014] Normalize the first user behavior vector and the cluster center to obtain a first unit vector corresponding to the first user behavior vector and a second unit vector corresponding to the cluster center;
[0015] Based on the first unit vector and the second unit vector, determine the clustering loss of the rights and interests prediction model, where the clustering loss of the rights and interests prediction model includes a decoupling loss and an alignment loss;
[0016] Based on the clustering loss of the rights and interests prediction model, with the goal of minimizing the overall loss of the rights and interests prediction model, the rights and interests prediction model is iteratively trained to obtain a target model.
[0017] In one embodiment, after determining the clustering loss of the rights and interests prediction model based on the first unit vector and the second unit vector, the method provided by the present disclosure includes:
[0018] Update the cluster center based on the clustering loss of the rights and interests prediction model and a preset algorithm;
[0019] Perform sequential enhancement processing on the first user behavior sequence to obtain a first user behavior enhanced sequence;
[0020] Based on the cluster center, the first user behavior sequence, and the first user behavior enhanced sequence, determine an intention-assisted contrast learning loss, where the intention-assisted contrast learning loss includes a behavior sequence contrast loss and an intention contrast loss;
[0021] The step of, based on the clustering loss of the rights and interests prediction model, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively training the rights and interests prediction model to obtain a target model, includes:
[0022] Based on the clustering loss of the rights and interests prediction model and the intention-assisted contrast learning loss, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively train the rights and interests prediction model to obtain a target model.
[0023] In one embodiment, determining the intention-assisted contrast learning loss based on the cluster center, the first user behavior sequence, and the first user behavior enhanced sequence includes:
[0024] Determine a first index corresponding to the first user behavior sequence and a second index corresponding to the first user behavior enhancement sequence based on the clustering center, the first user behavior sequence, and the first user behavior enhancement sequence;
[0025] Based on the first user behavior sequence and the first index or the first user behavior enhancement sequence and the second index, obtain a behavior fusion sequence and a behavior fusion enhancement sequence according to a first fusion strategy, where the first fusion strategy is used to fuse the potential intention of the user into the first user behavior sequence and the first user behavior enhancement sequence;
[0026] Determine a behavior sequence contrast loss based on the behavior fusion sequence and the behavior fusion enhancement sequence, and determine an intention contrast loss based on the behavior fusion sequence and the clustering center;
[0027] Sum the behavior sequence contrast loss and the intention contrast loss to obtain an intention-assisted contrast learning loss.
[0028] In one embodiment, after summing the behavior sequence contrast loss and the intention contrast loss to obtain an intention-assisted contrast learning loss, the method provided by the present disclosure includes:
[0029] Obtain a data representation, where the data representation includes a representation of the first user behavior sequence and a representation of the rights and interests information;
[0030] Connect the representation of the first user behavior sequence and the representation of the rights and interests information to obtain a combined representation;
[0031] Based on the combined representation, determine the next right and interest corresponding to the combined representation from the mapping relationship between the combined representation and the next right and interest, and determine a prediction loss corresponding to the next right and interest;
[0032] Taking the overall loss of the rights and interests prediction model to be minimized as the goal, iteratively train the rights and interests prediction model based on the clustering loss of the rights and interests prediction model and the intention-assisted contrast learning loss, to obtain a target model, including:
[0033] Taking the overall loss of the rights and interests prediction model to be minimized as the goal, iteratively train the rights and interests prediction model based on the clustering loss of the rights and interests prediction model, the intention-assisted contrast learning loss, and the prediction loss corresponding to the next right and interest, to obtain a target model.
[0034] A second aspect embodiment of the present disclosure proposes an intention learning device, and the device includes:
[0035] An acquisition unit, configured to acquire a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of a first user at corresponding times;
[0036] An embedding unit, configured to embed the first user behavior sequence into a latent space to obtain a first user behavior vector;
[0037] A training unit, configured to iteratively train the rights and interests prediction model based on the first user behavior vector with the goal of minimizing the overall loss of the rights and interests prediction model, to obtain a target model, where the target model is an end-to-end learnable clustering model;
[0038] A determination unit, configured to obtain a second user behavior sequence and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
[0039] A third aspect embodiment of the present disclosure provides an electronic device, including:
[0040] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect embodiment of the present disclosure.
[0041] A fourth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in the first aspect embodiment of the present disclosure.
[0042] A fifth aspect embodiment of the present disclosure provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect embodiment of the present disclosure.
[0043] In summary, the present disclosure provides an intention learning method, which includes: obtaining a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of a first user at corresponding times; embedding the first user behavior sequence into a latent space to obtain a first user behavior vector; iteratively training a rights and interests prediction model based on the first user behavior vector with the goal of minimizing the overall loss of the rights and interests prediction model, to obtain a target model, where the target model is an end-to-end learnable clustering model; obtaining a second user behavior sequence and using the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
[0044] According to the solution provided by the present disclosure, through the use of an end-to-end learnable clustering model based on the user's behavior sequence, overall end-to-end optimization is achieved, effectively clustering the user behavior data, providing more accurate user interest and intention information for the system, thereby improving the accuracy of the recommendation system. By introducing intention-assisted contrastive learning, the interaction between intention learning and behavior learning is enhanced, enabling better capture of the user's latent intentions, improving the performance and efficiency of the recommendation system, and solving the problems of falling into local optimal solutions and slow algorithm convergence speed caused by the step-by-step approach based on the EM framework in the related solutions.
[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0046] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0047] Figure 1 It is a flowchart of the intention learning method provided by an embodiment of the present disclosure;
[0048] Figure 2 It is a flowchart of the method for obtaining the first user behavior vector provided by an embodiment of the present disclosure;
[0049] Figure 3 It is a flowchart of the method for obtaining the target model provided by an embodiment of the present disclosure;
[0050] Figure 4 It is a flowchart of the method for determining the intention-assisted contrastive learning loss provided by an embodiment of the present disclosure;
[0051] Figure 5 It is a flowchart of the method for determining the intention-assisted contrastive learning loss provided by an embodiment of the present disclosure;
[0052] Figure 6 It is a flowchart of the method for determining the prediction loss corresponding to the next right provided by an embodiment of the present disclosure;
[0053] Figure 7 It is an architecture diagram of the intention learning method provided by an embodiment of the present disclosure;
[0054] Figure 8 It is a flowchart of the intention learning method provided by the application example of the present disclosure;
[0055] Figure 9 It is a structural diagram of the intention learning device provided by an embodiment of the present disclosure;
[0056] Figure 10 Schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present disclosure. Detailed implementation manners
[0057] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation to the present disclosure.
[0058] In related intent learning schemes, intent learning is usually based on the EM framework. The scheme alternately performs the expectation (E) step and the maximization (M) step during the learning process to gradually improve the model performance. However, due to the separation of the E step and the M step in the EM framework, each step can only be optimized within its respective local range and cannot be comprehensively coordinated, which leads to the problems of falling into local optimal solutions and slow convergence speed in related intent learning schemes.
[0059] The following briefly introduces the intent learning methods in related technologies:
[0060] The existing optimization paradigm for intent learning is mainly the EM framework. Under this framework, the intent learning problem is usually modeled as a probability model with latent variables, and the EM algorithm is widely used to optimize model parameters and infer latent variables. The basic idea of this framework is to alternately perform the expectation (E) step and the maximization (M) step during the learning process to gradually improve the model performance. Specifically, the EM framework usually includes two main steps: Expectation step (E): According to the estimated value of the current model parameters, calculate the posterior probability distribution of the latent variables. This step usually involves calculating the expected value so as to use these probabilities to update the model parameters in the M step. In intent learning, the E step infers the latent intent structure from the embedding space of user behaviors through a clustering algorithm. Maximization step (M): Using the posterior probability distribution of the latent variables calculated in the E step, maximize the log-likelihood function or other appropriate loss functions to update the model parameters. This step usually involves maximizing the expected value. In intent learning, the M step uses self-supervised learning methods to perform embedding learning on the user behavior sequence to further optimize the model parameters and make them better approximate the true latent distribution.
[0061] In the above scheme, the following defects exist:
[0062] First, in the E step, it is necessary to apply a clustering algorithm to the entire data, which easily leads to problems such as out-of-memory or long running time. In the case of a large amount of data, the computational complexity of the clustering algorithm increases sharply, which limits the scalability and practicality of the model.
[0063] Secondly, the separation between the expectation step and the maximization step in the EM framework leads to several problems: the inadequacy and mutual influence of the sub-optimization process: the separation of the E step and the M step causes each step to be optimized only within its respective local scope, rather than being comprehensively coordinated. This will make the overall optimization process insufficient, unable to fully utilize the information of the entire dataset to optimize the model, resulting in the final model performance being affected; the emergence of local optimal solutions. Due to the separation of the E step and the M step, the optimization process of each step is carried out independently. Such independence may cause the model to fall into local optimal solutions and fail to find the global optimal solution; the slowdown of the algorithm convergence speed. The optimization of the E step and the M step is difficult to coordinate, resulting in poor coordination of the entire optimization process, thus slowing down the convergence speed of the algorithm. In practical applications, this will lead to a long model training time and high resource consumption, affecting the practicality of the model.
[0064] To solve the defects in the related technologies, the present disclosure realizes overall end-to-end optimization by using an end-to-end learnable clustering model based on the user's behavior sequence, effectively clustering the user behavior data, providing more accurate user interest and intention information for the system, thereby improving the accuracy of the recommendation system. By introducing intention-assisted contrast learning, the interaction between intention learning and behavior learning is enhanced, enabling better capture of the user's potential intentions, improving the performance and efficiency of the recommendation system, and solving the problems of falling into local optimal solutions and slow algorithm convergence speed caused by the step-by-step approach based on the EM framework in the related solutions.
[0065] An intention learning method provided by an embodiment of the present disclosure can be applied to multiple fields such as personalized recommendation systems, dialogue systems, and intelligent assistants, such as an accurate recommendation system for rights and interests customers for product recommendation or marketing. The execution subject of the method can be any computing system or platform capable of processing and analyzing user behavior data and having the ability to implement machine learning or deep learning models.
[0066] The following further describes the present disclosure in detail with reference to the accompanying drawings and specific embodiments.
[0067] As Figure 1 shown, Figure 1 is a flowchart of the intention learning method provided by an embodiment of the present disclosure. The intention learning method provided by an embodiment of the present disclosure includes the following steps:
[0068] Step 101, obtain a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of the first user at the corresponding time;
[0069] In one embodiment, the number of first user behavior vectors can be multiple.
[0070] In one embodiment, the first user behavior sequence can be obtained by analyzing the usage of the rights and interests related application (APP) of the first user.
[0071] In one embodiment, the user set U, the product set V, and the set of historical behavior sequences of users are collected Among them, in the user set U, the mobile phone number is used as the unique identifier for each user. The product set V contains all rights and interests information, such as video rights, network disk rights, global access security inspection rights, VIP lounge rights, etc. Each type of right is represented by a unique identifier. The set of historical behavior sequences of users includes the behavior sequences related to the rights and interests of each user. The behaviors include browsing, watching, downloading, purchasing, etc. on the APP related to the rights and interests. Among them, S u represents the behavior sequence of the u-th user.
[0072] In one embodiment, the intention learning method further includes processing the above collected data for subsequent training of the model. Specifically, it mainly includes data cleaning, feature extraction, feature encoding, etc. to ensure the quality and integrity of the data. Among them, data cleaning is used to perform deduplication operations on possible duplicate records to ensure that each record in the dataset is unique, detect and process missing values in the data, analyze the reasons for data missing and perform targeted processing; feature extraction is used to extract meaningful features from the user's behavior data, such as the browsing times, watching duration, download times, etc. on different rights and interests APPs, and extract features from the attribute information of the rights and interests, such as the type, duration, usage rules, etc. of the rights and interests; feature encoding is used to convert categorical features, such as rights and interests types, user regions, etc., into numerical formats using label encoding methods. For numerical features, such as user behavior times, watching duration, etc., normalization or standardization, etc. can be performed to ensure that the numerical ranges of different features are consistent.
[0073] In one embodiment, the intention learning method further includes dividing the dataset into a training set, a validation set, and a test set according to a preset ratio (such as 7:2:1) for model training, parameter tuning, and evaluation.
[0074] Step 102, embed the first user behavior sequence into the latent space to obtain the first user behavior vector;
[0075] In one embodiment, the latent space refers to a low-dimensional and continuous representation space. In this space, different features of the original data are mapped into a set of numerical values, and the main structure and patterns of the original data can be captured through these numerical values.
[0076] In one embodiment, the behavior encoder can be used to embed the first user behavior sequence into the latent space to obtain the first user behavior vector.
[0077] In one embodiment, by embedding the first user behavior sequence into the latent space, a first user behavior vector is obtained, which can capture the latent behavior characteristics of the user, reflect the user's preferences and intentions, and thus provide a more accurate rights and interests recommendation service for each user.
[0078] Step 103: Based on the first user behavior vector, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively train the rights and interests prediction model to obtain a target model, where the target model is an end-to-end learnable clustering model.
[0079] In one embodiment, since the target model is a pre-trained model, when the user behavior sequence is obtained again, the next possible interested rights and interests corresponding to the user behavior sequence can be quickly obtained, thereby improving the speed and accuracy of the next possible interested rights and interests corresponding to the user behavior sequence.
[0080] In one embodiment, by determining that the target model is an end-to-end learnable clustering model, the internal information and data of the system can be better utilized to achieve overall end-to-end optimization, improve the performance and efficiency of the recommendation system, and at the same time avoid the problems of falling into local optimal solutions and slow algorithm convergence speed caused by the separation of the expectation step and the maximization step in the EM framework in the traditional intention learning scheme.
[0081] Step 104: Obtain a second user behavior sequence, and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the next possible interested rights and interests of the user.
[0082] In one embodiment, the second user behavior sequence is used for the prediction of the target model.
[0083] In one embodiment, the second user behavior sequence may be the same as or different from the aforementioned first user behavior sequence.
[0084] In one embodiment, by using the target model to determine the rights and interests corresponding to the second user behavior sequence, the speed and accuracy of obtaining the next possible interested rights and interests corresponding to the user behavior sequence are improved.
[0085] In one embodiment, as Figure 2 shown, embedding the first user behavior sequence into the latent space to obtain a first user behavior vector includes:
[0086] Step 201: Use an encoder to embed the first user behavior sequence into the latent space to obtain the behavior sequence embeddings of the user at different times.
[0087] In one embodiment, the encoder is based on the Transformer architecture.
[0088] In one embodiment, the mathematical expression of the first user behavior sequence is represented by the following formula:
[0089] E u = F(S u );
[0090] where represents the behavior sequence embedding of user u, d' is the number of dimensions of the features in the latent space, F refers to the behavior encoder, and |S u | represents the length of the behavior sequence of user u.
[0091] In one embodiment, since the behavior sequence lengths of different users are different, all users' behavior sequences can be preprocessed into sequences with the same time length by padding or truncating.
[0092] In one embodiment, the behavior encoder captures and summarizes the behaviors of each user at different times. In the precise recommendation system for privileged customers, the behaviors at different times can include: in the morning, the user checks the discount offers and limited-time promotions of the day; at noon, the user participates in exclusive membership activities; in the afternoon, the user browses the privileged information of specific categories, such as tourism, shopping, etc. By embedding these behaviors at different times into the latent space, the behavior encoder can capture the temporal patterns and features of the user's behavior, forming a behavior sequence embedding with time dependence. These embeddings can not only reflect the specific operations of the user at different time periods but also reveal the user's preferences and interest changes at different time periods.
[0093] Step 202: Aggregate the behavior sequence embeddings of the user at different times to obtain a first user behavior vector.
[0094] In one embodiment, the behavior sequence embeddings of the user at different times can be aggregated through a concatenation pooling function P to obtain a first user behavior vector. Specifically, the concatenation pooling function P concatenates the behavior sequence embeddings of the user at different times to form a long behavior vector.
[0095] In one embodiment, the mathematical expression of the first user behavior vector is represented by the following formula:
[0096]
[0097] where represents the behavior sequence embedding of the user at the i-th step, T refers to the length of the behavior sequence, and h u ∈R 1 ×Td′Denote the aggregated behavior embedding of user u, i.e., the first user behavior vector, and re - represent Td' as d.
[0098] In one embodiment, through encoding and aggregation, the behavior sequence embeddings of all users are obtained, H ∈ R |U|×d , H = [h 1 , h 2 , L, h i , L, h b T Denote the set of all user behavior vectors.
[0099] In one embodiment, by aggregating the behavior sequence embeddings of the user at different times, the first user behavior vector is obtained, which can summarize the behavior of each user at different times.
[0100] In one embodiment, as Figure 3 shown, based on the first user behavior vector, with the goal of minimizing the overall loss of the equity prediction model, the equity prediction model is iteratively trained to obtain the target model, including:
[0101] Step 301, initialize the cluster centers as learnable neural network parameters;
[0102] In one embodiment, the cluster centers are used to indicate the user intentions corresponding to the user behavior sequences.
[0103] In one embodiment, the learnable neural network parameters are tensors with gradients, and these cluster centers will be continuously updated during training, which can better reflect the potential patterns of user behavior.
[0104] In one embodiment, the mathematical expression of the cluster centers is represented by the following formula:
[0105] C ∈ R k×d ;
[0106] where C = [c 1 , c 2 , L, c i , L, c k T Denote the set of all cluster centers, c i denotes the i - th cluster center, k is the number of cluster centers, and d is the dimension of the cluster centers.
[0107] Step 302, normalize the first user behavior vector and the cluster centers to obtain the first unit vector corresponding to the first user behavior vector and the second unit vector corresponding to the cluster centers;
[0108] In one embodiment, for the first user behavior vector hi and the clustering center c i perform normalization processing to obtain a first unit vector corresponding to the first user behavior vector and a second unit vector corresponding to the clustering center
[0109] In one embodiment, determining the first unit vector corresponding to the first user behavior vector is represented by the following mathematical expression: where ||h i || 2 is the Euclidean norm of the first unit vector, specifically the square root of the sum of the squares of its respective elements.
[0110] In one embodiment, the second unit vector corresponding to the clustering center is represented by the following mathematical expression: where ||c i || 2 represents the Euclidean norm of the second unit vector corresponding to the clustering center, specifically the square root of the sum of the squares of its respective elements.
[0111] In one embodiment, a clustering center matrix is formed by all the normalized clustering centers Similarly, a first unit vector matrix is formed by all the normalized first user behavior vectors
[0112] In one embodiment, by performing normalization processing on the first user behavior vector and the clustering center, constraining the behavior embedding and the clustering center embedding to be distributed on a unit sphere improves the numerical stability of the model, the consistency of distance measurement, and the interpretability of the model.
[0113] Step 303, based on the first unit vector and the second unit vector, determine the clustering loss of the rights and interests prediction model, and the clustering loss of the rights and interests prediction model includes a decoupling loss and an alignment loss;
[0114] In one embodiment, the mathematical expression for determining the decoupling loss is represented by the following formula:
[0115]
[0116] where k represents the number of clusters (intentions), represents the clustering center c iThe unit vector, by minimizing the negative squared Euclidean distance between different cluster centers, enables better separation between different intentions (clusters), thereby reducing the overlap between clusters and improving the interpretability of the model. The time complexity and space complexity of the decoupling loss are O(k^2*d) and O(kd) respectively, and the number of user intentions is much smaller than the number of users (i.e., k << |U|). Therefore, the intention decoupling part does not incur significant time or space costs.
[0117] In one embodiment, the mathematical expression for determining the alignment loss is represented by the following formula:
[0118]
[0119] where b represents the batch size, specifically the number of users, represents the unit vector of the behavior sequence embedding h i By minimizing the distance between the behavior embedding and all cluster centers, similar behaviors are ensured to be clustered within the same cluster. The behavior embedding is pulled towards all cluster centers instead of the nearest one, which can avoid the confirmation bias problem and ensure the stability of the clustering process.
[0120] In one embodiment, the mathematical expression for the clustering loss of the rights and interests prediction model is represented by the following formula:
[0121] L cluster = L decoupling + L alignment ;
[0122] In one embodiment, the clustering loss is introduced to train the network and the cluster centers, so that similar user behavior sequence embeddings can be clustered together, while the centers between different clusters are as far apart as possible.
[0123] Step 304, based on the clustering loss of the rights and interests prediction model, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively train the rights and interests prediction model to obtain the target model.
[0124] In one embodiment, as Figure 4 shown, after determining the clustering loss of the rights and interests prediction model based on the first unit vector and the second unit vector, the intention learning method includes:
[0125] Step 401, update the cluster centers based on the clustering loss of the rights and interests prediction model and a preset algorithm;
[0126] In one embodiment, the preset algorithm refers to the backpropagation algorithm.
[0127] In one embodiment, the user's behavior sequences are embedded and grouped into various clusters, which represent the user's potential intents or interests. Among them, some clusters may represent the user's preference for discount offers, and other clusters may indicate the user's interest in exclusive events or membership privileges. The center c of each cluster i is a d-dimensional vector representing the features of this type of intent. The user behavior embedding h u is assigned to the nearest cluster according to its distance from each cluster center. In this way, similar behaviors are assigned to the same cluster, forming multiple clusters, and each cluster contains behavior embeddings with similar intents.
[0128] In one embodiment, through the backpropagation algorithm, the network parameters and cluster centers are continuously updated according to the magnitude of the clustering loss, so that the clustering loss gradually decreases and the behavior sequence embeddings gradually converge to appropriate clusters.
[0129] Step 402, perform sequential augmentation processing on the first user behavior sequence to obtain a first user behavior augmented sequence;
[0130] In one embodiment, the sequential augmentation processing at least includes processing such as masking, cropping, and reordering.
[0131] In one embodiment, the first user behavior augmented sequence is a new view of the first user behavior sequence. For example, the two views of the behavior sequence of user u are respectively represented as and
[0132] In one embodiment, by performing sequential augmentation processing on the first user behavior sequence to obtain a first user behavior augmented sequence, contrastive learning can be further performed between behavior sequences.
[0133] Step 403, based on the cluster center, the first user behavior sequence, and the first user behavior augmented sequence, determine an intent-assisted contrastive learning loss, where the intent-assisted contrastive learning loss includes a behavior sequence contrast loss and an intent contrast loss;
[0134] In one embodiment, based on the first user behavior sequence and the first user behavior augmented sequence, determine a behavior contrast loss, and based on the first user behavior sequence and the cluster center, determine an intent contrast loss.
[0135] In one embodiment, the behavior embeddings are respectively obtained through the behavior encoder F, representing the two views and of the behavior sequence of user u, corresponding behavior embeddings.
[0136] In one embodiment, the mathematical expression for determining the behavioral sequence contrast loss is represented by the following formula:
[0137]
[0138] Wherein, is the behavioral sequence contrast loss of user u, sim(g) represents the dot product similarity, neg represents the negative sample pair, and the same sequences with different augmentations are regarded as the positive sample pair, and other sample pairs are regarded as the negative sample pair.
[0139] In one embodiment, by minimizing the sequence contrast loss Similar behaviors are pulled together, and other behaviors are pushed apart from each other, thereby enhancing the representational ability of user behaviors. In the scenario of precise recommendation for premium customers, in this way, the behavior patterns of users can be captured, such as the use of premiums and interest behaviors of users at different time periods. At this time, each behavior embedding in the behavioral sequence represents the specific operation of the user at a certain moment (such as viewing a certain premium, using a certain service, etc.), and these embeddings can well summarize the behavioral characteristics of each user at different times.
[0140] In one embodiment, the mathematical expression for determining the behavioral intention contrast loss is represented by the following formula:
[0141]
[0142] Wherein, is the dual-view behavioral embedding of user u, represents all negative behavior-intention pairs in the pairings, and the behavior embedding and the corresponding nearest intention center are regarded as the positive pair, and other pairings are regarded as the negative pair.
[0143] In one embodiment, by minimizing the intention contrast loss, behaviors with the same intention are pulled together, while behaviors with different intentions are pushed apart. Through intention contrast learning, the behavioral embedding of the user can be better matched with its potential intention center. When the user uses a certain type of premium, its behavioral embedding will be closer to the intention center representing this type of premium, thereby improving the accuracy of recommendation.
[0144] In one embodiment, the intention-assisted contrast learning loss is determined by the behavioral sequence contrast loss and the intention contrast loss.
[0145] The clustering loss based on the premium prediction model aims to minimize the overall loss of the premium prediction model, and iteratively trains the premium prediction model to obtain the target model, including:
[0146] Based on the clustering loss of the premium prediction model and the intention-assisted contrast learning loss, aiming to minimize the overall loss of the premium prediction model, iteratively train the premium prediction model to obtain the target model.
[0147] In one embodiment, as Figure 5 shown, based on the clustering center, the first user behavior sequence, and the first user behavior enhancement sequence, determining the intent-assisted contrastive learning loss includes:
[0148] Step 501, based on the clustering center, the first user behavior sequence, and the first user behavior enhancement sequence, determining a first index corresponding to the first user behavior sequence and a second index corresponding to the first user behavior enhancement sequence;
[0149] In one embodiment, the updated clustering center C ∈ R k×d is used as a self-supervised signal.
[0150] In one embodiment, the mathematical expression for querying the index of the assigned cluster is expressed by the following formula:
[0151]
[0152] where c i ∈ R 1×d represents the embedding of the i-th cluster (intent) center.
[0153] In one embodiment, the index of is obtained in the same way as obtaining the index. of
[0154] Step 502, based on the first user behavior sequence and the first index or the first user behavior enhancement sequence and the second index, obtaining a behavior fusion sequence and a behavior fusion enhancement sequence according to a first fusion strategy, where the first fusion strategy is used to fuse the latent intent of the user into the first user behavior sequence and the first user behavior enhancement sequence;
[0155] In one embodiment, the first fusion strategy can be concatenation fusion or displacement fusion. Specifically, the mathematical expression of concatenation fusion is The mathematical expression of displacement fusion is where, on the left side of the equal sign, is the behavior fusion sequence, and on the right side of the equal sign, is the first user behavior sequence.
[0156] In one embodiment, the latent intent of the user is fused into the first user behavior enhancement sequence using the same fusion strategy as used.
[0157] Step 503: Determine the behavior sequence contrast loss based on the behavior fusion sequence and the behavior fusion enhancement sequence, and determine the intent contrast loss based on the behavior fusion sequence and the cluster center;
[0158] In one embodiment, substitute the behavior fusion sequence and the behavior fusion enhancement sequence into the expression of the behavior sequence contrast loss in the foregoing step 403 to obtain the final behavior sequence contrast loss.
[0159] In one embodiment, by determining the behavior sequence contrast loss based on the behavior fusion sequence and the behavior fusion enhancement sequence, after fusing the intent information into the user behavior, by minimizing L seq_cl Train the neural network. In the scenario of accurate recommendation for privileged customers, the fusion of intent information can more accurately capture the potential needs and interests of users. For example, the use of privileges and interest behaviors of users not only reflect their current needs, but also can determine the intent contrast loss based on the behavior fusion sequence and the cluster center, and further associate to potential interests through the intent center, thereby improving the accuracy of recommendations.
[0160] In one embodiment, by determining the intent contrast loss based on the behavior fusion sequence and the cluster center, the intent learning and sequence representation learning can be further coordinated.
[0161] Step 504: Sum the behavior sequence contrast loss and the intent contrast loss to obtain the intent-assisted contrast learning loss.
[0162] In one embodiment, the mathematical expression for determining the intent-assisted contrast learning loss is represented by the following formula:
[0163] L icl = L seq_cl + L intent_cl ;
[0164] where, L seq_cl and L intent_cl are respectively the behavior sequence contrast loss determined based on the behavior fusion sequence and the behavior fusion enhancement sequence, and the intent contrast loss determined based on the behavior fusion sequence, the behavior fusion enhancement sequence and the cluster center.
[0165] In one embodiment, by determining the intent-assisted contrast learning loss, the mutual promotion between behavior learning and clustering is further enhanced, which can effectively improve the accuracy of user behavior representation and the relevance of recommendation results in the accurate recommendation system for privileged customers, thus bringing positive effects. This method can not only improve the prediction performance of the model, but also enhance the user experience, increase user stickiness and satisfaction.
[0166] In one embodiment, such as Figure 6As shown, after summing the behavioral sequence contrast loss and the intention contrast loss to obtain the intention-assisted contrast learning loss, the intention learning method includes:
[0167] Step 601, obtain data representations, where the data representations include the representation of the first user's behavioral sequence and the representation of the rights and interests information;
[0168] In one embodiment, a Recurrent Neural Network (RNN) can be used as the user behavioral sequence model to extract the representation of the first user's behavioral sequence data, capture the temporal dependencies in the first user's behavioral sequence, and learn the behavioral patterns and preferences of the first user.
[0169] In one embodiment, a Convolutional Neural Network (CNN) can be used as the rights and interests information model to extract the representation of the rights and interests information and capture the local structures and features in the rights and interests information.
[0170] Step 602, connect the representation of the first user's behavioral sequence and the representation of the rights and interests information to obtain a combined representation;
[0171] In one embodiment, the data representations of the aforementioned trained user behavioral sequence model and rights and interests information model are combined. Specifically, the representation of the first user's behavioral sequence and the representation of the rights and interests information can be connected through a fully connected layer to obtain a combined representation.
[0172] Step 603, based on the combined representation, determine the next right and interest corresponding to the combined representation from the mapping relationship between the combined representation and the next right and interest, and determine the prediction loss corresponding to the next right and interest;
[0173] In one embodiment, the mapping relationship between the combined representation and the next right and interest can be embodied as a mapping relationship table, a mapping relationship curve, or a mapping probability distribution.
[0174] In one embodiment, taking the mapping relationship between the combined representation and the next right and interest as a mapping probability distribution as an example, an output layer is added on top of the combined model to predict the next right and interest that the first user may be interested in. Preferably, the output layer in this application is a fully connected layer that maps the combined representation to the probability distribution of the target right and interest, and takes the right and interest corresponding to the maximum probability as the next right and interest that the first user may be interested in.
[0175] In one embodiment, by calculating the difference between the prediction result of the next right and interest that the first user may be interested in and the true label, the prediction loss corresponding to the next right and interest is obtained to measure the prediction accuracy of the model.
[0176] Based on the clustering loss of the rights and interests prediction model and the intention-assisted contrastive learning loss, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively training the rights and interests prediction model to obtain a target model, including:
[0177] Based on the clustering loss of the rights and interests prediction model, the intention-assisted contrastive learning loss, and the prediction loss corresponding to the next right and interest, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively training the rights and interests prediction model to obtain a target model.
[0178] In one embodiment, as Figure 7 shown, Figure 7 As the architecture diagram of the intention learning method provided by the embodiments of the present disclosure, it can be seen that the overall loss of the rights and interests prediction model is jointly determined by the clustering loss, the intention-assisted contrastive learning loss, and the prediction loss corresponding to the next right and interest. Among them, the clustering loss is used to measure the understanding of the user's potential intention, the intention-assisted contrastive learning loss is used to measure the mutual cooperation between intention learning and behavior learning, and the prediction loss corresponding to the next right and interest is a common task in the recommendation system.
[0179] In one embodiment, the mathematical expression for determining the overall loss of the rights and interests prediction model is represented by the following formula:
[0180] L overall = L next_item + 0.1×L icl + α×L cluster
[0181] where, L next_item is the prediction loss corresponding to the next right and interest, L icl is the intention-assisted contrastive learning loss, L cluster is the clustering loss, α is a trade-off hyperparameter, and L overall is the overall loss of the rights and interests prediction model.
[0182] In one embodiment, combining the end-to-end learnable clustering model and the intention-assisted contrastive learning constitutes an end-to-end optimization framework for user intention learning, improving the performance and convenience of intention learning.
[0183] In summary, the solution provided by the present disclosure:
[0184] First, through the use of an end-to-end learnable clustering model based on the user's behavior sequence, overall end-to-end optimization is achieved, effectively clustering the user behavior data, providing more accurate user interest and intention information for the system, thus improving the accuracy of the recommendation system. By introducing intention-assisted contrast learning, the interaction between intention learning and behavior learning is enhanced, enabling better capture of the user's latent intentions, improving the performance and efficiency of the recommendation system, and solving the problems of getting stuck in local optimal solutions and slow algorithm convergence speed caused by step-by-step processing based on the EM framework in related solutions.
[0185] Secondly, by aggregating the behavior sequence embeddings of the user at different times, a first user behavior vector is obtained, which can summarize the behavior of each user at different times.
[0186] Thirdly, by determining the behavior sequence contrast loss based on the behavior fusion sequence and the behavior fusion enhancement sequence, and determining the intention contrast loss based on the behavior fusion sequence and the clustering center, the latent needs and interests of the user can be captured more accurately, thereby improving the accuracy of rights and interests recommendation.
[0187] The following uses an application example to further illustrate the intention learning method provided by the present disclosure:
[0188] As Figure 8 shown, Figure 8 is a schematic flowchart of the intention learning method provided by the application example of the present disclosure. The intention learning method provided by the application example of the present disclosure includes the following steps:
[0189] Step 801, obtain a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of the first user at corresponding times;
[0190] Step 802, use an encoder to embed the first user behavior sequence into the latent space to obtain the behavior sequence embeddings of the user at different times;
[0191] Step 803, perform aggregation processing on the behavior sequence embeddings of the user at different times to obtain a first user behavior vector;
[0192] Step 804, initialize the clustering center as learnable neural network parameters;
[0193] Step 805, perform normalization processing on the first user behavior vector and the clustering center to obtain a first unit vector corresponding to the first user behavior vector and a second unit vector corresponding to the clustering center;
[0194] Step 806: Based on the first unit vector and the second unit vector, determine the clustering loss of the equity prediction model, where the clustering loss of the equity prediction model includes decoupling loss and alignment loss;
[0195] Step 807: Based on the clustering loss of the equity prediction model and a preset algorithm, update the cluster centers;
[0196] Step 808: Perform sequential enhancement processing on the first user behavior sequence to obtain a first enhanced user behavior sequence;
[0197] Step 809: Based on the cluster centers, the first user behavior sequence, and the first enhanced user behavior sequence, determine a first index corresponding to the first user behavior sequence and a second index corresponding to the first enhanced user behavior sequence;
[0198] Step 810: Based on the first user behavior sequence and the first index or the first enhanced user behavior sequence and the second index, obtain a behavior fusion sequence and a behavior fusion enhanced sequence according to a first fusion strategy, where the first fusion strategy is used to fuse the potential intention of the user into the first user behavior sequence and the first enhanced user behavior sequence;
[0199] Step 811: Based on the behavior fusion sequence and the behavior fusion enhanced sequence, determine a behavior sequence contrast loss, and based on the behavior fusion sequence and the cluster centers, determine an intention contrast loss;
[0200] Step 812: Sum the behavior sequence contrast loss and the intention contrast loss to obtain an intention-assisted contrast learning loss;
[0201] Step 813: Obtain data representations, where the data representations include the representation of the first user behavior sequence and the representation of equity information;
[0202] Step 814: Concatenate the representation of the first user behavior sequence and the representation of equity information to obtain a combined representation;
[0203] Step 815: Based on the combined representation, determine the next equity corresponding to the combined representation from the mapping relationship between the combined representation and the next equity, and determine the prediction loss corresponding to the next equity;
[0204] Step 816: Based on the clustering loss of the equity prediction model, the intention-assisted contrast learning loss, and the prediction loss corresponding to the next equity, with the goal of minimizing the overall loss of the equity prediction model, perform iterative training on the equity prediction model to obtain a target model, where the target model is an end-to-end learnable clustering model;
[0205] Step 817: Obtain a second user behavior sequence, and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
[0206] To implement the intention learning method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide an intention learning device, as Figure 9 shown. Figure 9 FIG. is a schematic structural diagram of the intention learning device provided in the embodiments of the present disclosure. The intention learning device 900 includes:
[0207] An obtaining unit 901, configured to obtain a first user behavior sequence, where the first user behavior sequence is used to reflect the specific operations of the first user at corresponding times;
[0208] An embedding unit 902, configured to embed the first user behavior sequence into a latent space to obtain a first user behavior vector;
[0209] A training unit 903, configured to iteratively train the rights and interests prediction model based on the first user behavior vector with the goal of minimizing the overall loss of the rights and interests prediction model, to obtain a target model, where the target model is an end-to-end learnable clustering model;
[0210] A determining unit 904, configured to obtain a second user behavior sequence, and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
[0211] In one embodiment, the embedding unit 902 is specifically configured to:
[0212] Use an encoder to embed the first user behavior sequence into a latent space to obtain behavior sequence embeddings of the user at different times;
[0213] Perform an aggregation process on the behavior sequence embeddings of the user at different times to obtain a first user behavior vector.
[0214] In one embodiment, the training unit 903 is specifically configured to:
[0215] Initialize the cluster centers as learnable neural network parameters;
[0216] Perform a normalization process on the first user behavior vector and the cluster centers to obtain a first unit vector corresponding to the first user behavior vector and a second unit vector corresponding to the cluster centers;
[0217] Based on the first unit vector and the second unit vector, determine the clustering loss of the equity prediction model, where the clustering loss of the equity prediction model includes a decoupling loss and an alignment loss;
[0218] Based on the clustering loss of the equity prediction model, with the goal of minimizing the overall loss of the equity prediction model, perform iterative training on the equity prediction model to obtain a target model.
[0219] In one embodiment, the intent learning device 900 further includes an intent-assisted contrast learning loss determination unit, and the intent-assisted contrast learning loss determination unit is configured to:
[0220] Update the cluster center based on the clustering loss of the equity prediction model and a preset algorithm;
[0221] Perform sequential augmentation processing on the first user behavior sequence to obtain a first user behavior augmented sequence;
[0222] Based on the cluster center, the first user behavior sequence, and the first user behavior augmented sequence, determine an intent-assisted contrast learning loss, where the intent-assisted contrast learning loss includes a behavior sequence contrast loss and an intent contrast loss;
[0223] In one embodiment, the training unit 903 is specifically configured to:
[0224] Based on the clustering loss of the equity prediction model and the intent-assisted contrast learning loss, with the goal of minimizing the overall loss of the equity prediction model, perform iterative training on the equity prediction model to obtain a target model.
[0225] In one embodiment, the intent-assisted contrast learning loss determination unit is further configured to:
[0226] Based on the cluster center, the first user behavior sequence, and the first user behavior augmented sequence, determine a first index corresponding to the first user behavior sequence and a second index corresponding to the first user behavior augmented sequence;
[0227] Based on the first user behavior sequence and the first index or the first user behavior augmented sequence and the second index, obtain a behavior fusion sequence and a behavior fusion augmented sequence according to a first fusion strategy, where the first fusion strategy is used to fuse the latent intent of the user into the first user behavior sequence and the first user behavior augmented sequence;
[0228] Based on the behavior fusion sequence and the behavior fusion augmented sequence, determine a behavior sequence contrast loss, and based on the behavior fusion sequence and the cluster center, determine an intent contrast loss;
[0229] Sum the behavior sequence contrast loss and the intention contrast loss to obtain the intention-assisted contrast learning loss.
[0230] In one embodiment, the intention learning device 900 further includes a prediction loss determination unit, and the prediction loss determination unit is configured to:
[0231] Obtain a data representation, where the data representation includes a representation of a first user behavior sequence and a representation of rights and interests information;
[0232] Connect the representation of the first user behavior sequence and the representation of rights and interests information to obtain a combined representation;
[0233] Based on the combined representation, determine the next rights and interests corresponding to the combined representation from the mapping relationship between the combined representation and the next rights and interests, and determine the prediction loss corresponding to the next rights and interests;
[0234] In one embodiment, the training unit 903 is specifically configured to:
[0235] Based on the clustering loss of the rights and interests prediction model, the intention-assisted contrast learning loss, and the prediction loss corresponding to the next rights and interests, with the goal of minimizing the overall loss of the rights and interests prediction model, iteratively train the rights and interests prediction model to obtain a target model.
[0236] It should be noted that: when the intention learning device provided in the above embodiment performs intention learning, only the above-mentioned division of each program module is used for illustration. In practical applications, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the intention learning device is divided into different program modules to complete all or part of the above-described processing. In addition, the intention learning device provided in the above embodiment and the intention learning method embodiment provided in the present disclosure belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0237] Figure 10 This is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present disclosure. As Figure 10 shown, the electronic device 1000 includes at least one processor 1002; and a memory 1001 communicatively connected to at least one processor 1002; wherein, the memory 1001 stores instructions executable by at least one processor 1002, and the instructions are executed by at least one processor 1002 to implement the steps of the intention learning method in the embodiment of the present disclosure.
[0238] Optionally, the electronic device may specifically be the intention learning device in the embodiment of the present application, and the electronic device can implement the corresponding processes implemented by the intention learning device in each method of the embodiment of the present application. For the sake of brevity, it will not be elaborated here.
[0239] It can be understood that the electronic device further includes a communication interface 1003. Each component in the electronic device is coupled together through a bus system 1004. It can be understood that the bus system 1004 is used to implement connection and communication between these components. In addition to including a data bus, the bus system 1004 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 10 all kinds of buses are labeled as the bus system 1004.
[0240] It can be understood that the memory 1001 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 1001 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0241] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 1002. The processor 1002 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above methods can be completed by the integrated logic circuit in the hardware of the processor 1002 or instructions in the form of software. The above-mentioned processor 1002 may be a general-purpose processor, a DSP (Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1002 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the methods disclosed in the embodiments of the present invention, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, which is located in the memory 1001. The processor 1002 reads the information in the memory 1001 and combines its hardware to complete the steps of the foregoing methods.
[0242] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, ProgrammableLogic Device), complex programmable logic devices (CPLDs, Complex Programmable Logic Device), FPGAs (FieldProgrammable Gate Array, Field Programmable Gate Array), general-purpose processors, controllers, MCUs (Microcontroller Unit, Microcontroller Unit), microprocessors (Microprocessor), or other electronic components, and is used to execute the foregoing methods.
[0243] This embodiment also provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to enable the computer to execute the steps of the intention learning method in the embodiments of the present invention when executed.
[0244] Optionally, the computer-readable storage medium can be applied to the intention learning device in the embodiments of the present application, and the computer instructions enable the computer to execute the corresponding processes implemented by the intention learning device in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0245] An embodiment of the present disclosure also provides a computer program product, including a computer program, which implements the steps of the intention learning method provided by the embodiment of the present invention when executed by a processor.
[0246] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the displayed or discussed components can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0247] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0248] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0249] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0250] Alternatively, if the above-integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0251] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An intention learning method, characterized in that: include: Acquire a first user behavior sequence, where the first user behavior sequence is used to reflect specific operations of the first user at a corresponding time; Embedding the first user behavior sequence into a latent space to obtain a first user behavior vector; Based on the first user behavior vector, with the goal of minimizing the overall loss of the equity prediction model, iteratively training the equity prediction model to obtain a target model, where the target model is an end-to-end learnable clustering model; A second user behavior sequence is obtained, and the target model is used to determine the equity corresponding to the second user behavior sequence, where the equity corresponding to the second user behavior sequence is used to indicate the equity that the user may be interested in next.
2. The method according to claim 1, characterized in that The embedding the first user behavior sequence into a latent space to obtain a first user behavior vector includes: Embedding the first user behavior sequence into a latent space using an encoder to obtain embeddings of the user's behavior sequence at different times; Aggregate the behavior sequence embeddings of the user at different times to obtain a first user behavior vector.
3. The method according to claim 1, characterized in that The method of iteratively training the equity prediction model based on the first user behavior vector and taking minimizing the overall loss of the equity prediction model as a goal to obtain a target model includes: Initialize the cluster centers as learnable neural network parameters; Normalizing the first user behavior vector and the cluster center to obtain a first unit vector corresponding to the first user behavior vector and a second unit vector corresponding to the cluster center; Determine a clustering loss of a equity prediction model based on the first unit vector and the second unit vector, wherein the clustering loss of the equity prediction model includes a decoupling loss and an alignment loss; Based on the clustering loss of the equity prediction model, the equity prediction model is iteratively trained with the goal of minimizing the overall loss of the equity prediction model to obtain a target model.
4. The method according to claim 3, characterized in that After determining the clustering loss of the equity prediction model based on the first unit vector and the second unit vector, the method includes: Based on the clustering loss of the equity prediction model and a preset algorithm, updating the cluster center; Performing sequential enhancement processing on the first user behavior sequence to obtain a first user behavior enhanced sequence; Determine an intention-assisted contrastive learning loss based on the cluster center, the first user behavior sequence, and the first user behavior enhancement sequence, where the intention-assisted contrastive learning loss includes a behavior sequence contrast loss and an intention contrast loss; The clustering loss based on the equity prediction model is aimed at minimizing the overall loss of the equity prediction model, and the equity prediction model is iteratively trained to obtain a target model, including: Based on the clustering loss of the equity prediction model and the intention-assisted contrastive learning loss, the equity prediction model is iteratively trained with the goal of minimizing the overall loss of the equity prediction model to obtain a target model.
5. The method according to claim 4, characterized in that The determining, based on the cluster center, the first user behavior sequence, and the first user behavior enhancement sequence, the intention-assisted contrastive learning loss includes: Based on the cluster center, the first user behavior sequence and the first user behavior enhancement sequence, determining a first index corresponding to the first user behavior sequence and a second index corresponding to the first user behavior enhancement sequence; Based on the first user behavior sequence and the first index or the first user behavior enhancement sequence and the second index, obtaining a behavior fusion sequence and a behavior fusion enhancement sequence according to a first fusion strategy, wherein the first fusion strategy is used to fuse the user's potential intention into the first user behavior sequence and the first user behavior enhancement sequence; Determine a behavior sequence contrast loss based on the behavior fusion sequence and the behavior fusion enhancement sequence, and determine an intention contrast loss based on the behavior fusion sequence and the cluster center; The behavior sequence contrast loss and the intention contrast loss are summed to obtain the intention-assisted contrast learning loss.
6. The method according to claim 5, characterized in that After the behavior sequence contrast loss and the intention contrast loss are summed to obtain the intention-assisted contrast learning loss, the method includes: Acquire a data representation, the data representation comprising a representation of the first user behavior sequence and a representation of equity information; Connecting the representation of the first user behavior sequence and the representation of the equity information to obtain a combined representation; Based on the combined representation, determining a next equity corresponding to the combined representation from a mapping relationship between the combined representation and the next equity, and determining a predicted loss corresponding to the next equity; The clustering loss based on the equity prediction model and the intention-assisted contrastive learning loss are used to iteratively train the equity prediction model with the goal of minimizing the overall loss of the equity prediction model to obtain a target model, including: Based on the clustering loss of the equity prediction model, the intention-assisted contrastive learning loss and the prediction loss corresponding to the next equity, the equity prediction model is iteratively trained with the goal of minimizing the overall loss of the equity prediction model to obtain a target model.
7. An intention learning device, characterized in that: include: An acquisition unit, configured to acquire a first user behavior sequence, where the first user behavior sequence is used to reflect a specific operation of the first user at a corresponding time; an embedding unit, configured to embed the first user behavior sequence into a latent space to obtain a first user behavior vector; a training unit, configured to iteratively train the equity prediction model based on the first user behavior vector and with the goal of minimizing the overall loss of the equity prediction model, to obtain a target model, wherein the target model is an end-to-end learnable clustering model; The determination unit is used to obtain a second user behavior sequence and use the target model to determine the rights and interests corresponding to the second user behavior sequence, where the rights and interests corresponding to the second user behavior sequence are used to indicate the rights and interests that the user may be interested in next.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by a processor.