Sample sampling method and device, electronic equipment and computer readable storage medium

CN122594840APending Publication Date: 2026-08-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179591.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

但是上述方案均未能在跨域序列推荐场景下有效缓解假负样本问题,从而难以对硬负样本进行充分利用

Benefits of technology

[0035]This application embodiment obtains a behavior sequence generated by an interactive object interacting with at least one sample in the current domain sample set, where the current domain sample set includes a source domain sample set and a target domain sample set. Then, based on the behavior sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. Next, a sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Then, based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the target domain sample set, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Finally, at least one target negative sample is sampled from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity. This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences using target domain sample sets and source domain sample sets. It then trains a sequence recommendation model using these behavioral sequences as training samples. Based on this model, it determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, it comprehensively utilizes both relative popularity and source domain interest preferences to correct target domain interest preferences, thereby uncovering hard negative samples in the target domain sample set. This alleviates the false negative sample problem in negative sampling and improves the sampling quality of hard negative samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594840A_ABST
    Figure CN122594840A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a sample sampling method and device, electronic equipment and computer readable storage medium; the embodiments of the present application construct cross-domain sequence recommendation scene and the behavior sequence corresponding to cross-domain sequence recommendation scene, the behavior sequence corresponding to cross-domain sequence recommendation scene is used as training sample to train a sequence recommendation model, then, the interest preference of interactive object in source domain and target domain respectively for candidate negative sample in target domain sample set is determined according to the sequence recommendation model, at the same time, the relative popularity of candidate negative sample relative to target positive sample is determined, finally, the relative popularity and source domain interest preference are comprehensively utilized, the target domain interest preference is corrected to mine hard negative sample in target domain sample set, the false negative example problem in negative sampling is alleviated, and the sampling quality of hard negative sample is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically to a sample sampling method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Cross-domain sequence recommendation aims to model dynamic user preferences based on users' sequential behaviors across multiple domains, achieving better recommendation performance in the target domain by leveraging information from the source domain. Generally, cross-domain sequence recommendation systems understand user preferences based on observed implicit feedback (e.g., clicks, purchases). For this purpose, the recommendation model needs to be trained using high-quality positive and negative samples. However, since implicit feedback data typically only contains positive samples, employing negative sampling techniques to provide contrasting information for a more comprehensive understanding of user preferences is crucial.

[0003] In related technologies, uniform random sampling is typically performed from a set of items that the user has not interacted with, or additional information is used to assist in sampling. However, none of these solutions have been able to effectively mitigate the problem of spurious negative samples in cross-domain sequence recommendation scenarios, making it difficult to fully utilize hard negative samples. Summary of the Invention

[0004] This application provides a sample sampling method, apparatus, electronic device, and computer-readable storage medium, which can alleviate the problem of false negatives in the hard negative sample sampling process and improve the sampling quality of hard negative samples.

[0005] This application provides a sample collection method, including:

[0006] Obtain the sequence of behaviors generated by the interaction object interacting with at least one sample in the current domain sample set, wherein the current domain sample set includes a source domain sample set and a target domain sample set;

[0007] Based on the behavior sequence, at least one sample that has not interacted with the interaction object is selected from the target domain sample set to obtain candidate negative samples;

[0008] A sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence.

[0009] Based on the behavioral sequence, target positive samples corresponding to the candidate negative samples are selected from the target domain sample set, and the relative popularity of the candidate negative samples is determined, wherein the relative popularity indicates the difference in frequency of selection of the candidate negative samples relative to the target positive samples;

[0010] Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples.

[0011] Accordingly, embodiments of this application provide a sample sampling device, including:

[0012] The acquisition unit is used to acquire the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set, wherein the current domain sample set includes a source domain sample set and a target domain sample set.

[0013] A filtering unit is used to filter at least one sample that has not interacted with the interaction object from the target domain sample set based on the behavior sequence, so as to obtain candidate negative samples;

[0014] The prediction unit is used to predict the first preference probability of the interactive object in the target domain for the candidate negative sample and the second preference probability of the interactive object in the source domain for the candidate negative sample using a sequence recommendation model, wherein the sequence recommendation model is trained based on the behavior sequence.

[0015] The determining unit is configured to, based on the behavioral sequence, filter out the target positive sample corresponding to the candidate negative sample in the target domain sample set, and determine the relative popularity of the candidate negative sample, wherein the relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample;

[0016] A sampling unit is configured to sample at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity.

[0017] In some embodiments, the acquisition unit may be specifically used to acquire the source domain behavior sequence generated by the interaction object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence generated by the interaction object interacting with at least one sample in the target domain sample set; and to merge the source domain behavior sequence and the target domain behavior sequence to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set.

[0018] In some embodiments, the filtering unit may be specifically used to filter out target domain samples that do not have an interaction relationship with the interaction object from the target domain sample set based on the target domain behavior sequence, to obtain a non-interactive target domain sample set; and to select at least one non-interactive target domain sample from the non-interactive target domain sample set to obtain the candidate negative sample.

[0019] In some embodiments, the prediction unit may be specifically used to obtain sample feature data of the candidate negative sample in the current domain sample set; use the sequence recommendation model to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain; and determine the first preference probability and the second preference probability based on the sample feature data, the target domain preference data and the source domain preference data.

[0020] In some embodiments, the prediction unit may be specifically used to predict the target domain preference data of the interactive object in the target domain based on the target domain behavior sequence and using the sequence recommendation model; and to predict the source domain preference data of the interactive object in the source domain based on the source domain behavior sequence and using the sequence recommendation model.

[0021] In some embodiments, the prediction unit may be specifically used to fuse the sample feature data and the target domain preference data to obtain the first preference probability; and to fuse the sample feature data and the source domain preference data to obtain the second preference probability.

[0022] In some embodiments, the determining unit may be specifically used to obtain the positive sample popularity of the target positive sample and the negative sample popularity of the candidate negative sample, wherein the positive sample popularity indicates the frequency at which the target positive sample is selected and the negative sample popularity indicates the frequency at which the candidate negative sample is selected; compare the positive sample popularity and the negative sample popularity to obtain a popularity comparison result; and determine the relative popularity of the candidate negative sample based on the popularity comparison result.

[0023] In some embodiments, the determining unit may be specifically used to determine the relative popularity based on the difference between the negative sample popularity and the positive sample popularity when the popularity comparison result indicates that the popularity of the positive sample is less than the popularity of the negative sample; and to set the relative popularity to 0 when the popularity comparison result indicates that the popularity of the positive sample is greater than or equal to the popularity of the negative sample.

[0024] In some embodiments, the sampling unit may be specifically used to adjust the probability value of the first preference probability according to the second preference probability and the relative popularity to obtain an adjusted preference probability; and based on the adjusted preference probability, to select at least one target negative sample from the candidate negative samples.

[0025] In some embodiments, the sample sampling device may further include a training unit, which is used to use the target negative sample as the current training sample of the sequence recommendation model in the current training round; and to update the model parameters of the sequence recommendation model based on the current training sample to obtain the updated sequence recommendation model.

[0026] In some embodiments, the training unit may be specifically used to obtain a first sampling probability for the candidate negative sample and a second sampling probability for the target negative sample in the next training round of the current training round; based on the target sample sampling probability corresponding to the next training round, to select the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative sample or the target negative sample, wherein the target sample sampling probability includes at least one of the first sampling probability and the second sampling probability; and to update the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained.

[0027] In some embodiments, the training unit may be specifically configured to use the first sampling probability of the candidate negative sample in the current training round as a reference sampling probability, and determine the probability update parameter of the next training round for the reference sampling probability; calculate the first sampling probability of the candidate negative sample in the next training round based on the reference sampling probability and the probability update parameter; and calculate the second sampling probability of the target negative sample in the next training round based on the first sampling probability.

[0028] In some embodiments, the training unit may be specifically used to obtain the model training loss value corresponding to at least one training round before the next training round; and to determine the probability update parameter of the reference sampling probability based on the model training loss value.

[0029] In some embodiments, the training unit may be specifically used to obtain a random probability value, compare the random probability value with the first sampling probability; if the comparison result shows that the random probability value is less than the first sampling probability, select the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples; if the comparison result shows that the random probability value is greater than or equal to the first sampling probability, select the next training sample corresponding to the updated sequence recommendation model in the next training round from the target negative samples.

[0030] In some embodiments, the training unit may be specifically used to sample at least one target negative sample from the candidate negative samples based on the updated sequence recommendation model, so as to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round.

[0031] In some embodiments, the training unit may be specifically used to update the target negative sample in the current training round based on the next training sample to obtain the updated target negative sample; and use the updated target negative sample as the target negative sample corresponding to the next training round.

[0032] Furthermore, this application also provides an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to execute the sample sampling method provided in this application.

[0033] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the sample sampling method provided in embodiments of this application.

[0034] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the sample sampling methods provided in embodiments of this application.

[0035] This application embodiment obtains a behavior sequence generated by an interactive object interacting with at least one sample in the current domain sample set, where the current domain sample set includes a source domain sample set and a target domain sample set. Then, based on the behavior sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. Next, a sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Then, based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the target domain sample set, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Finally, at least one target negative sample is sampled from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity. This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences using target domain sample sets and source domain sample sets. It then trains a sequence recommendation model using these behavioral sequences as training samples. Based on this model, it determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, it comprehensively utilizes both relative popularity and source domain interest preferences to correct target domain interest preferences, thereby uncovering hard negative samples in the target domain sample set. This alleviates the false negative sample problem in negative sampling and improves the sampling quality of hard negative samples. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a schematic diagram illustrating an application scenario of the sample sampling method provided in the embodiments of this application;

[0038] Figure 2 This is a schematic flowchart of the sample sampling method provided in the embodiments of this application;

[0039] Figure 3 This is a schematic diagram of the process for determining the sequence recommendation model provided in the embodiments of this application;

[0040] Figure 4 This is a schematic diagram illustrating sample sampling using a sequence recommendation model, provided in an embodiment of this application.

[0041] Figure 5 This is a schematic diagram illustrating the utilization of target negative samples provided in an embodiment of this application;

[0042] Figure 6 This is a schematic diagram illustrating how the first and second sampling probabilities, as provided in the embodiments of this application, change with the training rounds;

[0043] Figure 7 This is another schematic diagram of the sample sampling method provided in the embodiments of this application;

[0044] Figure 8 This is a schematic diagram of the sample sampling device provided in an embodiment of this application;

[0045] Figure 9 This is another schematic diagram of the sample sampling device provided in the embodiments of this application;

[0046] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0047] The technical solutions described below, with reference to the accompanying drawings, will be clearly and completely described. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0048] Before introducing the technical solution of this application, the relevant knowledge of the technical solution of this application will be explained below.

[0049] Recommender systems are broadly classified into two categories: general recommendation and sequential recommendation. Simply put, the distinction lies in whether or not time sequence needs to be considered. The former treats user preferences as static, learning static representations of users and items, while the latter assumes that user preferences change dynamically over time, predicting the next item a user might like based on the sequence of interactions.

[0050] Sequential recommendation is a recommendation system paradigm that models the patterns of interactive behaviors and items over time to recommend relevant items to users. Sequential recommendation systems model sequences of interactive behaviors, learn changes in user interests, and thus predict the user's next action. This type of recommendation system has a time limit for evaluation; it is generally considered a successful recommendation if the user visits the recommended item within a certain period after the system's recommendation; recommendations exceeding this time are considered invalid.

[0051] Cross-Domain Sequential Recommendation (CDSR) refers to personalized recommendations across multiple different domains, leveraging users' behavioral sequences across these domains. Given a user's historical interaction data, CDSR aims to predict the item the user is most likely to interact with in the target domain at the next moment. This method addresses common issues in single-domain sequential recommendation, such as data attributes and the cold start problem for new users. By integrating data from multiple domains and utilizing users' behavioral history across these domains, CDSR improves the accuracy and coverage of recommendations.

[0052] Negative sampling is a technique commonly used in machine learning and natural language processing (NLP) to handle imbalanced data in recommender systems. By selecting a subset of items from micro-interactions as negative samples, it provides negative signals to the model during training, optimizing the model's training efficiency and effectiveness. Negative sampling helps recommender models better learn user preferences, thereby improving recommendation accuracy.

[0053] Hard negative samples are samples that have a high similarity to positive samples but are actually negative. Hard negative samples are difficult to distinguish from positive samples in the embedding space. This similarity may be due to the characteristics of the data itself or biases in the model during the learning process. By introducing hard negative samples, the model needs to better learn how to distinguish these difficult-to-distinguish samples, thereby enhancing its representational and generalization abilities.

[0054] False negatives refer to the incorrect sampling of potentially positive samples as negative samples, which leads to confusion during model training and performance degradation.

[0055] The difficulty level of a sample refers to how easily or how poorly the model learns from it. Difficult and easy samples can be categorized based on the model's predictive ability during training or testing. Difficult samples typically refer to those that the model struggles to classify correctly; these samples usually have high classification difficulty, possibly due to large intra-class or small inter-class differences in their features. Easy samples, on the other hand, are those that the model can easily and correctly classify because they have high discriminative power in their features. The distinction between difficult and easy samples is related to the model's current performance; as model performance improves, previously difficult samples may become easy.

[0056] This application provides a sample sampling method, apparatus, electronic device, and computer-readable storage medium. The sample sampling apparatus can be integrated into an electronic device, which may be a server or a user terminal, etc.

[0057] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud-preset databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN) acceleration services, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0058] Figure 1 A schematic diagram illustrating an application scenario of the sample sampling method provided in this application embodiment is shown. For example... Figure 1 As shown, taking an example where the sample sampling device is integrated into an electronic device, and the electronic device is a server, the server can obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set. The current domain sample set includes a source domain sample set and a target domain sample set. Based on the behavior sequence, at least one sample that has not interacted with the interaction object is selected from the target domain sample set to obtain candidate negative samples. A sequence recommendation model is used to predict the first preference probability of the interaction object for the candidate negative sample in the target domain and the second preference probability of the interaction object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the target domain sample set, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples.

[0059] It is understood that, in the specific embodiments of this application, there are sample sets from different domains and related data such as the behavioral sequences generated by the interaction of interactive objects with at least one sample in the sample set. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0060] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0061] This embodiment will be described from the perspective of a sample sampling device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices that can generate image files.

[0062] A sample sampling method includes: acquiring a behavior sequence generated by an interactive object interacting with at least one sample in a current domain sample set, the current domain sample set including a source domain sample set and a target domain sample set; based on the behavior sequence, filtering out at least one sample in the target domain sample set that has not interacted with the interactive object to obtain candidate negative samples; employing a sequence recommendation model to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain, the sequence recommendation model being trained based on the behavior sequence; based on the behavior sequence, filtering out target positive samples corresponding to the candidate negative samples in the target domain sample set, and determining the relative popularity of the candidate negative samples, the relative popularity indicating the frequency difference of the candidate negative samples being selected relative to the target positive samples; and sampling at least one target negative sample from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity.

[0063] Figure 2 A schematic flowchart of the sample sampling method provided in an embodiment of this application is shown. Figure 2 As shown, the specific process of this sample sampling method is as follows:

[0064] 101. Obtain the sequence of actions generated by the interaction object interacting with at least one sample in the current domain sample set, where the current domain sample set includes the source domain sample set and the target domain sample set.

[0065] Among them, behavioral sequences are one of the most important types of features in a recommendation system. Examples include a user's click sequence, viewing sequence, and so on. In this embodiment, a behavioral sequence can be understood as a sequence generated by an interactive object (user) performing selection operations on samples in a sample set (such as the current domain sample set, the source domain sample set, or the target domain sample set).

[0066] In deep learning, the source domain sample set can be understood as the set of samples corresponding to the source domain. The source domain can be understood as an existing dataset or domain used to train the model, containing a large number of labeled data samples. These labeled data samples can be used to build and train the model, enabling it to learn the mapping relationship between input data and output labels.

[0067] The target domain sample set can be understood as the set of samples corresponding to the target domain. The target domain can be understood as the domain where the model will be applied, containing a small number of data samples. Typically, there are only a few labeled samples in the target domain. Therefore, transfer learning is needed to apply the knowledge and features learned from the source domain to the target domain, allowing the model to better adapt to the features and data distribution of the target domain, thereby improving the model's performance on the target task.

[0068] The source and target domains can be determined based on the specific circumstances. For example, the source and target domains can be sub-domains within various application scenarios such as content platforms, social media, and e-commerce platforms. Taking content platforms as an example, user behavior information in the book category can be used to improve the training of recommendation models in the movie or music category. Here, the book category can be understood as the source domain, and the movie or music category as the target domain. Taking social media as an example, a user's interaction history on one application (such as a public account) can help another application (such as a mini-program) improve the accuracy of recommendations. Here, one application can be understood as the source domain, and the other application as the target domain. Taking e-commerce platforms as an example, by combining interaction data from the clothing category, better recommendations can be achieved in the personal care category. Here, the interaction data from the clothing category can be understood as the source domain sample set, and the relevant data from the personal care category can be understood as the target domain sample set.

[0069] The current domain sample set can be understood as a total sample set obtained by merging the source domain sample set and the target domain sample set across domains. Therefore, the current domain sample set can also be called a hybrid domain sample set that integrates the source domain sample set and the target domain sample set.

[0070] In this embodiment of the application, obtaining the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set may include: obtaining the source domain behavior sequence generated by the interaction object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence generated by the interaction object interacting with at least one sample in the target domain sample set; merging the source domain behavior sequence and the target domain behavior sequence to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set.

[0071] The sequence of behaviors generated by an interactive object interacting with at least one sample in the current domain sample set is obtained by merging the following two behavior sequences across domains: the source domain behavior sequence and the target domain behavior sequence. The source domain behavior sequence is the sequence of behaviors generated by the interactive object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence is the sequence of behaviors generated by the interactive object interacting with at least one sample in the target domain sample set.

[0072] After obtaining the source domain behavior sequence and the target domain behavior sequence, the source domain behavior sequence and the target domain behavior sequence can be fused in chronological order to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set. Therefore, the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set can also be understood as: the complete behavior sequence of the interaction object in the mixed domain sample set.

[0073] For ease of description, the source domain sample set is represented as The target domain sample set is represented as The interaction object is represented as u. For the interaction object u, its sequence of behaviors in the source domain (i.e., the source domain behavior sequence) is represented as: 'a' represents the sequence length of the source domain behavior sequence. Similarly, the behavior sequence of the interactive object 'u' in the target domain (i.e., the target domain behavior sequence) is represented as... , where b represents the sequence length of the target domain behavior sequence. The interaction is merged chronologically to obtain a complete sequence of behaviors of the interactive objects in the hybrid domain (i.e., a behavior sequence). To simplify sequence representation, the subscript u will be omitted below. After omission, S will be used. S replace To represent the source domain behavior sequence, S is used. T replace To represent the target domain behavior sequence, S is used instead of S. u To represent a sequence of behaviors (i.e., a mixed-domain sequence of behaviors).

[0074] 102. Based on the behavior sequence, at least one sample that has not interacted with the interaction object is selected from the target domain sample set to obtain candidate negative samples.

[0075] After obtaining the behavior sequence, at least one sample that has not interacted with the interaction object can be sampled from the target domain sample set based on the behavior sequence to obtain candidate negative samples.

[0076] Specifically, based on the behavior sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. This may include: based on the target domain behavior sequence, selecting target domain samples that do not have an interaction relationship with the interactive object from the target domain sample set to obtain a set of non-interactive target domain samples; and selecting at least one non-interactive target domain sample from the set of non-interactive target domain samples to obtain candidate negative samples.

[0077] As mentioned earlier, the behavioral sequence S u Source domain behavior sequence S S and target domain behavior sequence S TThe hybrid domain behavior sequence is obtained by fusing in chronological order. The process of determining candidate negative samples in this application embodiment can be understood as sampling at least one target domain sample from the target domain sample set that has not interacted with the interaction object. Therefore, the target domain behavior sequence S can be directly used as the basis for determining the candidate negative sample. T The target domain samples that have not interacted with the interaction object (i.e., have no interaction relationship with the interaction object) are selected from the target domain sample set to obtain the non-interactive target domain sample set; then at least one non-interactive target domain sample is sampled from the non-interactive target domain sample set, and the sampled non-interactive target domain sample can be used as a candidate negative sample to obtain the candidate negative sample set.

[0078] For ease of description, the set of candidate negative samples is denoted as M, M = {j1,j2,j3,j...} m-1 ,...,j m}, where m represents the number of candidate negative samples.

[0079] 103. A sequence recommendation model is used to predict the probability of the first preference of the interactive object for the candidate negative sample in the target domain and the probability of the second preference of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence.

[0080] A sequence recommendation model is a model that can be used to perform sequence recommendation tasks. The type of sequence recommendation model involved in the embodiments of this application can be selected according to the actual situation. For example, it can be a recurrent neural network model, a convolutional neural network model, a Transformer (a neural network architecture based on a self-attention mechanism), etc.

[0081] Figure 3 A schematic diagram illustrating the process of determining the sequence recommendation model provided in an embodiment of this application is shown. Figure 3 As shown, the sequence recommendation model in this application embodiment is based on the behavior sequence S u It is obtained through training. For example, an initial sequence recommendation model can be obtained and based on the behavior sequence S. u The initial sequence recommendation model is trained to obtain the next sequence recommendation model. The initial sequence recommendation model can be a recurrent neural network, a convolutional neural network, a Transformer (a neural network architecture based on self-attention), etc. Since the next sequence recommendation model is trained based on the initial model, its model type depends on the type of the initial model chosen.

[0082] Figure 4 This illustration shows a schematic diagram of sample sampling using a sequence recommendation model, as provided in an embodiment of this application. Figure 4As shown, after obtaining the sequence recommendation model, it can be used to predict the preference probability of an interactive object for candidate negative samples in both the target and source domains. For clarity, the preference probability of an interactive object for candidate negative samples in the target domain is called the first preference probability, and the preference probability of an interactive object for candidate negative samples in the source domain is called the second preference probability. The first preference probability can be understood as the probability that a candidate negative sample is selected by the interactive object in the target domain, and the second preference probability can be understood as the probability that a candidate negative sample is selected by the interactive object in the source domain.

[0083] It should be noted that candidate negative samples originate from the target domain sample set, meaning that the target domain sample set contains candidate negative samples. However, for the source domain sample set, the source domain sample set may or may not contain the candidate negative sample. Furthermore, candidate negative samples are those selected from the target domain sample set that have not interacted with the interactive object. However, for the source domain sample set, even if the source domain sample set contains candidate negative samples, those candidate negative samples may or may not have interacted with the interactive object within the source domain sample set.

[0084] The process of using a sequence recommendation model to predict the first preference probability of an interactive object for a candidate negative sample in the target domain and the second preference probability of an interactive object for a candidate negative sample in the source domain may include: obtaining sample feature data of the candidate negative sample in the current domain sample set; using a sequence recommendation model to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain; and determining the first preference probability and the second preference probability based on the sample feature data, the target domain preference data, and the source domain preference data.

[0085] The sample feature data of a candidate negative sample in the current domain sample set can be understood as the embedding representation of that candidate negative sample in the current domain sample set, which can be a feature vector. For any candidate negative sample j in the candidate negative sample set M, it can be represented by e. j This represents the sample feature data of j.

[0086] It should be understood that after training the initial sequence recommendation model based on the behavioral sequence to obtain the sequence recommendation model, the sample feature data corresponding to each sample in the current domain sample set can be obtained. Therefore, after obtaining the sequence recommendation model, the sample feature data of the candidate negative samples can be directly obtained.

[0087] After obtaining the sequence recommendation model, it can be used to predict the target domain preference data of the interacting object in the target domain, and the source domain preference data of the interacting object in the source domain. Target domain preference data can be obtained by providing the sequence recommendation model with the target domain behavior sequence S. TSource domain preference data can be obtained by providing the source domain behavior sequence S to the sequence recommendation model. S To obtain this information, specifically, a sequence recommendation model is used to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain. This can include: predicting the target domain preference data of the interactive object in the target domain based on the target domain behavior sequence; and predicting the source domain preference data of the interactive object in the source domain based on the source domain behavior sequence.

[0088] To facilitate differentiation, the target domain preference data is represented as h. T The source domain preference data is represented as h S .

[0089] h T =f rec (S T (1-1)

[0090] h S =f rec (S S (1-2)

[0091] Among them, f rec This represents a sequence recommendation model.

[0092] After determining the sample feature data, target domain preference data, and source domain preference data, the first preference probability and the second preference probability can be calculated based on these three. Specifically, determining the first preference probability and the second preference probability based on the sample feature data, target domain preference data, and source domain preference data can include: fusing the sample feature data and target domain preference data to obtain the first preference probability; and fusing the sample feature data and source domain preference data to obtain the second preference probability.

[0093] For ease of distinction, the probability of the first preference is expressed as r(S) T Let the second preference probability be represented as r(S,j), where r(S,j) is the probability of the second preference. S The process of fusing sample feature data with target domain preference data and source domain preference data to obtain the first preference probability and the second preference probability can be represented by the following formula.

[0094] r(S T ,j)=h T ·e j (2-1)

[0095] r(S S ,j)=h S ·e j (2-2)

[0096] 104. Based on the behavioral sequence, select the target positive samples corresponding to the candidate negative samples in the target domain sample set, and determine the relative popularity of the candidate negative samples. The relative popularity indicates the difference in frequency of selection of the candidate negative samples relative to the target positive samples.

[0097] Target positive samples can be understood as samples in the target domain sample set that have a historical interaction relationship with the interacting object, i.e., samples contained in the target domain behavior sequence, which can be used as label data in the training phase of the sequence recommendation model. Therefore, the process of selecting target positive samples corresponding to candidate negative samples in the target domain sample set based on the behavior sequence can be expressed as: Based on the target domain behavior sequence S T The process involves selecting the target positive sample corresponding to the candidate negative sample from the target domain sample set. For ease of description, the target positive sample is represented as...

[0098] In sequence recommendation, the sequence recommendation model can predict the item corresponding to the next action of an interactive object based on the current interaction behavior. Therefore, when predicting behaviors at different times, the target positive sample corresponding to the candidate negative sample may be different. For example, in the target domain behavior sequence... In this context, if the prediction is made based on the item corresponding to the first action of the interacting object, then the candidate negative sample corresponds to the target positive sample. That is If the prediction is made for the item corresponding to the second action of the interactive object, then the target positive sample corresponds to the candidate negative sample. That is Similarly, if we are predicting the item corresponding to the b-th action of an interactive object, then the target positive sample corresponding to the candidate negative sample is... That is etc.

[0099] After identifying the target positive sample corresponding to the candidate negative sample, the relative popularity of the candidate negative sample relative to the target positive sample can be determined based on the sample popularity of both the target positive sample and the candidate negative sample. Sample popularity can be understood as the frequency with which the sample appears in interactive behaviors, including the number of times the sample is selected by all interactive objects in the interactive behaviors corresponding to the target domain sample set. Sample popularity can also be understood as the total number of users who have interacted with that sample.

[0100] Specifically, determining the relative popularity of candidate negative samples may include: obtaining the positive sample popularity of the target positive sample and the negative sample popularity of the candidate negative sample, where the positive sample popularity indicates the frequency of the target positive sample being selected and the negative sample popularity indicates the frequency of the candidate negative sample being selected; comparing the positive sample popularity and the negative sample popularity to obtain a popularity comparison result; and determining the relative popularity of the candidate negative sample based on the popularity comparison result.

[0101] For ease of distinction, the sample popularity of the target positive sample is called the positive sample popularity, which is expressed as: The popularity of a candidate negative sample is called the negative sample popularity, denoted as O(j). Positive sample popularity indicates the frequency with which the target positive sample is selected, i.e., the number of times the target positive sample is selected by all interacting objects in the interaction behavior corresponding to the target domain sample set. Similarly, negative sample popularity indicates the frequency with which a candidate negative sample is selected, i.e., the number of times the candidate negative sample is selected by all interacting objects in the interaction behavior corresponding to the target domain sample set. Positive sample popularity can also indicate the total number of users who have interacted with the target positive sample within the target domain sample set; conversely, negative sample popularity indicates the total number of users who have interacted with the candidate negative sample within the target domain sample set.

[0102] After obtaining the popularity of positive and negative samples, the popularity of positive and negative samples can be compared to obtain a popularity comparison result. Then, the relative popularity of candidate negative samples can be determined based on this result. Specifically, determining the relative popularity of candidate negative samples based on the popularity comparison result can include: when the popularity comparison result indicates that the popularity of positive samples is less than that of negative samples, the relative popularity is determined based on the difference between the popularity of negative and positive samples; when the popularity comparison result indicates that the popularity of positive samples is greater than or equal to that of negative samples, the relative popularity is set to 0.

[0103] In this embodiment of the application, relative popularity can be expressed as The formula for calculating relative popularity is as follows:

[0104]

[0105] 105. Based on the first preference probability, the second preference probability, and the relative popularity, sample at least one target negative sample from the candidate negative samples.

[0106] After determining the first preference probability, the second preference probability, and the relative popularity, it can be used as a basis to determine whether the candidate negative sample can be used as the target negative sample. The target negative sample can be understood as a hard negative sample selected from the target domain sample set according to the sample sampling method provided in the embodiments of this application.

[0107] In this embodiment of the application, sampling at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity may include: adjusting the probability value of the first preference probability based on the second preference probability and the relative popularity to obtain the adjusted preference probability; and selecting at least one target negative sample from the candidate negative samples based on the adjusted preference probability.

[0108] For ease of description, the adjusted preference probability can be expressed as: Furthermore, the adjusted preference probability can be calculated using the following formula.

[0109]

[0110] α represents the weight of the second preference probability. The weighting coefficient, which reflects the relative popularity of candidate negative samples, can be defined as:

[0111]

[0112] Where τ is the temperature coefficient.

[0113] Combining equations (4) and (5), it can be seen that the adjusted preference probability is obtained by introducing the second preference probability and the relative popularity of the candidate negative sample relative to the target positive sample on the basis of the first preference probability. The second preference probability and the relative popularity are used to correct the first preference probability.

[0114] After calculating the adjusted preference probabilities of the candidate negative samples, at least one target negative sample can be selected from the candidate negative samples based on the adjusted preference probabilities. The target negative sample can be understood as a hard negative sample in the target domain sample set determined by the above sampling method. For ease of description, the target negative sample can be represented as... For example, the candidate negative sample with the highest adjusted preference probability can be used as the target negative sample. Where j∈M. Alternatively, the adjusted preference probabilities corresponding to each candidate negative sample can be sorted, and the top N candidate negative samples with the highest adjusted preference probabilities can be selected as the target negative samples.

[0115] On the one hand, introducing a second preference probability to correct the first preference probability can be understood as guiding the sampling process of hard negative samples in the target domain by utilizing the preferences of the interacting object in the source domain. This approach can help distinguish between hard negative samples and false negative samples in the target domain sample set. Specifically, compared to hard negative samples, false negative samples exhibit a higher probability of consistency with user preferences in both the target and source domains. For example, if a user likes the Harry Potter series of novels in the book domain, then in the movie domain, if there is no observed interaction between the user and the Harry Potter series of movies, the Harry Potter series of movies are more likely to be a false negative sample for that user.

[0116] Based on the scoring function of Equation (4), a candidate negative sample that obtains a high score needs to have a high preference prediction (high probability of first preference) in the target domain to ensure the difficulty of the sample. At the same time, the candidate negative sample has a low preference prediction (low probability of second preference) in the source domain to reduce the risk of sampling negative false cases.

[0117] On the other hand, since the popularity of an item is an important indicator of the difficulty of negative sampling, the methods used in related technologies generally statistically analyze the overall popularity of items in the candidate set from the perspective of the item itself, which is a static absolute popularity. Based on this static popularity, popular items are assigned a higher preference probability (sampling probability). However, from the perspective of the user (interaction object), different interaction objects have different preferences for popular items. If the interaction object itself prefers more popular items, then items with high popularity are more likely to be false negative samples rather than hard negative samples for that interaction object. In this regard, this application also introduces an adaptive relative popularity to correct the first preference probability, thereby adjusting the influence of the popularity of candidate negative samples in negative sampling. As shown in Equation (5), the obtained relative popularity can be converted into a weighting coefficient that reflects the relative popularity of candidate negative samples. Then, the weighting coefficient is applied to Equation (4) to adjust the first preference probability.

[0118] Combining equations (3) and (5), it can be seen that in the weighting coefficients Under the correction, for interactive objects that tend to interact with popular items, i.e., high... In this situation, candidate negative samples are less likely to be assigned high weight coefficients, and correspondingly, the adjusted preference probability corresponding to the candidate negative sample is smaller, indicating that the candidate negative sample is more likely to be a false negative sample. Conversely, if the interaction object tends to interact with unpopular items, i.e. In lower-value cases, candidate negative samples have a higher probability of being assigned higher weight coefficients, indicating that the candidate negative sample is more likely to be a true negative sample.

[0119] As described above, this embodiment of the application obtains a behavior sequence generated by an interactive object interacting with at least one sample in the current domain sample set, where the current domain sample set includes a source domain sample set and a target domain sample set. Then, based on the behavior sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. Next, a sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Then, based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the target domain sample set, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Finally, at least one target negative sample is sampled from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity. This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences using target domain sample sets and source domain sample sets. It then trains a sequence recommendation model using these behavioral sequences as training samples. Based on this model, it determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, it comprehensively utilizes both relative popularity and source domain interest preferences to correct target domain interest preferences, thereby uncovering hard negative samples in the target domain sample set. This alleviates the false negative sample problem in negative sampling and improves the sampling quality of hard negative samples.

[0120] In this embodiment of the application, after sampling at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity, the method may further include: using the target negative sample as the current training sample of the sequence recommendation model in the current training round; and updating the model parameters of the sequence recommendation model based on the current training sample to obtain the updated sequence recommendation model.

[0121] Figure 5 This illustration shows a schematic diagram of utilizing target negative samples according to an embodiment of this application. For example... Figure 5 As shown, after sampling the target negative sample from the candidate negative sample set, the target negative sample can be used to train the sequence recommendation model, so that the trained target sequence recommendation model can be better applied to the target domain for sequence recommendation.

[0122] It's important to note that training a sequence recommendation model involves multiple training rounds, with the model parameters updated after each round. For example, the sequence recommendation model can be used as the training object for the first training round. After the first round, the model parameters are updated, resulting in the updated sequence recommendation model. This updated model can then be used as the training object for the second training round, where training continues based on the updated model. After the second round, the model parameters are updated again. This process is repeated for a third round, until the model converges, yielding the target sequence recommendation model. Therefore, the updated sequence recommendation model can be understood as a transitional model corresponding to each stage of the sequence recommendation model's training process.

[0123] In this embodiment of the application, during the training of the sequence recommendation model, the target negative samples obtained by the above-described sample sampling method can be used as the current training samples in the current training round. Based on the current training samples, the model parameters of the sequence recommendation model are updated to obtain the updated sequence recommendation model. The current training round may be the first training round, the second training round, the third training round, and so on.

[0124] In this embodiment of the application, after updating the model parameters of the sequence recommendation model based on the current training samples to obtain the updated sequence recommendation model, the method may further include: obtaining the first sampling probability for candidate negative samples and the second sampling probability for target negative samples in the next training round of the current training round; selecting the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples or target negative samples according to the target sample sampling probability corresponding to the next training round, wherein the target sample sampling probability includes at least one of the first sampling probability and the second sampling probability; updating the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained.

[0125] During the training of the sequence recommendation model, the training samples may differ for different training rounds. For each training round, the corresponding training samples can be sampled from candidate negative samples or can directly use existing target negative samples as training samples.

[0126] Let's take the current training round as the first training round and the sequence recommendation model as the model training object for the first training round as an example. The current training sample corresponding to the current training round is the first training sample corresponding to the first training round. The first training sample is the target negative sample sampled from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity. The next training round is the second training round, and the next training sample corresponding to the next training round is the second training sample corresponding to the second training round.

[0127] After training the sequence recommendation model using the target negative sample to obtain the updated sequence recommendation model, the updated sequence recommendation model becomes the training object for the second training round. At this point, it is necessary to determine the source of the second training samples corresponding to the second training round. Specifically, we can first obtain the first sampling probability of the candidate negative sample and the second sampling probability of the target negative sample (the first training sample) for the second training round; then, based on the first and second sampling probabilities, we can select the second training samples corresponding to the second training round from the candidate negative sample or the first training sample.

[0128] The process of obtaining the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round of the current training round can include: using the first sampling probability for the candidate negative sample in the current training round as a reference sampling probability, and determining the probability update parameter for the reference sampling probability in the next training round; calculating the first sampling probability for the candidate negative sample in the next training round based on the reference sampling probability and the probability update parameter; and calculating the second sampling probability for the target negative sample in the next training round based on the first sampling probability.

[0129] In this embodiment of the application, the first sampling probability corresponding to the t-th training round can be expressed as ∈ t The first sampling probability corresponding to the (t+1)th training round can be expressed as ∈ t+1 t is a natural number, ∈ t+1 The calculation formula is as follows:

[0130] ∈ t+1 =max(∈ t -η t+1 ,∈ min Equation (6)

[0131] η t+1 Represents the probability update parameters corresponding to the (t+1)th training round, ∈ min The minimum probability value representing the first sampling probability, ∈ min Less than 1.

[0132] As shown in equation (6), the first sampling probability corresponding to the (t+1)th training round depends on the first sampling probability corresponding to the previous training round (the t-th training round) and the probability update parameter corresponding to the (t+1)th training round, and the first sampling probability corresponding to the (t+1)th training round is not less than the preset minimum probability value. Wherein, ∈ min You can set it according to the actual situation.

[0133] Figure 6 This diagram illustrates how the first and second sampling probabilities change with the number of training rounds in an embodiment of this application. Figure 6 As shown, when the current training round is the first training round, the first sampling probability of the candidate negative sample in the first training round can be represented as ∈1. The next training round is the second training round, and the first sampling probability of the candidate negative sample in the second training round can be represented as ∈2. For the first sampling probability ∈1 of the first training round, ∈1 can be set to 1, that is, the training sample corresponding to the first training round is sampled from the candidate negative sample.

[0134] As the training rounds change, ∈ t+1 Continuously updated. At the beginning of training, the sequence recommendation model explores and samples candidate negative samples more extensively to discover new hard negative samples. In the later stages of training, it makes more use of the hard negative samples that have already been explored for model training.

[0135] The process of determining the probability update parameters for the reference sampling probability in the next training round may include: obtaining the model training loss values ​​corresponding to at least one training round before the next training round; and determining the probability update parameters for the reference sampling probability based on the model training loss values.

[0136] For η t+1 The setting can be interpreted as follows: if the training loss in the t-th training round shows a significant decrease, it means that the sampled target negative samples provide valuable information and sufficient gradients. In this case, η t+1 η should be set to a larger value to increase the probability of training with the target negative sample. Conversely, if the training loss at the t-th training epoch does not change significantly, it indicates that the model needs to continue mining new target negative samples from the candidate negative samples. In this case, it is recommended to use a smaller η. t+1 Update accordingly. Based on the above idea, the training loss value from the most recent C rounds can be cached during training, where C is a natural number greater than 1, and η can be updated. t+1 The formula for calculation is defined as follows:

[0137]

[0138] Where η and λ are hyperparameters, L t This represents the loss condition corresponding to the t-th training round.

[0139] After determining the first sampling probability for the candidate negative sample in the next training round, the second sampling probability for the target negative sample in the next training round can be calculated based on the first sampling probability. In this embodiment, the second sampling probability corresponding to the t-th training round can be expressed as 1-∈ t Similarly, the second sampling probability corresponding to the (t+1)th training round can be expressed as 1 - ∈ t+1 .

[0140] After determining the first sampling probability and the second sampling probability corresponding to the next training round, at least one of the first sampling probability and the second sampling probability can be used as the target sample sampling probability of the next training round. Based on the target sample sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples or the target negative samples, so as to use the next training sample to train the updated sequence recommendation model.

[0141] In this embodiment of the application, the target sample sampling probability can be either a first sampling probability or a second sampling probability.

[0142] When the target sample sampling probability is the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples or the target negative samples according to the target sample sampling probability corresponding to the next training round. This may include: obtaining a random probability value, comparing the random probability value with the first sampling probability; and selecting the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples or the target negative samples according to the comparison result.

[0143] The random probability value can be any value between 0 and 1. The first sampling probability can also be any value between 0 and 1.

[0144] Based on the comparison results, the next training sample for the updated sequence recommendation model in the next training round is selected from the candidate negative samples or the target negative samples, including the following two cases:

[0145] The first scenario: If the comparison results show that the random probability value is less than the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples.

[0146] Let's continue with the example of the current training round being the first training round and the next training round being the second training round. If the first sampling probability corresponding to the second training round is 0.9 (90%), while the random probability is 0.8 (meaning the random probability is less than the first sampling probability), then we need to re-select the training samples for the updated sequence recommendation model (the sequence recommendation model obtained after the first training round) in the second training round from the candidate negative samples.

[0147] Specifically, selecting the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples can include: based on the updated sequence recommendation model, sampling at least one target negative sample from the candidate negative samples to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round.

[0148] For example, the updated sequence recommendation model can be reused to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. Then, based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the sample set of the target domain, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Finally, based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative sample, thereby obtaining the training sample corresponding to the updated sequence recommendation model in the second training round.

[0149] In this embodiment of the application, the step of sampling at least one target negative sample from the candidate negative samples based on the updated sequence recommendation model is the same as the aforementioned steps 103-105, except that the original sequence recommendation model is replaced with the updated sequence recommendation model. The specific execution method of the above steps can be referred to equations (1-1)-(7).

[0150] After sampling at least one target negative sample from the candidate negative samples to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round, the method may further include: updating the target negative sample in the current training round based on the next training sample to obtain the updated target negative sample; and using the updated target negative sample as the target negative sample corresponding to the next training round.

[0151] In this embodiment, after selecting new target negative samples from candidate negative samples according to the updated sequence recommendation model, the target negative samples corresponding to the first training round can be updated based on the new target negative samples to obtain the target negative samples corresponding to the second training round. In this embodiment, the target negative samples corresponding to each training round can be cached in the hard negative sample replay pool. The samples in this hard negative sample replay pool can be updated based on the target negative samples selected in each training round, thus ensuring that the samples in the hard negative sample replay pool are always high-quality hard negative samples in the current training round.

[0152] During the update process, the adjusted preference probabilities of each target negative sample can be compared, and the target negative samples with higher adjusted preference probabilities are retained. For example, if the target negative samples corresponding to the first training round include j1 and j2, the adjusted preference probability of j1 is 0.6, and the adjusted preference probability of j2 is 0.3. Based on the updated sequence recommendation model, the new target negative samples selected from the candidate negative samples include j1 and j3, with the adjusted preference probability of j1 becoming 0.5 and the adjusted preference probability of j3 becoming 0.4. Sorted from highest to lowest adjusted preference probability as: j1 (0.5), j3 (0.4), j2 (0.3), j1 and j3 can be used as the target negative samples corresponding to the second training round.

[0153] It should be noted that in the example above, the target negative sample corresponding to the first training round includes j1, and the new target negative sample obtained by the updated sequence recommendation model from the candidate negative samples also includes j1, but the adjusted preference probabilities of the two are different. This is because j1(0.6) is calculated using the sequence recommendation model corresponding to the first training round, while j1(0.5) is calculated using the updated sequence recommendation model (obtained by training the sequence recommendation model from the first training round). The model parameters of the sequence recommendation model and the updated sequence recommendation model are different, therefore, the adjusted preference probabilities obtained when predicting the same candidate negative sample are different.

[0154] During the update process, the adjusted preference probability of each target negative sample can be compared with the preset preference probability threshold. Target negative samples with an adjusted preference probability greater than or equal to the preset preference probability threshold are used as the target negative samples for the second training round.

[0155] The second scenario: If the comparison results show that the random probability value is greater than or equal to the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the target negative samples.

[0156] Let's take the current training epoch as the first training epoch and the next training epoch as the second training epoch as an example. If the first sampling probability corresponding to the second training epoch is 0.9 (90%), and the random probability value is 0.92, the random probability value is greater than the first sampling probability. In this case, at least one target negative sample can be directly selected from the target negative samples in the first training epoch to obtain the training samples for the second training epoch.

[0157] In this embodiment, after training the updated sequence recommendation model obtained in the first training round based on the training samples from the second training round, the updated sequence recommendation model obtained in the second training round can be obtained. Then, the first sampling probability for candidate negative samples and the second sampling probability for target negative samples in the third training round can be obtained. It should be noted that the target negative sample for the third training round is the target negative sample obtained after updating based on the target negative sample obtained in the second training round. Afterwards, according to the first sampling probability, the training samples corresponding to the third training round are selected from the candidate negative samples and the target negative samples obtained after updating based on the target negative samples obtained in the second training round. The updated sequence recommendation model obtained in the second training round is then trained based on the training samples corresponding to the third training round to obtain the updated sequence recommendation model obtained in the third training round. This process continues until the model converges, and finally, the target sequence recommendation model is obtained.

[0158] It should be understood that during the training of a sequence recommendation model, the training samples for each training epoch include not only target negative samples but also target positive samples. The target negative samples provide a contrast signal during model training, helping the trained sequence recommendation model to more comprehensively understand user preferences.

[0159] The model training process involving the sample sampling method provided in the embodiments of this application will be described below. The model training process may include the following steps:

[0160]

[0161]

[0162] As shown in steps 3-19, the model training strategy provided in this application embodiment can be viewed as an exploration-utilization balanced strategy based on course learning, which balances the exploration and utilization of hard negative samples in cross-domain sequence recommendation scenarios. Exploration aims to mine hard negative samples from a large number of uninterrupted candidate negative samples, while utilization refers to fully utilizing the already explored hard negative samples to train the sequence recommendation model. At the beginning of training, both focus on exploration, requiring the sequence recommendation model to come into contact with as many candidate negative samples as possible and mine hard negative samples for model training. As training progresses, the probability of utilizing previously obtained hard negative samples gradually increases, and model training places more emphasis on using the already explored hard negative samples to improve the model. Specifically, at the beginning of training, all hard negative samples need to come from the exploration of unobserved samples. Therefore, the initial exploration probability (the first sampling probability corresponding to the first training round) ∈1 = 1, while limiting the minimum exploration probability (the minimum probability value of the first sampling probability) ∈ min As the number of training rounds increases, ∈ t Gradually decrease the number of samples to reduce the probability of exploring new hard negative samples.

[0163] Therefore, this application provides high-quality hard negative samples for the training of the sequence recommendation model from the perspective of negative samples. This can help the sequence recommendation model capture the preferences of the interactive objects more comprehensively and effectively, so that the target sequence recommendation model after training can provide detailed and accurate personalized recommendation results for the interactive objects in the application stage, thereby improving the user experience.

[0164] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.

[0165] In this embodiment, the sample sampling device will be specifically integrated into an electronic device, which will be a server, as an example for explanation.

[0166] Figure 7 Another schematic flowchart of the sample sampling method provided in an embodiment of this application is shown. Figure 7 As shown, a sample sampling method is described below:

[0167] 201. The server obtains the action sequence generated by the interaction object interacting with at least one sample in the current domain sample set, which includes the source domain sample set and the target domain sample set.

[0168] For example, the server can obtain the source domain behavior sequence generated by the interaction object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence generated by the interaction object interacting with at least one sample in the target domain sample set; and merge the source domain behavior sequence and the target domain behavior sequence to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set.

[0169] 202. Based on the behavior sequence, the server filters out at least one sample from the target domain sample set that has not interacted with the interaction object, thus obtaining candidate negative samples.

[0170] For example, the server can filter out target domain samples that do not have an interaction relationship with the interactive object from the target domain sample set based on the target domain behavior sequence, and obtain a non-interactive target domain sample set; select at least one non-interactive target domain sample from the non-interactive target domain sample set to obtain candidate negative samples.

[0171] 203. The server uses a sequence recommendation model to predict the probability of the first preference of the interactive object for the candidate negative sample in the target domain and the probability of the second preference of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence.

[0172] For example, the server can obtain the sample feature data of the candidate negative sample in the current domain sample set; use a sequence recommendation model to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain; and determine the first preference probability and the second preference probability based on the sample feature data, the target domain preference data and the source domain preference data.

[0173] For example, a server can use a sequence recommendation model to predict the target domain preference data of an interactive object in the target domain based on the target domain behavior sequence; and use a sequence recommendation model to predict the source domain preference data of an interactive object in the source domain based on the source domain behavior sequence.

[0174] For example, the server can fuse sample feature data and target domain preference data to obtain a first preference probability; and fuse sample feature data and source domain preference data to obtain a second preference probability.

[0175] 204. Based on the behavior sequence, the server filters out the target positive samples corresponding to the candidate negative samples in the target domain sample set, and determines the relative popularity of the candidate negative samples. The relative popularity indicates the difference in frequency of the candidate negative samples being selected relative to the target positive samples.

[0176] For example, the server can obtain the positive sample popularity of the target positive sample and the negative sample popularity of the candidate negative sample. The positive sample popularity indicates the frequency of the target positive sample being selected, and the negative sample popularity indicates the frequency of the candidate negative sample being selected. The positive sample popularity and the negative sample popularity are compared to obtain the popularity comparison result. Based on the popularity comparison result, the relative popularity of the candidate negative sample is determined.

[0177] For example, the server can determine the relative popularity based on the difference between the popularity of negative and positive samples when the popularity comparison result indicates that the popularity of positive samples is less than that of negative samples; and set the relative popularity to 0 when the popularity comparison result indicates that the popularity of positive samples is greater than or equal to that of negative samples.

[0178] 205. The server samples at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity.

[0179] For example, the server can adjust the probability value of the first preference probability based on the second preference probability and the relative popularity to obtain the adjusted preference probability; based on the adjusted preference probability, at least one target negative sample can be selected from the candidate negative samples.

[0180] 206. The server uses the target negative sample as the current training sample for the sequence recommendation model in the current training round, and updates the model parameters of the sequence recommendation model based on the current training sample to obtain the updated sequence recommendation model.

[0181] For example, the server can use the sequence recommendation model as the training object for the first training epoch. After the first epoch of training, the model parameters of the sequence recommendation model are updated to obtain the updated sequence recommendation model. Then, the updated sequence recommendation model can be used as the training object for the second training epoch. That is, the second epoch of training continues based on the updated sequence recommendation model. After the second training epoch is completed, the model parameters of the updated sequence recommendation model are updated. Then, the third epoch of training continues based on the sequence recommendation model obtained from the second epoch, until the model converges to obtain the target sequence recommendation model.

[0182] 207. The server obtains the first sampling probability of the candidate negative sample and the second sampling probability of the target negative sample for the next training round of the current training round.

[0183] For example, the server can use the first sampling probability of the candidate negative sample in the current training round as the reference sampling probability and determine the probability update parameter for the reference sampling probability in the next training round; calculate the first sampling probability of the candidate negative sample in the next training round based on the reference sampling probability and the probability update parameter; and calculate the second sampling probability of the target negative sample in the next training round based on the first sampling probability.

[0184] For example, the server can obtain the model training loss value corresponding to at least one training epoch before the next training epoch; based on the model training loss value, it can determine the probability update parameter of the reference sampling probability.

[0185] 208. The server selects the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples or the target negative samples based on the target sample sampling probability corresponding to the next training round. The target sample sampling probability includes at least one of the first sampling probability and the second sampling probability.

[0186] For example, the server can obtain a random probability value and compare it with the first sampling probability. If the comparison result shows that the random probability value is less than the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples. If the comparison result shows that the random probability value is greater than or equal to the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the target negative samples.

[0187] For example, the server can sample at least one target negative sample from the candidate negative samples based on the updated sequence recommendation model to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round.

[0188] For example, the server can update the target negative sample in the current training round based on the next training sample to obtain the updated target negative sample; and use the updated target negative sample as the target negative sample for the next training round.

[0189] 209. The server updates the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained.

[0190] For example, the server can train the updated sequence recommendation model obtained in the first training round based on the training samples from the second training round, thus obtaining the updated sequence recommendation model obtained in the second training round. Then, it can continue to obtain the first sampling probability for candidate negative samples and the second sampling probability for target negative samples in the third training round. It should be noted that the target negative sample for the third training round is the target negative sample obtained after updating based on the target negative sample obtained in the second training round. Afterwards, according to the first sampling probability, the server selects the training samples corresponding to the third training round from the candidate negative samples and the target negative samples obtained after updating based on the target negative samples obtained in the second training round. The server then trains the updated sequence recommendation model obtained in the second training round based on the training samples corresponding to the third training round, thus obtaining the updated sequence recommendation model obtained in the third training round. This process continues until the model converges, finally yielding the target sequence recommendation model.

[0191] As described above, in this embodiment, the server obtains the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set. The current domain sample set includes a source domain sample set and a target domain sample set. Based on the behavior sequence, the server filters out at least one sample in the target domain sample set that has not interacted with the interaction object, obtaining candidate negative samples. The server uses a sequence recommendation model to predict the first preference probability of the interaction object in the target domain for the candidate negative samples and the second preference probability of the interaction object in the source domain for the candidate negative samples. Based on the behavior sequence, the server filters out the target positive samples corresponding to the candidate negative samples in the target domain sample set and determines the relative popularity of the candidate negative samples. The server samples from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity. The server generates at least one target negative sample; it uses the target negative sample as the current training sample for the sequence recommendation model in the current training round, and updates the model parameters of the sequence recommendation model based on the current training sample to obtain the updated sequence recommendation model; the server obtains the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round of the current training round; based on the target sample sampling probability corresponding to the next training round, the server selects the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative sample or the target negative sample, where the target sample sampling probability includes at least one of the first sampling probability and the second sampling probability; the server updates the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained.

[0192] This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences from target and source domain sample sets. The behavioral sequences from these scenarios are then used as training samples to train a sequence recommendation model. Based on this model, the model determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, by combining relative popularity and source domain interest preferences, the target domain interest preferences are corrected to uncover hard negative samples in the target domain sample set. This mitigates the false negative example problem in negative sampling and improves the sampling quality of hard negative samples. Furthermore, this approach provides high-quality hard negative samples for training the sequence recommendation model, enabling it to more comprehensively and effectively capture the preferences of interaction objects. This allows the trained sequence recommendation model to provide detailed and accurate personalized recommendation results to interaction objects during the application phase, thereby enhancing the user experience.

[0193] To better implement the above methods, this application also provides a sample sampling device, which can be integrated into a network device, such as a server or terminal. The terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0194] Figure 8 A schematic diagram of the sample sampling device provided in an embodiment of this application is shown. Figure 8 As shown, the sample sampling device may include an acquisition unit 301, a screening unit 302, a prediction unit 303, a determination unit 304, and a sampling unit 305, as follows:

[0195] (1) Obtain unit 301;

[0196] The acquisition unit 301 is used to acquire the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set, wherein the current domain sample set includes the source domain sample set and the target domain sample set.

[0197] For example, the acquisition unit 301 can be specifically used to acquire the source domain behavior sequence generated by the interaction object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence generated by the interaction object interacting with at least one sample in the target domain sample set; and merge the source domain behavior sequence and the target domain behavior sequence to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set.

[0198] (2) Filtering unit 302;

[0199] The filtering unit 302 is used to filter out at least one sample that has not interacted with the interaction object from the target domain sample set based on the behavior sequence, so as to obtain candidate negative samples.

[0200] For example, the filtering unit 302 can be used to filter out target domain samples that do not have an interaction relationship with the interactive object from the target domain sample set based on the target domain behavior sequence, so as to obtain a non-interactive target domain sample set; and select at least one non-interactive target domain sample from the non-interactive target domain sample set to obtain a candidate negative sample.

[0201] (3) Prediction unit 303;

[0202] The prediction unit 303 is used to predict the first preference probability of the interactive object in the target domain for the candidate negative sample and the second preference probability of the interactive object in the source domain for the candidate negative sample using a sequence recommendation model. The sequence recommendation model is trained based on the behavior sequence.

[0203] For example, prediction unit 303 can be used to obtain sample feature data of candidate negative samples in the current domain sample set; use a sequence recommendation model to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain; and determine the first preference probability and the second preference probability based on the sample feature data, the target domain preference data and the source domain preference data.

[0204] For example, prediction unit 303 can be used to predict the target domain preference data of the interactive object in the target domain based on the target domain behavior sequence and using a sequence recommendation model; and to predict the source domain preference data of the interactive object in the source domain based on the source domain behavior sequence and using a sequence recommendation model.

[0205] For example, prediction unit 303 can be used to fuse sample feature data and target domain preference data to obtain a first preference probability; and to fuse sample feature data and source domain preference data to obtain a second preference probability.

[0206] (4) Determine unit 304;

[0207] The determination unit 304 is used to filter out the target positive samples corresponding to the candidate negative samples in the target domain sample set based on the behavior sequence, and to determine the relative popularity of the candidate negative samples, wherein the relative popularity indicates the difference in frequency of the candidate negative samples being selected relative to the target positive samples.

[0208] For example, the determination unit 304 can be used to obtain the positive sample popularity of the target positive sample and the negative sample popularity of the candidate negative sample. The positive sample popularity indicates the frequency of the target positive sample being selected, and the negative sample popularity indicates the frequency of the candidate negative sample being selected. The positive sample popularity and the negative sample popularity are compared to obtain the popularity comparison result. Based on the popularity comparison result, the relative popularity of the candidate negative sample is determined.

[0209] For example, the determination unit 304 can be used to determine the relative popularity based on the difference between the popularity of the negative sample and the popularity of the positive sample when the popularity comparison result indicates that the popularity of the positive sample is less than the popularity of the negative sample; and to set the relative popularity to 0 when the popularity comparison result indicates that the popularity of the positive sample is greater than or equal to the popularity of the negative sample.

[0210] (5) Sampling unit 305;

[0211] The sampling unit 305 is used to sample at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability and the relative popularity.

[0212] For example, sampling unit 305 can be used to adjust the probability value of the first preference probability based on the second preference probability and the relative popularity to obtain the adjusted preference probability; and based on the adjusted preference probability, to select at least one target negative sample from the candidate negative samples.

[0213] Figure 9 Another schematic diagram of the sample sampling device provided in an embodiment of this application is shown. For example... Figure 9 As shown, the sample sampling device may also include a training unit 306.

[0214] (6) Training Unit 306;

[0215] Training unit 306 is used to take the target negative sample as the current training sample of the sequence recommendation model in the current training round; based on the current training sample, the model parameters of the sequence recommendation model are updated to obtain the updated sequence recommendation model.

[0216] For example, training unit 306 can be used to obtain the first sampling probability of the candidate negative sample and the second sampling probability of the target negative sample for the next training round in the current training round; based on the target sample sampling probability corresponding to the next training round, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative sample or the target negative sample, and the target sample sampling probability includes at least one of the first sampling probability and the second sampling probability; the model parameters of the updated sequence recommendation model are updated based on the next training sample until the target sequence recommendation model is obtained.

[0217] For example, training unit 306 can be used to take the first sampling probability of the candidate negative sample in the current training round as the reference sampling probability and determine the probability update parameter of the reference sampling probability in the next training round; calculate the first sampling probability of the candidate negative sample in the next training round based on the reference sampling probability and the probability update parameter; and calculate the second sampling probability of the target negative sample in the next training round based on the first sampling probability.

[0218] For example, training unit 306 can be used to obtain the model training loss value corresponding to at least one training round before the next training round; and based on the model training loss value, determine the probability update parameter of the reference sampling probability.

[0219] For example, training unit 306 can be used to obtain a random probability value and compare it with the first sampling probability. If the comparison result shows that the random probability value is less than the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples. If the comparison result shows that the random probability value is greater than or equal to the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the target negative samples.

[0220] For example, training unit 306 can be used to sample at least one target negative sample from the candidate negative samples based on the updated sequence recommendation model, so as to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round.

[0221] For example, training unit 306 can be used to update the target negative sample in the current training round based on the next training sample, and obtain the updated target negative sample; the updated target negative sample is then used as the target negative sample for the next training round.

[0222] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0223] As can be seen from the above, in this embodiment of the application, the acquisition unit 301 acquires the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set, the current domain sample set including the source domain sample set and the target domain sample set; then, the filtering unit 302, based on the behavior sequence, filters out at least one sample in the target domain sample set that has not interacted with the interaction object, obtaining candidate negative samples; then, the prediction unit 303 uses a sequence recommendation model to predict the first preference probability of the interaction object for the candidate negative sample in the target domain and the second preference probability of the interaction object for the candidate negative sample in the source domain, the sequence recommendation model being trained based on the behavior sequence; the determination unit 304, based on the behavior sequence, filters out the target positive sample corresponding to the candidate negative sample in the target domain sample set, and determines the relative popularity of the candidate negative sample, the relative popularity indicating the difference in frequency of selection of the candidate negative sample relative to the target positive sample; finally, the sampling unit 305 samples at least one target negative sample from the candidate negative samples according to the first preference probability, the second preference probability and the relative popularity. This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences using target domain sample sets and source domain sample sets. It then trains a sequence recommendation model using these behavioral sequences as training samples. Based on this model, it determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, it comprehensively utilizes both relative popularity and source domain interest preferences to correct target domain interest preferences, thereby uncovering hard negative samples in the target domain sample set. This alleviates the false negative sample problem in negative sampling and improves the sampling quality of hard negative samples.

[0224] This application also provides an electronic device, such as... Figure 10 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0225] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 10 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0226] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0227] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0228] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0229] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0230] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0231] The process involves obtaining a sequence of actions generated by an interactive object interacting with at least one sample in the current domain sample set, which includes a source domain sample set and a target domain sample set. Based on the action sequence, at least one sample in the target domain sample set that has not interacted with the interactive object is selected to obtain candidate negative samples. A sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the action sequence. Based on the action sequence, target positive samples corresponding to the candidate negative samples are selected from the target domain sample set, and the relative popularity of the candidate negative samples is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples.

[0232] For example, an electronic device can acquire a sequence of actions generated by an interactive object interacting with at least one sample in a current domain sample set, where the current domain sample set includes a source domain sample set and a target domain sample set. Based on the action sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. A sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. Based on the action sequence, target positive samples corresponding to the candidate negative samples are selected from the target domain sample set, and the relative popularity of the candidate negative samples is determined. Based on the first preference probability, the second preference probability, and the relative popularity, at least one target positive sample is sampled from the candidate negative samples. The process involves: labeling negative samples; using the target negative sample as the current training sample for the sequence recommendation model in the current training round, and updating the model parameters based on the current training sample to obtain the updated sequence recommendation model; obtaining the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round; selecting the next training sample for the updated sequence recommendation model in the next training round from the candidate negative sample or the target negative sample according to the target sample sampling probability corresponding to the next training round, where the target sample sampling probability includes at least one of the first and second sampling probabilities; updating the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained, and so on.

[0233] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0234] As described above, this embodiment of the application obtains a behavior sequence generated by an interactive object interacting with at least one sample in the current domain sample set, where the current domain sample set includes a source domain sample set and a target domain sample set. Then, based on the behavior sequence, at least one sample that has not interacted with the interactive object is selected from the target domain sample set to obtain candidate negative samples. Next, a sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Then, based on the behavior sequence, the target positive sample corresponding to the candidate negative sample is selected from the target domain sample set, and the relative popularity of the candidate negative sample is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Finally, at least one target negative sample is sampled from the candidate negative samples according to the first preference probability, the second preference probability, and the relative popularity. This scheme constructs cross-domain sequence recommendation scenarios and corresponding behavioral sequences using target domain sample sets and source domain sample sets. It then trains a sequence recommendation model using these behavioral sequences as training samples. Based on this model, it determines the interaction object's interest preferences for candidate negative samples in the target domain sample set within both the source and target domains. Simultaneously, it determines the relative popularity of candidate negative samples relative to target positive samples. Finally, it comprehensively utilizes both relative popularity and source domain interest preferences to correct target domain interest preferences, thereby uncovering hard negative samples in the target domain sample set. This alleviates the false negative sample problem in negative sampling and improves the sampling quality of hard negative samples.

[0235] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0236] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the sample sampling methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0237] The process involves obtaining a sequence of actions generated by an interactive object interacting with at least one sample in the current domain sample set, which includes a source domain sample set and a target domain sample set. Based on the action sequence, at least one sample in the target domain sample set that has not interacted with the interactive object is selected to obtain candidate negative samples. A sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the action sequence. Based on the action sequence, target positive samples corresponding to the candidate negative samples are selected from the target domain sample set, and the relative popularity of the candidate negative samples is determined. The relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample. Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples.

[0238] For example, the process involves obtaining a sequence of actions generated by an interactive object interacting with at least one sample in the current domain sample set, which includes a source domain sample set and a target domain sample set. Based on the action sequence, at least one sample in the target domain sample set that did not interact with the interactive object is selected to obtain candidate negative samples. A sequence recommendation model is used to predict the first preference probability of the interactive object in the target domain for the candidate negative samples and the second preference probability of the interactive object in the source domain for the candidate negative samples. Based on the action sequence, target positive samples corresponding to the candidate negative samples are selected in the target domain sample set, and the relative popularity of the candidate negative samples is determined. Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples. This process involves: using the target negative sample as the current training sample for the sequence recommendation model in the current training round; updating the model parameters based on the current training sample to obtain the updated sequence recommendation model; obtaining the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round; selecting the next training sample for the updated sequence recommendation model in the next training round from the candidate negative sample or the target negative sample according to the target sample sampling probability corresponding to the next training round, where the target sample sampling probability includes at least one of the first and second sampling probabilities; updating the model parameters of the updated sequence recommendation model based on the next training sample until the target sequence recommendation model is obtained, and so on.

[0239] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0240] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0241] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the sample sampling methods provided in the embodiments of this application, the beneficial effects that any of the sample sampling methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0242] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various alternative implementations of the data access aspect described above.

[0243] The foregoing has provided a detailed description of a sample sampling method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A sample sampling method, characterized in that, include: Obtain the sequence of behaviors generated by the interaction object interacting with at least one sample in the current domain sample set, wherein the current domain sample set includes a source domain sample set and a target domain sample set; Based on the behavior sequence, at least one sample that has not interacted with the interaction object is selected from the target domain sample set to obtain candidate negative samples; A sequence recommendation model is used to predict the first preference probability of the interactive object for the candidate negative sample in the target domain and the second preference probability of the interactive object for the candidate negative sample in the source domain. The sequence recommendation model is trained based on the behavior sequence. Based on the behavioral sequence, target positive samples corresponding to the candidate negative samples are selected from the target domain sample set, and the relative popularity of the candidate negative samples is determined, wherein the relative popularity indicates the difference in frequency of selection of the candidate negative samples relative to the target positive samples; Based on the first preference probability, the second preference probability, and the relative popularity, at least one target negative sample is sampled from the candidate negative samples.

2. The sample sampling method as described in claim 1, characterized in that, The acquisition of the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set includes: Obtain the source domain behavior sequence generated by the interaction object interacting with at least one sample in the source domain sample set, and the target domain behavior sequence generated by the interaction object interacting with at least one sample in the target domain sample set; The source domain behavior sequence and the target domain behavior sequence are merged to obtain the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set.

3. The sample sampling method as described in claim 2, characterized in that, The step of filtering at least one sample from the target domain sample set that has not interacted with the interaction object based on the behavior sequence to obtain candidate negative samples includes: Based on the target domain behavior sequence, target domain samples that do not have an interaction relationship with the interaction object are selected from the target domain sample set to obtain a non-interactive target domain sample set. At least one non-interactive target domain sample is selected from the set of non-interactive target domain samples to obtain the candidate negative sample.

4. The sample sampling method as described in claim 2, characterized in that, The step of using the sequence recommendation model to predict the first preference probability of the interactive object in the target domain for the candidate negative sample and the second preference probability of the interactive object in the source domain for the candidate negative sample includes: Obtain the sample feature data of the candidate negative sample in the current domain sample set; Using the sequence recommendation model, the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain are predicted. Based on the sample feature data, the target domain preference data, and the source domain preference data, the first preference probability and the second preference probability are determined.

5. The sample sampling method as described in claim 4, characterized in that, The step of using the sequence recommendation model to predict the target domain preference data of the interactive object in the target domain and the source domain preference data of the interactive object in the source domain includes: Based on the target domain behavior sequence, the sequence recommendation model is used to predict the target domain preference data of the interactive object in the target domain. Based on the source domain behavior sequence, the sequence recommendation model is used to predict the source domain preference data of the interactive object in the source domain.

6. The sample sampling method as described in claim 4, characterized in that, Determining the first preference probability and the second preference probability based on the sample feature data, the target domain preference data, and the source domain preference data includes: The sample feature data and the target domain preference data are fused to obtain the first preference probability; The sample feature data and the source domain preference data are fused to obtain the second preference probability.

7. The sample sampling method as described in claim 1, characterized in that, Determining the relative popularity of the candidate negative samples includes: Obtain the positive sample popularity of the target positive sample and the negative sample popularity of the candidate negative sample, wherein the positive sample popularity indicates the frequency of the target positive sample being selected and the negative sample popularity indicates the frequency of the candidate negative sample being selected. The popularity of the positive samples and the popularity of the negative samples are compared to obtain the popularity comparison results; Based on the popularity comparison results, the relative popularity of the candidate negative samples is determined.

8. The sample sampling method as described in claim 7, characterized in that, Determining the relative popularity of the candidate negative samples based on the popularity comparison results includes: When the popularity comparison result indicates that the popularity of the positive sample is less than the popularity of the negative sample, the relative popularity is determined based on the difference between the popularity of the negative sample and the popularity of the positive sample; When the popularity comparison result indicates that the popularity of the positive sample is greater than or equal to the popularity of the negative sample, the relative popularity is set to 0.

9. The sample sampling method as described in claim 1, characterized in that, The step of sampling at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity includes: Based on the second preference probability and the relative popularity, the probability value of the first preference probability is adjusted to obtain the adjusted preference probability; Based on the adjusted preference probability, at least one target negative sample is selected from the candidate negative samples.

10. The sample sampling method according to any one of claims 1-9, characterized in that, After sampling at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity, the method further includes: The target negative sample is used as the current training sample for the sequence recommendation model in the current training round. Based on the current training samples, the model parameters of the sequence recommendation model are updated to obtain the updated sequence recommendation model.

11. The sample sampling method as described in claim 10, characterized in that, After updating the model parameters of the sequence recommendation model based on the current training samples to obtain the updated sequence recommendation model, the method further includes: Obtain the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round of the current training round; Based on the target sample sampling probability corresponding to the next training round, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples or the target negative samples. The target sample sampling probability includes at least one of the first sampling probability and the second sampling probability. The model parameters of the updated sequence recommendation model are updated based on the next training sample until the target sequence recommendation model is obtained.

12. The sample sampling method as described in claim 11, characterized in that, The step of obtaining the first sampling probability for the candidate negative sample and the second sampling probability for the target negative sample in the next training round of the current training round includes: The first sampling probability of the candidate negative sample in the current training round is used as the reference sampling probability, and the probability update parameter for the next training round is determined based on the reference sampling probability. Based on the reference sampling probability and the probability update parameter, calculate the first sampling probability for the candidate negative sample in the next training round; Based on the first sampling probability, the second sampling probability for the target negative sample in the next training round is calculated.

13. The sample sampling method as described in claim 12, characterized in that, The step of determining the probability update parameters for the reference sampling probability in the next training round includes: Obtain the model training loss value for each of at least one training round preceding the next training round; Based on the model training loss value, the probability update parameters of the reference sampling probability are determined.

14. The sample sampling method as described in claim 11, characterized in that, The target sample sampling probability is the first sampling probability. The step of selecting the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples or the target negative samples based on the sampling probability of the target sample corresponding to the next training round includes: Obtain a random probability value and compare the random probability value with the first sampling probability; If the comparison result shows that the random probability value is less than the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the candidate negative samples. If the comparison result shows that the random probability value is greater than or equal to the first sampling probability, the next training sample corresponding to the updated sequence recommendation model in the next training round is selected from the target negative sample.

15. The sample sampling method as described in claim 14, characterized in that, The step of selecting the next training sample corresponding to the updated sequence recommendation model in the next training round from the candidate negative samples includes: Based on the updated sequence recommendation model, at least one target negative sample is sampled from the candidate negative samples to obtain the next training sample corresponding to the updated sequence recommendation model in the next training round.

16. The sample sampling method as described in claim 15, characterized in that, After sampling at least one target negative sample from the candidate negative samples to obtain the next training sample corresponding to the next training round based on the updated sequence recommendation model, the method further includes: Based on the next training sample, the target negative sample in the current training round is updated to obtain the updated target negative sample; The updated target negative sample is used as the target negative sample for the next training round.

17. A sample sampling device, characterized in that, include: The acquisition unit is used to acquire the behavior sequence generated by the interaction object interacting with at least one sample in the current domain sample set, wherein the current domain sample set includes a source domain sample set and a target domain sample set. A filtering unit is used to filter at least one sample that has not interacted with the interaction object from the target domain sample set based on the behavior sequence, so as to obtain candidate negative samples; The prediction unit is used to predict the first preference probability of the interactive object in the target domain for the candidate negative sample and the second preference probability of the interactive object in the source domain for the candidate negative sample using a sequence recommendation model, wherein the sequence recommendation model is trained based on the behavior sequence. The determining unit is configured to, based on the behavioral sequence, filter out the target positive sample corresponding to the candidate negative sample in the target domain sample set, and determine the relative popularity of the candidate negative sample, wherein the relative popularity indicates the difference in frequency of selection of the candidate negative sample relative to the target positive sample; A sampling unit is configured to sample at least one target negative sample from the candidate negative samples based on the first preference probability, the second preference probability, and the relative popularity.

18. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the sample sampling method according to any one of claims 1 to 16.

19. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the sample sampling method according to any one of claims 1 to 16.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the sample sampling method according to any one of claims 1 to 16.