A recommendation system optimization method and device based on negative sample perception

By optimizing the utilization of negative samples through batch-based negative sample sharing and dynamic reward adjustment mechanisms, the problems of low efficiency and high computational cost of negative sample integration in recommendation systems are solved, achieving efficient and flexible negative sample optimization and improving the accuracy and real-time performance of recommendation systems.

CN120509497BActive Publication Date: 2025-10-28UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510998334.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-28
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing recommender systems suffer from low efficiency, high computational and storage overhead, and poor flexibility in the process of negative sample integration and optimization, which particularly affects system performance and scalability when dealing with large-scale data.

Method used

We adopt a recommendation system optimization method based on negative sample awareness. We expand the negative sample pool through an intra-batch negative sample sharing mechanism and dynamically adjust the confidence of negative samples based on a dynamic reward marginal adjustment mechanism. We also optimize the model parameters by combining a lightweight sequence recommendation model.

Benefits of technology

It significantly reduces computational and storage overhead, improves the accuracy and flexibility of model recommendations, and can quickly adapt to dynamic data changes to meet the needs of real-time and large-scale data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509497B_ABST
    Figure CN120509497B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for optimizing a recommender system based on negative sample awareness. The optimization method includes: receiving negative sample optimization requests submitted by users or the system; parsing the requests to determine optimization objectives and related data; expanding the negative sample pool through an intra-batch negative sample sharing mechanism, sharing the calculation results of negative samples within the same training batch to reduce computational overhead; dynamically adjusting the impact of negative samples on model training based on the confidence level of the negative samples according to a dynamic reward marginal adjustment mechanism, wherein the confidence level is calculated by assisting a lightweight sequence recommendation model; updating model parameters according to the optimization results, and locally adjusting the model structure through adapter parameters. This invention effectively reduces additional computational overhead by using an intra-batch sharing mechanism when expanding the negative sample pool, avoiding negative impacts on other parts of the model, greatly reducing time and resource costs in the negative sample processing process, and maintaining excellent computational performance when processing large-scale data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of recommender system technology, specifically a recommender system optimization method and apparatus based on negative sample perception. Background Technology

[0002] With the widespread application of large-scale language models (LLMs) in recommender systems, they have demonstrated powerful knowledge representation and reasoning capabilities, especially in handling user preferences and recommendation accuracy. However, despite the excellent performance of LLMs in recommendation tasks, existing methods still face a series of challenges in the process of negative sample integration and optimization. In particular, how to efficiently integrate negative samples to improve the model's recommendation accuracy while avoiding the degradation of system performance due to excessive computational and memory overhead has become a key issue in current research.

[0003] To address these issues, existing technologies have primarily proposed the following methods:

[0004] (1) Traditional negative sample sampling methods: These methods help the model distinguish between positive and negative preferences by using random negative samples or similarity-based negative samples during training. However, these methods often ignore the differences between negative samples and simply treat all negative samples as equally important, thus affecting the recommendation accuracy and the model's learning efficiency. In addition, as the number of negative samples increases, the computational and storage costs also increase, becoming a bottleneck affecting the system's efficiency.

[0005] (2) Dynamic Negative Sample Adjustment Method: This type of method attempts to improve the model's ability to handle different negative samples by adaptively adjusting the contribution of negative samples. By introducing a dynamic reward mechanism, the influence of negative samples in model training can be adjusted according to their credibility and relevance. This method improves recommendation performance to some extent, but it relies on a large amount of computing and memory resources, especially when dealing with large-scale data, it still faces high computational complexity and memory requirements.

[0006] Although existing methods have addressed the negative sample optimization problem to some extent, they still have the following key limitations:

[0007] (1) Low efficiency of negative sample integration: Most existing methods rely on static negative sample pools, which cannot dynamically adjust the quality and influence of negative samples, causing the model to miss potential optimization opportunities when processing complex data.

[0008] (2) High computational and memory overhead: Traditional negative sample sampling and optimization methods have rapidly increased computational and storage overhead as the negative sample size increases, affecting the real-time performance and scalability of the system.

[0009] (3) Poor flexibility: Many methods cannot adapt to the dynamic changes of new data, especially when it is necessary to adjust the negative sample strategy for different users or different tasks, lacking sufficient flexibility and adaptability.

[0010] Overall, existing methods still fall short in terms of efficiency, flexibility, and generality in negative sample optimization. There is an urgent need for a technique that can efficiently integrate negative samples, reduce computational and storage overhead, and offer greater flexibility and scalability. Such a method would not only improve the accuracy of recommendation systems but also meet the demands of real-time processing and large-scale data processing. Summary of the Invention

[0011] This embodiment provides a method, apparatus, electronic device, and storage medium for optimizing a recommendation system based on negative sample perception, in order to solve the problem of high computational and storage overhead in traditional negative sample sampling methods in related technologies.

[0012] In a first aspect, embodiments of the present invention provide a method for optimizing a recommender system based on negative sample awareness, the method comprising:

[0013] Receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data;

[0014] The negative sample pool is expanded by a negative sample sharing mechanism within the batch, and the calculation results of negative samples within the same training batch are shared to reduce computational overhead.

[0015] Based on the dynamic reward marginal adjustment mechanism, the impact of negative samples on model training is dynamically adjusted according to their confidence level, wherein the confidence level is calculated by assisting a lightweight sequence recommendation model.

[0016] The model parameters are updated based on the optimization results, and the model structure is locally adjusted through the adapter parameters. The model training uses an optimization objective function.

[0017] In an optional embodiment, the intra-batch negative sample sharing mechanism includes:

[0018] In training batches In the middle, a shared set of negative samples The calculation results; among which, For the user's historical interaction sequence, As a positive sample, It is a set of negative samples randomly sampled from an unobserved interaction space;

[0019] Decomposition by function ; User historical interaction sequence Mapped to an intermediate representation, and the final log probability is generated, where, This indicates the sequence of user's historical interactions. The function mapped to the intermediate representation, and It is a function that generates the final logarithmic probability.

[0020] In an optional embodiment, the dynamic reward marginal adjustment mechanism includes:

[0021] Calculate negative samples User historical interaction sequence correlation score ;

[0022] The confidence level of negative samples is determined based on the scores, and the reward margin is dynamically adjusted.

[0023] ;

[0024] Among them, α controls the adjustment range. It is a confidence benchmark updated through momentum. It is the initial reward margin. For negative sample confidence, For the user's historical interaction sequence, For negative samples, For the negative sample set, Expressing expectations, Represents user sequence From batch Mid-sampling, This indicates that negative samples are sampled from the set of negative samples.

[0025] In an optional embodiment, the negative sample confidence score is calculated using the following formula:

[0026] ;

[0027] in, This represents the relevance score calculated using a lightweight sequence recommendation model.

[0028] In an optional embodiment, the optimization objective function is:

[0029] ;

[0030] in, Represents a given user's historical interaction sequence Predict positive samples In the negative sample set In the context of preference probability, γ represents the dynamic reward margin. For the sigmoid function, Expressing expectations, Indicates sample Sample from dataset D.

[0031] In an optional embodiment, the model parameter update step specifically includes:

[0032] Only update the adapter parameters, retain the frozen parameters of the model, in order to avoid interfering with the original performance of the model;

[0033] Efficiently solve parameter optimization problems using the stochastic gradient descent algorithm.

[0034] Compared with existing technologies, the beneficial effects of the recommendation system optimization method based on negative sample awareness of the present invention are as follows:

[0035] This invention provides a highly versatile and adaptable solution by expanding the negative sample pool to mitigate popularity bias or adjusting the weights of negative samples based on users' historical behavior. When expanding the negative sample pool, an intra-batch sharing mechanism effectively reduces additional computational overhead and avoids negative impacts on other parts of the model. The optimization process of model parameters is transformed into an efficient computational problem, solved using efficient algorithms such as stochastic gradient descent (SGD). This design significantly reduces the time and resource overhead in negative sample processing, enabling this invention to maintain excellent computational performance when handling large-scale data.

[0036] Secondly, embodiments of the present invention provide a recommender system optimization apparatus based on negative sample perception, comprising:

[0037] The negative sample optimization request module is configured to: receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data;

[0038] The intra-batch negative sample sharing module is configured to expand the negative sample pool through the intra-batch negative sample sharing mechanism and share the negative sample calculation results within the same training batch to reduce computational overhead.

[0039] The dynamic reward marginal adjustment module is configured to: dynamically adjust the impact of negative samples on model training based on the confidence level of the negative samples according to the dynamic reward marginal adjustment mechanism, wherein the confidence level is calculated by assisting the lightweight sequence recommendation model;

[0040] The model parameter update module is configured to update the model parameters based on the optimization results and locally adjust the model structure through the adapter parameters, wherein the model training adopts the optimization objective function.

[0041] Thirdly, embodiments of the present invention provide an electronic device, including a processor, a communication interface, a memory, and a bus, wherein the processor, the communication interface, and the memory communicate with each other through the bus, and the processor can call logical instructions in the memory to execute the steps of the method provided in the first aspect.

[0042] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the recommendation system optimization method based on negative sample perception as described in the first aspect.

[0043] Compared with the prior art, the beneficial effects of the recommendation system optimization device, electronic device and storage medium based on negative sample perception of the present invention are the same as those of the recommendation system optimization method based on negative sample perception described in the first aspect, so they will not be repeated here. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of the recommendation system optimization method based on negative sample perception in an embodiment of the present invention;

[0046] Figure 2 This is a structural block diagram of the recommendation system optimization device based on negative sample perception in an embodiment of the present invention;

[0047] Figure 3 This is a structural block diagram of the electronic device in an embodiment of the present invention. Detailed Implementation

[0048] To better understand the purpose, technical solution, and advantages of this application, the application is described and explained below in conjunction with the accompanying drawings and embodiments.

[0049] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0050] In recommendation systems based on large-scale language models (LLM), the effective utilization of negative samples is crucial for improving recommendation performance. To address the shortcomings of existing methods in expanding the negative sample pool and dynamically evaluating the impact of negative samples, this paper proposes a Negative Sample Aware Optimization (NAPO) framework.

[0051] This method addresses the scale and efficiency issues of negative sample ensembles through two key innovations:

[0052] (1) Intra-batch negative sample sharing: Expand the negative sample pool without increasing computational cost.

[0053] (2) Dynamic reward marginal adjustment: The model update is dynamically adjusted according to the confidence of negative samples, thereby optimizing the impact of negative samples on the model.

[0054] To better utilize negative sample information, the NAPO framework proposes a new optimization objective that combines intra-batch negative sample sharing with a dynamic impact assessment mechanism.

[0055] This invention provides a method for optimizing a recommender system based on negative sample awareness (NAPO method). Figure 1 This is a flowchart of the recommendation system optimization method based on negative sample awareness of the present invention, as shown below. Figure 1As shown, the process includes the following steps:

[0056] S100: Receive negative sample optimization requests submitted by users or the system, parse the requests to determine optimization targets and related data;

[0057] For example, a user or system submits a negative sample optimization request through a specific interface, specifying the target data to be optimized. For instance, a user might request to expand the negative sample pool or adjust the weights of negative samples based on the needs of the actual application. The content of the negative sample optimization request will be parsed by the system into a specific task type, providing target data input for subsequent parameter adjustments.

[0058] S200: Expand the negative sample pool through an intra-batch negative sample sharing mechanism to share the negative sample calculation results within the same training batch in order to reduce computational overhead.

[0059] Intra-batch negative sample sharing mechanisms include:

[0060] In training batches In the middle, a shared set of negative samples The calculation results; among which, For the user's historical interaction sequence, As a positive sample, It is a set of negative samples randomly sampled from an unobserved interaction space;

[0061] Decomposition by function ; User historical interaction sequence Mapping to an intermediate representation and generating the final log probability. Wherein, This indicates the sequence of user's historical interactions. The function mapped to the intermediate representation, and It is a function that generates the final logarithmic probability.

[0062] Specifically, after the negative sample optimization request is parsed, the NAPO framework optimizes the processing of negative samples through intra-batch negative sample sharing and dynamic reward adjustment mechanisms. Specifically, the framework shares the negative sample information calculated within the batch and dynamically adjusts the reward based on the credibility of each negative sample. In this way, the framework can efficiently integrate a large number of negative samples, improving the efficiency and accuracy of the model when handling complex tasks.

[0063] The intra-batch negative sample sharing strategy significantly expands the coverage of negative samples without introducing additional computational overhead by sharing the computational results of negative samples. A training batch is set up. ,in For the user's historical interaction sequence, As a positive sample, This is a set of negative samples randomly sampled from an unobserved interaction space. Based on the model's computational pattern, the log probability of the predicted outcome can be decomposed into two functions:

[0064] ;

[0065] in, This indicates the sequence of user's historical interactions. The function mapped to the intermediate representation, and It is a function that generates the final log probability. By sharing the function value, NAPO reduces the computational burden within a batch through the sharing of log probabilities, thereby improving the coverage of negative samples.

[0066] S300: Based on the dynamic reward marginal adjustment mechanism, the impact of negative samples on model training is dynamically adjusted according to their confidence level, where the confidence level is calculated by assisting the lightweight sequence recommendation model.

[0067] It should be noted that the dynamic reward marginal adjustment mechanism includes:

[0068] Calculate negative samples User historical interaction sequence correlation score ;

[0069] The confidence level of negative samples is determined based on the scores. Specifically, the confidence level of negative samples is calculated using the following formula:

[0070] ;

[0071] in, This represents the relevance score calculated using the lightweight sequence recommendation model (SASRec model).

[0072] Dynamically adjust the reward margin:

[0073] ;

[0074] Among them, α controls the adjustment range. It is a confidence benchmark updated through momentum. It is the initial reward margin. The confidence level for negative samples. For the user's historical interaction sequence, For negative samples, For the negative sample set, Expressing expectations, Represents user sequence From batch Mid-sampling, This indicates that negative samples are sampled from the set of negative samples.

[0075] In this embodiment, to further improve the utilization efficiency of negative samples, NAPO introduces a dynamic reward marginal adjustment mechanism. This mechanism adjusts the impact of negative samples on model updates based on their confidence level. For each negative sample... It uses a lightweight, auxiliary sequence recommendation model (such as SASRec) to calculate the user's historical interaction sequence. Relevance score:

[0076] ;

[0077] in, This represents the relevance score calculated using a lightweight sequence recommendation model. For the user's historical interaction sequence, This is a negative sample.

[0078] This score is used to adjust the confidence level of negative samples. Higher confidence levels indicate a greater contribution from negative samples, and the reward margin γ is adjusted accordingly. The dynamic reward margin γ can be dynamically adjusted based on the confidence level of negative samples, as shown below:

[0079] ;

[0080] Among them, α controls the adjustment range. It is a confidence benchmark updated through momentum. This is the initial reward margin. This dynamic adjustment mechanism ensures that the model can utilize negative samples more efficiently, improving recommendation performance.

[0081] NAPO significantly improves the utilization efficiency of negative samples by sharing negative samples within batches and dynamically adjusting the marginal reward, while expanding the coverage of negative samples without increasing computational overhead. The core advantage of this method lies in its ability to dynamically adjust the impact of negative samples on model training based on their confidence level, thereby achieving more efficient preference optimization.

[0082] S400. Update the model parameters based on the optimization results, and locally adjust the model structure through the adapter parameters. The model training adopts the optimization objective function.

[0083] The specific steps for updating model parameters include:

[0084] Only update the adapter parameters, retain the frozen parameters of the model, in order to avoid interfering with the original performance of the model;

[0085] Efficiently solve parameter optimization problems using the stochastic gradient descent algorithm.

[0086] The objective function to be optimized is:

[0087] ;

[0088] in, Represents a given user's historical interaction sequence Predict positive samples In the negative sample set In the context of preference probability, γ represents the dynamic reward margin. For the sigmoid function, Expressing expectations, Indicates sample Sample from dataset D.

[0089] In this embodiment, based on the negative sample sharing and dynamic adjustment results, the present invention directly updates the model parameters (adapter parameters). The adjustment of the adapter parameters only affects the local structure of the model, avoiding interference with frozen parameters and effectively preserving the model's original capabilities. This method not only accelerates the negative sample optimization process but also reduces potential performance degradation issues, enabling the model to quickly adapt to the demands of negative sample optimization requests.

[0090] After the parameters are updated, the model is comprehensively evaluated to check if its performance meets the requirements. The evaluation includes the model's performance on the negative sample optimization task (e.g., whether negative samples are effectively integrated) and the recommendation accuracy on other tasks. Based on the evaluation results, the parameter adjustment strategy or hyperparameter configuration is further optimized to ensure that the model maintains a high level of recommendation accuracy and stability after negative sample optimization.

[0091] The updated model is then deployed in real-world applications to handle real-time tasks. For example, in recommender systems, the model can immediately apply the latest negative sample optimization results to provide personalized recommendations to users; in generative tasks, the corrected model can output more accurate and reliable recommendations. Through this step, the NAPO framework achieves rapid response to dynamic data demands, meeting the application scenarios in industry and academia that require high real-time performance and efficiency.

[0092] For example, (1) Personalized Recommendation Systems: In personalized recommendation systems, user behavior data (such as browsing history, click logs, and purchase history) is widely used to train recommendation models. However, this data may contain outdated or irrelevant information. For instance, when a user requests optimization of the negative sample pool, NAPO can efficiently adjust the weights of the negative samples without retraining the entire recommendation model. Simultaneously, user behavior in recommendation systems is dynamic, and NAPO can adjust the model in real time to adapt to the latest user preferences, ensuring the accuracy and relevance of the recommendation results. Through this method, recommendation systems can not only optimize the utilization of negative samples but also maintain continuously optimized recommendation performance.

[0093] (2) Privacy Protection: In many applications involving personal data (such as medical diagnosis, financial analysis, and social media), protecting user privacy is a top priority. NAPO can effectively mitigate the impact of privacy data by expanding the negative sample pool, thereby avoiding the risk of data leakage. For example, in financial recommendation systems, users may request the removal of specific historical transaction data. NAPO can efficiently meet this need through precise negative sample adjustment and parameter updates, while ensuring that the model's performance is not significantly affected after optimization. Compared with traditional negative sample optimization methods, NAPO is faster and more computationally efficient, providing a highly efficient technical means for privacy protection.

[0094] (3) Generative Recommendation Models: In content generation and recommendation tasks, models may generate results containing erroneous information or inaccurate recommendations. For example, a recommendation system may generate irrelevant recommendations based on outdated user behavior data. Through NAPO, generative recommendation models can correct these inaccurate recommendations in a timely manner. For example, the model can adjust the impact of irrelevant negative samples, thereby improving the accuracy of recommendations and user experience. This method is not only applicable to personalized recommendation tasks, but can also be extended to multimodal recommendation applications (such as cross-platform recommendation and multi-category recommendation), achieving cross-domain flexibility.

[0095] As can be seen from the above application examples, the NAPO framework combines efficient computing power and flexible task adaptability in the negative sample optimization task of large-scale recommendation systems, significantly reducing computing costs while meeting diverse practical needs.

[0096] The NAPO method proposed in this invention demonstrates significant advantages and positive effects in solving the negative sample optimization problem in recommender systems. First, NAPO precisely adjusts the impact of negative samples by introducing intra-batch negative sample sharing and a dynamic reward adjustment mechanism, thereby quickly adapting to data changes without requiring additional model training. Compared with traditional methods, NAPO not only improves the recommendation accuracy of the model when dealing with the negative sample optimization problem but also significantly reduces computational costs and time overhead.

[0097] Experimental results show that NAPO outperforms state-of-the-art methods on both the Goodreads and LastFM datasets. Specifically, on the Goodreads dataset, NAPO demonstrates significant improvements over baseline models in both HitRatio@1 and Popularity Bias metrics. On the HitRatio@1 metric, NAPO achieves a 7.47% improvement over the S-DPO model, and reduces the Popularity Bias metric by 10.32%. These improvements indicate that NAPO possesses superior capabilities in capturing user preference patterns and optimizing the distribution of negative samples.

[0098] In summary, NAPO effectively combines intra-batch negative sample sharing and dynamic reward adjustment mechanisms to solve the negative sample optimization problem, which is difficult for traditional methods to handle. It not only improves the model's recommendation accuracy and robustness but also enhances computational efficiency, demonstrating its strong potential in large-scale recommender systems. The successful validation of NAPO shows that this method has broad applicability and practical value in dealing with negative sample optimization problems in dynamically changing and large-scale datasets.

[0099] Table 1: Recommendation performance of different methods

[0100]

[0101] This invention also provides a recommender system optimization device based on negative sample perception, which is used to implement the above-described method embodiments; details already described will not be repeated. The terms "module," "unit," and "subunit," etc., used below refer to combinations of software and / or hardware that perform a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation or a combination of software and hardware is also possible and contemplated.

[0102] like Figure 2 As shown, Figure 2 This is a structural block diagram of the recommendation system optimization device based on negative sample perception in this invention. The device includes:

[0103] The negative sample optimization request module 101 is configured to: receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data;

[0104] The batch negative sample sharing module 102 is configured to expand the negative sample pool through the batch negative sample sharing mechanism and share the negative sample calculation results within the same training batch to reduce computational overhead.

[0105] The dynamic reward marginal adjustment module 103 is configured to: dynamically adjust the impact of negative samples on model training based on the confidence level of the negative samples according to the dynamic reward marginal adjustment mechanism, wherein the confidence level is calculated by assisting the lightweight sequence recommendation model;

[0106] The model parameter update module 104 is configured to update the model parameters based on the optimization results and locally adjust the model structure through the adapter parameters, wherein the model training adopts an optimization objective function.

[0107] The recommendation system optimization device based on negative sample perception in this invention is used to implement the above method, so it will not be described in detail here.

[0108] Figure 3A structural block diagram of the electronic device provided in the embodiments of the present invention, such as... Figure 3 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute the following methods:

[0109] Receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data;

[0110] The negative sample pool is expanded by a negative sample sharing mechanism within the batch, and the calculation results of negative samples within the same training batch are shared to reduce computational overhead.

[0111] Based on the dynamic reward marginal adjustment mechanism, the impact of negative samples on model training is dynamically adjusted according to their confidence level, wherein the confidence level is calculated by assisting a lightweight sequence recommendation model.

[0112] The model parameters are updated based on the optimization results, and the model structure is locally adjusted through the adapter parameters. The model training uses an optimization objective function.

[0113] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.

[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing a recommender system based on negative sample awareness, characterized in that, The recommended system optimization method based on negative sample awareness includes: Receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data; The negative sample pool is expanded by a negative sample sharing mechanism within the batch, and the calculation results of negative samples within the same training batch are shared to reduce computational overhead. Based on the dynamic reward marginal adjustment mechanism, the impact of negative samples on model training is dynamically adjusted according to their confidence level, wherein the confidence level is calculated by assisting a lightweight sequence recommendation model. The model parameters are updated based on the optimization results obtained through the dynamic reward marginal adjustment mechanism and the intra-batch negative sample sharing mechanism, and the model structure is locally adjusted through the adapter parameters, wherein the model training adopts an optimization objective function; The dynamic reward marginal adjustment mechanism includes: Calculate negative samples User historical interaction sequence correlation score ; The confidence level of negative samples is determined based on the scores, and the reward margin is dynamically adjusted. ; Among them, α controls the adjustment range. It is a confidence benchmark updated through momentum. It is the initial reward margin. For negative sample confidence, For the user's historical interaction sequence, For negative samples, For the negative sample set, Expressing expectations, Represents user sequence From batch Mid-sampling, This indicates that negative samples are sampled from the set of negative samples; The confidence level of the negative sample is calculated using the following formula: ; in, This represents the relevance score calculated using a lightweight sequence recommendation model.

2. The recommendation system optimization method based on negative sample perception according to claim 1, characterized in that, The batch-based negative sample sharing mechanism includes: In training batches In the middle, a shared set of negative samples The calculation results, among which, For the user's historical interaction sequence, As a positive sample, It is a set of negative samples randomly sampled from an unobserved interaction space; Decomposition by function ; User historical interaction sequence Mapped to an intermediate representation, and the final log probability is generated, where, This indicates the sequence of user's historical interactions. The function mapped to the intermediate representation. It is a function that generates the final logarithmic probability.

3. The recommendation system optimization method based on negative sample perception according to claim 1, characterized in that, The optimization objective function is: ; in, Represents a given user's historical interaction sequence Predict positive samples In the negative sample set In the context of preference probability, γ represents the dynamic reward margin. For the sigmoid function, Expressing expectations, Indicates sample Sample from dataset D.

4. The recommendation system optimization method based on negative sample perception according to claim 1, characterized in that, The model parameter update steps specifically include: Only update the adapter parameters, retain the frozen parameters of the model, in order to avoid interfering with the original performance of the model; Solve the parameter optimization problem using the stochastic gradient descent algorithm.

5. A recommender system optimization device based on negative sample perception, characterized in that, include: The negative sample optimization request module is configured to: receive negative sample optimization requests submitted by users or the system, and parse the requests to determine optimization targets and related data; The intra-batch negative sample sharing module is configured to expand the negative sample pool through the intra-batch negative sample sharing mechanism and share the negative sample calculation results within the same training batch to reduce computational overhead. The dynamic reward marginal adjustment module is configured to: dynamically adjust the impact of negative samples on model training based on the confidence level of the negative samples according to the dynamic reward marginal adjustment mechanism, wherein the confidence level is calculated by assisting the lightweight sequence recommendation model; The model parameter update module is configured to update the model parameters based on the optimization results obtained through the dynamic reward marginal adjustment mechanism and the intra-batch negative sample sharing mechanism, and to locally adjust the model structure through the adapter parameters, wherein the model training adopts an optimization objective function; The dynamic reward marginal adjustment mechanism includes: Calculate negative samples User historical interaction sequence correlation score ; The confidence level of negative samples is determined based on the scores, and the reward margin is dynamically adjusted. ; Among them, α controls the adjustment range. It is a confidence benchmark updated through momentum. It is the initial reward margin. For negative sample confidence, For the user's historical interaction sequence, For negative samples, For the negative sample set, Expressing expectations, Represents user sequence From batch Mid-sampling, This indicates that negative samples are sampled from the set of negative samples; The confidence level of the negative sample is calculated using the following formula: ; in, This represents the relevance score calculated using a lightweight sequence recommendation model.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the recommendation system optimization method based on negative sample awareness as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the recommendation system optimization method based on negative sample awareness as described in any one of claims 1 to 4.