Recommendation system adjustment method and video recommendation system adjustment method

By obtaining client operation behavior data and matching adjustment strategies, and using reinforcement learning algorithms to optimize the recommendation system, the problem of poor user experience is solved, and higher recommendation accuracy and user trust are achieved.

CN120687633APending Publication Date: 2025-09-23HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410324799.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing recommendation systems suffer from poor data quality, algorithm modeling bias, and lack of real-time user interest modeling, leading to poor user experience, reduced user trust and engagement. In addition, the repair methods of manual strategy mechanisms are costly and do not differentiate between user experience issues.

Method used

By obtaining the client's operational behavior data, determining the target feedback type and matching the corresponding adjustment strategy, the recommendation system is adjusted, the reinforcement learning algorithm is used to optimize the user experience, and personalized adjustment strategies are designed to improve the accuracy of the recommendation system.

Benefits of technology

It improves the accuracy of the recommendation system, provides personalized and accurate recommendation results, reduces the occurrence of user experience problems, and increases user trust and platform revenue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687633A_ABST
    Figure CN120687633A_ABST
Patent Text Reader

Abstract

The invention discloses an adjusting method of a recommendation system and an adjusting method of a video recommendation system. The method comprises the steps that under the condition that recommendation data are sent to a client side through a recommendation system, operation behavior data, aiming at the recommendation data, of the client side are acquired; determining whether a target feedback type matched with the operation behavior data exists in the multiple feedback types or not; under the condition that the target feedback type exists in the multiple feedback types, determining a target adjustment strategy matched with the operation behavior data from at least one adjustment strategy corresponding to the target feedback type; and adjusting the recommendation system based on the target adjustment strategy. The technical problem of low accuracy of a recommendation system in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of adjustment of recommendation systems, and in particular to an adjustment method for a recommendation system and an adjustment method for a video recommendation system. Background Art

[0002] At present, the user experience in the recommendation system has deteriorated due to poor data quality, algorithm modeling deviations, lack of real-time evolution of user interest modeling, and overfitting of the model to training data. User experience issues will reduce users' trust and participation in the recommendation system. In order to improve the user experience of the recommendation system, the current main method of adjusting the recommendation system is to use artificial strategy mechanisms. The specific methods of artificial strategy mechanisms are strategy development and strategy experiments to observe whether the problem is solved and whether the solution of the problem will lead to loss of effect. However, the experience problems reported by users in the artificial strategy mechanism vary from person to person. For example, the diversity experience problem may not be a problem for some users, but the problem repair method of the artificial strategy mechanism will achieve indiscriminate processing through rules or strategies, resulting in a poor user experience.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a method for adjusting a recommendation system and a method for adjusting a video recommendation system, so as to at least solve the technical problem of low accuracy of the recommendation system in the related art.

[0005] According to one aspect of an embodiment of the present application, a method for adjusting a recommendation system is provided, comprising: when recommendation data is sent to a client via the recommendation system, obtaining operational behavior data of the client with respect to the recommendation data; determining whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when there is a target feedback type among the multiple feedback types, determining a target adjustment strategy that matches the operational behavior data from at least one adjustment strategy corresponding to the target feedback type; and adjusting the recommendation system based on the target adjustment strategy.

[0006] According to another aspect of an embodiment of the present application, a method for adjusting a video recommendation system is also provided, including: when a recommended video is sent to a client through the video recommendation system, obtaining the client's viewing behavior data for the recommended video; determining whether there is a target viewing question type that matches the viewing behavior data among multiple viewing question types; when there is a target viewing question type among multiple viewing question types, determining a target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing question type; and adjusting the video recommendation system based on the target adjustment strategy.

[0007] According to another aspect of an embodiment of the present application, an adjustment device for a recommendation system is provided, including: an acquisition module for acquiring operational behavior data of a client with respect to the recommendation data when the recommendation data is sent to the client via the recommendation system; a determination module for determining whether there is a target feedback type that matches the operational behavior data among multiple feedback types; a matching module for determining, when there is a target feedback type among multiple feedback types, a target adjustment strategy that matches the operational behavior data from at least one adjustment strategy corresponding to the target feedback type; and an adjustment module for adjusting the recommendation system based on the target adjustment strategy.

[0008] According to another aspect of an embodiment of the present application, an adjustment device for a video recommendation system is also provided, including: an acquisition module for acquiring the viewing behavior data of the client for the recommended video when the recommended video is sent to the client through the video recommendation system; a determination module for determining whether there is a target viewing problem type that matches the viewing behavior data among multiple viewing problem types; a matching module for determining a target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing problem type when the target viewing problem type exists among multiple viewing problem types; and an adjustment module for adjusting the video recommendation system based on the target adjustment strategy.

[0009] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes any one of the methods in the above embodiments when running.

[0010] According to another aspect of the embodiments of the present application, a computer terminal is further provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the method of each embodiment of the present application is executed when the program is running.

[0011] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present application.

[0012] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which implements the methods in various embodiments of the present application when executed by a processor.

[0013] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.

[0014] According to another aspect of the embodiments of the present application, a computer program is further provided, which implements the methods in various embodiments of the present application when executed by a processor.

[0015] Through the above steps, when the recommendation data is sent to the client through the recommendation system, the client's operational behavior data for the recommendation data is obtained; it is determined whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when the target feedback type exists among multiple feedback types, the target adjustment strategy that matches the operational behavior data is determined from at least one adjustment strategy corresponding to the target feedback type; the recommendation system is adjusted based on the target adjustment strategy, thereby improving the accuracy of the recommendation system; it is easy to notice that when the recommendation system sends recommendation data to the client, the client's operational behavior data for the recommendation data can be used to determine the client user's experience problem with the recommendation system, that is, the feedback type. For different experience problems, corresponding adjustment strategies can be set to adjust the recommendation system. By continuously repairing the experience problems, the recommendation accuracy of the recommendation system is improved, thereby solving the technical problem of low accuracy of the recommendation system in related technologies.

[0016] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for adjusting a recommendation system according to an embodiment of the present application;

[0019] Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application;

[0020] Figure 3 This is a structural block diagram of a service grid according to an embodiment of the present application;

[0021] Figure 4 is a flowchart of a method for adjusting a recommendation system according to Example 1 of the present application;

[0022] Figure 5 is a schematic diagram of feature extraction according to an embodiment of the present application;

[0023] Figure 62 is a schematic diagram of a feedback structure for fine-tuning a language model according to an embodiment of the present application;

[0024] Figure 7 2 is a schematic diagram of a feedback structure of a recommendation system according to an embodiment of the present application;

[0025] Figure 8 is a flowchart of a method for adjusting a video recommendation system according to Example 2 of the present application;

[0026] Figure 9 is a schematic diagram of an adjustment device of a recommendation system according to an embodiment of the present application;

[0027] Figure 10 is a schematic diagram of an adjustment device for a video recommendation system according to an embodiment of the present application;

[0028] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0032] Language Model Fine-tuning (LLM) reflection technology: LLM self-evaluates and corrects itself based on its output content to provide more accurate and reasonable answers.

[0033] Advantage Actor Critic (A2C) reinforcement learning algorithm: By separating the strategy (actor) and the value function (critic), it simultaneously learns the action distribution and value evaluation, and uses the advantage function (advantage function) to reduce the reward variance to improve the algorithm efficiency.

[0034] User experience issues: Inaccurate recommendations that do not meet user interests and needs or may cause user dissatisfaction occur during the recommendation process.

[0035] Recommendation reflection mechanism: refers to detecting user experience problems through the recommendation system and adjusting the probability distribution of actions needed to solve the experience problems based on user feedback to achieve self-optimization.

[0036] At present, recommendation systems generally have user experience problems, mainly including data quality and coverage, algorithm limitations, diversity and complexity of user behavior, and social and ethical factors.

[0037] Among them, data quality and coverage refer to the fact that the recommendation algorithm relies on user behavior data for recommendations. If there is bias in the data or the data is incomplete, it will affect the generalization of the model and ultimately affect the accuracy of the recommendation. Secondly, for new users or new products, there is a lack of sufficient data for recommendations. Finally, abnormal records or errors in user behavior data will interfere with the learning process of the recommendation algorithm.

[0038] Algorithm limitations mean that the model may overfit the training data, have poor generalization capabilities, and be unable to adapt to unseen data. Algorithm updates are performed under certain assumptions. For example, the duration estimation task assumes that the distribution of playback duration follows a Gaussian distribution, but these assumptions may not hold true in practice. The recommendation system may focus too much on user historical behavior and ignore the diversity and novelty of the recommendation results.

[0039] The diversity and complexity of user behavior means that users' interests and behaviors are highly dynamic and complex, and a single recommendation strategy is difficult to adapt to all situations.

[0040] Social and ethical factors mean that the recommendation system may confine users to an information cocoon, that is, only recommending content that is highly similar to the user's past behavior, limiting the user's vision. The recommendation results may unconsciously reflect or exacerbate the bias in the data, leading to unfair recommendations.

[0041] User experience issues can impact short-term platform revenue and may even impact a company's long-term reputation, user trust, and operating performance. Therefore, minimizing these issues is crucial. Traditionally, manual policy mechanisms have been used to address user experience issues. The problem-solving process involves developing a strategy, conducting A / B experiments, and observing whether the issue has been resolved and whether resolving it has resulted in a loss of effectiveness. However, the drawbacks are the long and costly iteration cycles required for fixing these issues. Furthermore, user feedback on user experience issues varies from user to user; for example, diverse user experience issues may not be a problem for some users. However, manual policy mechanisms, through rules or policies, address these issues indiscriminately.

[0042] Compared with the current industry's manual policy mechanism processing methods, this application proposes a self-closed-loop user experience problem processing solution to reduce the occurrence of user experience problems and provide users with more personalized and accurate recommendation results. This application needs to define several important user experience problems in combination with task requirements, and develop k problem repair strategies (actions) for each type of experience problem; based on the LLM reflection technology framework, the reflection mechanism construction solution of the recommendation system is designed, and the Advantage Actor Critic algorithm is used to design rewards with the goal of improving the consumption indicators of the recommendation system. The action distribution that needs to be executed for various experience problems is learned, and the experience problems are solved while improving the effectiveness indicators of the recommendation system.

[0043] Example 1

[0044] According to an embodiment of the present application, a method for adjusting a recommendation system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for adjusting a recommendation system according to an embodiment of the present application. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0046] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0047] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the methods in the above embodiments. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0049] The display may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0050] Figure 1 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the computer terminal 10 (or mobile device), but also as an exemplary block diagram of the server. In an optional embodiment, Figure 2 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of a computing node in the computing environment 201 . Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application, such as Figure 2 As shown, computing environment 201 includes multiple computing nodes (e.g., servers) (illustrated as 210-1, 210-2, ...) running on a distributed network. Each computing node includes local processing and memory resources, and end user 202 can remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in computing environment 201, representing services "A," "D," "E," and "H," respectively.

[0051] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to the ingress gateway 230. The ingress gateway 230 may include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in the computing environment 201).

[0052] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine, executing programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system (OS) so that multiple workloads can run on a single operating system instance.

[0053] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively, Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively, containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls network functions related to the service, such as routing and load balancing. Other services can also be equipped with similar Pods.

[0054] During operation, executing a user request from the end user 202 may require calling one or more services in the computing environment 201, and executing one or more functions of a service may require calling one or more functions of another service. Figure 2 As shown, service “A” 220 - 1 receives a user request from end user 202 from ingress gateway 230 , service “A” 220 - 1 may call service “D” 220 - 2 , and service “D” 220 - 2 may request service “E” 220 - 3 to perform one or more functions.

[0055] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.

[0056] In another optional embodiment, Figure 3 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of the service grid. Figure 3This is a structural diagram of a service grid according to an embodiment of the present application. Figure 3 As shown, the service grid 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them to run on different clusters / machines.

[0057] like Figure 3 As shown, the microservices may include application service instance A and application service instance B, which form the functional application layer of the service grid 300. In one embodiment, application service instance A runs in the form of container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs in the form of container / process 310 on machine / workload container group 316 (Pod).

[0058] In one implementation, application service instance A may be a product query service, and application service instance B may be a product ordering service.

[0059] like Figure 3 As shown, application service instance A and grid proxy (sidecar) 303 coexist in machine workload container group 314, while application service instance B and grid proxy 305 coexist in machine workload container 316. Grid proxy 303 and grid proxy 305 form the data plane layer (dataplane) of service grid 300. Grid proxy 303 and grid proxy 305 run as container / process 304 and container / process 306, respectively, and can receive requests 312 for product query services. Bidirectional communication is possible between grid proxy 303 and application service instance A, and between grid proxy 305 and application service instance B. Furthermore, bidirectional communication is possible between grid proxy 303 and grid proxy 305.

[0060] In one embodiment, the traffic of application service instance A is routed to the appropriate destination via grid proxy 303, and the network traffic of application service instance B is routed to the appropriate destination via grid proxy 305. It should be noted that the network traffic mentioned herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), the high-performance, general-purpose open source framework (Google Remote Procedure Call, gRPC), the open source in-memory data structure storage system (Redis), and other forms.

[0061] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in service mesh 300. Service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0062] like Figure 3 As shown, the service grid 300 also includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, and these services are hosted by a hosting control plane component 301 in a machine / workload container group (machine / Pod) 302. Figure 3 As shown, managed control plane component 301 communicates bidirectionally with mesh proxy 303 and mesh proxy 305. Managed control plane component 301 is configured to perform certain control and management functions. For example, managed control plane component 301 receives telemetry data transmitted by mesh proxy 303 and mesh proxy 305 and can further aggregate this telemetry data. Managed control plane component 301 also provides user-oriented application programming interfaces (APIs) to facilitate manipulation of network behavior and to provide configuration data to mesh proxy 303 and mesh proxy 305.

[0063] Under the above operating environment, this application provides Figure 4 The tuning method of the recommendation system shown. Figure 4 is a flowchart of a method for adjusting a recommendation system according to Example 1 of the present application, the method comprising:

[0064] Step S402: When recommendation data is sent to the client via the recommendation system, operation behavior data of the client with respect to the recommendation data is obtained.

[0065] The above-mentioned recommendation system can be applied to products with recommendation tasks, among which the recommendation system can include but is not limited to news recommendation system, picture and text recommendation system, video recommendation system, short drama recommendation system, advertising recommendation system, e-commerce recommendation system, content recommendation system, etc.

[0066] The client mentioned above may be a client used by a user, and the recommendation data sent by the recommendation system may be displayed through the client. The client may display an interactive interface, and the recommendation data recommended by the recommendation system may be displayed in the interactive interface.

[0067] The above-mentioned recommended data refers to the data displayed on the client through the recommendation system, wherein the recommended data includes but is not limited to text, video, image, live broadcast, etc., and the recommended data is not limited here.

[0068] The aforementioned operational behavior data may be operational behavior data related to recommended data generated by the client, wherein the operational behavior data may be behavioral data generated when a user views recommended data through the client, the operational behavior data may be behavioral log analysis of the client regarding recommended data, or the operational behavior data may be a survey file of the client regarding recommended data. When the operational behavior data may be behavioral data generated when a user views recommended data through the client, the operational behavior data may include but is not limited to clicking to play, pausing, dragging the video playback progress bar, liking, commenting, adding to favorites, expressing disinterest, forwarding, sharing, closing, clicking on ads, skipping ads, switching between landscape and portrait modes, etc. No limitation is imposed on the operational behavior data herein.

[0069] In an optional embodiment, recommendation data can be sent to the client through the recommendation system so that the user can view the recommendation data through the client. The user can perform relevant operations on the recommended data based on his or her own interests, preferences, and other information about the recommended data, such as likes, comments, and collections, thereby generating client operation behavior data for the recommended data.

[0070] Step S404, determining whether there is a target feedback type matching the operation behavior data among the multiple feedback types;

[0071] The above-mentioned multiple feedback types can be pre-defined user feedback types. The feedback type can also be expressed as negative feedback from the client user to the recommendation system, that is, the client user's experience problems with the recommendation system. The feedback type can also be expressed as positive feedback from the client user to the recommendation system. Among them, multiple feedback types that may appear in the application scenario can be pre-defined according to the actual application scenario of the recommendation system.

[0072] The above-mentioned multiple feedback types may include positive feedback type and negative feedback type, wherein the positive feedback type is used to indicate that the client gives positive feedback to the recommendation data, and the negative feedback type is used to indicate that the client gives negative feedback to the recommendation system.

[0073] The above-mentioned multiple feedback types may be manually preset or obtained through data mining. The setting method of the multiple feedback types is not limited here.

[0074] It should be noted that feedback can be positive behaviors: like, finish broadcast, forward, favorite, comment, or negative behaviors: feedback of lack of interest, quickly scrolling past the content, etc.

[0075] In an optional embodiment, the operational behavior data can be analyzed to determine whether it matches one of multiple pre-set feedback types. If the operational behavior data successfully matches a target feedback type among the multiple feedback types, it can be determined that the client user has an experience issue, and subsequent adjustment steps need to be implemented based on the target feedback type corresponding to the operational behavior data. If the operational behavior data does not successfully match multiple feedback types, it is determined that the client user does not have an experience issue, and the recommendation system can send the recommendation data to the client according to the normal process.

[0076] Taking the video recommendation system as an example, multiple feedback types may include but are not limited to user explicit feedback such as lack of interest in recommended content, forgotten interests, and information cocoon.

[0077] When the operation behavior data indicates that the user is not interested in the recommended questionnaire or does not press the interest button, it can be determined that the target feedback type corresponding to the operation behavior data is the user's explicit feedback that the user is not interested in the recommended content.

[0078] When the results of multiple consecutive refreshes of the operational behavior data are irrelevant to the user's long-term preference profile, for example, the user's long-term profile preference is film and television, but the recommendation results of two consecutive refreshes do not include content under the film and television category, then it is determined that the recommendation results of the recommendation system are irrelevant to the user's interests, and the target feedback type corresponding to the operational behavior data can be determined to be interest forgetting.

[0079] When the operational behavior data is over a period of time, the recommended content is the categories that the user has interacted with in the past. Since the recommendation model generally has the ability to remember and analyze, for high-frequency and deep users with rich behaviors, the results of the model recommendations are all categories of historical interactions. Users are trapped in the information cocoon with which they have interacted, and it is difficult for them to obtain hot spots within the terminal and newly produced content under other categories that have not been interacted with. It is determined that the recommendation results of the recommendation system focus too much on user interests, and it is difficult to achieve interest migration. It can be determined that the target feedback type corresponding to the operational behavior data is information cocoon.

[0080] Step S406 : When a target feedback type exists among the multiple feedback types, a target adjustment strategy matching the operation behavior data is determined from at least one adjustment strategy corresponding to the target feedback type.

[0081] The above-mentioned multiple feedback types may be respectively provided with corresponding adjustment strategies, and the number of the adjustment strategies may be one or more.

[0082] The at least one adjustment strategy mentioned above may be set manually or obtained through data mining. The setting method of the adjustment strategy is not limited here.

[0083] The above adjustment strategies can be applied to multiple stages of the recommendation system, for example, the recall stage, the sorting stage, and the rearrangement stage, and the multiple stages are not limited here. The types of adjustment strategies may include but are not limited to filtering, weight adjustment, and forced intervention. There is no specific limitation on the application stage and type of the adjustment strategy here, and it can be set according to the actual situation. The above filtering adjustment strategy can be to filter the recommended data that the user is not interested in and no longer push it to the user's client. The above weight adjustment strategy can be to adjust the weights of different types of recommended data in the recommendation system. The above forced intervention strategy can be to reduce the number of displays of categories to which the user is not interested, set a maximum number of displays, or simply not display the content under the category to which the user is not interested.

[0084] In an optional embodiment, when the target feedback type is determined and there is only one adjustment strategy, the adjustment strategy can be directly determined as the target adjustment strategy that matches the operational behavior data. Alternatively, when the target feedback type is determined and there are multiple adjustment strategies, the target adjustment strategy can be determined based on the degree of match between the multiple adjustment strategies and the operational behavior data. For example, if the operational behavior data indicates that the user is not interested in the recommended data, the adjustment strategy can be a filtering adjustment strategy that uses the recall layer to reduce the recall ratio of tags belonging to items that the user is not interested in.

[0085] Taking the case where the target feedback type is the user displaying feedback that they are not interested in the recommended content as an example, the first adjustment strategy (action1) can be to do nothing and keep the original recommendation; the second adjustment strategy (action2) can be a filtering adjustment strategy, in which the recall layer reduces the recall ratio of the category to which the user's items are not interested; the third adjustment strategy (action3) can be a filtering adjustment strategy, in which the recall layer reduces the recall ratio of the tag to which the user's items are not interested; the fourth adjustment strategy (action4) can be a weighting adjustment strategy, in which the sorting layer reduces the model score of the category to which the user's items are not interested; the fifth adjustment strategy (action5) can be a weighting adjustment strategy, in which the sorting layer reduces the model score of the tag to which the user's items are not interested; the sixth adjustment strategy (action6) can be a forced intervention adjustment strategy, in which the re-ranking layer reduces the number of displayed items in the category to which the user's items are not interested, sets a maximum number of displayed items, or directly does not display the content under the category to which the user is not interested.

[0086] Step S408: Adjust the recommendation system based on the target adjustment strategy.

[0087] In an optional embodiment, the recall phase, the sorting phase, or the rearrangement phase of the recommendation system may be adjusted through a target adjustment strategy so that the recommendation system can push recommendation data that meets the needs of the user to the client.

[0088] The recall stage represents the process of filtering out a portion of potentially interesting recommendation data from a large amount of candidate recommendation data based on the user's historical behavior and preferences in the recommendation system for subsequent sorting and recommendation.

[0089] The sorting stage represents sorting the recommended data filtered out in the recall stage according to certain algorithms and rules, making it more likely for users to see and click on the recommended data they are interested in.

[0090] The re-ranking stage represents the re-ranking of the sorted recommendation data based on the user's real-time behavior and feedback after the sorting stage to ensure that the recommendation data the user sees is more in line with his or her current interests and needs.

[0091] In another optional embodiment, after the recommendation system is adjusted based on the target adjustment strategy, the recommendation system can be continuously adjusted according to the client's operational behavior data on the recommendation data, and can also be periodically adjusted according to the client's operational behavior data on the recommendation data. There is no limit on the adjustment cycle and number of adjustments of the recommendation system here, and the adjustment method of the recommendation system can be determined according to actual needs.

[0092] The adjustment method of the recommendation system of this application can be applied to various popular recommendation systems, such as content recommendation, e-commerce recommendation, advertising recommendation, etc., which are not limited here. The application scenario of the recommendation system can be determined based on the products that actually need the recommendation system. This application first defines different types of experience problems and the repair strategies that need to be implemented for each type of experience problem by designing a reflection mechanism for the recommendation system; uses a reinforcement learning algorithm to design reward values ​​with the goal of improving the overall consumption index of the recommendation system, and continuously optimizes the repair strategies that need to be implemented for different types of experience problems.

[0093] Through the above steps, when the recommendation data is sent to the client through the recommendation system, the client's operational behavior data for the recommendation data is obtained; it is determined whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when the target feedback type exists among multiple feedback types, the target adjustment strategy that matches the operational behavior data is determined from at least one adjustment strategy corresponding to the target feedback type; the recommendation system is adjusted based on the target adjustment strategy, thereby improving the accuracy of the recommendation system; it is easy to notice that when the recommendation system sends recommendation data to the client, the client's operational behavior data for the recommendation data can be used to determine the client user's experience problem with the recommendation system, that is, the feedback type. For different experience problems, corresponding adjustment strategies can be set to adjust the recommendation system. By continuously repairing the experience problems, the recommendation accuracy of the recommendation system is improved, thereby solving the technical problem of low accuracy of the recommendation system in related technologies.

[0094] In the above embodiment of the present application, a target adjustment strategy that matches the operation behavior data is determined from multiple adjustment strategies corresponding to the target feedback type, including: obtaining object attribute information of the operation object corresponding to the client, and context information associated with the operation behavior data, wherein the context information at least includes device attribute information of the client, and environmental information corresponding to the moment when the operation behavior data occurs; performing feature extraction on the operation behavior data, object attribute information and context information to obtain state features corresponding to the operation behavior data; and using a policy network to determine the target adjustment strategy that matches the operation behavior data from multiple adjustment strategies corresponding to the target feedback type based on the state features.

[0095] The aforementioned object attribute information may be static user features, including but not limited to features that rarely change, such as the user's predicted age, permanent residence (province, city, county, etc.), and user profile features (occupation, demographic preferences, whether or not there are children). Object attribute information is not specifically limited here and can be determined based on actual usage scenarios.

[0096] The aforementioned context information may be a context feature (Context). The device attribute information in the context information may include, but is not limited to, the type of device used (tablet, mobile phone), the device operating system, and the application version number. The environmental information in the context information may include, but is not limited to, the time of the behavior (day of the week, hour, weekday, holiday), and weather information (rainy, cold snap, heavy snow, etc.). The context information is not specifically limited here and may be determined based on actual usage scenarios.

[0097] The above-mentioned state embedding may be obtained by extracting multiple features from the operation behavior data, object attribute information, and context information respectively, and then concatenating the multiple features.

[0098] The aforementioned policy network may be a pre-trained network for selecting a target adjustment strategy.

[0099] Optionally, when the recommendation system detects experience problems in user feedback, it can determine multiple adjustment strategies for target feedback types corresponding to the experience problems, and determine the target adjustment strategy that the current recommendation system needs to execute based on the extracted state features.

[0100] In an optional embodiment, the object attribute information of the client's corresponding operation and the context information associated with the operation behavior data can be obtained. The context information associated with the operation behavior data can be obtained by including the client's device attribute information and the environmental information corresponding to the time when the operation behavior data occurs. Clients corresponding to different device attribute information may have different operation habits, and different environmental information may also have different impacts on the operation behavior data. Different object attribute information may produce different intentions behind the operation behavior data for the same recommended data. Therefore, the operation behavior data is feature extracted in combination with the context information and the object attribute information. By extracting features from the operation behavior data with reference to various factors, more accurate state features can be obtained.

[0101] In the above embodiments of the present application, feature extraction is performed on the operation behavior data, object attribute information and context information to obtain state features corresponding to the operation behavior data, including: using a feature extraction model to extract features from the operation behavior data to obtain behavior features, wherein the feature extraction model at least includes an attention module; using an embedding layer to extract features from the object attribute information to obtain object attribute features; using an embedding layer to extract features from the context information to obtain context features; splicing the behavior features, object attribute features and context features to obtain splicing features; and performing matrix transformation on the splicing features to obtain state features.

[0102] The feature extraction model described above can be a bidirectional long short-term memory (Bi-LSTM) within a long short-term memory (LSTM) network. LSTM is a recurrent neural network-based algorithm that can be used to process both time series and sequence data and capture long-term dependencies. Bi-LSTM adds a backpropagation structure to LSTM to better capture interdependencies between items in a sequence. The feature extraction model here can also be other models used for feature extraction. The type of feature extraction model is not specifically limited and can be determined based on actual needs.

[0103] The above-mentioned feature extraction model includes at least an attention module (self-attention), in which the function of the attention module is to calculate the attention weight of the input behavior sequence, so as to pay more attention to important behavior information during the encoding process, thereby improving the representation effect of user behavior, so as to help the model better understand the user's behavior patterns and preferences.

[0104] The embedding layer described above converts discrete input data into a continuous, dense vector representation. This conversion transforms the original high-dimensional, sparse data into low-dimensional, dense data while preserving the semantic relationships between the data. This continuous representation better captures similarities and correlations between data, enabling better learning and inference within neural networks. In fields such as natural language processing and recommender systems, embedding layers are widely used to convert discrete data such as text and categories into continuous representations, improving model performance and efficiency.

[0105] The aforementioned operational behavior data can be a user-side behavior sequence. A behavior sequence can refer to different content in different projects. In the video recommendation project, a behavior sequence can include, but is not limited to, an exposure-but-no-click sequence, a click sequence within the last 15 minutes, a click sequence within the last 50 videos, a like sequence, a disinterest sequence, a forwarding sequence, and a completed video sequence. The 15 minutes and 50 videos here are merely examples and do not impose specific numerical limits. These values ​​can be adjusted based on specific scenario requirements.

[0106] In an optional embodiment, by adopting Bi-LSTM in the LSTM algorithm, the user-side behavior sequence can be bidirectionally represented to obtain the mutual dependence relationship between items in the behavior sequence, thereby obtaining the temporal information of the content in the behavior sequence and the feature representation of the item in the sequence, that is, to obtain the above-mentioned behavior characteristics, the two features (embedding) of the forward prediction value h1 and the reverse prediction value h2 can be element-wise added (emement-wise) to represent the user's real-time sequence information, thereby obtaining the behavior characteristics.

[0107] In an optional embodiment, the embedding layer can be used to extract features of object attribute information and context information to obtain object attribute features and context features. The behavioral features, object attribute features and context features can be spliced ​​to obtain spliced ​​features. By performing matrix transformation on the spliced ​​features, feature vectors from different sources can be integrated and transformed to obtain the final user-side representation feature vector, that is, the above-mentioned state features, which facilitates subsequent model training and prediction.

[0108] Figure 5 is a schematic diagram of feature extraction according to an embodiment of the present application, such as Figure 5 As shown, a bidirectional long short-term memory network and an attention module can be used to perform feature encoding on the operation behavior data, that is, the user's behavior sequence (Seq fea), to obtain behavior features. The embedding layer is used to encode the object attribute information (User Fea) and the context information (Context Fea) to obtain object attribute features and context features. The behavior features, object attribute features and context features can be spliced ​​to obtain spliced ​​features. The spliced ​​features can be matrix-transformed to obtain a representation vector on the user side, that is, a state feature.

[0109] In the above embodiment of the present application, a policy network is used to determine a target adjustment strategy that matches the operation behavior data from multiple adjustment strategies corresponding to the target feedback type based on state characteristics, including: using the policy network to determine the output probabilities of multiple adjustment strategies based on state characteristics, wherein the output probability is used to characterize the degree of matching between the multiple adjustment strategies and the operation behavior data; obtaining the adjustment strategy corresponding to the maximum output probability among the multiple adjustment strategies to obtain the target adjustment strategy.

[0110] In an optional embodiment, the state characteristics and multiple adjustment strategies can be matched to obtain the degree of matching between the state characteristics and the adjustment strategies, and the adjustment strategy with a higher degree of matching with the state characteristics can be selected from the multiple adjustment strategies as the target feature strategy, so that the user experience problems encountered can be fixed in a targeted manner through the target feature strategy, or the recommendation data that the user is interested in can be further strengthened, thereby improving the recommendation accuracy of the recommendation system.

[0111] In the above embodiment of the present application, the method also includes: when sending a first recommendation sample to the client through the recommendation system, obtaining a first training sample, wherein the first training sample includes: a first operation behavior sample for the first recommendation sample, object attribute information, and a first context sample associated with the first operation behavior sample; using the policy network to determine a training adjustment strategy based on the first training sample; adjusting the recommendation system based on the training adjustment strategy to obtain an adjusted recommendation system; when sending a second recommendation sample through the adjusted recommendation system, obtaining a second training sample and a target reward value for the second training sample, wherein the second training sample includes: a second operation behavior sample for the second recommendation sample, object attribute information, and a second context sample associated with the second operation behavior sample; evaluating the first training sample using the evaluation network to obtain a first evaluation result corresponding to the first training sample, and evaluating the second training sample using the evaluation network to obtain a second evaluation result corresponding to the second training sample; adjusting the network parameters of the policy network and the evaluation network based on the first evaluation result, the second evaluation result and the target reward value.

[0112] The above-mentioned policy network (Actor) is responsible for selecting the probability distribution of actions based on the current state.

[0113] The above-mentioned evaluation network (Critic) is a value function network responsible for evaluating the value of the state.

[0114] In an optional embodiment, user feedback on the recommendation system can be categorized into multiple categories, and remediation strategies (actions) for each category can be defined. Next, data can be collected, with samples collected in the form of tuples, and state features can be used as underlying input features (raw features). The collected data can be used to train the policy network and the evaluation network for n training epochs, which can be 5 to 10, but not limited to this number. The specific number of training epochs can be set based on actual needs. Based on the defined reward value, a temporal difference algorithm can be used to calculate the advantage function value to regress the evaluation network.

[0115] The first training sample mentioned above can be the initial state S, and the second training sample mentioned above can be the state S ′ .

[0116] The policy network can be used to determine the training adjustment policy (action) based on the first training sample, and the recommendation system can be adjusted based on the training adjustment policy to obtain the adjusted recommendation system. The adjusted recommendation system can be used to send the second recommendation sample and obtain the second training sample. In the first adjustment process, the training adjustment policy is determined by the user's feedback, but the user's feedback is very sparse compared to the exposure. Therefore, based on the user's initial feedback, λ2 (estimated reward) can be added once. action -Estimated reward) feedback process, so that this feedback action has a corresponding reward value.

[0117] The purpose of designing the reward value is to allow the policy network to find it through a reinforcement learning algorithm (Policy Gradient algorithm, abbreviated as PG). This method satisfies the recommendation system's estimated reward improvement and allows the action that observes the reward improvement to be handed over to the recommendation system for execution. Since the observed reward is sparse, a dense value called estimated reward is added to ensure that multiple actions in the training data have corresponding reward values.

[0118] The aforementioned reinforcement learning algorithm maximizes cumulative rewards by optimizing the policy function. Specifically, the PG algorithm updates the policy function's parameters using gradient ascent, making the probability of selecting each action proportional to the cumulative reward associated with that action. This allows the PG algorithm to learn a more optimal policy, resulting in higher rewards within the environment. The PG algorithm is often trained in conjunction with a deep learning model, such as using a neural network to represent the policy function and updating its parameters via backpropagation.

[0119] In an optional embodiment, the evaluation network can be used to evaluate the first training sample, and the first evaluation result can be determined by determining the advantage function value that can be generated by the combination of the first training sample and the training adjustment strategy, and the second evaluation result can be determined by the advantage function value that can be generated by the combination of the second training sample and the training adjustment strategy.

[0120] In the above embodiment of the present application, obtaining the target reward value of the second training sample includes: determining the first feedback reward value of the second recommended sample based on the target operation behavior data associated with the second recommended sample in the second operation behavior sample; determining the second feedback reward value of the second recommended sample based on the next operation behavior data located after the target operation behavior data in the second operation behavior sample; determining the first estimated reward value of the first recommended sample and the second estimated reward value of the second recommended sample; and summarizing the first feedback reward value, the second feedback reward value, the first estimated reward value and the second estimated reward value to obtain the target reward value.

[0121] The above target reward value can be determined based on two parts. The first part can be determined based on the user's immediate feedback information, and the second part can be determined based on the user feedback after the training adjustment strategy is executed.

[0122] The above-mentioned target reward value can be based on the target indicator for improving the recommendation system, wherein the target indicator can be an indicator that the recommendation system needs to improve, such as a consumption indicator, a browsing indicator, etc. There is no limitation on the indicator that the recommendation system needs to improve.

[0123] The next operation behavior data after the target operation behavior data can be determined based on the user's daily operation behavior data. Multiple operation behavior data of the user in a day can be pre-sorted in chronological order, and the operation behavior data marked with the adjustment policy is found as the target operation behavior data. The next operation behavior data adjacent to the target operation behavior data is the next operation behavior data after the target operation behavior data.

[0124] The first estimated reward value may be an estimated reward value when the training adjustment strategy is not executed, and the second estimated reward value may be an estimated reward value when the training adjustment strategy is executed to the maximum extent.

[0125] In an optional embodiment, the first feedback reward value of the second recommended sample can be determined based on the user's immediate feedback, and the second feedback reward value of the second recommended sample can be determined based on the impact feedback after the training adjustment strategy is executed. After the training adjustment strategy is executed, an improvement in user feedback can be observed, which means that it is considered necessary to increase the execution probability of the training adjustment strategy; conversely, the execution probability of the action decreases; the hyperparameter λ1 is the influence weight of the feedback reward value, which can be a number less than 1.0, but is not limited to this; the hyperparameter λ2 is the influence weight of the estimated reward value. The above-mentioned first estimated reward value and second estimated reward value are obtained, and the first feedback reward value, the second feedback reward value, the first estimated reward value, and the second estimated reward value are summarized to obtain the target reward value.

[0126] In the above embodiment of the present application, the first feedback reward value, the second feedback reward value, the first estimated reward value, and the second estimated reward value are summarized to obtain the target reward value, including: obtaining the difference between the second estimated reward value and the first estimated reward value to obtain the reward value error; and performing weighted sum processing on the first feedback reward value, the second feedback reward value, and the reward value error to obtain the target reward value.

[0127] In an optional embodiment, the difference between the second estimated reward value and the first estimated reward value can be obtained to obtain the reward value error. The target reward value can be obtained by weighting and processing the first feedback reward value, the second feedback reward value, and the reward value error according to a pre-set weight. The specific formula is as follows: reward = (Pv + Pvt) curr +λ1(Pv+Pvt) next +λ2(estimated reward action - estimated reward);

[0128] Among them, (Pv+Pvt) curr is the first feedback reward value, λ1(Pv+Pvt) next is the second feedback reward value, estimated reward action is the first estimated reward value, and estimatedreward is the second estimated reward value.

[0129] In the above embodiments of the present application, the network parameters of the policy network and the evaluation network are adjusted based on the first evaluation result, the second evaluation result and the target reward value, including: determining the evaluation error based on the first evaluation result, the second evaluation result and the target reward value; updating the network parameters of the evaluation network based on the evaluation error; and updating the network parameters of the policy network based on the evaluation error and the training adjustment strategy.

[0130] The above-mentioned evaluation error may be a mean square error, and the network parameters of the evaluation network may be updated according to the mean square error as a ladder.

[0131] Based on the evaluation error and training adjustment strategy, the activation function (Sofmax) or Gaussian branch function can be used to update the network parameters of the policy network.

[0132] In the above embodiment of the present application, the method also includes: when sending historical recommendation data to multiple clients through the recommendation system, obtaining historical feedback data of multiple clients for the historical recommendation data, wherein the historical feedback data includes at least one of the following: historical feedback results, historical operation data; analyzing the historical feedback data sent by at least one of the multiple clients to determine multiple feedback types; and generating at least one adjustment strategy corresponding to the feedback type.

[0133] The above-mentioned historical feedback results can be the results of whether users of multiple clients are interested in the historical recommendation data, or other feedback results. The historical operation data can be the behavioral operations performed by users of multiple clients on the historical recommendation data, such as likes, forwarding, blocking, etc., which are not limited here.

[0134] In an optional embodiment, when setting an adjustment strategy, historical feedback data of multiple clients for the historical recommendation data can be obtained when historical recommendation data is sent to multiple clients through the recommendation system. The historical feedback data sent by one or more clients among the multiple clients can be analyzed to obtain multiple feedback types of at least one client. Adjustment strategies corresponding to multiple feedback types can be pre-set so that the recommendation system can be adjusted subsequently through the adjustment strategy.

[0135] In the above embodiment of the present application, the method also includes: detecting the operation behavior data to determine whether the operation behavior data meets the preset adjustment conditions; when it is determined that the operation behavior data meets the preset adjustment conditions, determining the target adjustment strategy corresponding to the recommendation system based on the operation behavior data; when it is determined that the operation behavior data does not meet the preset adjustment conditions, prohibiting adjustment of the recommendation system.

[0136] The above preset adjustment conditions are used to indicate that there are experience problems in user feedback.

[0137] In an optional embodiment, the operational behavior data can be detected to determine whether there are user experience issues in the operational behavior data. Optionally, if the user chooses to dislike, block, etc., it is determined that there are user experience issues in the operational behavior data, and the target adjustment strategy is determined through negative feedback. Alternatively, if the user chooses to like, forward, comment, etc., it is determined that there are user experience issues in the operational behavior data, and the target adjustment strategy is determined through positive feedback. If the operational behavior data does not meet the preset adjustment conditions, there is no need to adjust the recommendation system, and the recommendation system can maintain the current state to recommend data to the client.

[0138] Figure 6 is a structural diagram of a feedback method for fine-tuning a language model according to an embodiment of the present application, such as Figure 6As shown, the user's feedback on the language model fine-tuning (LLM) can be used to generate feedback text through the LLM, using the LLM's fine-tuning instructions. Prompts can then be designed to correct the LLM's responses to questions. Self-reflection includes external feedback (obs / rewards) provided by the environment (environment) and internal feedback (evaluator). The reinforcement learning algorithm (Actor) can be adjusted based on the trajectory (trajectory) provided by the environment's rewards and the experience (experience) provided by self-feedback. The trajectory can be short-term memory, and the experience can be long-term memory. The trajectory can be determined by the environment's reward mechanism, and the experience can be determined based on the reflective text provided by the self-feedback. The reinforcement algorithm can be used to repair, or adjust, the language model.

[0139] Figure 7 This is a schematic diagram of a feedback structure of a recommendation system according to an embodiment of the present application. Figure 7 As shown, the recommendation system (Recsys) lacks the advantage of fine-tuning instruction following with a language model. Therefore, it is necessary to define different types of experience issues and the instructions to be executed for different experience issues, namely, repair strategies. Recommendation system errors can be defined based on user feedback on the recommendation system. In the recommendation system, errors are experience issues reported by users. After receiving user feedback on the recommendation system, the pre-defined repair strategy for the experience issue can be called, and the recommendation system can execute the selected repair strategy, and the system can obtain new user feedback. The reward value calculated based on user feedback serves as the target (ground truth) for action evaluation machine learning. The reflection module in the agent is responsible for learning user feedback for different actions (instructions) and increasing the execution probability of high-feedback instructions based on the estimated feedback.

[0140] Reflection is a self-assessment and correction technique that helps language models understand and generate text more accurately during fine-tuning. During language model fine-tuning, reflection helps the model examine its own output, evaluate it, and correct it to provide more reasonable and accurate answers.

[0141] During language model fine-tuning, the principles of reflection technology mainly include the following aspects:

[0142] 1. Self-evaluation: The language model evaluates its own performance based on its output, comparing the generated text with the standard answer. This self-evaluation allows the model to identify its shortcomings and make corrections and improvements.

[0143] 2. Correction strategy: After discovering its own problems, the language model will adopt corresponding correction strategies, such as adjusting model parameters, increasing training data, changing generation strategies, etc., to improve its own performance and generation quality.

[0144] 3. Self-learning: Through continuous self-evaluation and correction, the language model can gradually improve its performance and learn more accurate and reasonable text generation methods, thereby improving overall performance and effectiveness.

[0145] Through reflection technology, language models can continuously improve themselves during fine-tuning, providing more accurate and reasonable answers, thereby better meeting user needs and improving application effectiveness. This technology is of great significance in the fields of natural language processing and artificial intelligence, and it can bring more possibilities and opportunities to the development and application of language models.

[0146] Since the reflection technology in the recommendation system does not have the advantages of LLM instruction following, it is necessary to customize the problem discovery and instruction execution parts. To ensure the comprehensiveness and accuracy of the experience problems and repair strategies, this application provides a detailed description of the experience problems and repair strategies.

[0147] For experience issues, after the recommendation system sends out the content, it relies on user feedback to judge the quality of the current recommended content. Among them, the discovery mechanism of experience issues includes: behavioral analysis based on user behavior logs; explicit reporting by users through the client, such as questionnaires (to investigate satisfaction with the recommendation results and improvements), explicit feedback by clicking on the "not interested" button, and case feedback from key customers. It should be noted that the reflection mechanism is not limited to solving experience problems, but can also be used to strengthen user positive feedback. The description of this application introduces the recommendation reflection technology using experience issues as an example. Since the reflection module is improved based on the reward mechanism, instructions that can strengthen positive feedback are designed to recommend more content that users may give positive feedback. The reflection module will increase the execution probability of such instructions based on the reward mechanism.

[0148] Regarding the repair strategy: the repair strategy is defined from the perspective of the three-layer architecture of recall + sorting + re-arrangement of the recommendation system. The multi-layer instruction content corresponds to the experience problem. When a certain type of experience problem is detected, the repair strategy of the recall, sorting, and re-arrangement layer corresponding to the problem will be triggered; for example, if a user's recent swipe behavior hits the fast sliding type problem, the reflection module will obtain the classification and label information of the user's fast sliding content, and use this information to generate the recall, sorting, and re-arrangement layer strategy. In this case, the recall layer needs to reduce the recall ratio of content under the user's fast sliding content category.

[0149] This application defines the repair process of user experience problems as a reinforcement learning problem under the reflection technology framework, defines the reflection mechanism and recommendation system as Agent, regards the user as the environment, and represents the user-side behavior sequence and context characteristics when the user encounters various experience problems as state. Reward is defined as the improvement of the consumption index of the recommendation system. The entire algorithm process is modeled. When the user feedbacks the problem, the action that should be executed under this type of problem is recommended, which can bring about an improvement in the consumption index. This application can achieve that when different users feedback experience problems in different contexts, the experience problem repair action executed by the recommendation system may be different. The execution distribution of the action is obtained by optimizing offline data through A2C algorithm learning. After manually defining several major types of experience problems and the corresponding actions for each type of problem according to project requirements, the detection of experience problems and the execution of actions are automatically performed. There is no need to launch a strategy and observe the AB experiment results like the manual strategy mechanism repair method. This application saves the time cost and development cost of experimental iteration.

[0150] The recommendation reflection framework proposed in this application has strong scalability and can be flexibly and cost-effectively adapted to existing recommendation systems in the industry. It only needs to define the main experience problems faced by several recommendation systems and the corresponding repair actions in combination with the project, accumulate data online, use the A2C algorithm to train the model, and then deploy the service.

[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0152] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0153] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0154] Example 2

[0155] According to an embodiment of the present application, a method for adjusting a video recommendation system is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from this.

[0156] Figure 8 is a flow chart of a method for adjusting a video recommendation system according to Example 2 of the present application. Figure 8 As shown, the method includes the following steps:

[0157] Step S802: When a recommended video is sent to a client via the video recommendation system, obtaining viewing behavior data of the client for the recommended video;

[0158] The above-mentioned video recommendation system can be used to send recommended videos, recommended live broadcasts and other content to the client.

[0159] The above-mentioned viewing behavior data may refer to the behavior generated by the client on the recommended video, such as likes, reposts, comments, etc., and the viewing behavior data is not limited here.

[0160] Step S804, determining whether there is a target viewing problem type matching the viewing behavior data among the multiple viewing problem types;

[0161] The above-mentioned multiple viewing question types can be pre-set question types and are not specifically limited here.

[0162] Step S806 , when a target viewing problem type exists among the multiple viewing problem types, determining a target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing problem type;

[0163] Step S808: Adjust the video recommendation system based on the target adjustment strategy.

[0164] Through the above steps, when the recommended video is sent to the client through the video recommendation system, the client's viewing behavior data for the recommended video is obtained; it is determined whether there is a target viewing problem type that matches the viewing behavior data among multiple viewing problem types; when the target viewing problem type exists among multiple viewing problem types, the target adjustment strategy that matches the viewing behavior data is determined from at least one adjustment strategy corresponding to the target viewing problem type; the video recommendation system is adjusted based on the target adjustment strategy, thereby improving the accuracy of the recommendation system; it is easy to notice that when the recommendation system sends recommendation data to the client, the client's operational behavior data on the recommendation data can be used to determine the client user's experience problem with the recommendation system, that is, the feedback type. For different experience problems, corresponding adjustment strategies can be set to adjust the recommendation system. By continuously repairing the experience problems, the recommendation accuracy of the recommendation system is improved, thereby solving the technical problem of low accuracy of the recommendation system in related technologies.

[0165] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0166] Example 3

[0167] According to an embodiment of the present application, a device for adjusting a recommendation system for implementing the above-mentioned method for adjusting a recommendation system is also provided. Figure 9 is a schematic diagram of an adjustment device of a recommendation system according to an embodiment of the present application, such as Figure 9 As shown, the apparatus 900 includes: an acquisition module 902 , a determination module 904 , a matching module 906 , and an adjustment module 908 .

[0168] Among them, the acquisition module is used to obtain the client's operational behavior data for the recommended data when the recommendation data is sent to the client through the recommendation system; the determination module is used to determine whether there is a target feedback type that matches the operational behavior data among multiple feedback types; the matching module is used to determine the target adjustment strategy that matches the operational behavior data from at least one adjustment strategy corresponding to the target feedback type when the target feedback type exists among multiple feedback types; the adjustment module is used to adjust the recommendation system based on the target adjustment strategy.

[0169] It should be noted that the acquisition module 902, determination module 904, matching module 906, and adjustment module 908 correspond to steps S402 to S408 in Example 1. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules can also be run as part of the device in the computer terminal 10 provided in Example 1.

[0170] In the above embodiment of the present application, the matching module is also used to obtain object attribute information of the operation object corresponding to the client, and context information associated with the operation behavior data, wherein the context information at least includes the device attribute information of the client, and the environmental information corresponding to the moment when the operation behavior data occurs; feature extraction is performed on the operation behavior data, object attribute information and context information to obtain state features corresponding to the operation behavior data; and a target adjustment strategy that matches the operation behavior data is determined from multiple adjustment strategies corresponding to the target feedback type based on the state features using a policy network.

[0171] In the above embodiment of the present application, the matching module is also used to use the feature extraction model to extract features from the operation behavior data to obtain behavior features, wherein the feature extraction model at least includes an attention module; use the embedding layer to extract features from the object attribute information to obtain object attribute features; use the embedding layer to extract features from the context information to obtain context features; splice the behavior features, object attribute features and context features to obtain spliced ​​features; perform matrix transformation on the spliced ​​features to obtain state features.

[0172] In the above embodiment of the present application, the matching module is also used to use the policy network to determine the output probabilities of multiple adjustment strategies based on state characteristics, wherein the output probability is used to characterize the degree of matching between the multiple adjustment strategies and the operation behavior data; obtain the adjustment strategy corresponding to the maximum output probability among the multiple adjustment strategies to obtain the target adjustment strategy.

[0173] In the above embodiments of the present application, the device further includes: an evaluation module.

[0174] Among them, the acquisition module is also used for the first operation behavior sample, object attribute information, and the first context sample associated with the first operation behavior sample for the first recommendation sample; the determination module is also used to use the policy network to determine the training adjustment strategy based on the first training sample; the adjustment module is also used to adjust the recommendation system based on the training adjustment strategy to obtain the adjusted recommendation system; the acquisition module is also used to obtain the second training sample and the target reward value of the second training sample when sending the second recommendation sample through the adjusted recommendation system, wherein the second training sample includes: the second operation behavior sample, object attribute information, and the second context sample associated with the second operation behavior sample for the second recommendation sample; the evaluation module is also used to evaluate the first training sample using the evaluation network to obtain a first evaluation result corresponding to the first training sample, and to evaluate the second training sample using the evaluation network to obtain a second evaluation result corresponding to the second training sample; the adjustment module is also used to adjust the network parameters of the policy network and the evaluation network based on the first evaluation result, the second evaluation result and the target reward value.

[0175] In the above embodiment of the present application, the acquisition module is also used to determine the first feedback reward value of the second recommended sample based on the target operation behavior data associated with the second recommended sample in the second operation behavior sample; determine the second feedback reward value of the second recommended sample based on the next operation behavior data located after the target operation behavior data in the second operation behavior sample; determine the first estimated reward value of the first recommended sample and the second estimated reward value of the second recommended sample; and summarize the first feedback reward value, the second feedback reward value, the first estimated reward value and the second estimated reward value to obtain the target reward value.

[0176] In the above embodiment of the present application, the acquisition module is also used to obtain the difference between the second estimated reward value and the first estimated reward value to obtain a reward value error; and the first feedback reward value, the second feedback reward value and the reward value error are weighted and processed to obtain a target reward value.

[0177] In the above embodiments of the present application, the evaluation module is also used to determine the evaluation error based on the first evaluation result, the second evaluation result and the target reward value; update the network parameters of the evaluation network based on the evaluation error; and update the network parameters of the strategy network based on the evaluation error and the training adjustment strategy.

[0178] In the above embodiments of the present application, the device further includes: an analysis module and a generation module.

[0179] Among them, the acquisition module is also used to obtain historical feedback data of multiple clients for historical recommendation data when sending historical recommendation data to multiple clients through the recommendation system, wherein the historical feedback data includes at least one of the following: historical feedback results, historical operation data; the analysis module is used to analyze the historical feedback data sent by at least one client among the multiple clients to determine multiple feedback types; the generation module is used to generate at least one adjustment strategy corresponding to the feedback type.

[0180] In the above embodiments of the present application, the device further includes: a detection module and a prohibition module.

[0181] Among them, the detection module is used to detect the operation behavior data and determine whether the operation behavior data meets the preset adjustment conditions; the determination module is used to determine the target adjustment strategy corresponding to the recommendation system based on the operation behavior data when it is determined that the operation behavior data meets the preset adjustment conditions; the prohibition module is used to prohibit the adjustment of the recommendation system when it is determined that the operation behavior data does not meet the preset adjustment conditions.

[0182] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0183] Example 4

[0184] According to an embodiment of the present application, a video recommendation system adjustment device for implementing the above-mentioned video recommendation system adjustment method is also provided. Figure 10 is a schematic diagram of an adjustment device for a video recommendation system according to an embodiment of the present application. Figure 10 As shown, the apparatus 1000 includes: an acquisition module 1002 , a determination module 1004 , a matching module 1006 , and an adjustment module 1008 .

[0185] Among them, the acquisition module is used to obtain the client's viewing behavior data for the recommended video when the recommended video is sent to the client through the video recommendation system; the determination module is used to determine whether there is a target viewing problem type that matches the viewing behavior data among multiple viewing problem types; the matching module is used to determine the target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing problem type when the target viewing problem type exists among multiple viewing problem types; the adjustment module is used to adjust the video recommendation system based on the target adjustment strategy.

[0186] It should be noted that the acquisition module 1002, determination module 1004, matching module 1006, and adjustment module 1008 correspond to steps S802 to S808 in Example 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules can also be run as part of the device in the computer terminal 10 provided in Example 1.

[0187] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0188] Example 5

[0189] The embodiment of the present application may provide an electronic device, which may be any electronic device in a group of electronic devices. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.

[0190] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0191] In this embodiment, the computer terminal can execute the program code in the method.

[0192] Optionally, Figure 11 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 11 As shown, the electronic device A may include: one or more (only one is shown in the figure) processors 102, a memory 104, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module and a display.

[0193] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: when sending recommendation data to the client through the recommendation system, obtain the client's operational behavior data for the recommendation data; determine whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when there is a target feedback type among multiple feedback types, determine the target adjustment strategy that matches the operational behavior data from at least one adjustment strategy corresponding to the target feedback type; and adjust the recommendation system based on the target adjustment strategy.

[0195] Optionally, the processor may also execute the program code of the following steps: obtaining object attribute information of the operation object corresponding to the client, and context information associated with the operation behavior data, wherein the context information includes at least the device attribute information of the client, and the environmental information corresponding to the moment when the operation behavior data occurs; performing feature extraction on the operation behavior data, object attribute information and context information to obtain state features corresponding to the operation behavior data; and using a policy network to determine a target adjustment strategy that matches the operation behavior data from a plurality of adjustment strategies corresponding to the target feedback type based on the state features.

[0196] Optionally, the processor may also execute the program code of the following steps: performing feature extraction on the operation behavior data using a feature extraction model to obtain behavior features, wherein the feature extraction model includes at least an attention module; performing feature extraction on the object attribute information using an embedding layer to obtain object attribute features; performing feature extraction on the context information using an embedding layer to obtain context features; splicing the behavior features, object attribute features, and context features to obtain splicing features; and performing matrix transformation on the splicing features to obtain state features.

[0197] Optionally, the processor may also execute the program code of the following steps: using the policy network to determine the output probabilities of multiple adjustment strategies based on state characteristics, wherein the output probabilities are used to characterize the degree of matching between the multiple adjustment strategies and the operation behavior data; obtaining the adjustment strategy corresponding to the maximum output probability among the multiple adjustment strategies to obtain the target adjustment strategy.

[0198] Optionally, the processor may also execute the program code of the following steps: when sending a first recommendation sample to the client through the recommendation system, obtaining a first training sample, wherein the first training sample includes: a first operation behavior sample for the first recommendation sample, object attribute information, and a first context sample associated with the first operation behavior sample; using the policy network to determine a training adjustment strategy based on the first training sample; adjusting the recommendation system based on the training adjustment strategy to obtain an adjusted recommendation system; when sending a second recommendation sample through the adjusted recommendation system, obtaining a second training sample and a target reward value for the second training sample, wherein the second training sample includes: a second operation behavior sample for the second recommendation sample, object attribute information, and a second context sample associated with the second operation behavior sample; evaluating the first training sample using the evaluation network to obtain a first evaluation result corresponding to the first training sample, and evaluating the second training sample using the evaluation network to obtain a second evaluation result corresponding to the second training sample; adjusting the network parameters of the policy network and the evaluation network based on the first evaluation result, the second evaluation result and the target reward value.

[0199] Optionally, the processor may also execute the program code of the following steps: determining the first feedback reward value of the second recommended sample based on the target operation behavior data associated with the second recommended sample in the second operation behavior sample; determining the second feedback reward value of the second recommended sample based on the next operation behavior data following the target operation behavior data in the second operation behavior sample; determining the first estimated reward value of the first recommended sample and the second estimated reward value of the second recommended sample; summarizing the first feedback reward value, the second feedback reward value, the first estimated reward value and the second estimated reward value to obtain the target reward value.

[0200] Optionally, the processor may further execute the program code of the following steps: obtaining the difference between the second estimated reward value and the first estimated reward value to obtain a reward value error; performing weighted sum processing on the first feedback reward value, the second feedback reward value, and the reward value error to obtain a target reward value.

[0201] Optionally, the processor may also execute the program code of the following steps: determining the evaluation error based on the first evaluation result, the second evaluation result and the target reward value; updating the network parameters of the evaluation network based on the evaluation error; and updating the network parameters of the policy network based on the evaluation error and the training adjustment strategy.

[0202] Optionally, the processor may also execute the program code of the following steps: when sending historical recommendation data to multiple clients through the recommendation system, obtaining historical feedback data from multiple clients for the historical recommendation data, wherein the historical feedback data includes at least one of the following: historical feedback results, historical operation data; analyzing the historical feedback data sent by at least one of the multiple clients to determine multiple feedback types; and generating at least one adjustment strategy corresponding to the feedback type.

[0203] Optionally, the processor may also execute the program code of the following steps: detecting the operation behavior data to determine whether the operation behavior data satisfies the preset adjustment conditions; if it is determined that the operation behavior data satisfies the preset adjustment conditions, determining the target adjustment strategy corresponding to the recommendation system based on the operation behavior data; if it is determined that the operation behavior data does not satisfy the preset adjustment conditions, prohibiting adjustment of the recommendation system.

[0204] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: when sending a recommended video to the client through the video recommendation system, obtain the client's viewing behavior data for the recommended video; determine whether there is a target viewing question type that matches the viewing behavior data among multiple viewing question types; when there is a target viewing question type among multiple viewing question types, determine the target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing question type; and adjust the video recommendation system based on the target adjustment strategy.

[0205] According to an embodiment of the present application, when recommendation data is sent to a client through a recommendation system, the client's operational behavior data for the recommendation data is obtained; it is determined whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when there is a target feedback type among multiple feedback types, a target adjustment strategy that matches the operational behavior data is determined from at least one adjustment strategy corresponding to the target feedback type; the recommendation system is adjusted based on the target adjustment strategy, thereby improving the accuracy of the recommendation system; it is easy to notice that when the recommendation system sends recommendation data to the client, the client's operational behavior data for the recommendation data can be used to determine the client user's experience problem with the recommendation system, that is, the feedback type. For different experience problems, corresponding adjustment strategies can be set to adjust the recommendation system. By continuously repairing the experience problems, the recommendation accuracy of the recommendation system is improved, thereby solving the technical problem of low accuracy of the recommendation system in related technologies.

[0206] Those skilled in the art will understand that Figure 11The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 11 It does not limit the structure of the above electronic device. For example, the electronic device A may also include Figure 11 The present invention may include more or fewer components (such as a network interface, a display device, etc.) shown in the figure, or may have a different configuration than that shown in the figure.

[0207] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0208] Example 6

[0209] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.

[0210] Optionally, in this embodiment, the storage medium may be located in any electronic device in a group of electronic devices in a computer network, or in any mobile terminal in a group of mobile terminals.

[0211] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in the case of sending recommendation data to the client through the recommendation system, obtaining the client's operational behavior data for the recommendation data; determining whether there is a target feedback type that matches the operational behavior data among multiple feedback types; in the case of a target feedback type among multiple feedback types, determining a target adjustment strategy that matches the operational behavior data from at least one adjustment strategy corresponding to the target feedback type; and adjusting the recommendation system based on the target adjustment strategy.

[0212] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: obtaining object attribute information of the operation object corresponding to the client, and context information associated with the operation behavior data, wherein the context information includes at least the device attribute information of the client, and the environmental information corresponding to the moment when the operation behavior data occurs; performing feature extraction on the operation behavior data, object attribute information and context information to obtain state features corresponding to the operation behavior data; and using a policy network to determine a target adjustment strategy that matches the operation behavior data from a plurality of adjustment strategies corresponding to the target feedback type based on the state features.

[0213] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: using a feature extraction model to perform feature extraction on operation behavior data to obtain behavior features, wherein the feature extraction model includes at least an attention module; using an embedding layer to perform feature extraction on object attribute information to obtain object attribute features; using an embedding layer to perform feature extraction on context information to obtain context features; splicing behavior features, object attribute features and context features to obtain splicing features; performing matrix transformation on the splicing features to obtain state features.

[0214] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: using a policy network to determine the output probabilities of multiple adjustment strategies based on state characteristics, wherein the output probabilities are used to characterize the degree of matching between the multiple adjustment strategies and the operational behavior data; obtaining the adjustment strategy corresponding to the maximum output probability among the multiple adjustment strategies to obtain the target adjustment strategy.

[0215] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: when a first recommendation sample is sent to a client through a recommendation system, obtaining a first training sample, wherein the first training sample includes: a first operation behavior sample for the first recommendation sample, object attribute information, and a first context sample associated with the first operation behavior sample; using a policy network to determine a training adjustment strategy based on the first training sample; adjusting the recommendation system based on the training adjustment strategy to obtain an adjusted recommendation system; when a second recommendation sample is sent through the adjusted recommendation system, obtaining a second training sample and a target reward value for the second training sample, wherein the second training sample includes: a second operation behavior sample for the second recommendation sample, object attribute information, and a second context sample associated with the second operation behavior sample; evaluating the first training sample using an evaluation network to obtain a first evaluation result corresponding to the first training sample, and evaluating the second training sample using an evaluation network to obtain a second evaluation result corresponding to the second training sample; adjusting the network parameters of the policy network and the evaluation network based on the first evaluation result, the second evaluation result and the target reward value.

[0216] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: determining a first feedback reward value for the second recommended sample based on target operation behavior data associated with the second recommended sample in the second operation behavior sample; determining a second feedback reward value for the second recommended sample based on the next operation behavior data located after the target operation behavior data in the second operation behavior sample; determining a first estimated reward value for the first recommended sample and a second estimated reward value for the second recommended sample; and summarizing the first feedback reward value, the second feedback reward value, the first estimated reward value, and the second estimated reward value to obtain a target reward value.

[0217] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: obtaining the difference between the second estimated reward value and the first estimated reward value to obtain a reward value error; and weighting and processing the first feedback reward value, the second feedback reward value, and the reward value error to obtain a target reward value.

[0218] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: determining an evaluation error based on the first evaluation result, the second evaluation result and the target reward value; updating the network parameters of the evaluation network based on the evaluation error; and updating the network parameters of the policy network based on the evaluation error and the training adjustment strategy.

[0219] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: when sending historical recommendation data to multiple clients through the recommendation system, obtaining historical feedback data of the multiple clients for the historical recommendation data, wherein the historical feedback data includes at least one of the following: historical feedback results, historical operation data; analyzing the historical feedback data sent by at least one of the multiple clients to determine multiple feedback types; and generating at least one adjustment strategy corresponding to the feedback type.

[0220] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: detecting the operational behavior data to determine whether the operational behavior data meets the preset adjustment conditions; when it is determined that the operational behavior data meets the preset adjustment conditions, determining a target adjustment strategy corresponding to the recommendation system based on the operational behavior data; when it is determined that the operational behavior data does not meet the preset adjustment conditions, prohibiting adjustment of the recommendation system.

[0221] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: in the case of sending a recommended video to a client through a video recommendation system, obtaining the client's viewing behavior data for the recommended video; determining whether there is a target viewing question type among multiple viewing question types that matches the viewing behavior data; in the case of a target viewing question type among multiple viewing question types, determining a target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing question type; and adjusting the video recommendation system based on the target adjustment strategy.

[0222] According to an embodiment of the present application, when recommendation data is sent to a client through a recommendation system, the client's operational behavior data for the recommendation data is obtained; it is determined whether there is a target feedback type that matches the operational behavior data among multiple feedback types; when there is a target feedback type among multiple feedback types, a target adjustment strategy that matches the operational behavior data is determined from at least one adjustment strategy corresponding to the target feedback type; the recommendation system is adjusted based on the target adjustment strategy, thereby improving the accuracy of the recommendation system; it is easy to notice that when the recommendation system sends recommendation data to the client, the client's operational behavior data for the recommendation data can be used to determine the client user's experience problem with the recommendation system, that is, the feedback type. For different experience problems, corresponding adjustment strategies can be set to adjust the recommendation system. By continuously repairing the experience problems, the recommendation accuracy of the recommendation system is improved, thereby solving the technical problem of low accuracy of the recommendation system in related technologies.

[0223] Example 7

[0224] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.

[0225] Example 8

[0226] The embodiments of the present application further provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program that, when executed by a processor, implements the method provided in the embodiments above.

[0227] Example 9

[0228] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.

[0229] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0230] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0231] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0232] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0233] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0234] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0235] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for adjusting a recommendation system, characterized in that: include: When recommendation data is sent to a client through a recommendation system, obtaining operation behavior data of the client in response to the recommendation data; determining whether there is a target feedback type matching the operation behavior data among the multiple feedback types; In a case where the target feedback type exists among the multiple feedback types, determining a target adjustment strategy that matches the operation behavior data from at least one adjustment strategy corresponding to the target feedback type; The recommendation system is adjusted based on the target adjustment strategy.

2. The method according to claim 1, characterized in that Determining a target adjustment strategy that matches the operation behavior data from a plurality of adjustment strategies corresponding to the target feedback type includes: Obtaining object attribute information of an operation object corresponding to the client, and context information associated with the operation behavior data, wherein the context information at least includes device attribute information of the client and environment information corresponding to the moment when the operation behavior data occurs; Performing feature extraction on the operation behavior data, the object attribute information, and the context information to obtain state features corresponding to the operation behavior data; The target adjustment strategy that matches the operation behavior data is determined by using a strategy network based on the state characteristics from a plurality of adjustment strategies corresponding to the target feedback type.

3. The method according to claim 2, characterized in that The performing feature extraction on the operation behavior data, the object attribute information, and the context information to obtain state features corresponding to the operation behavior data includes: Performing feature extraction on the operation behavior data using a feature extraction model to obtain behavior features, wherein the feature extraction model at least includes an attention module; Extracting features from the object attribute information using an embedding layer to obtain object attribute features; Extracting features from the context information using the embedding layer to obtain context features; Splicing the behavior feature, the object attribute feature, and the context feature to obtain a spliced ​​feature; Performing matrix transformation on the splicing features to obtain the state features.

4. The method according to claim 2, characterized in that The utilizing strategy network to determine the target adjustment strategy that matches the operation behavior data from a plurality of adjustment strategies corresponding to the target feedback type based on the state feature includes: Determining output probabilities of the plurality of adjustment strategies based on the state characteristics using the strategy network, wherein the output probabilities are used to characterize the degree of matching between the plurality of adjustment strategies and the operation behavior data; An adjustment strategy corresponding to the maximum output probability among the multiple adjustment strategies is obtained to obtain the target adjustment strategy.

5. The method according to any one of claims 2 to 4, characterized in that The method further comprises: In a case where a first recommendation sample is sent to the client through the recommendation system, a first training sample is obtained, wherein the first training sample includes: a first operation behavior sample for the first recommendation sample, the object attribute information, and a first context sample associated with the first operation behavior sample; Determining a training adjustment strategy based on the first training sample using the strategy network; Adjusting the recommendation system based on the training adjustment strategy to obtain an adjusted recommendation system; When a second recommendation sample is sent through the adjusted recommendation system, obtaining a second training sample and a target reward value of the second training sample, wherein the second training sample includes: a second operation behavior sample for the second recommendation sample, the object attribute information, and a second context sample associated with the second operation behavior sample; Using the evaluation network to evaluate the first training sample, obtaining a first evaluation result corresponding to the first training sample, and using the evaluation network to evaluate the second training sample, obtaining a second evaluation result corresponding to the second training sample; Based on the first evaluation result, the second evaluation result, and the target reward value, network parameters of the policy network and the evaluation network are adjusted.

6. The method according to claim 5, characterized in that The obtaining of the target reward value of the second training sample includes: determining a first feedback reward value for the second recommended sample based on target operation behavior data associated with the second recommended sample in the second operation behavior sample; determining a second feedback reward value for the second recommended sample based on the next operation behavior data following the target operation behavior data in the second operation behavior sample; Determining a first estimated reward value for the first recommended sample and a second estimated reward value for the second recommended sample; The first feedback reward value, the second feedback reward value, the first estimated reward value, and the second estimated reward value are aggregated to obtain the target reward value.

7. The method according to claim 6, characterized in that The summing up the first feedback reward value, the second feedback reward value, the first estimated reward value, and the second estimated reward value to obtain the target reward value includes: Obtaining a difference between the second estimated reward value and the first estimated reward value to obtain a reward value error; A weighted sum process is performed on the first feedback reward value, the second feedback reward value, and the reward value error to obtain the target reward value.

8. The method according to claim 5, characterized in that The adjusting the network parameters of the policy network and the evaluation network based on the first evaluation result, the second evaluation result, and the target reward value includes: determining an evaluation error based on the first evaluation result, the second evaluation result, and the target reward value; updating the network parameters of the evaluation network based on the evaluation error; Based on the evaluation error and the training adjustment strategy, the network parameters of the policy network are updated.

9. The method according to any one of claims 2 to 4, characterized in that The method further comprises: In the case of sending historical recommendation data to multiple clients through the recommendation system, obtaining historical feedback data of the multiple clients for the historical recommendation data, wherein the historical feedback data includes at least one of the following: historical feedback results and historical operation data; Analyze historical feedback data sent by at least one of the multiple clients to determine the multiple feedback types; Generate the at least one adjustment strategy corresponding to the feedback type.

10. The method according to claim 1, characterized in that The method further comprises: Detecting the operation behavior data to determine whether the operation behavior data meets a preset adjustment condition; In the case where it is determined that the operation behavior data satisfies the preset adjustment condition, determining the target adjustment strategy corresponding to the recommendation system based on the operation behavior data; If it is determined that the operation behavior data does not meet the preset adjustment condition, adjusting the recommendation system is prohibited.

11. A method for adjusting a video recommendation system, characterized in that: include: When a recommended video is sent to a client through a video recommendation system, obtaining viewing behavior data of the client for the recommended video; determining whether there is a target viewing problem type among a plurality of viewing problem types that matches the viewing behavior data; In a case where the target viewing problem type exists among the multiple viewing problem types, determining a target adjustment strategy that matches the viewing behavior data from at least one adjustment strategy corresponding to the target viewing problem type; The video recommendation system is adjusted based on the target adjustment strategy.

12. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 11 when running.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 11.

14. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.