Content recommendation model training method and related device
Through Markov decision-making process and deep learning technology, the weight parameter search is optimized, and the problem of time-consuming random search algorithm is solved, which improves the efficiency and accuracy of the content recommendation model.
Patent Information
- Application Number
- CN202410017929.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the random search algorithm takes a long time to calculate the evaluation value of the multimedia content recommendation sequence, and the calculation is large, making it difficult to ensure the efficiency and effect of content recommendation.
Using the content recommendation model, through Markov decision-making process and deep learning technology, the first network is used to generate actions and evaluate action states through the second network, optimize the search process of weight parameters, and improve the efficiency and accuracy of the recommended model.
By optimizing the search process of weight parameters, the efficiency and accuracy of multimedia content recommendations are improved and the recommendation effect is improved.
Smart Images

Figure CN120256944A_ABST
Abstract
Description
Background Art
[0002] In a content recommendation scenario, the recommended order evaluation value of a multimedia content is determined by multiple value scores and the weight of each score. Among them, the multimedia content includes, but is not limited to, one or more media forms such as video, text, image, audio, etc. The multiple value scores include, but are not limited to, click score (pctr), duration score (preadtime), timeliness score (ptimebonus), and each value score corresponds to at least one score weight.
[0003] In the related art, the value of each score weight is determined by the Random Search algorithm. In the Random Search algorithm, all possible values of each score weight are traversed to obtain each candidate value combination, and each value combination is evaluated respectively through a set index. Through the evaluation results, the target value combination is selected from each candidate value combination, and then the target value combination is used to calculate the recommended order evaluation value.
[0004] However, the random search method evaluates each candidate value combination by traversing and exhausting, and the scale of the candidate value combination will increase exponentially with the increase of the number of weight parameters. Obviously, the random search method takes a long time to calculate and has a large amount of calculation, which is difficult to ensure the content recommendation efficiency and has a poor recommendation effect within a limited time. Summary of the Invention
[0005] The embodiments of the present application provide a training method for a content recommendation model and related devices to improve the recommendation efficiency and recommendation effect of multimedia content.
[0006] In a first aspect, the embodiments of the present application provide a method for training a content recommendation model. The content recommendation model includes: a first network for action generation and a second network for action state evaluation. The method includes:
[0007] For each group of training data, the content recommendation model is used to perform at least one state transition respectively, and the parameters are adjusted based on the obtained model losses. Each state transition includes:
[0008] Based on the current sample state of a group of training data, the first network is used to obtain an execution evaluation value set for each weight parameter. Each execution evaluation value set contains: the execution evaluation value of each candidate action corresponding to the corresponding weight parameter, and a candidate action represents a type of numerical adjustment operation.
[0009] Based on each execution evaluation value set respectively, the corresponding target action is screened from each candidate action, and based on the obtained target actions, the target value corresponding to each weight parameter is obtained.
[0010] Based on the obtained target values, rearrange each sample list in the set of training data respectively, and based on each rearranged list, obtain the next sample state;
[0011] Based on each set of execution evaluation values, combined with the current sample state, use the second network to obtain a predicted evaluation value, and based on the next sample state, use the second network to obtain a true evaluation value, and based on the true evaluation value and the predicted evaluation value, obtain a model loss.
[0012] In a second aspect, an embodiment of the present application provides a content recommendation model training device. The content recommendation model includes: a first network for action generation and a second network for action state evaluation. The device includes:
[0013] An evaluation unit, configured to perform at least one state transition on each set of training data by using the content recommendation model respectively; in each state transition, the evaluation unit is specifically configured to perform the following operations:
[0014] For each set of training data, based on the current sample state of a set of training data, use the first network to obtain a set of execution evaluation values for each weight parameter. Each set of execution evaluation values includes: the execution evaluation values of each candidate action corresponding to the corresponding weight parameter, and one candidate action represents a type of numerical adjustment operation;
[0015] Based on each set of execution evaluation values respectively, screen the corresponding target actions from the candidate actions, and based on the obtained target actions, obtain the target values corresponding to each weight parameter;
[0016] Based on the obtained target values, rearrange each sample list in the set of training data respectively, and based on each rearranged list, obtain the next sample state;
[0017] Based on each set of execution evaluation values, combined with the current sample state, use the second network to obtain a predicted evaluation value, and based on the next sample state, use the second network to obtain a true evaluation value, and based on the true evaluation value and the predicted evaluation value, obtain a model loss.
[0018] A parameter tuning unit, configured to perform parameter tuning based on the obtained model losses.
[0019] As a possible implementation, the set of training data includes: the object features of the sample object, each sample list and the content features of each sample content included therein;
[0020] When obtaining the next sample state based on each rearranged list, the evaluation unit is specifically configured to:
[0021] For each of the permutation lists, perform the following operations respectively: Based on the permutation order of the sample contents in a permutation list, screen out at least one reference content that meets the set content screening conditions from the sample contents in a permutation list;
[0022] Based on the content features of the obtained reference contents and in combination with the object features, obtain the next sample state of the set of training data.
[0023] As a possible implementation manner, when obtaining the execution evaluation value sets of the respective weight parameters by using the first network based on the current sample state of a set of training data, the evaluation unit specifically is used for:
[0024] If the current state transition is the first state transition, based on the content features included in the set of training data and the object features, obtain the current sample state of the set of training data, and input the object features into the first network to obtain the execution evaluation value sets of the respective weight parameters;
[0025] If the current state transition does not belong to the first state transition, directly input the current sample state of the set of training data into the first network to obtain the execution evaluation value sets of the respective weight parameters.
[0026] As a possible implementation manner, the set of training data further includes: the historical interaction information of the sample contents included in each of the sample lists;
[0027] When obtaining the next sample state of the set of training data based on the content features of the obtained reference contents and in combination with the object features, the evaluation unit specifically is used for:
[0028] Based on the historical interaction information in the set of training data and in combination with the recommendation order evaluation values respectively corresponding to the permutation lists, obtain the task evaluation values respectively corresponding to at least one target task; wherein, the recommendation order evaluation values are determined based on the respective target values;
[0029] Based on the obtained at least one task evaluation value, in combination with the object features and the content features of the obtained reference contents, obtain the next sample state of the set of training data.
[0030] As a possible implementation manner, when respectively rearranging each of the sample lists in the set of training data based on the obtained respective target values, the evaluation unit specifically is used for:
[0031] For each of the sample lists, respectively based on the set recommendation order evaluation method and in combination with the respective target values, obtain the recommendation order evaluation values of the sample contents in the corresponding sample list;
[0032] Based on the obtained evaluation values of each recommendation order, rearrange each of the sample lists respectively.
[0033] As a possible implementation manner, after rearranging each of the sample lists in the set of training data respectively based on the obtained target values, before obtaining the true evaluation value by using the second network based on the next sample state, the evaluation unit is further configured to:
[0034] Based on each rearranged sample list, test each task to obtain feedback information, where the feedback information is used to characterize: after transferring from the current sample state to the next sample state through the respective target actions, the task evaluation values of each task;
[0035] When obtaining the true evaluation value by using the second network based on the next sample state, the evaluation unit is specifically configured to:
[0036] Based on the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state, and based on the predicted evaluation value corresponding to the next sample state and the feedback information, obtain the true evaluation value.
[0037] As a possible implementation manner, when obtaining the predicted evaluation value corresponding to the next sample state by using the second network based on the next sample state, the evaluation unit is specifically configured to:
[0038] Based on the next sample state, use the first network to obtain the set of execution evaluation values corresponding to the next sample state for each of the weight parameters;
[0039] Based on the obtained set of execution evaluation values corresponding to the next sample state, in combination with the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state.
[0040] As a possible implementation manner, when obtaining the model loss based on the true evaluation value and the predicted evaluation value, the evaluation unit is specifically configured to:
[0041] Based on the predicted evaluation value, in combination with the gradients of the model parameters with respect to the respective target actions, obtain the first network loss; and, based on the true evaluation value and the predicted evaluation value, obtain the second network loss;
[0042] Based on the second network loss and the first network loss, obtain the model loss.
[0043] As a possible implementation manner, when obtaining the target value corresponding to each weight parameter based on the obtained target actions, the evaluation unit is specifically configured to:
[0044] For each of the weight parameters, perform the following operations respectively:
[0045] Based on the set update step, use the target action corresponding to a weight parameter to update the current value corresponding to the weight parameter, and obtain the target value corresponding to the weight parameter.
[0046] As a possible implementation manner, each set of training data includes the object features of the corresponding sample object. After adjusting the parameters based on the obtained model losses, the evaluation unit is further configured to:
[0047] If the content recommendation model meets the set model convergence condition, output the target values corresponding to each set of training data, and associatively store the object features included in each set of training data and their corresponding target values in the target storage area;
[0048] When the target object features of the target object are received, obtain the target values associated with the target object features from the target storage area according to the target object features.
[0049] As a possible implementation manner, the evaluation unit is further configured to:
[0050] Obtain the value score sets of each candidate recommendation information;
[0051] Based on the target values associated with the target object features, fuse the value scores in the value score sets of each candidate recommendation information respectively, and obtain the recommendation order evaluation values corresponding to each candidate recommendation information;
[0052] Based on the obtained recommendation order evaluation values, select the target recommendation information from the candidate recommendation information for recommendation.
[0053] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0054] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method in any of the above aspects.
[0055] Fifth aspect, an embodiment of the present application provides a computer program product, the program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program, so that the electronic device executes the steps of the method in any of the above aspects.
[0056] In an embodiment of the present application, in each round of iteration process, using the Markov decision process, at least one state transition is performed for each group of training respectively. By regarding the search problem of parameter values as a decision process, the search efficiency of each weight parameter in the multi-objective formula and the accuracy of content ranking can be improved, thereby improving the recommendation effect; in each state transition, a first network is used to generate an action, and a second network is used to generate an action state evaluation value. The action state evaluation value can measure the goodness or badness of adjusting the parameters using the target value in the current state, so that the model loss obtained subsequently based on the action state evaluation value can effectively adjust the model parameters, thereby improving the model performance.
[0057] Other features and advantages of the present application will be described in the subsequent description, and part of them will become obvious from the description, or will be understood by implementing the present application. The purpose and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0059] Figure 1 It is a schematic diagram of an application scenario provided in an embodiment of the present application;
[0060] Figure 2 It is a schematic diagram of the architecture of a content recommendation model provided in an embodiment of the present application;
[0061] Figure 3 It is a schematic diagram of the flow of a method for training a content recommendation model provided in an embodiment of the present application;
[0062] Figure 4A It is a schematic diagram of a subscription message provided in an embodiment of the present application;
[0063] Figure 4B It is a schematic diagram of each sample list provided in an embodiment of the present application;
[0064] Figure 5A It is a schematic diagram of the logic of the process for obtaining a current sample state provided in an embodiment of the present application;
[0065] Figure 5B It is a logical schematic diagram of another process for obtaining the current sample state provided in the embodiments of the present application;
[0066] Figure 6 It is a logical schematic diagram of a process for determining a target operation provided in the embodiments of the present application;
[0067] Figure 7 It is a logical schematic diagram of a process for determining a target value provided in the embodiments of the present application;
[0068] Figure 8 It is a logical schematic diagram of a process for determining a target value provided in the embodiments of the present application;
[0069] Figure 9 It is a logical schematic diagram of a process for determining the next sample state provided in the embodiments of the present application;
[0070] Figure 10 It is a logical schematic diagram of model input and model output provided in the embodiments of the present application;
[0071] Figure 11A It is a logical schematic diagram of a process for constructing a quadruple provided in the embodiments of the present application;
[0072] Figure 11B It is a logical schematic diagram of a state transition process provided in the embodiments of the present application;
[0073] Figure 12 It is a structural schematic diagram of a training device for a content recommendation model provided in the embodiments of the present application;
[0074] Figure 13 It is a structural schematic diagram of an electronic device provided in the embodiments of the present application. Detailed implementation manners
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.
[0076] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.
[0077] It can be understood that in the specific implementation of the present application, for data related to object features, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0078] Next, some terms involved in the embodiments of the present application will be described.
[0079] Session list: In the recommended scenario of subscription accounts, it is the longest list of exposed messages generated after a sample object enters the subscription account box once. For example, if after the sample object enters the subscription account box, messages ranked from the first to the fifth are exposed to the sample object, and then the sample object exits the subscription account box, and the messages ranked sixth and later are not exposed, then the length of the session list is five.
[0080] Send list: The list of messages actually sent to the sample object. For example, for a sample object, 10 messages are sent at a time. Although only the first five messages are exposed, the length of the send list is 10.
[0081] Area Under Curve (AUC) of the ROC curve: AUC is an index used to evaluate the performance of a model. The higher the value of AUC, the better the performance of the model. Exemplarily, AUC can be used, but is not limited to, evaluating the performance of a Click Through Rate (CTR) prediction model.
[0082] Group AUC (GAUC): GAUC is an improved version of AUC.
[0083] Offline environment: The offline environment means that the data set used for analysis and computational simulation is the exposed samples that have been collected, and the target prediction scores and label information such as whether to click and reading duration of each sample are known. In contrast, the online environment is the information push and consumption process where the exposure situation cannot be predicted and the label information is determined by real-time data.
[0084] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0086] Computer Vision (CV) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, monitoring, and measurement, and further performing graphic processing to make the computer process images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as Swin Transformer, Vision Transformer (ViT), Vision Model based on Mixture of Experts (V-MoE), and Masked Autoencoders (MAE) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, Optical Character Recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0087] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future. The large model technology has brought about a revolution in the development of speech technology. Pretrained models that follow the Transformer architecture, such as WavLM and UniSpeech, have powerful generalization and versatility and can excellently complete speech processing tasks in various directions.
[0088] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves disciplines such as computer science and mathematics. The important technology for model training in the field of artificial intelligence, the pretrained model, has evolved from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0089] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pretrained model is the latest development result of deep learning, integrating the above technologies.
[0090] Autonomous driving technology refers to the vehicle's ability to drive itself without the operation of a driver. It usually includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autonomous driving includes multiple development paths such as single-vehicle intelligence, vehicle-road cooperation, and networked cloud control. Autonomous driving technology has a wide range of application prospects. Currently, in addition to the fields of logistics, public transportation, taxis, and intelligent transportation, it will be further developed in the future.
[0091] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0092] The solution provided in the embodiments of this application relates to the machine learning technology of artificial intelligence, specifically to reinforcement learning technology. Reinforcement learning, also known as re-inforcement learning, evaluation learning or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent maximizes or achieves a specific goal through learning strategies during the interaction with the environment. The common model of reinforcement learning is the standard Markov decision process. Deep learning models can be used in reinforcement learning to form deep reinforcement learning.
[0093] Deep Reinforcement Learning is a subfield of reinforcement learning that uses deep learning techniques (such as deep neural networks) to represent and learn the generation of actions. Deep learning techniques enable agents to handle more complex and higher-dimensional state spaces and action spaces, thus solving more complex problems.
[0094] Markov Decision Process (MDP) is a mathematical model used to describe decision-making problems, which includes elements such as state, action, transition probability, reward, discount factor, cumulative reward, etc. MDP assumes that the system satisfies the Markov property, that is, the next state depends only on the current state and action, and is independent of previous states and actions.
[0095] MDP contains a set of interacting objects, namely the agent and the environment. The agent is the proxy for machine learning in MDP, which can perceive the state of the external environment to make decisions, act on the environment and adjust decisions through the feedback of the environment. The environment is the set of all things outside the agent in the MDP model, and its state will be affected by the agent's actions and the above changes can be fully or partially perceived by the agent. The environment may feedback corresponding rewards to the agent after each decision.
[0096] A state is a description of the environment. After an agent takes a certain action, the state will change in a certain way, and the evolution has the Markov property. The set of all states in an MDP is called the state space. The state space can be discrete or continuous. In an MDP, the state space is denoted by S, S = {s1, s2, ……, s r}, where s1, s2, ……, s r represent different states.
[0097] An action is a description of the agent's behavior and is the result of the agent's decision-making. The set of all possible actions in an MDP is the action space. The action space can be discrete or continuous. In an MDP, the action space is denoted by A, A = {a1, a2, ……, a r}, where a1, a2, ……, a r represent different states.
[0098] The transition probability describes the probability of transitioning from the current state to the next state given the current state and action. In an MDP, the transition probability is usually denoted by P, P = P(s'|s, a), where s and s' represent the current state and the next state respectively, and a represents the current action.
[0099] The reward (which can also be called the incentive) is the feedback from the environment to the agent when the agent transfers from one state to another by performing a certain action. In an MDP, the reward is usually denoted by R, R = R(s, a, s'), where s and s' represent the current state and the next state respectively, and a represents the current action. In the embodiments of this application, the reward can also be called feedback information.
[0100] The cumulative reward is the accumulation of rewards over time steps. The cumulative reward is usually calculated using the Action-State Value Function (Q function). The Q function is a function used in reinforcement learning to evaluate the expected reward for taking a certain action in a given state. The Q function is based on the state-action pair and is used to measure how good or bad it is to execute a certain action in a certain state. Based on the value of the Q function, the intelligent agent can select the optimal action to maximize the cumulative reward. The Q function is usually expressed as Q(s, a), where s represents the state and a represents the action. The value of the Q function represents the expected future cumulative reward after taking action a in state s. In reinforcement learning, the goal of the intelligent agent is to find a policy such that taking the corresponding action in each state can maximize the value of the Q function.
[0101] The discount factor is a value between 0 and 1, which is used to adjust the importance of future rewards. The closer the discount factor is to 1, the more the agent focuses on long-term rewards; the closer the discount factor is to 0, the more the agent focuses on short-term rewards.
[0102] Many algorithms in reinforcement learning (such as Q-learning and Deep Q-Networks) are optimized based on the Q function. The core idea of these algorithms is to gradually update and optimize the Q function through interactions with the environment and observed rewards.
[0103] In the embodiments of this application, the search problem of the value of each weight parameter in the multi-objective formula is regarded as a decision-making process. Through appropriate definitions of (state, action, policy, reward), a specific MDP mathematical model is obtained. Then, deep reinforcement learning technology is used to train and solve this decision-making problem.
[0104] Among them, the multi-objective formula refers to a formula used to calculate the evaluation value of the recommendation order of multimedia content. The multi-objective formula usually consists of multiple value scores and the weights of each score (i.e., weight parameters). That is to say, the evaluation value of the recommendation order of a multimedia content is determined by multiple value scores and the weights of each score. The multiple value scores include, but are not limited to: click score (pctr), duration score (preadtime), timeliness score (ptimebonus), and each value score corresponds to at least one score weight.
[0105] Exemplarily, score is used to represent the evaluation value of the recommendation order, and score is calculated using the following multi-objective formula: score = w1 * pctr + w2 * preadtime + w3 * ptimebonus, where w1 represents the score weight of the click score, w2 represents the score weight of the duration score, and w3 represents the score weight of the timeliness score.
[0106] In the related art, the values of the weights of each score are determined by a random search algorithm. In the random search algorithm, all possible values of the weights of each score are traversed to obtain each candidate value combination, and each value combination is evaluated respectively through a set index. Through the evaluation results, the target value combination is selected from each candidate value combination, and then the target value combination is used to calculate the evaluation value of the recommendation order.
[0107] However, the random search method evaluates each candidate value combination by traversing and exhausting, and the scale of the candidate value combination will increase exponentially with the increase in the number of weight parameters; obviously, the random search method takes a long time to calculate and has a large amount of calculation, which is difficult to ensure the content recommendation efficiency and has a poor recommendation effect within a limited time.
[0108] In the embodiments of the present application, in each round of iteration, using the Markov decision process, at least one state transition is performed for each group of training. By treating the search problem of parameter values as a decision process, the search efficiency of each weight parameter in the multi-objective formula and the accuracy of content ranking can be improved, thereby improving the recommendation effect. In each state transition, the first network is used to generate an action, and the second network is used to generate an action state evaluation value. The action state evaluation value can measure the goodness or badness of adjusting the parameters using the target value in the current state, so that the model loss obtained subsequently based on the action state evaluation value can effectively adjust the model parameters, thereby improving the model performance.
[0109] Referring to Figure 1 As shown, it is a schematic diagram of an application scenario provided in the embodiments of the present application. This application scenario includes a terminal device 110 and a server 120. The number of terminal devices 110 can be one or more. The number of servers 120 can also be one or more. The present application does not specifically limit the number of terminal devices 110 and servers 120.
[0110] In the embodiments of the present application, the terminal device 110 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. The terminal device 110 supports a client for content recommendation, and the server 120 is the background server corresponding to the client.
[0111] The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0112] The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this here.
[0113] In some embodiments, for each set of training data, the server 120 performs at least one state transition using the content recommendation model respectively, and adjusts the parameters based on the obtained model losses; wherein, each state transition includes: based on the current sample state of a set of training data, using the first network to obtain the execution evaluation value sets of each weight parameter, and each execution evaluation value set includes: the execution evaluation values of each candidate action corresponding to the corresponding weight parameter, and one candidate action represents a type of numerical adjustment operation; respectively based on each execution evaluation value set, screening the corresponding target actions from each candidate action, and based on the obtained target actions, obtaining the target values corresponding to each weight parameter; based on the obtained target values, rearranging each sample list in a set of training data respectively, and based on each rearranged list, obtaining the next sample state; based on each execution evaluation value set, combining the current sample state, using the second network to obtain the prediction evaluation value, and based on the next sample state, using the second network to obtain the true evaluation value, and based on the true evaluation value and the prediction evaluation value, obtaining the model loss.
[0114] It should be noted that the method flow provided in each embodiment of the present application can be executed by the server or the terminal device, or jointly executed by the server and the terminal device. Here, the server is mainly used as an example for introduction. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0115] In the embodiments of the present application, the training process of the content recommendation model is a process of performing multiple loop iterations using the training data, which mainly includes a model design stage, a data preparation stage, and an iterative training stage, which will be introduced separately below.
[0116] I. Model Design Stage
[0117] Refer to Figure 2 As shown, it is a schematic diagram of the architecture of a content recommendation model provided in the embodiments of the present application. The content recommendation model includes: a first network for action generation and a second network for action state evaluation. Among them, the first network can also be called an action generation network (Actor Network), and the second network can also be called an action state evaluation network (Critic Network).
[0118] The action generation network is used to obtain the execution evaluation value sets of the respective weight parameters corresponding to a certain sample state of a set of training data, and the obtained execution evaluation value sets are used to decide the target actions of the respective weight parameters in the sample state. For example, the action generation network is used to obtain the execution evaluation value sets of the respective weight parameters corresponding to the current sample state of a set of training data, and to obtain the execution evaluation value sets of the respective weight parameters corresponding to the next sample state of a set of training data.
[0119] The action state evaluation network is used to obtain the corresponding prediction evaluation values according to a certain sample state and the execution evaluation value sets of the respective weight parameters corresponding thereto. For example, the action state evaluation network can be used to obtain the corresponding prediction evaluation values according to the current sample state of a set of training data and the execution evaluation value sets of the respective weight parameters corresponding thereto, and to obtain the corresponding prediction evaluation values according to the next sample state of a set of training data and the execution evaluation value sets of the respective weight parameters corresponding thereto.
[0120] In the embodiments of the present application, the agent can handle more complex and higher-dimensional state spaces and action spaces, thereby solving more complex problems. The Actor Network and the Critic Network can be implemented using deep learning techniques. For example, the Actor Network and the Critic Network are implemented using deep neural networks. The reinforcement learning algorithm can utilize the memory and representation capabilities of the neural network model, and also pay more attention to long-term rewards during the search process, thereby improving the search efficiency of the multi-objective formula and the accuracy of message sorting and the recommendation effect.
[0121] As a possible implementation, in the Actor Network, three layers of Multi-Layer Perceptron (MLP) are connected to the softmax function. The input dimension of the Actor Network is [BatchSize, k], and the output dimension is [BatchSize, p, b], where BatchSize is the amount of data for one forward propagation, k is the dimension of the sample state, p is the number of weight parameters, and b is the number of candidate actions. Exemplarily, the input dimension of the Actor Network is [BatchSize, 73], and the output dimension is [BatchSize, 5, 3]. The value of BatchSize can be 16, that is, the amount of data for one forward propagation is 16, the feature dimension of the object features is 32, there are 5 weight parameters, and 3 candidate actions.
[0122] It should be noted that in the embodiments of the present application, the recommended order evaluation value is used to evaluate the recommended order of a multimedia content (during the training process, the multimedia content can also be referred to as the sample content). The recommended order evaluation value can be represented by a numerical value or a grade, and there is no limitation in this regard. In this article, only the numerical form is used as an example for illustration.
[0123] As a possible implementation manner, the recommended order evaluation value of a sample content is combined in a multiplicative form by value scores such as click score, duration score, timeliness score, and quality score. Among them, the click score is used to characterize the probability that the sample content is clicked, and the duration score can be obtained through a click-through rate prediction model; the duration score is used to characterize the probability that the sample content is read for a long time; the timeliness score is used to characterize the timeliness of the sample content, and the timeliness score can be calculated from information such as the mass sending time, exposure time, and message category of the sample content; the quality score is used to characterize the quality score of the sample content. It should be noted that in the actual application process, the click score, duration score, timeliness score, and quality score can also be calculated using rules.
[0124] Of course, calculating the recommended order evaluation value using the four evaluation values of click score, duration score, timeliness score, and quality score is only an example. In the actual application process, the recommended order evaluation value can also be composed of any one or more combinations of the above four evaluation values, but it is not limited to this.
[0125] In the embodiments of the present application, the calculation formula of the recommended order evaluation value can also be referred to as a multi-objective fusion formula. Exemplarily, the multi-objective fusion formula is represented by formula (1):
[0126]
[0127] Among them, fusion represents the recommended order evaluation value of content n, score0 represents the click score, score1 represents the duration score, score2 represents the timeliness score, score3 represents the quality score, and w1, w2, w3, w4, w5 represent weight parameters. It should be noted that in order to simplify the description of the solution, in the embodiments of the present application, only the example of the multi-objective fusion formula serving the optimization of the click task is used for illustration, that is, the introduction of the duration score, quality score, and timeliness score is used as an auxiliary task for the click task, and the ultimate goal is to optimize the click task. In the actual application process, the optimization task is not limited to the click task.
[0128] In the following, only the calculation of the recommended order evaluation value using formula (1) is used as an example for illustration, that is, calculating the respective values of w1, w2, w3, w4, w5.
[0129] In the embodiments of the present application, a complete action in the MDP can be defined as obtaining the target values of the weight parameters in the multi-objective fusion formula based on the current sample state, the output result of the Actor Network, and the discretized policy π. In this article, only the discretized policy π with a transition probability of 1 is taken as an example for illustration.
[0130] The output result of the Actor Network will be combined with the corresponding sample state and passed to the Critic Network. The output result of the Actor Network is also used to re-rank each issued sample list.
[0131] As a possible implementation, the structural design of the Critic Network is similar to that of the Actor Network. The Critic Network consists of three layers of MLP but does not connect to the softmax function. The Critic Network can be regarded as the implementation process of Q(s, a), and the output result of the Critic Network is the value of Q(s, a). Since the input of the Critic Network includes the state and the action, taking the dimension of s as 73 and the output result of the Actor Network as a 5×3 matrix as an example, the input dimension of the Critic Network is [BatchSize, 73 + 15 = 88], and the output dimension is [BatchSize, 1]. Here, S represents the dimension of the sample state, and the 1D output represents the value of the Q function (i.e., the predicted evaluation value).
[0132] The data processing process executed by the above model will be introduced in detail in the subsequent process, so it will not be elaborated here too much.
[0133] II. Data Preparation Phase
[0134] Data collection is of utmost importance in machine learning and can be said to be the most important link. The data preparation phase of the embodiments of the present application is mainly used to prepare several groups of training data. Each group of training data includes the following information:
[0135] (1) The object features of the sample object.
[0136] (2) Each sample list issued for the sample object, and each sample list contains each sample content. It should be noted that the number of sample contents included in each sample list can be the same or different, that is, the length of the sample list can be non-fixed.
[0137] (3) The content features of each sample content. Among them, the content features include but are not limited to one or more of click score, duration score, quality score, and timeliness score.
[0138] (4) Historical interaction information of each sample content. The historical interaction information includes but is not limited to click tags (used to represent whether the sample object clicks on the sample content), duration tags (used to represent whether the sample object browses the sample content, etc.).
[0139] In the scenario of public account message recommendation, the content type of the sample content included in each group of training data is a message; in the video recommendation scenario, the content type of the sample content included in each group of training data is a video.
[0140] That is to say, in the embodiments of the present application, for a sample object, each sample list sent to the sample object is obtained, and a group of training data is constructed according to the sent sample lists.
[0141] As a possible implementation manner, in the embodiments of the present application, when the object feature of a sample object is obtained, each sent list corresponding to the sample object is sampled from the sent log, and each sent list is used as each sample list, and a group of training data is constructed based on the object feature of the sample object, each sample list, the content feature of each sample content included in each sample list, and the historical interaction information of each sample content.
[0142] It should be noted that in the embodiments of the present application, the sample object can be a certain object or a certain type of object.
[0143] III. Iterative training stage
[0144] In the embodiments of the present application, after the training data is prepared, the constructed model can be trained using these training data.
[0145] In one implementation manner, the parameters and data required for the content recommendation model can be set according to the structure of the above model, including the parameters to be trained. After setting the hyperparameters such as batch size (batchsize), number of epochs (epoch), and learning rate (learning rate) respectively, the training starts.
[0146] For example, set the batchsize of the content recommendation model to 128, the epoch to 1000, and the learning rate to 0.0001, that is, perform iterative training 1000 times, and 128 quadruples will be learned in each iteration. Of course, the training parameters here are only a possible example, and can be adjusted according to requirements in actual situations.
[0147] In the embodiments of the present application, the learning of each batch can be regarded as a construction process of a quadruple (s, a, r, s'), where s represents the current sample state of a set of training data, a represents the set of execution evaluation values of each weight parameter obtained for the current sample state of a set of training data, r represents the reward, and s' represents the next sample state of a set of training data. For the specific construction process of the quadruple (s, a, r, s'), please refer to the following text (such as Figure 8 ).
[0148] In some embodiments, a set of training data can undergo state transitions multiple times (for example, 15 times). During each state transition process, a corresponding quadruple will be constructed. Taking 15 state transitions as an example, after 15 state transitions, a trajectory sample with 15 steps can also be formed, and the trajectory sample can be stored in the database. In this way, after performing multiple state transitions on each set of training data respectively, multiple trajectory samples can be obtained, and the construction process of the multiple trajectory samples can be used to guide the training of the model. It should be noted that in the embodiments of the present application, the number of state transitions of each set of training data can be the same or different, and there is no limitation on this.
[0149] Refer to Figure 3 As shown, it is a schematic flowchart of the content recommendation model training method provided by the embodiments of the present application. During the iterative training process, all training samples are divided into specified batches, and training is performed based on the training samples of each sub-batch. Since the steps performed during the training of each batch in each iteration process are similar, the training process of each set of training data in one batch is taken as an example for illustration here.
[0150] S301. Obtain the i-th set of training data in each set of training data.
[0151] Still taking the subscription number message recommendation scenario as an example, refer to Figure 4A As shown, it is the message interface of the "subscription number message" public account in the social software. The message cards pushed by this public account include those subscribed by the user himself / herself and those actively recommended by the social software platform. The social software platform can obtain the recommendation order evaluation value corresponding to each message based on one or more value scores and the values of each weight parameter, and perform sorting and recommendation on each message to be recommended based on the recommendation order evaluation value.
[0152] Refer to Figure 4BAs shown, assume that the i-th group of training data contains 10,000 issued lists. Use L to represent the issued list. Only take L1, L2, and L3 in the 10,000 issued lists as examples. L1 contains 10 messages, L2 contains 20 messages, and L3 contains 10 messages. The i-th group of training data also contains the content features corresponding to each message included in each issued list. The content features include click score, duration score, quality score, timeliness score, and the click label corresponding to each message. In addition, the i-th group of training data may also contain the object features corresponding to each message.
[0153] S302. Based on the current sample state of the i-th group of training data, use the first network to obtain the execution evaluation value set of each weight parameter. Each execution evaluation value set contains: the execution evaluation value of each candidate action corresponding to the corresponding weight parameter. A candidate action represents a type of numerical adjustment operation.
[0154] In the embodiments of the present application, to better implement personalized recommendation, the current sample state of the i-th group of training data can be described by combining static features and dynamic features. Use s1 to represent the current static feature, use s2 to represent the current dynamic feature, and use s to represent the current sample state, s = [s1, s2]. Similarly, use s1' to represent the next static feature, use s2' to represent the next dynamic feature, and use s' to represent the next sample state.
[0155] The dynamic feature is used to represent the multi-dimensional aggregation information of each sample list corresponding to the current multi-objective formula (formed by the values corresponding to each weight parameter). Exemplarily, for the current sample state of the i-th group of training data, the dynamic feature is used to represent the multi-dimensional aggregation information of the sample lists in the i-th group of training data. For the next sample state of the i-th group of training data, the dynamic feature is used to represent the multi-dimensional aggregation information of each re-arranged list.
[0156] The multi-dimensional aggregation information can be determined by one or more of, but not limited to, click score, duration score, quality score, and timeliness score. In this article, only the example of determining multi-dimensional aggregation information based on click score, duration score, quality score, and timeliness score is used for illustration.
[0157] The static features are the object features of the sample object. The static features include, but are not limited to, one or more of the following information: the age, gender, interest preference category, activity status, etc. of the sample object. As a possible implementation, the static features are discrete features. Using discrete static features as input can achieve personalized actions (i.e., different object features correspond to different actions). In addition, since the input space of the Actor Network is enumerable, after the model is trained offline, the value combinations of each static feature and their final decision-making actions (i.e., the final multi-objective formula) can be stored in the offline database in the form of key-value pairs, which is convenient for subsequent use.
[0158] In the embodiments of the present application, the dynamic features can be obtained in but not limited to the following possible ways:
[0159] The first way:
[0160] For each sample list in each sample list set, based on the arrangement order of each sample content in each sample list, at least one reference content that meets the set content screening conditions is selected from each sample content, and then based on the content features of the obtained reference contents, the dynamic feature s2 of the i-th group of training data is obtained.
[0161] Among them, the arrangement order of each sample content in a sample list is determined according to the current recommended order evaluation value of each sample content. The current recommended order evaluation value of each sample content is determined according to the current value setting of each weight parameter. If the current is the first state transition, then, according to the initial value settings of each weight parameter, the recommended order evaluation value is calculated. If the current is not the first state transition, then, according to the target value of each weight parameter obtained in the previous state transition, the recommended order evaluation value is calculated.
[0162] For example, in the first iteration of rearrangement, the initial values of w1, w2, w3, w4, and w5 are all 0.5. According to each initial value, using formula (1), the recommended order evaluation value of each sample content is calculated, and then, according to the recommended order evaluation value of each sample content, the arrangement order of each sample content is obtained.
[0163] The reference content that meets the set content screening conditions can be the sample content located at the set position (for example, the first sample content in each sample content, or the first three sample contents in each sample content). The reference content that meets the set content screening conditions can also be the sample content that meets the set sample content extraction interval, such as extracting one sample content as the reference content every five sample contents. Of course, the set content screening conditions are not limited to this, and no limitation is made thereto.
[0164] Based on the content features of each obtained reference content, the content features of each reference content can be directly concatenated to obtain the dynamic features of the i-th group of training data, but it is not limited to this.
[0165] The dynamic features of the i-th group of training data can also include the task evaluation value of the target task. The target task can include, but is not limited to, one or more of the click task, duration task, quality task, and timeliness task. The task evaluation value of a target task can be the GAUC of the current fusion score (i.e., the recommended order evaluation value) and the corresponding task label (indicated by historical interaction information). Taking the click task as an example, the dynamic features can also include the task evaluation value of the click task, and the task evaluation value of the click task can be the GAUC of the current fusion score and the click label.
[0166] Suppose, refer to Figure 5A As shown, the i-th group of training data contains N sample lists. As an example, the content features (click score, duration score, quality score, timeliness score) of the first message in the first sample list are used as the first 1-4 dimensional features of the dynamic feature s2; the content features (click score, duration score, quality score, timeliness score) of the first message in the second sample list are used as the first 5-8 dimensional features of the dynamic feature; similarly, according to the content features (click score, duration score, quality score, timeliness score) of the first message in the third sample list,..., the N-th sample list, the dynamic feature s2 is constructed in turn. Obviously, the feature dimension of the dynamic feature s2 is 4*N dimensions. On this basis, if the task evaluation value of the click task (i.e., the click task offline metric) is further combined, the feature dimension of the dynamic feature s2 is 4*N+1 dimensions. Suppose the value of N is 10, then the feature dimension of the dynamic feature s2 is 4*10+1 = 41 dimensions.
[0167] The second method: For some of the sample lists in each sample list, based on the arrangement order of each sample content in each sample list, at least one reference content that meets the set content screening conditions is selected from each sample content, and then based on the content features of the obtained reference contents, the dynamic feature s2 of the i-th group of training data is obtained. The selection of some sample lists can be preset or randomly selected, and there is no limitation on this. Since the second method is similar to the first method, it will not be elaborated here.
[0168] The third method: Feature fusion is performed on each sample content in each sample list respectively, and based on the sample features corresponding to each sample list, the dynamic feature s2 is obtained. Specifically, for each sample list in each sample list, based on the arrangement order of each sample content in the sample list and combined with the content features of each sample content in the sample list, the list feature corresponding to the sample list is obtained; Feature fusion is performed on the obtained list features corresponding to each sample list respectively (including but not limited to methods such as calculating the mean), and the dynamic feature s2 is constructed.
[0169] Refer to Figure 5B As shown, the third method can also be understood as obtaining the first 1-4 dimensional features of the dynamic feature s2 according to the average value of the content features of each sample list located at the first sample content respectively, obtaining the first 5-8 dimensional features of the dynamic feature s2 according to the average value of the content features of each sample list located at the second sample content respectively, and similarly, obtaining the dynamic feature s2 according to the average value of the content features of each sample content located at the corresponding position of each sample list.
[0170] Taking L1 as an example, L1 contains 10 sample contents, and the arrangement order of the 10 sample contents is as follows: Content 1, Content 2,..., Content 10. The content features (click score, duration score, quality score, timeliness score) of Content 1 are used as the first 1-4 dimensional features of the list feature; The content features (click score, duration score, quality score, timeliness score) of Content 2 are used as the first 5-8 dimensional features of the list feature; Similarly, according to the content features (click score, duration score, quality score, timeliness score) corresponding to Content 3,..., Content 10 respectively, the list feature of L1 is constructed in sequence. Obviously, the feature dimension of the list feature of L1 is 4 * 10 = 40 dimensions. Similarly, the list features corresponding to each issued list are obtained, and the list features corresponding to each issued list are averaged to obtain a 40-dimensional fusion feature, and then combined with the task evaluation value of the click task to obtain a 41-dimensional dynamic feature s2.
[0171] As a possible implementation method, whether it is the current state or the next state, the static feature can be mapped to a dense feature through the Embedding layer, and the dense feature is concatenated with the dynamic feature to obtain the corresponding sample state. For example, the dynamic feature is a 41-dimensional feature, and the dense feature is a 32-dimensional feature. The dense feature is concatenated with the dynamic feature to obtain a 73-dimensional sample state.
[0172] In the embodiments of the present application, based on the current sample state of the i-th group of training data, the first network is used to obtain the execution evaluation value set of each weight parameter, and there are but not limited to the following possible implementation methods:
[0173] Method 1: Directly input the current sample state of the i-th group of training data into the first network to obtain the execution evaluation value sets of the respective weight parameters.
[0174] Among them, the execution evaluation value sets of the respective weight parameters can be represented by a matrix. Each row in the matrix represents the execution evaluation value set corresponding to a weight parameter, and each column represents a candidate action.
[0175] For example, for 5 weight parameters (w1, w2, w3, w4, w5), input s t into the Actor Network to obtain a 5X3 matrix. Each row in the 5X3 matrix represents the execution evaluation values of addition, subtraction, or unchanged corresponding to a weight parameter. Each column in the matrix represents a candidate action (value increase operation, value decrease operation, or value unchanged operation). Suppose the first row of the matrix is [0.3, 0.5, 0.1], indicating that the execution evaluation value of the value increase operation corresponding to w1 is 0.3, the execution evaluation value of the value decrease operation corresponding to w1 is 0.5, and the execution evaluation value of the value unchanged operation corresponding to w1 is 0.1.
[0176] Method 2: In the embodiments of the present application, the object feature of the current sample state can also be used as the input of the Actor Network in the first state transition, and the current sample state can be used as the input of the Actor Network in non-first state transitions.
[0177] Specifically, if the current state transition is the first state transition, based on the respective content features and object features included in the i-th group of training data, obtain the current sample state of the i-th group of training data, and input the object feature into the Actor Network to obtain the execution evaluation value sets of the respective weight parameters;
[0178] If the current state transition does not belong to the first state transition, directly input the current sample state of a group of training data into the first network to obtain the execution evaluation value sets of the respective weight parameters.
[0179] That is to say, in the first state transition, input s1 into the Actor Network to obtain the execution evaluation value sets of the respective weight parameters. In subsequent state transitions, input s into the Actor Network to obtain the execution evaluation value sets of the respective weight parameters.
[0180] S303: Based on the respective execution evaluation value sets, screen the corresponding target actions from the respective candidate actions, and based on the obtained target actions, obtain the respective target values corresponding to the weight parameters.
[0181] As a possible implementation, in the embodiments of the present application, according to the discretization strategy π, based on each execution evaluation value set respectively, the corresponding target actions are screened from each candidate action. The discretization strategy π is a strategy used to discretize the action space. Exemplarily, the discretization strategy π can be to select the candidate action with the highest value of the corresponding execution evaluation value as the target action.
[0182] It should be noted that in the embodiments of the present application, only taking the transition probability as 1 as an example, the discretization strategy π is described, but it is not limited thereto.
[0183] Refer to Figure 6 As shown, for w1, the respective execution evaluation values corresponding to the three candidate actions (value increase operation, value decrease operation, value unchanged operation) are 0.3, 0.5, and 0.1. Among them, the value decrease operation has the highest value of the corresponding execution evaluation value. According to the discretization strategy π, the value decrease operation is used as the target action corresponding to w1; for w2, the respective execution evaluation values corresponding to the value increase operation, value decrease operation, and value unchanged operation are 0.1, 0.3, and 0.6. Among them, the value unchanged operation has the highest value of the corresponding execution evaluation value. According to the discretization strategy π, the value unchanged operation is used as the target action corresponding to w2; for w3, the respective execution evaluation values corresponding to the value increase operation, value decrease operation, and value unchanged operation are 0.6, 0.3, and 0.1. Among them, the value increase operation has the highest value of the corresponding execution evaluation value. According to the discretization strategy π, the value increase operation is used as the target action corresponding to w3; for w4, the respective execution evaluation values corresponding to the value increase operation, value decrease operation, and value unchanged operation are 0.5, 0.2, and 0.1. Among them, the value increase operation has the highest value of the corresponding execution evaluation value. According to the discretization strategy π, the value increase operation is used as the target action corresponding to w4; for w5, the respective execution evaluation values corresponding to the value increase operation, value decrease operation, and value unchanged operation are 0.1, 0.4, and 0.2. Among them, the value decrease operation has the highest value of the corresponding execution evaluation value. According to the discretization strategy π, the value decrease operation is used as the target action corresponding to w4; In summary, the respective target actions corresponding to w1, w2, w3, w4, and w5 are: value decrease operation, value unchanged operation, value increase operation, value increase operation, value decrease operation.
[0184] As a possible implementation, in the embodiments of the present application, based on the obtained respective target actions, the target values corresponding to the respective weight parameters are obtained, specifically including:
[0185] For each weight parameter, the following operations are respectively performed:
[0186] Based on the set update step size, using the weight parameter w iThe corresponding target action is used to update the current value of the weight parameter w i to obtain the target value of the weight parameter w i corresponding thereto.
[0187] Among them, the weight parameter w i can be any one of the weight parameters. The update step size can be a preset hyperparameter. In the first state transition, the current value of the weight parameter w i corresponding thereto is the set initial value. The initial values of the weight parameters can be the same or different.
[0188] Refer to Figure 7 As shown, the target actions corresponding to w1, w2, w3, w4, and w5 are respectively: a value reduction operation, a value unchanged operation, a value increase operation, a value increase operation, and a value reduction operation. Use pStep to represent the update step size of each iteration. Assume that the current values corresponding to w1, w2, w3, w4, and w5 are all 0.5 and pStep is 0.01. Then, the target value corresponding to w1 is w1 - pStep = 0.5 - 0.01 = 0.49, the target value corresponding to w2 is w2 = 0.5, the target value corresponding to w3 is w3 + pStep = 0.5 + 0.01 = 0.51, the target value corresponding to w4 is w4 + pStep = 0.5 + 0.01 = 0.51, and the target value corresponding to w5 is w5 - pStep = w5 - 0.01 = 0.49.
[0189] S304. Based on the obtained target values, respectively rearrange each sample list in the i-th group of training data, and based on each rearranged list, obtain the next sample state.
[0190] As a possible implementation manner, in the embodiments of the present application, based on the obtained target values, respectively rearranging each sample list in a group of training data includes:
[0191] For each sample list, respectively based on the set recommended order evaluation method, combined with each target value, obtain the recommended order evaluation value of each sample content in the corresponding sample list;
[0192] Based on the obtained recommended order evaluation values, respectively rearrange each sample list.
[0193] Among them, the set recommended order evaluation method can be to calculate the recommended order evaluation value using formula (1) mentioned above. For the sake of description, the rearranged sample list can also be referred to as a rearranged list.
[0194] That is to say, after obtaining each target value, a new multi-objective formula can be obtained based on the obtained target values, and based on the new multi-objective formula, the recommended order evaluation value of each sample content can be calculated, and then the list can be rearranged according to the calculated recommended order evaluation values of each sample content.
[0195] Refer to Figure 8 As shown, taking L1 as an example, before rearrangement, L1 sequentially includes: Message 1, Message 2, Message 3,..., Message 10. Assume that the target values are: w1 = 0.49, w2 = 0.5, w3 = 0.51, w4 = 0.51, w5 = 0.49. Substitute the target values into formula (1) to obtain a new multi-objective formula, and based on the new multi-objective formula, calculate the recommended order evaluation values of Message 1, Message 2, Message 3,..., Message 10 respectively. Assume that the recommended order evaluation values of Message 1, Message 2, Message 3,..., Message 10 represent the order after rearrangement of Message 1 - Message 10 as: Message 3, Message 1, Message 2,..., Message 10. Then, based on the obtained recommended order evaluation values, rearrange L1, and the rearranged list L1' sequentially includes: Message 3, Message 1, Message 2,..., Message 10.
[0196] In the embodiments of the present application, the process of obtaining the next sample state based on each rearranged list is similar to the process of obtaining the current sample state in the above text. The process of obtaining the next sample state can adopt but is not limited to the following several possible ways:
[0197] The first way: The i-th training data includes: the object features of the sample object, each sample list and the content features of each sample content included therein. When obtaining the next sample state based on each rearranged list, it specifically includes:
[0198] For each rearranged list, perform the following operations respectively: Based on the arrangement order of each sample content in a rearranged list, screen out at least one reference content that meets the set content screening conditions from each sample content in a rearranged list;
[0199] Based on the content features of the obtained reference contents and in combination with the object features, obtain the next sample state s' of a set of training data.
[0200] Among them, the reference content that meets the set content screening conditions can be the sample content located at the set position (for example, the first sample content among each sample content, or the first three sample contents among each sample content). The reference content that meets the set content screening conditions can also be the sample content that meets the set sample content extraction interval, such as extracting one sample content as the reference content every five sample contents. Of course, the set content screening conditions are not limited to this, and no limitation is made thereto.
[0201] As a possible implementation, based on the content features of each obtained reference content, the content features of each reference content can be directly concatenated to obtain the dynamic feature s2' of the i-th group of training data. The content features of each reference content can also be feature-fused (such as calculating the mean) to obtain the dynamic feature s2', and then, based on the dynamic feature s2' and the object feature s1, the next sample state s' can be obtained.
[0202] For example, refer to Figure 9 As shown, for each distribution list in the i-th group of training data, based on the arrangement order of each sample content included in each distribution list, the first sample content in each distribution list is used as the reference content. Then, the click score, duration score, quality score, and timeliness score of the reference content corresponding to the first distribution list are used as the first 4-dimensional features of s2'; the click score, duration score, quality score, and timeliness score of the reference content corresponding to the second distribution list are used as the 5-8 dimensional features of s2'; similarly, according to the click score, duration score, quality score, and timeliness score of the reference content corresponding to each distribution list, the dynamic feature s2' is constructed. After that, based on the dynamic feature s2' and the object feature s1, the next sample state s' is obtained.
[0203] As a possible implementation, the i-th group of training data further includes: the historical interaction information of each sample content included in each sample list. Then, based on the obtained content features of each reference content and combined with the object feature, the next sample state of a group of training data is obtained, including:
[0204] Based on the historical interaction information in the i-th group of training data and combined with the recommendation order evaluation value corresponding to each permutation list, at least one task evaluation value corresponding to each target task is obtained; wherein, the recommendation order evaluation value is determined based on each target value;
[0205] Based on the obtained at least one task evaluation value, combined with the object feature and the obtained content features of each reference content, the next sample state of the i-th group of training data is obtained.
[0206] As an example, when obtaining the next sample state of the i-th group of training data based on the obtained at least one task evaluation value, combined with the object feature and the obtained content features of each reference content, specifically, based on the obtained at least one task evaluation value, combined with the content features of each reference content, the dynamic feature s2' can be obtained, and then, based on the dynamic feature s2' and the object feature s1, the next sample state s' can be obtained.
[0207] When obtaining the dynamic feature s2' based on at least one obtained task evaluation value and combining the content features of each reference content, the at least one obtained task evaluation value and the content features of each reference content can be directly concatenated to obtain the dynamic feature s2' of the i-th group of training data. It is also possible to perform feature fusion (such as calculating the mean) on the content features of each reference content and combine at least one task evaluation value to obtain the dynamic feature s2'.
[0208] Taking the target task as the click task as an example, still referring to Figure 9 As shown, based on the click labels included in each historical interaction information in the i-th group of training data and combining the recommendation order evaluation values corresponding to each re-ranked list, the task evaluation value corresponding to the click task (i.e., the click task offline metric) is obtained. Then, based on the task evaluation value corresponding to the click task and combining the object feature s1 and the content features of each obtained reference content (click score, duration score, quality score, timeliness score), the next sample state s' of the i-th group of training data is obtained.
[0209] The second method: For some of the re-ranked lists among the re-ranked lists, based on the arrangement order of each sample content in each re-ranked list, at least one reference content that meets the set content screening conditions is selected from each sample content. Then, based on the content features of each obtained reference content, the dynamic feature s2' of the i-th group of training data is obtained. The selection of some re-ranked lists can be pre-set or randomly selected, and there is no limitation on this. Since the second method is similar to the first method, it will not be elaborated here.
[0210] The third method: For each re-ranked list among the re-ranked lists, based on the arrangement order of each sample content in this re-ranked list and combining the content features of each sample content in this re-ranked list, the list feature corresponding to this sample list is obtained; based on the list features corresponding to each obtained sample list, the dynamic feature s2' is constructed; then, based on the constructed dynamic feature s2' and combining the object feature s1, the next sample state s' is obtained.
[0211] Among them, the arrangement order of each sample content in the re-ranked list is determined according to the recommendation order evaluation values of each of the N sample contents. The list feature corresponding to a re-ranked list can also include the task evaluation index of the target task.
[0212] Obviously, the next sample state s' is composed of the static feature s1 and the dynamic feature s2'. The way to obtain the dynamic feature s2' is similar to the way to obtain the dynamic feature s2, and it will not be elaborated here.
[0213] S305. Based on each set of execution evaluation values, combined with the current sample state, use the second network to obtain a predicted evaluation value, and based on the next sample state, use the second network to obtain a true evaluation value, and obtain a model loss based on the true evaluation value and the predicted evaluation value.
[0214] In the embodiment of the present application, the output result a of the Actor Network (i.e., each set of execution evaluation values) and the current sample state s are input into the Critic Network to obtain a predicted evaluation value Q(s, a).
[0215] As a possible implementation, before obtaining the true evaluation value based on the next sample state using the second network, it is also possible to test each task based on the rearranged list of samples to obtain feedback information, where the feedback information is used to represent: after transferring from the current sample state to the next sample state through each target action, the task evaluation value of each task.
[0216] Among them, each task includes but is not limited to one or more of click tasks, duration tasks, quality tasks, and timeliness tasks. The task evaluation value of each task can be represented by but is not limited to GAUC.
[0217] Exemplarily, a weighted sum of the task evaluation values of each task is obtained to obtain feedback information, and the weights of the task evaluation values of each task can be set as needed.
[0218] For example, the feedback information r = 1 * GAUC(click task) + 0.1 * GAUC(duration task) + 0.1 * GAUC(quality task) + 0.1 * GAUC(timeliness task). Among them, GAUC(click task) represents the task evaluation value of the click task, GAUC(duration task) represents the task evaluation value of the duration task, GAUC(quality task) represents the task evaluation value of the quality task, and GAUC(timeliness task) represents the task evaluation value of the timeliness task.
[0219] As a possible implementation, in the embodiment of the present application, when obtaining the true evaluation value based on the next sample state using the second network, it can be but is not limited to the following methods:
[0220] Based on the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state, and obtain the true evaluation value based on the predicted evaluation value corresponding to the next sample state and the feedback information.
[0221] Among them, based on the next sample state, using the second network to obtain the predicted evaluation value corresponding to the next sample state specifically includes:
[0222] Based on the next sample state, use the first network to obtain the set of execution evaluation values corresponding to the next sample state for each weight parameter.
[0223] Based on the obtained set of execution evaluation values corresponding to each next sample state, and in combination with the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state.
[0224] Refer to Figure 10 As shown, in the embodiment of the present application, s’ is input into the Actor Network to obtain the corresponding output result a’ (i.e., the set of execution evaluation values corresponding to the next sample state for each weight parameter). Then, a’ and s’ are input into the Critic Network to obtain the predicted evaluation value Q’(s’, a’) corresponding to the next sample state.
[0225] In the embodiment of the present application, the model loss can specifically adopt the network loss metrics of the Actor Network and the Critic Network respectively. The first network loss is used to characterize the network loss of the Actor Network, and the second network loss is used to characterize the network loss of the Critic Network. Specifically, when obtaining the model loss based on the true evaluation value and the predicted evaluation value, it can be implemented in but not limited to the following ways:
[0226] Based on the predicted evaluation value, in combination with the gradient of each target action with respect to the model parameters, obtain the first network loss; and, based on the true evaluation value and the predicted evaluation value, obtain the second network loss; based on the second network loss and the first network loss, obtain the model loss.
[0227] Exemplarily, the first network loss is represented by formula (2):
[0228]
[0229] where Actor Loss represents the first network loss, Q(s,a) represents the predicted evaluation value corresponding to the current sample state, represents the gradient of the action (the action generated by the Actor Network) with respect to the network parameters, E[·] represents the expected value, and E[·] is used to calculate the average loss of all possible state transitions. Among them, the action gradient is logarithmically processed to facilitate gradient calculation.
[0230] Exemplarily, the second network loss is represented by formula (3):
[0231] Critic Loss = E[(Q(s, a)-(r + γ * Q’(s’, a’))) 2 Formula (3)
[0232] Among them, Critic Loss represents the second network loss, r represents the feedback information, r + γ * Q'(s', a') represents the true evaluation value, γ represents the discount factor, γ is used to measure the importance of future rewards, E[·] represents the expected value, and E[·] is used to calculate the average loss of all possible state transitions.
[0233] S306. Determine whether the state transition end condition is satisfied. If so, execute S307; otherwise, continue the state transition for the i-th group of training data, that is, jump to S302.
[0234] Among them, the state transition end condition includes but is not limited to the number of state transitions reaching a threshold.
[0235] S307. Determine whether the preset batch size is reached. If so, execute S308; otherwise, perform at least one state transition for the (i + 1)-th group of training data in each group of training data, that is, jump to S301.
[0236] S308. Adjust the parameters based on the obtained model losses.
[0237] If the above conditions are satisfied, it is determined that the content recommendation model has met the convergence condition (that is, the preset batch size is reached), and the training ends; otherwise, it is determined that the content recommendation model has not met the convergence condition, then the model parameters need to be further adjusted, and the adjusted content recommendation model is used to enter the next iteration of training.
[0238] In the embodiments of the present application, the total loss value obtained from the outputs of the first network and the second network is used to adjust the entire model. In this way, both the first network and the second network can learn, thereby further improving the recommendation accuracy of the entire model.
[0239] As a possible implementation, if the content recommendation model meets the set model convergence condition, output the respective target values corresponding to each group of training data, and store the object features included in each group of training data and their respective corresponding target values in an associated manner in the target storage area;
[0240] After receiving the target object features of the target object, obtain the respective target values associated with the target object features from the target storage area according to the target object features.
[0241] Among them, the target storage area includes but is not limited to a database. Exemplarily, the target object features of the target object can be carried in the data query request. After receiving the data query request, obtain the associated respective target values from the database according to the target object features carried in the data query request.
[0242] That is to say, when the content recommendation model has met the convergence condition, the target values of the associated storage object features and each weight coefficient are correlated, so that during online inference, the corresponding values can be directly obtained from the database, thereby greatly reducing the computational overhead of online real-time request inference.
[0243] For example, after the model converges, the object feature A of object A and the target values of its corresponding five weight coefficients are associated and stored in the database. In this way, when the object feature A is received, the target values of the five weight coefficients associated with the object feature A are obtained from the database according to the object feature A.
[0244] Furthermore, it is also possible to obtain the value score sets of each candidate recommendation information; based on the target values associated with the target object feature, each value score in the value score sets of each candidate recommendation information is fused respectively to obtain the recommendation order evaluation values corresponding to each candidate recommendation information; based on the obtained recommendation order evaluation values, the target recommendation information is selected from each candidate recommendation information for recommendation.
[0245] Among them, when selecting the target recommendation information for recommendation, the top K candidate recommendation information can be selected based on the obtained recommendation order evaluation values as the target recommendation information. The candidate recommendation information includes but is not limited to subscription number messages, videos, audios, picture texts, etc.
[0246] For example, assume that there are M candidate recommendation information, each candidate recommendation information corresponds to a value score set, and each value score set contains a click score, a duration score, a quality score, and a timeliness score. Based on the target values of the five weight coefficients associated with the object feature A, using formula (1), each value score in the value score sets of each candidate recommendation information is fused respectively to obtain the recommendation order evaluation values corresponding to each candidate recommendation information. Then, based on the obtained M recommendation order evaluation values, the first 10 candidate recommendation information are selected from the M candidate recommendation information as the target recommendation information. Next, the present application will be described in conjunction with a specific embodiment.
[0247] Refer to Figure 11A As shown, it is a logical schematic diagram of a quadruple construction process provided in an embodiment of the present application.
[0248] First, receive the issued original data, which contains the relevant data of multiple sample lists in a batch. Specifically, the original data contains the context of each sample content in each sample list in a batch (such as object portrait information), value scores (such as one or more of click score, duration score, quality score, timeliness score, etc.), and labels (such as one or more of click label, duration label, etc.). The feature dimensions of the original data are determined according to the batch size, the number of sample lists, the number of contents contained in each sample list, the context of each content, the value scores, and the dimensions of the labels. It should be noted that since the number of contents contained in each sample list can be the same or different, the number of contents is variable-length.
[0249] Next, calculate the current sample state (s). The current sample state (s) of the sample list is composed of scoring statistical information of multiple sample lists and various offline metrics, etc. Among them, the scoring statistical information includes: click score, duration score, quality score, timeliness score, and each offline metric includes click task offline metric, duration task offline metric, quality task offline metric, and timeliness task offline metric.
[0250] Next, according to the current sample state, obtain a. Specifically, input s1 into the Actor Network to obtain the output result a. a is a 5×3 matrix, and each row of the matrix is used to represent the execution evaluation value of each weight parameter.
[0251] Next, obtain the respective target values according to a. Specifically, based on each set of execution evaluation values, filter the corresponding target actions from each candidate action, and based on the obtained target actions, obtain the target values corresponding to each weight parameter.
[0252] Next, perform rearrangement according to the respective target values. Among them, rearrangement refers to offline simulation rearrangement. Offline simulation rearrangement can be implemented in two ways. The first way is the fine-rank score rearrangement simulation. Specifically, for multiple sample lists, based on the set recommendation order evaluation method, combined with each target value, obtain the respective recommendation order evaluation values of each sample content in the corresponding sample list. Based on the obtained recommendation order evaluation values, rearrange each sample list respectively. The second way is that the background provides a gateway simulation interface. According to each target value, use the interface to rearrange multiple sample lists respectively.
[0253] Next, based on the rearranged multiple sample lists, calculate the next sample state (s’). s’ is composed of scoring statistical information of the rearranged multiple sample lists and various offline metrics, etc.
[0254] Finally, according to the next sample state, calculate r, so as to obtain the quadruple (s, a, r, s’).
[0255] Obviously, in the part of state modeling for the production of the quadruple here, only the part related to dynamic features is involved. The reason for this design is that the recommendation system is a system with feedback delay and it is difficult to obtain the real-time changes of object features under the new multi-objective formula in real time. Therefore, when modeling the state, the dynamic features on the object side (such as the number of clicks of the object in a recent period of time) are avoided. When generating the quadruple, only the "dynamic features" of the statistical information of the issued list that can be statistically counted offline in real time are adjusted, so as to ensure the real-time nature of the processing. Refer to Figure 11B As shown, it is a logical schematic diagram of a state transition process provided in an embodiment of the present application.
[0256] First, when the environment of the reinforcement learning obtains an object feature s1, according to the object feature s1, each issued list corresponding to the object feature s1 (32 dimensions) is screened out from the issued logs. Among them, the sample content included in each issued list is a video. For each issued list, the click score, duration score, quality score, and timeliness score of each video in the issued list are saved offline.
[0257] In the first state transition, according to the average values of the click scores, duration scores, quality scores, and timeliness scores of the first video in each issued list, the first 1-4 dimensional features of the dynamic feature (s2) are obtained. According to the average values of the click scores, duration scores, quality scores, and timeliness scores of the second video in each issued list, the first 4-8 dimensional features of s2 are obtained. Assume that the issued list usually contains 10 videos. Then, similarly, according to the average values of the click scores, duration scores, quality scores, and timeliness scores of the videos in the 10 positions in each issued list respectively, 40-dimensional s2 is obtained. Furthermore, in combination with the task evaluation value of the click task, 41-dimensional s2 is obtained. Based on the 41-dimensional s2 and the 32-dimensional s1, the current sample state (s) is obtained.
[0258] Then, s1 is input into the Actor Network to obtain an output result a, where a is a 5×3 matrix, and each row of the matrix is used to represent the execution evaluation value of each weight parameter. Then, based on each set of execution evaluation values respectively, the corresponding target actions are screened out from each candidate action, and based on the obtained target actions, the target values corresponding to each weight parameter are obtained.
[0259] Next, based on the obtained target values, each sample list is rearranged respectively, and the average click score, average duration score, average quality score, and average timeliness score of the videos at 10 positions in each rearranged list are used respectively. Combining with the task evaluation value of the current click task, s2’ is obtained. Then, based on s1 and s2’, the next sample state (s’) is obtained, s’ = [s1, s2’]. In addition, based on the task evaluation values of the click task, duration task, quality task, and timeliness task respectively, feedback information (r) is obtained.
[0260] Next, a and s are input into the Critic Network to obtain the predicted evaluation value (Q(s, a)) corresponding to the current sample state.
[0261] Next, s’ is input into the Actor Network to obtain the output result a’, and a’ and s’ are input into the Critic Network to obtain the predicted evaluation value (Q’(s’, a’)) corresponding to the next sample state.
[0262] Next, based on Q(s, a), r, and Q’(s’, a’), the first network loss and the second network loss can be calculated respectively, and based on the first network loss and the second network loss, the model loss of the first state transition is obtained. At this time, the first quadruple (s, a, r, Q) is constructed.
[0263] In the second state transition process, since in the first state transition process, the output result a’ corresponding to s’ was output using the Actor Network, therefore, the output result a’ directly obtained by inputting s’ into the Actor Network can be used. Based on each execution evaluation value set represented by a’, the corresponding target actions are screened from each candidate action, and based on the obtained target actions, the target values corresponding to each weight parameter are obtained.
[0264] Next, based on the obtained target values, each sample list is rearranged respectively, and the average click score, average duration score, average quality score, and average timeliness score of the videos at 10 positions in each rearranged list are used respectively. Combining with the task evaluation value of the current click task, s2” is obtained. Then, based on s1 and s2”, the next sample state (s”) is obtained, s” = [s1, s2”]. In addition, based on the task evaluation values of the click task, duration task, quality task, and timeliness task respectively, feedback information (r’) is obtained.
[0265] Next, input s” into the Actor Network to obtain the output result a”, and input a” and s” into the Critic Network to obtain the predicted evaluation value corresponding to the next sample state (Q”(s”, a”)).
[0266] Next, based on Q”(s”, a”), r’, and Q’(s’, a’), the first network loss and the second network loss can be calculated respectively, and based on the first network loss and the second network loss, the model loss of the second state transition is obtained. At this time, the first quadruple (s’, a’, r’, Q’(s’, a’)) is constructed.
[0267] During the third state transition, according to the output result a” obtained by inputting s” into the Actor Network, based on each execution evaluation value set represented by a”, the corresponding target actions are screened from each candidate action, and based on the obtained target actions, the target values corresponding to each weight parameter are obtained.
[0268] Next, based on the obtained target values, each sample list is rearranged respectively, and based on the average of the click scores, the average of the duration scores, the average of the quality scores, and the average of the timeliness scores of the videos located at 10 positions in each rearranged list, combined with the task evaluation value of the current click task, s2”’ is obtained, and then based on s1 and s2”’, the next sample state (s”’) is obtained, s”’ = [s1, s2”’]. In addition, based on the task evaluation values of the click task, the duration task, the quality task, and the timeliness task respectively, feedback information (r”) is obtained.
[0269] Next, input s”’ into the Actor Network to obtain the output result a”’, and input a”’ and s”’ into the Critic Network to obtain the predicted evaluation value corresponding to the next sample state (Q”’(s”’, a”’)).
[0270] Next, based on Q”’(s”’, a”’), r”, and Q”(s”, a”), the first network loss and the second network loss can be calculated respectively, and based on the first network loss and the second network loss, the model loss of the second state transition is obtained. At this time, the third quadruple (s”, a”, r”, Q”’(s”’, a”’)) is constructed.
[0271] Similarly, after 8 state transitions, 8 quadruples are obtained, and training is performed for the next set of training data.
[0272] Based on the same inventive concept, an embodiment of the present application provides a training device for a content recommendation model. As Figure 12As shown in the figure, it is a schematic structural diagram of a training device 1200 for a content recommendation model. The content recommendation model includes: a first network for action generation and a second network for action state evaluation. The device includes:
[0273] An evaluation unit 1201, configured to perform at least one state transition on each group of training data respectively by using the content recommendation model. In each state transition, the evaluation unit 1201 is specifically configured to perform the following operations:
[0274] For each group of training data, based on the current sample state of a group of training data, use the first network to obtain an execution evaluation value set for each weight parameter. Each execution evaluation value set includes: the execution evaluation value of each candidate action corresponding to the corresponding weight parameter, and one candidate action represents a type of numerical adjustment operation.
[0275] Based on each execution evaluation value set respectively, screen the corresponding target actions from the candidate actions, and based on the obtained target actions, obtain the target values corresponding to each weight parameter.
[0276] Based on the obtained target values, rearrange each sample list in the group of training data respectively, and based on each rearranged list, obtain the next sample state.
[0277] Based on each execution evaluation value set, combine the current sample state, use the second network to obtain a predicted evaluation value, and based on the next sample state, use the second network to obtain a true evaluation value, and based on the true evaluation value and the predicted evaluation value, obtain the model loss.
[0278] A parameter tuning unit 1202, configured to perform parameter tuning based on the obtained model losses.
[0279] As a possible implementation manner, the group of training data includes: the object features of the sample object, each sample list, and the content features of each sample content included therein.
[0280] When obtaining the next sample state based on each rearranged list, the evaluation unit 1201 is specifically configured to:
[0281] For each rearranged list, perform the following operations respectively: based on the arrangement order of each sample content in a rearranged list, screen at least one reference content that meets the set content screening condition from each sample content in a rearranged list.
[0282] Based on the content features of the obtained reference contents, combine the object features to obtain the next sample state of the group of training data.
[0283] As a possible implementation, when obtaining the execution evaluation value sets of the respective weight parameters by using the first network based on the current sample state of a set of training data, the evaluation unit 1201 is specifically configured to:
[0284] If the current state transition is the first state transition, based on each content feature and the object feature included in the set of training data, obtain the current sample state of the set of training data, and input the object feature into the first network to obtain the execution evaluation value sets of the respective weight parameters;
[0285] If the current state transition does not belong to the first state transition, directly input the current sample state of the set of training data into the first network to obtain the execution evaluation value sets of the respective weight parameters.
[0286] As a possible implementation, the set of training data further includes: historical interaction information of each sample content included in each sample list;
[0287] When obtaining the next sample state of the set of training data by combining the content features of the obtained reference contents and the object feature, the evaluation unit 1201 is specifically configured to:
[0288] Based on the historical interaction information in the set of training data, combine the recommendation order evaluation values corresponding to the respective rearrangement lists to obtain the task evaluation values corresponding to at least one target task; wherein, the recommendation order evaluation value is determined based on the respective target values;
[0289] Based on the obtained at least one task evaluation value, combine the object feature and the content features of the obtained reference contents to obtain the next sample state of the set of training data.
[0290] As a possible implementation, when rearranging each sample list in the set of training data based on the obtained respective target values, the evaluation unit 1201 is specifically configured to:
[0291] For each sample list, respectively, based on a set recommendation order evaluation method, combine the respective target values to obtain the recommendation order evaluation values of each sample content in the corresponding sample list;
[0292] Based on the obtained recommendation order evaluation values, rearrange each sample list respectively.
[0293] As a possible implementation, after rearranging each sample list in the set of training data based on the obtained respective target values, before obtaining the true evaluation value by using the second network based on the next sample state, the evaluation unit 1201 is further configured to:
[0294] Based on each rearranged sample list, each task is tested to obtain feedback information, and the feedback information is used to characterize: after transferring from the current sample state to the next sample state through the respective target actions, the task evaluation values of each task;
[0295] When obtaining the true evaluation value by using the second network based on the next sample state, the evaluation unit 1201 is specifically configured to:
[0296] Based on the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state, and based on the predicted evaluation value corresponding to the next sample state and the feedback information, obtain the true evaluation value.
[0297] As a possible implementation manner, when obtaining the predicted evaluation value corresponding to the next sample state by using the second network based on the next sample state, the evaluation unit 1201 is specifically configured to:
[0298] Based on the next sample state, use the first network to obtain the set of execution evaluation values corresponding to the next sample state of each weight parameter;
[0299] Based on the obtained set of execution evaluation values corresponding to the next sample state, in combination with the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state.
[0300] As a possible implementation manner, when obtaining the model loss based on the true evaluation value and the predicted evaluation value, the evaluation unit 1201 is specifically configured to:
[0301] Based on the predicted evaluation value, in combination with the gradients of the model parameters by the respective target actions, obtain the first network loss; and, based on the true evaluation value and the predicted evaluation value, obtain the second network loss;
[0302] Based on the second network loss and the first network loss, obtain the model loss.
[0303] As a possible implementation manner, when obtaining the target value corresponding to each weight parameter based on the obtained respective target actions, the evaluation unit 1201 is specifically configured to:
[0304] For each weight parameter, perform the following operations respectively:
[0305] Based on a set update step size, use the target action corresponding to a weight parameter to update the current value corresponding to the weight parameter, and obtain the target value corresponding to the weight parameter.
[0306] As a possible implementation, each set of training data includes the object features of the corresponding sample object. After parameter adjustment based on the obtained model losses, the evaluation unit 1201 is further configured to:
[0307] If the content recommendation model meets the set model convergence condition, output the target values corresponding to each group of training data, and associatively store the object features included in each group of training data and their corresponding target values in a target storage area;
[0308] When the target object features of a target object are received, according to the target object features, obtain the target values associated with the target object features from the target storage area.
[0309] As a possible implementation, the evaluation unit 1201 is further configured to:
[0310] Obtain the value score sets of each candidate recommendation information;
[0311] Based on the target values associated with the target object features, respectively fuse the value scores in the value score sets of each candidate recommendation information to obtain the recommendation order evaluation values corresponding to each candidate recommendation information;
[0312] Based on the obtained recommendation order evaluation values, select target recommendation information from the candidate recommendation information for recommendation.
[0313] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0314] Regarding the device in the above embodiments, the specific manners in which each unit executes requests have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0315] Those skilled in the art of the present technical field can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0316] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. In one embodiment, the electronic device may be a server or a terminal device. Refer to Figure 13 As shown, it is a schematic structural diagram of a possible electronic device provided in an embodiment of the present application. Figure 13 In this, the electronic device 1300 includes: a processor 1310 and a memory 1320.
[0317] Among them, the memory 1320 stores a computer program executable by the processor 1310. By executing the instructions stored in the memory 1320, the processor 1310 can execute the steps of the above content recommendation model training method.
[0318] The memory 1320 may be a volatile memory, such as a random-access memory (RAM); the memory 1320 may also be a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1320 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1320 may also be a combination of the above memories.
[0319] The processor 1310 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1310 is used to implement the above content recommendation model training method when executing the computer program stored in the memory 1320.
[0320] In some embodiments, the processor 1310 and the memory 1320 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0321] In the embodiments of the present application, the specific connection medium between the above processor 1310 and the memory 1320 is not limited. In the embodiments of the present application, taking the connection between the processor 1310 and the memory 1320 through a bus as an example, the bus is described in thick lines in Figure 13 The connection manners between other components are only for illustrative purposes and are not to be taken as a limitation. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 13 only one thick line is used to describe it in, but it does not describe that there is only one bus or one type of bus.
[0322] Based on the same inventive concept, embodiments of the present application provide a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above content recommendation model training method. In some possible implementation manners, various aspects of the content recommendation model training method provided by the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the above content recommendation model training method. For example, the electronic device can execute as Figure 3 the steps shown in
[0323] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (Compact Disk Read Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0324] The program product of the embodiments of the present application can adopt a CD-ROM and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a computer program, and the computer program can be used by or in combination with a command execution system, apparatus, or device.
[0325] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a computer program for use by or in combination with a command execution system, apparatus, or device.
[0326] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0327] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these modifications and variations.
Claims
1. A training method for a content recommendation model, characterized in that, The content recommendation model includes: a first network for action generation and a second network for action state evaluation. The method includes: For each group of training data, perform at least one state transition using the content recommendation model respectively, and adjust the parameters based on the obtained model losses of each time. Wherein, each state transition includes: Based on the current sample state of a group of training data, use the first network to obtain the execution evaluation value sets of each weight parameter. Each execution evaluation value set includes: the execution evaluation values of each candidate action corresponding to the corresponding weight parameter. A candidate action represents a type of numerical adjustment operation. Based on each execution evaluation value set respectively, screen the corresponding target actions from the candidate actions, and based on the obtained target actions, obtain the target values corresponding to each weight parameter respectively. Based on the obtained target values, rearrange each sample list in the group of training data respectively, and based on each rearranged list, obtain the next sample state. Based on each execution evaluation value set, combine the current sample state, use the second network to obtain the predicted evaluation value, and based on the next sample state, use the second network to obtain the true evaluation value, and based on the true evaluation value and the predicted evaluation value, obtain the model loss.
2. The method according to claim 1, wherein The group of training data includes: the object features of the sample object, each sample list and the content features of each sample content included therein. The obtaining the next sample state based on each rearranged list includes: For each rearranged list, perform the following operations respectively: based on the arrangement order of each sample content in a rearranged list, screen at least one reference content that meets the set content screening conditions from each sample content in a rearranged list. Based on the content features of the obtained reference contents, combine with the object features to obtain the next sample state of the group of training data.
3. The method according to claim 2, wherein The obtaining the execution evaluation value sets of each weight parameter respectively by using the first network based on the current sample state of a group of training data includes: If the current state transition is the first state transition, then based on each content feature and the object feature included in the group of training data, obtain the current sample state of the group of training data, and input the object feature into the first network to obtain the execution evaluation value sets of each weight parameter. If the current state transition does not belong to the first state transition, directly input the current sample state of the group of training data into the first network to obtain the execution evaluation value sets of each weight parameter.
4. The method according to claim 2, wherein The group of training data further includes: the historical interaction information of each sample content included in each sample list. The obtaining the next sample state of the group of training data by combining the content features of the obtained reference contents with the object features includes: Based on the historical interaction information in the group of training data, combine with the recommendation order evaluation values corresponding to each rearranged list to obtain the task evaluation values corresponding to at least one target task respectively. Wherein, the recommendation order evaluation value is determined based on the obtained target values. Based on at least one task evaluation value obtained, in combination with the object features and the content features of each piece of reference content obtained, obtain the next sample state of the set of training data.
5. The method according to any one of claims 1 to 4, characterized in that, The rearranging of each sample list in the set of training data based on the obtained target values for each includes: For each sample list, respectively, based on a set recommended order evaluation method, in combination with the target values for each, obtain the recommended order evaluation value for each sample content in the corresponding sample list; Based on the obtained recommended order evaluation values, rearrange each sample list respectively.
6. The method according to any one of claims 1-4, characterized in that, After the rearranging of each sample list in the set of training data based on the obtained target values for each, and before obtaining the true evaluation value by using the second network based on the next sample state, further include: Based on the rearranged sample lists, test each task to obtain feedback information, where the feedback information is used to characterize: after transferring from the current sample state to the next sample state through the target actions for each, the task evaluation values of each task; The obtaining of the true evaluation value by using the second network based on the next sample state includes: Based on the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state, and based on the predicted evaluation value corresponding to the next sample state and the feedback information, obtain the true evaluation value.
7. The method according to claim 6, wherein The obtaining of the predicted evaluation value corresponding to the next sample state by using the second network based on the next sample state includes: Based on the next sample state, use the first network to obtain the set of execution evaluation values corresponding to the next sample state for each of the weight parameters; Based on the obtained set of execution evaluation values corresponding to the next sample state, in combination with the next sample state, use the second network to obtain the predicted evaluation value corresponding to the next sample state.
8. The method according to any one of claims 1 to 4, characterized in that The obtaining of the model loss based on the true evaluation value and the predicted evaluation value includes: Based on the predicted evaluation value, in combination with the gradient of the model parameters by the target actions for each, obtain the first network loss; and, based on the true evaluation value and the predicted evaluation value, obtain the second network loss; Based on the second network loss and the first network loss, obtain the model loss.
9. The method according to any one of claims 1-4, characterized in that, The obtaining of the target value corresponding to each weight parameter based on the obtained target actions for each includes: For each weight parameter, respectively perform the following operations: Based on a set update step size, use the target action corresponding to a weight parameter to update the current value corresponding to the weight parameter to obtain the target value corresponding to the weight parameter.
10. The method according to any one of claims 1-4, characterized in that, Each set of training data contains the object features of the corresponding sample object. After parameter adjustment based on the obtained model losses for each, further include: If the content recommendation model meets the set model convergence condition, output the target values corresponding to each set of training data, and associate and store the object features included in each set of training data and their corresponding target values for each in a target storage area; After receiving the target object features of the target object, according to the target object features, obtain the respective target values associated with the target object features from the target storage area.
11. The method according to 10, characterized in that, It further includes: Obtain the value score sets of the respective candidate recommendation information; Based on the respective target values associated with the target object features, fuse the respective value scores in the value score sets of the respective candidate recommendation information to obtain the respective recommendation order evaluation values corresponding to the respective candidate recommendation information; Based on the obtained respective recommendation order evaluation values, select the target recommendation information from the respective candidate recommendation information for recommendation.
12. A training device for a content recommendation model, characterized in that, The content recommendation model includes: a first network for action generation and a second network for action state evaluation, and the device includes: An evaluation unit for performing at least one state transition respectively on each group of training data by using the content recommendation model; wherein, in each state transition, the evaluation unit is specifically used to perform the following operations: For each group of training data, based on the current sample state of a group of training data, use the first network to obtain the respective execution evaluation value sets of the respective weight parameters, and each execution evaluation value set includes: the respective execution evaluation values of the respective candidate actions corresponding to the corresponding weight parameters, and one candidate action represents a type of numerical adjustment operation; Based on the respective execution evaluation value sets respectively, screen the corresponding target actions from the respective candidate actions, and based on the obtained respective target actions, obtain the respective target values corresponding to the respective weight parameters; Based on the obtained respective target values, rearrange each sample list in the group of training data respectively, and based on each rearranged list, obtain the next sample state; Based on the respective execution evaluation value sets, in combination with the current sample state, use the second network to obtain a predicted evaluation value, and based on the next sample state, use the second network to obtain a true evaluation value, and based on the true evaluation value and the predicted evaluation value, obtain a model loss; A parameter adjustment unit for performing parameter adjustment based on the obtained respective model losses.
13. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, It includes a computer program, and the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the method according to any one of claims 1 to 11.