Model adjustment method and apparatus

By predicting user interaction behavior and optimizing parameters of the rewriting model, the problem of poor model rewriting effect in the existing technology is solved, and continuous optimization of the rewriting effect and improvement of click-through rate are achieved.

WO2025201217A1PCT designated stage Publication Date: 2025-10-02VIVO MOBILE COMM CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084270
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-03-24
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

The model rewriting effect in existing technologies is poor and lacks an effective evaluation mechanism, resulting in insignificant improvement in click-through rate.

Method used

By obtaining the original information, rewriting it using the rewriting model, evaluating the rewriting effect in combination with the user interaction behavior prediction model, optimizing the rewriting model parameters based on the evaluation results, and adjusting the model using the proximal strategy optimization algorithm.

Benefits of technology

The click-through rate prediction accuracy and rewriting effect of the rewriting model were improved, the deviation of manual evaluation was reduced, and the continuous optimization of the model rewriting effect was achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084270_02102025_PF_FP_ABST
    Figure CN2025084270_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a model adjustment method and apparatus, relating to the technical field of information. The model adjustment method comprises: acquiring first information (101); using a first rewriting model to rewrite the first information to obtain rewritten second information (102); performing user interaction behavior prediction on the basis of the second information to obtain predicted interaction information (103); and on the basis of the interaction information, adjusting parameters of the first rewriting model to obtain a second rewriting model (104).
Need to check novelty before this filing date? Find Prior Art

Description

Model adjustment method and device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202410356220.1 filed on March 27, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application belongs to the field of information technology, and specifically relates to a model adjustment method and device thereof. Background Art

[0004] When searching for information, users usually browse based on the title information displayed on the search interface and click on the title content that interests them.

[0005] In the existing technology, in order to increase the user click-through rate of title content, the experience of operational experts is often used to manually write some eye-catching titles to increase the click-through rate; some also use large models to rewrite titles to achieve the purpose of attracting user clicks, but due to the lack of an effective evaluation mechanism for the rewriting results, the rewriting effect of the model is poor. Summary of the Invention

[0006] The purpose of the embodiments of the present application is to provide a model adjustment method and device thereof, which can solve the problem of poor rewriting effect of existing models.

[0007] In a first aspect, an embodiment of the present application provides a model adjustment method, the method comprising:

[0008] Obtaining first information;

[0009] rewriting the first information using a first rewriting model to obtain rewritten second information;

[0010] Predicting user interaction behavior based on the second information to obtain predicted interaction information;

[0011] Based on the interaction information, the parameters of the first rewriting model are adjusted to obtain a second rewriting model.

[0012] In a second aspect, an embodiment of the present application provides a model adjustment device, comprising:

[0013] An acquisition module, configured to acquire first information;

[0014] a rewriting module, configured to rewrite the first information using a first rewriting model to obtain rewritten second information;

[0015] a prediction module, configured to predict user interaction behavior based on the second information and obtain predicted interaction information;

[0016] An adjustment module is used to adjust the parameters of the first rewriting model based on the interaction information to obtain a second rewriting model.

[0017] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0019] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.

[0020] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect.

[0021] In the embodiment of the present application, first information is obtained; the first information is rewritten using a first rewriting model to obtain rewritten second information; user interaction behavior is predicted based on the second information to obtain predicted interaction information; and based on the interaction information, the parameters of the first rewriting model are adjusted to obtain a second rewriting model. In this way, by predicting user interaction behavior based on the information rewritten by the rewriting model, an evaluation mechanism for the rewriting results is introduced, and the rewriting model is optimized and adjusted based on the evaluation results, which can improve the rewriting effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIG1 is a flow chart of a model adjustment method provided in an embodiment of the present application;

[0023] FIG2 is a flowchart of online request and data collection provided by an embodiment of the present application;

[0024] FIG3 is a schematic diagram of a click-through rate model training process provided in an embodiment of the present application;

[0025] FIG4 is a schematic diagram of a rewriting model training process provided in an embodiment of the present application;

[0026] FIG5 is a structural diagram of a model adjustment device provided in an embodiment of the present application;

[0027] FIG6 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0028] FIG7 is a hardware structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0030] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0031] The model adjustment method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0032] Please refer to FIG1 , which is a flow chart of a model adjustment method provided in an embodiment of the present application. As shown in FIG1 , the method includes the following steps:

[0033] Step 101: Obtain first information.

[0034] The above-mentioned first information can be original, unrewritten information, for example, it can be the original search result in a search scenario, or it can be the original recommendation result in an information recommendation scenario. In particular, the first information can be the title of each search result data / each recommendation data.

[0035] Optionally, the first information is the title of the recommended content.

[0036] That is, the embodiment of the present application can be applied to rewrite the recommended content title in the content recommendation scenario to increase the probability of users clicking, liking, sharing, and other interactive behaviors on the recommended content title. The above-mentioned acquisition of the first information can be obtaining the title of the recommended content.

[0037] Step 102: rewrite the first information using a first rewriting model to obtain rewritten second information.

[0038] The first rewriting model may be a general model for rewriting title information, such as a large language model. When used as a rewriting model, it may be provided with a rewriting function prompt.

[0039] In this step, the first information may be input into the first rewriting model, and the first information may be appropriately rewritten using the first rewriting model to obtain a rewriting result output by the first rewriting model, that is, the second information.

[0040] For example, the first information is the article title “How to Make Sushi”, and the rewritten second information is “Learn how to easily make professional-grade sushi at home!”. The rewritten content can be more attractive to users to click and view.

[0041] Optionally, step 102 includes:

[0042] The first information, user information and usage scenario information are input into the first rewriting model, and the first information, the user information and the usage scenario information are integrated through the first rewriting model to rewrite the first information.

[0043] In some embodiments, user information and usage scenario information can also be used as input parameters of the first rewriting model. The usage scenario information can be information such as usage environment, usage time, and APP usage, so that the first rewriting model can integrate the original content, user characteristics, and usage scenario characteristics for rewriting, ensuring that the rewriting results are more in line with the user's habits in a specific usage scenario, that is, it can improve the click-through rate of the rewriting results and other interactive behavior probabilities.

[0044] Step 103: Predict user interaction behavior based on the second information to obtain predicted interaction information.

[0045] In an embodiment of the present application, a mechanism for evaluating the rewriting effect of the rewriting model can be introduced, that is, the rewriting result of the first rewriting model can be evaluated, specifically predicting user interaction behavior with the second information, that is, predicting the possibility of the user interacting with the second information, such as predicting whether the user will click on the second information to view it, and even predicting whether the user will like or share it after viewing it. The specific prediction mechanism can be based on the user's historical interaction behavior data, such as judging whether the second information belongs to the type of information that the user frequently clicks on based on the type of information that the user frequently clicks on. If so, it is predicted that the user is more likely to click on the second information; otherwise, it is predicted that the user is less likely to click on the second information.

[0046] In this step, by predicting user interaction behavior for the second information, predicted interaction information can be obtained. The interaction information can be characterized by interaction probability, possibility, etc. When the second information includes multiple rewriting results, the interaction information can include the interaction probability for each rewriting result.

[0047] Optionally, step 103 includes:

[0048] An evaluation model is used to predict user interaction behavior for the second information, wherein the evaluation model is trained using rewritten sample data, and the rewritten sample data includes rewritten data rewritten by the first rewriting model and user interaction behavior data on the rewritten data.

[0049] In some embodiments, the rewriting data of the rewriting model and the user interaction result data can be used as rewriting sample data to train an evaluation model, and then the evaluation model is used to predict the rewriting results of the rewriting model, thereby improving the effectiveness and accuracy of the rewriting model evaluation mechanism.

[0050] Specifically, rewriting sample data can be collected first. As shown in Figure 2, the original content is input into the rewriting model, which then outputs "rewritten content" that is more attractive to users. Here, the original content, or input, is represented by the symbol x, and the rewriting model is represented by the symbol π. The content output by the rewriting model is represented by the symbol y, so the rewriting process can be expressed as y = π(x, u), where u represents the specific user, that is, the user to whom the "rewritten content" is displayed. In some embodiments, the impact of user information on rewriting can be considered, and user characteristics can be incorporated into the rewriting model as input parameters to make the rewriting results more targeted. Subsequent user interactions after viewing the "rewritten content," such as exposure without clicks, clicks, likes, and shares, can serve as user interaction data. After collecting data on the rewritten content and user interactions with it, it can be used as training and optimization samples for subsequent evaluation models.

[0051] Optionally, the interactive behavior data includes at least one of the following:

[0052] Behavior data of not clicking on the rewritten data, behavior data of clicking on the rewritten data, behavior data of liking the rewritten data, and behavior data of sharing the rewritten data.

[0053] In other words, the collected user interaction behavior data can include behavior data on users who did not click on the rewritten content (i.e., exposure but not click), as well as behavior data on users who clicked, liked, and shared the rewritten content. Rewritten content that was exposed but not clicked can serve as negative samples, while rewritten content that was clicked, liked, and shared can serve as positive samples. By collecting comprehensive user interaction behavior data as training samples for the evaluation model, the evaluation accuracy of the model can be improved.

[0054] After obtaining the rewritten sample data, the rewritten sample data can be cleaned to obtain cleaned samples. The most basic samples include negative samples that were exposed but not clicked and positive samples that were clicked by users.

[0055] Finally, the cleaned rewritten sample data can be used to train an evaluation model, using the rewritten content as the model input and interactive behavior as the model output. Through iterative training using a large amount of sample data, an evaluation model can be developed that predicts user interactive behavior based on the model rewrite results. This evaluation model can estimate the probability of user interaction with the "rewritten content," such as click-through rate (CTR). Therefore, this evaluation model can also be called a CTR model. The CTR model training process is shown in Figure 3. The evaluation model can be a regression model, a deep learning model, a machine learning model, or any other model type; there are no specific model types required. This evaluation model can be defined as r.

[0056] It should be noted that the above-mentioned evaluation model, in addition to being a click-through rate model, can also be a single-objective or multi-objective model trained using any user interaction behavior. For example, interaction behavior can include data such as likes and shares.

[0057] Through this implementation, the accuracy of the evaluation model can be guaranteed, and the effectiveness of the evaluation model in evaluating the rewriting results of the rewriting model can be guaranteed, which in turn helps to accurately optimize the rewriting model.

[0058] Step 104: Based on the interaction information, adjust the parameters of the first rewriting model to obtain a second rewriting model.

[0059] In an embodiment of the present application, the rewriting model can be optimized based on the evaluation results so that the rewriting results of the optimized rewriting model are more attractive to user interaction. Specifically, the interaction information can be used to guide appropriate adjustments to the parameters of the first rewriting model. For example, if the interaction information predicts a low probability of the user interacting with the second information, it indicates that the current rewriting result is not very attractive to the user, and the model parameters need to be optimized and fine-tuned. After the adjustment, an updated rewriting model, i.e., the second rewriting model, can be obtained.

[0060] In the embodiment of the present application, the first rewriting model refers to the rewriting model before optimization and adjustment, and the second rewriting model refers to the rewriting model after optimization and adjustment.

[0061] Optionally, step 102 includes:

[0062] Inputting the first information into a first rewriting model, and selecting N rewriting results output by the first rewriting model using a beam search strategy, wherein the second information includes the N rewriting results, where N is a positive integer greater than 1;

[0063] The step 103 includes:

[0064] Performing user interaction behavior prediction on each of the N rewriting results to obtain N pieces of interaction information predicted for each of the N rewriting results;

[0065] The step 104 includes:

[0066] Based on the N pieces of interaction information, parameters of the first rewriting model are adjusted.

[0067] In some embodiments, during the training or fine-tuning of the rewriting model, a portion of the original content x exposed to the user online, i.e., the first information, can be randomly sampled, and after passing through the rewriting model π, the top N rewriting results y output by the model can be selected using the BeamSearch strategy. i , that is, the N most probable rewriting results, the whole is represented by y i =π(x,u), where i = 1, 2, ..., n. Here, the data set composed of (x, y, u) can be defined as D PPO .

[0068] In this way, the user interaction behavior prediction can be performed on the N rewriting results respectively during the prediction, that is, for each y i The estimated user interaction probability, such as click-through rate, can be expressed as r(y i ,u). Now the optimization goal of the entire model becomes to optimize the rewriting model π so that the output result y of π corresponds to the estimated click rate r(y,u) score as high as possible. Therefore, we can i The corresponding estimated click rate r(y i ,u), optimize and adjust the parameters of the rewritten model π.

[0069] The training process of the above rewritten model can be shown in Figure 4.

[0070] Through this implementation, the training optimization data of the rewriting model can be enriched to ensure the optimized training effect of the rewriting model.

[0071] Optionally, step 104 includes:

[0072] Calculating a loss value of the first rewriting model using a proximal policy optimization (PPO) algorithm based on the interaction information, the first rewriting model, and an initial model of the first rewriting model;

[0073] Based on the loss value, parameters of the first rewriting model are adjusted.

[0074] In some embodiments, the interaction information, the first rewriting model and the initial model of the first rewriting model can be combined to calculate the model loss value using the PPO algorithm, wherein the initial model of the first rewriting model is the initial unadjusted model. For example, if it is a large language model, π 0 Specifically, a first loss value can be calculated based on the interaction information, the first rewriting model, and the initial model. A second loss value can be calculated based on the first rewriting model. The two loss values ​​are then combined in a weighted manner to form a total loss value. Finally, the parameters of the first rewriting model can be adjusted based on the loss values.

[0075] Thus, in this embodiment, by adopting the PPO strategy to optimize and adjust the rewriting model and limiting the amplitude of the strategy update, it is possible to ensure that the algorithm maintains high sample efficiency while also ensuring good stability.

[0076] Optionally, the interaction information includes an interaction probability, and the loss value includes a first part and a second part, the first part being the expected value of the difference between the interaction probability and the first value in the first data set, the second part being the product of the expected value of the logarithm of the first rewritten model in the second data set and a first preset hyperparameter γ, and the first value being the product of the logarithm of the ratio of the first rewritten model to the initial model and a second preset hyperparameter β;

[0077] The first data set includes the first information and the second information, the second data set includes third information and fourth information, and the fourth information is information obtained by manually rewriting the third information.

[0078] In some embodiments, the loss value of the first rewriting model can be calculated using the following loss formula 1, that is, the loss of the optimized π model is defined as:

[0079] Among them, β and γ are hyperparameters ranging from 0 to 1, and D sft It is a high-quality rewritten dataset, that is, manually rewritten data to avoid model deviation, π 0 It is a fixed initialization rewrite model, DPPO It is a data set composed of (x, y, u), which includes the original content, user information and rewritten content. r(y, u) is the estimated click rate of the rewritten content. Indicates that D PPO The expected value in the dataset, Indicates that in D sft The expected value in the dataset. The evaluation model used in the above formula incorporates user features u and fully utilizes user interaction behavior to optimize the rewritten model π.

[0080] Through this implementation, the model loss value can be accurately calculated based on the above loss formula, thereby ensuring the fine-tuning optimization effect of the rewritten model based on the loss value.

[0081] It should be noted that in the embodiment of the present application, the optimization and adjustment of the rewriting model can be a continuous process, that is, the click-through rate evaluation model can be continuously used to predict the model rewriting results, and the rewriting model can be optimized and adjusted based on the predicted results to continuously improve the click-through rate of the model rewriting results.

[0082] This embodiment of the application proposes a method for implementing large-scale model rewriting to improve recommendation click-through rate. It uses massive user interaction behaviors to optimize the large-scale model rewriting results, reducing the cost of manual annotation and the existing effect deviation, and aligning the rewriting goals with the goal of improving click-through rate. The main improvements are as follows:

[0083] 1) Use a click-through rate prediction model to evaluate the rewriting effect of the large model and reduce the bias of manual evaluation.

[0084] 2) Use the PPO training method to align the rewriting results of the large model and the evaluation results of the click-through rate evaluation model, so that the rewriting target and the click-through rate target remain consistent.

[0085] 3) Use online user interaction behavior to continuously optimize and train the click-through rate evaluation model, thereby driving the continuous optimization of the rewrite model. Make full use of user interaction behavior data to optimize the model process and combine it with user profile data to achieve a virtuous cycle of the entire chain.

[0086] The model adjustment method in the embodiments of the present application obtains first information; rewrites the first information using a first rewriting model to obtain rewritten second information; predicts user interaction behavior based on the second information to obtain predicted interaction information; and adjusts the parameters of the first rewriting model based on the interaction information to obtain a second rewriting model. Thus, by predicting user interaction behavior based on the information rewritten by the rewriting model, an evaluation mechanism for the rewriting results is introduced, and the rewriting model is optimized and adjusted based on the evaluation results, thereby improving the model rewriting effect.

[0087] The model adjustment method provided in the embodiment of the present application can be executed by a model adjustment device. In the embodiment of the present application, the model adjustment device provided in the embodiment of the present application is described by taking the execution of the model adjustment method by the model adjustment device as an example.

[0088] Please refer to FIG5 , which is a schematic diagram of the structure of a model adjustment device provided in an embodiment of the present application. As shown in FIG5 , the model adjustment device 500 includes:

[0089] An acquisition module 501 is configured to acquire first information;

[0090] A rewriting module 502, configured to rewrite the first information using a first rewriting model to obtain rewritten second information;

[0091] Prediction module 503, configured to predict user interaction behavior based on the second information to obtain predicted interaction information;

[0092] The adjustment module 504 is configured to adjust the parameters of the first rewriting model based on the interaction information to obtain a second rewriting model.

[0093] Optionally, the prediction module 503 is used to predict user interaction behavior on the second information using an evaluation model, wherein the evaluation model is trained using rewritten sample data, and the rewritten sample data includes rewritten data rewritten by the first rewriting model and user interaction behavior data on the rewritten data.

[0094] Optionally, the adjustment module 504 includes:

[0095] a calculation unit, configured to calculate a loss value of the first rewriting model using a proximal policy optimization (PPO) algorithm based on the interaction information, the first rewriting model, and an initial model of the first rewriting model;

[0096] An adjusting unit is used to adjust the parameters of the first rewriting model based on the loss value.

[0097] Optionally, the interaction information includes an interaction probability, and the loss value includes a first part and a second part, the first part being an expected value of a difference between the interaction probability and the first value in the first data set, the second part being a product of an expected value of a logarithm of the first rewritten model in the second data set and a first preset hyperparameter, and the first value being a product of a logarithm of a ratio of the first rewritten model to the initial model and a second preset hyperparameter;

[0098] The first data set includes the first information and the second information, the second data set includes third information and fourth information, and the fourth information is information obtained by manually rewriting the third information.

[0099] Optionally, the interactive behavior data includes at least one of the following:

[0100] Behavior data of not clicking on the rewritten data, behavior data of clicking on the rewritten data, behavior data of liking the rewritten data, and behavior data of sharing the rewritten data.

[0101] Optionally, the rewriting module 502 is configured to input the first information into a first rewriting model, and select N rewriting results output by the first rewriting model using a beam search strategy, wherein the second information includes the N rewriting results, where N is a positive integer greater than 1.

[0102] The prediction module 503 is used to predict user interaction behaviors for each of the N rewriting results, and obtain N pieces of interaction information predicted for each of the N rewriting results;

[0103] The adjustment module 504 is configured to adjust parameters of the first rewriting model based on the N pieces of interaction information.

[0104] Optionally, the first information is the title of the recommended content.

[0105] Optionally, the rewriting module 502 is configured to input the first information, user information, and usage scenario information into the first rewriting model, and rewrite the first information by fusing the first information, the user information, and the usage scenario information through the first rewriting model.

[0106] The model adjustment device 500 in the embodiment of the present application obtains first information; rewrites the first information using a first rewriting model to obtain rewritten second information; predicts user interaction behavior based on the second information to obtain predicted interaction information; and adjusts the parameters of the first rewriting model based on the interaction information to obtain a second rewriting model. In this way, by predicting user interaction behavior based on the information rewritten by the rewriting model, an evaluation mechanism for the rewriting results is introduced, and the rewriting model is optimized and adjusted based on the evaluation results, the rewriting effect of the model can be improved.

[0107] The model adjustment device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or a device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not make specific limitations.

[0108] The model adjustment device in the embodiment of the present application can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0109] The model adjustment device provided in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 1 to 4, and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0110] Optionally, as shown in Figure 6, an embodiment of the present application also provides an electronic device 600, including a processor 601 and a memory 602, and the memory 602 stores a program or instruction that can be run on the processor 601. When the program or instruction is executed by the processor 601, the various steps of the above-mentioned model adjustment method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0111] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0112] FIG7 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0113] The electronic device 700 includes but is not limited to components such as a radio frequency unit 701 , a network module 702 , an audio output unit 703 , an input unit 704 , a sensor 705 , a display unit 706 , a user input unit 707 , an interface unit 708 , a memory 709 , and a processor 710 .

[0114] Those skilled in the art will appreciate that the electronic device 700 may further include a power source (e.g., a battery) to power various components. The power source may be logically connected to the processor 710 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The electronic device structure shown in FIG7 does not limit the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components or arrange the components differently, which will not be described in detail here.

[0115] The processor 710 is configured to:

[0116] Obtaining first information;

[0117] rewriting the first information using a first rewriting model to obtain rewritten second information;

[0118] Predicting user interaction behavior based on the second information to obtain predicted interaction information;

[0119] Based on the interaction information, the parameters of the first rewriting model are adjusted to obtain a second rewriting model.

[0120] Optionally, the processor 710 is further used to predict user interaction behavior on the second information using an evaluation model, wherein the evaluation model is trained using rewritten sample data, and the rewritten sample data includes rewritten data rewritten by the first rewriting model and user interaction behavior data on the rewritten data.

[0121] Optionally, the processor 710 is further configured to:

[0122] Calculating a loss value of the first rewriting model using a proximal policy optimization (PPO) algorithm based on the interaction information, the first rewriting model, and an initial model of the first rewriting model;

[0123] Based on the loss value, parameters of the first rewriting model are adjusted.

[0124] Optionally, the interaction information includes an interaction probability, and the loss value includes a first part and a second part, the first part being an expected value of a difference between the interaction probability and the first value in the first data set, the second part being a product of an expected value of a logarithm of the first rewritten model in the second data set and a first preset hyperparameter, and the first value being a product of a logarithm of a ratio of the first rewritten model to the initial model and a second preset hyperparameter;

[0125] The first data set includes the first information and the second information, the second data set includes third information and fourth information, and the fourth information is information obtained by manually rewriting the third information.

[0126] Optionally, the interactive behavior data includes at least one of the following:

[0127] Behavior data of not clicking on the rewritten data, behavior data of clicking on the rewritten data, behavior data of liking the rewritten data, and behavior data of sharing the rewritten data.

[0128] Optionally, the processor 710 is further configured to:

[0129] Inputting the first information into a first rewriting model, and selecting N rewriting results output by the first rewriting model using a beam search strategy, wherein the second information includes the N rewriting results, where N is a positive integer greater than 1;

[0130] Performing user interaction behavior prediction on each of the N rewriting results to obtain N pieces of interaction information predicted for each of the N rewriting results;

[0131] Based on the N pieces of interaction information, parameters of the first rewriting model are adjusted.

[0132] Optionally, the first information is the title of the recommended content.

[0133] Optionally, the processor 710 is further configured to input the first information, user information, and usage scenario information into the first rewriting model, and rewrite the first information by fusing the first information, the user information, and the usage scenario information through the first rewriting model.

[0134] The electronic device 700 in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 1 to 4 and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0135] It should be understood that in an embodiment of the present application, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042, and the graphics processor 7041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes a touch panel 7071 and at least one of other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0136] The memory 709 can be used to store software programs and various data. The memory 709 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 may include a volatile memory or a non-volatile memory, or the memory 709 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 709 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0137] Processor 710 may include one or more processing units. Optionally, processor 710 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 710.

[0138] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned model adjustment method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0139] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0140] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned model adjustment method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0141] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0142] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned model adjustment method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0143] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0144] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0145] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A model adjustment method, comprising: Obtaining first information; rewriting the first information using a first rewriting model to obtain rewritten second information; Predicting user interaction behavior based on the second information to obtain predicted interaction information; Based on the interaction information, the parameters of the first rewriting model are adjusted to obtain a second rewriting model.

2. The method according to claim 1, wherein The adjusting the parameters of the first rewriting model based on the interaction information includes: Calculating a loss value of the first rewriting model using a proximal policy optimization (PPO) algorithm based on the interaction information, the first rewriting model, and an initial model of the first rewriting model; Based on the loss value, parameters of the first rewriting model are adjusted.

3. The method according to claim 2, wherein: The interaction information includes an interaction probability, and the loss value includes a first part and a second part, the first part being the expected value of the difference between the interaction probability and the first value in the first data set, the second part being the product of the expected value of the logarithm of the first rewritten model in the second data set and a first preset hyperparameter, and the first value being the product of the logarithm of the ratio of the first rewritten model to the initial model and a second preset hyperparameter; The first data set includes the first information and the second information, the second data set includes third information and fourth information, and the fourth information is information obtained by manually rewriting the third information.

4. The method according to any one of claims 1 to 3, wherein The step of rewriting the first information using the first rewriting model to obtain rewritten second information includes: Inputting the first information into a first rewriting model, and selecting N rewriting results output by the first rewriting model using a beam search strategy, wherein the second information includes the N rewriting results, where N is a positive integer greater than 1; The step of predicting user interaction behavior based on the second information to obtain predicted interaction information includes: Performing user interaction behavior prediction on each of the N rewriting results to obtain N pieces of interaction information predicted for each of the N rewriting results; The adjusting the parameters of the first rewriting model based on the interaction information includes: Based on the N pieces of interaction information, parameters of the first rewriting model are adjusted.

5. The method according to any one of claims 1 to 3, wherein The rewriting of the first information by using the first rewriting model includes: The first information, user information and usage scenario information are input into the first rewriting model, and the first information, the user information and the usage scenario information are integrated through the first rewriting model to rewrite the first information.

6. A model adjustment device comprising: An acquisition module, configured to acquire first information; a rewriting module, configured to rewrite the first information using a first rewriting model to obtain rewritten second information; a prediction module, configured to predict user interaction behavior based on the second information and obtain predicted interaction information; An adjustment module is used to adjust the parameters of the first rewriting model based on the interaction information to obtain a second rewriting model.

7. The model adjustment device according to claim 6, wherein: The adjustment module includes: a calculation unit, configured to calculate a loss value of the first rewriting model using a proximal policy optimization (PPO) algorithm based on the interaction information, the first rewriting model, and an initial model of the first rewriting model; An adjusting unit is used to adjust the parameters of the first rewriting model based on the loss value.

8. The model adjustment device according to claim 7, wherein: The interaction information includes an interaction probability, and the loss value includes a first part and a second part, the first part being the expected value of the difference between the interaction probability and the first value in the first data set, the second part being the product of the expected value of the logarithm of the first rewritten model in the second data set and a first preset hyperparameter, and the first value being the product of the logarithm of the ratio of the first rewritten model to the initial model and a second preset hyperparameter; The first data set includes the first information and the second information, the second data set includes third information and fourth information, and the fourth information is information obtained by manually rewriting the third information.

9. The model adjustment device according to any one of claims 6 to 8, wherein: The rewriting module is configured to input the first information into a first rewriting model, and select N rewriting results output by the first rewriting model using a beam search strategy, wherein the second information includes the N rewriting results, where N is a positive integer greater than 1; The prediction module is used to predict user interaction behaviors for the N rewriting results respectively, and obtain N interaction information predicted for the N rewriting results respectively; The adjustment module is used to adjust the parameters of the first rewriting model based on the N pieces of interaction information.

10. The model adjustment device according to any one of claims 6 to 8, wherein: The rewriting module is used to input the first information, user information and usage scenario information into the first rewriting model, and rewrite the first information by fusing the first information, the user information and the usage scenario information through the first rewriting model.

11. An electronic device comprising a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

12. A readable storage medium storing a program or instruction, wherein the program or instruction is executed by a processor to implement the steps of the method according to any one of claims 1 to 5.

13. A chip comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps of the method according to any one of claims 1 to 5.

14. A computer program product, wherein the program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method according to any one of claims 1 to 5.

15. An electronic device, configured to perform the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Query rewriting service optimization method and related device

    CN116955410A

  • Text rewriting method and device, terminal equipment and storage medium

    CN117172229A

  • Method and device for model training, equipment and storage medium

    CN117689003A

  • Training method and query method of query word rewriting model and related products

    CN117743505A

  • Model adjustment method and device

    CN118260395A