Language understanding model training method and device, electronic equipment, computer readable storage medium and computer program product

CN122839006APending Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365261.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

相关技术常常通过公开的训练集或是人工撰写的互动信息训练语言理解模型,模型在训练过程中针对不同的互动信息使用相同的训练方式,缺乏针对重要信息的侧重;同时,针对不同对象使用相同的互动信息进行训练,导致训练后的语言理解模型输出的互动信息内容单一,缺乏针对性

Benefits of technology

[0023]在训练第一语言理解模型,得到第二语言理解模型时,通过将基于推荐信息的评价指标确定的第一权重与第一互动信息的第一特征向量进行融合,在第二特征向量中融入了评价指标,基于第一互动信息和第二特征向量训练第一语言理解模型,第一权重使得第一语言理解模型在学习第一互动信息的第一特征向量时,既能够全面地基于第一互动信息进行学习,也能够通过评价指标对不同的第一互动有所侧重,增强了模型针对高评价指标对应的互动信息的关注和理解,提高了模型训练的全面性和针对性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839006A_ABST
    Figure CN122839006A_ABST
Patent Text Reader

Abstract

This application provides a training method, apparatus, electronic device, computer-readable storage medium, and computer program product for a language understanding model. The method includes: determining a first weight of first interactive information based on evaluation metrics of recommendation information, wherein the first interactive information represents the interaction performed in response to the recommendation information; fusing the first weight and a first feature vector of the first interactive information to obtain a second feature vector of the first interactive information; training a first language understanding model based on the first interactive information and the second feature vector to obtain a second language understanding model; generating second interactive information by calling the second language understanding model based on the recommendation information and object feature information of the object to be recommended; and training the second language understanding model based on the second interactive information and feedback information of the object to be recommended in response to the second interactive information to obtain a third language understanding model. This application improves the relevance of the trained language understanding model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for training a language understanding model. Background Technology

[0002] Against the backdrop of rapid development in digital advertising, interactive messages related to recommendations are gradually becoming crucial touchpoints for brands to enhance user engagement and drive marketing conversions. The demand for diversity, marketing orientation, and personalization in interactive messages is becoming increasingly urgent. They need to maintain a high degree of alignment with different user preferences while simultaneously considering creativity, interactivity, and commercial value. Related technologies often train language understanding models using publicly available training sets or manually written interactive messages. During training, the same training methods are used for different interactive messages, lacking focus on key information. Furthermore, using the same interactive messages for different users results in the trained language understanding models outputting monotonous and untargeted interactive messages. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for training a language understanding model, which can improve the relevance of the trained language understanding model.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a method for training a language understanding model, the method comprising:

[0006] Based on the evaluation metrics of the recommendation information, a first weight of the first interactive information is determined, wherein the first interactive information represents the interaction performed in response to the recommendation information.

[0007] The first weight and the first feature vector of the first interaction information are fused to obtain the second feature vector of the first interaction information;

[0008] A first language understanding model is trained based on the first interaction information and the second feature vector to obtain a second language understanding model;

[0009] Based on the recommendation information and the object feature information of the object to be recommended, the second language understanding model is invoked to generate the second interactive information;

[0010] Based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, the second language understanding model is trained to obtain the third language understanding model.

[0011] This application provides a training device for a language understanding model, the device comprising:

[0012] The weight determination module is used to determine the first weight of the first interactive information based on the evaluation index of the recommendation information, wherein the first interactive information represents the interaction performed in response to the recommendation information.

[0013] The data fusion module is used to fuse the first weight and the first feature vector of the first interaction information to obtain the second feature vector of the first interaction information.

[0014] The first training module trains a first language understanding model based on the first interaction information and the second feature vector to obtain a second language understanding model.

[0015] The information generation module is used to generate second interactive information by calling the second language understanding model based on the recommendation information and the object feature information of the object to be recommended;

[0016] The second training module is used to train the second language understanding model based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, so as to obtain the third language understanding model.

[0017] This application provides an electronic device, the electronic device comprising:

[0018] Memory is used to store executable instructions or computer programs.

[0019] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the training method for the language understanding model provided in the embodiments of this application.

[0020] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the training method of the language understanding model provided in this application when executed by a processor.

[0021] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the training method for the language understanding model provided in this application.

[0022] The embodiments of this application have the following beneficial effects:

[0023] When training the first language understanding model to obtain the second language understanding model, the first weight determined by the evaluation index based on recommendation information is fused with the first feature vector of the first interaction information, and the evaluation index is incorporated into the second feature vector. The first language understanding model is trained based on the first interaction information and the second feature vector. The first weight enables the first language understanding model to learn comprehensively based on the first interaction information when learning the first feature vector of the first interaction information, and also to focus on different first interactions through the evaluation index. This enhances the model's attention to and understanding of the interaction information corresponding to high evaluation indexes, and improves the comprehensiveness and relevance of model training.

[0024] When training a second language understanding model to obtain a third language understanding model, second interaction information is generated based on recommendation information and object feature information of the object to be recommended. The second language understanding model is then trained based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information. This allows the trained third language understanding model to simultaneously consider object feature information and the feedback information of the object to be recommended in response to the second interaction information. The feedback information reflects the preference of the object to be recommended in response to the second interaction information. By training the model with feedback information, the model can be continuously adjusted to align with object preferences, thereby improving the third language understanding model's relevance and generalization ability for different objects to be recommended. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the architecture of the language understanding model training system 100 provided in the embodiments of this application;

[0026] Figure 2 This is a schematic diagram of the structure of the server 200 provided in the embodiments of this application;

[0027] Figure 3A This is a schematic diagram of the first process of the training method for the language understanding model provided in the embodiments of this application;

[0028] Figure 3B This is a schematic diagram of the second process of the training method for the language understanding model provided in the embodiments of this application;

[0029] Figure 3C This is a schematic diagram of the third process of the training method for the language understanding model provided in the embodiments of this application;

[0030] Figure 3D This is a schematic diagram of the fourth process of the training method for the language understanding model provided in the embodiments of this application;

[0031] Figure 3E This is a schematic diagram of the fifth step in the training method of the language understanding model provided in the embodiments of this application;

[0032] Figure 3F This is a schematic diagram of the sixth step in the training method of the language understanding model provided in the embodiments of this application;

[0033] Figure 3G This is a schematic diagram of the seventh step in the training method of the language understanding model provided in the embodiments of this application;

[0034] Figure 3H This is the eighth flowchart of the training method for the language understanding model provided in the embodiments of this application;

[0035] Figure 3I This is a ninth flowchart illustrating the training method for the language understanding model provided in this application embodiment;

[0036] Figure 3J This is a schematic diagram of the tenth step of the training method for the language understanding model provided in this application embodiment;

[0037] Figure 3K This is a schematic diagram of the eleventh step in the training method of the language understanding model provided in the embodiments of this application;

[0038] Figure 3L This is a schematic diagram of the twelfth step of the training method for the language understanding model provided in the embodiments of this application;

[0039] Figure 3M This is a schematic diagram of the thirteenth step of the training method for the language understanding model provided in the embodiments of this application;

[0040] Figure 3N This is the fourteenth flowchart of the training method for the language understanding model provided in the embodiments of this application;

[0041] Figure 3O This is the fifteenth flowchart illustrating the training method for the language understanding model provided in this application embodiment;

[0042] Figure 3P This is the sixteenth flowchart of the training method for the language understanding model provided in the embodiments of this application;

[0043] Figure 4A This is a schematic diagram illustrating the principle of the training classification model provided in the embodiments of this application;

[0044] Figure 4B This is a schematic diagram illustrating the principle of training a first language understanding model provided in an embodiment of this application;

[0045] Figure 4C This is a schematic diagram illustrating the principle of training a second language understanding model provided in an embodiment of this application;

[0046] Figure 5A This is a schematic diagram of the first process for generating advertising copy using the related technologies provided in this application embodiment;

[0047] Figure 5B This is a schematic diagram of the second process for generating advertising copy using the related technologies provided in the embodiments of this application;

[0048] Figure 6 This is an application illustration of displaying advertisement reviews provided in an embodiment of this application;

[0049] Figure 7 This is a schematic diagram of the process for generating advertising reviews that conform to user preferences, provided in an embodiment of this application.

[0050] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0053] In the following description, the terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0054] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0055] Unless otherwise specified, "at least one" as used below refers to one or more cases, and "multiple" can refer to two or more cases.

[0056] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0057] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0058] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0059] 1) Responding to: used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0060] 2) Human-computer interaction interface, which is used to provide human-computer interaction functions / display recommendation information and interactive information.

[0061] For example, graphical user interfaces (GUIs) include augmented reality (AR) interfaces, virtual reality (VR) interfaces, voice user interfaces (VUIs), interactive projection interfaces (using projection technology to display information on a flat surface), eye-tracking interfaces (interfaces controlled by detecting the user's gaze), holographic interfaces (three-dimensional holograms formed by projecting images using holographic projection technology, allowing users to see stereoscopic images without wearing special glasses), multimodal interfaces (interfaces that combine multiple interaction methods, such as tactile, visual, and auditory interaction), and brain-machine interfaces (BMIs).

[0062] 3) Recommended information refers to any type of content or service provided by the recommendation system to users, aiming to meet their interests and needs. This can include, but is not limited to, articles, products, services, videos, music, and news. Taking advertising recommendation scenarios as an example, recommended information can be the description of advertisements.

[0063] 4) Evaluation metrics are quantitative standards used to assess the quality and effectiveness of recommendation information. Evaluation metrics help the recommendation system understand whether the recommendation information effectively attracts users and how users react to it. Common evaluation metrics include:

[0064] Click-Through Rate (CTR): The ratio of the number of times a user clicks on a recommendation to the number of times the recommendation is displayed.

[0065] Conversion Rate (CVR): The ratio of the number of times a user completes the expected action (such as purchase or registration) after clicking on a recommendation to the number of clicks.

[0066] Cost Per Mille (CPM): The cost per thousand impressions of recommended information.

[0067] Comment reply count (Reply, R): The number of replies from users to interactive messages generated from recommended information.

[0068] Negative Feedback Rate (NF): The ratio of the number of times a user expresses dislike or negative feedback on a recommendation to the number of times the recommendation is displayed.

[0069] 5) Interactive information refers to content generated by the publisher of the recommended information or the recommendation system, designed to communicate and interact with users. Interactive information can be text, images, videos, or other forms of content. It is typically designed to encourage user participation, increase user engagement, or guide users to take specific actions. Interactive information can be direct invitations, questions, prompts, or promotional information, aiming to stimulate user interest and prompt users to respond to the recommended information, such as by clicking, commenting, or sharing.

[0070] 6) The user to be recommended refers to the user who receives and interacts with recommendation information in the recommendation system. The user can be an individual, group, or organization, and interacts with the recommendation system through various channels (such as websites, applications, social media, etc.). The characteristics and behavioral data of the user to be recommended are used for personalized recommendations to improve the relevance and effectiveness of the recommendations.

[0071] The interactive behaviors of the target user in response to the recommendation information include, but are not limited to:

[0072] Engage actively: Click on recommended information, read content, watch videos, purchase products, participate in discussions, comment, like, share, etc.

[0073] Negative interactions include ignoring recommendations, turning off recommendations, marking content as uninteresting, and reporting content.

[0074] 7) Object feature information refers to data describing the attributes and behaviors of the object to be recommended. Object feature information helps the recommendation system understand user needs and preferences, thereby providing more personalized recommendations. Object feature information includes, but is not limited to, at least one of the following:

[0075] Statistical characteristics: age, gender, education level, geographical location, etc.

[0076] Behavioral characteristics: browsing history, search history, purchase behavior, click preferences, dwell time, conversion categories, etc.

[0077] Psychological characteristics: risk preference, price sensitivity, brand loyalty, interests, values, etc.

[0078] 8) Feedback information refers to the reactions and evaluations of the target user to the recommendation information and interaction information. Feedback information helps the recommendation system understand the effectiveness of the recommendation interaction information and user needs, thereby optimizing the recommendation algorithm. Feedback information is divided into positive and negative feedback. Taking the advertising scenario as an example, if a user replies to the interaction information with a positive comment, it represents positive feedback; if a user closes or complains about the advertisement, it represents negative feedback.

[0079] 9) Emotional polarity refers to the quantitative representation of the emotional state or attitude contained in feedback information. It reflects a user's subjective feelings and evaluations of a topic, product, service, event, etc. Emotional polarity is usually divided into the following categories:

[0080] Positive Sentiment: Represents positive emotions such as positivity, satisfaction, liking, and appreciation.

[0081] Negative Sentiment: Indicates negative emotions such as dissatisfaction, dislike, and criticism.

[0082] Neutral Sentiment: This refers to a statement of fact or a neutral viewpoint that does not carry a strong emotional bias.

[0083] 10) A reference point refers to one or more benchmarks used to compare and adjust model outputs to better match user preferences. A reference point is a benchmark used to measure changes in model performance, representing the average probability difference between the interactive information generated by the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration.

[0084] 11) The value function represents the cumulative reward expected to be obtained after taking a certain action in a given state. In machine learning, the value function can be used to characterize the quality of interactive information generated by the model. The value function is used to characterize the quality of interactive information generated by the second language understanding model in the nth iteration update.

[0085] Related technologies often train language understanding models using publicly available training sets or manually written interactive information, resulting in the output of interactive information from the trained language understanding models being monotonous and lacking interactivity and personalization.

[0086] Based on the above analysis, the applicant found that the training methods of language understanding models in related technologies cannot output personalized interactive information for different users after training. In order to address the above problems, this application provides a training method for language understanding models, which can improve the relevance of the trained language understanding models.

[0087] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. The following will describe exemplary applications when the electronic device is implemented as a server.

[0088] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the language understanding model training system 100 provided in the embodiments of this application. In order to support the training application of a language understanding model, the terminal (terminal 400-1 and terminal 400-2 are shown as examples) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0089] Server 200 is used to train the first language understanding model multiple times to obtain the third language understanding model, so as to generate interactive information related to the object to be recommended, and send it to terminal 400 for display on the human-computer interaction interface 410 of terminal 400.

[0090] The language understanding model training method provided in this application can be applied to various scenarios that require the generation of interactive information through the language understanding model, such as advertising recommendations and intelligent customer service. Examples are given below.

[0091] Taking an advertising recommendation scenario as an example, the terminal sends advertising information (recommendation information) and user (the object to be recommended) feature information (object feature information). The server generates corresponding interactive information based on the trained third language understanding model and sends it to the terminal.

[0092] Taking a game scenario as an example, the terminal sends information such as newly launched virtual props and virtual clothing, as well as the user's characteristic information, to the server. The server generates corresponding interactive information based on the trained third language understanding model and feeds it back to the user through the terminal to attract the user to purchase virtual props or virtual clothing.

[0093] Taking intelligent customer service scenarios as an example, the terminal sends questions or messages entered by users through websites or mobile applications. The server generates corresponding interactive information based on a trained third-language understanding model and feeds it back to the user through the terminal, thus realizing human-computer dialogue.

[0094] Taking a multimedia scenario as an example, the terminal sends descriptive text for the multimedia content, and the server generates interactive information corresponding to the descriptive text based on a trained third-language understanding model, and then feeds it back to the user through the terminal.

[0095] Taking the education scenario as an example, the terminal sends learning materials or course content, as well as student information, to the server. The server generates interactive information corresponding to the descriptive text based on the trained third language understanding model, and then feeds it back to the student through the terminal.

[0096] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0097] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the server 200 provided in the embodiments of this application. Figure 2 The server 200 shown includes at least one processor 210, memory 230, and at least one network interface 220. The various components of server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 240.

[0098] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0099] The memory 230 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 230 may optionally include one or more storage devices physically located away from the processor 210.

[0100] The memory 230 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 230 described in this application embodiment is intended to include any suitable type of memory.

[0101] In some embodiments, memory 230 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0102] Operating system 231 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0103] The network communication module 232 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0104] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A training device 233 for a language understanding model stored in memory 230 is shown. This device can be software in the form of programs or plugins, and includes the following software modules: a weight determination module 2331, a data fusion module 2332, a first training module 2333, an information generation module 2334, and a second training module 2335. These modules are logically connected and can therefore be arbitrarily combined or further divided according to their implemented functions. The functions of each module will be described below.

[0105] In some embodiments, the terminal or server can implement the language understanding model training method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as social media APPs or instant messaging APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0106] The training method of the language understanding model provided in this application will be described in conjunction with exemplary applications and implementations of the server provided in the embodiments of this application.

[0107] See Figure 3A , Figure 3A This is a schematic diagram of the first flowchart of the training method for the language understanding model provided in this application embodiment, with the server as the main body, and combining... Figure 3A The steps shown are explained.

[0108] In step 101, a first weight of the first interactive information is determined based on the evaluation index of the recommendation information, wherein the first interactive information represents the interaction performed in response to the recommendation information.

[0109] In some embodiments, the evaluation metric can be the feedback data collected from the target audience regarding the recommendation information after the recommendation information and the first interactive information are displayed on the corresponding platform or webpage. The evaluation metric is then extracted from the feedback data and used to evaluate the dissemination effect of the recommendation information. The first weight represents the degree of influence of the first interactive information on the evaluation metric. The first interactive information can be displayed in the recommendation area of ​​the recommendation information, or it can be displayed in the recommendation information in the form of bullet comments or cards. The display position and method of the first interactive information are not limited here.

[0110] For example, in an advertising scenario, the recommended information can be an advertisement, and the first interactive information can be an ad review, used to describe or promote the ad in the ad's review section. Evaluation metrics can include at least one of the following: click-through rate (CTR), conversion rate (CVR), cost per mille (CPM), number of comments replied (R), and negative feedback rate (NF).

[0111] In some embodiments, see Figure 3B , Figure 3BThis is a schematic diagram of the second process of the training method for the language understanding model provided in the embodiments of this application. Figure 3A Step 101, "Determine the first weight of the first interaction information based on the evaluation metrics of the recommendation information," can be used when there are multiple evaluation metrics. Figure 3B Steps 1011 to 1012 are implemented, and the details are explained below.

[0112] In step 1011, a second weight is determined for each evaluation indicator, wherein the second weight characterizes the importance of each evaluation indicator.

[0113] In some embodiments, the second weight can be manually set based on experience values. For example, the importance of multiple evaluation indicators can be determined according to business needs, and different second weights can be assigned to different evaluation indicators in descending order of importance.

[0114] For example, taking click-through rate (CTR), conversion rate (CTR), and cost per thousand impressions (CPM) as evaluation metrics, in an advertising scenario, if multiple evaluation metrics are sorted in descending order of importance, the result of the descending sort can be: CTR, conversion rate, CPM. Therefore, the second weights corresponding to CTR, conversion rate, and CPM can be set to 0.5, 0.3, and 0.2, respectively.

[0115] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the third process of the training method for the language understanding model provided in the embodiments of this application. Figure 3B Step 1011, "Determining the second weight for each evaluation indicator," can be achieved through... Figure 3C Steps 10111 to 10114 are implemented, and the details are explained below.

[0116] In step 10111, for each evaluation index, the initial coefficients of the evaluation index are obtained, and the partial derivatives of the objective function with respect to the evaluation index are determined. The objective function is pre-constructed based on multiple evaluation indexes and the coefficients of each evaluation index.

[0117] In some embodiments, the product of each evaluation index and its coefficient can be determined, and the products of multiple evaluation indices can be summed to obtain the objective function. Partial derivatives can be obtained by taking the partial derivative of the objective function with respect to the initial coefficients of the evaluation indices.

[0118] For example, if the initial coefficients of evaluation index a1 are k1, the initial coefficients of evaluation index a2 are k2, the initial coefficients of evaluation index a3 are k3, and the objective function is F, then the partial derivative of the objective function with respect to evaluation index a1 is: The partial derivative of the objective function with respect to the evaluation index a2 is The partial derivative of the objective function with respect to the evaluation index a3 is

[0119] In step 10112, an iterative function is used to determine the initial coefficients based on the partial derivatives.

[0120] In some embodiments, for each initial coefficient, during the first iteration, the product of a preset learning rate and the partial derivative corresponding to the initial coefficient is determined, and the difference between the initial coefficient and the product is used as the iteration function.

[0121] Following the example in step 10111 above, if the learning rate is μ, then with respect to the initial coefficients k1, the partial derivatives of the learning rate with respect to the initial coefficients k1 are... The product is The iteration function is The partial derivative of the learning rate with respect to the initial coefficients k2. The product is The iteration function is With respect to the initial coefficients k3, the partial derivative of the learning rate with respect to the initial coefficients k3 is... The product is The iteration function is

[0122] In step 10113, the iterative function is iteratively calculated to obtain the coefficients of the evaluation index after each iteration.

[0123] In some embodiments, if there are a total of M iterations, the coefficient after the m-th iteration, where m is a positive integer less than or equal to M, can be obtained by performing the following process: determining the product of the preset learning rate and the partial derivatives corresponding to the coefficients calculated after the (m-1)-th iteration, and taking the difference between the coefficients calculated after the (m-1)-th iteration and the product as the coefficient after the m-th iteration.

[0124] For example, the coefficient k calculated after the (m-1)th iteration. m-1 The corresponding partial derivative is The product of the preset learning rate and the partial derivatives of the coefficients calculated in the (m-1)th iteration is: Determine the difference between the coefficients and the product after the (m-1)th iteration. k m As the coefficient after the m-th iteration.

[0125] In step 10114, when the objective function converges, the coefficient of each evaluation index calculated in the last iteration is used as the second weight of each evaluation index.

[0126] In some embodiments, the objective function can be considered to have converged when any of the following conditions are met: a preset number of iterations is reached, or the change in the objective function is less than a change threshold.

[0127] Following the example of step 10113 above, if the preset number of iterations is M, then when m = M, k m The coefficient of the evaluation index after the last iteration is used as the second weight of the evaluation index; or, if the change threshold is 0.002, and the change between the objective function of the current iteration and the objective function of the previous iteration is 0.001, then the current iteration is taken as the last iteration, and the coefficient of the evaluation index after the last iteration is used as the second weight of the evaluation index.

[0128] In this embodiment, the second weight can be automatically adjusted through an iterative process to optimize the objective function, reducing the time cost of manual adjustment. The weight adjustment is based on the data and the partial derivatives of the objective function, making the weights more objective and data-driven. The iterative process typically exhibits convergence, guaranteeing that the optimal or near-optimal second weight will be found under certain conditions.

[0129] See also Figure 3B In step 1012, multiple evaluation indicators are weighted and fused based on the second weight to obtain the first weight of the first interactive information.

[0130] For example, if the second weight of evaluation index a1 is The second weight of evaluation index a2 is The second weight of evaluation index a3 is The first weight of the first interactive information is obtained by weighting and fusing multiple evaluation indicators based on the second weight.

[0131] This application's embodiments, by weighting and fusing multiple evaluation indicators, can more comprehensively reflect the combined impact of the first interactive information on the evaluation indicators. The second weight can be adjusted according to the importance of different evaluation indicators, making the weight allocation more reasonable and consistent with actual conditions. It can be flexibly adjusted according to different application scenarios and evaluation objectives, thus possessing broad applicability.

[0132] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the training method for the language understanding model provided in this application embodiment. Before step 101 "determine the first weight of the first interaction information based on the evaluation index of the recommendation information", the following steps are executed: Figure 3D Steps 201 to 204 are explained in detail below.

[0133] In step 201, candidate interaction information is selected from multiple original interaction information of the recommendation information according to preset rules.

[0134] Here, the multiple original interactive information in the recommendation information can be manually written interactive information or interactive information generated by a pre-trained language understanding model.

[0135] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the fifth step in the training method of the language understanding model provided in the embodiments of this application. Figure 3D Step 201, "Filtering candidate interaction information from multiple original interaction messages in the recommendation information according to preset rules," can be achieved through... Figure 3E Steps 2011 to 2013 were implemented, and the details are explained below.

[0136] In step 2011, a portion of the original interactive information that meets the first preset condition is selected from multiple original interactive information. The first preset condition includes at least one of the following: the length of the original interactive information is less than a length threshold, the sentence structure of the original interactive information is a preset sentence structure, and the symbols in the original interactive information are compliant.

[0137] For example, if the length threshold is 10, the original interactive information A is "Want to try this apple?", and the length of the original interactive information A is 8. The original interactive information B is "This apple is rotten, it's so bad, we have to throw it away", and the length of the original interactive information B is 21. Then, the original interactive information A is selected as part of the original interactive information. If the preset sentence structure is a question, the original interactive information A is a question, and the original interactive information B is an affirmative sentence. Then, the original interactive information A is selected as part of the original interactive information. If the compliant symbols include: smiley face, heart, flower, and gift symbols, the original interactive information C includes a smiley face, and the original interactive information D includes an angry symbol. Then, the original interactive information C is selected as part of the original interactive information.

[0138] In step 2012, non-compliant words are removed from a portion of the original interactive information.

[0139] In some embodiments, non-compliant words can be identified from a portion of the original interaction information using a pre-trained classification model, and then removed.

[0140] For example, a classification model can be trained by performing the following steps: obtaining an initialized classification model; obtaining word samples and ground truth class labels, where the ground truth class labels represent whether a word sample is an inappropriate word; calling the initialized classification model based on the word samples to obtain predicted class labels; determining the class loss for the ground truth class labels and predicted class labels; updating the parameters of the initialized classification model based on the class loss to obtain a pre-trained classification model.

[0141] For example, see Figure 4A , Figure 4A This is a schematic diagram illustrating the principle of training a classification model provided in this application embodiment. Based on word samples, the convolutional layer of the initialized classification model is used for feature extraction to obtain the feature vector of the word samples. Based on the feature vector of the word samples, the fully connected layer of the initialized classification model is used for classification processing to obtain the predicted class label. The class loss between the true class label and the predicted class label is determined through a loss function. The class loss is backpropagated to update the parameters of the initialized classification model. This process of iteratively calculating the class loss and updating the parameters continues until the global loss converges, at which point the iteration process stops, forming a pre-trained classification model. The classification model can be a Transformer model, a Logistic Regression model, a Multilayer Perceptron (MLP), a Support Vector Machine (SVM), a Random Forest, etc. The loss function can be a mean squared error loss function, a cross-entropy loss function, a multi-label classification loss function, or a triplet loss function. Backpropagation is implemented using the backpropagation algorithm, calculating the gradient of each neuron from the output layer to the input layer, and updating the weights and biases of the neurons based on the gradient. Gradient descent is used to continuously update the parameters, thereby reducing the loss value. Gradient descent can employ various gradient descent algorithms, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent.

[0142] In step 2013, candidate interaction information is selected from the updated portion of the original interaction information according to multiple preset dimensions.

[0143] Here, candidate interaction information can satisfy some or all of the preset dimensions. In advertising scenarios, these multiple dimensions can be product category, industry, optimization goal, etc. For example, if the product category is electrical appliances, the industry is home appliances, and the optimization goal is to improve positive reviews, then candidate interaction information can be selected from these three dimensions. For instance, a candidate interaction message could be "This TV is so high definition!"

[0144] This application's embodiments ensure the quality and compliance of interactive information by filtering and deleting non-compliant content. Candidate interactive information is selected based on preset dimensions, making the interactive information more targeted and relevant. Selecting interactive information that meets optimization goals helps improve positive feedback and user satisfaction. By filtering and selecting, the processing of invalid interactive information is reduced, saving human and material resources.

[0145] See also Figure 3D In step 202, the candidate interaction information is rewritten to obtain the target interaction information.

[0146] In some embodiments, see Figure 3F , Figure 3F This is a schematic diagram of the sixth process of the training method for the language understanding model provided in the embodiments of this application. Figure 3D Step 202, "Rewriting the candidate interaction information to obtain the target interaction information," can be achieved through... Figure 3F Steps 2021 to 2023 will be implemented, as detailed below.

[0147] In step 2021, the candidate interaction information is style-transformed to obtain candidate interaction information in multiple styles.

[0148] In some embodiments, multiple text samples are collected, covering various styles; the text samples are labeled, with each label representing the style of the text sample; text features are extracted for each style of text sample; the text features are vectorized to obtain text feature vectors; a machine learning model is trained using the labeled text samples to obtain a trained machine learning model, which is used to convert text samples from a first style to a second style. Based on candidate interaction information, the trained machine learning model is invoked to obtain candidate interaction information for multiple styles.

[0149] Examples of styles include formal, humorous, and friendly. Annotation can be done manually or using automated text classification techniques. Text features include word selection, sentence structure, sentiment, and tone. Vectorization can be achieved using any of the following methods: Bag of Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), Word Embedding, or Long Short-Term Memory (LSTM). Machine learning models can include Recurrent Neural Networks (RNNs), LSTMs, Transformers, etc.

[0150] In step 2022, for each style of candidate interactive information, syntactic analysis is performed on the candidate interactive information to obtain the syntactic structure of the candidate interactive information.

[0151] In some embodiments, firstly, candidate interaction information is segmented to obtain multiple lexical units (words or phrases); each lexical unit is tagged with its part of speech, i.e., the part of speech of each lexical unit in the sentence is determined; using syntactic dependency analysis, the dependency relationships between lexical units in the sentence are determined based on the part of speech of the lexical units; the syntactic components of the sentence are analyzed; based on the dependency relationships and syntactic components, the syntactic structure of the sentence is constructed. The syntactic structure can be a tree structure, where each node represents a syntactic component, and the connections between nodes represent the dependency relationships between nodes; the syntactic structure is represented in a structured manner, such as a dependency tree, syntactic component tree, or sequence labeling, to facilitate subsequent structural reorganization and the addition of rhetorical devices.

[0152] For example, multiple lexical units include any type of word or phrase. The part of speech of a lexical unit in a sentence includes nouns, verbs, adjectives, etc. Dependency relations describe the grammatical relationships between words, such as subject-verb relations, verb-object relations, and modification relations. The syntactic components of a sentence include subject, predicate, object, attributive, adverbial, and complement.

[0153] In step 2023, the candidate interaction information is restructured based on the syntactic structure to obtain the target interaction information.

[0154] In some embodiments, the candidate interactive information is converted from active to passive voice based on the identified syntactic structure. For example, sentence components are rearranged to convert active voice to passive voice or vice versa. For the candidate interactive information after voice conversion, the positions of sentence components are moved to change the rhythm, emphasis, or rhetorical effect of the sentence. The updated candidate interactive information is then used to add rhetorical devices such as metaphor and parallelism to the sentence based on context awareness to obtain the target interactive information.

[0155] For example, the active voice sentence "The cat catches the mouse" can be transformed into the passive voice sentence "The mouse is caught by the cat." Moving sentence components can involve moving an adverbial phrase from the beginning to the end to emphasize time or condition; for example, "In the storm, we continued forward," becomes "We continued forward, in the storm." Metaphor requires understanding the sentence's context and theme, mapping one concept to another, such as "Time is money." Parallelism requires repeating the same sentence structure to enhance rhythm and emphasis, such as "Faster, Higher, Stronger."

[0156] This application's embodiments, through style transfer and structural reorganization, can generate interactive information in various styles and forms, increasing content diversity and appeal. The style and content of interactive information can be adjusted according to different target audiences and scenarios to better meet specific needs and expectations. Generating various interactive information automatically can save human and material resources and improve work efficiency.

[0157] See also Figure 3D In step 203, the first language understanding model is invoked based on the recommendation information to generate reference interaction information.

[0158] In some embodiments, natural language processing technology is used to parse the recommendation information and extract key information; based on the recommendation information and the target object receiving the recommendation information, a matching interactive information template is selected; in the selected interactive information template, the slots that need to be filled are identified; the key information is used to fill these slots to generate candidate interactive information; and a first language understanding model is invoked to adjust the sentiment of the candidate interactive information according to the business requirements for the interactive information.

[0159] For example, key information includes product name, features, promotional information, brand, etc. Slots include product name, features, promotional information, etc. The system can extract the third feature of the recommendation information and the object feature information of the object to be recommended. The third feature is vectorized to obtain a third feature vector, and the object feature information is vectorized to obtain an object feature vector. The third feature vector and the object feature vector are concatenated to form a fused feature vector. Template features of the interactive information template are extracted, and the template features are vectorized to obtain a template feature vector. The similarity between the fused feature vector and the template feature vector is determined. Interactive information templates corresponding to template feature vectors with similarity greater than a similarity threshold are used as the interactive information templates matching the recommendation information and the object to be recommended. The similarity can be determined using any of the following algorithms: cosine similarity, Euclidean distance, Manhattan distance, or Pearson correlation coefficient. Taking Euclidean distance as an example, if the fused feature vector is (x1, y1) and the template feature vector is (x2, y2), the distance between the fused feature vector and the template feature vector is determined, i.e. Used as a similarity measure.

[0160] In step 204, the pre-trained discriminator is invoked based on the target interaction information and the reference interaction information to obtain the first interaction information.

[0161] Here, a pre-trained discriminator is used to evaluate the quality of the interaction information. The discriminator can be trained on a large amount of labeled data, learning to distinguish between high-quality and low-quality interaction information. The discriminator can be a binary classification model to determine whether the quality of the interaction information meets a certain standard, or a regression model to predict the quality score of the interaction information.

[0162] In some embodiments, see Figure 3G , Figure 3G This is a schematic diagram of the seventh process of the training method for the language understanding model provided in the embodiments of this application. Figure 3DStep 204, "Based on the target interaction information and the reference interaction information, call the pre-trained discriminator to obtain the first interaction information," can be achieved through... Figure 3G Steps 2041 to 2043 are implemented, and the details are explained below.

[0163] In step 2041, based on the pre-trained discriminator, a first quality parameter of the target interaction information is determined, and a second quality parameter of the reference interaction information is determined.

[0164] In some embodiments, features are extracted from both the target interaction information and the reference interaction information. The extracted features are then input into a pre-trained discriminator model for inference, and the quality parameters corresponding to the target interaction information and the reference interaction information are output respectively.

[0165] For example, extracted features may include lexical features (such as vocabulary richness and terminology usage), syntactic features (such as sentence length and complexity), semantic features (such as sentiment and semantic coherence), and stylistic features (such as formality and affinity). Quality parameters can be probability scores, scores, or category labels. When the quality parameter is a probability score, it represents the probability that the interactive information is of high quality. When the quality parameter is a score, it can be compared with predefined scoring criteria to determine the quality level of the interactive information.

[0166] In step 2042, the maximum mass parameter is selected from the first mass parameter and the second mass parameter.

[0167] For example, if the first quality parameter of the target interactive information A is 0.8 and the second quality parameter of the reference interactive information A is 0.5, then the first quality parameter of 0.8 is taken as the maximum quality parameter.

[0168] In step 2043, the interaction information corresponding to the maximum quality parameter is used as the first interaction information.

[0169] Following step 2042 above, if the maximum quality parameter is the first quality parameter of the target interactive information A, then the target interactive information A is taken as the first interactive information.

[0170] This application's embodiments evaluate the quality of interactive information through a pre-trained discriminator, ensuring that the generated interactive information meets certain standards. Automated selection of interactive information with the highest quality parameters reduces manual intervention and improves efficiency. The data-driven evaluation and selection process allows for a more objective assessment of the quality of interactive information.

[0171] See also Figure 3A In step 102, the first weight and the first feature vector of the first interaction information are fused to obtain the second feature vector of the first interaction information.

[0172] In some embodiments, the product of a first weight and a first feature vector of the first interaction information is determined, and the product is used as a second feature vector of the first interaction information.

[0173] For example, if the first weight is k and the first feature vector of the first interaction information is c, then the product is k·c, and k·c is used as the second feature vector of the first interaction information.

[0174] In step 103, a first language understanding model is trained based on the first interaction information and the second feature vector to obtain a second language understanding model.

[0175] In some embodiments, see Figure 3H , Figure 3H This is the eighth flowchart of the training method for the language understanding model provided in the embodiments of this application. Figure 3A Step 103 can be achieved through Figure 3H Steps 1031 to 1033 are implemented, and the details are explained below.

[0176] In step 1031, the second feature vector is masked to obtain the masked feature vector.

[0177] In some embodiments, at least one element is randomly selected from the second feature vector and replaced with a specific marker, such as a mask marker, or set to 0.

[0178] For example, if the second feature vector is (d, e, f, g), and element e is replaced with [mask], then the mask feature vector is (d, mask, f, g).

[0179] In step 1032, sentence pairs are collected from the first interactive information, wherein each sentence pair includes at least two sentences.

[0180] In some embodiments, the first interactive information is segmented into sentences to obtain multiple sentences, and at least any two sentences from the multiple sentences are combined into a sentence pair.

[0181] For example, identify clause markers in the first interactive information, such as "," ".", "?", and "!". Treat the consecutive content before and after the clause marker as a sentence. At least two sentences in a sentence pair can be consecutive and semantically related, without limitation here.

[0182] In step 1033, the first language understanding model is self-supervised and trained based on sentence pairs and mask feature vectors to obtain the second language understanding model.

[0183] Here, self-supervised learning is a type of unsupervised learning that utilizes an auxiliary task (pretext) to mine supervisory information from large-scale unsupervised data. This constructed supervisory information is then used to train the network, enabling it to learn representations valuable for downstream tasks. The training tasks for self-supervised learning include Masked Language Model (MLM) and Next Sentence Prediction (NSP). MLM randomly selects portions of content for masking, allowing the model to predict the masked words based on the context. NSP provides sentence pairs, enabling the model to determine whether the sentences in the pairs are consecutive.

[0184] The self-supervised training in this application does not require a large amount of labeled data, reducing the cost and time of data annotation. By masking and predicting the masked parts, the model learns to remain robust even with partially missing information. Sentence pairs enable the model to better understand the semantic relationships and contextual connections between sentences. Self-supervised training helps the model to understand language structure and semantics more deeply, improving its performance in various natural language processing tasks.

[0185] In some embodiments, see Figure 3I , Figure 3I This is the ninth flowchart illustrating the training method for the language understanding model provided in this application embodiment. Figure 3H Step 1033 can be achieved through Figure 3I Steps 10331 to 10335 are implemented, and the details are explained below.

[0186] In step 10331, the mask feature vector, sentence pairs, and first real labels are combined into training samples, wherein the first real labels represent whether the labeled sentence pairs are continuous in description.

[0187] Here, the mask feature vector represents the sentence representation with missing information, the sentence pair provides contextual information, and the first true label indicates whether the labeled sentence pair is continuous in description. If it is continuous, it is labeled as 1, and if it is not continuous, it is labeled as 0.

[0188] In step 10332, the first language understanding model is invoked based on the training samples to obtain the predicted feature vector corresponding to the mask feature vector and the first predicted label.

[0189] In some embodiments, the convolutional layer and fully connected layer of the first language understanding model are invoked based on the training samples to obtain the predicted feature vector corresponding to the mask feature vector and the first predicted label, wherein the first predicted label represents whether the predicted sentence pairs are continuous in description.

[0190] For example, see Figure 4B , Figure 4B This is a schematic diagram illustrating the principle of training a first language understanding model according to an embodiment of this application. Based on training samples, the convolutional layer of the first language understanding model is invoked to extract features from each sentence in the sentence pair, obtaining a sentence vector for each sentence in the sentence pair. Prediction processing is then performed through the fully connected layer of the first language understanding model to obtain the predicted feature vector corresponding to the mask feature vector. Finally, classification processing is performed through the fully connected layer to obtain the first predicted label.

[0191] In step 10333, a first loss is determined based on the predicted feature vector and the second feature vector, wherein the first loss is the loss corresponding to the first language understanding model when performing the mask prediction task.

[0192] In some embodiments, a first loss between the predicted feature vector and the second feature vector is determined by a loss function.

[0193] Following the example in step 1031 above, the loss function can be the mean squared error loss function, the cross-entropy loss function, the multi-label classification loss function, and the triplet loss function. Taking the mean squared error loss function as an example, if the second feature vector is (d, e, f, g) and the predicted feature vector is (d′, e′, f′, g′), for each first element in the second feature vector, determine the corresponding second element in the predicted feature vector. For example, the first element d in the second feature vector corresponds to the second element d′ in the predicted feature vector, the first element e in the second feature vector corresponds to the second element e′ in the predicted feature vector, the first element f in the second feature vector corresponds to the second element f′ in the predicted feature vector, and the first element g in the second feature vector corresponds to the second element g′ in the predicted feature vector. Then, determine the difference between each first element and its corresponding second element. For example, the difference between the first element d and the second element d′ is dd′, the difference between the first element e and the second element e′ is ee′, the difference between the first element f and the second element f′ is ff′, and the difference between the first element g and the second element g′ is gg′. Finally, determine the square of each difference, for example, (dd′). 2 、(ee′) 2 、(ff′) 2 、(gg′) 2 Determine the sum of the squares corresponding to multiple differences, for example, the sum S = (dd'). 2 +(ee′) 2 +(ff′) 2 +(gg′) 2 Determine the ratio of the sum to the number of elements in the first element, for example, S / 4, as the first loss.

[0194] In step 10334, a second loss is determined based on the first predicted label and the first true label, wherein the second loss is the loss corresponding to the first language understanding model when performing the sentence prediction task.

[0195] In some embodiments, a second loss between the first predicted label and the first true label is determined by a loss function, or the difference between the first predicted label and the second predicted label is used as the second loss.

[0196] For example, the loss function can be the mean squared error loss function, the cross-entropy loss function, the multi-label classification loss function, and the triplet loss function.

[0197] In step 10335, the parameters of the first language understanding model are updated based on the first loss and the second loss to obtain the second language understanding model.

[0198] In some embodiments, the first loss and the second loss are backpropagated to update the parameters of the first language understanding model. The process of calculating the loss and updating the parameters is repeated multiple times until the global loss converges, at which point the iteration process stops, and a second language understanding model is formed.

[0199] For example, backpropagation is implemented using the backpropagation algorithm, which calculates the gradient of each neuron from the output layer to the input layer and updates the neuron's weights and biases based on the gradients. Gradient descent is used to continuously update the parameters, thus reducing the loss value. Various gradient descent algorithms can be used, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent.

[0200] This application's embodiments, by simultaneously performing mask prediction and sentence continuity prediction, enable the model to learn language representations from multiple perspectives, improving its comprehensive comprehension ability. Training with unlabeled data reduces reliance on labeled data, saving labeling costs. The mask prediction task helps the model learn more robust feature representations, improving its performance across various tasks. The use of sentence pairs helps the model better understand contextual information, enhancing its capabilities in tasks such as dialogue and discourse comprehension.

[0201] See also Figure 3A In step 104, based on the recommendation information and the object feature information of the object to be recommended, the second language understanding model is invoked to generate second interactive information.

[0202] Here, the object to be recommended is the object that receives the recommendation information. The object's characteristic information includes at least one of the following: age, gender, location, click preference, dwell time, conversion category, risk preference, and price sensitivity.

[0203] In some embodiments, see Figure 3J , Figure 3J This is a schematic diagram of the tenth process of the training method for the language understanding model provided in the embodiments of this application. Figure 3A Step 104 can be achieved through Figure 3J Steps 1041 to 1042 are implemented, and the details are explained below.

[0204] In step 1041, prompt words are constructed based on the recommendation information and the object feature information of the object to be recommended. The prompt words are used to generate information based on the recommendation information and the object feature information of the object to be recommended.

[0205] For example, taking an advertising scenario, the recommended information is apple juice, and the target user's characteristics are "Age: 20, Gender: Female, likes to click on food ads, likes friendly and lively advertising slogans". The recommended information and target user characteristics are combined into prompt words, such as "There is an apple juice ad now. Based on the following user information, generate appropriate ad comments: Age: 20, Gender: Female, likes to click on food ads, likes friendly and lively advertising slogans".

[0206] In step 1042, the second language understanding model is invoked based on the prompt words, so that the second language understanding model generates second interactive information corresponding to the recommendation information and the object feature information of the object to be recommended.

[0207] In some embodiments, the prompt word is input into the encoder of the second language understanding model. The encoder analyzes the relationship between the various parts of the prompt word through a multi-head self-attention mechanism to generate an encoded representation containing rich semantic information. The encoded representation is then passed to the decoder of the second language understanding model. The decoder predicts the next word based on the encoded representation output by the encoder and the already generated part of the advertising text. At each time step, the decoder outputs a probability distribution representing the likelihood of the next word. The second-cause model selects the word with the highest probability based on the probability distribution to generate the second interactive information. The above process is iterated until the model generates a complete second interactive information or reaches a predefined maximum length.

[0208] Following the example in step 1041, the second interactive message could be: "Newly launched apple juice, with a sweet and refreshing taste, is the perfect choice for young people. Come and try it!"

[0209] This application's embodiments combine recommendation information and object feature information to generate more personalized interactive information, better meeting user needs and preferences. Utilizing a second language understanding model to automatically generate secondary interactive information reduces manual writing workload and improves efficiency. The generated interactive information maintains consistency and relevance with recommendation information and object feature information, contributing to enhanced recommendation effectiveness.

[0210] See also Figure 3A In step 105, based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, the second language understanding model is trained to obtain the third language understanding model.

[0211] Here, the feedback information of the target audience to the second interactive information includes positive feedback and negative feedback. Taking the advertising scenario as an example, the second interactive information is an ad review. Positive feedback can be replying to the ad review or clicking on the ad corresponding to the ad review (recommendation information). Negative feedback can be expressing disinterest in the ad review or ad click.

[0212] In some embodiments, see Figure 3K , Figure 3K This is a schematic diagram of the eleventh step in the training method of the language understanding model provided in the embodiments of this application. Figure 3A Step 105 can be achieved through Figure 3K Steps 1051 to 1054 are implemented, and the details are explained below.

[0213] In step 1051, the prompt word and the second interactive information are combined into a training sample, and the feedback information is used as the second true label, wherein the second true label represents the true emotional polarity of the object to be recommended to the second interactive information.

[0214] Here, emotional polarity refers to the existence of paired, opposing attributes in emotional experiences, manifested as affirmation and negation, positive and negative states, etc. The true emotional polarity of the target audience towards the second interactive information can be either like or dislike, with liking corresponding to a second true label of 1 and disliking corresponding to a second true label of 0.

[0215] In step 1052, the second language understanding model is invoked based on the training samples to obtain the second predicted label, wherein the second predicted label represents the predicted sentiment polarity of the object to be recommended to the second interactive information.

[0216] In some embodiments, the second predicted label is obtained by calling the convolutional and fully connected layers of the second language understanding model based on the training samples.

[0217] For example, see Figure 4C , Figure 4C This is a schematic diagram illustrating the principle of training a second language understanding model according to an embodiment of this application. Based on training samples, the convolutional layer of the second language understanding model is invoked to extract features from the training samples, obtaining sample feature vectors. Based on these sample feature vectors, the fully connected layer of the second language understanding model is invoked to classify the sample feature vectors, obtaining a second predicted label.

[0218] In step 1053, a third loss is determined based on a preset loss function for the second true label and the second predicted label, wherein the third loss characterizes the difference between the true sentiment polarity and the predicted sentiment polarity of the object to be recommended to the second interaction information.

[0219] In some embodiments, a third loss between the second true label and the second predicted label is determined by a loss function.

[0220] For example, the loss function can be the mean squared error loss function, the cross-entropy loss function, the multi-label classification loss function, and the triplet loss function.

[0221] In some embodiments, see Figure 3L , Figure 3L This is a schematic diagram of the twelfth step of the training method for the language understanding model provided in this application embodiment. In the case of N iterations of updating the second language understanding model, where N is the total number of iterations, before step 1053, for the nth iteration update, where n is a positive integer less than or equal to N, the following steps are executed: Figure 3L Steps 301 to 303 are explained in detail below.

[0222] In step 301, a reference point is determined based on the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration. The reference point represents the average probability difference between the interactive information generated by the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration.

[0223] Here, the second language understanding model that has not been iteratively updated refers to the language understanding model that has not undergone any iterative updates. The probability difference between the second language understanding model after each iterative update and the second language understanding model that has not been iteratively updated is determined. The sum of the probability differences corresponding to multiple iterative updates is determined, and the ratio of this sum to the total number of iterative updates M is determined. This ratio is used as the average probability difference, i.e., the reference point.

[0224] In some embodiments, see Figure 3M , Figure 3M This is a schematic diagram of the thirteenth step of the training method for the language understanding model provided in the embodiments of this application. Figure 3L Step 301 can be achieved through Figure 3M Steps 3011 to 3014 are implemented, and the details are explained below.

[0225] In step 3011, a first probability distribution of the first candidate interaction information is determined, wherein the first candidate interaction information is interaction information generated based on the second language understanding model updated in the nth iteration.

[0226] Here, the first probability distribution refers to the distribution of probabilities of different first candidate interaction information generated by the second language understanding model in the nth iteration update.

[0227] In step 3012, a second probability distribution of the second candidate interaction information is determined, wherein the second candidate interaction information is interaction information generated based on the un-updated second language understanding model.

[0228] Here, the second probability distribution refers to the distribution of probabilities of different second candidate interaction information generated by the second language understanding model that has not been iteratively updated.

[0229] In step 3013, a divergence operation is performed on the first probability distribution and the second probability distribution to obtain the difference value between the first probability distribution and the second probability distribution.

[0230] In some embodiments, a divergence measure is used to calculate the difference between a first probability distribution and a second probability distribution. A higher divergence value indicates a greater difference between the first and second probability distributions; a lower divergence value indicates a more similar first and second probability distribution.

[0231] For example, divergence measures can use relative entropy (Kullback-Leibler Divergence, KL divergence) or JS divergence (Jensen-Shannon Divergence). Taking KL divergence as an example, determine the ratio of the first probability distribution to the second probability distribution; take the natural logarithm of each ratio to obtain the logarithm value; determine the product of the logarithm value and the first probability distribution; determine the sum of the products corresponding to all interactive events to obtain the difference between the first and second probability distributions. The difference value is non-negative, and is zero only when the first and second probability distributions are identical.

[0232] In step 3014, the expected value of the difference is determined and used as a reference point.

[0233] Here, the expected value (i.e., the reference point) reflects the average difference between the probability distribution of interactive information generated by the second language understanding model after multiple iterations and the probability distribution of interactive information generated by the second language understanding model after no iterations.

[0234] This application's embodiments can evaluate the model's improvement effect by comparing the probability distribution of interaction information generated by the model before and after iterative updates. The divergence operation of the probability distribution helps quantify the uncertainty of the model's generated interaction information, aiding in understanding the model's confidence level. The expected value, as a reference point, can help determine the direction and effect of model iterative updates, guiding further model optimization.

[0235] See also Figure 3L In step 302, a value function is determined based on a reference point, wherein the value function is used to characterize the quality of the interactive information generated by the second language understanding model in the nth iteration update.

[0236] In some embodiments, see Figure 3N , Figure 3N This is the fourteenth flowchart of the training method for the language understanding model provided in the embodiments of this application. Figure 3L Step 302, "Determining the value function based on the reference point," can be achieved through... Figure 3N Steps 3021 to 3025 are implemented, and the details are explained below.

[0237] In step 3021, the ratio of the first probability distribution and the second probability distribution is determined, and the logarithm of the ratio is determined. The logarithm represents the difference in the probability of generating the third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration.

[0238] In some embodiments, the logarithm of the ratio reflects the difference in probability between the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration when generating third interactive information (i.e., interactive information with specified conditions). If the logarithm is greater than 0, it means that the second language understanding model updated in the nth iteration has a higher probability of generating third interactive information than the second language understanding model that has not been updated in the previous iteration; if the logarithm is less than 0, it means that the second language understanding model updated in the nth iteration has a lower probability of generating third interactive information than the second language understanding model that has not been updated in the previous iteration.

[0239] For example, if the first probability distribution is P and the second probability distribution is Q, then the ratio of the first probability distribution to the second probability distribution is P / Q, and the logarithm of the ratio is logP / Q.

[0240] In step 3022, a first difference between the logarithm and the reference point is determined, and a second difference between the reference point and the logarithm is determined. The first difference represents the difference between the probability difference and the reference point. The probability difference represents the difference in the probability of generating the third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration. The second difference represents the opposite of the first difference.

[0241] In some embodiments, the first difference reflects the difference between the probability of the second language understanding model updated in the nth iteration and the probability of the second language understanding model not updated in the nth iteration generating the third interactive information and the average probability difference calculated above. Conversely, the second difference reflects the difference between the average probability difference and the probability difference between the second language understanding model updated in the nth iteration and the probability of the second language understanding model not updated in the nth iteration generating the third interactive information.

[0242] Following the example of step 3021 above, if the reference point is R, then the first difference between the logarithm and the reference point is logP / QR, and the second difference between the reference point and the logarithm is R-logP / Q.

[0243] In step 3023, a first mapping is performed on the first difference to obtain a first mapping value, and a second mapping is performed on the second difference to obtain a second mapping value. The first mapping value represents the strength of the positive feedback of the object to be recommended to the interactive information, and the second mapping value represents the strength of the negative feedback of the object to be recommended to the interactive information.

[0244] In some embodiments, the first difference and the second difference can be mapped using a logistic function, with both the first difference and the second difference mapped to values ​​between 0 and 1, to reflect the strength of the feedback from the object to be recommended to the interactive information.

[0245] Following the example of step 3022 above, if σ represents a logic function, the first difference is mapped to the first mapping value σ(logP / QR) through the logic function σ, and the second difference is mapped to the second mapping value σ(R-logP / Q).

[0246] In step 3024, a first product of the positive coefficient and the first mapping value is determined, and a second product of the negative coefficient and the second mapping value is determined. The positive coefficient represents a preset coefficient when the object to be recommended provides positive feedback to the interactive information, and the negative coefficient represents a preset coefficient when the object to be recommended provides negative feedback to the interactive information.

[0247] In some embodiments, the first mapping value corresponding to the interaction information that provides positive feedback to the object to be recommended is positively weighted using a positive coefficient, and the second mapping value corresponding to the interaction information that provides negative feedback to the object to be recommended is negatively weighted using a negative coefficient, so as to adjust the intensity of the feedback.

[0248] Following the example of step 3023 above, if the positive coefficient is γ1 and the negative coefficient is γ2, then the first product of the positive coefficient and the first mapping value is γ1·σ(logP / QR), and the second product of the negative coefficient and the second mapping value is γ2·σ(R-logP / Q).

[0249] In step 3025, the first product and the second product are combined to form a value function.

[0250] Following the example of step 3024 above, the value function can be expressed as follows: when the object to be recommended provides positive feedback on the interaction information, the value function is T = γ1·σ(logP / QR); when the object to be recommended provides negative feedback on the interaction information, the value function is T = γ2·σ(R-logP / Q).

[0251] This application's embodiments, through mapping and multiplication operations, can quantify the strength of positive and negative feedback from the recommended object to interactive information, providing valuable feedback information for the recommendation system. The introduction of positive and negative coefficients allows for adjustment of the strength of positive and negative feedback according to actual needs, making the value function more reflective of the actual feedback situation. The value function comprehensively considers both positive and negative feedback, providing a comprehensive evaluation metric that helps the recommendation system make more accurate recommendation decisions.

[0252] See also Figure 3L In step 303, a preset loss function is constructed based on the second true label and the value function.

[0253] In some embodiments, see Figure 3O , Figure 3O This is the fifteenth flowchart of the training method for the language understanding model provided in the embodiments of this application. Figure 3L Step 303, "Constructing a pre-defined loss function based on the second true label and the value function," can be achieved through... Figure 3O Steps 3031 to 3032 are implemented, and the details are explained below.

[0254] In step 3031, a third difference between the second true label and the value function is determined.

[0255] Following the example of step 3025 above, if the second true label is β, then the third difference between the second true label and the value function is β-T.

[0256] In step 3032, the third difference is combined with a preset expectation function to form a preset loss function.

[0257] Following the example in step 3031 above, the expectation function can be the mean squared error (MSE), cross-entropy loss function, etc. If the expectation function is represented by E, then the loss function can be Z = E(β-T).

[0258] The loss function in this application provides a method for quantifying model performance, helping to evaluate the gap between model predictions and true labels. The value of the loss function guides the direction of model optimization; by minimizing the loss function, the model can improve its prediction accuracy. Different expectation functions can adapt to different task requirements, making the loss function flexible and adaptable.

[0259] See also Figure 3K In step 1054, the parameters of the second language understanding model are updated based on the third loss to obtain the third language understanding model.

[0260] In some embodiments, the third loss is backpropagated to update the parameters of the second language understanding model. The process of calculating the third loss and updating the parameters is repeated multiple times until the global loss converges, at which point the iteration process stops, thus forming the third language understanding model.

[0261] For example, backpropagation is implemented using the backpropagation algorithm, which calculates the gradient of each neuron from the output layer to the input layer and updates the neuron's weights and biases based on the gradients. Gradient descent is used to continuously update the parameters, thus reducing the loss value. Various gradient descent algorithms can be used, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent.

[0262] This application's embodiments train a second language understanding model to predict the sentiment polarity of the target audience to second interactive information, thereby enabling the second language understanding model to acquire sentiment analysis capabilities. Utilizing the feedback information from the target audience as true labels, the model can continuously learn and optimize, improving its prediction accuracy. The third language understanding model can then provide more personalized recommendation information based on the sentiment polarity of the target audience, improving user satisfaction.

[0263] In some embodiments, see Figure 3P , Figure 3P This is the sixteenth flowchart of the training method for the language understanding model provided in this application embodiment. After step 105, the following steps are executed: Figure 3P Steps 106 and 107 are explained in detail below.

[0264] In step 106, target recommendation information and target object feature information of the target object are obtained.

[0265] Here, the target object is the object that receives the target recommendation information. The target object's characteristic information includes at least one of the following: age, gender, location, click preference, dwell time, conversion category, risk preference, and price sensitivity.

[0266] In step 107, the third language understanding model is invoked based on the target recommendation information and the target object feature information to generate target interaction information corresponding to the target recommendation information.

[0267] Here, based on the target recommendation information and the target object feature information, the convolutional layer of the third language understanding model is called to vectorize the target recommendation information to obtain the target recommendation vector, and the target object feature information is vectorized to obtain the target object vector. The target recommendation vector and the target object vector are concatenated to form the target fusion vector. Based on the target fusion vector, the fully connected layer of the third language understanding model is called to generate the target interaction information corresponding to the target recommendation information.

[0268] This application's embodiments combine target recommendation information and target object feature information to generate more personalized target interaction information, thereby increasing target object participation and enhancing user-recommendation system interaction. By analyzing target object feedback on interaction information, the recommendation algorithm can be further optimized, improving the accuracy and effectiveness of recommendations.

[0269] The following will describe exemplary applications of the embodiments of this application in advertising recommendation scenarios.

[0270] Against the backdrop of the rapid development of digital advertising, ad reviews (i.e., interactive information) are gradually becoming an important touchpoint for brands to enhance user engagement and drive marketing conversions. Compared to simply focusing on ad copy generation or comment bots, ad reviews have higher demands for diversity, marketing orientation, and personalization. They need to maintain a high degree of alignment with the preferences of different users while also considering creativity, interactivity, and commercial value. Large-scale model language understanding and generation capabilities, such as Bidirectional Encoder Representations from Transformers (BERT), provide new technological support for achieving automated and personalized ad review writing.

[0271] See Figure 5A , Figure 5A This is a schematic diagram of the first process of generating advertising copy using the related technologies provided in this application embodiment. Related technologies are typically based on the natural language generation capabilities of pre-trained language models, such as the Large Language Model Meta AI (Llama), to provide initial advertising creative copy for brands or marketing teams. The workflow includes: general corpus training, where the pre-trained model is trained using a large-scale general corpus to acquire and master a wide range of language knowledge. Due to the high cost of pre-training, open-source models are generally used directly as the model base; fine-tuning with limited domain data, where the pre-trained model is fine-tuned based on limited advertising text or brand information, for example, using supervised fine-tuning (SFT) or proximal policy optimization (PPO) algorithms to adjust its language style and brand tone in the advertising copy; and inference deployment for copy generation, where a model with lower memory requirements is obtained after quantification. Based on input keywords or themes, it can output copy that meets advertising needs, assisting marketers in secondary editing and modification. This technology can lower the barrier to entry for writing advertising copy and reduce manual operation costs, but its output content is usually still based on a unified template or fixed style, lacking a deep understanding of user opinions and interactive characteristics in the commenting scenario, and is difficult to meet the needs of advertising comments that require strong interaction and refined positioning.

[0272] See Figure 5B , Figure 5B This is a schematic diagram of the second process for generating advertising copy using the related technology provided in this application embodiment. The related technology provides intelligent replies in the comment section to enhance user interaction and platform activity. This typically manifests as a "comment robot" automatically replying to user questions or comments in the comment section. Its workflow includes: first, fine-tuning the model using large-scale dialogue data to enable it to understand and respond to multi-turn dialogue contexts; then, at the application level, combining several keyword recognition and topic management mechanisms to provide standardized answers to popular topics or common questions; some solutions use simple profile information of the target audience to provide slightly different answers to different users, enhancing personalization. While this related technology can attract user participation and interaction to some extent, its main purpose is to improve the activity and basic interaction quality of the comment section. It lacks a deep alignment with advertising marketing goals and commercial value, and cannot systematically address further needs such as diverse expressions and advertising effectiveness measurement.

[0273] The datasets used by these technologies for training large models primarily come from general, publicly available web resources, lacking specialized corpora for advertising review scenarios. While chatbot technology based on comment sections exists, its technical goals focus on everyday conversational interactions, failing to meet the needs of marketing-oriented comment generation. These technologies suffer from limited data scale, reliance on manual writing leading to high production costs, inconsistent quality due to a lack of systematic quality assessment standards, and a lack of diverse styles. They employ an equal sample weighting strategy, ignoring the guiding value of actual advertising review performance data. Finally, their uniform output model makes it difficult to adapt to the diverse needs of different user groups.

[0274] This application's embodiments construct a three-level data augmentation mechanism, improving the scale and quality of training data (i.e., first interaction information) through a progressive processing flow of filtering, rewriting, and synthesis. A dynamic sample weighting mechanism is proposed, establishing a multi-dimensional effect index fusion model to achieve differentiated value assessment of training samples, resolving the alignment issue between model optimization direction and business objectives. A user-segment-based preference modeling system is constructed, combined with an improved reinforcement learning alignment algorithm, enabling precise personalized generation of advertising reviews.

[0275] In social advertising scenarios, the interaction between users and ads in related technologies is relatively simple, often failing to fully leverage the social potential of ads. (See also...) Figure 6 , Figure 6This is an application illustration of displaying advertisement comments provided in the embodiments of this application. The first comment (i.e., interactive information) of the advertisement (i.e., recommendation information) has social attributes. It deeply integrates the advertisement comment area with artificial intelligence technology, activates the social value of the advertisement, and enhances the interactive experience between users and between users and advertisements.

[0276] On social media or related interactive platforms, comment sections are the primary venues for users to exchange opinions, share interests, and share information. However, advertising formats using these technologies often struggle to integrate into social interactions, resulting in a lack of deep communication between advertisers and users. This application's embodiments, starting from the social needs of users and the marketing needs of advertisers, intelligently generate and personalize the first comment in the comment section, making the advertising content more closely aligned with users' actual interests and creating a truly personalized interactive model.

[0277] This application's embodiments generate customized comment content for different types of users. Based on the user's (i.e., the target audience) background information (i.e., target characteristic information) and advertising information (i.e., recommendation information), the copywriting style and content direction of the initial ad review are automatically matched and optimized. Through this intelligent and dynamic method of generating initial ad reviews, a higher level of social appeal can be created in the comment section, increasing user attention and engagement with the advertising content.

[0278] By cultivating a genuine, engaging, and sustainable interactive atmosphere in the comment section, users can be effectively incentivized to leave comments and share further, creating a positive cycle on a social level. More interaction leads to more valuable data accumulation, which in turn helps the model continuously learn and optimize, ultimately improving the overall effectiveness of advertising and brand awareness. Furthermore, a more active comment section strengthens user engagement with the platform, enhancing its position in information dissemination and social interaction, creating a synergistic effect of value creation between advertising and the platform.

[0279] See Figure 7 , Figure 7This is a flowchart illustrating the process of generating ad reviews that match user preferences, as provided in this application embodiment. The language understanding model training method provided in this application embodiment, centered on the core needs of ad review generation, constructs a three-in-one systematic technical framework of "data construction - model optimization - preference alignment." Through the synergistic effect of multi-level technical modules, it achieves a comprehensive improvement in the quality, effectiveness, and personalization of ad reviews. The domain-adaptive training data construction process serves as the technical foundation, focusing on the large-scale production of high-quality training data in ad review scenarios, addressing the core pain points of traditional solutions such as data scarcity and monotonous styles. The supervised fine-tuning training process based on dynamic sample weighting serves as the optimization engine, achieving alignment between model capabilities and advertiser's core objectives through sample value quantification driven by commercial performance indicators. The preference alignment model based on the Kahneman-Tversky Optimisation (KTO) algorithm serves as the decision-making center, dynamically adjusting the generation strategy by combining user segmentation characteristics and the risk perception mechanism of prospect theory, ensuring a high degree of adaptation between review content and user needs.

[0280] Data construction provides high-quality input for model optimization, which in turn improves the decision accuracy of preference alignment. The feedback data generated by preference alignment feeds back into data construction and model updates, ultimately enabling the continuous evolution of the advertising review generation system.

[0281] The domain-adaptive training data construction process adopts a three-level data construction workflow, including three parts: data screening, data rewriting, and data synthesis. The specific implementation is as follows: The data screening part constructs a multi-dimensional quality evaluation system, including: formal compliance indicators, such as comment length (10-50 characters) (i.e., the length of the original interactive information is less than the length threshold), sentence structure (the proportion of interrogative / exclamatory sentences) (i.e., the sentence structure of the original interactive information is a preset sentence structure), and emoticon usage standards (i.e., the symbols in the original interactive information are compliant); semantic compliance indicators, such as detecting non-compliant words and sensitive content through a large model classifier (i.e., deleting non-compliant words from some of the original interactive information); and diversity assurance mechanisms, such as stratified sampling according to product category, industry, and optimization goal (i.e., selecting candidate interactive information from the updated portion of the original interactive information according to multiple preset dimensions).

[0282] The data rewriting section implements a differentiated rewriting strategy, specifically as follows: style transfer, such as switching between preset styles like formal / humorous / friendly; sentence optimization, such as structural reorganization based on deep syntactic analysis to achieve active / passive conversion, component shifting, etc., and adding rhetorical devices such as metaphor and parallelism in combination with context awareness.

[0283] The data synthesis part establishes a dual-model collaborative generation mechanism, and the specific implementation is as follows: In the generation stage, a pre-trained language understanding model is used to generate comments, and multiple marketing script templates are preset; in the screening stage, a quality discriminator (i.e., a pre-trained discriminator) is constructed to compare the generated comments with high-quality comments written by humans, and the comments with higher quality (i.e., quality parameters) are selected as part of the training data.

[0284] By leveraging the feedback data from existing manually written reviews in the advertising system, we can construct multi-dimensional performance evaluation metrics (i.e., evaluation indicators), such as immediate feedback metrics: click-through rate (CTR) and conversion rate (CVR); business value metrics: cost per thousand impressions (CPM); interaction metrics: number of comment replies (R); and risk metrics: negative feedback rate (NF).

[0285] A weighted function is constructed using multidimensional performance evaluation indicators, as shown in the following formula (1):

[0286]

[0287] Where μ represents the mean and σ represents the variance. α, β, and γ are user-defined coefficients (i.e., the second weight), log is the logarithmic operation, max is used to select the maximum value from multiple numbers, and Weight represents the weighting function, used to balance the importance of each indicator in the weighting formula. These coefficients can be adjusted according to specific business needs and data analysis results to ensure that the weighted result reasonably reflects the actual contribution of each indicator. The optimization goal of the weighting function is to increase the ad conversion probability of CTR*CVR while improving the ad's CPM and user response interest through appropriate ad reviews.

[0288] First, manually written comments (i.e., the first interactive information) are uploaded online. Multi-dimensional metrics (i.e., evaluation metrics) of the advertisements corresponding to the comments are obtained. The weights of the manually written comments (i.e., the first weights) are calculated using these metrics. The language understanding model (i.e., the first language understanding model) is then trained in a self-supervised manner using the weighted manually written comments to obtain the trained language understanding model (i.e., the second language understanding model), thereby improving the model's performance in generating advertisement comments.

[0289] When aligning preferences in the model, the first step is to segment users and collect preference data. For user segmentation, a hierarchical clustering algorithm is used to construct a multi-dimensional user (i.e., the object to be recommended) feature space (i.e., object feature information) based on the following characteristics: statistical features (age, gender, location), behavioral features (click preferences, dwell time, conversion category), and psychological features (risk preference, price sensitivity). For preference data collection, prompts are designed to embed user information (i.e., object feature information) and advertising information (i.e., recommendation information). The large model (i.e., the second language understanding model) trained above generates personalized advertising reviews (i.e., second interaction information). After these generated advertising reviews are deployed online, user feedback data on these reviews (i.e., feedback information) is collected for subsequent preference model training.

[0290] This application employs the KTO algorithm for preference alignment. This algorithm does not require paired preference data; it only needs to provide a good or bad binary signal (i.e., feedback information) for each (prompt, response) pair. Here, prompt refers to the content input to the large model, response refers to the advertisement review output by the large model, and the good or bad binary signal refers to the user's feedback. Compared to algorithms such as SFT and PPO, the KTO algorithm requires less data and can more efficiently utilize limited training data. The KTO algorithm does not require paired preference data, simplifying the data collection process. The KTO algorithm is more suitable for handling datasets with imbalanced positive and negative sample classes, improving the model's performance across different classes.

[0291] The KTO algorithm is based on prospect theory, which aims to explain people's decision-making behavior when faced with uncertainty and risk, especially why people's decisions often deviate from the expected utility theory in traditional economics. In advertising review generation, user reactions to advertising reviews can also be explained by prospect theory, specifically including the following core concepts:

[0292] Reference Point: Users typically judge the attractiveness of ad reviews based on changes relative to a certain reference point. In an advertising system, users use previously viewed reviews as a reference point to evaluate the quality of new reviews.

[0293] Loss aversion: Users are more sensitive to losses than to equivalent gains. For example, generating inappropriate ad reviews can significantly increase negative emotions in users, while obtaining positive feedback such as clicks and conversions is relatively difficult.

[0294] Diminishing marginal utility (DMP): As gains or losses increase, marginal utility or marginal pain decreases. In ad reviews, generating one or two high-quality words or phrases can significantly increase positive user feedback, while even a few inappropriate expressions can lead to a sharp increase in negative feedback.

[0295] Assume that users judge the relative quality of new ad reviews based on all previously seen ad reviews, rather than their absolute quality. Setting the reference point as the expected reward under the optimal strategy simplifies to the KL divergence between the optimal and reference strategies. The reference point can be determined using the following formula (2):

[0296] z0 = E x~D [KL(π θ (y|x)||π ref (y|x))] (2)

[0297] Where x represents prompt (advertisements and user information), y represents response (advertisement comments), and π θ (y|x) represents the new large model (i.e., the large model being trained in this round) (i.e., the second language understanding model updated in the nth iteration), π ref (y|x) represents the large model obtained from the above training (i.e., the second language understanding model that has not been iteratively updated). In actual calculation, D represents a batch (data in batch), E represents expectation operation, KL represents divergence operation, and z0 represents the reference point.

[0298] The original value function in prospect theory contains an exponential term, which has high computational complexity and is difficult to optimize. Therefore, a logistic function is used instead to achieve a sensitive response to small changes. First, the logarithm of the ratio of the large model being trained in this round to the large model already trained is determined, as shown in the following formula (3). Then, the value function is determined by the logarithm and the reference point, as shown in the following formula (4):

[0299]

[0300] Where log represents the logarithmic operation, π θ (y|x) represents the new large model (i.e., the large model currently being trained), π ref (y|x) represents the large model obtained from the above training, λ F (i.e., the positive coefficient) and λ U (That is, the negative coefficients) represent the loss weights for positive and negative samples, respectively, σ is the logistic function, ω is the adjustment parameter, and y ~ y desirable|x represents the probability that, given input x, output y is a comment liked by the user, where y ~ y undesirable |x represents the probability that, given input x, output y is a comment that the user dislikes.

[0301] The loss function is constructed by combining the label and the value function, as shown in the following formula (5):

[0302]

[0303] Among them, L KTO (π θ ,π ref ) represents the loss function (i.e., the preset loss function), λ y The label (i.e., the second true label) takes a value of 0 or 1. Samples in the training data with positive feedback such as clicks, conversions, likes, and comments are labeled as 1, and the rest are labeled as 0. That is, x = <ad information, user information, ad comments>, y = <1 / 0>. Based on the large model obtained from the above training, and guided by the loss function above, further training is performed to obtain the final pre-trained large language model (i.e., the third language understanding model), which is used to generate ad comments (i.e., target interaction information) corresponding to user preferences.

[0304] This application's embodiments, through precise comment generation, not only enable users to receive comments that better match their interests and needs, but also reduce annoying or repetitive invalid information, thereby effectively lowering the negative feedback rate. More creative and emotionally resonant interactions in the comment section can significantly improve the platform's overall activity and user retention rate. The automated generation and filtering of ad comments drastically reduces the time and manpower costs of manual copywriting, review, and management. Simultaneously, dynamic iteration based on business data enables continuous and stable optimization of ad comment quality, further reducing the costs of later maintenance and iteration. By deeply exploring the potential of large models in aligning natural language generation with user preferences, a novel solution is provided for ad comment scenarios that combines diverse expression, precise marketing, and personalized interaction. This enhances the commercial conversion and placement value of advertisements while strengthening user engagement and positive experience in the comment section, achieving a win-win situation for both the advertising platform and users.

[0305] The following description continues to illustrate the exemplary structure of the language understanding model training device 233 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules in the training device 233 for the language understanding model stored in the memory 230 may include:

[0306] The weight determination module 2331 is used to determine the first weight of the first interactive information based on the evaluation index of the recommendation information, wherein the first interactive information represents the interaction performed on the recommendation information.

[0307] The data fusion module 2332 is used to fuse the first weight and the first feature vector of the first interaction information to obtain the second feature vector of the first interaction information.

[0308] The first training module 2333 is used to train the first language understanding model based on the first interaction information and the second feature vector to obtain the second language understanding model.

[0309] The information generation module 2334 is used to generate second interactive information by calling the second language understanding model based on the recommendation information and the object feature information of the object to be recommended.

[0310] The second training module 2335 is used to train the second language understanding model based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, so as to obtain the third language understanding model.

[0311] In some embodiments, the weight determination module 2331 is further configured to determine a second weight for each evaluation indicator, wherein the second weight characterizes the importance of each evaluation indicator; and to perform weighted fusion of multiple evaluation indicators based on the second weight to obtain a first weight for the first interactive information.

[0312] In some embodiments, the weight determination module 2331 is further configured to: obtain the initial coefficients of the evaluation index for each evaluation index; determine the partial derivatives of the objective function with respect to the evaluation index, wherein the objective function is pre-constructed based on multiple evaluation indexes and the coefficients of each evaluation index; determine the iterative function based on the partial derivatives; perform iterative calculation on the iterative function to obtain the coefficients of the evaluation index after each iteration calculation; and when the objective function converges, use the coefficients of each evaluation index after the last iteration calculation as the second weight of each evaluation index.

[0313] In some embodiments, the weight determination module 2331 is further configured to: filter candidate interaction information from multiple original interaction information of the recommendation information according to preset rules; rewrite the candidate interaction information to obtain target interaction information; call the first language understanding model based on the recommendation information to generate reference interaction information; and call the pre-trained discriminator based on the target interaction information and the reference interaction information to obtain the first interaction information.

[0314] In some embodiments, the weight determination module 2331 is further configured to select a portion of the original interactive information that meets a first preset condition from a plurality of original interactive information, wherein the first preset condition includes at least one of the following: the length of the original interactive information is less than a length threshold, the sentence structure of the original interactive information is a preset sentence structure, and the symbols in the original interactive information are compliant; non-compliant words are deleted from the portion of the original interactive information; and candidate interactive information is selected from the updated portion of the original interactive information according to a plurality of preset dimensions.

[0315] In some embodiments, the weight determination module 2331 is further configured to perform style conversion on the candidate interaction information to obtain candidate interaction information of multiple styles; perform syntactic analysis on the candidate interaction information for each style to obtain the syntactic structure of the candidate interaction information; and restructure the candidate interaction information based on the syntactic structure to obtain the target interaction information.

[0316] In some embodiments, the weight determination module 2331 is further configured to determine a first quality parameter of the target interaction information and a second quality parameter of the reference interaction information based on a pre-trained discriminator; select the maximum quality parameter from the first quality parameter and the second quality parameter; and use the interaction information corresponding to the maximum quality parameter as the first interaction information.

[0317] In some embodiments, the first training module 2333 is further configured to mask the second feature vector to obtain a masked feature vector; collect sentence pairs from the first interaction information, wherein each sentence pair includes at least two sentences; and perform self-supervised training on the first language understanding model based on the sentence pairs and the masked feature vector to obtain a second language understanding model.

[0318] In some embodiments, the first training module 2333 is further configured to combine the mask feature vector, sentence pairs, and first ground truth labels into training samples, wherein the first ground truth labels characterize whether the labeled sentence pairs are continuous in description; call the first language understanding model based on the training samples to obtain the predicted feature vector corresponding to the mask feature vector and the first predicted label; determine the first loss based on the predicted feature vector and the second feature vector, wherein the first loss is the loss corresponding to the first language understanding model when performing the mask prediction task; determine the second loss based on the first predicted label and the first ground truth label, wherein the second loss is the loss corresponding to the first language understanding model when performing the sentence prediction task; and update the parameters of the first language understanding model based on the first loss and the second loss to obtain the second language understanding model.

[0319] In some embodiments, the information generation module 2334 is further configured to construct prompt words based on the recommendation information and the object feature information of the object to be recommended, wherein the prompt words are used to generate information based on the recommendation information and the object feature information of the object to be recommended; and to call the second language understanding model based on the prompt words, so that the second language understanding model generates second interactive information corresponding to the recommendation information and the object feature information based on the recommendation information and the object feature information of the object to be recommended.

[0320] In some embodiments, the second training module 2335 is further configured to combine the prompt word and the second interaction information into training samples, and use the feedback information as the second true label, wherein the second true label represents the true emotional polarity of the object to be recommended to the second interaction information; call the second language understanding model based on the training samples to obtain a second predicted label, wherein the second predicted label represents the predicted emotional polarity of the object to be recommended to the second interaction information; determine a third loss between the second true label and the second predicted label based on a preset loss function, wherein the third loss represents the difference between the true emotional polarity and the predicted emotional polarity of the object to be recommended to the second interaction information; and update the parameters of the second language understanding model based on the third loss to obtain a third language understanding model.

[0321] In some embodiments, the second training module 2335 is further configured to, when the second language understanding model is updated N times (where N is the total number of iterations), and for the nth iteration update (where n is a positive integer less than or equal to N), perform the following processing: determine a reference point based on the second language understanding model updated in the nth iteration and the second language understanding model not updated in the nth iteration, wherein the reference point characterizes the average probability difference between the interactive information generated by the second language understanding model updated in the nth iteration and the second language understanding model not updated in the nth iteration; determine a value function based on the reference point, wherein the value function is used to characterize the quality of the interactive information generated by the second language understanding model updated in the nth iteration; and construct a preset loss function based on the second true label and the value function.

[0322] In some embodiments, the second training module 2335 is further configured to: determine a first probability distribution of the first candidate interaction information, wherein the first candidate interaction information is interaction information generated based on the second language understanding model updated in the nth iteration; determine a second probability distribution of the second candidate interaction information, wherein the second candidate interaction information is interaction information generated based on the second language understanding model not updated in the previous iteration; perform a divergence operation on the first probability distribution and the second probability distribution to obtain the difference value between the first probability distribution and the second probability distribution; and determine the expected value of the difference value, using the expected value as a reference point.

[0323] In some embodiments, the second training module 2335 is further configured to determine the ratio of the first probability distribution and the second probability distribution, determine the logarithm of the ratio, the logarithm representing the difference in the probability of generating third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model not updated in the previous iteration; determine a first difference between the logarithm and a reference point, and determine a second difference between the reference point and the logarithm, wherein the first difference represents the difference between the probability difference and the reference point, the probability difference represents the difference in the probability of generating third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model not updated in the previous iteration, and the second difference represents the opposite of the first difference; A first mapping is applied to the first difference to obtain a first mapped value, and a second mapping is applied to the second difference to obtain a second mapped value. The first mapped value represents the strength of the positive feedback from the target object to the interactive information, and the second mapped value represents the strength of the negative feedback from the target object to the interactive information. A first product of the positive coefficient and the first mapped value is determined, and a second product of the negative coefficient and the second mapped value is determined. The positive coefficient represents a preset coefficient when the target object provides positive feedback to the interactive information, and the negative coefficient represents a preset coefficient when the target object provides negative feedback to the interactive information. The first product and the second product are combined to form a value function.

[0324] In some embodiments, the second training module 2335 is further configured to determine a third difference between the second true label and the value function; and combine the third difference with a preset expectation function to form a preset loss function.

[0325] In some embodiments, the second training module 2335 is further configured to acquire target recommendation information and target object feature information of the target object; and to call a third language understanding model based on the target recommendation information and target object feature information to generate target interaction information corresponding to the target recommendation information.

[0326] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the language understanding model training method described in this application.

[0327] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the training method for the language understanding model provided in this application. For example, ... Figure 3A The training method for the language understanding model is shown.

[0328] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0329] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0330] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0331] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0332] In summary, by using the embodiments of this application to train a first language understanding model to obtain a second language understanding model, the first weights determined by the evaluation index based on recommendation information are fused with the first feature vector of the first interaction information. The evaluation index is also incorporated into the second feature vector. The first language understanding model is trained based on the first interaction information and the second feature vector. The first weights enable the first language understanding model to learn comprehensively based on the first interaction information while also allowing for targeted learning of different first interactions through the evaluation index. This enhances the model's focus on and understanding of interaction information corresponding to high evaluation indices, thereby improving the comprehensiveness of model training. And specificity; when training the second language understanding model to obtain the third language understanding model, second interaction information is generated based on the recommendation information and the object feature information of the object to be recommended. Based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, the second language understanding model is trained. This allows the trained third language understanding model to simultaneously consider the object feature information and the feedback information of the object to be recommended in response to the second interaction information. The feedback information reflects the preference of the object to be recommended in response to the second interaction information. By training the model with feedback information, the model can be continuously adjusted to align with the object preferences, thereby improving the specificity and generalization ability of the third language understanding model for different objects to be recommended.

[0333] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for training a language understanding model, characterized in that, The method includes: Based on the evaluation metrics of the recommendation information, a first weight of the first interactive information is determined, wherein the first interactive information represents the interaction performed in response to the recommendation information. The first weight and the first feature vector of the first interaction information are fused to obtain the second feature vector of the first interaction information; A first language understanding model is trained based on the first interaction information and the second feature vector to obtain a second language understanding model; Based on the recommendation information and the object feature information of the object to be recommended, the second language understanding model is invoked to generate the second interactive information; Based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, the second language understanding model is trained to obtain the third language understanding model.

2. The method according to claim 1, characterized in that, The first weight characterizes the degree of influence of the first interaction information on the evaluation index. When there are multiple evaluation indicators, determining the first weight of the first interaction information based on the evaluation index of recommendation information includes: A second weight is determined for each of the evaluation indicators, wherein the second weight characterizes the importance of each of the evaluation indicators; The first weight of the first interactive information is obtained by weighting and fusing multiple evaluation indicators based on the second weight.

3. The method according to claim 2, characterized in that, Determining the second weight for each of the evaluation indicators includes: For each of the evaluation indicators, the initial coefficients of the evaluation indicators are obtained, and the partial derivatives of the objective function with respect to the evaluation indicators are determined, wherein the objective function is pre-constructed based on multiple evaluation indicators and the coefficients of each of the evaluation indicators; The iterative function that determines the initial coefficients based on the partial derivatives; The iterative function is iteratively calculated to obtain the coefficients of the evaluation index after each iteration. When the objective function converges, the coefficient of each evaluation index after the last iteration is used as the second weight of each evaluation index.

4. The method according to claim 1, characterized in that, The process of training a first language understanding model based on the first interaction information and the second feature vector to obtain a second language understanding model includes: Mask the second feature vector to obtain the masked feature vector; Sentence pairs are collected from the first interactive information, wherein each sentence pair includes at least two sentences; The first language understanding model is self-supervised and trained based on the sentence pairs and the mask feature vectors to obtain the second language understanding model.

5. The method according to claim 4, characterized in that, The step of performing self-supervised training on the first language understanding model based on the sentence pairs and the mask feature vectors to obtain the second language understanding model includes: The mask feature vector, the sentence pair, and the first real label are combined into training samples, wherein the first real label represents whether the labeled sentence pair is continuous in description; Based on the training samples, the first language understanding model is invoked to obtain the predicted feature vector corresponding to the mask feature vector, and the first predicted label; A first loss is determined based on the predicted feature vector and the second feature vector, wherein the first loss is the loss corresponding to the first language understanding model when performing a mask prediction task; A second loss is determined based on the first predicted label and the first true label, wherein the second loss is the loss corresponding to the first language understanding model when performing the sentence prediction task; The parameters of the first language understanding model are updated based on the first loss and the second loss to obtain the second language understanding model.

6. The method according to claim 1, characterized in that, The step of generating second interactive information by calling the second language understanding model based on the recommendation information and the object feature information of the object to be recommended includes: Prompt words are constructed based on the recommendation information and the object feature information of the object to be recommended, wherein the prompt words are used to generate information based on the recommendation information and the object feature information of the object to be recommended; The second language understanding model is invoked based on the prompt word, so that the second language understanding model generates second interactive information corresponding to the recommendation information and the object feature information of the object to be recommended, based on the recommendation information and the object feature information of the object to be recommended.

7. The method according to claim 6, characterized in that, The process of training the second language understanding model based on the second interaction information and the feedback information of the target user in response to the second interaction information to obtain the third language understanding model includes: The prompt word and the second interaction information are combined to form a training sample, and the feedback information is used as a second true label, wherein the second true label represents the true emotional polarity of the object to be recommended to the second interaction information; The second language understanding model is invoked based on the training samples to obtain a second predicted label, wherein the second predicted label represents the predicted sentiment polarity of the object to be recommended to the second interaction information; A third loss is determined based on a preset loss function for the second true label and the second predicted label, wherein the third loss characterizes the difference between the true sentiment polarity and the predicted sentiment polarity of the object to be recommended in relation to the second interaction information; The parameters of the second language understanding model are updated based on the third loss to obtain the third language understanding model.

8. The method according to claim 7, characterized in that, In the case of updating the second language understanding model N times, where N is the total number of iterations, before determining the third loss based on the preset loss function to determine the second true label and the second predicted label, the method further includes: For the nth iteration update, where n is a positive integer less than or equal to N, perform the following processing: A reference point is determined based on the second language understanding model updated in the nth iteration and the second language understanding model not updated in the nth iteration, wherein the reference point represents the average probability difference between the interactive information generated by the second language understanding model updated in the nth iteration and the second language understanding model not updated in the nth iteration. A value function is determined based on the reference point, wherein the value function is used to characterize the quality of the interactive information generated by the second language understanding model in the nth iteration update; The preset loss function is constructed based on the second true label and the value function.

9. The method according to claim 8, characterized in that, The determination of the reference point based on the second language understanding model updated in the nth iteration and the second language understanding model not updated in the nth iteration includes: A first probability distribution of the first candidate interaction information is determined, wherein the first candidate interaction information is interaction information generated based on the second language understanding model updated in the nth iteration; Determine a second probability distribution for the second candidate interaction information, wherein the second candidate interaction information is interaction information generated based on the second language understanding model that has not been iteratively updated; Perform divergence calculation on the first probability distribution and the second probability distribution to obtain the difference value between the first probability distribution and the second probability distribution; Determine the expected value of the difference and use the expected value as a reference point.

10. The method according to claim 9, characterized in that, The determination of the value function based on the reference point includes: Determine the ratio of the first probability distribution to the second probability distribution, and determine the logarithm of the ratio. The logarithm represents the difference in the probability of generating the third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration. A first difference between the logarithm and the reference point is determined, and a second difference between the reference point and the logarithm is determined, wherein the first difference represents the difference between the probability difference and the reference point, the probability difference represents the difference in the probability of generating the third interactive information between the second language understanding model updated in the nth iteration and the second language understanding model that has not been updated in the previous iteration, and the second difference represents the opposite of the first difference; The first difference is mapped to a first mapping value, and the second difference is mapped to a second mapping value. The first mapping value represents the strength of the positive feedback of the object to be recommended to the interactive information, and the second mapping value represents the strength of the negative feedback of the object to be recommended to the interactive information. A first product of a positive coefficient and the first mapping value is determined, and a second product of a negative coefficient and the second mapping value is determined, wherein the positive coefficient represents a preset coefficient when the object to be recommended provides positive feedback to the interaction information, and the negative coefficient represents a preset coefficient when the object to be recommended provides negative feedback to the interaction information; The first product and the second product are combined to form a value function.

11. The method according to claim 8, characterized in that, The construction of the preset loss function based on the second true label and the value function includes: Determine the third difference between the second true label and the value function; The third difference is combined with the preset expectation function to form the preset loss function.

12. The method according to any one of claims 1 to 11, characterized in that, Before determining the first weight of the first interaction information based on the evaluation metrics of the recommendation information, the method further includes: Candidate interaction information is selected from multiple original interaction information in the recommendation information according to preset rules; The candidate interaction information is rewritten to obtain the target interaction information; Based on the recommended information, the first language understanding model is invoked to generate reference interaction information; Based on the target interaction information and the reference interaction information, a pre-trained discriminator is invoked to obtain the first interaction information.

13. The method according to claim 12, characterized in that, The step of filtering candidate interaction information from multiple original interaction information of the recommendation information according to preset rules includes: Select a portion of the original interactive information that meets a first preset condition from the plurality of original interactive information, wherein the first preset condition includes at least one of the following: the length of the original interactive information is less than a length threshold, the sentence structure of the original interactive information is a preset sentence structure, and the symbols in the original interactive information are compliant; Remove non-compliant words from the aforementioned original interactive information; Candidate interaction information is selected from the updated portion of the original interaction information according to multiple preset dimensions.

14. The method according to claim 12, characterized in that, The step of rewriting the candidate interaction information to obtain the target interaction information includes: The candidate interaction information is style-transformed to obtain candidate interaction information in multiple styles; For each style of candidate interaction information, perform syntactic analysis on the candidate interaction information to obtain the syntactic structure of the candidate interaction information; Based on the syntactic structure, the candidate interaction information is restructured to obtain the target interaction information.

15. The method according to claim 12, characterized in that, The step of calling a pre-trained discriminator based on the target interaction information and the reference interaction information to obtain the first interaction information includes: Based on the pre-trained discriminator, a first quality parameter of the target interaction information is determined, and a second quality parameter of the reference interaction information is determined. Select the maximum mass parameter from the first mass parameter and the second mass parameter; The interaction information corresponding to the maximum quality parameter is used as the first interaction information.

16. The method according to any one of claims 1 to 15, characterized in that, After training the second language understanding model based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information to obtain the third language understanding model, the method further includes: Obtain target recommendation information and target object feature information; Based on the target recommendation information and the target object feature information, the third language understanding model is invoked to generate target interaction information corresponding to the target recommendation information.

17. A training device for a language understanding model, characterized in that, The device includes: The weight determination module is used to determine the first weight of the first interactive information based on the evaluation index of the recommendation information, wherein the first interactive information represents the interaction performed in response to the recommendation information. The data fusion module is used to fuse the first weight and the first feature vector of the first interaction information to obtain the second feature vector of the first interaction information. The first training module trains a first language understanding model based on the first interaction information and the second feature vector to obtain a second language understanding model. The information generation module is used to generate second interactive information by calling the second language understanding model based on the recommendation information and the object feature information of the object to be recommended; The second training module is used to train the second language understanding model based on the second interaction information and the feedback information of the object to be recommended in response to the second interaction information, so as to obtain the third language understanding model.

18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the training method of the language understanding model according to any one of claims 1 to 16.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the training method of the language understanding model according to any one of claims 1 to 16.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the training method of the language understanding model according to any one of claims 1 to 16.