Viewpoint generation method and device

By applying reinforcement learning algorithms and low-resource sample annotation strategy in view generation, the problems of low efficiency and insufficient diversity of view generation in the existing technology are solved, and the effect of improving the quality and efficiency of view generation under limited resources is achieved.

CN119961450APending Publication Date: 2025-05-09BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311466288.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, the generation of views of manual marking methods is inefficient, requiring more human resources to be invested, and the mining cost is high; while the zero-sample zero shot method is aimed at no training data, resulting in insufficient diversity of views of the mining, poor accuracy and quality, and cannot meet actual needs.

Method used

A reinforcement learning algorithm is adopted and a candidate view generation model is built using a low-resource sample annotation strategy, and the initial model of view generation is determined, and the efficiency and quality of view generation is improved through a small number of high-quality sample annotations.

Benefits of technology

Under the premise of limited sample resources, the efficiency and quality of view generation will be steadily improved, which is more suitable for the needs of product improvement and market decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961450A_ABST
    Figure CN119961450A_ABST
Patent Text Reader

Abstract

The invention discloses a viewpoint generation method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps of obtaining to-be-reasoned text data in response to a received viewpoint generation request; inputting the to-be-reasoned text data into a pre-constructed viewpoint generation model to obtain viewpoints corresponding to the to-be-reasoned text data, the viewpoint generation model being obtained by performing reinforcement learning training on an initial viewpoint generation model, the initial viewpoint generation model is obtained according to a candidate text observation set determined by the candidate viewpoint generation model and a preset text-to-text migration neural network model. According to the embodiment, the initial viewpoint generation model is determined by using the candidate viewpoint generation model constructed by using the reinforcement learning algorithm and the low-resource sample labeling strategy, so that the viewpoint generation efficiency and quality are stably improved on the premise of limited sample resources, and product improvement and market decision making are better served.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for generating opinions. Background Art

[0002] With the promotion and application of network technology and the development of the information industry, data information as a new type of resource has been increasingly valued for its potential value. In the field of product consumption, by collecting user comments on products, pre-sales and after-sales consultation information and other feedback information, and analyzing and mining these feedback information, we can obtain user opinions, which are of great significance to product improvement and market decision-making deployment. There are currently two main methods for generating opinion mining. One is to generate opinions through manual labeling based on expert and business expertise; the other is to generate opinions based on zero-shot generation of traditional deep learning.

[0003] In the process of implementing the present invention, the inventors found that the prior art has the following problems:

[0004] The opinions obtained by manual labeling are inefficient, require a lot of human resources, and have high mining costs; and zero-shot itself is oriented towards untrained data, which will lead to insufficient diversity of mined opinions, poor accuracy and quality, and cannot meet actual needs. Summary of the invention

[0005] In view of this, an embodiment of the present invention provides a method and device for generating opinions. The initial model for generating opinions is determined by using a reinforcement learning algorithm and a candidate opinion generation model constructed using a low-resource sample annotation strategy. The initial model for generating opinions is determined by annotating a small amount of high-quality samples. This achieves a steady improvement in the efficiency and quality of opinion generation under the premise of limited sample resources, thereby better serving product improvement and market decision-making.

[0006] To achieve the above object, according to one aspect of an embodiment of the present invention, a method for generating opinions is provided, comprising:

[0007] In response to receiving the opinion generation request, obtaining text data to be inferred;

[0008] The text data to be inferred is input into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

[0009] Optionally, before inputting the text data to be inferred into a pre-built opinion generation model, the method further includes: using a pre-acquired text data set as a training sample, using a reinforcement learning algorithm with opinion generation enhancement to train an initial opinion generation model, and constructing the opinion generation model.

[0010] Optionally, before using the pre-acquired text data set as a training sample, the method also includes: training the pre-trained model using a contrastive learning framework to obtain a text vector extraction model; based on the text vector extraction model, using a greedy algorithm to deduplicate the text data set to obtain a de-redundant core text set, so that the core text set can be used as the text data set.

[0011] Optionally, before using a reinforcement learning algorithm with opinion generation enhancement to train the initial opinion generation model, the method also includes: extracting a first text set from the text data set, constructing a candidate opinion generation model based on the first text set in combination with a preset text-to-text migration neural network model, and obtaining a corresponding candidate text opinion set based on the candidate opinion generation model; using the candidate text opinion set to train the text-to-text migration neural network model to obtain the initial opinion generation model.

[0012] Optionally, a first text set is extracted from the text data set, and a candidate opinion generation model is constructed based on the first text set and in combination with a preset text-to-text migration neural network model, and a corresponding candidate text opinion set is obtained based on the candidate opinion generation model, including: extracting a specified number of texts from the first text set, marking the opinion of each of the texts, and obtaining a first text opinion set; using the text opinions in the first text opinion set and the context text opinions of the text opinions to train a preset text-to-text migration neural network model to obtain a candidate opinion generation model; inputting the remaining text of the first text set into the candidate opinion generation model to obtain the opinion corresponding to the remaining text, and using the second text opinion set consisting of the remaining text and the opinion corresponding to the remaining text, and the first text opinion set as the candidate text opinion set.

[0013] Optionally, before taking the second text opinion set and the first text opinion set as candidate text opinion sets, the method further includes: verifying the second text opinion set and deleting text opinions in the second text opinion set that fail the verification.

[0014] Optionally, a pre-acquired text data set is used as a training sample, and a reinforcement learning algorithm with enhanced opinion generation is used to train the initial opinion generation model to construct the opinion generation model, including: extracting a second text set from the text data set, constructing a reward text opinion set, and determining a reward model in combination with the initial opinion generation model; extracting the remaining part from the text data set as a third text set, using the reward model as an evaluation model for reinforcement learning, and using the initial opinion generation model as an action model for the reinforcement learning; based on the third text set, combined with a loss function with enhanced opinion generation, the evaluation model and the action model are updated and trained to obtain the opinion generation model after the initial opinion generation model is trained.

[0015] Optionally, a second text set is extracted from the text data set, a reward text opinion set is constructed, and an initial model is generated in combination with the opinion. The reward model is determined, including: for each second text in the second text set, a specified number of second opinions are generated through the opinion generation initial model, opinion pairs for each second text are constructed, and the quality of the opinion pairs is labeled according to preset rules to obtain a reward text opinion set; an initial reward model is established based on the opinion generation initial model, and a reward model is obtained by training the initial reward model using the reward text opinion set.

[0016] Optionally, based on the third text set, in combination with a loss function with enhanced opinion generation, the evaluation model and the action model are updated and trained, including: for each text in the third text set, generating an opinion and an experience value for each text through the evaluation model and the action model, storing each text, the opinion of each text and the experience value in an experience pool, and determining the optimal opinion of each text based on the experience value; based on the texts in the experience pool and the optimal opinion of the texts, updating and training the action model and the evaluation model by calculating the action loss function, the evaluation loss function, and the loss function with enhanced opinion generation of the reinforcement learning.

[0017] Optionally, based on the text in the experience pool and the optimal viewpoint of the text, the action model and the evaluation model are updated and trained by calculating the action loss function, evaluation loss function, and loss function of enhanced viewpoint generation of the reinforcement learning, including: according to a preset number of training times, the current text and the corresponding viewpoint are obtained from the experience pool, the output probability of each word in the current viewpoint and the value of the current viewpoint are obtained in combination with the action model and the evaluation model, and the action loss function and evaluation loss function of the reinforcement learning are calculated in combination with the experience value corresponding to the current text in the experience pool; according to the current text and the optimal viewpoint corresponding to the current text, the output probability of the next word in the optimal viewpoint is calculated under the premise that the current text is input, to obtain the loss function of enhanced viewpoint generation; according to the action loss function, evaluation loss function, and loss function of enhanced viewpoint generation of the reinforcement learning, the action model and the evaluation model are updated and trained until the preset number of training times is reached.

[0018] Optionally, after obtaining the viewpoints corresponding to the text data to be inferred, the method further includes: using the reward model to score the viewpoints corresponding to the text data to be inferred, and storing the viewpoints with scores greater than a preset scoring threshold into a viewpoint library.

[0019] According to a second aspect of an embodiment of the present invention, there is provided a device for generating a viewpoint, including:

[0020] A text acquisition module, configured to acquire text data to be inferred in response to receiving a viewpoint generation request;

[0021] The opinion generation module is used to input the text data to be inferred into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

[0022] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for generating a point of view, comprising:

[0023] one or more processors;

[0024] a storage device for storing one or more programs,

[0025] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.

[0026] According to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method provided by the first aspect of the embodiment of the present invention is implemented.

[0027] An embodiment of the invention has the following advantages or beneficial effects: by responding to receiving an opinion generation request, the text data to be inferred is obtained; the text data to be inferred is input into a pre-constructed opinion generation model to obtain the opinion corresponding to the text data to be inferred, the opinion generation model is obtained by performing reinforcement learning training on the opinion generation initial model, the opinion generation initial model is a technical solution obtained based on the candidate text observation set determined by the candidate opinion generation model and the preset text-to-text migration neural network model, the reinforcement learning algorithm is used to determine the opinion generation initial model by using the candidate opinion generation model constructed by the low-resource sample annotation strategy, the opinion generation initial model is determined by a small amount of high-quality sample annotations, and the efficiency and quality of opinion generation are steadily improved under the premise of limited sample resources, so as to better serve product improvement and market decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0029] Figure 1 is a schematic diagram of the main process of the method for generating opinions according to an embodiment of the present invention;

[0030] Figure 2 is a flowchart of a method for generating opinions according to an embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of main modules of a device for generating viewpoints according to an embodiment of the present invention;

[0032] Figure 4 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;

[0033] Figure 5 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0034] It should be noted that in the technical solution of the present invention, the collection / collection, updating, analysis, use, transmission, storage and other aspects of user personal information involved are in compliance with the provisions of relevant laws and regulations, are used for legal and reasonable purposes, are not shared, disclosed or sold outside of these legal uses, and are subject to supervision and management by national regulatory authorities. Necessary measures should be taken for user personal information to selectively block the use or access to personal information data to prevent illegal access to such personal information data, ensure that persons who have access to personal information data comply with the provisions of relevant laws and regulations, and ensure the security of user personal information. In addition, once these user personal information data are no longer needed, the risks should be minimized by limiting or even prohibiting data collection and / or deleting data.

[0035] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0036] Among the existing opinion generation methods, the manual labeling method has low opinion generation efficiency, requires a lot of human resources, and has high mining costs; while zero-shot itself is oriented towards untrained data, which will lead to insufficient diversity of mined opinions and low accuracy, which cannot meet actual needs.

[0037] In order to solve the above problems existing in the prior art, the present invention proposes a method for generating opinions, inputting the text data to be inferred into a pre-built opinion generation model to generate corresponding opinions, and the opinion generation model is mainly obtained by training the initial opinion generation model with a reinforcement learning algorithm, wherein the initial opinion generation model is obtained according to the candidate text observation set determined by the candidate opinion generation model and the preset text-to-text migration neural network model. By using the reinforcement learning algorithm and the candidate opinion generation model constructed by the low-resource sample annotation strategy to determine the initial opinion generation model, and determining the initial opinion generation model by annotating a small amount of high-quality samples, the efficiency and quality of opinion generation are steadily improved under the premise of limited sample resources, so as to better serve product improvement and market decision-making.

[0038] In the introduction to the embodiments of the present invention, the terms and their meanings are as follows:

[0039] Dropout: is a neural network training technique that randomly turns off neurons to reduce overfitting;

[0040] Prompt: is a technique that leverages the knowledge of a pre-trained model by adding additional text or vectors to the input to guide the model to complete a specific task;

[0041] Cosine similarity: also called cosine similarity, measures the similarity between two vectors by calculating the cosine value of the angle between them.

[0042] Figure 1 is a schematic diagram of the main process of the method for generating opinions according to an embodiment of the present invention, such as Figure 1 As shown, the method for generating opinions in the embodiment of the present invention includes the following steps S101 to S102.

[0043] Step S101: In response to receiving a viewpoint generation request, obtaining text data to be inferred.

[0044] Specifically, in the field of product consumption, with the development and promotion of Internet technology, more and more product trading businesses have expanded to online, and users and merchants conduct trading activities through online communication. Merchants can integrate online users' demand for products or feedback information, extract users' opinions from them, and provide rich and reliable information support for market demand decisions and product improvements. Specific feedback information can be user demand information before sales, comment information after receipt, post-sales usage information, and question and answer information sent by intended users. These text information can be used as text data to be inferred in embodiments of the present invention. When the system receives a request to generate an opinion, the text data information to be inferred is extracted from the information repository according to the extraction information in the request to generate the opinion.

[0045] Step S102: input the text data to be inferred into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

[0046] Specifically, the opinion generation model of the embodiment of the present invention is trained using a reinforcement learning algorithm. At the same time, taking into account the high cost and low efficiency of existing manual labeling methods, a candidate opinion generation model constructed using a low-resource sample annotation strategy is used to determine the initial opinion generation model, and the learning ability of the model is fully utilized. A rich set of candidate text opinions is obtained through a small amount of high-quality text opinion annotations, so as to obtain an initial opinion generation model with better performance.

[0047] According to one embodiment of the present invention, before inputting the text data to be inferred into a pre-built opinion generation model, the method further includes: using a pre-acquired text data set as a training sample, using a reinforcement learning algorithm with opinion generation enhancement to train an initial opinion generation model, and constructing the opinion generation model.

[0048] Specifically, the training sample set plays an important role in model training. The quality and quantity of the training sample set have a direct and inevitable relationship with the effect of model training. Rich and diverse training samples can train and build a high-quality opinion generation model. The embodiment of the present invention ensures the richness of text samples from the data source, and pre-randomly obtains text sets from multiple data sources such as user comments, intended user questions and answers, and customer service questions and answers. The text data in the pre-acquired text set is segmented and divided into sentences according to punctuation marks to obtain a text data set. The rich and diverse text data sets are used as model training samples, and considering that in the training process of reinforcement learning, the quality effect curve of generating opinions is an oscillating optimization process. In order to reduce excessive oscillations in the direction of the difference, the embodiment of the present invention will input the opinions with the best quality in the training samples into the model again, introduce the loss function of opinion generation enhancement to train the model in a good direction, optimize the training process of the model, and can more quickly and efficiently mine high-quality opinions of the text data to be inferred. The reinforcement learning algorithm with the loss function of opinion generation enhancement is used to train the initial model of opinion generation to build a high-quality opinion generation model.

[0049] According to another embodiment of the present invention, before using a pre-acquired text data set as a training sample, the method further includes: training the pre-trained model in a contrastive learning framework to obtain a text vector extraction model; based on the text vector extraction model, using a greedy algorithm to perform redundancy processing on the text data set to obtain a de-redundant core text set, and using the core text set as the text data set.

[0050] Specifically, considering that multiple data sources are comments or feedback on a specified product, there will inevitably be a large number of semantically similar texts in the collected text data set, and randomly acquired texts will also have text redundancy, so it is best to perform redundancy removal on the above-obtained text vector set. In order to perform text clustering, the distance between the vectors of semantically similar texts is as close as possible, and the distance between the vectors of semantically unrelated texts is as far apart as possible, and the pre-acquired text data set needs to be encoded into a text vector. The embodiment of the present invention adopts the unsupervised training method of the contrastive learning framework Simcse, uses the pre-trained roberta model as the encoder, and trains it using the pre-trained text sample set collected in advance. Based on the principle that the two dropout data of each sample are positive samples (considered to be the same text), and the other samples in each batch are negative samples (considered not to be the same sample), the loss function of the training is determined as:

[0051]

[0052] Where N is the number of batch samples, h i is the vector of the [CLS] character output by the i-th sample encoder after the first dropout. is the vector of the [CLS] character output by the i-th sample encoder after the second dropout, sim represents the cosine similarity. τ is the temperature hyperparameter, which is set to 0.8 in this scheme, and the text vector extraction model is obtained by the gradient descent method.

[0053] Furthermore, with the help of the above-mentioned text vector extraction model, based on the principle that the semantics of two texts are similar and the cosine distance of the text vectors output by the text vector extraction model is also close, a text set is found in the text vector set so that the cosine similarity distance between the text vectors in this set is large, and text redundancy removal can be performed. This solution uses a greedy algorithm to solve. First, the above-mentioned text data set D1 is input into the vector extraction model S to obtain the text vector set E; a text d is randomly extracted from D1 to initialize the core text set C and the minimum similarity min_dis, and d is added to C. According to the text vector E of d d , calculate E d The cosine similarity dis of E is used, and dis is used as the initial value of min_dis. The number of texts to be extracted is used as the number of iterations. For each iteration, the text d corresponding to the maximum value of min_dis is selected and added to the C set to obtain the text vector E of the text d. d , and then calculate E dThe cosine similarity dis with E is selected, and the minimum value between min_dis and dis is updated, that is, min_dis = min(min_dis, dis), until the preset number of iterative updates is reached. At this time, the text in the C set is the core text set for de-redundant text. In this way, the core text set is used as the text dataset for subsequent model training, which can not only build the opinion generation model more efficiently, but also improve the text diversity in the fixed-size text dataset.

[0054] According to one embodiment of the present invention, before using a reinforcement learning algorithm with opinion generation enhancement to train the initial opinion generation model, the method also includes: extracting a first text set from the text data set, constructing a candidate opinion generation model based on the first text set in combination with a preset text-to-text migration neural network model, and obtaining a corresponding candidate text opinion set based on the candidate opinion generation model; using the candidate text opinion set to train the text-to-text migration neural network model to obtain the initial opinion generation model.

[0055] Specifically, the embodiment of the present invention makes full use of a text data set with rich viewpoints, obtains rich training data from the data source level, and based on the Chinese T5-base text-to-text migration neural network model, has experts manually define corresponding viewpoints for a small amount of text in the first text set, constructs a candidate viewpoint generation model, and then uses the candidate viewpoint generation model to generate a rich and diverse set of candidate text viewpoints; finally, the T5-base text-to-text migration neural network model is trained with the candidate text set to obtain the viewpoint generation initial model T52.

[0056] According to another embodiment of the present invention, a first text set is extracted from the text data set, a candidate opinion generation model is constructed based on the first text set and in combination with a preset text-to-text migration neural network model, and a corresponding candidate text opinion set is obtained based on the candidate opinion generation model, including: extracting a specified number of texts from the first text set, marking the opinion of each of the texts, and obtaining a first text opinion set; using the text opinions in the first text opinion set and the context text opinions of the text opinions to train a preset text-to-text migration neural network model to obtain a candidate opinion generation model; inputting the remaining text of the first text set into the candidate opinion generation model to obtain the opinion corresponding to the remaining text, and using the second text opinion set consisting of the remaining text and the opinion corresponding to the remaining text, and the first text opinion set as the candidate text opinion set.

[0057] Specifically, a specified number of texts are extracted from the first text set. The number of texts extracted here is generally determined to be a relatively small number according to the number of text data sets. For example, if the number of text data sets in the embodiment of the present invention is in the millions, then the number extracted here is 500. This small set of text sets can be manually defined by experts to define the viewpoints of each text, and obtain a high-precision and high-quality first text viewpoint set corresponding to the text and the viewpoint. According to the neural network model of Chinese T5-base text-to-text migration, in order to adapt the T5 model, prompt learning is adopted, according to two prompts: 1. Viewpoint generation task: What viewpoints can be extracted from the {s} sentence? 2. Viewpoint generation task: If the viewpoint {g} can be extracted from the {s} sentence, then what viewpoint can be extracted from the {s1} sentence? Use the text viewpoint in the first text viewpoint set and the context text viewpoint of the text viewpoint to form a prompt data pair (prompt, viewpoint) as the training data of the T5-base model, input it into the T5-base model for training, and obtain the model candidate viewpoint generation model T51. The remaining texts of the first text set are then input into the constructed candidate opinion generation model T51 to obtain the opinions corresponding to each of the remaining texts. These remaining texts and their corresponding opinions constitute the second text opinion set. The above-mentioned high-precision and high-quality first text opinion set and the currently acquired second text opinion set are combined to form a candidate text opinion set.

[0058] According to yet another embodiment of the present invention, before using the second text opinion set and the first text opinion set as candidate text opinion sets, the method further includes: verifying the second text opinion set and deleting text opinion sets in the second text opinion set that fail the verification.

[0059] Specifically, in order to further improve the quality of the second text opinion set, the embodiment of the present invention delivers the second text opinion set to annotators before using the second text opinion set and the first text opinion set as candidate text opinion sets. The annotators then make judgments and verify the second text opinion set. This manual verification method reduces the difficulty and improves the efficiency compared to defining opinions based on text. Compared to defining opinions by experts for a large amount of text, it also reduces labor costs while ensuring the quality of opinions.

[0060] According to another embodiment of the present invention, a pre-acquired text data set is used as a training sample, and a reinforcement learning algorithm with enhanced opinion generation is used to train an initial opinion generation model to construct the opinion generation model, including: extracting a second text set from the text data set, constructing a reward text opinion set, and determining a reward model in combination with the initial opinion generation model; extracting the remaining part from the text data set as a third text set, using the reward model as an evaluation model for reinforcement learning, and using the initial opinion generation model as an action model for the reinforcement learning; according to the third text set, combined with a loss function with enhanced opinion generation, updating and training the evaluation model and the action model to obtain the opinion generation model after the initial opinion generation model is trained.

[0061] Specifically, considering that the initial model for generating opinions will inevitably have other uncontrollable problems or interference factors, in order to ensure the quality of generated opinions, an embodiment of the present invention adds a reward model, which is used to judge the quality of generated opinions; further, using the above-determined initial model for generating opinions and the reward model, according to the third text set, the reinforcement learning actor action model and the critic evaluation model are trained through a loss function with opinion generation enhancement to obtain a trained opinion generation model.

[0062] According to another embodiment of the present invention, a second text set is extracted from the text data set, a reward text opinion set is constructed, an initial model is generated in combination with the opinion, and a reward model is determined, including: for each second text in the second text set, a specified number of second opinions are generated through the opinion generation initial model, opinion pairs of each second text are constructed, and the quality of the opinion pairs is labeled according to preset rules to obtain a reward text opinion set; an initial reward model is established based on the opinion generation initial model, and a reward model is obtained by training the initial reward model using the reward text opinion set.

[0063] Specifically, the embodiment of the present invention extracts 2000 second texts from the text data set to form a second text set, uses the above-trained viewpoints to generate the initial model T52, generates 5 second viewpoints for each second text, and combines the generated 5 second viewpoints in pairs to form 10 viewpoint pairs (viewpoint 1, viewpoint 2). Finally, according to the rule that the viewpoint quality of viewpoint 1 is higher than the viewpoint quality of viewpoint 2, these viewpoint pairs are manually standardized, and the viewpoint pairs that meet the rule are marked as 1, and the viewpoint pairs that do not meet the rule are marked as 0. These marked viewpoint pairs constitute the reward text viewpoint set. The initial reward model generates the encoder encoder part of the initial model based on the viewpoint, and then inputs the vector output by the encoder of the first word [CLS] of the text into a fully connected layer LM with an output dimension of 1 to form a complete initial reward model. The two viewpoints of the viewpoint pair marked as 0 in the above-marked reward text viewpoint set are swapped to ensure that the quality of the previous viewpoint is higher than the quality of the next viewpoint. The training data is obtained in the format of (prompt+viewpoint 1, prompt+viewpoint 2), and the reward model is finally obtained by calculating the following loss function.

[0064]

[0065] Among them, g1 is the vector of prompt+viewpoint 1 output [CLS] through T52 encoder, g2 is the vector of prompt+viewpoint 2 output [CLS] through T52 encoder, and 1m is the output of the fully connected layer LM.

[0066] According to another embodiment of the present invention, based on the third text set, in combination with a loss function with enhanced opinion generation, the evaluation model and the action model are updated and trained, including: for each text in the third text set, generating an opinion and an experience value for each text through the evaluation model and the action model, storing each text, the opinion of each text and the experience value in an experience pool, and determining the optimal opinion of each text based on the experience value; based on the texts in the experience pool and the optimal opinion of the texts, updating and training the action model and the evaluation model by calculating the action loss function, the evaluation loss function, and the loss function with enhanced opinion generation of the reinforcement learning.

[0067] Specifically, the embodiment of the present invention trains the evaluation model and the action model efficiently and with high quality by using a reinforcement learning method with enhanced opinion generation. The specific reinforcement learning training process is divided into two stages. The first stage is the sampling stage. For each text d in the third text set, the action model A, that is, the initial model for opinion generation, is used to generate the opinion g corresponding to d. The text d and the corresponding opinion g are input into the A model to obtain the output probability action_probs of each word of opinion g; the text d and the corresponding opinion g are then input into the T52 model to obtain the output baseline probability base_action_probs of each word of opinion g. In this process, the T52 model does not participate in the training and only provides a baseline probability; similarly, the text d and the corresponding opinion g are input into the evaluation model C to obtain the value output by the evaluation model C for the opinion g, and then the text d and the corresponding opinion g are input into the evaluation model C to obtain the value output by the evaluation model C for the opinion g. The opinion g is input into the evaluation model reward model R, and the r output by the reward model R for the opinion g is obtained. In this process, the R model does not participate in the training, but only provides a benchmark score. In addition, G{} is used to store the optimal opinion corresponding to each text d in the third text set in each iteration, and each text in the third text set is input into the T52 model to obtain the initial opinion as the initial value of G. The optimal opinion gz and text d corresponding to text d are extracted from G and input into the above-mentioned reward model R to obtain the optimal opinion r_gz, and G is updated by the value of r and r_gz. If the r value of opinion g is greater than r_gz, the g opinion is updated to G to obtain the optimal opinion set G updated in real time determined by the experience values ​​r and r_gz. Based on action_probs and base_action_probs, the reward value reward can be calculated:

[0068]

[0069] where coeT = 0.01; and the advantage value adv:

[0070] adv i =reward i -value i :

[0071] Finally, the text d, the corresponding opinion g and the output probability action_probs, reward value reward, advantage value adv and value value in each experience value are stored in the experience pool EX_Buffer.

[0072] The text d, opinion g, corresponding experience value in the experience pool and the corresponding optimal opinion gz obtained by sampling in the first stage are applied to the second stage model training stage of reinforcement learning. The action model and evaluation model are trained and updated by calculating the action loss function, evaluation loss function and loss function enhanced by opinion generation. The action model after training and updating is the opinion generation model.

[0073] According to another embodiment of the present invention, based on the text in the experience pool and the optimal viewpoint of the text, by calculating the action loss function, evaluation loss function, and loss function of enhanced viewpoint generation of the reinforcement learning, the action model and the evaluation model are updated and trained, including: according to a preset number of training times, the current text and the corresponding viewpoint are obtained from the experience pool, the output probability of each word in the current viewpoint and the value of the current viewpoint are obtained in combination with the action model and the evaluation model, and the action loss function and evaluation loss function of the reinforcement learning are calculated in combination with the experience value corresponding to the current text in the experience pool; according to the current text and the optimal viewpoint corresponding to the current text, the output probability of the next word in the optimal viewpoint is calculated under the premise that the current text is input, to obtain the loss function of enhanced viewpoint generation; according to the action loss function, the evaluation loss function, and the loss function of enhanced viewpoint generation of the reinforcement learning, the action model and the evaluation model are updated and trained until the preset number of training times is reached.

[0074] Specifically, in the current training phase of reinforcement learning, according to the preset number of training times, the current text d and the corresponding opinion g are obtained from the above experience pool EX_Buffer, d and g are input into the action model A, and the output probability action_probs of each word in opinion g is obtained. Then, the current text d and the corresponding opinion g are input into the evaluation model C to obtain the value value of opinion g, and then the experience values ​​ex_action_probs and ex_adv corresponding to the sampling phase of the current text d are pulled from the experience pool, and the action loss function is calculated using (action_probs, ex_action_probs, ex_adv):

[0075]

[0076]

[0077] Among them clip eps is a hyperparameter, which is set to 0.4 in this solution. Similarly, the experience values ​​ex_value and ex_reward of the sampling phase corresponding to the current text d are pulled from the experience pool, and the evaluation loss function is calculated using (value, ex_value, ex_reward):

[0078]

[0079]

[0080]

[0081] f clip (x i )=ex_value(x i )+min(clip eps ,max(-clip eps ,value(x i )-ex_value(x i )));

[0082] Where value(x i ) is the current critic evaluation model for the xth i The output value of the sample, ex_value(x i ) is the critic evaluation model for the xth i The output value of the sample, clip eps To truncate the hyperparameter, it is set to 0.4 in this embodiment, and ex_reward is the reward value calculated for the kth sampling.

[0083] Furthermore, in addition to calculating the action loss function and evaluation loss function of reinforcement learning, the embodiment of the present invention also makes full use of the optimal viewpoint of the text d, introduces the loss function of viewpoint generation enhancement to train the model in a good direction, and inputs the optimal viewpoint in the training sample into the action model A to optimize the training process of the model. According to the current text d, the corresponding optimal viewpoint gz is obtained from the optimal viewpoint set G, and (d, gz) is used to input into the action model A, and the output probability of the next word in the optimal viewpoint is calculated. The specific viewpoint generation enhancement loss function is:

[0084]

[0085] Where N is the number of batch samples, T is the number of opinion label tokens, p(y′ i,j |xi,y′ i,1:j-1 ) is the probability of the jth token of the i-th sample output by the actor action model A, x i is the input sentence, that is, the current text d, y′ i,1:j-1 is the word sequence that has been output by the optimal viewpoint gz of the i-th sample, y′ i,j The next word in the vocabulary sequence that has been output by the optimal view gz for the i-th sample.

[0086] Finally, according to the action loss function, the evaluation loss function, and the enhanced loss function of the viewpoint generation, the loss function of the reinforcement learning of the embodiment of the present invention is obtained: loss = actor loss +critic loss +1e-4*gloss, use the loss function loss to update the action model A and the evaluation model C. When the preset number of training times is reached, the updated action model A becomes the opinion generation model.

[0087] According to another embodiment of the present invention, after obtaining the opinions corresponding to the text data to be inferred, the method further includes: using the reward model to score the opinions corresponding to the text data to be inferred, and storing opinions with scores greater than a preset scoring threshold into an opinion library.

[0088] Specifically, after the text data to be inferred is input into a pre-constructed opinion generation model and the opinions corresponding to the text data to be inferred are obtained, in order to further ensure the quality of the opinions, the embodiment of the present invention uses the above-mentioned reward model to score the generated opinions, filters out low-scoring opinions, and stores opinions with scores greater than a preset scoring threshold into the opinion library, thereby obtaining high-quality opinions of the text to be inferred.

[0089] Figure 2It is a flow chart of the opinion generation method of an embodiment of the present invention. The text data in the text collection obtained from multiple data sources such as comments, after-sales, and questions and answers are segmented and divided into sentences according to punctuation marks to obtain a text data set. Based on the pre-trained roberta model, the pre-trained sample set of text collected in advance is used to obtain a text vector extraction model through the unsupervised training method of the comparative learning framework Simcse, and then the greedy algorithm is used to remove redundancy from the text data set to obtain the core text set coreset. The first text set is extracted from the core text set after redundancy removal, and a specified small number of texts are extracted from the first text set to accurately mark the opinions. Based on the accurate first text opinion set obtained by opinion marking, prompt learning is used to determine the text opinions in the first text opinion set through prompt design, and the context text opinions of the text opinions, and then the candidate opinion model is trained in combination with the neural network model T5-base of text-to-text migration to obtain a candidate opinion generation model. The second text opinion set obtained by the candidate opinion generation model is verified by using the manual annotation and discrimination method, and the text opinions that fail the verification are deleted. The text opinions that pass the verification and the above-mentioned accurate first text opinion set form a candidate text opinion set, and the candidate text opinion set is used to train T5-base to obtain the initial opinion generation model. The second text set is extracted from the core text set after redundancy removal, and the reward text opinion set is constructed, and then the reward model is determined; the reward model is used as the evaluation model of reinforcement learning, and the initial opinion generation model is used as the action model of reinforcement learning. Combined with the above-determined reward model, the evaluation model and the action model are trained by the reinforcement learning algorithm with opinion generation enhancement to obtain the opinion generation model after the initial opinion generation model is trained. The text data to be inferred is input into the reasoning service constructed by the opinion generation model to obtain the corresponding opinions, and then the reward model is used to score the generated opinions, filter out low-scoring opinions, and review high-scoring opinions into the opinion library.

[0090] The opinion generation method proposed in the embodiment of the present invention is based on rich data sources, adopts a low-resource sample annotation strategy, accurately annotates a small number of opinions in a text vector set, utilizes the rich knowledge contained in a large model, combines the training method of reinforcement learning, and the scoring and screening mechanism of the reward model, and introduces the idea of ​​enhancing the generation of opinions for the optimal opinions in reinforcement learning, which greatly optimizes the opinion generation model and ensures the generation of high-quality and diversified opinions.

[0091] Figure 3 Schematic diagram of the main modules of the apparatus for generating opinions according to an embodiment of the present invention. Figure 3 As shown, the apparatus 300 for generating opinions mainly includes a text acquisition module 301 and an opinion generation module 302 .

[0092] A text acquisition module 301, configured to acquire text data to be inferred in response to receiving a viewpoint generation request;

[0093] The opinion generation module 302 is used to input the text data to be inferred into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

[0094] According to one embodiment of the present invention, the apparatus 300 for generating opinions further includes an opinion generation model construction module (not shown in the figure), which is used to: before inputting the text data to be inferred into a pre-constructed opinion generation model, use a pre-acquired text data set as a training sample, and use a reinforcement learning algorithm with opinion generation enhancement to train the initial opinion generation model to construct the opinion generation model.

[0095] According to another embodiment of the present invention, the apparatus 300 for generating opinions further includes a redundancy removal module (not shown in the figure), which is used to: before using a pre-acquired text data set as a training sample, train a pre-trained model by means of a contrastive learning framework to obtain a text vector extraction model; according to the text vector extraction model, use a greedy algorithm to perform redundancy removal processing on the text data set to obtain a redundancy-free core text set, and use the core text set as the text data set.

[0096] According to another embodiment of the present invention, the apparatus 300 for generating opinions further includes an opinion generation initial model construction model (not shown in the figure), which is used for: extracting a first text set from the text data set before training the opinion generation initial model using a reinforcement learning algorithm with opinion generation enhancement, constructing a candidate opinion generation model based on the first text set in combination with a preset text-to-text migration neural network model, and obtaining a corresponding candidate text opinion set based on the candidate opinion generation model; and using the candidate text opinion set to train the text-to-text migration neural network model to obtain an opinion generation initial model.

[0097] According to another embodiment of the present invention, the opinion generation initial model construction model (not shown in the figure) is also used to: extract a specified number of texts from the first text set, mark the opinion of each of the texts, and obtain a first text opinion set; use the text opinions in the first text opinion set, and the context text opinions of the text opinions, to train a preset text-to-text migration neural network model to obtain a candidate opinion generation model; input the remaining text of the first text set into the candidate opinion generation model to obtain the opinion corresponding to the remaining text, and use the second text opinion set consisting of the remaining text and the opinion corresponding to the remaining text, and the first text opinion set as the candidate text opinion set.

[0098] According to another embodiment of the present invention, the apparatus 300 for generating opinions further includes a second text opinion set verification module (not shown in the figure), which is used to verify the second text opinion set and the first text opinion set before using the second text opinion set and the first text opinion set as candidate text opinion sets, and delete the text opinions in the second text opinion set that fail the verification.

[0099] According to another embodiment of the present invention, the opinion generation model construction module (not shown in the figure) is also used to: extract a second text set from the text data set, construct a reward text opinion set, and determine the reward model in combination with the opinion generation initial model; extract the remaining part from the text data set as a third text set, use the reward model as an evaluation model for reinforcement learning, use the opinion generation initial model as an action model for the reinforcement learning, and update the evaluation model and the action model based on the third text set in combination with a loss function with opinion generation enhancement to obtain the opinion generation model after the opinion generation initial model is trained.

[0100] According to another embodiment of the present invention, the opinion generation model construction module (not shown in the figure) is also used to: for each second text in the second text set, generate a specified number of second opinions through the opinion generation initial model, construct opinion pairs for each second text, and label the quality of the opinion pairs according to preset rules to obtain a reward text opinion set; establish an initial reward model based on the opinion generation initial model, and obtain a reward model by training the initial reward model using the reward text opinion set.

[0101] According to another embodiment of the present invention, the opinion generation model construction module (not shown in the figure) is also used to: for each text in the third text set, generate the opinion and experience value of each text through the evaluation model and the action model, store each text, the opinion of each text and the experience value in an experience pool, and determine the optimal opinion of each text according to the experience value; based on the text in the experience pool and the optimal opinion of the text, update and train the action model and the evaluation model by calculating the action loss function, evaluation loss function and opinion generation enhanced loss function of the reinforcement learning.

[0102] According to another embodiment of the present invention, the opinion generation model construction module (not shown in the figure) is also used to: obtain the current text and the corresponding opinion from the experience pool according to a preset number of training times, obtain the output probability of each word in the current opinion and the value of the current opinion in combination with the action model and the evaluation model, and calculate the action loss function and evaluation loss function of the reinforcement learning in combination with the experience value corresponding to the current text in the experience pool; obtain the opinion generation enhanced loss function based on the current text and the optimal opinion corresponding to the current text by calculating the output probability of the next word in the optimal opinion under the premise that the current text is input; update and train the action model and the evaluation model based on the action loss function of the reinforcement learning, the evaluation loss function, and the opinion generation enhanced loss function until the preset number of training times is reached.

[0103] According to another embodiment of the present invention, the device 300 for generating opinions further includes a reward scoring module (not shown in the figure), which is used to: after obtaining the opinions corresponding to the text data to be inferred, use the reward model to score the opinions corresponding to the text data to be inferred, and store opinions with scores greater than a preset scoring threshold into the opinion library.

[0104] Figure 4 is an exemplary system architecture diagram to which embodiments of the present invention can be applied.

[0105] like Figure 4 As shown, system architecture 400 may include terminal devices 401, 402, 403, network 404 and server 405. Network 404 is used to provide a medium for communication links between terminal devices 401, 402, 403 and server 405. Network 404 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0106] Users can use terminal devices 401, 402, 403 to interact with server 405 via network 404 to receive or send messages, etc. Terminal devices 401, 402, 403 can be installed with various communication client applications, such as opinion generation applications, etc. (only as an example).

[0107] The terminal devices 401 , 402 , and 403 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0108] Server 405 may be a server that provides various services, such as a background management server that provides support for the generation of opinions performed by users using terminal devices 401, 402, and 403 (only as an example). In response to receiving a request for generating an opinion, the background management server may obtain text data to be inferred; input the text data to be inferred into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, etc. The opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, which is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model, and feed back the processing results (such as opinion data, etc. - only as an example) to the terminal device.

[0109] It should be noted that the method for generating opinions provided in the embodiment of the present invention is generally executed by the server 405 , and accordingly, the device for generating opinions is generally arranged in the server 405 .

[0110] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0111] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Figure 5 The terminal device or server shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0112] like Figure 5As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0113] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage section 508 as needed.

[0114] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present invention are executed.

[0115] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination thereof.

[0116] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0117] The units involved in the embodiments of the present invention may be implemented by software or hardware. The units described may also be arranged in a processor, for example, it may be described as follows: a processor includes a text acquisition module and a viewpoint generation module.

[0118] The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, a text acquisition module can also be described as a "module for acquiring text data to be inferred in response to receiving an opinion generation request."

[0119] On the other hand, the present invention also provides a computer-readable medium, which may be included in the device described in the embodiment; or may exist independently without being assembled into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the device includes: in response to receiving a viewpoint generation request, obtaining text data to be inferred; inputting the text data to be inferred into a pre-built viewpoint generation model to obtain the viewpoint corresponding to the text data to be inferred, wherein the viewpoint generation model is obtained by performing reinforcement learning training on the viewpoint generation initial model, and the viewpoint generation initial model is obtained according to the candidate text observation set determined by the candidate viewpoint generation model and the preset text-to-text migration neural network model.

[0120] The technical solution according to the embodiment of the present invention has the following advantages or beneficial effects: by responding to receiving a request for opinion generation, the text data to be inferred is obtained; the text data to be inferred is input into a pre-constructed opinion generation model to obtain the opinion corresponding to the text data to be inferred, the opinion generation model is obtained by performing reinforcement learning training on the initial opinion generation model, and the initial opinion generation model is a technical solution obtained based on the candidate text observation set determined by the candidate opinion generation model and the preset text-to-text migration neural network model, which realizes the use of reinforcement learning algorithm and the candidate opinion generation model constructed by the low-resource sample annotation strategy to determine the initial opinion generation model, and determines the initial opinion generation model through the annotation of a small amount of high-quality samples, which realizes the steady improvement of the efficiency and quality of opinion generation under the premise of limited sample resources, thereby better serving product improvement and market decision-making.

[0121] The specific implementation methods described herein do not constitute limitations on the scope of protection of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A method for generating opinions, characterized in that: include: In response to receiving the opinion generation request, obtaining text data to be inferred; The text data to be inferred is input into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

2. The method according to claim 1, characterized in that Before inputting the text data to be inferred into the pre-built opinion generation model, the method further includes: The pre-acquired text data set is used as a training sample, and a reinforcement learning algorithm with opinion generation enhancement is used to train the initial opinion generation model to construct the opinion generation model.

3. The method according to claim 2, characterized in that Before using the pre-acquired text data set as a training sample, the method further includes: The pre-trained model is trained using a contrastive learning framework to obtain a text vector extraction model; According to the text vector extraction model, a greedy algorithm is used to perform redundancy removal processing on the text data set to obtain a redundancy-free core text set, and the core text set is used as the text data.

4. The method according to claim 2, characterized in that: Before training the initial opinion generation model using a reinforcement learning algorithm with opinion generation enhancement, the method further includes: Extracting a first text set from the text data set, constructing a candidate viewpoint generation model based on the first text set in combination with a preset text-to-text migration neural network model, and obtaining a corresponding candidate text viewpoint set based on the candidate viewpoint generation model; The candidate text opinion set is used to train the neural network model of text-to-text migration to obtain an initial opinion generation model.

5. The method according to claim 4, characterized in that Extracting a first text set from the text data set, constructing a candidate viewpoint generation model based on the first text set in combination with a preset text-to-text migration neural network model, and obtaining a corresponding candidate text viewpoint set based on the candidate viewpoint generation model, including: Extracting a specified number of texts from the first text set, and marking opinions on each of the texts to obtain a first text opinion set; Using the text opinions in the first text opinion set and the context text opinions of the text opinions, training a preset neural network model for text-to-text migration to obtain a candidate opinion generation model; The remaining text of the first text set is input into the candidate opinion generation model to obtain the opinions corresponding to the remaining text, and a second text opinion set consisting of the remaining text and the opinions corresponding to the remaining text and the first text opinion set are used as the candidate text opinion set.

6. The method according to claim 5, characterized in that Before using the second text opinion set and the first text opinion set as candidate text opinion sets, the method further includes: The second text opinion set is verified, and text opinions in the second text opinion set that fail the verification are deleted.

7. The method according to claim 2, characterized in that Using the pre-acquired text dataset as a training sample, a reinforcement learning algorithm with enhanced opinion generation is used to train the initial opinion generation model, and the opinion generation model is constructed, including: Extracting a second text set from the text data set, constructing a reward text opinion set, generating an initial model based on the opinion, and determining a reward model; The remaining part is extracted from the text data set as the third text set, the reward model is used as the evaluation model of reinforcement learning, and the initial model for generating opinions is used as the action model of the reinforcement learning. According to the third text set and in combination with a loss function with enhanced opinion generation, the evaluation model and the action model are updated and trained to obtain the opinion generation model after the initial model for generating opinions is trained.

8. The method according to claim 7, characterized in that Extracting a second text set from the text data set, constructing a reward text opinion set, generating an initial model in combination with the opinion, and determining a reward model, including: For each second text in the second text set, generating a specified number of second opinions by using the opinion generation initial model, constructing opinion pairs for each second text, and marking the quality of the opinion pairs according to a preset rule to obtain a reward text opinion set; An initial reward model is established based on the initial opinion generation model, and a reward model is obtained by training the initial reward model using the reward text opinion set.

9. The method according to claim 7, characterized in that: According to the third text set, in combination with a loss function with enhanced opinion generation, updating and training the evaluation model and the action model includes: For each text in the third text set, generating a viewpoint and an experience value of each text through the evaluation model and the action model, storing each text, the viewpoint of each text and the experience value in an experience pool, and determining the optimal viewpoint of each text according to the experience value; Based on the texts in the experience pool and the optimal viewpoints of the texts, the action model and the evaluation model are updated and trained by calculating the action loss function, the evaluation loss function of the reinforcement learning, and the loss function of viewpoint generation enhancement.

10. The method according to claim 9, characterized in that Based on the text in the experience pool and the optimal viewpoint of the text, by calculating the action loss function, the evaluation loss function of the reinforcement learning, and the loss function of viewpoint generation enhancement, the action model and the evaluation model are updated and trained, including: According to a preset number of training times, the current text and the corresponding viewpoint are obtained from the experience pool, the output probability of each word in the current viewpoint and the value of the current viewpoint are obtained by combining the action model and the evaluation model, and the action loss function and the evaluation loss function of the reinforcement learning are calculated by combining the experience value corresponding to the current text in the experience pool; According to the current text and the optimal viewpoint corresponding to the current text, by calculating the output probability of the next word in the optimal viewpoint under the premise that the current text is input, a loss function for viewpoint generation enhancement is obtained; According to the action loss function of the reinforcement learning, the evaluation loss function, and the enhanced loss function generated by the viewpoint, the action model and the evaluation model are updated and trained until the preset number of training times is reached.

11. The method according to claim 7, characterized in that After obtaining the viewpoint corresponding to the text data to be inferred, the method further includes: The reward model is used to score the opinions corresponding to the text data to be inferred, and opinions with scores greater than a preset scoring threshold are stored in an opinion library.

12. A device for generating opinions, characterized in that: include: A text acquisition module, configured to acquire text data to be inferred in response to receiving a viewpoint generation request; The opinion generation module is used to input the text data to be inferred into a pre-built opinion generation model to obtain the opinion corresponding to the text data to be inferred, wherein the opinion generation model is obtained by performing reinforcement learning training on an initial opinion generation model, and the initial opinion generation model is obtained based on a candidate text observation set determined by a candidate opinion generation model and a preset text-to-text migration neural network model.

13. A mobile electronic device terminal, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 11.

14. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.