Unsupervised Opinion Summarization Generation Method and System Based on an Adversarial Framework
The M2A framework addresses the inflexibility of domain-specific metadata-dependent summarization by using a model-agnostic adversarial model with natural language inference, improving accuracy and efficiency in generating opinion summaries.
Patent Information
- Application Number
- CN202310784057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-29
AI Technical Summary
The existing summary method of opinion is highly dependent on metadata in specific fields, lacks flexibility and versatility, and is difficult to effectively apply in different fields.
Using an adversarial framework-based unsupervised opinion summary generation method, an M2A model is constructed, including an unsupervised summary generator and discriminator, and the model is trained through cross-entropy loss function and natural language inference to improve the universality and robustness of the model.
It significantly improves the category accuracy and emotional accuracy of opinion summary, reduces the hallucination problem of generating summary, and improves the generation efficiency.
Smart Images

Figure CN116775860B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing and artificial intelligence, relates to the technical field of text summary generation, and mainly relates to an unsupervised opinion summary generation method and system based on an adversarial framework. Background Art
[0002] Artificial intelligence is a science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. With the surge in data, significant improvements in computing power, and the emergence of new algorithms, especially deep learning algorithms, artificial intelligence theory and technology are becoming increasingly mature, and the application fields are also expanding. It has initially met the conditions for implementation in various fields.
[0003] Text summarization is a traditional natural language processing task. Text summarization aims to convert text or text collections into short summaries containing key information. Text summarization technology is also one of the key technologies to improve people's efficiency in obtaining effective information in the era of information explosion. With the rapid development of online communication platforms, opinion summaries in text summaries have attracted widespread attention. The main scenarios of opinion summaries are texts that describe personal opinions and preferences, such as product reviews, blog content, or social media texts. In this task, it is crucial to identify the most important aspects of a product or event and the most prominent emotions of users, and combine them to generate a concise and fluent summary that conforms to the facts. Amazon, as a traditional opinion summary dataset, contains a large number of product-related reviews. Current methods mainly focus on unsupervised generative scenarios. Generative summaries generate summaries by understanding the meaning of the original text. In fact, they imitate the way humans write summaries. They may use words in the original text or new words (words that do not appear in the original text) to express them. Compared with the extractive method, the words used are more flexible.
[0004] Many existing opinion summarization methods focus on metadata such as sentiment polarity and aspect information in specific fields, and even directly use sentiment tags and method information to make the model more effective. These methods are constrained by the availability of metadata in specific fields and lack flexibility and versatility. Therefore, how to overcome the use of metadata in specific fields to make the model more flexible has become an active research direction for technicians in this field. Summary of the invention
[0005] The present invention aims at the problem that the existing technology is too dependent on metadata in specific fields, and provides an unsupervised opinion summary generation method and system based on an adversarial framework. First, a data set is constructed, and test data is selected from the database, and the data is preprocessed, and the data for verification and testing are divided to construct an unsupervised opinion summary data set; then, M is constructed. 2 A model, select two generation summary models as summary generators to embody M2 The model - independence of A, constructing an abstract discriminator for the model based on the method of natural - language inference; then training the abstract discriminator first, and then training M as a whole 2 A model, constructing a loss function and an optimizer; finally, evaluating M 2 A model. The present invention significantly improves the category accuracy and sentiment accuracy in generating opinion abstracts, and greatly improves the working efficiency of abstract generation.
[0006] To achieve the above - mentioned purpose, the technical solution adopted by the present invention is: an unsupervised opinion - abstract generation method based on an adversarial framework, including the following steps:
[0007] S1, dataset construction: Selecting data to be measured in the database, pre - processing the data, dividing the data for verification and testing, and constructing an unsupervised opinion - abstract dataset;
[0008] S2, M 2 A model construction: The model generates an adversarial network GAN, which consists of an unsupervised abstract generator and an abstract discriminator based on natural - language inference;
[0009] S3, M 2 A model training: First training the abstract discriminator, and then training M as a whole 2 A model; when training the abstract discriminator, using a cross - entropy loss function; in the overall training of M 2 A model, the loss function for constructing the abstract generator is based on the objective function of VAE, and parameter updates are performed by minimizing the sum of all sub - task loss functions, and an optimizer is constructed;
[0010] S4, M 2 A model evaluation: The evaluation indexes of the model include ROUGE indexes, class - accuracy indexes, and sentiment - accuracy indexes.
[0011] As an improvement of the present invention, Coop and Copycat are used as the abstract generators in step S2, where
[0012] Coop as the abstract generator includes: a generator encoder, a projector, a generator decoder, and an inferencer. The generator encoder uses the final hidden state obtained by a bidirectional long - short - term memory network as the representation of the comment sequence; in the projector, the hidden state is projected through an affine transformation to obtain the distribution of latent variables; resampling on the obtained distribution, passing through the generator decoder to obtain the latent variables of the comment, using a long - short - term memory network to obtain the reconstructed comment, and finally through the inferencer, selecting the best abstract based on the input - output word overlap between the generated abstract and the comment input in the abstract selector;
[0013] The Copycat as an abstract generator includes: a generator encoder, an inference network, and a generator decoder. The generator encoder embeds the words of the review into a GRU encoder to calculate the hidden state, then calculates the Gaussian distribution through the inference network, and finally the generator decoder calculates the abstract distribution and finally generates the abstract.
[0014] As an improvement of the present invention, when the Coop is used as an abstract generator, 3 inference networks are used to obtain e, z i 、z s respectively, where e represents the latent variable of the review set R e ,z i represents the latent variable of a single review r i ,z s represents the latent variable of the abstract;
[0015] First, calculate the weight α of each word in each input review ij ,and calculate the weighted sum corresponding to each word according to the weight α ij ,so as to obtain the intermediate representation of e ; use an affine transformation to obtain the parameters for calculating the Gaussian distribution of e:
[0016]
[0017]
[0018] where A e 、 G e 、 are the parameters of the affine transformation, and μ φ (R e ) and σ φ (R e ) represent the mean and variance of the Gaussian distribution of e respectively; use reparameterization to obtain e from this Gaussian distribution ,and then splice e with the hidden state of the last word of the generator encoder and obtain the Gaussian distribution of z i through an affine transformation
[0019]
[0020]
[0021] where T i represents the number of words in the review r i ,A z 、 G z 、 are the parameters of the affine transformation, μ φ (z i ) and σ φ (z i ) represent the mean and variance of the Gaussian distribution of z i respectively, and z is obtained from this Gaussian distribution using reparameterization i ;
[0022] The Gaussian distribution of z is obtained by applying an affine transformation to e s where A
[0023]
[0024]
[0025] where A s 、 G s 、 are the parameters of the affine transformation, μ φ (z s ) and σ φ (z s ) represent the mean and variance of the Gaussian distribution of z s respectively, and z is obtained from this Gaussian distribution using reparameterization s .
[0026] As another improvement of the present invention, ESIM is used as the abstract discriminator in step S2. The ESIM includes an encoder layer, an interaction layer, an inference synthesis layer, and an output layer. Among them, a bidirectional long short-term memory network is used as the encoder, and local inference modeling is performed in the interaction layer to obtain local inference information, and then the local inference information is mixed and transformed through average pooling and max pooling to obtain the output of the final abstract discriminator.
[0027] As another improvement of the present invention, in the interaction layer of the abstract discriminator, local inference modeling is performed to calculate the local inference attention weights of each pair and Specific calculations are as follows: Specifically, the calculations are as follows:
[0028]
[0029]
[0030] where refers to the hidden state of the premise at the i-th moment, refers to the hidden state of the hypothesis at the j-th moment, represents 's weighted hidden state, represents The weighted hidden state, represents the hypothesis premise p k , represents the hypothesis o k The number of, k represents the number of hypotheses or premises;
[0031] For local reasoning And the results of their bitwise subtraction and bitwise multiplication are respectively concatenated to obtain local reasoning information and
[0032] As another improvement of the present invention, the cross-entropy loss function of the training abstract discriminator is as follows:
[0033]
[0034] Where for each discriminator's final representation v k , v k′ represents the positive sample, v j represents the negative sample in the same training sample block, sim() represents the cosine similarity function, n represents the size of the number of samples batch-size taken in one training during the training process, τ is the temperature coefficient, represents the indicator function.
[0035] As another improvement of the present invention, the method is implemented based on the deep learning framework Pytorch. During the training process, the learning rate of Copycat as the abstract generator is 0.0005, the learning rate of Coop as the abstract generator is 0.001, the learning rate of the abstract discriminator is 0.0004, the optimizer is Adam, the temperature setting of beam search is 0.1, and the regularization method uses Dropout to prevent overfitting.
[0036] To achieve the above object, the technical solution adopted by the present invention is also: an unsupervised opinion abstract generation system based on an adversarial framework, including a computer program, characterized in that: when the computer program is executed by a processor, it implements the steps of any one of the above methods.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] (1) The present invention uses a model-independent and metadata-independent adversarial framework for unsupervised opinion summarization, which is the first attempt to use an adversarial framework in the field of unsupervised opinion summarization.
[0039] (2) The present invention uses unsupervised contrastive learning to retrain the natural language inference discriminator, which can align the features of the generation task with the knowledge in the discriminator without using metadata and annotations, improving the generality of the opinion summary model.
[0040] (3) The present invention uses two different opinion summary models as the summary generator part, and the effects are significantly improved compared with the original model, which reflects the flexibility of the present invention and enhances the robustness of the opinion summary model.
[0041] (4) The experimental results of the present invention on the Amazon review dataset show that when evaluating the faithfulness of the summary in terms of category accuracy and sentiment accuracy, the present invention is significantly superior to other state-of-the-art baseline works, reducing the hallucination problem of the generated summary. Description of the Drawings
[0042] Figure 1 is the flowchart of the steps of the method of the present invention;
[0043] Figure 2 is the flowchart of the steps of constructing an unsupervised opinion summary dataset in step S1 of Embodiment 1 of the present invention;
[0044] Figure 3 is for step S2 of the method of the present invention to construct M 2 A model flowchart;
[0045] Figure 4 is for step S3 of the method of the present invention to train M 2 A model flowchart;
[0046] Figure 5 is for step S4 of the method of the present invention to evaluate M 2 A model schematic diagram;
[0047] Figure 6 is the M constructed in step S2 of the method of the present invention 2 A model schematic diagram
[0048] Figure 7 is for using Coop as the training and inference flowchart of the M 2 A model summary generator in Embodiment 1 of the present invention;
[0049] Figure 8 is for using Copycat as the training and inference flowchart of the M 2 A model summary generator in Embodiment 1 of the present invention. Detailed Embodiments
[0050] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0051] Embodiment 1
[0052] Unsupervised Opinion Summarization Generation Method Based on an Adversarial Framework. In this embodiment, taking the acquisition of effective review summaries in Amazon product reviews as an example, as Figure 1 shown, the method includes the following steps:
[0053] Step S1: Construct an unsupervised opinion summary dataset.
[0054] Select data of four product categories in the Amazon product review dataset, preprocess the data, divide the product data for verification and testing, and construct an unsupervised opinion summary dataset, as Figure 2 shown.
[0055] In this embodiment, an unsupervised opinion summarization task is performed on the Amazon product review dataset. Each review in this dataset is attached with a 5-star rating. In this embodiment, review data of four categories are selected from the dataset: electronic devices; clothing, shoes, and jewelry; home and kitchen; health and personal care.
[0056] Preprocess the data. Only product data with reviews of 20 to 70 words each and more than 10 reviews will be used, and those particularly popular products will be excluded to avoid affecting model training. For evaluation, the present invention uses 60 products, and each product has 3 manually created summaries, among which 32 products are used for verification and 28 for testing.
[0057] Step S2: Construct the M 2 A model. The M 2 A model, whose full English name is Model-agnostic and Metadata-free Adversarial Framework, meaning a model-agnostic and metadata-free adversarial framework.
[0058] Step S2 is used to construct a summary generator and a summary discriminator; select two summary generation models as the summary generator respectively to reflect the model-agnostic property of M 2 A, and construct the summary discriminator of the model based on the method of natural language inference, as Figure 3 shown. 2.1 Construct the overall framework of the M 2 A model
[0059] The M 2 A model proposed by the present invention is similar to the generative adversarial network GAN and consists of an unsupervised summary generator and a summary discriminator based on natural language inference. Specifically, as Figure 6 shown.
[0060] The review set Re is input into the summary generator to generate a summary s. The summary s and the review set Re are segmented into sentences and input into the data recombination module. Each sentence in each summary s is paired with each review R eThe sentences in it form a pair of premises p k and hypothesis o k All the "premise-hypothesis pairs" are input into the NLI discriminator to evaluate whether the hypothesis can be inferred from the premise.
[0061] 2.2 Construct M 2 The abstract generator in model A
[0062] Since the method of the present invention is model-independent, any abstract generation model can be selected as M 2 The abstract generator of model A. In this implementation, two representative abstract generation models are selected as the abstract generators: Coop and Copycat.
[0063] Using Coop as the abstract generator, as Figure 7 shown:
[0064] a) Generator encoder: The hidden state h i finally obtained by using the bidirectional long short-term memory network is used as the representation of the comment sequence r i where i represents the comment number in the comment R e r i represents the i-th comment in the comment set R e h i represents the hidden state of the i-th comment at a certain moment in the bidirectional long short-term memory network.
[0065] b) Projector: Project the hidden state h i through an affine transformation to obtain the distribution of the latent variable, the mean μ(i) and variance σ(i) of the distribution. Where A, G, b A b G are the parameters of the affine transformation.
[0066] μ(i) = Ah i +b A
[0067] logσ(i) = Gh i +b G
[0068] c) Generator decoder: Resample on the obtained distribution to obtain the latent variable z i of the comment, and use the long short-term memory network to obtain the reconstructed comment r i '.
[0069] d) Inference engine: Select the best abstract through the input-output word overlap between the generated abstract and the comments input in the abstract selector.
[0070] The structure using Copycat as the abstract generator is as Figure 8As shown below:
[0071] a) Generator Encoder: Embed the words w of the comment ij Input a GRU encoder to calculate the hidden state h ij . Where i represents the comment number in the comment R e , j represents the word sequence number in a single comment r i , w ij represents the j-th word in the comment r i , h ij represents the hidden state of w ij .
[0072] b) Inference Network: Use 3 inference networks to obtain e, z i , z s , respectively, with the aim of calculating the Gaussian distribution. Where e represents the latent variable of the comment set R e , z i represents the latent variable of a single comment r i , z s represents the latent variable of the summary.
[0073] First, calculate the weight α of each word in each input comment ij , and calculate the weighted sum corresponding to each word according to the weight α ij to obtain the intermediate representation of e Use an affine transformation to obtain the parameters for finding the Gaussian distribution of e:
[0074]
[0075]
[0076] where A e , G e , are the parameters of the affine transformation, μ φ (R e ) and σ φ (R e ) represent the mean and variance of the Gaussian distribution of e respectively. Use reparameterization to obtain e from this Gaussian distribution . Then concatenate e with the hidden state of the last word of the generator encoder and obtain the Gaussian distribution of z i through an affine transformation
[0077]
[0078]
[0079] where T i denotes the number of words in the review r i , A z , G z , are the parameters of the affine transformation, and μ φ (z i ) and σ φ (z i ) represent the mean and variance of the Gaussian distribution of z i respectively. z i is obtained from this Gaussian distribution using reparameterization.
[0080] The Gaussian distribution of z s is obtained by applying an affine transformation to e
[0081]
[0082]
[0083] where A s , G s , are the parameters of the affine transformation, and μ φ (z s ) and σ φ (z s ) represent the mean and variance of the Gaussian distribution of z s respectively. z s is obtained from this Gaussian distribution using reparameterization.
[0084] c) Generator - decoder: It consists of a GRU decoder with an attention mechanism and a pointer - generator network, which calculates the distribution p θ (r i |z i , R -i ) of the input review, retaining the details in the input review set, and calculates the distribution p θ (s|z s , R e ) of the summary s, and finally generates the summary s. Here, R -i is the set of reviews in R e excluding the review r i .
[0085] 2.3M 2 The summary discriminator in the A model
[0086] ESIM, whose full name is Enhanced Sequential Inference Model, is called Enhanced Sequential Inference Model in Chinese. It is a chained LSTM sequential inference model for natural language inference.
[0087] The present invention uses ESIM as an abstract discriminator:
[0088] a) Encoder layer: Use a bidirectional long short-term memory network as the encoder. The abstract s and the comment set R output by the abstract generator of this framework e After passing through the data reconstruction module, they are input into the encoder to construct the hidden state of the premise at the i-th moment and the hidden state of the hypothesis at the j-th moment
[0089] b) Interaction layer: Local inference modeling, calculate each pair of and local inference attention weights The specific calculation is as follows:
[0090]
[0091]
[0092] where represents the weighted hidden state of , represents the weighted hidden state of . Among them, represents the hypothesis premise p k , represents the hypothesis o k the number of, i represents the i-th moment of the premise hidden state, j represents the j-th moment of the hypothesis hidden state, and k represents the number of the hypothesis or premise.
[0093] For local inference and the results of their element-wise subtraction and element-wise multiplication are concatenated respectively,, to obtain local inference information and
[0094] c) Inference synthesis layer and output layer: In order to mix local inference information, and are input into another bidirectional long short-term memory network inference part to calculate the premise sentence representation and the hypothesis sentence representation and and are transformed through average pooling and max pooling to obtain the output v of the final abstract discriminator k , the specific calculation is as follows:
[0095] v k = [v p,ave ; v p,max ; v o,ave ; o h,max
[0096] Among them, ave represents average pooling, and max represents max pooling.
[0097] Finally, by sending v k into the multi-layer perceptron classifier, the overall relationship is obtained.
[0098] Step S3: Train the M 2 A model.
[0099] In this step, first train the abstract discriminator to make it have a certain discrimination ability, and then train the M 2 A model as a whole, construct a loss function, and construct an optimizer, as Figure 4 shown.
[0100] 3.1 Abstract discriminator training
[0101] The method of the present invention needs to first train the abstract discriminator to make it have a certain discrimination ability. First, pre-train the ESI M model on the MNLI dataset to let the model remember the relevant knowledge of inferring sentence entailment relationships. Then, it is necessary to re-train on the abstract dataset using the method of unsupervised contrast learning to solve the problem of domain inconsistency between the abstract generator and the abstract discriminator. Extract sentence pairs in the Amazon dataset as training data for unsupervised contrast learning through the Next Sentence Prediction (NSP) criterion or the N-gram Similarity (SS) criterion.
[0102] For the unsupervised contrast learning part of training the discriminator, the cross-entropy loss function is adopted, and the formula is as follows:
[0103]
[0104] Among them, for the final representation v of each discriminator k , v k′ represents the positive sample, v j represents the negative sample in the same training sample block, sim() represents the cosine similarity function, n represents the size of the number of samples batch-size taken in one training during the training process, τ is the temperature coefficient, represents the indicator function.
[0105] When training, using this loss function can make the representations of positive and negative samples in the vector space farther apart, which helps to enhance the discrimination ability of the abstract discriminator.
[0106] 3.2 M 2 A Model Training
[0107] In the abstract generator part of the present invention, Coop or Copycat is adopted, and the loss functions for constructing the generators in this part are all based on the objective function of the traditional VAE.
[0108] Coop as the loss function of the generator:
[0109]
[0110] Where β is the hyperparameter for constraining the KL divergence, and θ, φ are the parameters of the model encoder and decoder.
[0111] Copycat as the loss function of the generator:
[0112]
[0113] The method of the present invention introduces a loss When the abstract is segmented into sentences before entering the discriminator, penalizing particularly short sentences at the sentence level often has good effects in the early training process.
[0114] In order to make the generated abstract more reasonable and well-founded, for each sentence in the abstract, it is required that there is at least one sentence in the comment to support it. Using the cross-entropy loss function, the loss function for constructing the abstract discriminator is:
[0115]
[0116] Where L s is the number of sentences in the abstract s, and y i is the label of the premise-hypothesis pair, and the default setting is 1.
[0117] Thus, in the process of training the MA total model, the objective function of the model is the sum of the loss functions of all subtasks, and the parameters are updated by minimizing this sum of the objective function. 2 The overall objective function of the model of the present invention
[0118] is the sum of the loss functions of all subtasks, specifically: is the sum of the loss functions of all subtasks, specifically:
[0119]
[0120] For the part of constructing the optimization function, all the deep learning models used in this embodiment are implemented based on the deep learning framework Pytorch. During the training process, the learning rate of Copycat as the summary generator is set to 0.0005, the learning rate of Coop as the summary generator is set to 0.001, the learning rate of the summary discriminator is set to 0.0004, the optimizer is Adam, the temperature of beam search is set to 0.1, and the regularization method uses Dropout to prevent overfitting.
[0121] Step S4: Evaluate M 2 Model A
[0122] The evaluation metrics in this step are divided into three parts: ROUGE metrics, class accuracy metrics, and sentiment accuracy metrics, comprehensively evaluating the faithfulness of the generated opinion summary, as Figure 5 shown
[0123] ROUGE metrics are common evaluation metrics in the field of summary generation. ROUGE calculates the corresponding score by comparing the summary generated by the model with the reference summary (usually generated manually) to evaluate the quality of the method.
[0124] To better evaluate the faithfulness of the summary, the class accuracy metric and the sentiment accuracy metric are used to judge whether the product information and the sentiment polarity of the generated summary are correct.
[0125] For the class accuracy metric, the present invention trains a BERT-based classification model to predict the input class. The number of classes is 4, corresponding to the four classes of the Amazon dataset selected by the present invention: electronic devices; clothing, shoes, and jewelry; home and kitchen; health and personal care. Each summary is divided into sentences to predict the class.
[0126] For the sentiment accuracy metric, the present invention uses the BERT-based interface of the existing sentiment analysis model to calculate the sentiment accuracy of the overall summary. The gold label is the average rating of the relevant reviews (the rating given for each review in the Amazon dataset is close to the sentiment polarity).
[0127] The evaluation metric finally constructed by the present invention combines the three category metrics and adopts a more comprehensive model selection strategy rather than only based on ROUGE-L. First, in the evaluation stage, select the top 5 models ranked by ROUGE-L. Then, select the model with the best class accuracy as the final result.
[0128] The model of the present invention has achieved better results than the advanced models on the Amazon dataset, as shown in Table 1 specifically.
[0129] Table 1: Experimental results of summaries on Amazon
[0130]
[0131] As can be seen from the above table, the model of the present invention has been compared with some recent advanced methods. The experimental results show that the method of the present invention has been greatly improved, especially in terms of the category accuracy and sentiment accuracy of generating opinion summaries, indicating that the model of the present invention has a significant improvement in the effectiveness of unsupervised opinion summaries.
[0132] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. An unsupervised opinion summary generation method based on an adversarial framework, characterized in that The steps are as follows: S1. Dataset construction: Select the data to be tested in the database, preprocess the data, divide the data for verification and testing, and construct an unsupervised opinion summary dataset; S2, M 2 A model construction: The model is a generative adversarial network (GAN), which consists of an unsupervised abstract generator and an abstract discriminator based on natural language inference; among them, Coop and Copycat are used as the abstract generators, Coop, as a summary generator, includes: a generator encoder, a projector, a generator decoder, and an inferencer. The hidden state finally obtained by the generator encoder using a bidirectional long short-term memory network is used as the representation of the comment sequence; in the projector, the hidden state is projected through an affine transformation to obtain the distribution of latent variables; resampling is performed on the obtained distribution, and through the generator decoder, the latent variables of the comment are obtained, the long short-term memory network is used to obtain the reconstructed comment, and finally through the inferencer, the input-output word overlap between the generated summary and the comment input in the summary selector is used to select the best summary; Copycat, as a summary generator, includes: a generator encoder, an inference network, and a generator decoder. The word embedding of the comment is input into a GRU encoder by the generator encoder to calculate the hidden state, then the Gaussian distribution is calculated through the inference network, and finally the summary distribution is calculated by the generator decoder and the summary is finally generated; ESIM is used as a summary discriminator. The ESIM includes an encoder layer, an interaction layer, an inference synthesis layer, and an output layer; among them, a bidirectional long short-term memory network is used as the encoder, and local inference modeling is performed in the interaction layer to obtain local inference information, and then the local inference information is mixed, and the final output of the summary discriminator is obtained through average pooling and max pooling transformations; S3, M 2 A model training: First train the abstract discriminator, and then train M as a whole 2 A model; when training the abstract discriminator, use the cross-entropy loss function; train M as a whole 2 In the A model, the loss functions for constructing the abstract generator are all based on the objective function of VAE, and the parameters are updated by minimizing the combination of all subtask loss functions to construct an optimizer; S4, M 2 Evaluation of Model A: The evaluation metrics of the model include ROUGE metrics, class accuracy metrics, and sentiment accuracy metrics.
2. The unsupervised opinion summary generation method based on an adversarial framework according to claim 1, wherein: When the Coop is used as a summary generator, three inference networks are used to obtain e and z respectively. i , z s , where e represents the latent variable of the comment set R e , z i represents the latent variable of a single comment r i , z s represents the latent variable of the summary; First, calculate the weight α of each word in each input comment ij , and according to the weight α ij calculate the weighted sum corresponding to each word to obtain the intermediate representation of e Use an affine transformation to obtain the parameters for the Gaussian distribution of e: where A e , G e , are the parameters of the affine transformation, and μ φ (R e ) and σ φ (R e ) represent the mean and variance of the Gaussian distribution of e respectively; e is obtained from this Gaussian distribution using reparameterization, and then e is concatenated with the hidden state of the last word of the generator encoder to obtain the Gaussian distribution of z through an affine transformation i where T i denotes the number of words in the comment r i , A z , G z , are the parameters of the affine transformation, μ φ (z i ) and σ φ (z i ) represent the mean and variance of the Gaussian distribution of z i respectively, and z i is obtained from this Gaussian distribution using reparameterization; Obtain z by applying an affine transformation to e s with a Gaussian distribution Among them, A s , G s , are the parameters of the affine transformation, and μ φ (z s ) and σ φ (z s ) represent the mean and variance of the Gaussian distribution of z s respectively, and z s is obtained from this Gaussian distribution using reparameterization.
3. The unsupervised opinion summary generation method based on an adversarial framework according to claim 2, characterized in that: In the interaction layer of the abstract discriminator, for local inference modeling, calculate the local inference attention weights for each pair of and as follows: The specific calculation is as follows: Among them, refers to the hidden state premised at the i-th moment, refers to the hidden state hypothesized at the j-th moment, represents the weighted hidden state of represents the weighted hidden state of indicates the hypothesized premise p k , indicates the hypothesized o k the number of, k represents the number of hypotheses or premises; For local reasoning And splice the results of their bitwise subtraction and bitwise multiplication respectively to obtain local reasoning information and 4. The unsupervised opinion summary generation method based on an adversarial framework according to claim 1, wherein: The cross-entropy loss function for training the summary discriminator is as follows: Among them, for the final representation v of each discriminator k , v k′ represents the positive sample, v j represents the negative sample in the same training sample block, sim() represents the cosine similarity function, n represents the size of the number of samples batch-size taken in one training during the training process, τ is the temperature coefficient, represents the indicator function.
5. The unsupervised opinion summary generation method based on an adversarial framework according to claim 1, wherein: The method is implemented based on the deep learning framework Pytorch. During the training process, the learning rate of Copycat as a summary generator is 0.0005, the learning rate of Coop as a summary generator is 0.001, the learning rate of the summary discriminator is 0.0004, the optimizer is Adam, the temperature of beam search is set to 0.1, and the regularization method uses Dropout to prevent overfitting.
6. An unsupervised opinion summary generation system based on an adversarial framework, including a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-5 above.