Film review emotion intelligent discrimination method based on improved GRU
By improving the GRU model to Mem_GRU, adding memory cell structure and optimizing loss function, the problem that the GRU model is difficult to capture contextual relationships in film review sentiment analysis is solved, improving the accuracy of film review classification, and helping producers better understand audience feedback.
Patent Information
- Application Number
- CN202510419555.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing GRU model is difficult to effectively capture the key emotional factors in contextual relationships in the emotional analysis of film reviews, resulting in inaccurate classification results of film reviews and affecting the producer's insight into audience preferences and dissatisfaction.
Based on the GRU model, the Mem_GRU model is improved and the memory cell structure similar to LSTM is added, and the information flow is controlled through forgetting gate, input gate, reset gate, update gate and output gate, and the binary cross entropy loss function optimization model is adopted.
It enhances the ability to handle complex and long-term dependencies in film reviews, improves the accuracy of emotional analysis of film reviews, and allows producers to more accurately understand the audience's preferences and dissatisfaction with the movie, providing more valuable references for the subsequent promotion and creation of the movie.
Smart Images

Figure CN120336959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a method for intelligent discrimination of movie review sentiment based on an improved GRU. Background Art
[0002] In the current field of movie review sentiment analysis, neural networks are widely used in sentiment analysis tasks; in the prior art, the CNN-LR model has been proposed, and experiments are carried out by crawling movie-related review data on the Douban movie website. Although the evaluation indexes are improved to a certain extent, there are still limitations. When the GRU model is used to process movie review sentiment analysis, due to its lack of memory function itself, it is difficult to effectively capture the key sentiment factors in the context relationship when analyzing multiple movie reviews of a certain reviewer. This results in an inability to accurately classify whether a certain comment is positive or negative, and ultimately has a greater impact on the classification result of the comment.
[0003] These defects of the prior art make it difficult for the accuracy of movie review sentiment analysis to meet the actual needs. For example, producers cannot accurately understand the preferences and dissatisfaction of audiences with movies through analyzing movie reviews, which in turn affects the decision-making of subsequent movie promotion, improvement, and creation of future works. Therefore, a method for intelligent discrimination of movie review sentiment based on an improved GRU is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art, and a method for intelligent discrimination of movie review sentiment based on an improved GRU is proposed.
[0005] The method for intelligent discrimination of movie review sentiment based on an improved GRU includes the following steps:
[0006] S1. Using web crawler technology to obtain movie review data from a movie platform;
[0007] S2. Preprocessing the obtained movie review data;
[0008] S3. Randomly dividing the preprocessed movie review data into a training set and a test set according to a ratio of 1:1;
[0009] S4. Constructing a Mem_GRU model based on the GRU model, and training the Mem_GRU model using the training set;
[0010] S5. Using the test set to test the trained Mem_GRU model, and evaluating the performance of the Mem_GRU model according to the accuracy rate, precision rate, recall rate, and F1 value.
[0011] Preferably, in the step S1, the movie review data includes the sentiment label of the movie review, the review time, and the detailed review content.
[0012] Preferably, in the step S2, the preprocessing of the movie review data includes word segmentation, stop word removal, and word vector construction.
[0013] Preferably, in the step S4, constructing the Mem_GRU model includes adding a memory cell structure similar to LSTM to the GRU model, introducing a new cell state c t , and implementing the functions of the model by designing a forget gate f t , an input gate i t , a reset gate r t , an update gate Z t and an output gate i t ;
[0014] The forget gate f t controls how much information is forgotten from the memory unit c t , and its calculation formula is:
[0015] f t = σ(W f x t + U f h t-1 + b f );
[0016] Design an input gate i t to control how much new information can enter the memory unit C t , and its calculation formula is:
[0017] i t = σ(W i x t + U i h t-1 + b i );
[0018] The candidate cell state of the new information entering the memory unit The calculation formula is:
[0019]
[0020] Use the reset gate r t to control whether the candidate state ( depends on the previous moment state h t-1 , and its calculation formula is:
[0021]
[0022] Then, the update gate Z t determines how much historical information to retain, and its calculation formula is:
[0023]
[0024] Design an output gate o t , and the calculation formula is:
[0025] o t = σ(W o x t + U o h t-1 + b o );
[0026] Finally, the output calculation formula is:
[0027] y t = o t ⊙ h t .
[0028] Preferably, in the step S4, a binary cross-entropy loss function is used to optimize the Mem_GRU model, and the formula of the binary cross-entropy loss function is:
[0029]
[0030] where it is assumed that there are N samples, and the true label of the i-th sample is y i , and the predicted probability that the model belongs to the positive class for this sample is Then the binary cross-entropy loss of this batch of samples is the average value of the losses of individual samples.
[0031] Compared with the existing technology, the advantages of the present invention are as follows:
[0032] 1. The Mem_GRU model proposed by the present invention is an improvement based on the GRU model. The newly added memory unit structure is similar to that of the LSTM, and its structure is more complex and perfect, capable of better processing long-term dependence information. When dealing with the movie review sentiment analysis task, the Mem_GRU model enhances the ability to handle complex long-term dependence relationships in movie reviews.
[0033] 2. The Mem_GRU model proposed by the present invention is used for intelligent discrimination of movie review sentiment, enabling producers to more accurately understand the preferences and dissatisfaction of audiences towards movies, providing more valuable references for subsequent movie promotion, improvement, and creation of future works. At the same time, it also helps production parties formulate more efficient marketing strategies, expanding the influence and box office revenue of movies. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of the present invention.
[0035] Figure 2 is a structural diagram of the Mem_GRU model in the present invention.
[0036] Figure 3 is a comparison chart of the training amounts of different models on the W dataset and the IMDB dataset in the present invention. Detailed implementation manners
[0037] To make the technical means, creative features, achieved purposes and functions of the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation manners.
[0038] Referring to Figure 1-2 As shown, the method for intelligent discrimination of movie review sentiment based on improved GRU includes the following steps:
[0039] S1. Use web crawler technology to obtain movie review data from movie platforms;
[0040] S2. Preprocess the obtained movie review data;
[0041] S3. Randomly divide the preprocessed movie review data into a training set and a test set according to a ratio of 1:1;
[0042] S4. Build a Mem_GRU model based on the GRU model, and use the training set to train the Mem_GRU model;
[0043] S5. Use the test set to test the trained Mem_GRU model, and evaluate the performance of the Mem_GRU model according to accuracy, precision, recall rate and F1 value.
[0044] In the step S1, the movie review data includes sentiment labels (positive, negative, neutral) of the movie review, review time and detailed review content.
[0045] In the step S2, preprocessing the movie review data includes word segmentation, removing stop words and constructing word vectors, and the Word2Vec method is used to construct word vectors.
[0046] In the step S4, building the Mem_GRU model includes adding a memory cell structure similar to LSTM to the GRU model, introducing a new cell state c t , and realizing the function of the model by designing a forgetting gate f t , an input gate i t , a reset gate r t , an update gate Z t and an output gate i t ;
[0047] The forgetting gate f t controls how much information is forgotten from the memory unit c t , and the calculation formula is:
[0048] f t =σ(W f x t +U f h t-1 +bf )
[0049] Design an input gate i t to control how much new information can enter the memory cell C t , and the calculation formula is:
[0050] i t = σ(W i x t + U i h t-1 + b i )
[0051] The candidate cell state for new information to enter the memory cell The calculation formula is:
[0052]
[0053] Use the reset gate r t to control whether the candidate state ( depends on the previous state h t-1 , and the calculation formula is:
[0054]
[0055] Then, the update gate Z t determines how much historical information to retain, and the calculation formula is:
[0056]
[0057] Design an output gate o t , and the calculation formula is:
[0058] o t = σ(W o x t + U o h t-1 + b o )
[0059] Finally, the output calculation formula is:
[0060] y t = o t ⊙ h t .
[0061] In step S4, the binary cross-entropy loss function is used to optimize the Mem_GRU model, and the binary cross-entropy loss function formula is:
[0062]
[0063] Assume there are N samples, and the true label of the i-th sample is y i, the predicted probability of the model that the sample belongs to the positive class is Then the binary cross-entropy loss of this batch of samples is the average of the losses of individual samples.
[0064] Embodiment
[0065] S1. In practical applications, taking a certain movie as an example, first use web crawler technology to obtain a large amount of review data about the movie from a movie platform. Send an HTTP request using the requests library in Python, perform URL encoding in combination with the urllib.parse library, and then use the etree module in the lxml library to parse the HTML page to obtain the sentiment tags (positive, negative, neutral) of the reviews, the review time, and the detailed review content;
[0066] S2. Preprocess the obtained original data, including operations such as word segmentation, stop word removal, and construction of word vectors;
[0067] S3. Randomly divide the preprocessed data into a training set and a test set according to a ratio of 1:1;
[0068] S4. Under the TensorFlow 2.18.0 and keras frameworks, use the Python 3.12.8 language to build a Mem_GRU model in the PyCharm Community Edition 2024.1.1 integrated development environment, use the training set to train the model, and continuously optimize the model during the training process by adjusting the model parameters;
[0069] S5. After the training is completed, use the test set to test the Mem_GRU model, evaluate the model performance according to evaluation indicators such as accuracy, precision, recall rate, and F1 value. If the performance of the Mem_GRU model does not meet the expectations, the model parameters can be adjusted or the amount of training data can be increased, and then train and test again until a satisfactory effect is achieved.
[0070] At the same time, this experiment also adds the GRU model and the CNN_LR model to conduct tests together, obtains the experimental results of each model on the W dataset and the IMDB dataset respectively, and the results of different models on the accuracy, precision, recall rate, and F1 value evaluation indicators are as follows in the table:
[0071] Table 1 Evaluation indicators of different models on the W dataset
[0072]
[0073] Table 2 Evaluation indicators of different models on the IMDB dataset
[0074]
[0075] The experimental results show that on the W dataset, compared with the GRU model, the Mem_GRU model has an 18% increase in accuracy, a 12% increase in precision, an 18% increase in recall, and a 21% increase in F1 value; compared with the CNN_LR model, the Mem_GRU model has a 4% increase in accuracy, a 3% increase in precision, a 4% increase in recall, and a 4% increase in F1 value. On the publicly available IMDB dataset, compared with the GRU model, all indicators of the Mem_GRU model have also improved, with accuracy, precision, recall, and F1 value all increasing by 2%. Compared with the CNN_LR model, the Mem_GRU model has a 7% increase in accuracy, a 6% increase in precision, a 7% increase in recall, and a 7% increase in F1 value.
[0076] In terms of structure, the newly added memory unit structure of the Mem_GRU model is similar to that of the LSTM. Compared with the GRU model, its structure is more complex and perfect, and it can better handle long-term dependence information. Functionally, when dealing with the movie review sentiment analysis task, the Mem_GRU model enhances the ability to handle complex long-term dependence relationships in movie reviews, which enables producers to more accurately understand the preferences and dissatisfaction of the audience towards movies, provides more valuable references for the subsequent publicity, improvement, and creation of future works of movies, and also helps the production side formulate more efficient marketing strategies to expand the influence and box office revenue of movies.
[0077] As is known by common technical knowledge, the present invention can be implemented by other embodiments that do not depart from its spiritual essence or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and are not the only ones. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.
Claims
1. An intelligent discriminant method for movie review sentiment based on improved GRU, characterized in that: It includes the following steps: S1. Obtain movie review data from a movie platform using web crawler technology; S2. Preprocess the obtained movie review data; S3. Randomly divide the preprocessed movie review data into a training set and a test set at a ratio of 1:1; S4. Construct a Mem_GRU model based on the GRU model and use the training set to train the Mem_GRU model; S5. Use the test set to test the trained Mem_GRU model and evaluate the performance of the Mem_GRU model according to the accuracy rate, precision rate, recall rate, and F1 value.
2. The method for intelligent discrimination of movie review sentiment based on improved GRU according to claim 1, characterized in that: In the step S1, the movie review data includes the sentiment label of the movie review, the review time, and the detailed review content.
3. The method for intelligent discrimination of movie review sentiment based on improved GRU according to claim 1, characterized in that: In the step S2, preprocessing the movie review data includes word segmentation, removing stop words, and constructing word vectors.
4. The method for intelligent discrimination of movie review sentiment based on improved GRU according to claim 1, characterized in that: In the step S4, constructing the Mem_GRU model includes adding a memory cell structure similar to LSTM to the GRU model and introducing a new cell state c t , and realizing the functions of the model by designing a forgetting gate f t , an input gate i t , a reset gate r t , an update gate Z t and an output gate o t ; Forget gate f t Controls how much information is forgotten from the memory cell c t The calculation formula is as follows: f t = σ(W f x t + U f h t-1 + b f ); Design an input gate i t to control how much new information can enter the memory cell C t , and the calculation formula is: i t = σ(W i x t + U i h t-1 + b i ); Candidate cell state for new information to enter the memory unit The calculation formula is as follows: Using the reset gate r t to control the candidate state ( whether it depends on the previous state h t-1 , the calculation formula is: Then, update the gate Z t Determine how much historical information to retain, and the calculation formula is: Design an output gate o t , and the calculation formula is: o t = σ(W o x t + U o h t-1 + b o ); The final output calculation formula is: y t = o t ⊙ h t 。 5. The method for intelligent discrimination of movie review sentiment based on improved GRU according to claim 1, characterized in that: In the step S4, a binary cross-entropy loss function is used to optimize the Mem_GRU model, and the binary cross-entropy loss function formula is: Suppose there are N samples, and the true label of the i-th sample is y i , and the predicted probability that the model assigns this sample to the positive class is Then the binary cross-entropy loss of this batch of samples is the average of the losses of individual samples.