Fine granularity sentiment analysis method for automobile comments
By segmenting, word segmentation, word vectorization, similar word clustering and high-frequency phrase annotation, the problems of fine-grained feature extraction and sentiment analysis of automobile reviews are solved, and precise feature extraction and sentiment analysis of complex user reviews are achieved.
Patent Information
- Application Number
- CN202311773606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
AI Technical Summary
It is difficult for prior art to perform fine-grained feature extraction and sentiment analysis of automotive reviews, especially when dealing with the colloquial and diverse features of user language descriptions and the wide variety of automotive features.
By segmenting short sentences and word segmentation of the training corpus in the training corpus, word vectorization is performed, similar words are clustered, high-frequency phrases are extracted and classified and annotated, short sentences are classified and sentiment analysis is performed. The specific steps include dividing short sentences based on punctuation marks, using word2vec technology to perform word vectorization, calculating the vector cosine value of the subject word for clustering similar words, extracting high-frequency phrases and annotating classification, classifying short sentences and performing sentiment analysis through the sentiment analysis model.
It realizes accurate feature extraction and fine-grained sentiment analysis of automobile reviews, which can effectively process complex user review data and provide detailed user feedback and sentiment analysis results.
Smart Images

Figure CN120218073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sentiment analysis of automotive reviews, and more specifically, to a method for fine-grained sentiment analysis of automotive reviews. Background Art
[0002] For vehicle manufacturers, obtaining complete user feedback to extract the advantages and disadvantages of their own products and competitor models from the user's perspective is a very important analysis task. Traditional methods such as after-sales problem collection and questionnaire surveys are limited by the amount of information and efficiency, and it is difficult to obtain very complete user information. Online car purchase reviews containing rich information provide possibilities for this analysis. The reviews not only include the user's preference levels for various automotive modules, such as engines, transmissions, interior and exterior decorations, etc., but also include descriptions of various features in these modules, such as the comfort, aesthetics, safety, etc. of the seats. How to perform fine-grained feature extraction and sentiment analysis on user reviews has great research significance. However, the user's language descriptions are colloquial and diverse, and at the same time, there are a large number of automotive features, which adds great difficulties to this task. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for accurately extracting features from automotive reviews and then performing fine-grained sentiment analysis.
[0004] To achieve the above purpose, the fine-grained sentiment analysis method for automotive reviews of the present invention includes the following steps:
[0005] S100, segment the training corpus in the training corpus into short sentences and perform word segmentation;
[0006] S200, perform word vectorization on the segmented words;
[0007] S300, perform clustering of similar words for the subject words;
[0008] S400, extract high-frequency word groups from the training corpus in the training corpus and classify and label them;
[0009] S500, classify the short sentences; and
[0010] S600, perform sentiment analysis on the classified short sentences.
[0011] In an embodiment of the above fine-grained sentiment analysis method for automotive reviews, the step S100 includes: segmenting the training corpus into short sentences by punctuation marks.
[0012] In an embodiment of the above fine-grained sentiment analysis method for automotive reviews, the step S100 includes: establishing a dictionary of indivisible words, and performing word segmentation on the segmented short sentences based on the dictionary of indivisible words.
[0013] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, the step S100 includes: establishing a stop word dictionary, and performing word segmentation on the segmented short sentences based on the stop word dictionary.
[0014] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, the step S200 includes: using the word2vec technology to vectorize all the segmented words in the training corpus.
[0015] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, the step S300 includes: calculating the cosine values of the vectors of the topic words and the vectors of other segmented words in the training corpus, and statistically finding the larger values among them to perform similar word clustering.
[0016] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, the step S400 includes:
[0017] S410, extracting high-frequency words from the short sentence to form word groups;
[0018] S420, for a short sentence containing multiple high-frequency words, reducing the number of features in the word group by calculating the co-occurrence probability of a single segmented word and other segmented words;
[0019] S430, labeling the determined high-frequency word groups with their classifications.
[0020] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, the step S500 includes: for the short sentence to be analyzed, if it contains a high-frequency word group that has been labeled, directly classify it according to its label; if there is no corresponding high-frequency word group, search for the labeled word group that is most similar to its vector and classify it according to its label.
[0021] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, for all the high-frequency words in the word group, calculate the maximum cosine distance between them and all the segmented words in the short sentence to be classified, sum up this maximum cosine distance and divide it by the number of high-frequency words in the word group to finally obtain the average maximum cosine distance to obtain the similarity between the word group and the short sentence.
[0022] In one embodiment of the above-mentioned fine-grained sentiment analysis method for automotive reviews, set the difference between the number of high-frequency words in the word group and the number of segmented words in the short sentence as the similarity correction coefficient. The larger this difference is, the smaller the similarity correction coefficient between the two is.
[0023] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but it is not intended to limit the present invention. Description of the Drawings
[0024] Figure 1Step diagram of the fine-grained sentiment analysis method for automotive reviews of the present invention;
[0025] Figure 2 Classification flowchart of the fine-grained sentiment analysis method for automotive reviews of the present invention. Detailed implementation manners
[0026] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments to further understand the purpose, solution and efficacy of the present invention, but it is not intended to limit the protection scope of the appended claims of the present invention.
[0027] References to "one embodiment", "embodiment", "example embodiment", etc. in the specification mean that the described embodiment may include specific features, structures or characteristics, but not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures or characteristics in an embodiment, it has been shown that combining such features, structures or characteristics into other embodiments is within the knowledge of those skilled in the art, whether or not explicitly described.
[0028] The purpose of the present invention is to provide a method for accurately extracting features from automotive reviews and then performing fine-grained sentiment analysis, as Figure 1 shown, including the following steps:
[0029] S100, segment short sentences and perform word segmentation on the training corpus in the training corpus. This step cleans and segments the training corpus, including splitting the training corpus into short sentences based on punctuation marks, and performing word segmentation on the short sentences based on the indivisible word dictionary and the stop word dictionary.
[0030] S200, perform word vectorization on the segmented words. The present invention uses the existing word embedding method Word2vec for word vectorization.
[0031] S300, perform similar word clustering on the topic words. Based on word vectorization, determine the similar words of the topic words. For example, taking'seat' as the topic word, that is, analyzing user reviews related to seats, it is necessary to determine the similar words of'seat', such as 'front row', 'cushion', 'backrest', etc.
[0032] S400, extract high-frequency word groups from the training corpus in the training corpus and classify and label them. For the sentences in the training corpus containing the topic word and its similar words, extract their high-frequency words and form high-frequency word groups. For example: 'back row' + 'crossed legs', and label its classification: 'back row space'.
[0033] S500 classifies short sentences. For the corpus to be analyzed, if it contains labeled phrases, it is directly classified; if there are no corresponding phrases, it searches for the existing labeled phrases with the closest vectors and classifies them according to their labels.
[0034] S600 performs sentiment analysis on the classified short sentences. For each phrase, a certain amount of sentences are searched in the corpus, and their sentiment (positive or negative) is labeled. The existing sentiment analysis model snownlp is trained. Then the corpus to be analyzed is classified by attribute, and sentiment analysis is performed through the trained snownlp model to obtain its sentiment score. For each classification, the relevant number of comments and sentiment scores are counted.
[0035] Specifically, combined with Figure 1 and Figure 2 In an embodiment of the fine-grained sentiment analysis method for automotive reviews of the present invention, the step S100 includes: splitting the training corpus into short sentences by punctuation marks. First, the training corpus is split into short sentences by punctuation marks such as ',', '.', '!'. For example, 'I really like the appearance of this car, the interior is also very luxurious, and the driving experience is also very good' will be split into three short sentences: 'I really like the appearance of this car', 'The interior is also very luxurious', and 'The driving experience is also very good'.
[0036] The step S100 includes: establishing an indivisible word dictionary, and the split short sentences are segmented based on the indivisible word dictionary; and establishing a stop word dictionary, and the split short sentences are segmented based on the stop word dictionary.
[0037] Establishing an indivisible word dictionary includes setting specific words in the automotive field to avoid being split and losing their original meanings. Such as 'lane keeping', 'four-way camera', etc., and forming a file. Based on the indivisible word dictionary, a word segmentation tool is used for preliminary word segmentation. The word frequency of the word segmentation result is counted, and the result is output to a file. Words that do not affect semantic recognition such as 'ah', 'ne', 'actually' are selected as stop words in the file and a file is formed.
[0038] Read the indivisible word and stop word dictionaries, and perform word segmentation through a word segmentation tool. For example, 'I really like the vehicle appearance' will be segmented into 'I / very much / like / vehicle / appearance', and the part-of-speech of each segmented word is saved.
[0039] According to the word segmentation result, the indivisible word dictionary and the stop word dictionary are improved in multiple rounds until the word segmentation result is concise and the semantics are clear.
[0040] In an embodiment of the fine-grained sentiment analysis method for automotive reviews of the present invention, step S200 includes performing word vectorization on all word segments in the training corpus using the word2vec technique. It is difficult for a computer to directly calculate and analyze text, and it needs to be first converted into a digital form. The present invention uses the existing word2vec technique to perform word vectorization on all word segments in the training corpus. At the same time, word2vec can also ensure that the word vectorization results of words with similar semantics are similar.
[0041] Step S300 includes calculating the cosine values of the vectors of the topic words and the vectors of other word segments in the training corpus, and counting the larger values among them for similar word clustering. As an integrator of civilian industries, an automobile has a large number of internal systems, resulting in a huge number of user comment dimensions, and it is difficult to screen out relevant comments through a single topic word. Taking the topic word'seat' as an example, the actual user comments on the'seat' are likely not to contain the word segment'seat'. For example: 'There is no heating function in the front row', 'The seat cushion is too short', 'The backrest angle is very comfortable'. In order to ensure that all relevant comments of the topic word can be recognized, it is necessary to establish a similar word clustering of the topic word. According to the characteristic of word2vec that the word vectorization results of words with similar semantics are similar, calculate the cosine values of the vectors of the topic word and the vectors of other word segments in the training corpus, and count the larger values among them, that is, the word segments with semantics similar to the seat, such as 'front row', 'cushion', 'backrest', 'cover', etc. are the similar words of the topic word'seat'.
[0042] Step S400 includes:
[0043] S410, for the short sentences in the training corpus, extract the high-frequency words to form word groups, such as 'front row' + 'comfortable'. Take the word groups formed by the extracted high-frequency words as a whole, perform word frequency statistics, and screen out those with high word frequencies.
[0044] S420, for the case where a short sentence contains multiple high-frequency words, appropriately reduce the number of features in the word group by calculating the co-occurrence probability of a single word segment and other word segments. That is, delete the word segments with a low co-occurrence probability with other word segments to avoid too many word segments and too large dimensions in the word group, resulting in too many word groups extracted later and difficult manual annotation.
[0045] S430, label the determined high-frequency word groups with their classifications. For example, 'front row' + 'comfortable', its classification is 'front row comfort'.
[0046] The step S500 includes: for the short sentence to be analyzed, if it contains a word group that has been labeled, directly classify it according to its label; if there is no corresponding word group, search for the labeled word group that is most similar to its vector, and classify it according to its label.
[0047] Among them, the method for searching for the most similar labeled phrase to its vector is as follows: for all high-frequency words in the phrase, calculate the maximum cosine distance between them and all the segmented words in the short sentence to be classified, sum up this maximum cosine distance and divide it by the number of high-frequency words in the phrase, and finally obtain the average maximum cosine distance to obtain the similarity between the phrase and the short sentence.
[0048] To avoid misclassification, a threshold is set for the similarity, and short sentences with a similarity less than this threshold are not classified.
[0049] Furthermore, set the difference between the number of high-frequency words in the phrase and the number of segmented words in the short sentence as the similarity correction coefficient. The greater the difference between the number of high-frequency words in the phrase and the number of segmented words in the short sentence, the smaller the similarity correction coefficient of the two. Specifically, map this difference to a certain interval, such as 0.9 - 1.0, as this coefficient. Finally, use the product of the average maximum cosine distance and the similarity correction coefficient between the phrase and the short sentence as the final similarity result.
[0050] Such as Figure 2 As shown, the classification process of the fine-grained sentiment analysis method for automotive reviews of the present invention is: screen the short sentences in the training corpus with topic words and similar words, perform word frequency statistics on the screened short sentences to determine high-frequency words, form phrases with the high-frequency words and label the classification, and classify the short sentences to be analyzed. Finally, perform sentiment analysis based on features and summarize the results.
[0051] Of course, the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A fine-grained sentiment analysis method for automotive reviews, characterized in that, It includes the following steps: S100, segment the training corpus in the training corpus into short sentences and perform word segmentation; S200, vectorize the words after word segmentation; S300, cluster similar words for the subject words; S400, extract high-frequency word groups from the training corpus in the training corpus and classify and label them; S500, classify the short sentences; and S600, perform sentiment analysis on the classified short sentences.
2. The fine-grained sentiment analysis method for automotive reviews according to claim 1, wherein The step S100 includes: segmenting the training corpus into short sentences by punctuation marks.
3. The fine-grained sentiment analysis method for automotive reviews according to claim 2, wherein The step S100 includes: establishing an indivisible word dictionary, and segmenting the short sentences after segmentation based on the indivisible word dictionary.
4. The fine-grained sentiment analysis method for automotive reviews according to claim 2, wherein The step S100 includes: establishing a stop word dictionary, and segmenting the short sentences after segmentation based on the stop word dictionary.
5. The fine-grained sentiment analysis method for automotive reviews according to claim 1, wherein The step S200 includes: using the word2vec technology to vectorize all the words after segmentation in the training corpus.
6. The fine-grained sentiment analysis method for automotive reviews according to claim 5, wherein The step S300 includes: calculating the cosine value of the vector of the subject word and the vectors of other words after segmentation in the training corpus, and counting the larger values among them to perform similar word clustering.
7. The fine-grained sentiment analysis method for automotive reviews according to claim 2, wherein The step S400 includes: S410, extracting high-frequency words from the short sentences to form word groups; S420, for short sentences containing multiple high-frequency words, reducing the number of features in the word groups by calculating the joint probability of the co-occurrence of a single word after segmentation and other words after segmentation; S430, labeling the classified high-frequency word groups with their classifications.
8. The fine-grained sentiment analysis method for automotive reviews according to claim 7, wherein The step S500 includes: for the short sentences to be analyzed, if they contain the high-frequency word groups that have been labeled, directly classify them according to their labels; if there are no corresponding high-frequency word groups, search for the labeled word groups that are most similar to their vectors and classify them according to their labels.
9. The fine-grained sentiment analysis method for automotive reviews according to claim 8, characterized in that For all the high-frequency words in the word groups, calculate the maximum cosine distance between them and all the words after segmentation in the short sentences to be classified, sum up this maximum cosine distance and divide it by the number of high-frequency words in the word groups, and finally obtain the average maximum cosine distance to obtain the similarity between the word groups and the short sentences.
10. The fine-grained sentiment analysis method for automotive reviews according to claim 9, characterized in that Set the difference between the number of high-frequency words in the word groups and the number of words after segmentation in the short sentences as the similarity correction coefficient. The larger this difference is, the smaller the similarity correction coefficient between the two is.