Statistical analysis method based on automobile consumption complaints
By using a statistical analysis method based on text mining in automobile consumption complaints, the consumption problems are automatically refined and classified, and the problem of excessive warnings and reliance on manual judgments in the existing technology is solved, and more efficient and accurate warnings are achieved.
Patent Information
- Application Number
- CN202510116527.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems in automobile consumer complaints that warnings are too broad and rely too much on manual judgment, which leads to the warning value of warnings that are not large enough and are prone to cover up when facing large complaints.
A statistical analysis method based on automobile consumer complaints is adopted, including acquisition and pre-processing of automobile consumer complaint data, construction of LDA theme model for automobile consumer complaint texts, text classification and statistical analysis of text classification results. Automatically refine and classify consumption problems through text mining technology, reduce manual dependence and improve the accuracy of consumption prompts and warnings.
It has achieved in-depth exploration and detailed classification of automobile consumer complaints, reduced the dependence on manual judgment, improved the accuracy and efficiency of consumer warnings, and effectively discovered hot areas and problems of complaints.
Smart Images

Figure CN119988612A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of statistical analysis algorithms, and in particular to a statistical analysis method based on automobile consumer complaints. Background Art
[0002] There are many methods for text extraction, recognition and classification, such as statistical learning-based methods, machine learning-based methods, deep learning-based methods, etc. Although there are many methods, to study text extraction, recognition and classification methods based on a specific field, it is necessary to deeply analyze the text data sources and data characteristics in the field and select appropriate text mining methods.
[0003] Timely issuing consumption reminders and warnings to consumers is an important application direction of consumer rights protection data analysis. Judging from the automobile consumption reminders and warnings issued by the Market Supervision Bureau in the past, there are two problems: First, the consumption reminders and warnings are too broad and the value of the reminders and warnings is not great enough. The Market Supervision Bureau publishes the number of hot complaints based on the classification of goods and services and the categories of complaints. Often, due to the coarse classification granularity and the fact that the registration personnel classify too many "other" categories in pursuit of speed when registering, the various types of consumer problems distributed according to the standards and specifications are too general; second, although the consumer problems warned to consumers are detailed, they rely too much on manual judgment. In response to the first problem, in order to improve the value of the reminders and warnings, the market supervision bureaus in various places use manual judgment and add detailed consumer problem analysis content. However, the situation of using manual judgment can still cope with a small number of complaints, but it is very easy to generalize in the face of a large number of complaints. Summary of the invention
[0004] In view of the above problems, the present invention provides a statistical analysis method based on automobile consumption complaints.
[0005] The technical solution adopted is a statistical analysis method based on automobile consumer complaints, which includes the following steps:
[0006] S1. Acquisition and preprocessing of automobile consumer complaint data;
[0007] S2. Construct LDA topic model for automobile consumption complaint text;
[0008] S3. Classification of automobile consumer complaint texts;
[0009] S4. Statistical analysis of automobile consumer complaint text classification results.
[0010] Optionally, S1 includes the following steps:
[0011] S11. Collection of automobile consumption complaint data;
[0012] S12. Cleaning of automobile consumption complaint data;
[0013] S13. Noise filtering of automobile consumption complaint data;
[0014] S14. Perform Chinese word segmentation on the complaint text information in the automobile consumption complaint data.
[0015] Optionally, in S11, automobile-related consumer complaints in the target province are obtained from the national 12315 complaint reporting platform through a database table;
[0016] In S12, a regular expression is used to remove irrelevant characters, punctuation marks, numbers, and special symbols in the complaint question obtained in text S11, and the main content of the text is retained;
[0017] In S13, the repeated, meaningless or irrelevant words and sentences in the main body of the text in S12 were removed;
[0018] In S14, the complaint text in S13 is preprocessed to establish a user automobile consumption complaint key segmentation set and a corresponding automobile failure key segmentation set.
[0019] Optionally, S2 includes the following steps:
[0020] S21. Text data import;
[0021] S22. Construct corpus dictionary;
[0022] S23. Form a corpus bag-of-words model;
[0023] S24. Setting topic modeling parameters;
[0024] S25. Output topic modeling results.
[0025] Optionally, in S25, the topic modeling result includes a number of topics that match the parameter settings, and each of the topics includes multiple topic words and topic word weights.
[0026] Optionally, S3 includes the following steps:
[0027] S31. Preprocess the text, remove punctuation marks and spaces in the text, automatically segment the text, and remove stop words;
[0028] S32. Extracting feature values of text by frequency analysis and part-of-speech analysis of relevant text;
[0029] S33. Use Bunch type to describe the complaint text;
[0030] S34. Assigning weights to keywords using TF-IDF values to construct a keyword weight matrix;
[0031] S35. Use the Naive Bayes algorithm to classify text.
[0032] Optionally, in S34, a training set and a test set are generated;
[0033] In S35, the constructed naive Bayes classifier is first trained with the data in the training set, and then filtered according to the characteristics of the training process in the test set data processing.
[0034] The benefits of the present invention include:
[0035] 1. Use text mining technology to conduct in-depth research on complaint issues, automatically refine and classify consumer issues, minimize manual reliance, and improve the accuracy of consumer warnings.
[0036] 2. It can compare the number of complaints and year-on-year growth rates of automobile consumer complaints in various regions, identify hot complaint areas, and based on text mining technology, automatically and regularly sort out hot issues of consumer complaints, display the proportion of hot complaint issues, and support drilling down by issue to view the specific content of consumer complaints. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Flowchart for LDA topic modeling;
[0038] Figure 2 This is a technical framework diagram for automatic text classification;
[0039] Figure 3 This is a distribution map of automobile consumption complaints. DETAILED DESCRIPTION
[0040] The following describes the implementation of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific implementations, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0041] It should be noted that the illustrations provided in the following embodiments are only used to schematically illustrate the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0042] A statistical analysis method based on automobile consumption complaints includes the following steps:
[0043] S1. Acquisition and preprocessing of automobile consumer complaint data;
[0044] S2. Construct LDA topic model for automobile consumption complaint text;
[0045] S3. Classification of automobile consumer complaint texts;
[0046] S4. Statistical analysis of automobile consumer complaint text classification results.
[0047] The purpose of this design is to use text mining technology to conduct in-depth research on complaint issues, automatically refine and classify consumer issues, minimize manual reliance, and improve the accuracy of consumer warnings.
[0048] Meanwhile, in the specific implementation, S1 includes the following steps:
[0049] In S11, automobile-related consumer complaints in the target province are obtained from the national 12315 complaint reporting platform through database tables;
[0050] In S12, a regular expression is used to remove irrelevant characters, punctuation marks, numbers, and special symbols in the complaint question obtained in text S11, and the main content of the text is retained;
[0051] In S13, the repeated, meaningless or irrelevant words and sentences in the main body of the text in S12 were removed;
[0052] In S14, the complaint text in S13 is preprocessed to establish a user automobile consumption complaint key segmentation set and a corresponding automobile failure key segmentation set.
[0053] like Figure 1 As shown, in S2, the following steps are included:
[0054] S21. Text data import;
[0055] S22. Construct corpus dictionary;
[0056] S23. Form a corpus bag-of-words model;
[0057] S24. Setting topic modeling parameters;
[0058] S25. Output topic modeling results.
[0059] In a specific embodiment, taking part of the complaint text data obtained in step S1 as an example, the details are shown in Table 1:
[0060] Serial number Complaint text 1 The complainant stated that the car he purchased here in September last year frequently had problems starting and he suspected that the vehicle had quality problems. Please investigate and verify the complaint in a timely manner. 2 The consumer complained that the merchant had broken the glass of the consumer's car and demanded compensation for the loss. 3 The complainant claimed that he bought car headlights from the store, but the merchant refused to handle the issue during the warranty period. 4 The car purchased in December 2019 had quality problems and the engine had to be repaired twice, so the consumer requested a refund. 5 The complainant had paid a deposit for the car but the merchant refused to refund the deposit saying that the price of the car had increased and he demanded that the merchant return the deposit.
[0061] Based on the contents of Table 1, stop words are removed from these texts and segmented to obtain segmented data, see Table 2:
[0062] Serial number Corpus 1 'complainant','purchase','car','frequent','occur','misfire','problem','vehicle','existence','quality','problem','complaint','timely','investigation','verification' 2 'consumer','complaint','merchant','consumer','car','glass','broken','consumer','demand','compensation','loss' 3 'complainant','purchase','car','headlight','warranty period','merchant','reject','handle' 4 'purchase','car','occurrence','quality','problem','engine','failure','repair','consumer','request','return' 5 'complainant','car','deposit','merchant','car','price increase','merchant','not','refund','deposit','request','merchant','refund','deposit'
[0063] Import the corpus in Table 2 and generate a corpus dictionary. The corpus dictionary gives each word a unique ID. The dictionary is as follows:
[0064] 'complainant':0,'purchase':1,'car':2,'frequent':3,'appear':4,'no fire':5,'vehicle':6,'existence':7,'quality':8,'complaint':9,'timely':10,'investigation':11,'verification':12,'consumer':13,'merchant':14,'glass':15,'rotten':16,'request':17,'compensation':18,'loss':19,'headlight':20,'warranty period':21,'rejected':22,'handled':23,'problem':24,'engine':25,'fault':26,'repair':27,'return':28,'deposit':29,'price increase':30,'refund':31
[0065] The corpus data is converted into a bag-of-words type according to the word ID and word frequency. In the brackets, the first one is the word ID, and the second one is the word frequency in this sentence, as shown in Table 3 below:
[0066] Serial number Corpus 1 [(0, 1), (1, 1), (2, 1), (3, 1), (4, 1), (5, 1), (24,2)(6, 1), (7, 1), (8, 1), (9, 1)(10,1),(11,1),(12,1)] 2 [(13, 3), (9, 1), (14, 1), (2, 1), (15, 1), (16, 1), (17,1)(18, 1), (19, 1)] 3 [(0, 1), (1, 1), (2, 1), (20, 1), (21, 1), (14, 1), (22,1)(23, 1)] 4 [(1, 1), (2, 1), (4, 1), (8, 1), (24, 1), (25, 1), (26,1)(13, 1), (17, 1),(28,1)] 5 [(0, 1), (2, 2), (29, 3), (14, 3), (30, 1), (22, 1), (31,2)(17, 1)]
[0067] Finally, set the topic modeling parameters, including the number of topic types, the number of topic keywords, and the number of training times.
[0068] At the same time, the topic modeling results include multiple topics, and the number of topics is determined by the parameter settings. Each topic contains multiple subject words and the weight of the subject words. The weight represents the probability that the subject word belongs to the topic. According to the setting of the keyword parameters, the results are sorted by the size of the weight to filter out the possible subject words under the topic.
[0069] In this embodiment, the model parameters are set to have 5 topics and 10 keywords, and the partial results are as follows:
[0070] (1) '0.037*"Maintenance" + 0.030*"Consumption card" + 0.022*"After-sales service" + 0.019*"Service attitude" + 0.019*"Fee dispute" + 0.018*"Repair refused" + 0.018*"License plate" + 0.017*"Repair damage" + 0.017*"Accessories" + 0.086*"Insurance claim"');
[0071] (2) '0.109*"Deposit is not refundable" + 0.097*"Mandatory insurance" + 0.062*"Used car" + 0.040*"Expenses" + 0.036*"Certificate of conformity" + 0.034*"Mandatory consumption" + 0.090*"Contract" + 0.037*"Overpricing" + 0.017*"Failure to fulfill commitment" + 0.015*"Invoice problem"';
[0072] (3) '0.146*"car mass" + 0.129*"engine" + 0.030*"car glass" + 0.030*"car paint" + 0.028*"key" + 0.112*"tire" + 0.107*"gearbox" + 0.098*"car door" + 0.086*"steering wheel" + 0.115*"battery"'.
[0073] Combining the process and service characteristics of automobile consumer services with the results of topic modeling, the extracted keywords are summarized and analyzed, and it can be concluded that the themes of automobile complaints mainly include the following aspects:
[0074] Topic 1, the extracted keywords include maintenance, service attitude, repair and other words. According to the keyword analysis, it can be concluded that the main complaints of consumers include related problems arising from the process of car maintenance and repair. Therefore, car repair services have become a major content of consumer complaints.
[0075] Topic 2, the extracted keywords include deposit, insurance, contract, invoice and other keywords. Through the keywords, it can be determined that the main contents of consumer complaints include non-refund of deposits, compulsory insurance, failure to fulfill promises, etc. Therefore, automobile consumption transactions have become a major content of consumer complaints.
[0076] Topic 3, the extracted keywords include engine, car quality, tire, door, etc. These words mainly describe car quality issues and are also the topics with more complaint data.
[0077] like Figure 2 As shown, S3 includes the following steps:
[0078] S31. Preprocess the text, remove punctuation marks and spaces in the text, automatically segment the text, and remove stop words;
[0079] S32. Extracting feature values of text by frequency analysis and part-of-speech analysis of relevant text;
[0080] S33. Use Bunch type to describe the complaint text;
[0081] S34. Assigning weights to keywords using TF-IDF values to construct a keyword weight matrix;
[0082] S35. Use the Naive Bayes algorithm to classify text.
[0083] In the specific implementation, this step mainly includes the extraction of text features and the design of classifiers. In terms of text feature extraction, vector space models are often used to represent and describe text. The features of each dimension of the vector space model can be directly obtained by word segmentation or word frequency statistics algorithm to obtain vocabulary, or the features of identification ability can be further extracted based on the vocabulary, such as information gain, word frequency, mutual information, TFIDF matrix, etc.
[0084] Since the vector space model obtained by directly using vocabulary as features has too large a dimension and is not conducive to solving, features are usually further extracted based on vocabulary to construct a vector space model. Among them, TFIDF features are the most widely used.
[0085] In terms of classifier design, with the rapid development of machine learning technology, there are many classifiers available. However, based on the characteristics of text classification, the commonly used classifiers in the field of text classification include K-nearest neighbor, support vector machine and naive Bayes. K-nearest neighbor and support vector machine classification methods work better when processing small samples, while the naive Bayes method is more adaptable to text classification and is suitable for text classification of sample sets of different sizes. Therefore, the naive Bayes classification method is widely used in the field of text classification.
[0086] Text classification is mainly divided into two processes: training process and testing process. The main goal of the training process is to obtain a text classifier based on the existing training data. The main steps include:
[0087] (1) Preprocess the text to remove punctuation marks and spaces, automatically segment the text, and remove stop words;
[0088] (2) Extracting feature values of texts through word frequency analysis and part-of-speech analysis of relevant texts;
[0089] (3) Use Bunch type to describe the complaint text;
[0090] (4) Assign weights to keywords using TF-IDF values and construct a keyword weight matrix;
[0091] (5) The Naive Bayes algorithm is used to classify the text. The data from the test process is processed in the same way and filtered according to the characteristics of the training process. Finally, the trained classifier is used to classify the text.
[0092] When performing feature extraction, TF-IDF (Term Frequency-InversDocument Frequency) is a weighting technology commonly used in information processing and data mining. This technology uses a statistical method to calculate the importance of a word in the entire corpus based on the number of times the word appears in the text and the document frequency in the entire corpus. Its advantage is that it can filter out some common but insignificant words while retaining important words that affect the entire text. The calculation method is shown in the following formula.
[0093]
[0094] Among them, tfidf represents the product of word frequency tfi,j and inverted text word frequency idfi. The larger the TF-IDF value, the more important the feature word is to the text.
[0095] TF (Term Frequency) indicates the frequency of a keyword appearing in the entire article.
[0096] IDF (InversDocument Frequency) means calculating the inverse document frequency. The document frequency refers to the number of times a keyword appears in all articles in the entire corpus. The inverse document frequency is also called the inverse document frequency. It is the inverse of the document frequency and is mainly used to reduce the effect of some common words in all documents that have little impact on the document. The following formula is the calculation formula for TF word frequency.
[0097]
[0098] Among them, ni,j is the number of times the feature word appears in the text, which is the number of all feature words in the text dj. The result of the calculation is the word frequency of a feature word. The following formula is the calculation formula of IDF.
[0099]
[0100] Among them, |D| represents the total number of texts in the corpus, which means the number of texts containing the feature word ti. In order to prevent the word from not existing in the corpus, that is, the denominator is 0, use as the denominator.
[0101] category Keywords Car breakdown 'car','quality','problem','engine','failure','repair','consumer','request','return'
[0102] For the Bayesian classifier part, the Bayesian network is a classifier learned from a series of sample instances with category labels. For a given instance X, represented by a feature vector (a1, a2, ..., an), the Bayesian network uses the following equation for classification:
[0103]
[0104] Where: C is the set of possible categories c, c(x) is the category of x predicted by the Bayesian network classifier. For known categories, it is assumed that all attribute conditions are independent of each other. The Bayesian network classifier based on the attribute condition independence assumption is called the naive Bayesian classifier. The naive Bayesian algorithm is the simplest form of the Bayesian network classifier and is one of the most widely used algorithms in the field of classifiers. The expression of the naive Bayesian classifier is:
[0105]
[0106] The probability p(c) and conditional probability p(ai|c) mentioned above can be calculated by the following equations:
[0107]
[0108] Where: n is the total number of training instances; l is the total number of categories; ni is the possible value of the ith attribute; cj is the category of the jth instance; aji is the ith attribute of the jth instance; δ(●) is a binary function, if the two parameter values are the same, the function value is 1, otherwise it is 0.
[0109] like Figure 3 As shown in the data, 5,505 automobile complaint information in a certain province was sorted out and integrated from the five aspects of automobile failure, sales transaction, maintenance service, violation of laws and regulations, and car ordering and car delivery. It was found that the number of passenger complaints in these three aspects was the largest. Among them, automobile failure accounted for 2,541, accounting for 46.2%; followed by sales transaction with 2,040, accounting for 37.1%; and then maintenance service with 470, accounting for 8.5%.
[0110] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A statistical analysis method based on automobile consumption complaints, characterized in that: The following steps are involved: S1. Acquisition and preprocessing of automobile consumer complaint data; S2. Construct LDA topic model for automobile consumption complaint text; S3. Classification of automobile consumer complaint texts; S4. Statistical analysis of automobile consumer complaint text classification results.
2. The statistical analysis method based on automobile consumption complaints according to claim 1 is characterized in that: S1 includes the following steps: S11. Collection of automobile consumption complaint data; S12. Cleaning of automobile consumption complaint data; S13. Noise filtering of automobile consumption complaint data; S14. Perform Chinese word segmentation on the complaint text information in the automobile consumption complaint data.
3. The statistical analysis method based on automobile consumption complaints according to claim 2 is characterized in that: In S11, automobile-related consumer complaints in the target province are obtained from the national 12315 complaint reporting platform through database tables; In S12, a regular expression is used to remove irrelevant characters, punctuation marks, numbers, and special symbols in the complaint question obtained in text S11, and the main content of the text is retained; In S13, the repeated, meaningless or irrelevant words and sentences in the main body of the text in S12 were removed; In S14, the complaint text in S13 is preprocessed to establish a user automobile consumption complaint key segmentation set and a corresponding automobile failure key segmentation set.
4. The statistical analysis method based on automobile consumption complaints according to claim 1 is characterized in that: S2 includes the following steps: S21. Text data import; S22. Construct corpus dictionary; S23. Form a corpus bag-of-words model; S24. Setting topic modeling parameters; S25. Output topic modeling results.
5. The statistical analysis method based on automobile consumption complaints according to claim 4 is characterized in that: In S25, the topic modeling result includes the number of topics that match the parameter settings, and each of the topics includes multiple topic words and topic word weights.
6. The statistical analysis method based on automobile consumption complaints according to claim 1 is characterized in that: S3 includes the following steps: S31. Preprocess the text, remove punctuation marks and spaces in the text, automatically segment the text, and remove stop words; S32. Extracting feature values of text by frequency analysis and part-of-speech analysis of relevant text; S33. Use Bunch type to describe the complaint text; S34. Assigning weights to keywords using TF-IDF values to construct a keyword weight matrix; S35. Use the Naive Bayes algorithm to classify text.
7. The statistical analysis method based on automobile consumption complaints according to claim 6 is characterized in that: In S34, a training set and a test set are generated; In S35, the constructed naive Bayes classifier is first trained with the data in the training set, and then filtered according to the characteristics of the training process in the test set data processing.