A Fine-Grained Automotive Review Quality Assessment Method Based on Size Model Collaboration
By combining adaptive K-means clustering and multi-task learning neural networks with preprocessing and a multidimensional quality assessment model, the problem of fine-grained evaluation in automotive review quality assessment is solved, achieving efficient and accurate multidimensional quality assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 湖南工商大学
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-17
AI Technical Summary
Existing automotive review quality assessment technologies struggle to achieve fine-grained evaluations, the evaluation granularity doesn't match user concerns, the data scale and computational costs are high, the quality of review content is inconsistent, and existing screening and evaluation systems are susceptible to noise interference.
An adaptive K-means clustering algorithm and a pre-trained multi-task learning neural network are used, combined with preprocessing and a multi-dimensional quality assessment model. Through clustering, length level division and information density prediction, readability and sentiment scoring are integrated to improve assessment accuracy and interpretability.
It significantly improves the accuracy and interpretability of car review quality assessment, can more accurately reflect the multi-dimensional quality of reviews, reduces noise interference, and lowers computational costs.
Smart Images

Figure CN122196188B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a method for evaluating the quality of fine-grained car reviews based on size-model collaboration. Background Technology
[0002] Driven by the digital wave, the automotive consumer market has fully entered the "word-of-mouth era." Various vertical media platforms have accumulated hundreds of millions of genuine car owner reviews, covering the entire lifecycle of users from car selection and purchase to long-term ownership, containing key information on product design, performance, user experience, and service quality. For potential consumers, these reviews have become an important reference for purchasing decisions; for OEMs and dealers, they are core data assets for understanding user needs, identifying product defects, and optimizing operational strategies. However, faced with the massive, continuously growing, and highly unstructured review data, existing word-of-mouth assessment technologies still face three prominent problems in practical applications: First, the evaluation granularity is seriously mismatched with the users' real concerns: a large amount of detailed information with clear direction is buried in the average value of macro indicators and the statistics of emotional polarity, which makes it difficult to support the OEM's targeted engineering improvements and also makes it difficult to build an interpretable and comparable multi-dimensional reputation profile for users.
[0003] Secondly, the contradiction between the explosion of data scale and the high cost of computing is prominent: large model inference costs are high and response latency is large. If high-computing large models are used to analyze millions of existing comments and massive daily incremental data one by one, it will bring unbearable computing power and time costs to enterprises. On the other hand, relying solely on traditional lightweight models makes it difficult to guarantee the accuracy of analysis in complex scenarios such as long texts, implicit emotions, and complex semantics.
[0004] Third, the quality of comments varies greatly, and the existing screening and evaluation system is easily affected by noise: the existing system usually relies on crude signals such as star rating, number of likes, and number of replies for sorting and screening, which is easy to manipulate and difficult to reflect the true quality of the comments in terms of information richness, logical consistency and readability. Summary of the Invention
[0005] Therefore, it is necessary to provide a fine-grained automotive review quality assessment method based on size model collaboration, including: S1: Acquire and preprocess car review text data to construct a raw car review dataset consisting of preprocessed review text; S2: Based on the A-TFIDF values of each word in the preprocessed comment text, the adaptive K-means clustering algorithm is used to cluster the preprocessed comment text belonging to any first-level indicator to obtain the second-level indicators under the corresponding first-level indicator; S3: Divide each preprocessed comment text belonging to any first-level indicator into length levels to obtain the length level of the corresponding preprocessed comment text; S4: Based on the pre-trained multi-task learning neural network information density prediction model, the preliminary information richness index of any pre-processed comment text is predicted. The preliminary information richness index is the number of secondary indicators covered in the corresponding pre-processed comment text. S5: Based on the length level and the preliminary information richness index, determine the information density level of the corresponding preprocessed comment text, and distribute any preprocessed comment text to the set of pre-trained language models applied to predict readability scores and sentiment scores based on the information density level, so as to obtain the readability score and sentiment score of the corresponding preprocessed comment text respectively. S6: Based on the information density level, the readability score, and the sentiment score, integrate them into a primary indicator and / or a multi-dimensional quality assessment result corresponding to the preprocessed comment text.
[0006] Beneficial Effects: This method preprocesses car review text data to obtain preprocessed review texts; it uses an adaptive K-means clustering algorithm to cluster preprocessed review texts belonging to any primary indicator, obtaining secondary indicators under the primary indicator; it divides the length levels of each preprocessed review text and predicts the preliminary information richness index of the preprocessed review text based on a pre-trained multi-task learning neural network information density prediction model; based on the length level and the preliminary information richness index, it determines the information density level of the corresponding preprocessed review text, and distributes any preprocessed review text to a set of pre-trained language models applied to predict readability scores and sentiment scores, respectively, to obtain the readability score and sentiment score of the corresponding preprocessed review text, and then integrates the primary indicator and / or the multi-dimensional quality assessment results corresponding to the preprocessed review text. This method significantly improves the accuracy and interpretability of car reputation quality assessment. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a flowchart of the fine-grained automotive review quality assessment method based on size model collaboration in this application. Detailed Implementation
[0009] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0010] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0011] like Figure 1 As shown, this embodiment provides a method for evaluating the quality of fine-grained car reviews based on size-model collaboration, including: S1: Acquire and preprocess car review text data to construct the original car review dataset consisting of preprocessed review text.
[0012] Specifically, the steps include: This embodiment retrieves text data of all car reviews for several target electric vehicle brands from the review details pages of automotive vertical websites. Specifically, targeting the top ten best-selling electric vehicle brands, a web crawler is built to collect data from the review details pages of the automotive vertical website "Autohome," capturing the unique resource locator (review_url) for each review and parsing it to obtain the review's metadata and content data.
[0013] Parsing data from any given car review text dataset yields the following: review ID, review post tag icon path, user rating for primary metrics, text review content containing primary metrics, vehicle model name, purchase price information, purchase time, purchase location, and vehicle usage information. The parsed data fields are shown in Table 1. Table 1 Data Field Mapping Table ; Preprocessing of car review text data includes: Remove HTML tags and garbled text from the car review text data, and discard the car review text data that does not contain primary indicators; Parse out the total vehicle mileage value in the vehicle usage information corresponding to the retained automotive review text data, and map the corresponding total vehicle mileage value to a preset mileage range; the specific preset mileage ranges are: "less than 2000", "greater than or equal to 2000 and less than 5000", "greater than or equal to 5000 and less than 10000", "greater than or equal to 10000 and less than 30k", "greater than or equal to 30000", with the unit being meters; According to the city names recorded in the purchase location corresponding to the retained automotive review text data, classify the corresponding purchase locations into geographical labels including the east or the middle or the west, and the south or the north; Parse the path of the comment post tag icon corresponding to the retained automotive review text data into a text tag type; Store any one of the processed automotive review text data in a JSON format file according to the electric vehicle brand and model to obtain the corresponding preprocessed review text; Integrate all the preprocessed review texts to construct the original automotive review dataset.
[0014] S2: According to the A-TFIDF values of each word in the preprocessed review text, use the adaptive K-means clustering algorithm to cluster the preprocessed review texts belonging to any one of the primary indicators to obtain the secondary indicators corresponding to the primary indicators.
[0015] Specifically, this step includes: S2.1: Denoise the text comment content belonging to any one of the primary indicators to construct a comment subset corresponding to the primary indicator; the calculation formula for the set of interfering words to be removed is: ; where, represents the set of interfering words under the th primary indicator , represents the frequency of occurrence of the th word in the set , represents the frequency of occurrence of the th word in the original automotive review dataset , represents the specificity threshold, and through this formula, identify and filter out the words that lack discrimination under specific indicators. The purpose of this step is to eliminate those "interfering words" (i.e., stop words, general words such as "of", "already", "feel", etc.) that appear frequently but are not helpful for distinguishing specific indicators from the original comments, so as to reduce data noise and improve the purity of subsequent feature mining.
[0016] S2.2: Calculate the A-TFIDF value of any word based on its frequency of occurrence within the comment subset corresponding to the primary metric; the formula for calculating the A-TFIDF value of a word is: ; ; ; in, Indicates the first Each primary indicator Next The first comment subset A-TF-IDF values of each word Indicates the first The first comment subset Weighted word frequency of each word Indicates the first-level indicator The total number of sub-sets of comments below Indicates the first-level indicator A collection of subsets of comments below. Indicates containing the first A subset of comments for each word in the set The proportion in This represents the Sigmoid activation function. Indicates the first The word in the first-level indicator The attribute specificity factor in Indicates the first The first comment subset The word frequency of each word Indicates the first The part-of-speech weight coefficient of each word. Indicates the first The word in the set Frequency of occurrence in Indicates the first-level indicator A collection of subsets of comments below. Indicates the first The word in the set Frequency of occurrence in This represents the smoothing term, which is a preset, minimal positive constant used to prevent the denominator from being zero when calculating the attribute specificity factor (ASF), thus ensuring the numerical stability of the calculation process. This step aims to calculate the importance weight of each word under a specific metric, considering not only word frequency but also introducing the "attribute specificity factor" to increase the weight of words that only frequently appear in the current metric (such as the correspondence between "stuttering" and "driving experience").
[0017] S2.3: Combine the A-TFIDF values of the top m words with the largest values in the comment subset corresponding to the primary indicator in descending order to form the comment vector of the comment subset corresponding to the primary indicator.
[0018] S2.4: Based on the average distance between the comment vectors of any comment subset corresponding to a primary indicator and the comment vectors of all other comment subsets belonging to the same primary indicator, and the average distance between the comment vectors of comment subsets in different primary indicators, calculate the silhouette coefficient of the corresponding comment subset. The formula for calculating the silhouette coefficient is: ; in, Indicates the first The silhouette coefficients of a subset of comments. Indicates the first The average distance between a subset of comments and the comment vectors of all other subsets of comments belonging to the same first-level metric. Indicates the first The average distance between a subset of comments and the comment vectors of each subset of comments in different first-level indicators can be calculated using conventional distance calculation formulas such as Euclidean distance.
[0019] S2.5: With the objective of maximizing the silhouette coefficient of the comment subset corresponding to the primary indicator, an adaptive K-means clustering algorithm is used for each comment subset to obtain the optimal number of clusters corresponding to the primary indicator and the corresponding number of clusters; the optimal number of clusters is calculated as follows: ; in, Indicates the first The total number of secondary indicators under each primary indicator is the optimal cluster number. Indicates the first-level indicator The total number of subsets of comments. This step aims to use the K-means algorithm to automatically cluster the comment text into several clusters, each cluster representing a potential secondary indicator topic (e.g., all comments discussing "backspace" would be clustered together).
[0020] S2.6: Calculate the intra-cluster importance and inter-cluster discriminancy of each word in each cluster, and retain the words whose intra-cluster importance and inter-cluster discriminancy are both greater than the corresponding thresholds. The retained words are the secondary indicators. This step aims to extract the most representative feature words from each cluster to define the name of the secondary indicator represented by that cluster and to construct the corresponding corpus.
[0021] It is worth noting that Cluster Internal Importance (CII) is a metric... w In its own clusterC The core representative of [the concept / method]. Its calculation formula is: ; in, Indicator In clusters Intra-cluster importance, Indicator In clusters word frequency in Indicator In clusters Inverse document frequency in Indicator In the full dataset Average word frequency in.
[0022] Cluster Inter-Distinction (CID): a metric w Compared to the unique discriminative power of other clusters, it avoids semantic overlap between indicators. Its calculation formula is: ; in, Indicator In clusters Inter-cluster discrimination in Indicates the absence of clusters Other clusters besides, Indicator In clusters Inverse document frequency in Indicator In the full dataset The average inverse document frequency in the data.
[0023] S3: Divide each preprocessed comment text belonging to any first-level indicator into length levels to obtain the length level of the corresponding preprocessed comment text.
[0024] Specifically, the steps include: Remove invalid characters from each preprocessed comment text to obtain the corresponding number of valid words. This step aims to calculate the "value" length of each comment by removing invalid characters (such as consecutive punctuation marks and emoticons), providing an accurate basis for subsequent judgment of comment length.
[0025] Calculate the 33rd percentile of the effective word count of all preprocessed comment texts belonging to any one of the primary indicators, and use it as the first boundary value corresponding to the primary indicator, as the boundary point between short text and medium-length text; Calculate the 66th percentile of the effective word count of all preprocessed comment texts belonging to any one of the primary indicators, and use it as the second boundary value corresponding to the primary indicator, as the boundary point between medium-length and long texts; Based on the first and second boundary values, the length levels of each preprocessed comment text corresponding to the first-level indicator are determined, as expressed by: ; in, This represents the length partitioning function. Indicates the first The first primary indicator A preprocessed comment text, Indicates preprocessed comment text Effective word count, Indicates the first The first dividing value corresponding to each primary indicator Indicates the first The second boundary value corresponding to each primary indicator This indicates that the length level is short. This indicates a medium length rating. The length level is indicated as "long." This dynamic classification formula eliminates the differences in length benchmarks between different indicators, ensuring relative fairness in the grading. This step aims to dynamically set the standards for "short," "medium," and "long" based on the commenting habits of different indicators (e.g., comments on "appearance" are usually shorter, while comments on "driving experience" are usually longer), avoiding a "one-size-fits-all" approach.
[0026] S4: Based on the pre-trained multi-task learning neural network information density prediction model, the preliminary information richness index of any pre-processed comment text is predicted. The preliminary information richness index is the number of secondary indicators covered in the corresponding pre-processed comment text.
[0027] In this embodiment, the pre-training process of the multi-task learning neural network information density prediction model includes: A standard sample set is sampled from the original dataset of the car reviews; For any primary metric, each preprocessed comment text in the standard sample set is manually labeled. Each label includes a label for the number of secondary metrics and a label for the probability of occurrence of the secondary metrics. This step aims to prepare high-quality data for training the "information density prediction model." Information density refers to the number of effective secondary metrics (such as "back row space" or "acceleration push-back feeling") contained in a comment.
[0028] Each preprocessed comment text is processed by the first pre-trained language model to generate a corresponding sentence vector.
[0029] Each preprocessed comment text is used to generate a corresponding TF-IDF vector based on the index corpus.
[0030] This step aims to transform unstructured comment text into high-dimensional numerical vectors using statistical methods to capture the explicit feature distribution of comments under a specific primary metric. The specific process includes the following sub-steps: Build a dedicated corpus and feature dictionary for each primary indicator: Extract all text belonging to this metric from the full preprocessed comment dataset and construct a metric-specific corpus. .right Perform global word frequency statistics and select the top M high-frequency words (excluding common stop words) as the feature dictionary for this indicator. .
[0031] Term Frequency (TF) calculation: For any preprocessed comment text in the standard sample set... d Calculate each feature word in the feature dictionary Frequency of occurrence in the text. To eliminate the influence of text length differences, normalization is used: ; in, Indicating characteristic words Preprocessing comment text word frequency in Indicating characteristic words Preprocessing comment text The number of times it appears in Indicating characteristic words Preprocessing comment text The number of times it appears in the text.
[0032] Inverse Document Frequency (IDF) Calculation: Based on an Indicator-Specific Corpus Calculate each feature word in the feature dictionary Importance weights. IDF reflects the scarcity of terms in the corpus, and is calculated using the following formula: ; in, Indicating characteristic words In the indicator-specific corpus Inverse document frequency in Indicator-specific corpus The total number of documents in the document. Indicator-specific corpus Contains feature words The number of documents.
[0033] Indicator-specific TF-IDF vector generation: The calculated... Value and Multiplying the values yields the feature words. Preprocessing comment text China's target indicators The feature weight values. Traversing the feature dictionary. The weight of each word is calculated sequentially, ultimately generating an M-dimensional numerical vector of the comment text. : .
[0034] Extract and combine the word count, sentence count, average sentence length, and punctuation density of any preprocessed comment text to generate corresponding structural statistical features.
[0035] The sentence vector, TF-IDF vector, and structural statistical features corresponding to the same preprocessed comment text are integrated to form a high-dimensional combined feature vector for the corresponding preprocessed comment text. This step aims to transform the original text into a numerical vector that can be understood by a computer, integrating statistical features, semantic features, and domain knowledge features to comprehensively describe the information content of the comment.
[0036] The high-dimensional combined feature vectors of each preprocessed comment text are passed through a multi-task learning neural network information density prediction model to output the corresponding prediction labels. The prediction labels include the predicted value of the number of secondary indicators and the predicted value of the probability of occurrence of secondary indicators. A loss function is calculated based on the predicted label and the labeled label of any preprocessed comment text. The information density prediction model of the multi-task learning neural network is optimized based on the loss function until the loss function converges, thus obtaining the pre-trained information density prediction model of the multi-task learning neural network.
[0037] In this embodiment, the multi-task learning neural network aims to capture the general semantic features of text through a shared parameter layer, and utilizes task-specific heads to achieve discrete prediction of the quantity of secondary indicators and continuous prediction of their occurrence probabilities. Its specific structure and parameter settings are as follows: Model hierarchy structure description: Input Layer: The receiving dimension is... High-dimensional combined feature vectors: ;in, For sentence vector dimension, For TF-IDF vector dimensions, This represents the dimension of structural statistical features.
[0038] Shared Representation Layers: Composed of multiple fully connected neural networks, with each layer's output mapped as follows: ;in, Represents the ReLU activation function. This indicates batch normalization processing. Represents a fully connected neural network The weight, Represents a fully connected neural network The bias, This represents a high-dimensional combined feature vector.
[0039] Task-Specific Headers: Quantitative prediction branch: The output layer uses the Softmax activation function to calculate the quantity of secondary indicators. n probability distribution: ;in, This represents the Softmax function. This indicates the weights of the output layer of the quantity prediction branch. This indicates the bias of the output layer of the quantity prediction branch.
[0040] Indicator identification branch: The output layer uses the Sigmoid activation function for calculation. K The probability vector of each secondary indicator. p : ;in, This represents the Sigmoid activation function. The indicator identifies the weights of the branch output layer. This indicates the bias of the output layer of the indicator identification branch.
[0041] It should be noted that the loss function used in the quantity prediction branch is the weighted cross-entropy loss function, while the loss function used in the indicator identification branch is the multi-label binary cross-entropy loss function. The loss function described in this embodiment is obtained by weighted summation of the weighted cross-entropy loss function and the multi-label binary cross-entropy loss function. The expression of the loss function is as follows: ; in, Represents the loss function. This represents the weighted cross-entropy loss function. This represents the multi-label binary cross-entropy loss function; , These represent the weight balance coefficients for the first task and the second task, respectively, both initially set to 0.5.
[0042] This step aims to design and train a multi-task learning neural network information density prediction model, simultaneously completing two tasks: "predicting the number of indicators" and "identifying specific indicators," thereby improving prediction accuracy by leveraging the correlation between tasks. The model is based on a shared feature representation layer and employs a hard parameter sharing mechanism. It includes at least a quantity prediction branch for outputting the distribution of the number of secondary indicators, and an indicator identification branch for outputting the probability of occurrence of each secondary indicator. The number of secondary indicators is estimated by summing the multi-label prediction results from the indicator identification branch.
[0043] As a further limitation of this embodiment, this step aims to train the model using labeled data and introduce a mechanism to evaluate the model's "confidence" in the prediction results, ensuring the reliability of the output. Supervised training of the model is performed using labeled data. A Monte Carlo Dropout mechanism is introduced during the prediction phase to calculate the variance of the prediction results as an uncertainty estimate, ensuring high-confidence output.
[0044] Using a trained model, batch predictions are performed on all unlabeled comments to obtain a "content value" score for each comment. The trained model is then used to infer the number of predicted secondary indicators for each comment from the entire set of unlabeled data. The preliminary information richness index is defined as the number of secondary indicators covered in the preprocessed comment text, serving as the core quantitative indicator for measuring the information density of the comment.
[0045] Furthermore, the step of sampling a standard sample set from the original dataset of car reviews includes: Calculate the proportion of the first sample of preprocessed review text belonging to various geographic tags in the original dataset of car reviews; Calculate the proportion of the second sample of preprocessed review texts belonging to various mileage ranges in the original dataset of car reviews; Calculate the proportion of third-sample preprocessed review texts belonging to various text tags (normal, featured, and max-level featured) in the original dataset of car reviews; This step aims to analyze the distribution of original comments by region, driving habits, and post quality, providing baseline data for subsequent "proportional sampling".
[0046] Based on the proportion of each of the first sample, the proportion of the second sample, and the proportion of the third sample, a three-dimensional quota sampling matrix is constructed; Based on the preset total number of standard samples, the proportion of the first sample, the proportion of the second sample, and the proportion of the third sample, preprocessed review texts that simultaneously belong to any one type of geographic label, any one type of mileage range, and any one type of text label are extracted from the original dataset of car reviews as the corresponding cell data in the three-dimensional quota sampling matrix; specifically: Based on the distribution ratios of the above three dimensions, a three-dimensional quota sampling matrix is constructed. Each cell of the matrix Preset target quantity representing a specific region, a specific mileage range, and a specific label type. : ; in, The total number of samples in the preset standard sample set is 10,000. Indicates geographic tags The proportion of the first sample Indicates mileage range The proportion of the second sample, Represents text label The third sample proportion. This step aims to establish a "reduced" database structure based on the statistical results of the above steps, ensuring that this smaller database can truly reflect the characteristics of big data.
[0047] The process continues until the data volume of each cell reaches the corresponding preset target quantity. All cell data is then aggregated to form the standard sample set, where the total preset target quantity is less than the total number of preprocessed comment texts in the original car review dataset. This step aims to randomly select comments that meet the criteria from a massive dataset according to the quota set in the matrix, ultimately synthesizing a highly representative "standard sample set."
[0048] The process involves iterating through the original car review dataset D, mapping each data point to a corresponding matrix cell based on its attributes. Random sampling without replacement is performed within each cell until the preset target number of samples for that cell is reached. If a cell has insufficient samples, they are supplemented from neighboring cells or generated using oversampling techniques. Finally, all cell data are aggregated to form a standard sample set that preserves the original data distribution characteristics.
[0049] S5: Based on the length level and the preliminary information richness index, determine the information density level of the corresponding preprocessed comment text, and distribute any preprocessed comment text to the set of pre-trained language models used to predict readability scores and sentiment scores based on the information density level, so as to obtain the readability score and sentiment score of the corresponding preprocessed comment text respectively.
[0050] It is worth noting that determining the information density level of the corresponding preprocessed comment text based on the length level and the preliminary information richness index includes: Based on the preliminary information richness index, the secondary index density level of the corresponding preprocessed comment text is determined. This step aims to combine the two dimensions of "length" and "density" of the comment to define the comprehensive status of the comment, in order to prepare for subsequent triage decisions.
[0051] In this embodiment, the grading criteria for length level and secondary index density level are shown in Table 2. Table 2 Grading Standards for Length Grades and Secondary Index Density Grades ; Based on the length level and the secondary index density level, the information density level of the corresponding preprocessed comment text is determined, and the expression is: ; ; in, The information density level classification function represents the classification function. Indicates the first The first primary indicator Preprocessed comment text Length class, Indicates the first The first primary indicator Preprocessed comment text The secondary indicator density level, This indicates a high information density level. This indicates that the information density level is normal. This indicates a low information density level. This indicates that the density level of the secondary indicator is high. This indicates that the density level of the secondary indicator is medium. This indicates that the density level of the secondary indicator is low. Indicates that, Indicates the first The first primary indicator Preprocessed comment text The preliminary information richness index value, Indicates the first The total number of secondary indicators under each primary indicator. This indicates rounding down to the nearest integer.
[0052] In this embodiment, the information density level classification is shown in Table 3; Table 3. Classification of Information Density Levels Long High / Mid Hard 5 In-depth long-term test, detailed driving log Long Low Normal 3 rambling, nonsensical, and copied-pasted advertorials Medium High Hard 5 Concise and insightful summary of key pain points Medium Mid Normal 3 Real feedback from ordinary users Medium Low Easy (Low) 1 General evaluation (e.g., "The space is okay") Short High Normal 3 Minimalist list of core parameters Short Mid / Low Easy (Low) 1 Random comments, meaningless spam Furthermore, a three-tiered model pool is constructed, corresponding to "high," "normal," and "low" information density levels, respectively, to support the aforementioned routing pattern. The system pre-configures at least five general-purpose pre-trained language models with different parameter sizes, and sorts them in ascending order of parameter size to form an ordered model set. In one exemplary implementation, the five models can be Qwen / Qwen2.5-7B-Instruct, Qwen / Qwen2-7B-Instruct, THUDM / glm-4-9b-chat, deepseek-ai / DeepSeek-V2.5, and internlm / internlm2_5-7b-chat, with their corresponding order corresponding to the ascending order of parameter size. The range of models selected based on different complexity routing modes is as follows: Easy: Only call the model with the fewest parameters. To reason; Normal: Calls the three models with smaller parameters. Integrate its output results; Hard: Call all five models Perform ensemble reasoning to achieve higher evaluation accuracy.
[0053] Specifically, the step of distributing any preprocessed comment text to a set of pre-trained language models used to predict readability scores and sentiment scores based on the information density level includes: When the information density level of the preprocessed comment text is high, the corresponding preprocessed comment text is input into five second pre-trained language models with different parameter sizes. Each second pre-trained language model outputs the corresponding readability score and sentiment score. The five readability scores are weighted and summed to obtain the readability score of the preprocessed comment text. The average of the five sentiment scores is taken to obtain the sentiment score of the preprocessed comment text. When the information density level of the preprocessed comment text is normal, the corresponding preprocessed comment text is input into three second pre-trained language models with different parameter sizes. Each second pre-trained language model outputs the corresponding readability score and sentiment score. The three readability scores are weighted and summed to obtain the readability score of the preprocessed comment text. The mean of the three sentiment scores is taken to obtain the sentiment score of the preprocessed comment text. When the information density level of the preprocessed comment text is low, the corresponding preprocessed comment text is input into a second pre-trained language model, and the second pre-trained language model outputs the readability score and sentiment score of the preprocessed comment text respectively.
[0054] It is worth noting that the readability score is calculated as follows: ; ; in, This represents the readability score, where N represents the number of second pre-trained language models. Indicates the first Consistency confidence weights of the second pre-trained language model Indicates the first The readability score output by the second pre-trained language model. Indicates the first The mean absolute deviation between the readability scores output by one second pre-trained language model and those output by other second pre-trained language models. This represents the sensitivity coefficient, used to control the sensitivity of occasional poor model scores among models to changes in weights. This step aims to synthesize the scores of multiple models, with models with high credibility receiving larger weights and models with low credibility (random scoring) receiving smaller weights. For readability scores (regression tasks), a weighted averaging strategy based on model consistency is used for integration.
[0055] The formula for calculating the sentiment score is: ; in, Indicates the emotional score. Indicates the first The sentiment scores output by the second pre-trained language model. This step aims to integrate the sentiment judgments of multiple models, following the majority opinion but considering probability confidence. For sentiment analysis (classification task), a probability averaging method (Soft Voting) is used for ensemble.
[0056] Based on this, according to the preprocessed comment text rating levels The numerical scale (1-5 points) is used for each sentiment level. Assign corresponding scores Calculate the probability-weighted expected value of the sentiment score. : ; This yields a continuous sentiment score ranging from 1 to 5. Simultaneously, predictive entropy is calculated as an indicator of uncertainty in sentiment analysis.
[0057] In this embodiment, the scoring criteria for readability are shown in Table 4, and the scoring criteria for sentiment are shown in Table 5. Table 4 Readability Scoring Criteria 1 point Unrighteous Garbled text, <5 characters, illogical statements "11111", "Top top top", "Not bad" 2 points Hollow / Vague Only adjectives, no details, generic phrases "It has a large interior space, a nice exterior, and a luxurious interior." 3 points Qualified / Specific Mentioning specific parts, and the sentences are fluent. "There's enough legroom in the back, and the trunk can fit two suitcases." 4 points Enrich / Scene Includes specific scenarios, data, and reference objects. "It wasn't crowded when three 180cm tall adults sat in it, and it took 500 kilometers to get back to my hometown." 5 points Essence / In-depth Competitive product comparison, extreme testing, and exquisite writing style "Compared to the Model Y, the suspension is softer, handling speed bumps crisply, but wind noise is slightly louder at 120 km / h..." Table 5. Affectiveness Scoring Criteria 1 point Strongly negative Angry, disappointed, and extremely critical Regret, being advised to quit, severe abnormal noise, malfunction 2 points Generally negative Dissatisfaction, complaints, specific shortcomings Not very good, a bit noisy, uses a lot of gas, and the workmanship is so-so. 3 points Neutral / Objective State the facts without obvious bias It's okay, average, acceptable, and at a normal level. 4 points Generally front Satisfaction, approval, specific advantages Good, sufficient, comfortable, fuel-efficient, and pleasing to the eye. 5 points Strongly positive Surprise, recommendation, high praise Exceeded expectations, perfect, most satisfying, YYDS S6: Based on the information density level, the readability score, and the sentiment score, integrate them into a primary indicator and / or a multi-dimensional quality assessment result corresponding to the preprocessed comment text.
[0058] Specifically, the steps include: Align the user rating values for the primary metrics in the preprocessed comment text with the corresponding sentiment scores in terms of dimensions; this step aims to bring the star ratings given by users (e.g., 5 stars) and the sentiment scores given by the model (e.g., 5 points) to the same dimension.
[0059] Calculate the aligned user rating values The sentiment score aligned with the The absolute difference between the values yields the alignment index of the preprocessed comment text; the calculation formula is: ; in, Indicates preprocessed comment text Alignment metrics Indicates the highest value. This represents the minimum value. This metric is used to measure the logical consistency between the content of the comment text and the user rating.
[0060] when When the text description is highly consistent with the score, it indicates that the text description and the score are very consistent.
[0061] when When this occurs, it indicates a serious logical contradiction (such as being irrelevant or ironic), and the text should be downgraded in subsequent evaluations.
[0062] This step aims to determine whether a user is "hypocritical" (e.g., giving a high score but then using abusive language), serving as an important reference for judging the quality of the comment.
[0063] The information density level, readability score, and alignment index are each normalized. The normalized readability score, normalized alignment index, and normalized information density level of the preprocessed comment text under any one primary index are then weighted and summed to obtain the comprehensive quality score of the preprocessed comment text under the corresponding primary index. The calculation formula is: ; in, Indicates preprocessed comment text The overall quality score ranges from [0,1]. Indicates preprocessed comment text The normalized information density level, Indicates preprocessed comment text The normalized readability score, Indicates preprocessed comment text The normalized alignment index, , , These are adjustable weight parameters.
[0064] The weighted arithmetic mean of the overall quality scores of the preprocessed comment text under all primary indicators is calculated to obtain the multidimensional quality assessment result of the preprocessed comment text; the calculation formula is: ; ; in, Indicates preprocessed comment text The corresponding multidimensional quality assessment results, Indicates preprocessed comment text The corresponding quality weights.
[0065] And / or, calculate the weighted arithmetic mean of the overall quality scores of all preprocessed comment texts under any one primary indicator to obtain the multidimensional quality assessment result corresponding to the primary indicator.
[0066] In this embodiment, statistical verification of label consistency is also included: This step aims to statistically verify that "maximum-level featured posts score the highest, followed by featured posts, and ordinary posts score the lowest." The average of the overall quality scores for the comment level is calculated for each of the three tag sets: ; ; ; In one exemplary embodiment, based on the original dataset of car reviews, data is extracted from the general post set. Collection of Featured Posts Collection of max-level featured posts N comments are randomly selected from each set to form a validation subset. A comprehensive quality score is then calculated for each comment using the method described above. and in the selection , , In this case, the average score of the samples is calculated for each of the three types of labels. , and When N=1000, the experimental results show that the overall quality scores of the entire user comments corresponding to the max-level featured posts, featured posts, and ordinary posts roughly fall within the ranges of 0.78~0.86, 0.66~0.74, and 0.44~0.50, respectively, with a representative mean of approximately... , and ,satisfy This relationship allows us to statistically verify that higher-level labels correspond to higher overall quality scores.
[0067] Furthermore, in the practical application of the same embodiment, a typical phenomenon reflecting the value of fine-grained evaluation was observed: selecting posts that are at the maximum level of excellence. Featured Posts and regular posts Each review is analyzed, and its overall quality score under the seven primary indicators is calculated. For example, if the seven primary indicators are space, driving experience, range, exterior, interior, value for money, and intelligent features, then the scores of the three reviews under each primary indicator can be expressed as follows: Table 6. Examples of quality scores for reviews of different tag levels under seven primary indicators. ; The results above show that: on the one hand, there are max-level, high-quality posts. The scores for indicators such as "driving experience," "range," and "intelligent features" were significantly higher than those for featured and regular posts, maintaining the highest level in the overall cross-indicator quality score of the entire review; on the other hand, in certain primary indicators such as "space" and "value for money," the regular posts... The score was even significantly higher than that of the highest-level featured posts, and the featured posts... Scores on metrics such as "appearance" may be slightly higher than those of top-tier featured posts. Simultaneously, it can be observed that featured posts and regular posts score 0 on some primary metrics such as "battery life" and "intelligent features," indicating that the corresponding comments completely fail to cover these aspects. These observations demonstrate that while higher-level tags do indeed correspond to higher overall quality scores, relying solely on a single tag is insufficient to characterize the structural differences between primary metrics. Introducing a fine-grained evaluation framework for primary metric quality scores can reveal the hidden strengths and weaknesses behind the tags, thereby further demonstrating the practical application value of this fine-grained evaluation mechanism.
[0068] The method shown in this embodiment can be applied to content management systems, intelligent recommendation systems, and user feedback analysis systems of automotive reputation communities, providing real-time or batch review quality assessment services through API interfaces.
[0069] The fine-grained automotive review quality assessment method based on size model collaboration provided in this embodiment has the following beneficial effects: First, we collected full-scale review data from automotive reputation communities. Focusing on seven primary indicators—space, driving experience, range, exterior, interior, cost-effectiveness, and intelligent features—we constructed a secondary, fine-grained indicator system covering the entire data set using attribute-enhanced A-TFIDF and adaptive K-means clustering. A standard sample set was then constructed using quota sampling, taking into account geographical distribution, mileage range, and tag proportions. Second, based on the standard sample set, we trained a multi-task learning neural network information density prediction model to predict the number of secondary indicators in each review, obtaining an information richness index. Combining this with text length grading results, an intelligent routing mechanism allocated reviews to large language model pools with different computing power levels according to task complexity. Under the architecture of "simple tasks handled by a small number of small models, complex tasks handled collaboratively by multiple large models," readability assessment and sentiment analysis were completed. Third, we constructed comprehensive quality scores and quality weights at the individual review and primary indicator levels. Then, we calculated the comprehensive quality score across the seven primary indicators for tagged reviews and conducted platform tag consistency verification experiments. Experimental results show that, overall, the comprehensive quality scores of max-level featured posts, featured posts, and ordinary posts follow a monotonic relationship. In practical applications, it was observed that some ordinary posts scored better than max-level featured posts in indicators such as space and cost-effectiveness. However, there were cases where featured posts and ordinary posts were completely uncovered in indicators such as range and intelligence. This proves that introducing fine-grained first-level indicator quality scores can reveal the hidden advantages and disadvantages behind the labels, significantly improving the accuracy and interpretability of car reputation quality assessment.
[0070] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for evaluating the quality of fine-grained car reviews based on size-model collaboration, characterized in that, include: S1: Acquire and preprocess car review text data to construct a raw car review dataset consisting of preprocessed review text; S2: Based on the A-TFIDF values of each word in the preprocessed comment text, an adaptive K-means clustering algorithm is used to cluster the preprocessed comment text belonging to any primary indicator, obtaining secondary indicators under the corresponding primary indicator, including: S2.1: Denoise the text comments belonging to any one of the primary indicators and construct the comment subset corresponding to the primary indicator; S2.2: Calculate the A-TFIDF value of any word based on its frequency of occurrence within the comment subset corresponding to the primary indicator; the formula for calculating the A-TFIDF value of a word is: ; ; ; in, Indicates the first Each primary indicator Next The first comment subset A-TF-IDF values of each word Indicates the first The first comment subset Weighted word frequency of each word Indicates the first-level indicator The total number of sub-sets of comments below Indicates the first-level indicator A collection of subsets of comments below. Indicates containing the first A subset of comments for each word in the set The proportion in This represents the Sigmoid activation function. Indicates the first The word in the first-level indicator The attribute specificity factor in Indicates the first The first comment subset The word frequency of each word Indicates the first Part-of-speech weight coefficient of each word Indicates the first The word in the set Frequency of occurrence in Indicates the first-level indicator A collection of subsets of comments below. Indicates the first The word in the set Frequency of occurrence in Indicates the smoothing term; S2.3: Combine the A-TFIDF values of the top m words with the largest values in the comment subset corresponding to the primary indicator in descending order to form the comment vector of the comment subset corresponding to the primary indicator; S2.4: Based on the average distance between the comment vectors of any comment subset corresponding to a primary indicator and the comment vectors of all other comment subsets belonging to the same primary indicator, and the average distance between the comment vectors of each comment subset in different primary indicators, calculate the silhouette coefficient of the corresponding comment subset; the formula for calculating the silhouette coefficient is: ; in, Indicates the first The silhouette coefficients of a subset of comments. Indicates the first The average distance between a subset of comments and the comment vectors of all other subsets of comments belonging to the same first-level metric. Indicates the first The average distance between a subset of comments and the comment vectors of each subset of comments in different first-level metrics; S2.5: With the objective of maximizing the silhouette coefficient of the comment subset corresponding to the primary indicator, an adaptive K-means clustering algorithm is used for each comment subset to obtain the optimal number of clusters corresponding to the primary indicator and the corresponding number of clusters; S2.6: Calculate the intra-cluster importance and inter-cluster discriminancy of each word in each cluster, and retain the words whose intra-cluster importance and inter-cluster discriminancy are both greater than the corresponding thresholds. The retained words are the secondary index; the formula for calculating intra-cluster importance is: ; in, Indicator In clusters Intra-cluster importance, Indicator In clusters word frequency in Indicator In clusters Inverse document frequency in Indicator In the full dataset Average word frequency in; The formula for calculating inter-cluster discrimination is: ; in, Indicator In clusters Inter-cluster discrimination in Indicates the absence of clusters Other clusters besides, Indicator In clusters Inverse document frequency in Indicator In the full dataset Average inverse document frequency in; S3: Divide each preprocessed comment text belonging to any first-level indicator into length levels to obtain the length level of the corresponding preprocessed comment text; S4: Based on the pre-trained multi-task learning neural network information density prediction model, the preliminary information richness index of any pre-processed comment text is predicted. The preliminary information richness index is the number of secondary indicators covered in the corresponding pre-processed comment text. S5: Based on the length level and the preliminary information richness index, determine the information density level of the corresponding preprocessed comment text, and distribute any preprocessed comment text to the set of pre-trained language models applied to predict readability scores and sentiment scores based on the information density level, so as to obtain the readability score and sentiment score of the corresponding preprocessed comment text respectively. S6: Based on the information density level, the readability score, and the sentiment score, integrate them into a primary indicator and / or a multi-dimensional quality assessment result corresponding to the preprocessed comment text.
2. The method according to claim 1, characterized in that, S1 includes: Extract text data of all car reviews for several target electric vehicle brands from the product detail pages of automotive vertical websites; Parse the following information from any given car review text data: review ID, review post tag icon path, user rating value for primary indicators, text review content containing primary indicators, car model name, purchase price information, purchase time, purchase location, and vehicle usage information. Preprocessing of car review text data includes: Remove HTML tags and garbled text from the car review text data, and discard the car review text data that does not contain primary indicators; Parse the total mileage value of the vehicle in the vehicle usage information corresponding to the retained car review text data, and map the corresponding total mileage value to a preset mileage range; Based on the city names recorded in the car purchase locations corresponding to the retained car review text data, the corresponding car purchase locations are classified into geographical tags including East or Central or West, and South or North. The retained car review text data is parsed into text tag type, corresponding to the review post tag icon path. Store any one of the processed car review text data as a JSON format file according to the electric vehicle brand and model to obtain the corresponding preprocessed review text; Integrate all the preprocessed review texts to construct the original dataset of car reviews.
3. The method according to claim 1, characterized in that, S3 include: Remove invalid characters from each preprocessed comment text to obtain the corresponding number of valid words; Calculate the 33rd percentile of the effective word count of all preprocessed comment texts belonging to any one of the primary indicators, and use it as the first boundary value corresponding to the primary indicator; Calculate the 66th percentile of the effective word count of all preprocessed comment texts belonging to any one of the primary indicators, and use it as the second boundary value corresponding to the primary indicator; Based on the first and second boundary values, the length levels of each preprocessed comment text corresponding to the first-level indicator are determined, as expressed by: ; in, This represents the length partitioning function. Indicates the first The first primary indicator A preprocessed comment text, Indicates preprocessed comment text Effective word count Indicates the first The first dividing value corresponding to each primary indicator Indicates the first The second boundary value corresponding to each primary indicator This indicates that the length level is short. This indicates a medium length rating. This indicates that the length level is long.
4. The method according to claim 1, characterized in that, The pre-training process of a multi-task learning neural network information density prediction model includes: A standard sample set is sampled from the original dataset of the car reviews; For any primary indicator, each preprocessed comment text in the standard sample set is manually labeled, and each label includes a label for the number of secondary indicators and a label for the probability of occurrence of secondary indicators. Each preprocessed comment text is processed by the first pre-trained language model to generate a corresponding sentence vector; Each preprocessed comment text is used to generate a corresponding TF-IDF vector based on the index corpus. Extract and combine the word count, sentence count, average sentence length, and punctuation density of any preprocessed comment text to generate the corresponding structural statistical features; The sentence vector, the TF-IDF vector, and the structural statistical features corresponding to the same preprocessed comment text are integrated to form a high-dimensional combined feature vector of the corresponding preprocessed comment text. The high-dimensional combined feature vectors of each preprocessed comment text are passed through a multi-task learning neural network information density prediction model to output the corresponding prediction labels. The prediction labels include the predicted value of the number of secondary indicators and the predicted value of the probability of occurrence of secondary indicators. A loss function is calculated based on the predicted label and the labeled label of any preprocessed comment text. The information density prediction model of the multi-task learning neural network is optimized based on the loss function until the loss function converges, thus obtaining the pre-trained information density prediction model of the multi-task learning neural network.
5. The method according to claim 4, characterized in that, The process of sampling a standard sample set from the original dataset of car reviews includes: Calculate the proportion of the first sample of preprocessed review text belonging to various geographic tags in the original dataset of car reviews; Calculate the proportion of the second sample of preprocessed review texts belonging to various mileage ranges in the original dataset of car reviews; Calculate the proportion of third samples of preprocessed review texts belonging to various text tags in the original dataset of car reviews; Based on the proportion of each of the first sample, the proportion of the second sample, and the proportion of the third sample, a three-dimensional quota sampling matrix is constructed; Based on the preset total number of standard samples, the proportion of the first sample, the proportion of the second sample, and the proportion of the third sample, each preprocessed review text that simultaneously belongs to any type of geographical label, any type of mileage range, and any type of text label is extracted from the original dataset of car reviews as the corresponding cell data in the three-dimensional quota sampling matrix. The process continues until the amount of data in each cell reaches the corresponding preset target quantity. All cell data are then aggregated to form the standard sample set. The total value of the preset target quantity is less than the total number of preprocessed review texts in the original dataset of car reviews.
6. The method according to claim 3, characterized in that, The process of determining the information density level of the corresponding preprocessed comment text based on the length level and the preliminary information richness index includes: The secondary index density level of the corresponding preprocessed comment text is determined based on the preliminary information richness index. Based on the length level and the secondary index density level, the information density level of the corresponding preprocessed comment text is determined, and the expression is: ; ; in, The information density level classification function represents the classification function. Indicates the first The first primary indicator Preprocessed comment text Length class, Indicates the first The first primary indicator Preprocessed comment text The secondary indicator density level, This indicates a high information density level. This indicates that the information density level is normal. This indicates a low information density level. This indicates that the density level of the secondary indicator is high. This indicates that the density level of the secondary indicator is medium. This indicates that the density level of the secondary indicator is low. Indicates that, Indicates the first The first primary indicator Preprocessed comment text The preliminary information richness index value, Indicates the first The total number of secondary indicators under each primary indicator. This indicates rounding down to the nearest integer.
7. The method according to claim 6, characterized in that, The step of distributing any preprocessed comment text to a set of pre-trained language models used to predict readability scores and sentiment scores based on the information density level includes: When the information density level of the preprocessed comment text is high, the corresponding preprocessed comment text is input into five second pre-trained language models with different parameter sizes. Each second pre-trained language model outputs the corresponding readability score and sentiment score. The five readability scores are weighted and summed to obtain the readability score of the preprocessed comment text. The average of the five sentiment scores is taken to obtain the sentiment score of the preprocessed comment text. When the information density level of the preprocessed comment text is normal, the corresponding preprocessed comment text is input into three second pre-trained language models with different parameter sizes. Each second pre-trained language model outputs the corresponding readability score and sentiment score. The three readability scores are weighted and summed to obtain the readability score of the preprocessed comment text. The mean of the three sentiment scores is taken to obtain the sentiment score of the preprocessed comment text. When the information density level of the preprocessed comment text is low, the corresponding preprocessed comment text is input into a second pre-trained language model, and the second pre-trained language model outputs the readability score and sentiment score of the preprocessed comment text respectively.
8. The method according to claim 2, characterized in that, S6 include: Align the user rating values for the primary metrics in the preprocessed comment text with the corresponding sentiment scores in terms of dimensions; The absolute difference between the aligned user rating value and the aligned sentiment score is calculated to obtain the alignment index of the preprocessed comment text. The information density level, the readability score, and the alignment index are normalized respectively. The normalized readability score, the normalized alignment index, and the normalized information density level of the preprocessed comment text under any one primary index are weighted and summed to obtain the comprehensive quality score of the preprocessed comment text under the corresponding primary index. The weighted arithmetic mean of the overall quality scores of the preprocessed comment text under all primary indicators is calculated to obtain the multidimensional quality assessment results of the preprocessed comment text. And / or, calculate the weighted arithmetic mean of the overall quality scores of all preprocessed comment texts under any one primary indicator to obtain the multidimensional quality assessment result corresponding to the primary indicator.