Comment classification method and device, equipment, storage medium and program product
By extracting features, performing topic clustering and frequency statistics on user review data, and automatically screening and sorting them, we can solve the problem of low efficiency in manual screening, achieve timely and efficient after-sales problem handling, and improve user experience.
Patent Information
- Application Number
- CN202510844112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-23
AI Technical Summary
In the existing technology, the screening of user review data requires manual processing, which leads to a waste of manpower and material resources and low accuracy, and fails to respond to user needs in a timely manner, thereby reducing the user experience.
By performing feature extraction, topic clustering and frequency statistics on comment data, the demand intensity classification of each topic information is determined, and automated screening and priority sorting are achieved.
No manual screening is required, which saves manpower and material resources, and can promptly and quickly handle after-sales issues with strong demand, thus improving user experience.
Smart Images

Figure CN120687616A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a comment classification method, apparatus, device, storage medium, and program product. Background Art
[0002] As competition in the automotive market intensifies, after-sales service has become a key factor influencing consumer satisfaction and brand loyalty. High-quality after-sales service not only improves customer satisfaction but also enhances a brand's market competitiveness. In this context, user reviews of vehicle usage provide a valuable indicator of the user experience. For example, reviews could include comments about the engine's lack of power, unclear photos from the vehicle's camera, or poor brake performance.
[0003] Usually, staff will manually screen user review data and select review data that needs priority after-sales processing to improve the vehicle user experience.
[0004] However, the above method requires manual screening of comment data, which wastes manpower and material resources. In addition, the accuracy of manual screening is not high and it cannot respond to user needs in a timely manner, thereby reducing the user experience. Summary of the Invention
[0005] This disclosure provides a review classification method that can solve the technical problems existing in related technologies. The technical solution is as follows:
[0006] In a first aspect, a comment classification method is provided, the method comprising:
[0007] Performing feature extraction on the acquired multiple comment data to obtain a feature vector corresponding to each comment data;
[0008] Performing topic clustering processing on the feature vectors corresponding to the plurality of comment data to obtain topic information corresponding to each comment data;
[0009] Performing frequency statistical processing on the topic data corresponding to the plurality of comment data to obtain the frequency of occurrence of each topic information;
[0010] Based on the topic information corresponding to each review data, the occurrence frequency of each topic information and the classification model, the demand strength classification corresponding to each topic information is determined, wherein the demand strength classification corresponding to the topic information is used to indicate the after-sales processing priority level of the topic information.
[0011] In a possible implementation, the feature extraction process is performed on each of the acquired multiple comment data to obtain a feature vector corresponding to each comment data, including:
[0012] Determining a topic feature vector corresponding to each comment data based on the plurality of comment data and a topic feature extraction model;
[0013] Determining a semantic feature vector corresponding to each comment data based on the plurality of comment data and a semantic feature extraction model;
[0014] Based on the topic feature vector and the semantic feature vector corresponding to each comment data, a feature vector corresponding to each comment data is determined.
[0015] In a possible implementation, obtaining the feature vector corresponding to each comment data based on the topic feature vector corresponding to each comment data and the semantic feature vector corresponding to each comment data includes:
[0016] The topic feature vector and the semantic feature vector corresponding to each comment data are concatenated to obtain a feature vector corresponding to each comment data.
[0017] In a possible implementation, the method further includes:
[0018] Determining, based on the plurality of comment data and the sentiment intensity model, a sentiment intensity value corresponding to each comment data, wherein the sentiment intensity value corresponding to the comment data is used to indicate the sentiment intensity degree and sentiment polarity of the comment data;
[0019] The determining, based on the plurality of review data, the topic information corresponding to each review data, the occurrence frequency of each topic information, and the classification model, of the demand intensity classification corresponding to each topic information includes:
[0020] Based on the topic information corresponding to each comment data, the occurrence frequency of each topic information, the sentiment intensity value corresponding to each comment data and the classification model, the demand intensity classification corresponding to each topic information is determined.
[0021] In a possible implementation, before performing feature extraction on the acquired plurality of comment data to obtain a feature vector corresponding to each comment data, the method further includes:
[0022] Get multiple initial review data;
[0023] Text preprocessing is performed on the multiple initial comment data to obtain the comment data.
[0024] In a possible implementation, the text preprocessing includes at least one of text deduplication processing, compression and word removal processing, short sentence deletion processing, pure symbol deletion processing, pure English deletion processing, text word segmentation processing, and stop word removal processing.
[0025] In a second aspect, a comment classification device is provided, the device comprising:
[0026] A feature extraction module is used to perform feature extraction processing on the acquired multiple comment data respectively to obtain a feature vector corresponding to each comment data;
[0027] A clustering module, configured to perform topic clustering processing on the feature vectors corresponding to the plurality of comment data to obtain topic information corresponding to each comment data;
[0028] A statistical module is used to perform frequency statistical processing on the topic data corresponding to the plurality of comment data to obtain the frequency of occurrence of each topic information;
[0029] A classification module is used to determine the demand intensity classification corresponding to each topic information based on the topic information corresponding to each review data, the frequency of occurrence of each topic information and the classification model, wherein the demand intensity classification corresponding to the topic information is used to indicate the after-sales processing priority level of the topic information.
[0030] In a possible implementation, the feature extraction module is configured to:
[0031] Determining a topic feature vector corresponding to each comment data based on the plurality of comment data and a topic feature extraction model;
[0032] Determining a semantic feature vector corresponding to each comment data based on the plurality of comment data and a semantic feature extraction model;
[0033] Based on the topic feature vector and the semantic feature vector corresponding to each comment data, a feature vector corresponding to each comment data is determined.
[0034] In a possible implementation, the feature extraction module is configured to:
[0035] The topic feature vector and the semantic feature vector corresponding to each comment data are concatenated to obtain a feature vector corresponding to each comment data.
[0036] In a possible implementation, the apparatus further includes a determining module configured to:
[0037] Determining, based on the plurality of comment data and the sentiment intensity model, a sentiment intensity value corresponding to each comment data, wherein the sentiment intensity value corresponding to the comment data is used to indicate the sentiment intensity degree and sentiment polarity of the comment data;
[0038] The classification module is used to:
[0039] Based on the topic information corresponding to each comment data, the occurrence frequency of each topic information, the sentiment intensity value corresponding to each comment data and the classification model, the demand intensity classification corresponding to each topic information is determined.
[0040] In a possible implementation, the apparatus further includes a preprocessing module, configured to:
[0041] Get multiple initial review data;
[0042] Text preprocessing is performed on the multiple initial comment data to obtain the comment data.
[0043] In a possible implementation, the text preprocessing includes at least one of text deduplication processing, compression and word removal processing, short sentence deletion processing, pure symbol deletion processing, pure English deletion processing, text word segmentation processing, and stop word removal processing.
[0044] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the comment classification method.
[0045] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the operations performed by the comment classification method.
[0046] In a fifth aspect, a computer program product is provided, wherein the computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the comment classification method.
[0047] The beneficial effects brought about by the technical solution provided by the present disclosure are as follows: the solution mentioned in the present disclosure can determine the subject information corresponding to each review data through feature extraction processing and subject clustering processing, and can also determine the demand intensity classification corresponding to each subject information through a classification model. In this way, the staff can prioritize each subject information according to the demand intensity classification corresponding to different subject information, and handle the after-sales problems corresponding to each subject information according to the priority ranking. In this way, there is no need to manually screen and process the review data, which saves manpower and material resources. In addition, the staff can directly handle the after-sales problems corresponding to each subject information according to the demand intensity classification, and can promptly and quickly handle after-sales problems with stronger demand, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0049] Figure 1 This is a flowchart of an evaluation classification method provided by an embodiment of the present disclosure;
[0050] Figure 2 This is a flowchart of an evaluation classification method provided by an embodiment of the present disclosure;
[0051] Figure 3 is a structural diagram of an evaluation and classification device provided by an embodiment of the present disclosure;
[0052] Figure 4 This is a structural block diagram of a terminal provided by an embodiment of the present disclosure;
[0053] Figure 5 This is a structural block diagram of a server provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0055] The present disclosure provides a review classification method that can be implemented by a computer device, which can be a terminal or a server, and a terminal can be a desktop computer, a notebook computer, a tablet computer, a mobile phone, etc.
[0056] A computer device may include a processor, memory, communication components, etc.
[0057] The processor can be a central processing unit (CPU), which can be used to read instructions and process data, for example, performing feature extraction processing on the multiple comment data obtained to obtain the feature vector corresponding to each comment data, performing topic clustering processing on the feature vectors corresponding to the multiple comment data to obtain the topic information corresponding to each comment data, performing frequency statistical processing on the topic data corresponding to the multiple comment data to obtain the frequency of occurrence of each topic information, determining the demand intensity classification corresponding to each topic information, and so on.
[0058] The memory may be any volatile memory or non-volatile memory, such as a solid state disk (SSD) or dynamic random access memory (DRAM). The memory may be used for data storage, for example, storing a plurality of acquired review data, storing a feature vector corresponding to each acquired review data, storing topic information corresponding to each acquired review data, storing a demand intensity classification corresponding to each determined topic information, and so on.
[0059] The communication component may be a wired network connector, a wireless fidelity (WiFi) module, a Bluetooth module, a cellular network communication module, etc. The communication component may be used to perform data transmission with other devices.
[0060] The comment classification method provided by the embodiment of the present disclosure can be applied to a vehicle. A host is provided on the vehicle, and the user can comment on the use of the vehicle through a comment application provided on the host. The host obtains the comment data (or initial comment data) filled in by the user and can send it to the background server. The background server classifies the multiple comment data obtained to obtain the demand intensity classification of different subject information corresponding to the multiple comment data, and then provides priority sorting of each subject information and corresponding after-sales improvement methods based on the demand intensity classification, thereby realizing timely optimization of the vehicle and improving the user's driving experience.
[0061] Of course, in addition to being applied to vehicles, the review classification method can also be applied to other reasonable review scenarios, such as classifying usage reviews of application programs, classifying usage reviews of mechanical equipment, classifying usage reviews of electrical appliances, etc., and the embodiments of the present disclosure do not limit this.
[0062] Below, taking the vehicle application scenario as an example, the review classification method provided by the embodiment of the present disclosure is introduced in detail.
[0063] Figure 1 and Figure 2 This is a flowchart of a review classification method provided by an embodiment of the present disclosure. Figure 1 and Figure 2 , the embodiment includes:
[0064] 101. Perform feature extraction on the acquired multiple comment data to obtain a feature vector corresponding to each comment data.
[0065] During implementation, when the preset conditions are met, multiple review data can be obtained. The review data is the user's evaluation of the vehicle's usage, for example, the engine power is not strong, the music software in the car has a comprehensive range of songs, the windows are difficult to control, etc. The review data can reflect the user's dissatisfaction or satisfaction after using the vehicle.
[0066] After obtaining the plurality of comment data, feature extraction may be performed on each of the plurality of comment data to obtain a feature vector corresponding to each comment data, where the feature vector is a semantic representation of the corresponding comment data.
[0067] In one possible implementation, the aforementioned preset condition may be that the time since the last review classification reaches a preset period, and / or the current time reaches the preset review classification time of the target project. The preset period may be one month, 20 days, etc., and the preset review classification time of the target project may be one month, two months, three months, etc. after the target project is launched, which is not specifically limited in the present embodiment.
[0068] It is understood that when the preset condition "the time since the last comment classification reaches the preset period" is triggered, the comment data obtained is the comment data entered by the user after the last comment classification. When the preset condition "the current time reaches the preset comment classification time of the target project" is triggered, the comment data obtained is the comment data entered by the user since the target was launched.
[0069] In a possible implementation, a method for performing feature extraction processing on a plurality of comment data respectively may be: determining a feature vector corresponding to each comment data based on the plurality of comment data and a feature extraction model.
[0070] In implementation, multiple review data may be input into a trained feature extraction model respectively to obtain a feature vector corresponding to each review data output.
[0071] In the embodiment of the present disclosure, the feature extraction model can be any reasonable deep learning model, for example, it can be a semantic feature extraction model, etc.
[0072] In another possible implementation, the method of performing feature extraction processing on multiple comment data separately can also be: based on multiple comment data and a topic feature extraction model, determining the topic feature vector corresponding to each comment data; based on multiple comment data and a semantic feature extraction model, determining the semantic feature vector corresponding to each comment data; based on the topic feature vector and the semantic feature vector corresponding to each comment data, determining the feature vector corresponding to each comment data.
[0073] During implementation, multiple review data can be input into the trained topic feature extraction model respectively to obtain the topic feature vector corresponding to each review data. The topic feature vector corresponding to the review data focuses on representing the topic meaning of the review data.
[0074] It is also possible to input multiple comment data into the trained semantic feature extraction model to obtain the semantic feature vector corresponding to each comment data. The semantic feature vector corresponding to the comment data focuses on representing the meaning of the entire sentence of the comment data.
[0075] Then, for each comment data, the topic feature vector corresponding to the comment data and the semantic feature vector corresponding to the comment data can be fused to obtain the feature vector corresponding to the comment data. The feature vector contains both the topic meaning of the comment data and the semantic meaning of the comment data, and can more accurately reflect the problems pointed out by the comment data, thereby improving the accuracy of comment classification.
[0076] In the disclosed embodiment, the topic feature extraction model can be an LDA (Latent Dirichlet Allocation, topic analysis model) topic model. LDA is a classic method for document modeling in natural language processing. Its core principle is to regard each document as a mixed distribution of multiple topics, and each topic is composed of a probability distribution of a set of words. LDA assumes that each word in a document is generated by the topic to which it belongs, and each topic is generated by the words in the vocabulary set according to a certain probability distribution. By iteratively optimizing the document-topic distribution and the topic-vocabulary distribution, LDA can maximize the generation probability of words in a document, thereby revealing the potential topic structure in the document set.
[0077] In the LDA topic model, the review data can be composed of word sequences w = {w1,w2,...,w n}, and the entire corpus consists of a series of review texts D = {d1, d2, ..., d m}, where n represents the number of word segments for each comment data, and m represents the number of comment texts in the corpus. The probability relationship between the potential topic and the word w calculated by the LDA topic model can be expressed as:
[0078]
[0079] Among them, z i Represents word w in the document i Corresponding topic.
[0080] Of course, the topic feature extraction model may also be a learning model of other structures, which is not limited in the embodiments of the present disclosure.
[0081] In the embodiments of the present disclosure, the semantic feature extraction model can be a BERT (Bidirectional Encoder Representations from Transformers) model or a SBert (Sentence Bidirectional Encoder Representations from Transformers, a feature extraction model improved based on the BERT model), etc. Of course, the semantic feature extraction model can also be any other reasonable feature extraction model, which is not limited in the embodiments of the present disclosure.
[0082] When training the semantic feature extraction model, a twin network architecture can be adopted, that is, two weight-sharing semantic feature extraction models are used to encode the input comment data respectively, and a pooling operation is performed on the output of all tokens in the last layer to obtain the predicted semantic feature vector corresponding to the comment data, and then the vector similarity (such as cosine similarity, etc.) between the predicted semantic feature vectors of the two comment data is calculated. Then, based on the actual similarity of the two comment data, the calculated vector similarity and the loss function, the loss value is calculated to train the semantic feature extraction model.
[0083] Furthermore, a fusion method for determining the feature vector corresponding to each comment data based on the topic feature vector and semantic feature vector corresponding to each comment data can be: separately splicing the topic feature vector and semantic feature vector corresponding to each comment data to obtain the feature vector corresponding to each comment data.
[0084] In implementation, for each comment data, the topic feature vector and the semantic feature vector corresponding to the comment data may be directly concatenated to obtain the feature vector corresponding to each comment data.
[0085] Alternatively, for each comment data, the topic feature vector and semantic feature vector corresponding to the comment data can be added together to obtain the feature vector corresponding to each comment data, and so on. Of course, other reasonable fusion methods can also be used, and the embodiments of the present disclosure do not limit this.
[0086] In a possible implementation, the comment data may be initial comment data input by a user into a host of the vehicle, and the initial comment data may also include text input data and / or text data corresponding to voice input data.
[0087] In another possible implementation, the comment data may be data obtained by pre-processing the initial comment data input by the user into the vehicle host. Before performing feature extraction on the obtained multiple comment data, the following processing may be performed:
[0088] Acquire multiple initial review data; perform text preprocessing on the multiple initial review data to obtain review data.
[0089] In implementation, the initial comment data may contain a lot of meaningless text content. For example, the user may repeat a sentence multiple times when inputting, or when the user outputs a comment through voice, the text data corresponding to the voice will also contain many repeated modal particles, etc.
[0090] Therefore, after obtaining the initial comment data, text preprocessing can be performed on it to delete meaningless text in the initial comment data, so as to obtain comment data that more accurately expresses the comment issue.
[0091] In one possible implementation, text preprocessing may include at least one of text deduplication, compression and word removal, short sentence deletion, pure symbol deletion, pure English deletion, text word segmentation, and stop word removal. Of course, it may also include other text preprocessing, which is not limited in the embodiments of the present disclosure.
[0092] Among them, text deduplication processing refers to: deleting part of multiple initial comment data that are completely repeated by the same user within a short period of time, and only retaining one of the initial comment data.
[0093] Compression word removal processing means: since vehicle users have two input modes, voice and typing, when using review applications, repeated words may appear in the same sentence of user feedback, such as "my my my...", "car map map map...", etc., which affect the semantic expression of the entire sentence. Therefore, meaningless repeated words such as "my" and "map" are compressed.
[0094] Short sentence deletion refers to deleting initial comment data that is too short. Extremely short sentences can express little information and are often caused by "incomplete input" or "misoperation" by the user. Therefore, a lower limit for sentence length is set during text preprocessing, and initial comment data with a sentence length shorter than this lower limit is deleted to ensure the integrity of the semantic expression of the corpus.
[0095] Pure symbol deletion processing refers to: using the regular matching method to delete the meaningless short comments of pure symbols and pure English in the initial comment data. Because pure symbol garbled text is meaningless, the initial comment data of pure symbols is deleted.
[0096] The pure English deletion process means that if the review classification method provided by the embodiment of the present disclosure is targeted at the domestic car market, the various algorithms used are all in Chinese, so the pure English text is deleted.
[0097] Text segmentation processing means that the initial comment data can be segmented to obtain segmentation information of each initial text data, that is, the comment data includes segmentation information.
[0098] Stop word removal refers to removing some meaningless words such as auxiliary words and modal particles from the initial comment data, as these words will interfere with subsequent analysis operations such as frequency statistics.
[0099] Typically, one or more of the above text preprocessing steps can enable the obtained comment data to more accurately express the issues raised by users, thereby simplifying subsequent calculations and improving classification accuracy.
[0100] 102. Perform topic clustering processing on the feature vectors corresponding to the multiple comment data to obtain topic information corresponding to each comment data.
[0101] In implementation, after obtaining the feature vector corresponding to each comment data in multiple comment data, topic clustering processing can be performed based on these multiple feature vectors, that is, these multiple feature vectors are matched to the corresponding topic information respectively, so as to obtain the topic information corresponding to each feature vector, and then obtain the topic information corresponding to each comment data.
[0102] Among them, the subject information may include host hardware, body hardware, ecological software, host interconnection, camera, power system, fuel consumption tank, etc.
[0103] Among them, host hardware can include audio, central control, passenger screen, sight light screen, return button, chip, black screen, etc. Body hardware can include the body, trunk, door buttons, hard buttons, luggage rack, tires, brakes, wipers, airbags, seats, rearview mirrors, etc. Ecosystem software can include navigation applications, various video applications, various music applications, various shopping applications, scene modes, application management, personal center, map applications, etc. Host interconnection can include various interconnection application software. Cameras can include panoramic cameras, radar, cameras, facial recognition devices, panoramic images, etc. The power system can include the engine, coolant, transmission, etc. The fuel tank can include the fuel tank, throttle, fuel consumption, etc.
[0104] In addition to the multiple subject information listed above, other subject information may also be included, which is not limited in the embodiments of the present disclosure.
[0105] In the embodiment of the present disclosure, the above-mentioned topic clustering process can be any reasonable clustering algorithm, for example, it can be a K-means algorithm and the like.
[0106] 103. Perform frequency statistical processing on the topic data corresponding to the multiple comment data to obtain the frequency of occurrence of each topic information.
[0107] In practice, after the review data are categorized by topic through the above step 102, the occurrence frequencies of the plurality of topic information can be calculated, that is, the occurrence frequencies of each topic information in the plurality of review data can be calculated.
[0108] For example, for a certain topic information, you can first calculate the number of comment data that contain the topic information in these multiple comment data, record it as the first number, and then calculate the ratio between the first number and the total number of these multiple comment data. This ratio is the frequency of occurrence of the topic information.
[0109] 104. Based on the topic information corresponding to each review data, the occurrence frequency of each topic information and the classification model, determine the demand intensity classification corresponding to each topic information.
[0110] In implementation, the topic information corresponding to the multiple comment data and the frequency of occurrence of the multiple topic information can be input into the classification model to obtain the demand intensity classification corresponding to each output topic information. The demand intensity classification corresponding to the topic information is used to indicate the after-sales processing priority of the topic information.
[0111] Among them, the demand intensity classification can include basic needs, expected needs, attractive needs and irrelevant needs.
[0112] The topic information corresponding to basic needs is a topic that users are highly concerned about. This type of implementation may not inspire users to agree or share, but if such needs are not met, users are likely to generate and express strong negative emotions.
[0113] The topic information corresponding to the expected needs is also a topic that users are highly concerned about. The realization of such needs will stimulate users' approval or sharing. If such needs are not met, the intensity of the emotions expressed by users will be lower than that of basic needs.
[0114] The topic information corresponding to the charm demand is a topic that users pay general attention to, and the controversy level of user feedback fluctuates greatly, which is a type of demand that is in a swing position.
[0115] The topic information corresponding to irrelevant needs is topics that users pay less attention to, and comments that have little or no relevance to the actual product.
[0116] The above classification model is used to obtain the demand intensity classification corresponding to each topic information. Then, these review data are distributed to each after-sales professional department of the automobile company according to the category of the topic information. The professional department formulates improvement strategies based on the urgency of the demand intensity classification and the category of the demand, and realizes an "end-to-end" solution from user review data to after-sales problems.
[0117] After achieving end-to-end demand identification, the automated approach described above can be used to regularly cycle through newly added review data, achieving targeted tracking of demand issues. Regular statistics are collected for each topic and each demand intensity category, enabling automated evaluation of the after-sales department's resolution effectiveness.
[0118] In the embodiment of the present disclosure, the classification model can be any reasonable algorithm model or machine learning model, such as the Kano model, etc., and the embodiment of the present disclosure is not limited to this.
[0119] In one possible implementation, before determining the demand intensity classification, the following processing can also be performed: based on multiple comment data and a sentiment intensity model, determine the sentiment intensity value corresponding to each comment data, wherein the sentiment intensity value corresponding to the comment data is used to indicate the sentiment intensity degree and sentiment polarity of the comment data.
[0120] In implementation, multiple comment data may be input into the sentiment intensity model respectively to obtain the sentiment intensity value corresponding to each comment data as output. The sentiment intensity value may range from 0 to 1.
[0121] 0 represents extremely negative sentiment, 0.5 represents neutral sentiment, and 1 represents extremely positive sentiment. Values between 0 and 0.5 represent negative sentiment, with values closer to 0 indicating a stronger negativity. Values between 0.5 and 1 represent positive sentiment, with values closer to 1 indicating a stronger positive sentiment.
[0122] Furthermore, when determining the demand intensity classification, the following processing can be performed: based on the topic information corresponding to each comment data, the frequency of occurrence of each topic information, the emotional intensity value corresponding to each comment data and the classification model, determine the demand intensity classification corresponding to each topic information.
[0123] The specific processing method can be: the topic information corresponding to the comment data, the frequency of occurrence of multiple topic information, and the emotional intensity values corresponding to multiple comment data can be input into the classification model for classification, and the demand intensity classification corresponding to each topic information can be obtained as the output.
[0124] Alternatively, the specific processing method may be: based on the sentiment intensity value corresponding to each comment data and the subject information corresponding to each comment data, an average sentiment score corresponding to each subject information is calculated. The average sentiment score corresponding to the subject information is used to represent the degree of emotion of the user in the comment data towards the subject information. If the average sentiment score is low, it means that the user's emotion towards the subject information is not strong, and the priority of after-sales processing can be lowered; if the average sentiment score is high, it means that the user's emotion towards the subject information is strong, and the priority of after-sales processing can be higher. In this way, major user issues can be resolved in a timely manner.
[0125] The method for calculating the average sentiment score can be: for each topic information, add the sentiment intensity values corresponding to one or more comment data corresponding to the topic information to obtain a first value, and then calculate the ratio of the first value to the number of comment data corresponding to the topic information. This ratio is the average sentiment score corresponding to the topic information.
[0126] After obtaining the average sentiment score corresponding to each topic information based on the above method, the average sentiment score corresponding to each topic information can be determined through the Kano model. The processing method of the Kano formula is shown in the following formula (2).
[0127]
[0128] Wherein, t is the topic information, K(t) is the demand intensity classification corresponding to the topic information t, f(t) is the frequency of occurrence corresponding to the topic information t, S(t) is the average sentiment score corresponding to the topic information t, α, β, δ, γ, and θ are all custom values, and their values can be set according to needs or actual conditions, and are not specifically limited in the embodiments of the present disclosure.
[0129] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0130] The solution mentioned in the embodiment of the present disclosure can determine the subject information corresponding to each review data through feature extraction processing and subject clustering processing, and can also determine the demand intensity classification corresponding to each subject information through a classification model. In this way, the staff can prioritize each subject information according to the demand intensity classification corresponding to different subject information, and handle the after-sales problems corresponding to each subject information according to the priority ranking. In this way, there is no need to manually screen the review data, saving human and material resources, and the staff can directly handle the after-sales problems corresponding to each subject information according to the demand intensity classification, and can promptly and quickly handle after-sales problems with stronger demand, thereby improving the user's driving experience of the vehicle.
[0131] The embodiment of the present disclosure provides a device for classifying comments, which may be the computer device in the above embodiment, such as Figure 3 As shown, the device includes:
[0132] The feature extraction module 320 is used to perform feature extraction processing on the acquired multiple comment data respectively to obtain a feature vector corresponding to each comment data;
[0133] A clustering module 330 is configured to perform topic clustering processing on the feature vectors corresponding to the plurality of comment data to obtain topic information corresponding to each comment data;
[0134] The statistics module 340 is used to perform frequency statistics on the topic data corresponding to the plurality of comment data to obtain the frequency of occurrence of each topic information;
[0135] The classification module 360 is used to determine the demand strength classification corresponding to each topic information based on the topic information corresponding to each review data, the frequency of occurrence of each topic information and the classification model, wherein the demand strength classification corresponding to the topic information is used to indicate the after-sales processing priority level of the topic information.
[0136] In a possible implementation, the feature extraction module 320 is configured to:
[0137] Determining a topic feature vector corresponding to each comment data based on the plurality of comment data and a topic feature extraction model;
[0138] Determining a semantic feature vector corresponding to each comment data based on the plurality of comment data and a semantic feature extraction model;
[0139] Based on the topic feature vector and the semantic feature vector corresponding to each comment data, a feature vector corresponding to each comment data is determined.
[0140] In a possible implementation, the feature extraction module 320 is configured to:
[0141] The topic feature vector and the semantic feature vector corresponding to each comment data are concatenated to obtain a feature vector corresponding to each comment data.
[0142] In a possible implementation, the apparatus further includes a determining module 350, configured to:
[0143] Determining, based on the plurality of comment data and the sentiment intensity model, a sentiment intensity value corresponding to each comment data, wherein the sentiment intensity value corresponding to the comment data is used to indicate the sentiment intensity degree and sentiment polarity of the comment data;
[0144] The classification module 360 is used to:
[0145] Based on the topic information corresponding to each comment data, the occurrence frequency of each topic information, the sentiment intensity value corresponding to each comment data and the classification model, the demand intensity classification corresponding to each topic information is determined.
[0146] In a possible implementation, the apparatus further includes a preprocessing module 310, configured to:
[0147] Get multiple initial review data;
[0148] Text preprocessing is performed on the multiple initial comment data to obtain the comment data.
[0149] In a possible implementation, the text preprocessing includes at least one of text deduplication processing, compression and word removal processing, short sentence deletion processing, pure symbol deletion processing, pure English deletion processing, text word segmentation processing, and stop word removal processing.
[0150] It should be noted that the review classification device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the review classification. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the review classification device provided in the above embodiment and the review classification method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0151] Figure 4 The following is a block diagram of a terminal 400 according to an exemplary embodiment of the present disclosure. The terminal may be the computer device described in the aforementioned embodiments. The terminal 400 may be a smartphone, a tablet computer, an MP3 player (moving picture experts group audio layer III), an MP4 player (moving picture experts group audio layer IV), a laptop computer, or a desktop computer. The terminal 400 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.
[0152] Typically, the terminal 400 includes a processor 401 and a memory 402 .
[0153] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (digital signal processing), FPGA (field-programmable gate array), or PLA (programmable logic array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (central processing unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (graphics processing unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (artificial intelligence) processor, which is used to process computing operations related to machine learning.
[0154] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the review classification method provided in the method embodiment of the present disclosure.
[0155] In some embodiments, terminal 400 may optionally include a peripheral device interface 403 and at least one peripheral device. Processor 401, memory 402, and peripheral device interface 403 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 403 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 404, a display screen 405, a camera 406, an audio circuit 407, a positioning component 408, and a power supply 409.
[0156] The peripheral device interface 403 can be used to connect at least one I / O (input / output)-related peripheral device to the processor 401 and the memory 402. In some embodiments, the processor 401, the memory 402, and the peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 401, the memory 402, and the peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0157] The radio frequency circuit 404 is used to receive and transmit RF (radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 404 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (wireless fidelity) networks. In some embodiments, the radio frequency circuit 404 may also include circuits related to NFC (near field communication), which is not limited in this disclosure.
[0158] Display screen 405 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. If display screen 405 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 405. These touch signals can be input as control signals to processor 401 for processing. Display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 405, located on the front panel of terminal 400. In other embodiments, there can be at least two display screens 405, located on different surfaces of terminal 400 or in a foldable design. In still other embodiments, display screen 405 can be a flexible display, located on a curved or foldable surface of terminal 400. Display screen 405 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 405 can be made of materials such as LCD (liquid crystal display) and OLED (organic light-emitting diode).
[0159] The camera assembly 406 is used to capture images or videos. Optionally, the camera assembly 406 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (virtual reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0160] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 401 for processing, or input into the radio frequency circuit 404 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each disposed at different locations on the terminal 400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.
[0161] The positioning component 408 is used to locate the current geographic location of the terminal 400 to implement navigation or LBS (location-based service). The positioning component 408 can be a positioning component based on the GPS (global positioning system), Beidou system, Greninja system or Galileo system.
[0162] Power supply 409 is used to power various components in terminal 400. Power supply 409 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 409 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0163] In some embodiments, the terminal 400 further includes one or more sensors 410 , including but not limited to: an acceleration sensor 411 , a gyroscope sensor 412 , a pressure sensor 413 , a fingerprint sensor 414 , an optical sensor 415 , and a proximity sensor 416 .
[0164] Accelerometer 411 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by terminal 400. For example, accelerometer 411 can be used to detect the components of gravity acceleration along the three coordinate axes. Processor 401 can control display screen 405 to display the user interface in either a landscape or portrait view based on the gravity acceleration signal collected by accelerometer 411. Accelerometer 411 can also be used to collect game or user motion data.
[0165] The gyroscope sensor 412 can detect the orientation and rotation angle of the terminal 400. It can also work with the accelerometer 411 to collect the user's 3D movements of the terminal 400. Based on the data collected by the gyroscope sensor 412, the processor 401 can implement the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0166] The pressure sensor 413 can be provided on the side frame of the terminal 400 and / or below the display screen 405. When the pressure sensor 413 is provided on the side frame of the terminal 400, it can detect the user's gripping signal of the terminal 400, and the processor 401 can perform left-hand or right-hand recognition or shortcut operations based on the gripping signal collected by the pressure sensor 413. When the pressure sensor 413 is provided below the display screen 405, the processor 401 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 405. Operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0167] The fingerprint sensor 414 is used to collect the user's fingerprint. The processor 401 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 414, or the fingerprint sensor 414 identifies the user's identity based on the collected fingerprint. When the user's identity is recognized as a trusted identity, the processor 401 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 414 can be set on the front, back, or side of the terminal 400. When a physical button or manufacturer logo is provided on the terminal 400, the fingerprint sensor 414 can be integrated with the physical button or manufacturer logo.
[0168] Optical sensor 415 is used to detect ambient light intensity. In one embodiment, processor 401 can control the display brightness of display screen 405 based on the ambient light intensity detected by optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of display screen 405 is increased; when the ambient light intensity is low, the display brightness of display screen 405 is decreased. In another embodiment, processor 401 can also dynamically adjust the shooting parameters of camera assembly 406 based on the ambient light intensity detected by optical sensor 415.
[0169] Proximity sensor 416, also known as a distance sensor, is typically located on the front panel of terminal 400. Proximity sensor 416 is used to detect the distance between the user and the front of terminal 400. In one embodiment, when proximity sensor 416 detects that the distance between the user and the front of terminal 400 is gradually decreasing, processor 401 controls display screen 405 to switch from the screen-on state to the screen-off state. When proximity sensor 416 detects that the distance between the user and the front of terminal 400 is gradually increasing, processor 401 controls display screen 405 to switch from the screen-off state to the screen-on state.
[0170] Those skilled in the art will understand that Figure 4 The structure shown in the figure does not constitute a limitation on the terminal 400, and the terminal 400 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0171] Figure 5 is a schematic diagram of the structure of a server provided in an embodiment of the present disclosure. The server 500 may vary significantly due to different configurations or performance, and may include one or more processors (central processing units, CPUs) 501 and one or more memories 502. The memories 502 store at least one instruction, which is loaded and executed by the processor 501 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which are not described in detail here.
[0172] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to implement the review classification method in the above embodiment. The computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be a ROM (read-only memory), RAM (random access memory), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0173] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0174] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the "comment data" and "initial comment data" involved in this disclosure are all obtained with full authorization.
[0175] The above description is merely an optional embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. A review classification method, characterized in that: The method comprises: Performing feature extraction on the acquired multiple comment data to obtain a feature vector corresponding to each comment data; Performing topic clustering processing on the feature vectors corresponding to the plurality of comment data to obtain topic information corresponding to each comment data; Performing frequency statistical processing on the topic data corresponding to the plurality of comment data to obtain the frequency of occurrence of each topic information; Based on the multiple review data, the topic information corresponding to each review data, the occurrence frequency of each topic information and the classification model, the demand strength classification corresponding to each topic information is determined, wherein the demand strength classification corresponding to the topic information is used to indicate the after-sales processing priority level of the topic information.
2. The method according to claim 1, characterized in that The feature extraction process is performed on each of the acquired comment data to obtain a feature vector corresponding to each comment data, including: Determining a topic feature vector corresponding to each comment data based on the plurality of comment data and a topic feature extraction model; Determining a semantic feature vector corresponding to each comment data based on the plurality of comment data and a semantic feature extraction model; Based on the topic feature vector and the semantic feature vector corresponding to each comment data, a feature vector corresponding to each comment data is determined.
3. The method according to claim 2, characterized in that The obtaining, based on the topic feature vector corresponding to each comment data and the semantic feature vector corresponding to each comment data, the feature vector corresponding to each comment data includes: The topic feature vector and the semantic feature vector corresponding to each comment data are concatenated to obtain a feature vector corresponding to each comment data.
4. The method according to claim 1, wherein The method further comprises: Determining, based on the plurality of comment data and the sentiment intensity model, a sentiment intensity value corresponding to each comment data, wherein the sentiment intensity value corresponding to the comment data is used to indicate the sentiment intensity degree and sentiment polarity of the comment data; The determining, based on the plurality of review data, the topic information corresponding to each review data, the occurrence frequency of each topic information, and the classification model, of the demand intensity classification corresponding to each topic information includes: Based on the topic information corresponding to each comment data, the occurrence frequency of each topic information, the sentiment intensity value corresponding to each comment data and the classification model, the demand intensity classification corresponding to each topic information is determined.
5. The method according to claim 1, wherein Before performing feature extraction processing on the acquired plurality of comment data to obtain a feature vector corresponding to each comment data, the method further includes: Get multiple initial review data; Text preprocessing is performed on the multiple initial comment data to obtain the comment data.
6. The method according to claim 5, characterized in that The text preprocessing includes at least one of text deduplication processing, compression and word removal processing, short sentence deletion processing, pure symbol deletion processing, pure English deletion processing, text word segmentation processing, and stop word removal processing.
7. A comment classification device, characterized in that: The device comprises: A feature extraction module is used to perform feature extraction processing on the acquired multiple comment data respectively to obtain a feature vector corresponding to each comment data; A clustering module, configured to perform topic clustering processing on the feature vectors corresponding to the plurality of comment data to obtain topic information corresponding to each comment data; A statistical module is used to perform frequency statistical processing on the topic data corresponding to the plurality of comment data to obtain the frequency of occurrence of each topic information; A classification module is used to determine the demand strength classification corresponding to each topic information based on the multiple review data, the topic information corresponding to each review data, the frequency of occurrence of each topic information and the classification model, wherein the demand strength classification corresponding to the topic information is used to indicate the after-sales processing priority level of the topic information.
8. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the operation performed by the comment classification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the operation performed by the comment classification method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the comment classification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Consumer demand analysis method and device based on multi-objective optimization algorithm
CN119313409A
GCN-based two-stage feature construction and link prediction method and device
CN119886943A
User demand comprehensive analysis method based on online user comment data
CN120146699A
Text classification and sentimentization with visualization
US20200257762A1