Data Aggregation Method Applied to the Digital Intelligence Management System of E-commerce Customer Service
Through the combination of AdaBoost algorithm and SVM algorithm, the word2vec model is used to extract comment text features, solving the problem of potential value identification in e-commerce comment data, and achieving efficient data aggregation and value mining.
Patent Information
- Application Number
- CN202410823655.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-06-25
AI Technical Summary
It is difficult for existing technology to effectively identify and mine the potential value in e-commerce review data, especially in the processing of massive review data in the era of big data.
AdaBoost algorithm combined with SVM algorithm is used to obtain multiple weak classifiers through iterative training, and the word2vec model is used to extract the feature vectors of the comment text, perform value classification, and identify valuable and worthless comment texts.
It improves the accuracy of identifying the value of e-commerce review texts, improves the efficiency of data aggregation, and can effectively explore the potential value in comment data.
Smart Images

Figure CN118673146B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data resource services, and in particular to a data aggregation method applied to an e-commerce customer service digital management system. Background Art
[0002] At present, the world is in the era of big data. The potential commercial and scientific value of data discovered through data mining, machine learning and other technologies has attracted widespread attention from all walks of life. In the era of big data, big data resource services are widely attached to the Internet.
[0003] With the rapid development of the Internet, e-commerce has also risen rapidly. In an era where almost everyone is shopping online, competition between major e-commerce companies and different merchants on the same platform has become increasingly fierce. In addition to being a carrier of information for feedback on product information and communication with stores, users' online reviews are more importantly an important reference for new buyers and an important reference for merchants to improve their services and products. Product reviews contain a lot of valuable information. On the one hand, consumers can understand the reputation of products through product reviews and make corresponding purchasing decisions; on the other hand, manufacturers can use reviews to discover problems with products and improve product quality.
[0004] In the era of big data, the review data of different online shopping platforms has formed a massive situation, and it has become impossible to collect and identify them manually. In the digital management system of e-commerce customer service, a scientific solution is urgently needed to assist users in data analysis and then explore the hidden value in the review data.
[0005] Based on this, Chinese patent CN103778245B discloses a method and device for identifying user comments, which includes: obtaining N target user comments, extracting the user ID of the target user comment, the number of characters contained in the target user comment, and the first M characters of the target user comment, where the user ID is a user identification code with a fixed number of digits and in numerical format, N>1, M>1; according to key=A / 10K+B+C, calculating N key values corresponding to the N target user comments, and recording the number of occurrences of each key value in the N key values, where A is the user ID of the target user comment, B is the number of characters contained in the target user comment, C is the encoding value of the first M characters of the target user comment in numerical format, K is a preset value, 0≤K<the number of digits of the user ID; judging whether the number of occurrences of each key value reaches the preset value, and determining the target user comment corresponding to the key value whose number of occurrences reaches the preset value as a variant repeated comment, the operation steps are simple, the amount of calculation is small, and the recognition efficiency of user comments is high.
[0006] However, the above disclosed method for identifying user comments still has the technical problem of being unable to tap the potential value of user comments. Specifically, the prior art mainly extracts the ID of the user who posted the target user comment, the number of characters contained in the target user comment, and the first M characters of the target user comment, obtains N key values corresponding to the N target comments according to the key value calculation formula, and determines the category of the target user comment according to the number of occurrences of each key value in the N key values. Although the application of the existing technical solution can determine the category of user comments, the valuable content in the user comments involving e-commerce service content, product technology improvements, etc. cannot be identified.
[0007] Specifically, the support vector machine algorithm, or SVM algorithm, is an algorithm that classifies data in a supervised learning manner. In addition to being applicable to linear classification problems, the algorithm can also be applied to nonlinear classification problems. The principle of the SVM algorithm is to obtain feature vectors from sample data and map these feature vectors to points in a high-dimensional space. Then, by solving the sample data, the super-flat interface or a dividing line with the largest distance between two different categories of data is found, and the data is divided into different types of data based on the obtained dividing interface or dividing line. After the data classification is completed, the super-flat interface or dividing line can still be used as a reference to complete the relevant classification of the newly added data points.
[0008] The word2vec model can process natural language and transform it into vectors that can be recognized by computers. Therefore, the core of the model is the process of transforming words or sentences into vectors. In this algorithm model, all words are represented in the form of vectors, so that the correlation analysis between words is converted into a measure of the relationship between their vectors, thereby analyzing and mining the relationship between words. This model is essentially a method based on word clustering, which can be applied to a variety of different scenarios such as semantic analysis of words and sentiment analysis of sentences.
[0009] The AdaBoost algorithm is a commonly used algorithm in the Boosting series of algorithms. In the Boosting series of algorithms, the algorithm used by each learner can be different or the same algorithm. If it is the same algorithm, the learner can be different from others by setting different parameters. Summary of the invention
[0010] Based on this, it is necessary to provide a data aggregation method applied to the digital management system of e-commerce customer service to address the technical problem of how to aggregate data on e-commerce platforms to identify the potential value of data.
[0011] A data aggregation method applied to an e-commerce customer service digital management system comprises the following steps:
[0012] S1: First, use the AdaBoost algorithm with the SVM algorithm as the basic classifier algorithm. Through iterative training, obtain the first classifier and the second classifier;
[0013] S2: Use the first classifier to process the review text data, that is, convert all letters in the review sample data to lowercase letters uniformly, and then use the Jieba word segmentation to segment the review text to obtain all word sets Wn of each review text;
[0014] S3: Use the first classifier to extract features from the review text data, that is, use the word2vec model to obtain the word vectors Vn corresponding to each word in all word sets Wn, and then sum and average the word vectors Vn of all words in the review text to obtain the text vector Vt. The calculation formula is as follows:
[0015]
[0016] S4: Use the first SVM classification algorithm in the first classifier to perform value classification processing on the obtained text vector Vt to obtain valuable review texts and valueless review texts respectively;
[0017] S5: Use the second classifier to process the valuable review texts and valueless review texts obtained in step S4 according to the preset weight values, and then divide the sample data into Chinese texts and English texts according to the language;
[0018] S6: For Chinese texts, first, segment the text data and simultaneously obtain the part-of-speech analysis corresponding to each word; then divide the part-of-speech into six categories, and count the number of times for each category of part-of-speech respectively; after the counting is completed, calculate the proportion of each category of words according to the total number of words contained in the six categories of part-of-speech of the text, form a vector as the text vector, and then use the second SVM classification algorithm for classification; finally, judge to obtain valuable review texts and valueless review texts;
[0019] S7: For English texts, according to the pre-organized word list covering a preset amount of English words, use space word segmentation to segment the English text. The text after word segmentation can be expressed as Wn. Then, compare each Wi with the pre-organized word list one by one, and obtain the proportion RW of the words in the English text in the word list; after the calculation of the proportion of English words is completed, calculate the proportion RC of letters in the English text, that is, the ratio of the number of letters in the text to the length of the text; after the calculation of the two proportions is completed, use the vector composed of RW and RC to represent the English text, and use the third SVM classification algorithm to classify the English text; finally, judge to obtain valuable review texts and valueless review texts.
[0020] Further, in the iterative training process of step S1, it specifically includes the following steps:
[0021] S11: First, initialize each sample in the training data to the same weight value; then, use the first classifier to train a weak classifier to obtain valuable comment samples and worthless comment samples.
[0022] S12: Next, according to the classification results of step S11, adjust the weight values of the samples, that is, reduce the sample weights of the valuable comment samples in the first classifier; at the same time, increase the sample weights of the worthless comment samples in the first classifier.
[0023] S13: Input the training data set obtained in step S12 into the second classifier for training; by increasing the weight values of the previously obtained worthless comment samples, make them the samples that the next classifier focuses on; through repeated learning and continuously corresponding to adjust the sample weights, finally obtain a strong classifier.
[0024] Further, the optimization method for the second SVM classification algorithm is: First, adopt the Gaussian kernel function as the kernel function in the SVM model, encode the SVM parameters by using the real number encoding method, calculate the fitness values of each part-of-speech category in the valuable comment text according to the fitness function, and then perform genetic algorithm operations on the sample data to generate the next generation of sub-sample data; after multiple iterations, the optimal part-of-speech category in the sample can be obtained; finally, decode the optimal part-of-speech category in the sample data to obtain the optimal SVM parameters.
[0025] Further, the optimization method for the second SVM classification algorithm is: In the initial stage of iteration, the particle swarm algorithm randomly generates a group of particles, that is, the feasible solutions of the model. The particles in the population update their own positions by tracking the individual extreme values and the global extreme values, and finally find the global optimal solution.
[0026] Further, the specific steps of the optimization method for the second SVM classification algorithm include:
[0027] S21: Initialize each part-of-speech category in the valuable comment text sample data and the corresponding parameters of the device.
[0028] S22: Calculate the fitness value of each part-of-speech category.
[0029] S23: Update the individual extreme value and the global extreme value of each part-of-speech category.
[0030] S24: Update the speed and position of each part-of-speech category.
[0031] S25: Determine whether the maximum number of iterations has been reached. If not, return to step S22; if so, enter the next step.
[0032] S26: The part-of-speech class with the optimal fitness is the parameter of the optimal second SVM classification algorithm.
[0033] In summary, the data aggregation method of the present invention applied to the digital intelligent management system of e-commerce customer service takes the AdaBoost algorithm as the core and the SVM algorithm as the basic classifier algorithm to complete the identification and classification of whether the review text is valuable. Among them, the purpose of the first classifier of the AdaBoost algorithm is to distinguish mathematical formulas or data in a specific format. In this classifier, first, the Jieba word segmentation is used to segment the text content of the review. Based on the text segmentation result; then, the word2vec model is used to obtain the text vector reflecting the characteristics of the review text; then, the SVM classification algorithm is used to obtain the judgment of whether the text is a mathematical formula or data in a specific format; then, the second classifier is used to classify the review text data that is not a mathematical formula or in a specific format in the first classifier. The second classifier makes full use of the characteristics of language text. First, the language of the review text is judged. If it is Chinese, the Chinese feature extraction method is used. If it is English, the English feature extraction method is used; finally, the SVM model is used to classify the data after feature extraction, so as to obtain the final judgment result of whether the review text is valuable. In the existing technology, although machine learning can identify the readability of text, the ability of a single classifier in machine learning is limited, and it often fails to make full use of the computing power of the computer. Therefore, the present invention uses a combination of multiple weak classifiers to identify and classify the value of review texts, so as to improve the classification effect of the overall text data. Among them, the SVM algorithm can be used in binary classification and multi-classification problems. This algorithm has good robustness and strong generalization ability for unknown data. Especially when the amount of data is not too large, it has a more superior performance compared with other classification learning algorithms. Therefore, the data aggregation method of the present invention applied to the digital intelligent management system of e-commerce customer service solves the technical problem of how to aggregate data on the e-commerce platform to identify the potential value of the data. Description of the Drawings
[0034] Figure 1 It is a flowchart of the data aggregation method of the present invention applied to the digital intelligent management system of e-commerce customer service. Detailed Embodiments
[0035] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0036] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0037] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0038] In the present invention, unless otherwise clearly specified and defined, the terms such as "mounted", "connected", "connected to", "fixed" and the like should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0039] In the present invention, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0040] It should be noted that when an element is referred to as "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are only for illustrative purposes and do not represent the only implementation.
[0041] Specifically, whether the comments in the e-commerce platform are valuable is a prerequisite for aggregating and sorting its data. System-generated comments, meaningless duplicate data, and some garbled characters have no use value. In order to identify valuable and valueless comments, it is necessary to determine the value of the pre-collected e-commerce review texts.
[0042] Generally, the statements of valuable comments are relatively smooth, the grammar is relatively correct, and the writing structure has a certain logic. For system-generated comments or garbled characters randomly entered by users, since the text is mostly specific content of meaningless repetition or garbled characters, the grammar and logical structure of the text do not conform to the inherent rules. Therefore, this type of comment has no value for collection and sorting. Therefore, the AdaBoost algorithm can be used to determine the value of the comments, with the SVM algorithm as the basic classifier.
[0043] According to the existing technology, a classifier classifies a set of sample data according to the sample characteristics. When determining the value of comments, the classifier can also be used to distinguish which samples are meaningful and which are meaningless. In the process of classifying the value of review texts, the text features are the main basis for distinguishing whether the data is valuable. The features of valuable review texts and valueless review texts are different. For example, it can be judged whether the text has value by judging whether there are garbled characters, mechanically repeated content, whether the text is composed of sentences, the part-of-speech tagging in the sentences, the length of the sentences, and the connection symbols of the text.
[0044] Please refer to Figure 1 , the data aggregation method of the present invention applied to the e-commerce customer service digital management system includes the following steps:
[0045] S1: First, use the AdaBoost algorithm with the SVM algorithm as the basic classifier algorithm to obtain the first classifier and the second classifier through iterative training;
[0046] S2: Process the review text data using the first classifier, that is, convert all letters in the review sample data to lowercase letters, and then perform word segmentation on the review text using Jieba word segmentation to obtain all word sets Wn of each review text;
[0047] S3: Extract features from the review text data using the first classifier, that is, use the word2vec model to obtain the word vectors Vn corresponding to each word in all word sets Wn, and then sum and average the word vectors Vn of all words in the review text to obtain the text vector Vt. The calculation formula is as follows:
[0048]
[0049] S4: Use the first SVM classification algorithm in the first classifier to perform value classification on the obtained text vector Vt to obtain valuable review texts and worthless review texts respectively;
[0050] S5: Use the second classifier to process the valuable review texts and worthless review texts obtained in step S4 according to the preset weight values, and then divide the sample data into Chinese texts and English texts according to the language;
[0051] S6: For Chinese texts, first, perform word segmentation on the text data and simultaneously obtain the part-of-speech analysis corresponding to each word; then divide the parts of speech into six categories and count the number of times for each category of part of speech respectively; after the counting is completed, calculate the proportion of each category of words according to the total number of words contained in the six categories of parts of speech of the text, form a vector as the text vector, and then use the second SVM classification algorithm for classification; finally, judge to obtain valuable review texts and worthless review texts;
[0052] S7: For English texts, according to the pre-organized word list covering a preset number of English words, use space word segmentation to perform word segmentation on the English text. The text after word segmentation can be expressed as Wn. Then, compare each Wi with the pre-organized word list one by one and obtain the proportion RW of the words in the English text in the word list; after the calculation of the proportion of English words is completed, calculate the proportion RC of letters in the English text, that is, the ratio of the number of letters in the text to the length of the text; after the calculation of the two proportions is completed, use the vector composed of RW and RC to represent the English text, and use the third SVM classification algorithm to classify the English text; finally, judge to obtain valuable review texts and worthless review texts.
[0053] Specifically, in the iterative training process of the foregoing step S1, it specifically includes the following steps:
[0054] S11: First, initialize each sample in the training data to have the same weight value; then, use the first classifier to train a weak classifier to obtain correctly classified samples and incorrectly classified samples;
[0055] S12: Next, according to the classification result of the previous step, the weight value of the sample is adjusted, that is, the weight of the sample correctly classified in the first classifier is reduced; at the same time, the weight of the sample incorrectly classified in the first classifier is increased;
[0056] S13: The training data set obtained in the above steps is then input into the second classifier for training; thus, those samples that are not correctly classified become the samples that are focused on in the next classifier by increasing the weights. After repeated learning and continuous adjustment of sample weights, a strong classifier is finally obtained.
[0057] Specifically, the present invention is applied to the value determination process of the data aggregation method of the digital management system of e-commerce customer service. The purpose of the first classifier of the AdaBoost algorithm is to distinguish mathematical formulas or data in a specific format. In this classifier, first, the text content of the comment is segmented using the Jieba word segmentation process, based on the text segmentation processing result; then, the word2vec model is used to obtain a text vector reflecting the characteristics of the comment text; then, the SVM classification algorithm is used to determine whether the text is a mathematical formula or data in a specific format; then, the non-mathematical formula and non-specific format comment text data in the first classifier are classified using the second classifier. The second classifier makes full use of the characteristics of the language text. First, the language of the comment text is judged. If it is Chinese, the Chinese feature extraction method is used, and if it is English, the English feature extraction method is used; finally, the SVM model is used to classify the data after feature extraction, so as to obtain the final judgment result of whether the comment text is valuable.
[0058] More specifically, in the aforementioned data aggregation method, the first classifier has a data processing part, a text feature extraction part, and a first SVM algorithm part. Among them, in order to improve the accuracy of classification and reduce the interference caused by inconsistent capitalization of English letters; during data processing, it is selected to uniformly change the letters in the sample data to lowercase letters, and then, the Jieba segmentation processing is used to segment the text to obtain all the word sets Wn of each text. Then, feature extraction is performed on the text in the text feature extraction part. The purpose is to find a vector so that it can reflect the features of the text as much as possible. To achieve this goal, first, based on the word set Wn, the word2vec model is used to obtain the word vector Vn corresponding to each word. The positional relationship of the word vectors can reflect the semantic correlation degree between words. Then, the word vectors of all the words in the text are added up and averaged to obtain the text vector Vt. The calculation formula is as shown in the aforementioned step S3. In this way, a text vector that not only retains the semantics of all the words in the sentence but also contains the comprehensive semantics of this text can be generated. This word vector can be used to represent the vector of this text. That is to say, the word2vec model replaces the complex process of feature extraction through summation and averaging. Therefore, even if the length of the text is not fixed, the word2vc model can still be used. After the review text undergoes the above feature extraction process, a text vector containing the comprehensive semantics of the review text is generated. After obtaining this text vector, the first SVM classification algorithm is used to perform value classification processing on the obtained text vector, and valuable text and valueless text can be obtained.
[0059] Furthermore, since there are some defects in the SVM algorithm, this algorithm is sensitive to missing samples and is prone to overfitting problems when the data samples are few. Therefore, it is necessary to improve the SVM algorithm applied in the aforementioned steps.
[0060] Specifically, the genetic algorithm is a random global search method that simulates the biological evolution process in nature. The genetic algorithm selects individuals with high fitness by simulating the law of survival of the fittest in nature, and moreover, the genetic operators can be used for crossover and mutation operations to generate new individuals. The genetic algorithm generates better individuals by continuously improving the fitness of individuals in the population during the iterative process and finally obtains the optimal individual. Due to the global searchability and parallelism of the genetic algorithm, the genetic algorithm can be combined with the support vector machine to optimize the SVM parameters: First, the Gaussian kernel function is used as the kernel function in the SVM model, and the SVM parameters are encoded by using real number coding. According to the fitness function, the fitness values of each part-of-speech category in the valuable review text are calculated, and then the genetic algorithm operation is performed on the sample data to generate the next-generation sub-sample data. After multiple iterations, the optimal part-of-speech category in the sample can be obtained. Finally, the optimal part-of-speech category in the sample data is decoded to obtain the optimal SVM parameters.
[0061] Furthermore, the particle swarm optimization algorithm is an optimization algorithm based on swarm intelligence. In the initial stage of iteration, the particle swarm optimization algorithm randomly generates a group of particles, namely the feasible solutions of the model. The particles in the population update their own positions by tracking the individual extreme value and the global extreme value, and finally find the global optimal solution. The SVM parameter optimization algorithm based on the particle swarm optimization algorithm is as follows:
[0062] S21: Initialize each part-of-speech class in the valuable review text sample data and the corresponding parameters of the device;
[0063] S22: Calculate the fitness value of each part-of-speech class;
[0064] S23: Update the individual extreme value and the global extreme value of each part-of-speech class;
[0065] S24: Update the velocity and position of each part-of-speech class;
[0066] S25: Determine whether the maximum number of iterations is reached. If not, return to step S22; if so, proceed to the next step;
[0067] S26: The part-of-speech class with the optimal fitness is the parameter of the optimal second SVM classification algorithm.
[0068] In summary, the data aggregation method of the present invention applied to the digital intelligent management system of e-commerce customer service takes the AdaBoost algorithm as the core and the SVM algorithm as the basic classifier algorithm to complete the identification and classification of whether the review text is valuable. Among them, the purpose of the first classifier of the AdaBoost algorithm is to distinguish mathematical formulas or data in a specific format. In this classifier, first, the Jieba tokenizer is used to tokenize the text content of the review. Based on the result of the text tokenization; then, the word2vec model is used to obtain the text vector reflecting the characteristics of the review text; then, the SVM classification algorithm is used to obtain the judgment on whether the text is a mathematical formula or data in a specific format; then, the second classifier is used to classify the review text data that is not a mathematical formula or in a specific format in the first classifier. The second classifier makes full use of the characteristics of language texts. First, the language of the review text is judged. If it is Chinese, the Chinese feature extraction method is used. If it is English, the English feature extraction method is used; finally, the SVM model is used to classify the data after feature extraction, so as to obtain the final judgment result on whether the review text is valuable. In the existing technology, although machine learning can identify the readability of texts, the ability of a single classifier in machine learning is limited, and it often fails to make full use of the computing power of the computer. Therefore, the present invention uses a combination of multiple weak classifiers to identify and classify the value of review texts, thereby improving the classification effect of the overall text data. Among them, the SVM algorithm can be used in binary classification and multi-classification problems. This algorithm has good robustness and strong generalization ability for unknown data. Especially when the amount of data is not too large, it has a more excellent performance compared with other classification learning algorithms. Therefore, the data aggregation method of the present invention applied to the digital intelligent management system of e-commerce customer service solves the technical problem of how to aggregate data on e-commerce platforms to identify the potential value of the data.
[0069] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0070] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A data aggregation method applied to an e-commerce customer service digital management system, characterized in that: It includes the following steps: S1: First, the AdaBoost algorithm with the SVM algorithm as the basic classifier algorithm is used to obtain the first classifier and the second classifier through iterative training; The iterative training includes the following steps: S11: First, initialize each sample in the training data to have the same weight value; then, use the first classifier to train a weak classifier to obtain valuable comment samples and worthless comment samples; S12: Next, according to the classification result of step S11, the weight value of the sample is adjusted, that is, the sample weight of the valuable comment sample in the first classifier is reduced; at the same time, the sample weight of the worthless comment sample in the first classifier is increased; S13: Input the training data set obtained in step S12 into the second classifier for training; increase the weight of the worthless review sample obtained above so that it becomes the sample that the next classifier focuses on; after repeated learning, the sample weight is adjusted repeatedly accordingly, and finally a strong classifier is obtained; S2: Use the first classifier to process the comment text data, that is, convert all letters in the comment sample data into lowercase letters, and then use the Jieba word segmentation process to segment the comment text to obtain the set of all words Wn for each comment text; S3: Use the first classifier to extract features from the comment text data, that is, use the word2vec model to obtain the word vector Vn corresponding to each word in the word set Wn, and then add the word vectors Vn of all words in the comment text and take the average to obtain the text vector Vt. The calculation formula is as follows: In the above formula, Vi refers to the word vector Vn corresponding to the i-th word in all word sets Wn; n refers to the sum of the number of word vectors Vn corresponding to all words in all word sets Wn; S4: using the first SVM classification algorithm in the first classifier to perform value classification processing on the obtained text vector Vt, and obtaining valuable comment text and worthless comment text respectively; S5: using the second classifier to process the valuable review text and the worthless review text obtained in step S4 according to a preset weight value, and then classifying the sample data into Chinese text and English text according to the language; S6: For Chinese text, first, the text data is segmented and the corresponding part of speech analysis of each word is obtained at the same time; then the parts of speech are divided into six categories, and the number of times of each category of parts of speech is counted; after the statistics are completed, the proportion of each category of words is calculated according to the total number of words contained in the six categories of parts of speech of the text, and the formed vector is used as the vector of the text, and then the second SVM classification algorithm is used for classification; finally, it is determined to obtain valuable comment text and worthless comment text; S7: For English text, according to a pre-organized vocabulary of English words covering a preset amount, the English text is segmented by space segmentation. The text after segmentation can be represented as Wn. Then, Wi is compared one by one to see if they are in the pre-organized vocabulary, and the proportion RW of the words in the English text in the vocabulary is obtained. After the English word proportion is calculated, the proportion RC of letters in the English text is calculated, that is, the ratio of the number of letters in the text to the length of the text. After the two proportions are calculated, the vector composed of RW and RC is used to represent the English text, and the third SVM classification algorithm is used to classify the English text. Finally, it is determined whether the valuable comment text and the worthless comment text are obtained.
2. The data aggregation method applied to the digital management system of e-commerce customer service according to claim 1 is characterized by: The optimization method for the second SVM classification algorithm is as follows: first, a Gaussian kernel function is used as the kernel function in the SVM model, the SVM parameters are encoded using real number coding, the fitness value of each type of part of speech in the valuable comment text is calculated according to the fitness function, and then the sample data is operated by a genetic algorithm to generate the next generation of sub-sample data; after multiple iterations, the optimal part of speech class in the sample can be obtained; finally, the optimal part of speech class in the sample data is decoded to obtain the optimal SVM parameters.
3. The data aggregation method applied to the digital management system of e-commerce customer service according to claim 1 is characterized by: The optimization method for the second SVM classification algorithm is as follows: at the beginning of the iteration, the particle swarm algorithm randomly generates a group of particles, namely, the feasible solution of the model. The particles in the population update their own positions by tracking individual extreme values and global extreme values, and finally find the global optimal solution.
4. The data aggregation method applied to the digital intelligent management system of e-commerce customer service according to claim 3 is characterized by: The specific steps of the optimization method for the second SVM classification algorithm include: S21: Initialize the parameters corresponding to each part of speech class in the valuable comment text sample data; S22: Calculate the fitness value of each part-of-speech class; S23: Update the individual extreme value and global extreme value of each part-of-speech class; S24: Update the speed and position of each part of speech class; S25: Determine whether the maximum number of iterations has been reached. If not, return to step S22; if yes, proceed to the next step; S26: The part-of-speech class with the best fitness is the optimal parameter of the second SVM classification algorithm.
Citation Information
Patent Citations
A method and device for identifying user comments
CN103778245B
Method and device for identifying user comments
CN103778245A
Multi-feature-fusion Chines-text classification method based on Attention neural network
CN108460089A