Network data mining method and device, storage medium and electronic equipment
By integrating multiple stages and algorithms, this method solves the problems of low efficiency and insufficient model generalization ability in traditional network data mining methods, and achieves efficient and accurate sentiment analysis and data value mining, adapting to the dynamic changes of network data.
Patent Information
- Application Number
- CN202511042961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional network data mining methods are inefficient and lack model generalization ability, making it difficult to cope with complex data and diverse network scenarios, and the value of the data is not fully explored.
Through multi-stage collaborative design and multi-algorithm fusion strategy, including data acquisition, preprocessing, feature engineering, text sentiment analysis model construction, model optimization and updating, and the fusion of TF-IDF feature extraction, K-means clustering, support vector machine and surface fitting, and Naive Bayes model, the entire process is optimized in a closed loop.
It improves feature quality and the accuracy of sentiment analysis, achieves dynamic adaptability to network data and synergistic efficiency across the entire process, ensures the timeliness and robustness of the model, and fully explores the value of the data.
Smart Images

Figure CN120910339A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network data mining, in particular to a network data mining method and device, a storage medium and an electronic device. BACKGROUND
[0002] At present, the scale of network data is growing exponentially, and various network platforms (such as e-commerce platforms, social media, online forums, etc.) have accumulated a large amount of user-generated content, which contains valuable information such as user preferences, emotional tendencies, and demand feedback. Traditional network data processing methods have obvious limitations: on the one hand, manual analysis of massive data is extremely inefficient and difficult to meet real-time and large-scale needs; on the other hand, a single algorithm model is easily affected by data noise and feature redundancy when dealing with complex data, resulting in a large deviation in the analysis results.
[0003] Specifically, in the field of text sentiment analysis, early methods rely on single feature extraction (such as simple word frequency statistics) and single classification models (such as only using support vector machines or naive Bayes), lacking deep mining of feature correlation and collaborative optimization between models. At the same time, the data preprocessing process lacks a dynamic adjustment mechanism, making it difficult to cope with dynamic changes in data distribution; the model training and application are disconnected, and there is no closed-loop optimization, resulting in insufficient model generalization ability and difficulty in adapting to diverse network data scenarios. In addition, the isolated application of different algorithms makes it difficult to fully exploit the value of data, and it is difficult to achieve full-process collaborative efficiency from data collection to result application. SUMMARY
[0004] The present application provides a network data mining method and device, a storage medium and an electronic device, which effectively breaks through the limitations of traditional network data mining through multi-link collaborative design and multi-algorithm fusion strategy.
[0005] To achieve the above purpose, the present application adopts the following technical solutions: The network data mining method comprises: S1: Determine the network data mining target and the data source range; clearly define the specific target and determine the data source range covering website page data, social media platform data, online forum data, and e-commerce platform data according to the target, and obtain the data through public API interface or network crawler technology conforming to robots protocol; S2: Data collection and preliminary storage; use network crawler tools or API interfaces to collect data within the range defined in S1, and store them in a relational database containing data unique identifier, data content, data source, and collection time fields; S3: Data preprocessing; clean, integrate, transform, and normalize the original data stored in the relational database; S4: Feature engineering; based on the pre-processed data in S3, extract TF-IDF features, construct new features such as evaluation period, use K-means clustering to obtain cluster labels as new features, evaluate feature importance through random forest and retain high importance features; S5: Text sentiment analysis model construction and training; construct a support vector machine model, combine high importance features evaluated by random forest to modify support vector machine output through surface fitting model, train the model for sentiment orientation classification; S6: User sentiment orientation analysis and result feedback; correlate the sentiment orientation classification results in S5 with the user ID and evaluation time after preprocessing in S3, and return to S3 for reprocessing if the results deviate greatly; S7: Report generation; according to the analysis results of S6, organize a network data mining report containing data sources, processing process, analysis results, conclusions and suggestions.
[0006] In this specification, the method of network data mining further includes S8: model optimization and updating; according to the report results of S7, increase the corresponding sentiment type training samples, adjust the order of surface fitting to retrain the model, and periodically collect new data to repeat S3 to S7 to realize dynamic updating.
[0007] In this specification, the method of network data mining further includes S9: multi-model fusion analysis; introduce Naive Bayes model, fuse its prediction probability and the probability distribution converted from the sentiment score modified by surface fitting in S5 to get the final sentiment orientation.
[0008] In this specification, the method of network data mining further includes S10: algorithm collaborative interaction verification; calculate the difference between the accuracy of the fusion model and the average accuracy of each single model, and if the difference is positive, it means that the fusion is effective.
[0009] In this specification, the method of network data mining further includes S11: collaborative application of clustering results and sentiment analysis; correlate the cluster labels in S4 with the sentiment analysis results in S5, and calculate the proportion of each cluster sentiment, and feed the features of the cluster with high negative proportion back to S4 to adjust feature selection.
[0010] In this specification, the method of network data mining further includes S12: feature dynamic updating mechanism; combine the results of S11 and the needs of S8, periodically re-execute the feature engineering step of S4, update the feature set and ensure compatibility with the original feature set.
[0011] In this specification, the method of network data mining further includes S13: dynamic adjustment of multi-model fusion weight, based on the performance score calculated from the accuracy, precision and recall of the fusion model in S9 on new data, if the score decreases by more than 5%, the fusion weight is re-determined.
[0012] The network data mining device applies the network data mining method in any one of the above, and the network data mining device comprises: A target and range defining module is configured to determine a network data mining target and define a data source range, the data source comprising website page data, social media platform data, network forum data and e-commerce platform data, and the data is obtained through a public API interface or a network crawler technology in accordance with a robots protocol; A data collection and storage module is configured to collect data within the range defined by the target and range defining module through a network crawler tool or an API interface, and store the original data in a relational database comprising a data unique identifier, data content, data source and collection time field; A data preprocessing module is configured to clean, integrate, convert and normalize the original data stored by the data collection and storage module; A feature engineering module is configured to extract TF-IDF features, construct evaluation period features, obtain cluster labels as new features through K-means clustering, and retain high importance features through random forest evaluation of feature importance based on the data processed by the data preprocessing module; A sentiment analysis model module is configured to construct a support vector machine model, combine high importance features obtained by the feature engineering module, correct the output of the support vector machine through a surface fitting model, and train the model to realize sentiment tendency classification; A sentiment analysis and feedback module is configured to correlate and analyze the sentiment tendency results obtained by the sentiment analysis model module with the user ID and evaluation time processed by the data preprocessing module, and feed back to the data preprocessing module for reprocessing if the results are significantly deviated; A report generation module is configured to generate a network data mining report comprising data sources, processing procedures, analysis results, conclusions and suggestions based on the analysis results of the sentiment analysis and feedback module.
[0013] A computer readable storage medium, the computer readable storage medium has a computer program stored thereon, the computer program is executed by a processor to implement the network data mining method in any one of the above.
[0014] An electronic device comprising a processor and a memory, the memory has a computer program stored thereon, and the processor executes the computer program to implement the network data mining method in any one of the above.
[0015] In summary, the present application has at least the following beneficial effects: Deep optimization of feature engineering: combining TF-IDF feature extraction, K-means clustering, and random forest feature importance evaluation, not only preserves the key information of text data, but also mines the internal structure of data through clustering, reduces redundancy through feature importance screening, and improves feature quality, laying a more reliable foundation for subsequent model analysis.
[0016] Accuracy improvement of sentiment analysis: the collaborative application of support vector machine and surface fitting, using surface fitting to correct the initial prediction results, making up for the linear limitations of single models; combined with the probability fusion of Naive Bayes model, further balancing the advantages of different models, significantly improving the accuracy of sentiment tendency classification, especially the recognition ability of neutral sentiment and other fuzzy categories is enhanced.
[0017] Dynamic adaptability of the whole process: through feature dynamic updating mechanism, model fusion weight adjustment, abnormal sentiment pattern detection and other links, the whole data mining process can respond to data distribution changes and application requirements in real time, ensuring the timeliness and robustness of the model, effectively dealing with the dynamics and complexity of network data.
[0018] Multi-link synergistic effect: the association analysis of clustering results and sentiment analysis, the feedback of abnormal pattern detection on data quality, and the whole process closed-loop optimization design, realize the linkage of data collection, preprocessing, feature engineering, model training, result application, make the data value be deeply mined in the circulation, provide more comprehensive and reliable basis for decision-making. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0020] Figure 1 The schematic diagram of the network data mining method involved in the present application.
[0021] Figure 2 The schematic diagram of the data preparation and preprocessing process involved in the present application.
[0022] Figure 3 The schematic diagram of the feature engineering process involved in the present application. DETAILED DESCRIPTION
[0023] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0024] The following disclosure provides many different implementations or examples for carrying out different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the embodiments of the present invention. Furthermore, reference numerals and / or reference letters may be repeated in different examples of the embodiments of the present invention; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various implementations and / or arrangements discussed.
[0025] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] like Figure 1 As shown, this embodiment provides a method for network data mining, including: S1: Define the target and scope of data sources for web data mining; clarify specific targets and, based on these targets, define the scope of data sources covering website page data, social media platform data, online forum data, and e-commerce platform data; acquire data through public API interfaces or web crawler technologies that comply with the robots.txt protocol. S2: Data Acquisition and Initial Storage; Use web crawlers or API interfaces to collect data within the scope defined in S1 and store it in a relational database containing fields for unique data identifiers, data content, data source, and acquisition time. S3: Data preprocessing; cleaning, integrating, transforming, and reducing raw data stored in relational databases; S4: Feature Engineering; Based on the data preprocessed in S3, TF-IDF features are extracted, new features such as evaluation time periods are constructed, K-means clustering is used to obtain cluster labels as new features, and random forest is used to evaluate feature importance and retain highly important features. S5: Text sentiment analysis model construction and training; construct a support vector machine model, combine the high importance features evaluated by random forest with the surface fitting model to correct the support vector machine output, and train the model for sentiment classification; S6: User sentiment analysis and result feedback; Correlate the sentiment classification results of S5 with the user ID and evaluation time after preprocessing in S3. If the result deviation is large, return to S3 for reprocessing. S7: report generation; based on the analysis result of S6, the network data mining report containing data source, processing procedure, analysis result, conclusion and suggestion is arranged.
[0027] In some implementations, the network data mining method further comprises S8: model optimization and update; according to the report result of S7, the model is retrained by adding the training sample of the corresponding sentiment type and adjusting the order of the surface fitting, and the dynamic update is realized by repeating S3 to S7 by regularly collecting new data.
[0028] In some implementations, the network data mining method further comprises S9: multi-model fusion analysis; the naive Bayes model is introduced, the prediction probability thereof is fused with the probability distribution converted from the sentiment score corrected by the surface fitting in S5, and the final sentiment tendency is obtained.
[0029] In some implementations, the network data mining method further comprises S10: algorithm collaborative interaction verification; the difference between the accuracy of the fusion model and the average value of the accuracy of each single model is calculated, and if the difference is positive, it means that the fusion is effective.
[0030] In some implementations, the network data mining method further comprises S11: collaborative application of clustering result and sentiment analysis; the cluster label of S4 is associated with the sentiment analysis result of S5, the proportion of each cluster sentiment is counted, and the features of the cluster with high negative proportion are fed back to S4 to adjust the feature selection.
[0031] In some implementations, the network data mining method further comprises S12: feature dynamic update mechanism; the feature engineering step of S4 is re-executed regularly in combination with the result of S11 and the requirement of S8, the feature set is updated and the compatibility with the original feature set is ensured.
[0032] In some implementations, the network data mining method further comprises S13: dynamic adjustment of multi-model fusion weight, the performance score is calculated based on the accuracy, precision and recall of the fusion model in S9 on new data, and if the score decreases by more than 5%, the fusion weight is re-determined.
[0033] In some implementations, the network data mining method further comprises S14: abnormal sentiment mode detection, based on the analysis result of S6, the difference distance between the latest evaluation feature vector and the historical evaluation feature vector of a user is calculated, and if it does not conform to the regular rule, it is marked as abnormal and manually reviewed.
[0034] In some implementations, the network data mining method further comprises S15: full-process collaborative optimization closed loop, each link of S1 to S14 is integrated, the data collection of S2 is optimized through the result of S14, the feature engineering of S4 is optimized through the result of S11, the model of S5 is adjusted and optimized through S13, and the performance of each link is regularly evaluated and optimized.
[0035] The technical concept of the application is as follows: S1: Define the target of network data mining and the scope of data sources Clearly define the specific target of network data mining, such as mining user behavior preferences, analyzing network public opinion trends, and identifying potential customers. According to the target, define the data source range, covering website page data, social media platform data (microblog, WeChat public number, Douyin, etc.), network forum data (Zhihu, Douban group, etc.), and e-commerce platform data (TaoBao, Jingdong commodity evaluation, etc.). These data are obtained through public Application Programming Interface (API); for websites without public interfaces, use Web Crawler technology, strictly follow the website robots protocol (Robots Exclusion Protocol) to obtain data legally.
[0036] Specific processing process: If the target is to mine the evaluation sentiment tendency of users of an e-commerce platform on a certain type of commodity, the data source is set to all user evaluation content of this type of commodity on the platform. Through the platform's public commodity evaluation API interface, the interface request parameters include commodity category number, page number, etc. After obtaining the returned JSON format data, parse out the user ID, evaluation content, evaluation time, score, etc. For the evaluation content not covered by the interface, use the Web Crawler program that meets the robots protocol to obtain, and also parse out the relevant fields. Store all the obtained data to ensure that each piece of data information is complete.
[0037] S2: Data collection and preliminary storage Use Web Crawler tools or API interfaces to collect network data in the data source range defined in S1. Store the collected raw data in the raw database, which uses a relational database (such as MySQL). The data table structure includes data unique identifier (ID), data content (Content), data source (Source), collection time (Collection_Time), etc.
[0038] Specific processing process: For the evaluation data of a certain type of commodity on an e-commerce platform determined in S1, use Python to write a Web Crawler program. When obtaining data through the API interface, send the request according to the parameter format required by the interface, parse the returned JSON format data, extract the user ID, evaluation content, etc. and store them in the "raw commodity evaluation table" of the MySQL database. For data obtained by the crawler, parse and store them in the same table to ensure that the data unique identifier is not repeated, i.e. the ID of each data in the table is unique.
[0039] S3: Data preprocessing The original data stored in the original database by S2 is pre-processed, including data cleaning, data integration, data conversion and data reduction.
[0040] Data cleaning: remove duplicate data, handle missing values, correct outliers.
[0041] Duplicate data removal: compare data unique identifiers (IDs), if there are data with the same ID, only keep the earliest one.
[0042] Missing value processing: data with missing content field is directly deleted; data with missing score field is filled with the average score of all evaluations of the product. Suppose the evaluation scores of a product are , where represents the score of the ith evaluation, n is the total number of evaluations, and the average score is calculated as: ; For example, a product has 5 evaluations with scores of 5, 4, 5, 3, and missing, n=4 (after removing missing values), then =(5+4+5+3) / 4=17 / 4=4.25, fill the missing score with 4.25.
[0043] Outlier processing: identify abnormal score values through boxplot method, set the lower quartile of score data as , the upper quartile as , the interquartile range as , and the abnormal value as less than or greater than , replace the abnormal score value with the average score .
[0044] Data integration: combine data from different sources but related to form a unified data set. For example, combine the evaluation data of the same product obtained through API interface and crawler to ensure consistent fields.
[0045] Data conversion: convert data into a form suitable for mining, such as converting text type evaluation content into computable numerical type data through Bag-of-Words model.
[0046] Data reduction: reduce data volume, such as retaining features related to evaluation sentiment tendency through feature selection.
[0047] Specific processing process: for the data of "original commodity evaluation table", first execute duplicate data removal, query the same ID data in the table and delete duplicate items. Then process the missing values, count the missing values of each field, and get the cleaned data after processing according to the above rules. When data integration, because the data comes from the same commodity, no additional merging operation is needed. In data conversion, use jieba segmentation tool to split Chinese evaluation content into words, build a bag-of-words model, and represent each evaluation content as a vector. Each dimension of the vector corresponds to a word, and the value is the number of times the word appears in the evaluation content. When data reduction, calculate the mutual information between words and sentiment tendency, and keep the top 1000 words with high mutual information as features.
[0048] In summary, the data preparation and preprocessing process is as shown in Figure 2 .
[0049] S4: Feature engineering Based on the data preprocessed in S3, feature engineering is performed, including feature extraction, feature construction, clustering processing and random forest feature importance evaluation.
[0050] 1. Feature extraction: valuable features are extracted from the preprocessed data. For text data, in addition to the word frequency features obtained by the bag-of-words model in S3, TF-IDF (Term Frequency-Inverse Document Frequency) features are also extracted. The calculation formula of TF-IDF is: ; Where, represents the frequency of word j in the i-th evaluation content, that is, , is the number of times word j appears in the i-th evaluation, is the total number of times all words appear in the i-th evaluation; represents the inverse document frequency of word j, , N is the total number of evaluation contents, is the number of evaluation contents containing word j.
[0051] 2. Feature construction: new features are constructed based on domain knowledge, such as constructing "evaluation period" features (morning, afternoon, evening) according to evaluation time.
[0052] 3. Clustering processing: K-means clustering algorithm is used to cluster the evaluation content data preprocessed in S3, and similar evaluation contents are classified into one class, which provides a basis for subsequent feature selection.
[0053] Model construction: The goal of K-means clustering algorithm is to partition m samples into k clusters, so that samples within a cluster have high similarity, and samples between clusters have low similarity. Let the cluster centers be where represents the center of the t-th cluster, and d is the feature dimension. The objective function is to minimize the sum of squared Euclidean distances from all samples to their assigned cluster centers: ; where represents the t-th cluster, represents the sample and the squared Euclidean distance between the cluster center .
[0054] Model training: Randomly select k samples as initial cluster centers; calculate the Euclidean distance of each sample to each cluster center, and assign the sample to the nearest cluster; recalculate the center of each cluster, which is the average of all samples in the cluster; repeat the above assignment and update steps until the cluster centers no longer change or the maximum number of iterations is reached.
[0055] Model application: Apply the trained clustering model to all evaluation content data to obtain the cluster label of each evaluation content, which is added as a new feature to the feature set.
[0056] 4. Random forest feature importance evaluation: Build a random forest model to evaluate the importance of the extracted features, and retain the features with higher importance.
[0057] Model construction: Random forest is composed of multiple decision trees , where T is the number of decision trees. For classification problems, the final prediction result is the majority vote of all decision tree prediction results.
[0058] Model training: Randomly sample with replacement T bootstrap sample sets from the training set, and each sample set is used to train a decision tree; when building each decision tree at each node, randomly select f features from all features, and select the optimal split feature based on the Gini index; the decision tree grows to the maximum depth without pruning.
[0059] Model application: Evaluate feature importance by calculating the average reduction of feature in all decision trees on node impurity (i.e., the reduction of Gini index), and the calculation formula is: ; where represents the set of nodes in the t-th decision tree that use feature j for splitting, represents the reduction of Gini index caused by using feature j for splitting node n. Retain the top p features in terms of importance.
[0060] Specific processing process: use S3 pretreatment evaluation content data, calculate the TF value and IDF value of each word in each evaluation, get TF-IDF feature vector. For example, "good" appears twice in the first evaluation, and the total number of evaluation words is 10, then =2 / 10=0.2; the total number of evaluations N=1000, and the evaluations containing "good" are 200, then , so . For the evaluation time field, judge the time period according to the time value and mark, such as 9:30 marked as "morning".
[0061] When clustering, set k=5, take TF-IDF feature vector as input, and randomly select 5 samples as cluster center. Calculate the Euclidean distance from each sample to 5 cluster centers, such as sample The distance from cluster center is , and is assigned to the cluster with the smallest distance. When updating the cluster center, , where is the number of samples in the tth cluster. After 100 iterations, the stable cluster center is obtained, and the cluster label of each evaluation content is taken as a new feature.
[0062] When training the random forest, set T=100, f= (d is the total number of features). For each bootstrap sample set, train a decision tree and calculate the importance of each feature. For example, feature j splits at node n of the tth tree, the Gini index of node n is Gini(n), and the Gini indexes of the left and right child nodes after splitting are Gini(left) and Gini(right), then: , where |n| is the number of samples in node n. The importance of feature j is the average value of all . Keep the top 80% of the features with the highest importance. In summary, the feature engineering process is shown in Figure 3 .
[0063] S5: Text sentiment analysis model construction and training Construct a text sentiment analysis model based on support vector machine (SVM, Support Vector Machine) and surface fitting, which is used for sentiment classification (positive, negative, neutral) of evaluation content.
[0064] Support vector machine model construction: the support vector machine model aims to find the optimal hyperplane to separate samples of different categories. For linearly separable cases, the optimal hyperplane satisfies: ; ; where, is the normal vector of hyperplane, d is the feature dimension; b is the bias term; is the feature vector of the i-th sample; is the class label of the i-th sample, represents positive sentiment, represents negative sentiment, represents neutral sentiment (linearly inseparable case is handled by introducing slack variable).
[0065] The objective function is to minimize subject to , where is the slack variable, used to handle linearly inseparable problem, introduce a penalty parameter C>0, the objective function becomes: ; where m is the number of training samples. By solving the Lagrange Multiplier Method, the dual problem is obtained: ; subject to , where is the Lagrange multiplier.
[0066] SVM model training: Take the TF-IDF feature vector obtained by S4 as the input feature, and the artificially labeled sentiment orientation (positive, negative, neutral) as the label, divide the training set and test set in the proportion of 7:3. Cross-validation method is used to select the optimal penalty parameter C and kernel function parameter (if nonlinear kernel function such as radial basis function RBF is used). Train the SVM model with the training set.
[0067] SVM model application: The trained model is used to predict the sentiment orientation of new evaluation content.
[0068] Specific processing process: The TF-IDF feature vector of S4 and the artificially labeled label form a data set, which is randomly shuffled and divided into training set and test set in the proportion of 7:3. Use the SVM module in the scikit-learn library, set the kernel function as the linear kernel function, and determine the penalty parameter C=1.0 through 5-fold cross-validation. When training, according to the dual problem, the optimal and b are solved. For example, for the training sample , the Lagrange multiplier is calculated, then (where j is the sample index satisfying 0< <C). When applied, the feature vector of the new evaluation content , calculate , the result greater than 0 is positive emotion, less than 0 is negative emotion, close to 0 is neutral emotion.
[0069] Surface fitting model construction: the output results of support vector machine are corrected by surface fitting model to improve the prediction accuracy. The output of support vector machine is (i.e. ), the input of surface fitting model is and the top 2 features with the highest importance obtained by random forest evaluation , , and the output is the corrected sentiment score . Quadratic surface fitting is adopted, and the model expression is:
[0070] ; where is the surface fitting coefficient.
[0071] Surface fitting model training: the true score corresponding to the manually labeled sentiment tendency (positive is 1, negative is -1, and neutral is 0) is used as the target value, and the least square method is used to estimate the fitting coefficient. Let the training sample be ( is the true score), and the loss function is: ; Take the partial derivative of L with respect to and set it to 0 to get a system of linear equations, and solve to get the coefficient estimate.
[0072] Surface fitting model application: input the output of support vector machine and features , into the surface fitting model to get the corrected score , and determine the sentiment tendency according to : > 0.3 is positive, <-0.3 is negative, otherwise it is neutral.
[0073] Specific processing process: the feature vector filtered by random forest in S4 and the manually labeled label are combined to form a data set, which is divided into training set and test set according to 7:3. Train the support vector machine model to get and b, calculate . Select the top two features in random forest, such as "quality" and "price" related TF-IDF features, as , .
[0074] Real score during curve fitting training According to the label setting. For each training sample, calculate , , and the quadratic term and cross term, construct the equation group. For example, for 3 samples, 3 equations can be obtained, and the coefficients are solved by matrix operation to . Assuming that the solution is =0.1, =0.8, =0.05, etc., the model is substituted to obtain . When applied, calculate , , ) for new samples, substitute the curve fitting model to obtain , and judge the emotional tendency.
[0075] S6: User sentiment tendency analysis and result feedback Correlate the sentiment tendency results predicted by the sentiment analysis model (support vector machine + curve fitting) in S5 with the user ID, evaluation time, etc. after pre-processing in S3, and analyze the sentiment tendency distribution of different user groups at different times. For example, calculate the proportion of positive evaluations of each user, the number of positive evaluations in different time periods, etc. The analysis results are fed back to the data preprocessing stage (S3). If it is found that the sentiment analysis results of a certain type of user's evaluation data are significantly biased, it may be that the feature extraction is not accurate enough during preprocessing, and S3 needs to be returned for reprocessing related to feature engineering.
[0076] Specific processing process: Add a "sentiment tendency" field to the pre-processed data set, and fill in the results predicted by S5. Group by user ID, and calculate the proportion of positive evaluations for each user , where is the number of positive evaluations of user u, is the total number of evaluations of user u. Group by evaluation time period, and calculate the number of positive evaluations, negative evaluations, and neutral evaluations in each period. If the proportion of positive evaluations of a certain user is significantly different from the manual judgment, for example, the model predicts that the user's overall satisfaction with the product is less than 50%, check the preprocessing of the user's evaluation content, including whether the segmentation is accurate (such as whether the positive words are mistakenly split into meaningless words), whether the clustering results are reasonable (such as whether the user's evaluation is divided into unrelated clusters), and whether the random forest feature importance evaluation misses key features (such as the user frequently mentioned specific positive words not being retained), return to S3 for reprocessing of data preprocessing and S4 feature engineering.
[0077] S7: Report generation According to the analysis results of S6, the analysis conclusion is arranged into a network data mining report, including data sources, processing process, analysis results, conclusions and suggestions, etc.
[0078] Specific processing process: detailed description of the specific implementation of each step in S1 to S6, including the specific platform and acquisition method of the data source, the specific operation of cleaning and conversion in data preprocessing, TF-IDF calculation in feature engineering, clustering process, key parameters and results of random forest feature importance evaluation, parameter selection of support vector machine in sentiment analysis model, and coefficient estimation value of surface fitting, etc. The analysis results part presents the sentiment tendency distribution data of different user groups (such as the proportion of users with more than 80% positive evaluation accounting for 30% of the total number of users), and the sentiment tendency data in different time periods (such as the proportion of positive evaluation in the evening 8-10 o'clock is higher than that in other time periods). The conclusion part summarizes the overall sentiment tendency of users to this type of goods, and the suggestion part proposes specific measures to improve the quality of goods and optimize logistics services for the user group with negative sentiment tendency combined with the high-frequency negative words (such as "poor quality" and "slow logistics") in their evaluation.
[0079] S8: Model optimization and update According to the report results of S7 and the actual application requirements, the sentiment analysis model in S5 is optimized. For example, if the neutral sentiment recognition accuracy after surface fitting is still low (such as less than 70%), the number of training samples of neutral sentiment can be increased, and the support vector machine model and the surface fitting model can be retrained; or the order of surface fitting can be adjusted (such as using cubic surface fitting), and the model expression is:
[0080]
[0081] ; where is the cubic surface fitting coefficient, which is re-estimated by least squares method. At the same time, new network data (S2) is collected regularly (such as every week), and the process of S3 to S7 is repeated to realize the dynamic update of the model.
[0082] Specific processing process: if the S7 report shows that the neutral sentiment recognition accuracy is 65%, then collect an additional 500 evaluation data labeled as neutral sentiment, combine it with the original training set, and re-divide the training set and test set. According to the process of S5, retrain the support vector machine model to get a new and b, and then re-estimate the cubic curve fitting coefficients three times with new training samples. After pre-processing the newly collected data, the updated model is used to predict the sentiment tendency. The prediction results are compared with the manually labeled results to calculate the accuracy, precision, recall and other indicators. If the indicators improve (e.g., the neutral sentiment accuracy improves to 75%), the updated model is adopted.
[0083] S9: Multi-model fusion analysis To further improve the accuracy of sentiment analysis, a naive Bayes model is introduced and fused with the support vector machine-cubic curve fitting model in S5.
[0084] Naive Bayes model construction: based on Bayes' theorem, it is assumed that the features are independent of each other. For sentiment classification problems, let the sentiment categories be , and the feature vector be , where is the value of the kth feature. The posterior probability is calculated as follows: ; Since is the same for all categories, only needs to be maximized during classification. Since the features are independent, .
[0085] Naive Bayes model training: the TF-IDF feature vector of S4 and the manually labeled labels are used to train the naive Bayes model to calculate the prior probability P(c) (the proportion of the number of samples of a certain sentiment category to the total number of samples) and the conditional probability (the probability of the occurrence of the kth feature under a certain sentiment category).
[0086] Model fusion: the prediction results of the support vector machine model and the prediction results of the naive Bayes model are weighted and fused. The calculation formula of the fused sentiment tendency is as follows: ; where and are the weights of the support vector machine model and the naive Bayes model, respectively, and + =1. Through testing on the validation set, it is determined that =0.6, =0.4; and are the probabilities of the two models predicting that the sample belongs to category c.
[0087] Specific processing process: the feature vector and the label of S4 are used to train the naive Bayes model to calculate the prior probability , The number of samples in the category c, m is the total number of samples, the conditional probability The prediction probability is obtained by maximum likelihood estimation. For a new sample, the prediction probability is obtained by two models respectively, and the probability of the largest category is taken as the final sentiment tendency according to the fusion formula.
[0088] In some embodiments, model fusion: the prediction probability of the naive Bayes model The sentiment score after surface fitting correction is fused. First, the is converted into a probability distribution , and the conversion formula is: ; ; ; Where is an indicator function, which is 1 when the condition in the parentheses is true, and 0 otherwise.
[0089] The probability distribution after fusion is: ; Where is the fusion weight, which is determined by the validation set (such as =0.3). The final sentiment tendency is the category with the largest probability after fusion.
[0090] Specific processing process: for a new sample , the prediction probability of naive Bayes is obtained respectively =0.6, =0.2, =0.2, the sentiment score after surface fitting correction =0.5. Then , =0, =0.378. After fusion , , , the final prediction is positive sentiment.
[0091] In some embodiments, the weighted fusion result of support vector machine (SVM) and naive Bayes (NB) can be further fused with the sentiment score after surface fitting correction. This multi-level fusion can integrate the advantages of different models (the classification boundary description ability of SVM, the probability estimation advantage of NB, and the nonlinear correction ability of surface fitting), and further improve the accuracy of sentiment tendency prediction. The specific fusion logic, implementation steps and formulas are as follows: Core idea 1. Output form uniformity: the weighted fusion result of SVM and NB is the probability distribution of sentiment categories (such as the probabilities of positive, negative, and neutral); the sentiment score after surface fitting correction is a continuous value (0.3 ), which needs to be converted into a probability distribution to ensure consistency in the output form.
[0092] 2. Secondary weighted fusion: the probability distribution of the previous step and the probability distribution converted after surface fitting are weighted and fused to obtain the final sentiment tendency probability by introducing a new fusion weight.
[0093] 3. Dynamic optimization of weight: the fusion weight is dynamically adjusted according to the performance (such as accuracy and F1 score) of each model on the validation set to ensure that the superior model obtains a higher weight.
[0094] Specific implementation steps and formulas Step 1: Uniform output form - convert surface fitting score to probability distribution The sentiment score after surface fitting correction is a continuous value (such as >0.3 is positive, <-0.3 is negative, and otherwise is neutral), which needs to be converted into a probability distribution (c represents the sentiment category). The conversion formula is based on the sigmoid function (to simulate the monotonicity of the probability distribution): ; ; ; where: k is the scaling factor (such as k=2 to enhance the impact of the score on the probability); is the indicator function (1 if the condition is true, otherwise 0); is the minimum value (such as =10 -5 to avoid the probability being 0, which would cause an abnormal calculation in the subsequent calculation).
[0095] Step 2: Initial fusion probability of SVM and NB Let the weighted fusion probability of SVM and NB be (which has been defined in S9): ; where is the initial fusion weight (such as =0.3), , are the prediction probabilities of NB and SVM, respectively.
[0096] Step 3: Secondary fusion - fusion with curved surface fitting probability Introducing new fusion weights (0 <1), fuse with and to get the final probability : ; The final sentiment tendency is the class with the highest probability: ; Step 4: Determination of fusion weights Weights are optimized through the validation set, aiming to maximize the F1 score (a combination of precision and recall) of the final fusion model: ; Where , the optimal value is solved by grid search (e.g. =0.1,0.2,...,0.9).
[0097] Interaction correlation and fusion logic 1. Information transmission chain: SVM output → curved surface fitting correction to → converted to ; At the same time, SVM and NB fusion into ; Finally, and are fused through to .
[0098] (The output of the previous model is directly used as the input of the subsequent fusion, forming a closed-loop interaction).
[0099] 2. Advantage complementation: SVM is good at high-dimensional feature classification, and NB is good at probability estimation. The fusion of the two enhances the stability of basic classification; Curved surface fitting captures the nonlinear relationship between SVM output and key features, correcting system errors; Secondary fusion further integrates the information of the two paths, reducing the bias of a single model.
[0100] Example Take the evaluation "filter core is expensive and replacement is troublesome, but CADR value meets the standard" as an example: 1. SVM output s=-0.2 (close to neutral), NB prediction =0.6, initial fusion =0.55 =0.3
[0101] 2. Curve fitting with s=-0.2 and "filter" "CADR value" feature correction =0.1, converted to =0.7.
[0102] 3. Validation set optimization =0.4, finally =0.4 x 0.55 + 0.6 x 0.7 = 0.64, predicted as neutral (consistent with manual annotation).
[0103] In summary, this multi-layer fusion can effectively integrate multi-model information through structured interaction formula and weight optimization, significantly improving the accuracy and robustness of sentiment prediction.
[0104] S10: Algorithm coordination interaction verification By calculating the coordination interaction index between each algorithm, the effectiveness of algorithm fusion is verified. Define the coordination gain G as the difference between the accuracy of the fusion model and the average accuracy of each single model: ; Where is the accuracy of the fusion model, , , are the accuracies of the support vector machine, naive Bayes, and curve fitting correction model respectively. If G>0, it means that the algorithm fusion is effective, and the fusion model can be used; otherwise, the fusion weight needs to be adjusted or other fusion methods need to be selected.
[0105] Specific processing process: calculate the accuracy of each model on the test set, assuming =0.85, =0.78, =0.75, =0.80, then: , indicating that the algorithm fusion is effective.
[0106] S11: Cooperative application of clustering results and sentiment analysis The cluster labels obtained from K-means clustering in S4 are correlated with the sentiment analysis results in S5 to explore the distribution patterns of sentiment tendencies within different clusters. For example, the proportion of positive, negative, and neutral sentiment evaluations in each cluster is statistically analyzed. If the proportion of negative sentiment in a certain cluster is significantly higher than that in other clusters (e.g., exceeding 60%), the common characteristics of the evaluation content in that cluster (such as frequently occurring negative words, similar evaluation periods, etc.) are analyzed. These characteristics are then fed back into the random forest feature importance evaluation stage in S4 as a basis for adjusting feature selection.
[0107] Specific processing steps: For each cluster Calculate the number of positive evaluations. Number of negative evaluations Neutral rating number Total number of evaluations The proportion of each emotion is then... , , If cluster of =0.65, significantly higher than the average of other clusters. If the value is 0.2, the top 10 most frequent words in the evaluation content of this cluster are extracted. It is found that negative words such as "damaged" and "difficult to return or exchange" appear frequently. The features corresponding to these words are given higher weights in the random forest feature importance evaluation, and the features are re-selected.
[0108] S12: Feature Dynamic Update Mechanism Based on the collaborative analysis results of S11 and the model optimization requirements of S8, a dynamic feature update mechanism is established. Periodically (e.g., every two weeks), based on newly collected data and cluster-sentiment collaborative analysis results, the feature engineering steps of S4 (including TF-IDF calculation, clustering, and random forest feature importance assessment) are re-executed to update the feature set. The updated feature set must be compatible with the original feature set; that is, the new feature set must contain the top 50% of the most important features from the original feature set to ensure the continuity of model application.
[0109] Specific processing steps: After preprocessing the newly collected data, the TF-IDF features are recalculated, K-means clustering is performed (using the previous k values) to obtain new cluster labels, and the importance of the new features is evaluated using random forest. Let the top 50% of the features in the original feature set be... The new features are those that rank in the top 80% by importance. The final updated feature set is For example, if the original feature set has 1000 features, For the first 500, after the new feature set is evaluated For the first 800, The union of the original top 500 features and the 800 new features (about 1000 after removing duplicate features) is used for subsequent model training and application.
[0110] S13: Dynamic adjustment of multi-model fusion weights Based on the prediction performance (such as accuracy, precision) of the fusion model on new data in S9, the fusion weights are dynamically adjusted Set performance evaluation indicators Where Acc is the accuracy, Prec is the precision, and Rec is the recall. If the Score decreases by more than 5% from the previous period, the optimal fusion weights are determined again through grid search (The search range is 0.1-0.9, with a step size of 0.1), so that the new Score reaches a maximum.
[0111] Specific processing process: Assuming that the Score of the previous period is 0.85 and the current period is 0.78, the decrease is about 8.2%, and the adjustment is needed. The Scores of the fusion models are calculated respectively, and if =0.4, the Score=0.83 is the maximum, then the fusion weights are updated to 0.4, and the new fusion probability formula is .
[0112] S14: Abnormal sentiment pattern detection Based on the results of user sentiment analysis in S6, an abnormal sentiment pattern detection model is built to identify sentiment data that does not conform to the regular rules. For example, a user's historical evaluations are all positive (positive evaluations account for more than 90%), but the latest evaluation is predicted to be negative, and the feature vector of this evaluation is significantly different from the average feature vector of the user's historical evaluations (Euclidean distance exceeds the set threshold, such as threshold =2.5), which is marked as an abnormal sentiment pattern and needs to be manually reviewed for evaluation content to determine whether it is a data collection error or a malicious evaluation.
[0113] Specific processing process: For user u, calculate the average value of its historical evaluation feature vector (excluding the latest evaluation), and the latest evaluation feature vector is , then the difference distance D= . If the user u's historical positive evaluation ratio >0.9, the latest evaluation sentiment is negative, and D> =2.5, it is marked as abnormal. After manual review finds that the evaluation is a malicious screen content, it is removed from the data set, and the user's sentiment tendency statistics are updated.
[0114] S15: Full-process collaborative optimization closed loop Integrate all stages from S1 to S14 to form a closed-loop collaborative optimization process. Optimize data acquisition quality through abnormal sentiment pattern detection results in S14 (S2), optimize feature engineering through clustering-sentiment co-analysis results in S11 (S4), optimize the sentiment analysis model through fusion weight adjustment in S13 (S5), and ensure model timeliness through the model update mechanism in S8. Regularly (e.g., monthly) evaluate the performance of each stage of the closed loop, calculating the deviation between the output results of each stage and the expected goals (e.g., deviation in data acquisition completeness). ,like If the value is greater than 0.1, the data collection strategy will be optimized to continuously improve the overall effect of network data mining.
[0115] Specific processing steps: Monthly statistical data collection completeness; if the target is to collect 10,000 data entries, 8,500 entries were actually collected. =1-0.85=0.15>0.1, then adjust the crawling frequency of the web crawler (e.g., change from once per hour to once every 30 minutes) or increase the concurrency of API interface requests; evaluate the accuracy of the sentiment analysis model, and if it decreases by 3% compared to the previous month, update the model according to the S8 process; through the collaborative optimization of each link, stabilize the accuracy of the data mining results at over 85%, and achieve an abnormal data detection rate of over 90%.
[0116] Taking the analysis of user reviews for "home air purifiers" on a certain e-commerce platform as an example, the specific application of the above-mentioned network data mining scheme involves the following steps: S1-S3: Data Preparation and Preprocessing The clear objective was to uncover users' emotional preferences and core evaluation dimensions regarding air purifiers. Nearly 100,000 review data points from the past six months were collected via platform API and compliant web scraping and stored in a MySQL database. In the preprocessing stage, 320 duplicate data points were removed, and for the 850 missing rating data points, the average rating for the category was used. Fill in the blanks; identify 120 outlier ratings (e.g., ratings of 1 point but all positive comments) using box plots, and then use the same method... Correction: The evaluation content is segmented using jieba, and after filtering out stop words, a bag-of-words model is constructed, retaining the features of the top 1000 words with the highest mutual information.
[0117] S4: Feature Engineering (including Clustering and Random Forest) TF-IDF calculation: Taking the review "This air purifier has low noise, long battery life, and is worth buying" as an example, "low noise" appears once, and the total number of words in the review is 8, therefore... =1 / 8=0.125; If there are 2000 reviews including "low noise" and a total of 100,000 reviews, then Therefore =0.125×3.91≈0.489.
[0118] K-means clustering: Set k=5, input TF-IDF feature vector, converge after 50 iterations. Cluster 1 concentrates on words like "silent" "noise" (noise-related), cluster 2 frequently appears "battery life" "power consumption" (battery life-related), use cluster labels as new features.
[0119] Random Forest Feature Importance: Build 100 decision trees and calculate feature importance. "Filter life" "CADR value" "price" rank top three, keep top 800 features.
[0120] S5: Sentiment Analysis (take SVM + surface fitting as an example, actual subsequent S9 can use ) SVM Model: Train with 70,000 data, penalty parameter C=1.0, linear kernel function. For the evaluation "filter replacement frequency, not worth the price", the feature vector is calculated (Preliminary prediction of negative).
[0121] Surface Fitting Correction: Take top 2 features of random forest (filter life TF-IDF), (price TF-IDF), combine SVM output s=-1.2, substitute into quadratic surface model: ; Calculate =-1.5, confirm negative sentiment.
[0122] S6-S7: Analysis and Report After associating user IDs, it is found that users aged 30-40 are sensitive to "silent" demand (72% of positive evaluations), while users over 50 are more concerned about "operation convenience" (65% of negative evaluations mention "complicated buttons"). The report suggests optimizing product detail page display for different age groups.
[0123] S8-S15: Optimization and Coordination Multi-model fusion: Naive Bayes and surface fitting model fusion, weight =0.3, negative sentiment recognition accuracy from 82% to 89%.
[0124] Anomaly detection: A user's last 50 evaluations are all positive, the latest evaluation is predicted to be negative, and the Euclidean distance D=3.2 =2.5 between its feature vector and historical mean is calculated. Artificial review confirms it as a malicious evaluation and removes it.
[0125] Dynamic update: 20000 new data every month, retrain the model, keep the core features such as "filter life" to ensure the continuity of analysis.
[0126] Through the whole process of cooperation, the final realization of user sentiment recognition accuracy is more than 90%, providing data support for platform commodity operation.
[0127] Each step and algorithm plays an irreplaceable core role in network data mining scheme, and the specific contribution is as follows: S1 (determine the target and data source) is the starting point of the scheme. By defining the mining target (such as analyzing the sentiment tendency of air purifier users) and defining the data range (e-commerce platform evaluation data), it provides direction for all subsequent operations, ensures that data collection and analysis do not deviate from the core demand, and avoids wasting resources on irrelevant data.
[0128] S2 (data collection and storage) is the "source" guarantee of data. Through API interface and compliant crawler, the original data is obtained and stored in the database, providing complete and traceable data source for subsequent processing, ensuring the accessibility and integrity of data, and is the material basis for the implementation of the whole scheme.
[0129] S3 (data preprocessing) solves the "messy" problem of raw data by cleaning duplicate data, handling missing values and outliers, and improves data quality. For example, fill in the missing score, correct the abnormal score, so that the data is more in line with the analysis standard; data conversion and reduction will convert text data into a computable form, reduce data volume, and clear obstacles for subsequent feature engineering and model training, ensuring that the data input to the subsequent steps is reliable and effective.
[0130] S4 (feature engineering) is the key bridge connecting data and model. TF-IDF extracts important information from text, quantifying the importance of words in evaluation; K-means clustering classifies similar evaluations into one category, mining the potential structure of evaluation content, and cluster labels enrich the data dimension as new features; Random forest feature importance evaluation selects the most discriminant features, reduces redundant information, reduces model complexity, improves the training efficiency and prediction accuracy of subsequent models, and ensures that the features input to the model are the most valuable.
[0131] S5 (sentiment analysis model) is the "core analyzer" of the scheme. Support vector machine realizes preliminary sentiment classification by finding the optimal hyperplane, providing a basis for sentiment tendency judgment; surface fitting combines important features to modify the SVM results, making up for the shortcomings of single model, improving the accuracy of sentiment prediction, and making the sentiment tendency judgment more accurate.
[0132] S6 (user sentiment analysis and feedback) associates sentiment analysis results with user attributes, mines the sentiment patterns of different user groups, and discovers problems in preprocessing and feature engineering through a feedback mechanism, forming a "analysis-feedback-optimization" cycle to continuously improve the quality of previous steps.
[0133] S7 (report generation) systematically presents the entire mining process and results, providing clear and actionable conclusions and recommendations for decision-makers, realizing the practical application value of data mining, and allowing the analysis results to directly guide business practice (such as product optimization on e-commerce platforms).
[0134] S8 (model optimization and updating) addresses model performance issues (such as low neutral sentiment recognition rate) by increasing samples, adjusting model structure (such as changing the order of surface fitting), and regularly updating the model with new data to ensure its timeliness and adaptability, allowing the model to continuously respond to new data and scenarios.
[0135] S9 (multi-model fusion) combines the advantages of Naive Bayes and Support Vector Machine-Surface Fitting models to obtain better prediction results, further improving the accuracy of sentiment recognition and reducing the limitations of single models.
[0136] S10 (algorithm coordination and interactive verification) verifies the effectiveness of model fusion by calculating the synergy gain to ensure that the performance of the fused model is better than that of a single model, providing a scientific basis for model selection and avoiding ineffective fusion operations.
[0137] S11 (clustering and sentiment analysis coordination) associates clustering results with sentiment orientation to discover the sentiment characteristics of different clusters, providing a basis for feature selection and allowing feature engineering to more effectively retain important features related to sentiment.
[0138] S12 (dynamic feature update) regularly updates the feature set with new data and collaborative analysis results while ensuring compatibility between new and old features, introducing new information while maintaining the continuity of model application, allowing features to adapt to changes in data.
[0139] S13 (dynamic adjustment of fusion weights) adjusts the fusion weights based on the performance of the model on new data to ensure that the fused model always maintains optimal state and improves the robustness of the model.
[0140] S14 (abnormal sentiment pattern detection) identifies sentiment data that does not conform to the norm (such as malicious reviews) to ensure data quality and avoid interference from abnormal data, improving the reliability of the analysis.
[0141] S15 (full-process collaborative optimization closed loop) integrates each link, through the interaction and optimization of each step (such as abnormal detection optimizing data collection, clustering analysis optimizing feature engineering), forms a continuously improving closed loop, ensures the overall performance of the whole scheme continuously improving, and realizes long-term, stable network data mining effect.
[0142] The network data mining device applies the network data mining method in any one of the above embodiments, and the network data mining device comprises: A target and range defining module is configured to determine a network data mining target and define a data source range, wherein the data source comprises website page data, social media platform data, network forum data, and e-commerce platform data, and the data is obtained through a public API interface or a network crawler technology conforming to a robots protocol; A data collection and storage module is configured to collect data within the range defined by the target and range defining module through a network crawler tool or an API interface, and store the raw data in a relational database comprising a data unique identifier, data content, data source, and collection time field; A data preprocessing module is configured to clean, integrate, convert, and normalize the raw data stored by the data collection and storage module; A feature engineering module is configured to extract TF-IDF features, construct evaluation period features, obtain cluster labels as new features by K-means clustering, and retain high importance features by evaluating feature importance through a random forest based on the data processed by the data preprocessing module; A sentiment analysis model module is configured to construct a support vector machine model, correct the output of the support vector machine through a surface fitting model in combination with the high importance features obtained by the feature engineering module, and train the model to realize sentiment tendency classification; A sentiment analysis and feedback module is configured to correlate and analyze the sentiment tendency results obtained by the sentiment analysis model module with the user ID and evaluation time processed by the data preprocessing module, and feed back to the data preprocessing module for reprocessing if the results deviate greatly; A report generation module is configured to generate a network data mining report comprising data sources, processing procedures, analysis results, conclusions, and suggestions based on the analysis results of the sentiment analysis and feedback module.
[0143] A computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the network data mining method in any one of the above embodiments.
[0144] An electronic device comprising a processor and a memory, wherein the memory has a computer program stored thereon, and the processor implements the network data mining method in any one of the above embodiments when executing the computer program.
[0145] The above-described embodiments are intended to illustrate, not to limit, the present application, and thus variations in the example values or substitution of equivalent elements are still within the scope of the present application.
[0146] From the above detailed description, it can be seen that the present application can achieve the aforementioned objects, and thus meets the requirements of the Patent Law.
[0147] Although the preferred embodiments of the present application have been described, those skilled in the art who have the benefit of the basic inventive concept can make additional changes and modifications to the embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application. The above description is merely the preferred embodiments of the present application, and is not intended to limit the present application. It should be noted that any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
[0148] It should be noted that the above description of the flow is merely for example and illustration, and does not limit the scope of the present application. Those skilled in the art can make various modifications and changes to the flow under the guidance of the present application. However, these modifications and changes are still within the scope of the present application.
[0149] The above has described the basic concept, and it is obvious that the above-mentioned disclosure of the present application is only as an example and does not constitute a limitation to the present application for those skilled in the art after reading this application. Although it is not explicitly stated here, those skilled in the art can make various modifications, improvements and modifications to the present application. Such modifications, improvements and modifications are suggested in the present application, so such modifications, improvements and modifications are still within the spirit and scope of the exemplary embodiments of the present application.
[0150] Meanwhile, the present application uses specific words to describe the embodiments of the present application. For example, "one embodiment", "an embodiment", and / or "some embodiments" means a certain feature, structure or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned in different positions in the specification does not necessarily refer to the same embodiment. In addition, some features, structures or characteristics in one or more embodiments of the present application can be properly combined.
[0151] Computer program code for carrying out operations of various aspects of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python and the like, conventional procedural programming languages, such as the C programming language, Visual Basic, Fortran 2103, Perl, COBOL 2102, PHP, ABAP, dynamic programming languages, such as Python, Ruby and Groovy, and / or other programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any form of network, such as a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider), or cloud computing environment, or a service such as Software as a Service (SaaS).
[0152] Accordingly, it should be noted that, in the description of embodiments of the present application, references have been made to acts and features which can occur in any order and / or concurrently. This is because embodiments of the application can have multiple features and / or acts and some of these can occur before, after or concurrently with other features and / or acts of the application. Similarly, it should be noted that, in the interests of simplifying the disclosure of the present application and helping with the understanding of one or more embodiments of the application, in the foregoing description of embodiments of the present application, various features can have been grouped together or described in a single embodiment, figure or description of an embodiment. However, the mere fact that different features or embodiments of the application can have been grouped together in the foregoing should not be interpreted to mean that the application requires more features than are explicitly recited in each claim. Rather, inventive subject matter can lie in fewer than all features of a single foregoing disclosed embodiment.
Claims
1. A method of network data mining, characterized by, Comprise: S1: Determine the network data mining target and data source range definition; clear specific target and define the data source range of website page data, social media platform data, network forum data, e-commerce platform data according to the target, obtain the data through public API interface or network crawler technology conforming to robots protocol; S2: Data collection and preliminary storage; use network crawler tool or API interface to collect data within the range defined in S1, and store it in a relational database containing data unique identifier, data content, data source, collection time field; S3: Data preprocessing; clean, integrate, convert and reduce the original data stored in the relational database; S4: Feature engineering; Based on the preprocessed data in S3, extract TF-IDF features, construct evaluation period and other new features, use K-means clustering to get cluster label as new feature, evaluate feature importance by random forest and retain high importance features; S5: Text sentiment analysis model construction and training; build support vector machine model, combine high importance features evaluated by random forest to modify support vector machine output through surface fitting model, train model for sentiment classification; S6: User sentiment analysis and result feedback; Correlate the sentiment classification results in S5 with the user ID, evaluation time after S3 preprocessing, and return S3 for reprocessing if the result deviation is large; S7: Report generation; According to the analysis results of S6, sort the network data mining report containing data source, processing process, analysis results, conclusion and suggestion.
2. The method of network data mining according to claim 1, characterized in that, Also includes S8: Model optimization and update; according to the report results of S7, increase the corresponding sentiment type training sample, adjust the order of surface fitting to retrain the model, and periodically collect new data to repeat S3 to S7 to realize dynamic update.
3. The method of network data mining according to claim 2, wherein, Also includes S9: Multi-model fusion analysis; introduce Naive Bayes model, fuse its prediction probability with the probability distribution converted from the sentiment score modified by surface fitting in S5 to get the final sentiment tendency.
4. The method of network data mining according to claim 3, wherein, Also includes S10: Algorithm collaborative interaction verification; calculate the difference between the accuracy of the fusion model and the average accuracy of each single model, if the difference is positive, it means that the fusion is effective.
5. The method of network data mining according to claim 4, characterized in that, Also includes S11: Cooperative application of clustering results and sentiment analysis; correlate the cluster label of S4 with the sentiment analysis result of S5, and calculate the proportion of each cluster sentiment, and feed back the features of the cluster with high negative proportion to S4 to adjust the feature selection.
6. The method of network data mining according to claim 5, wherein, Also includes S12: Dynamic feature updating mechanism; combine the results of S11 and the needs of S8, periodically reexecute the feature engineering step of S4, update the feature set and ensure compatibility with the original feature set.
7. The method of network data mining according to claim 6, characterized in that, Also includes S13: Dynamic adjustment of multi-model fusion weight, based on the performance score calculated by accuracy, precision and recall of the fusion model in S9 on new data, if the score decreases by more than 5%, the fusion weight is determined again.
8. An apparatus for network data mining, characterized by The network data mining device comprises: A target and range defining module is configured to determine a network data mining target and define a data source range, the data source including website page data, social media platform data, network forum data, and e-commerce platform data, and the data is obtained through a public API interface or a network crawler technology in compliance with a robots protocol; A data collection and storage module is configured to collect data within the range defined by the target and range defining module through a network crawler tool or an API interface, and store the raw data in a relational database including a data unique identifier, data content, data source, and collection time field; A data preprocessing module is configured to clean, integrate, convert, and normalize the raw data stored by the data collection and storage module; A feature engineering module is configured to extract TF-IDF features based on the data processed by the data preprocessing module, construct evaluation period features, obtain cluster labels as new features through K-means clustering, assess feature importance through a random forest, and retain high importance features; A sentiment analysis model module is configured to construct a support vector machine model, combine the high importance features obtained by the feature engineering module, correct the output of the support vector machine through a surface fitting model, train the model to achieve sentiment tendency classification, and train the model to achieve sentiment tendency classification; A sentiment analysis and feedback module is configured to correlate the sentiment tendency results obtained by the sentiment analysis model module with the user ID and evaluation time processed by the data preprocessing module, and feed back to the data preprocessing module for reprocessing if the results deviate greatly; A report generation module is configured to generate a network data mining report including data sources, processing procedures, analysis results, conclusions, and suggestions based on the analysis results of the sentiment analysis and feedback module.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the network data mining method of any one of claims 1 to 7.
10. An electronic device, comprising: The processor and the memory are included, and the memory stores a computer program, and the processor executes the computer program to implement the network data mining method of any one of claims 1 to 7.