Information classification processing method based on big data
Through the information classification method based on big data, the timing feature extraction and dynamic clustering framework is used, combined with information entropy analysis and user interaction behavior data, the classification threshold and confidence are dynamically adjusted, and the classification performance degradation under dynamic data is solved, and more efficient information classification is achieved.
Patent Information
- Application Number
- CN202510351595.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120277439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information classification, and particularly to an information classification and processing method based on big data. Background Art
[0002] Existing information classification technologies have obvious limitations in processing dynamic data. In today's rapidly changing digital environment, the characteristics of data are not static, but will change significantly over time. This dynamic nature is particularly prominent in many application scenarios. Take social media platforms as an example. The content generated by users shows a high degree of dynamism and diversity. Information such as text, pictures, or videos posted by users not only updates continuously in content, but also the characteristics such as language style, topic popularity, and sentiment tendency are constantly evolving. For example, a popular topic will quickly emerge and trigger a large number of discussions in a short period of time, and then quickly cool down, and the language expression and sentiment tendency of users during the discussion will also change accordingly. This rapid change makes it difficult for traditional information classification algorithms to effectively cope with. Traditional methods usually rely on static model training. Once the data characteristics change, the classification performance of the model will drop significantly. To adapt to these changes, traditional classification algorithms often need to retrain the model regularly, but this is not only time-consuming and laborious, but also leads to a further decline in classification accuracy during the model update period. This limitation in processing dynamic data seriously restricts the improvement of information classification technology in terms of real-time performance, accuracy, and adaptability, especially in application scenarios that require rapid response and precise decision-making, this problem is particularly prominent. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide an information classification and processing method based on big data to solve at least one of the above technical problems.
[0004] To achieve the above object, an information classification and processing method based on big data includes the following steps:
[0005] Step S1: Obtain an initial information data set; extract temporal features from the initial information data set to obtain multi-dimensional temporal feature representation data;
[0006] Step S2: Construct a hierarchical clustering framework based on the multi-dimensional temporal feature representation data to obtain initial classification framework data; calculate a dynamic cohesion coefficient for the initial classification framework data to obtain a clustering tightness parameter; perform adaptive clustering on the initial classification framework data according to the clustering tightness parameter to obtain dynamic classification basic data;
[0007] Step S3: Perform information entropy analysis on the dynamic classification basic data to obtain a classification threshold adjustment parameter; perform dynamic classification optimization based on the classification threshold adjustment on the dynamic classification basic data according to the classification threshold adjustment parameter to obtain adaptive classification decision data;
[0008] Step S4: Obtain user interaction behavior data; perform correlation analysis on the user interaction behavior data and the adaptive classification decision data to obtain a classification quality evaluation parameter; calculate the confidence attenuation coefficient of each category according to the classification quality evaluation parameter to obtain a clustering category confidence adjustment matrix;
[0009] Step S5: Apply the clustering category confidence adjustment matrix to the adaptive classification decision data to obtain adjusted classification result data; obtain the adjusted classification result data of multiple time windows, and perform weighted synthesis on the adjusted classification result data of multiple time windows to obtain the final information classification decision data.
[0010] Through the extraction of time series features, the present invention can capture the characteristics of data evolving over time, thereby providing a dynamic feature representation for subsequent classification, and solving the problem of the decline in classification performance caused by the change of data features in traditional methods. Through the adaptive clustering strategy based on the hierarchical clustering framework and the dynamic cohesion coefficient, the clustering structure can be flexibly adjusted to adapt to the dynamics of the data, enhancing the real-time performance and adaptability of the classification. Further, through the information entropy analysis and the dynamic classification optimization step based on the classification threshold adjustment, the classification threshold can be dynamically adjusted, the classification decision can be optimized, and the accuracy and stability of the classification can be improved. Through the correlation analysis combined with the user interaction behavior data, the classification quality can be evaluated in real time and the confidence can be adjusted, so that the classification result is more in line with the actual application requirements, enhancing the practicality and accuracy of the classification. Finally, through the weighted synthesis of multiple time windows, the data characteristics of different periods can be comprehensively considered, further improving the timeliness and reliability of the classification result, so that in the application scenarios that require rapid response and accurate decision-making, the overall performance of the information classification technology is significantly improved. Brief Description of the Drawings
[0011] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description with reference to the following drawings:
[0012] Figure 1 The step flow diagram of the information classification processing method based on big data in an embodiment is shown.
[0013] Figure 2 The detailed step flow diagram of step S16 in an embodiment is shown.
[0014] Figure 3 The detailed step flow diagram of step S5 in an embodiment is shown. Detailed Description of the Embodiment
[0015] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0016] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0017] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.
[0018] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides an information classification and processing method based on big data, including the following steps:
[0019] Step S1: Obtain an initial information data set; extract time series features from the initial information data set to obtain multi-dimensional time series feature representation data;
[0020] Step S2: Construct a hierarchical clustering framework based on the multi-dimensional time series feature representation data to obtain initial classification framework data; calculate the dynamic cohesion coefficient of the initial classification framework data to obtain a clustering tightness parameter; perform adaptive clustering on the initial classification framework data according to the clustering tightness parameter to obtain dynamic classification basic data;
[0021] Step S3: Perform information entropy analysis on the dynamic classification basic data to obtain a classification threshold adjustment parameter; perform dynamic classification optimization based on the classification threshold adjustment on the dynamic classification basic data according to the classification threshold adjustment parameter to obtain adaptive classification decision data;
[0022] Step S4: Obtain user interaction behavior data; perform correlation analysis on the user interaction behavior data and the adaptive classification decision data to obtain classification quality evaluation parameters; calculate the confidence decay coefficients for each category according to the classification quality evaluation parameters to obtain a clustering category confidence adjustment matrix;
[0023] Step S5: Apply the clustering category confidence adjustment matrix to the adaptive classification decision data to obtain adjusted classification result data; obtain the adjusted classification result data for multiple time windows, and perform weighted synthesis on the adjusted classification result data for multiple time windows to obtain the final information classification decision data.
[0024] In this embodiment, first, an initial information dataset is obtained from the database of the social media platform, which includes the user's behavior logs, such as posting time, comment content, liked objects, etc. The Python Pandas library is used to load the data, and time series features are extracted from the initial information dataset. By calculating features such as the behavior frequency and active time period of each user, multi-dimensional time series feature representation data is obtained. For example, calculate the posting frequency, comment quantity, and like behavior of each user in different time periods to generate a multi-dimensional dataset containing time series features. The linkage function in the Scipy library is used to adopt a bottom-up hierarchical clustering method, treating each user as an initial cluster, and gradually merging the clusters with the highest similarity to generate initial classification framework data. Then, the dynamic cohesion coefficient of the initial classification framework data is calculated. By calculating the cohesion degree of each hierarchical node and the separation degree of adjacent clustering categories, the clustering tightness parameter is obtained. According to the clustering tightness parameter, the initial classification framework data is adaptively clustered, and the hierarchical structure of the clustering is adjusted to obtain dynamic classification basic data. The information entropy formula is used to evaluate the data in combination with the within-class distribution characteristics and boundary clarity, quantify the uncertainty of the classification, and obtain the classification threshold adjustment parameter. According to the classification threshold adjustment parameter, the dynamic classification of the dynamic classification basic data is optimized based on the classification threshold adjustment. By adjusting the class attribution of the samples, the classification result is optimized to obtain adaptive classification decision data. User interaction behavior data is obtained, such as the interaction and following relationships between users. The Scikit-learn library is used to perform correlation analysis on the user interaction behavior data and the adaptive classification decision data. By calculating the correlation between features, the classification quality evaluation parameter is obtained. According to the classification quality evaluation parameter, the confidence decay coefficient of each category is calculated to obtain the clustering category confidence adjustment matrix. The adaptive classification decision data is applied to the clustering category confidence adjustment matrix to obtain the adjusted classification result data. Finally, the adjusted classification result data of multiple time windows is obtained, such as the data of the past week, month, and three months. The Python Pandas library is used to perform weighted synthesis on the adjusted classification result data of multiple time windows. According to the weights of the time windows (such as the weight of the past week is 0.6, the weight of one month is 0.3, and the weight of three months is 0.1), the weighted average is calculated to obtain the final information classification decision data.
[0025] Preferably, step S1 includes the following steps:
[0026] Step S11: Obtain the initial information dataset and perform timestamp correction on the initial information dataset to obtain the time series marked information dataset;
[0027] Specifically, the initial information dataset can be obtained from the platform database. This dataset contains the text content published by users and their corresponding timestamps. Use the open-source data processing tool Apache Flink to correct the timestamps of the initial information dataset. Through Flink's time window operation, convert the timestamps to a unified UTC format and correct the time deviation caused by incorrect user time zone settings or data transmission delays. For example, for data records with a timestamp deviation exceeding 1 minute, use Flink's Watermark mechanism for latency processing and correction. After timestamp correction, a time-series marked information dataset is obtained.
[0028] Step S12: Set time window parameters according to the time-series marked information dataset to obtain a window configuration scheme, and perform time window slicing on the time-series marked information dataset according to the window configuration scheme to obtain a preliminary time window set;
[0029] Specifically, taking the user text data of a social media platform as an example, the time-series marked information dataset contains the timestamps and text content of the texts published by users. Set the time window parameter to one window per hour. Use the Pandas library in Python to perform time window slicing on the time-series marked information dataset. Through the resample method of Pandas, slice the data at 1-hour time intervals to obtain a preliminary time window set. The data within each time window contains all the text records published by users during that time period. For example, for the discussion of a popular topic, time window slicing can clearly show the change trend of topic popularity in different time periods.
[0030] Step S13: Check the data sparsity of the preliminary time window set to obtain the window validity evaluation result;
[0031] Specifically, when processing the user text data of a social media platform, the preliminary time window set contains the text records published by users within each time window. Use the NumPy library in Python to check the data sparsity of the preliminary time window set. The specific operation is to calculate the ratio of the number of text records within each time window to the expected number of data points. If the ratio is lower than the set threshold (such as 0.6), it is considered that the data in this time window is sparse, which is caused by the decline in topic popularity or the decrease in user activity. For example, in a scenario where the topic popularity is gradually decreasing, the number of text records in some time windows is less than 60% of the expected value, and these time windows will be marked as sparse windows.
[0032] Step S14: Optimize and adjust the preliminary time window set according to the window validity evaluation result to obtain a dynamic information sequence;
[0033] Specifically, in the processing of user text data on a social media platform, according to the window validity evaluation results, it is found that the data sparsity of some time windows is relatively high, which is caused by the decline in topic popularity or data loss. The interpolation method is used to fill the data in the sparse time window. The specific operation is to use the interpolate function in the SciPy library of Python and select the linear interpolation method to estimate the missing data points. For example, for the situation where the number of text records in a time window is insufficient, virtual text records are generated through interpolation. At the same time, combined with the popular topic tags and user activity data on the social media platform, the sparse time window is supplemented, and finally a dynamic information sequence is obtained.
[0034] Step S15: Perform a fast Fourier transform on the dynamic information sequence to obtain the original data for spectral analysis, and denoise and smooth the original data for spectral analysis to obtain the frequency-domain feature data;
[0035] Specifically, the fast Fourier transform (FFT) can be performed on the dynamic information sequence using the NumPy library of Python. The specific operation is to call the numpy.fft.fft function to convert the time-domain signal (such as the text publishing frequency) into a frequency-domain signal to obtain the original data for spectral analysis. The signal module in the SciPy library of Python is used to denoise and smooth the original data for spectral analysis. The Butterworth low-pass filter is selected, and the cut-off frequency is set to 0.1 Hz to filter out the high-frequency noise components, and finally the smoothed frequency-domain feature data is obtained. For example, when analyzing the frequency change of user-posted texts, the periodic change of topic popularity can be clearly identified through FFT.
[0036] Step S16: Extract the time-series features from the initial information data set according to the frequency-domain feature data to obtain the multi-dimensional time-series feature representation data.
[0037] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S16.
[0038] Through timestamp correction, the present invention ensures the accuracy of data in the time dimension. By introducing the time window segmentation and sparsity check mechanisms, it can effectively identify and process the uneven distribution and sparse regions in the data, avoiding classification bias caused by data quality problems. The introduction of frequency-domain analysis further enhances the ability to mine the time-series features of data, and can more accurately capture the periodic and trend changes of data, thereby providing richer feature support for dynamic data classification.
[0039] Preferably, step S16 includes the following steps:
[0040] Step S161: Extract semantic features from the dynamic information sequence to obtain semantic feature data, and evaluate the importance of frequency domain feature data and semantic feature data to obtain a feature weight matrix;
[0041] Specifically, the natural language processing library SpaCy in Python can be used to process the text in the dynamic information sequence. Through SpaCy's pre-trained model (such as en_core_web_sm), the text is tokenized, part-of-speech tagged, and dependency syntactic analyzed to extract semantic features such as keywords, topic distribution, and sentiment tendency. For example, for a tweet about "environmental protection", the extracted keywords include "environmental protection", "garbage classification", "sustainable development", and the sentiment tendency is "positive". Use the random forest algorithm in the Scikit-learn
[0042] library to evaluate the importance of frequency domain feature data and semantic feature data. By setting n_estimators = 100 and max_depth = 10, calculate the importance score of each feature and generate a feature weight matrix. For example, the importance score of the sentiment tendency feature is 0.35, the importance score of the keyword feature is 0.40, and the importance score of the frequency domain feature is 0.25.
[0043] Step S162: Obtain a historical information dataset, and perform feature recognition and classification on the historical information dataset to obtain a historical frequency domain feature set and a historical semantic feature set;
[0044] Specifically, a historical information dataset for the past year can be obtained from the database of the social media platform, including the text content posted by users and their timestamps. Use the Pandas library in Python to preprocess the historical data and extract features such as timestamps, keywords, and sentiment tendency. Classify the historical data through the K-means clustering algorithm, and group similar text content into the same category. For example, classify the text about "environmental protection" into categories such as "garbage classification", "energy conservation and emission reduction", "environmental protection activities", etc., to obtain a historical frequency domain feature set and a historical semantic feature set. The frequency domain feature set reflects the change in the text posting frequency, while the semantic feature set contains information such as keywords and sentiment tendency.
[0045] Step S163: Identify the feature validity of the historical frequency domain feature set and the historical semantic feature set respectively to obtain a frequency domain feature validity evaluation dataset and a semantic feature validity evaluation dataset;
[0046] Specifically, the effectiveness of features in the historical frequency-domain feature set and the historical semantic feature set can be identified to evaluate which features contribute significantly to classification. The mutual information method in the Scikit-learn library is used to evaluate the effectiveness of features. For example, for the historical frequency-domain feature set, the mutual information value between each frequency-domain feature and the classification label is calculated, and a threshold of 0.1 is set. Features below this threshold are considered ineffective. For the historical semantic feature set, the mutual information values between the keyword and sentiment tendency features and the classification label are also calculated. After evaluation, it is found that the keyword "garbage classification" under the "environmental protection" theme has a high effectiveness, with a mutual information value of 0.25, while the keyword "sustainable development" has a low effectiveness, with a mutual information value of 0.08 and is marked as an ineffective feature. Finally, a frequency-domain feature effectiveness evaluation data set and a semantic feature effectiveness evaluation data set are obtained.
[0047] Step S164: Construct a timeline for the historical information data set to obtain a historical data timeline, and perform effectiveness-time analysis on the frequency-domain feature effectiveness evaluation data set and the semantic feature effectiveness evaluation data set based on the historical data timeline to obtain first effectiveness-time relationship data and second effectiveness-time relationship data;
[0048] Specifically, the Pandas library is used to arrange the historical data in chronological order to construct a historical data timeline. Effectiveness-time analysis is performed on the frequency-domain feature effectiveness evaluation data set and the semantic feature effectiveness evaluation data set based on the historical data timeline. For example, analyzing the change in the effectiveness of the keyword "garbage classification" over time, it is found that its mutual information value is 0.30 at the beginning of the year but drops to 0.15 at the end of the year. Through this analysis, first effectiveness-time relationship data (frequency-domain features) and second effectiveness-time relationship data (semantic features) are obtained.
[0049] Step S165: Construct a differential timeliness model for the historical frequency-domain feature set and the historical semantic feature set according to the first effectiveness-time relationship data and the second effectiveness-time relationship data to obtain a first feature timeliness model and a second feature timeliness model;
[0050] Step S166: Use the first feature timeliness model and the second feature timeliness model to evaluate the timeliness of the data of the corresponding feature categories in the frequency-domain feature data and the semantic feature data to obtain a feature timeliness coefficient set;
[0051] Specifically, the constructed timeliness model can be used to evaluate the timeliness of data corresponding to feature categories in the frequency-domain feature data and semantic feature data. For example, for the keyword "garbage classification" in the current data, the second feature timeliness model (stepwise decay model) is used to calculate its timeliness coefficient. Suppose the time from the current time to the time point when the feature validity decreases is 3 months, and the initial validity is 0.30, then the timeliness coefficient is 0.15. In this way, a set of feature timeliness coefficients is obtained.
[0052] Step S167: Construct a non-linear feature decay model based on the feature weight matrix and the set of feature timeliness coefficients to obtain an information time-series decay model;
[0053] Specifically, a non-linear feature decay model can be constructed according to the feature weight matrix and the set of feature timeliness coefficients. For example, for a feature, its weight is 0.40 and its timeliness coefficient is 0.15, then its decayed weight is 0.40 × 0.15 = 0.06. Use the NumPy library in Python to implement the non-linear decay formula, such as the decayed weight = weight × (1 - e -α × timeliness coefficient), where α is an adjustment parameter set to 2, and finally an information time-series decay model is obtained.
[0054] Step S168: Perform weighted fusion on the frequency-domain feature data and semantic feature data according to the information time-series decay model to obtain multi-dimensional time-series feature representation data.
[0055] Specifically, the frequency-domain feature data and semantic feature data can be weighted and fused according to the information time-series decay model. For example, for a sample, its frequency-domain feature weight is 0.25 and its decayed weight is 0.05; its semantic feature weight is 0.40 and its decayed weight is 0.10. Through weighted fusion, multi-dimensional time-series feature representation data is obtained. For example, the final feature vector contains the decayed frequency-domain feature value and semantic feature value, such as [0.05, 0.10]. These feature vectors can more accurately reflect the dynamic characteristics of the data.
[0056] Through the combination of semantic feature extraction and frequency-domain feature extraction, the present invention takes into account both the time-series characteristics and semantic content of the data, enabling a more comprehensive capture of the internal features of the data. By introducing feature importance evaluation and a weight matrix, the model can dynamically adjust its weights according to the actual contributions of the features, further optimizing the utilization efficiency of the features. Through timeline analysis of historical data and construction of a timeliness model, it is possible to effectively evaluate the changing trend of features over time, thereby assigning a timeliness coefficient to the features and solving the classification bias problem caused by outdated historical data in traditional methods. Through weighted fusion of features using a non-linear feature decay model, the adaptability of the model to dynamic data feature changes is further enhanced, enabling the classification results to more accurately reflect the real-time characteristics of the data.
[0057] Preferably, the construction of a differential timeliness model for the historical frequency-domain feature set and the historical semantic feature set according to the first validity-time relationship data and the second validity-time relationship data is specifically as follows:
[0058] When the first validity-time relationship data shows that the validity of the features in the historical information dataset decays exponentially over time, the feature timeliness model of the historical frequency-domain feature set is set as an exponential decay model, thereby obtaining the first feature timeliness model;
[0059] Specifically, for example, on a social media platform, assume that what is concerned is the discussion popularity (frequency-domain feature) of a certain popular movie among users. Through data collection, the number of posts related to this movie released by users in the past year is obtained. Use the Pandas library in Python to sort these data by time, and calculate the ratio of the number of posts at each time point to the initial popularity (the number of posts in the first week after the movie's release). Use the Matplotlib library to plot the curve of these ratios over time, and it is found that the curve shows an obvious exponential decay trend. Use the curve fitting function of the SciPy library to quantify this trend, and select the exponential decay model: validity(t) = initial validity × e -λt ; where t is the time (in weeks) and λ is the decay coefficient. Through fitting, λ = 0.1 is obtained. Therefore, the feature timeliness model of the historical frequency-domain feature set is set as an exponential decay model, obtaining the first feature timeliness model.
[0060] When the first validity-time relationship data shows that the validity of the features in the historical information dataset suddenly drops after a specific time point, the feature timeliness model of the historical frequency-domain feature set is set as a stepwise decay model, thereby obtaining the second feature timeliness model;
[0061] Specifically, for example, on a social media platform, assume that the focus is on the discussion popularity during a major sports event (such as the World Cup). Through data collection, the number of relevant posts published by users during and after the event is obtained. Use the Pandas library to sort the data by time, and calculate the ratio of the number of posts at each time point to the peak during the event. Use the Matplotlib library to plot the curve of these ratios over time, and it is found that the curve drops suddenly after the event (such as the first day after the final) and remains at a relatively low level. Select a step decay model to quantify this change:
[0062] When t < t0, effectiveness(t) = initial effectiveness;
[0063] When t ≥ t0, effectiveness(t) = initial effectiveness × α;
[0064] where t0 is the first day after the end of the event, and α is the ratio of the effectiveness after the decline. Through analysis, α = 0.2 is set. Therefore, the feature timeliness model of the historical frequency domain feature set is set as a step decay model to obtain the second feature timeliness model.
[0065] When the second effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset decays exponentially over time, then the feature timeliness model of the historical semantic feature set is set as an exponential decay model, thus obtaining the first feature timeliness model;
[0066] Specifically, for example, in the text analysis of a social media platform, assume that the focus is on the semantic features of keywords in a certain technical field (such as "blockchain"). Through data collection, posts related to "blockchain" published by users in the past two years are obtained, and semantic features such as the occurrence frequency and sentiment tendency of the keywords are extracted. Use the SpaCy library in Python for keyword extraction and sentiment analysis, and calculate the keyword importance score at each time point (such as the weighted sum of keyword frequency and sentiment tendency). Use the Matplotlib library to plot the curve of these scores over time, and it is found that the curve shows an exponential decay trend. Use the curve fitting function of the SciPy library to quantify this trend, and select the exponential decay model: effectiveness(t) = initial effectiveness × e -λt where t is the time (in months), and λ is the decay coefficient. Through fitting, λ = 0.05 is obtained. Therefore, the feature timeliness model of the historical semantic feature set is set as an exponential decay model to obtain the first feature timeliness model.
[0067] When the second effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset drops suddenly after a specific time point, then the feature timeliness model of the historical semantic feature set is set as a step decay model, thus obtaining the second feature timeliness model.
[0068] Specifically, for example, in the text analysis of a social media platform, assume that the focus is on the semantic features of keywords (such as "wedding" and "wedding dress") related to a specific event (such as the wedding of a certain star). Through data collection, relevant posts published by users before and after the event are obtained, and semantic features such as the occurrence frequency and sentiment tendency of the keywords are extracted. The SpaCy library in Python is used for keyword extraction and sentiment analysis, and the keyword importance score at each time point is calculated. The Matplotlib library is used to plot the curve of these scores over time, and it is found that the curve suddenly drops in the first month after the event and remains at a relatively low level. A step decay model is selected to quantify this change:
[0069] When t < t0, effectiveness(t) = initial effectiveness;
[0070] When t ≥ t0, effectiveness(t) = initial effectiveness × α;
[0071] where t0 is the first month after the event, and α is the effectiveness ratio after the decline. Through analysis, α = 0.1 is set. Therefore, the feature timeliness model of the historical semantic feature set is set as a step decay model to obtain the second feature timeliness model.
[0072] Through the construction of a differential timeliness model, for the different effectiveness change laws of historical frequency domain features and semantic features, an exponential decay model and a step decay model are respectively designed. This can accurately reflect the change trend of features over time and avoid the inadaptability of a single model to the change laws of different features. The exponential decay model can effectively capture the trend of features gradually weakening over time, while the step decay model can accurately identify the mutation of features after a specific time point. Through this differential processing, the model can more realistically reflect the timeliness change of features, thereby providing a more accurate basis for feature timeliness evaluation.
[0073] Preferably, the construction of the hierarchical clustering framework based on the multi-dimensional time series feature representation data in step S2 is specifically as follows:
[0074] Calculate the Pearson correlation coefficient for the multi-dimensional time series feature representation data to obtain a Pearson correlation coefficient matrix, and calculate the mutual information value for the multi-dimensional time series feature representation data to obtain a mutual information matrix. Based on the Pearson correlation coefficient matrix and the mutual information matrix, comprehensively evaluate the feature correlation of the multi-dimensional time series feature representation data to obtain a feature correlation heat map;
[0075] Specifically, correlation analysis can be performed on multi-dimensional time series feature representation data. Use the Pandas library in Python to load the data, and calculate the Pearson correlation coefficient matrix through the numpy.corrcoef() function in the NumPy library to evaluate the linear correlation between features. At the same time, use the mutual_info_regression function in the Scikit-learn library to calculate the mutual information matrix to evaluate the non-linear correlation between features. Combine the Pearson correlation coefficient matrix and the mutual information matrix, and use the Seaborn library to draw a feature correlation heatmap to intuitively display the correlation strength between features. For example, set the Pearson correlation coefficient threshold to 0.7 and the mutual information value threshold to 0.3, and mark the highly correlated feature pairs.
[0076] Mark redundant feature pairs on the feature correlation heatmap according to the preset feature correlation threshold to obtain a feature correlation matrix, which characterizes the strength of the dependence relationship between each feature dimension in the multi-dimensional time series feature representation data;
[0077] Specifically, according to the preset feature correlation threshold (Pearson correlation coefficient > 0.7 or mutual information value > 0.3), redundant features in the feature correlation heatmap can be marked. Use the Pandas library to traverse the heatmap data, mark the highly correlated feature pairs, and generate a feature correlation matrix. This matrix clearly characterizes the strength of the dependence relationship between each feature dimension in the multi-dimensional time series feature representation data. For example, if the "keyword frequency" and "sentiment tendency" show a correlation coefficient of 0.8 in the Pearson correlation coefficient matrix, it indicates that these two features are highly correlated, and one of them can be marked as a redundant feature.
[0078] Perform feature dimensionality reduction on the multi-dimensional time series feature representation data according to the feature correlation matrix to obtain dimensionality-reduced feature space data, and perform outlier detection on the dimensionality-reduced feature space data to obtain an outlier sample label set;
[0079] Specifically, feature dimensionality reduction can be performed on the multi-dimensional time series feature representation data according to the feature correlation matrix. Use the principal component analysis (PCA) method in the Scikit-learn library to reduce the original feature space to a new feature space that retains 90% of the variance, and obtain dimensionality-reduced feature space data. Perform outlier detection on the dimensionality-reduced data. Use the Isolation Forest algorithm to detect outlier samples through the IsolationForest class and mark the outlier sample set. For example, set contamination = 0.05, indicating that the proportion of outlier samples in the data is about 5%. Through the above steps, identify and mark the outliers caused by data entry errors or abnormal user behaviors.
[0080] Perform anomaly processing on the data in the dimensionality-reduced feature space according to the anomaly sample marking set to obtain purified feature data;
[0081] Specifically, the data operation function of the Pandas library can be used to remove the samples marked as anomalies from the dataset or replace them with reasonable values such as the median to obtain the purified feature data. For example, for the anomaly samples identified in outlier detection, the median of the same feature is selected for replacement.
[0082] Perform distance measurement on the purified feature data to obtain a feature space distance matrix, and construct a nearest neighbor relationship network based on the feature space distance matrix to obtain a sample association graph;
[0083] Specifically, the pairwise_distances function in the Scikit-learn library can be used to calculate the feature space distance matrix, and the Euclidean distance is selected as the measurement standard. Construct a nearest neighbor relationship network based on the feature space distance matrix. Use the NetworkX library to construct a sample association graph by calculating the connection strength between samples (such as distance-based similarity). For example, set the connection strength threshold to the 25% quantile of the median of the distance matrix, and consider the samples with a distance less than this threshold as nearest neighbor nodes and establish connections. Through the above steps, the obtained sample association graph can intuitively display the similarity and relevance between samples.
[0084] Construct an initial hierarchical clustering tree based on the sample association graph and the purified feature data to obtain initial classification framework data; wherein, the specific construction process of the initial hierarchical clustering tree is as follows:
[0085] Adopt a bottom-up hierarchical clustering strategy, regard each sample in the purified feature data as an independent cluster, gradually merge the most similar clusters, calculate the distance between samples in the feature space distance matrix and the connection strength in the sample association graph when calculating the similarity between clusters, record the hierarchical structure formed by each merge operation, and finally generate the initial classification framework data representing the internal classification hierarchical relationship of the data.
[0086] Specifically, a bottom-up hierarchical clustering strategy can be adopted, using the linkage function in the Scipy library, regarding each sample as an independent cluster, and gradually merging the most similar clusters. When calculating the similarity between clusters, consider both the distance between samples in the feature space distance matrix and the connection strength in the sample association graph. By recording each merge operation, a hierarchical structure is formed, and finally the initial classification framework data representing the internal classification hierarchical relationship of the data is generated. For example, set the method for calculating the similarity between clusters to "ward", generate a hierarchical clustering tree through the linkage function, and use the dendrogram function to visualize the hierarchical structure. Through the above steps, the obtained initial classification framework data can clearly display the internal classification hierarchical relationship of the data.
[0087] Through the comprehensive evaluation of the Pearson correlation coefficient and the mutual information value, the present invention can comprehensively analyze the correlation between features. The generated feature correlation heat map intuitively reflects the strength of the dependence relationship between each feature dimension. Based on this, feature dimensionality reduction not only effectively reduces data redundancy, but also retains key information and improves computational efficiency. At the same time, outlier detection and anomaly handling further purify the data and enhance the robustness of the model to noise and outliers. By constructing a sample association map and an initial hierarchical clustering tree, the inherent classification hierarchical relationship of the data can be dynamically reflected. Adopting a bottom-up hierarchical clustering strategy and combining the comprehensive measurement of the distance and connection strength between samples makes the clustering process more flexible and adaptable, and can effectively handle complex and changing data structures.
[0088] Preferably, the calculation of the dynamic cohesion coefficient for the initial classification framework data in step S2 is specifically as follows:
[0089] Calculate the cohesion degree of each hierarchical node in the initial classification framework data to obtain a clustering cohesion index; wherein, the clustering cohesion index characterizes the tightness degree of each hierarchical clustering.
[0090] Specifically, the hierarchical clustering tool in the Scipy library of Python can be used to calculate the cohesion degree of each hierarchical node in the initial classification framework data. The cohesion degree can be measured by calculating the average distance between samples within the cluster. The smaller the distance, the higher the cohesion degree. For example, for a certain hierarchical node, it contains 100 user behavior samples. By calculating the average Euclidean distance between these samples, the cohesion degree value is obtained as 0.2. This value reflects the tightness degree of the clustering of this hierarchical node.
[0091] Calculate the separation degree of adjacent clustering categories in the initial classification framework data to obtain a clustering interval index; wherein, the clustering interval index characterizes the clear boundary degree between clustering categories.
[0092] Specifically, the silhouette_score function in the Scikit-learn library can be used to calculate the separation degree of adjacent clustering categories in the initial classification framework data. The separation degree is measured by the silhouette coefficient. The higher the silhouette coefficient, the clearer the boundary between clustering categories. For example, for two adjacent clustering categories, the calculated result of the silhouette coefficient is 0.6, indicating that the boundary between these two categories is relatively clear. This value is used as the clustering interval index.
[0093] Evaluate the cohesion of the initial classification framework data according to the clustering cohesion index and the clustering interval index to obtain an initial cohesion coefficient;
[0094] Specifically, when processing social media user behavior data, the cohesion evaluation of the initial classification framework data can be combined with the clustering cohesion index and the clustering interval index. By calculating the weighted sum of the cohesion degree and the separation degree, the initial cohesion coefficient is obtained. For example, if the weights of the cohesion degree and the separation degree are set to 0.6 and 0.4 respectively, for a certain hierarchical node, its cohesion degree is 0.2, and the separation degree of adjacent categories is 0.6, then the initial cohesion coefficient is 0.6×0.2 + 0.4×0.6 = 0.36. This value reflects the cohesion of this hierarchical node.
[0095] Identify the temporal change trend of the multi-dimensional time series feature representation data to obtain the dynamic index of the feature distribution;
[0096] Specifically, the Pandas library in Python can be used to perform time series analysis on the data, and calculate the change rate of each feature within different time windows. For example, for the feature of user activity, calculate its change rate in the past week, month, and three months. By plotting the change rate curve, it is found that this feature shows an obvious downward trend in the past month, and the change rate is -15%. This change rate is used as the dynamic index of the feature distribution.
[0097] Evaluate the stability of the dynamic index of the feature distribution to obtain the distribution stability evaluation result; among them, the distribution stability evaluation result is either that the feature distribution is stable or that the feature distribution changes violently;
[0098] Specifically, the NumPy library in Python can be used to calculate the standard deviation of the dynamic index to evaluate the stability of the feature distribution. For example, for the change rate of user activity, its standard deviation in the past three months is 0.05. Set the stability threshold to 0.1. When the standard deviation is less than this threshold, it is considered that the feature distribution is stable; otherwise, it is considered that the feature distribution changes violently. In this example, since the standard deviation is 0.05, which is less than the threshold of 0.1, the evaluation result is that the feature distribution is stable.
[0099] When the distribution stability evaluation result is that the feature distribution is stable, linearly adjust the initial cohesion coefficient to obtain the first dynamic cohesion coefficient;
[0100] Specifically, when the distribution stability evaluation result is that the feature distribution is stable, then choose to linearly adjust the initial cohesion coefficient. Assume that the initial cohesion coefficient is 0.36, and set the linear adjustment coefficient to 1.2. Through a simple multiplication operation, the first dynamic cohesion coefficient is obtained as 0.36×1.2 = 0.432. This adjusted cohesion coefficient reflects the clustering tightness in the case of a stable feature distribution.
[0101] When the distribution stability evaluation result is that the feature distribution changes violently, non-linearly adjust the initial cohesion coefficient to obtain the second dynamic cohesion coefficient;
[0102] Specifically, when the distribution stability evaluation result is that the feature distribution changes drastically, the initial condensation coefficient can be non-linearly adjusted. Using the Math library in Python, the initial condensation coefficient is adjusted through an exponential function. Assume the initial condensation coefficient is 0.36, and the non-linear adjustment formula is set as the adjusted condensation coefficient = 1 - e -初始凝聚系数 . Through calculation, the second dynamic condensation coefficient is obtained as 1 - e -0.36 ≈0.302. This adjusted condensation coefficient reflects the clustering tightness in the case of drastic changes in the feature distribution.
[0103] Take the first dynamic condensation coefficient or the second dynamic condensation coefficient as the dynamic condensation coefficient, and perform a clustering threshold conversion on the dynamic condensation coefficient to obtain the clustering tightness parameter.
[0104] Specifically, the dynamic condensation coefficient (whether linearly adjusted or non-linearly adjusted) can be converted into a clustering tightness parameter. The clustering tightness parameter is used to determine the final clustering threshold. For example, assume the adjusted dynamic condensation coefficient is 0.432 (stable situation) or 0.302 (drastic change situation), and the conversion formula for the clustering tightness parameter is set as the clustering tightness parameter = dynamic condensation coefficient × 100. Therefore, the corresponding clustering tightness parameters are 43.2 or 30.2 respectively.
[0105] Through the calculation of the cohesion degree and separation degree, the present invention can accurately evaluate the tightness of clustering and the clarity of the boundary between categories, thereby providing a quantitative index for the clustering effect. By combining the recognition of the time series change trend and the stability evaluation, the dynamic condensation coefficient can be flexibly adjusted according to the stability of the feature distribution to adapt to the dynamic changes of the data. This can better handle the stable or drastic changes in the feature distribution, so as to maintain the accuracy and reliability of the clustering results in different situations. Finally, the clustering tightness parameter obtained by converting the dynamic condensation coefficient provides a key basis for subsequent adaptive clustering, further improving the flexibility and adaptability of the classification method, and enabling it to more effectively handle the complexity of dynamic data.
[0106] Preferably, the information entropy analysis of the dynamic classification basic data in step S3 is specifically as follows:
[0107] Statistically analyze the member distribution of each clustering category in the dynamic classification basic data to obtain the within-class distribution characteristics;
[0108] Specifically, the Pandas library in Python can be used to count the distribution of user behavior data in each category, such as the number of samples in each category, the mean and standard deviation of the sample features, etc. Suppose there are three clustering categories, representing high, medium, and low user activity levels respectively. Through statistics, it is found that category 1 (high activity) contains 1000 samples, with an average activity level of 8 (out of 10), and a standard deviation of 0.5; category 2 (medium activity) contains 800 samples, with an average activity level of 5, and a standard deviation of 1; category 3 (low activity) contains 600 samples, with an average activity level of 2, and a standard deviation of 0.8. These statistical results are used as the within-class distribution characteristics.
[0109] Identify the boundary regions for adjacent clustering categories in the dynamic classification basic data to obtain boundary clarity evaluation data;
[0110] Specifically, the silhouette_samples function in the Scikit-learn library can be used to calculate the silhouette coefficient for each sample, thereby identifying the boundary regions. The closer the silhouette coefficient is to 0, the closer the sample is to the category boundary. For example, for the boundary region between category 1 and category 2, the silhouette coefficient threshold is set to 0.1. Through calculation, it is found that 100 samples have a silhouette coefficient less than 0.1, and these samples are considered boundary region samples. The identification results of the boundary regions are used as the boundary clarity evaluation data.
[0111] Quantify the classification uncertainty of the dynamic classification basic data based on the within-class distribution characteristics and the boundary clarity evaluation data to obtain classification information entropy data;
[0112] Specifically, the information entropy formula can be used to calculate the uncertainty of each clustering category. For example, for category 1, its member distribution is relatively concentrated (small standard deviation), and there are few boundary region samples, so the information entropy is low, indicating low classification uncertainty; while for category 2, its member distribution is more dispersed (large standard deviation), and there are more boundary region samples, so the information entropy is high, indicating high classification uncertainty. Through the above steps, the classification information entropy data is obtained.
[0113] Perform simulated classification on the classification information entropy data under different classification threshold conditions to obtain different condition classification result datasets, identify the threshold-sensitive interval and the threshold-stable interval for the different condition classification result datasets to obtain threshold interval evaluation data, and generate a classification threshold - performance curve based on the threshold interval evaluation data;
[0114] Specifically, the NumPy library of Python can be used to generate a series of classification thresholds (such as 0.1, 0.2, ..., 0.9), and the data can be reclassified according to these thresholds. For example, when the threshold is 0.5, the samples in the boundary region between class 1 and class 2 are reassigned to class 1; when the threshold is 0.7, these samples are reassigned to class 2. By simulating classification multiple times, a classification result dataset under different conditions is obtained. The Matplotlib library is used to plot the performance curve of the classification results to identify the threshold-sensitive interval (such as 0.4 - 0.6) and the threshold-stable interval (such as 0.7 - 0.9). These threshold interval evaluation data are used to generate the classification threshold - performance curve.
[0115] Calculate the classification threshold adjustment parameter according to the classification information entropy data and the classification threshold - performance curve to obtain the classification threshold adjustment parameter.
[0116] Specifically, by analyzing the classification threshold - performance curve, the threshold interval with the optimal performance (such as 0.7 - 0.9) can be selected, and the average classification information entropy within this interval can be calculated. For example, assume that when the threshold is 0.8, the classification information entropy is the lowest, indicating the minimum classification uncertainty. Calculate the classification threshold adjustment parameter as the ratio of the current threshold to the average information entropy, such as In this way, the classification threshold adjustment parameter is obtained.
[0117] Through member distribution statistics and boundary region identification, the present invention can accurately capture the intra-class features and the clarity of the inter-class boundary, thereby comprehensively evaluating the certainty and ambiguity of classification. The classification threshold - performance curve generated based on the classification information entropy data and threshold sensitivity analysis further reveals the changing trend of the classification effect under different threshold conditions. By optimizing the classification threshold, the accuracy and adaptability of the classification decision are improved. Especially when dealing with complex and changing dynamic data, the misclassification rate can be effectively reduced.
[0118] Preferably, in step S3, the dynamic classification optimization based on the classification threshold adjustment for the dynamic classification basic data is specifically as follows:
[0119] Re-evaluate the class attribution of each sample in the dynamic classification basic data according to the classification threshold adjustment parameter to obtain the preliminary classification evaluation data of the samples;
[0120] Specifically, the class membership of each sample can be re-evaluated according to the classification threshold adjustment parameter. Using the Pandas library in Python, combined with the adjusted classification threshold (e.g., 0.8), the classification result of each sample is re-judged. For example, for a sample, its current classification is "high activity", but according to the new classification threshold, its feature value is closer to the "medium activity" category. Therefore, the class of this sample is re-evaluated as "medium activity" and recorded as the preliminary classification evaluation data of the sample.
[0121] Extract the sample feature values from the dynamic classification basic data, and use the preset classification model to calculate the Euclidean distance from each sample feature value to the decision boundary of each category, generating sample-boundary distance data;
[0122] Specifically, a classification model in the Scikit-learn library (such as logistic regression or support vector machine) can be used to calculate the Euclidean distance from each sample feature value to the decision boundary of each category. For example, for a sample, its distances to the decision boundaries of the "high activity" and "medium activity" categories are 0.2 and 0.3 respectively. These distance values are recorded as sample-boundary distance data.
[0123] When the sample-boundary distance data is less than the preset boundary threshold, the feature weight coefficient of the corresponding sample feature value is increased according to the preset weight adjustment strategy, obtaining the sample classification discrimination condition;
[0124] Specifically, it can be assumed that the preset boundary threshold is 0.25. When the sample-boundary distance is less than this threshold, it indicates that the sample is close to the decision boundary and the classification uncertainty is relatively high. At this time, according to the preset weight adjustment strategy (such as increasing the weight coefficient by 0.5), the sample feature values close to the boundary are adjusted. For example, for a sample, its boundary distance to the "high activity" category is 0.2, which is less than the threshold of 0.25. Therefore, the weight coefficient of its feature value is increased from 1 to 1.5.
[0125] Based on the dynamic classification basic data, calculate the membership probability of each sample in each category according to the sample classification discrimination condition, obtaining the class membership probability set;
[0126] Specifically, a classification model in the Scikit-learn library (such as random forest or Bayesian classifier) can be used to calculate the probability distribution of each sample feature value in each category. For example, for a sample, its membership probabilities in the "high activity", "medium activity", and "low activity" categories are 0.45, 0.35, and 0.20 respectively. These probability values are recorded as the class membership probability set.
[0127] When the difference between the highest class probability value and the second-highest class probability value in the class membership probability concentration of a sample is less than 0.15, the sample is marked as a high-uncertainty sample, and both the class labels corresponding to its highest probability and second-highest probability are retained in the preliminary sample classification evaluation data;
[0128] Specifically, when the difference between the highest class probability value and the second-highest class probability value of a sample is less than 0.15, the sample is marked as a high-uncertainty sample. For example, for a sample, its membership probabilities in the "high activity" and "medium activity" classes are 0.45 and 0.35 respectively, and the probability difference is 0.10, which is less than the threshold of 0.15. Therefore, the sample is marked as a high-uncertainty sample, and both the class labels corresponding to its highest probability and second-highest probability (i.e., "high activity" and "medium activity") are retained in the preliminary sample classification evaluation data. This marking process ensures the special treatment of uncertain samples.
[0129] According to the Mahalanobis distance between the sample feature values and the class feature center vectors of each class and the position of the sample feature values in the within-class distribution, a confidence score value is calculated for the classification decision of each sample. The range of this score value is from 0 to 1, and the score calculation formula is: where Z is the confidence score value, r is the distance from the sample to the class center, R is the class radius, and c is the within-class feature consistency index, and finally a threshold-optimized classification result is generated; among them, the threshold-optimized classification result includes the sample identifier, the optimized class membership, and the corresponding confidence score;
[0130] Specifically, the confidence score value for the classification decision of each sample can be calculated according to the Mahalanobis distance between the sample feature values and the class feature center vectors of each class and the position of the sample in the within-class distribution. Assume that the class radius is 1.5, the distance from the sample to the class center is 1.0, and the within-class feature consistency index is 0.8. According to the formula: The calculated confidence score value is 0.6. The range of this score value is from 0 to 1, which reflects the reliability of the sample classification decision. The finally generated threshold-optimized classification result includes the sample identifier, the optimized class membership, and the corresponding confidence score.
[0131] Perform post-processing rule optimization on the threshold-optimized classification result to obtain rule-enhanced classification data, and perform consistency verification on the rule-enhanced classification data to obtain a consistency check report;
[0132] Specifically, the Pandas library of Python can be used to adjust the classification results in combination with preset business rules (for example, the user activity cannot be lower than a certain threshold). For example, if a sample is classified as "high activity", but its actual activity is lower than the set threshold (such as 0.7), then its category is adjusted to "medium activity". Subsequently, consistency verification is performed on the rule-enhanced classification data, and the confusion matrix and classification report tools in the Scikit-learn library are used to generate a consistency check report. The report shows metrics such as classification accuracy, recall rate, and F1 score.
[0133] Adjust the rule-enhanced classification data according to the consistency check report to obtain adaptive classification decision data.
[0134] Specifically, the rule-enhanced classification data can be adjusted according to the consistency check report. For example, if the consistency check report shows that the classification accuracy of certain categories is low, then the classification threshold is further adjusted or the rules are optimized. The finally obtained adaptive classification decision data can dynamically adapt to the changes in the data and meet the business requirements. For example, after adjustment, the classification accuracy of user behavior data is increased from 85% to 90%, the recall rate is increased from 80% to 88%, and the F1 score is increased from 82% to 89%.
[0135] Through the re-evaluation of the sample category attribution by adjusting the parameters based on the classification threshold and combining the distance analysis from the sample to the decision boundary, the present invention can accurately identify the samples close to the classification boundary and adjust their feature weights, thereby enhancing the discrimination of the classification. At the same time, through uncertainty evaluation, multiple category labels are retained for samples with high uncertainty, avoiding the misclassification risk caused by a single decision. The reliability of the classification decision is further quantified through the confidence scoring mechanism, providing an intuitive credibility assessment for the classification results. Finally, the post-processing rule optimization and consistency verification ensure the stability and consistency of the classification results, enabling the classification system to better adapt to the changes in dynamic data and significantly improving the adaptability and robustness of the classification.
[0136] Preferably, step S4 includes the following steps:
[0137] Step S41: Obtain user interaction behavior data, and perform category hierarchy induction extraction on the adaptive classification decision data to obtain classification architecture data;
[0138] Specifically, user interaction behavior data can be obtained from the platform database, loaded using the Pandas library in Python, and the category hierarchy of the adaptive classification decision data can be inductively extracted. By analyzing the clustering results, classification system structure data at different levels can be extracted. For example, users can be divided into "active users", "ordinary users", and "inactive users", and further subdivided into subcategories such as "active users - high interaction" and "active users - low interaction".
[0139] Step S42: Calculate the correlation degree of multi-dimensional cross features based on the classification system structure data and user interaction behavior data to obtain category matching degree scoring data;
[0140] Specifically, the mutual_info_classif function in the Scikit-learn library can be used to calculate the mutual information value between user behavior features (such as the number of posts, comment frequency, number of likes) and the classification system structure. For example, the mutual information value between the "active user" category and the "number of posts" feature is 0.4, and the mutual information value between the "comment frequency" feature is 0.35. Through the above steps, the matching degree scoring data between each category and user behavior features can be obtained.
[0141] Step S43: Perform a trend dynamic evolution simulation on the category matching degree scoring data according to a preset time window to obtain category effectiveness evolution data;
[0142] Specifically, the Pandas library in Python can be used to group the category matching degree scoring data by time window and calculate the average value of the category matching degree scoring within each time window. For example, for the "active user" category, the matching degree score in the first week is 0.4, 0.42 in the second week, and 0.38 in the third week. By plotting a time series graph and observing the change trend of the category matching degree score over time, category effectiveness evolution data can be obtained. These data reflect the stability and changes of the category within different time windows.
[0143] Step S44: Conduct association rule mining based on the category effectiveness evolution data and user interaction behavior data to obtain behavior-classification association rule data;
[0144] Specifically, the apriori and association_rules functions in the MLxtend library can be used to mine the association rules between user behavior and categories. For example, setting the minimum support to 0.1 and the minimum confidence to 0.5, the confidence of the rule "If the number of posts a user publishes per week exceeds 10, then the user belongs to the 'active user' category" is 0.8. Through the above steps, behavior-classification association rule data can be obtained, which is used to evaluate the internal relationship between user behavior and categories.
[0145] Step S45: Quantify the cognitive pattern differences of the adaptive classification decision data based on the behavior-classification association rule data to obtain the classification balance quality index data, and perform a consistency comparison between the classification balance quality index data and the adaptive classification decision data to obtain the classification quality evaluation parameters;
[0146] Specifically, the Pandas library in Python can be used to calculate the difference between the actual distribution and the association rule prediction distribution of user behavior characteristics in each category. For example, for the "active user" category, the standard deviation of the actual post quantity distribution is 2, while the standard deviation predicted by the association rule is 1.5, and the difference is 0.5. Through the above steps, the classification balance quality index data is obtained. A consistency comparison is performed between the classification balance quality index data and the adaptive classification decision data to generate the classification quality evaluation parameters. For example, the accuracy of the consistency comparison result is calculated to be 90%, the recall rate is 85%, and the F1 score is 87%.
[0147] Step S46: Calculate the confidence attenuation coefficient for each category according to the classification quality evaluation parameters to obtain the clustering category confidence adjustment matrix.
[0148] Specifically, the NumPy library in Python can be used to calculate the confidence attenuation coefficient for each category according to the classification quality evaluation parameters (such as accuracy and recall rate). For example, for the "active user" category, the accuracy of its classification quality evaluation parameters is 90% and the recall rate is 85%, and the calculated confidence attenuation coefficient is 0.9. In this way, the clustering category confidence adjustment matrix is generated for adjusting the confidence of each category. For example, the confidence of the "active user" category in the adjustment matrix is adjusted from 0.9 to 0.81 (0.9×0.9).
[0149] Through inductively extracting the classification architecture data, the present invention can clearly reflect the hierarchical relationship of classification. Through the calculation of the multi-dimensional cross-feature correlation degree and the category matching degree score, the correlation between user behavior and classification results is further quantified, which can dynamically reflect the real needs of users. Through the trend dynamic evolution simulation and association rule mining, the dynamic changes of category effectiveness are captured, revealing the internal law between user behavior and classification results. Through the quantification of cognitive pattern differences and the calculation of classification quality evaluation parameters, the accuracy and consistency of classification results can be evaluated in real time, and the classification confidence is adjusted accordingly.
[0150] Preferably, step S5 includes the following steps:
[0151] Step S51: Re-calculate the confidence using the clustering category confidence adjustment matrix for the adaptive classification decision data to obtain the preliminary adjusted classification data;
[0152] Specifically, the confidence of each sample can be recalculated using a clustering category confidence adjustment matrix. In the specific operation, the adaptive classification decision data and the clustering category confidence adjustment matrix are loaded using the Pandas library in Python. For each sample, the corresponding confidence decay coefficient is found according to its belonging category, and its initial confidence is multiplied by this coefficient. For example, a sample has an initial confidence of 0.9, belongs to the category of "active users", and the corresponding confidence decay coefficient is 0.9. The recalculated confidence is 0.81. Through the above steps, the preliminary adjusted classification data is obtained.
[0153] Step S52: Re-evaluate the classification boundary based on the preliminary adjusted classification data to obtain boundary-optimized classification data;
[0154] Specifically, a classification model in the Scikit-learn library (such as a support vector machine or logistic regression) can be used to recalculate the classification boundary in combination with the adjusted confidence data. For example, for the two categories of "active users" and "ordinary users", by analyzing the adjusted confidence distribution, it is found that the confidence of some samples is close to the boundary value (such as 0.5). By retraining the classification model and adjusting the classification boundary, the boundary can be made more adaptable to the actual distribution of the data. For example, the adjusted classification boundary classifies samples with a confidence less than 0.45 as "ordinary users" and samples with a confidence greater than or equal to 0.45 as "active users". Through the above steps, the boundary-optimized classification data is obtained.
[0155] Step S53: Perform anomaly detection and correction on the boundary-optimized classification data to obtain adjusted classification result data;
[0156] Specifically, the Isolation Forest algorithm can be used to detect abnormal samples through the IsolationForest class in the Scikit-learn library. Set contamination = 0.05, indicating that the proportion of abnormal samples is about 5%. For the detected abnormal samples, they are corrected in combination with the business logic. For example, if a sample is misclassified as an "active user", but its actual behavioral characteristics (such as the number of posts and interaction frequency) are lower than the standard of "active users", it is reclassified as an "ordinary user". In this way, the adjusted classification result data is obtained.
[0157] Step S54: Obtain the adjusted classification result data for multiple time windows, and perform time window weight assignment on the adjusted classification result data for multiple time windows to obtain time window weight data;
[0158] Specifically, it can be assumed that three time windows are selected: the past week, the past month, and the past three months. Using the Pandas library in Python, weights are assigned according to the length of the time window and the freshness of the data. For example, the weight for the past week is set to 0.6, the weight for the past month is 0.3, and the weight for the past three months is 0.1. In this way, the time window weight data is obtained.
[0159] Step S55: Perform weighted fusion on the adjusted classification result data of multiple time windows according to the time window weight data to obtain the fused classification decision data;
[0160] Specifically, the NumPy library can be used to multiply the classification result of each time window by the corresponding weight and calculate the weighted average. For example, for a sample, its classification confidence levels in the past week, the past month, and the past three months are 0.8, 0.7, and 0.6 respectively, and the corresponding weights are 0.6, 0.3, and 0.1. The confidence level after weighted fusion is 0.8×0.6 + 0.7×0.3 + 0.6×0.1 = 0.75. Through the above steps, the fused classification decision data is obtained.
[0161] Step S56: Modify the rules for the fused classification decision data to obtain the rule-optimized classification data, and perform a final consistency check on the fused classification decision data based on the rule-optimized classification data to obtain the final information classification decision data.
[0162] Specifically, the Pandas library in Python can be used to adjust the classification result in combination with the preset business rules (such as the user activity cannot be lower than a certain threshold). For example, if a sample is classified as an "active user", but its actual activity is lower than the set threshold (such as 0.7), then its category is adjusted to an "ordinary user". Perform a final consistency check on the data after rule modification. Use the confusion matrix and classification report tools in the Scikit-learn library to generate a consistency check report. The report shows metrics such as classification accuracy, recall rate, and F1 score, which are used to evaluate the consistency and stability of the classification result. For example, after rule modification, the classification accuracy increases from 90% to 92%, the recall rate increases from 85% to 88%, and the F1 score increases from 87% to 89%. The finally obtained classification result can better adapt to the changes in the data and meet the business requirements.
[0163] The present invention recalculates the confidence of classification decision data through a clustering category confidence adjustment matrix, which can dynamically adjust the confidence of classification results to make them more in line with the actual data characteristics. The classification accuracy is further optimized by re-evaluating the classification boundary and correcting anomaly detection, reducing misclassification in the boundary area and the interference of abnormal data. The weighted fusion strategy of multiple time windows comprehensively considers the characteristic changes of data in different periods, assigns appropriate weights to different time windows, and thus more comprehensively reflects the dynamic characteristics of the data. Finally, through rule correction and consistency verification, the stability and accuracy of the classification results are further ensured.
[0164] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be embraced within the present invention.
[0165] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. An information classification and processing method based on big data, characterized in that, It includes the following steps: Step S1: Obtain the initial information dataset; extract the time series features from the initial information dataset to obtain the multi-dimensional time series feature representation data; Step S2: Construct a hierarchical clustering framework based on the multi-dimensional time series feature representation data to obtain the initial classification framework data; Calculate the dynamic cohesion coefficient for the initial classification framework data to obtain the clustering tightness parameter; perform adaptive clustering on the initial classification framework data according to the clustering tightness parameter to obtain the dynamic classification basic data; Step S3: Conduct information entropy analysis on the dynamic classification basic data to obtain the classification threshold adjustment parameter; perform dynamic classification optimization based on the classification threshold adjustment on the dynamic classification basic data according to the classification threshold adjustment parameter to obtain the adaptive classification decision data; Step S4: Obtain the user interaction behavior data; conduct correlation analysis on the user interaction behavior data and the adaptive classification decision data to obtain the classification quality evaluation parameter; calculate the confidence decay coefficient for each category according to the classification quality evaluation parameter to obtain the clustering category confidence adjustment matrix; Step S5: Apply the clustering category confidence adjustment matrix to the adaptive classification decision data to obtain the adjusted classification result data; obtain the adjusted classification result data for multiple time windows, and perform weighted synthesis on the adjusted classification result data for multiple time windows to obtain the final information classification decision data.
2. The information classification and processing method based on big data according to claim 1, wherein, Step S1 includes the following steps: Step S11: Obtain the initial information dataset, and perform timestamp correction on the initial information dataset to obtain the time series marked information dataset; Step S12: Set the time window parameters according to the time series marked information dataset to obtain the window configuration scheme, and perform time window splitting on the time series marked information dataset according to the window configuration scheme to obtain the preliminary time window set; Step S13: Check the data sparsity of the preliminary time window set to obtain the window effectiveness evaluation result; Step S14: Optimize and adjust the preliminary time window set according to the window effectiveness evaluation result to obtain the dynamic information sequence; Step S15: Perform fast Fourier transform on the dynamic information sequence to obtain the original spectrum analysis data, and perform denoising and smoothing on the original spectrum analysis data to obtain the frequency domain feature data; Step S16: Extract the time series features from the initial information dataset according to the frequency domain feature data to obtain the multi-dimensional time series feature representation data.
3. The information classification and processing method based on big data according to claim 2, wherein Step S16 includes the following steps: Step S161: Extract the semantic features from the dynamic information sequence to obtain the semantic feature data, and perform feature importance evaluation on the frequency domain feature data and the semantic feature data to obtain the feature weight matrix; Step S162: Obtain the historical information dataset, and perform feature recognition and classification on the historical information dataset to obtain the historical frequency domain feature set and the historical semantic feature set; Step S163: Perform feature effectiveness recognition on the historical frequency domain feature set and the historical semantic feature set respectively to obtain the frequency domain feature effectiveness evaluation dataset and the semantic feature effectiveness evaluation dataset; Step S164: Construct a timeline for the historical information dataset to obtain a historical data timeline, and perform effectiveness-time analysis on the frequency-domain feature effectiveness evaluation dataset and the semantic feature effectiveness evaluation dataset respectively based on the historical data timeline to obtain first effectiveness-time relationship data and second effectiveness-time relationship data; Step S165: Construct a differential timeliness model for the historical frequency-domain feature set and the historical semantic feature set according to the first effectiveness-time relationship data and the second effectiveness-time relationship data to obtain a first feature timeliness model and a second feature timeliness model; Step S166: Use the first feature timeliness model and the second feature timeliness model to perform timeliness evaluation on the data of the corresponding feature categories in the frequency-domain feature data and the semantic feature data to obtain a feature timeliness coefficient set; Step S167: Construct a non-linear feature attenuation model according to the feature weight matrix and the feature timeliness coefficient set to obtain an information time series attenuation model; Step S168: Perform weighted fusion on the frequency-domain feature data and the semantic feature data according to the information time series attenuation model to obtain multi-dimensional time series feature representation data.
4. The information classification and processing method based on big data according to claim 3, characterized in that The construction of the differential timeliness model for the historical frequency-domain feature set and the historical semantic feature set according to the first effectiveness-time relationship data and the second effectiveness-time relationship data is specifically as follows: When the first effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset decays exponentially over time, the feature timeliness model of the historical frequency-domain feature set is set as an exponential decay model, thereby obtaining the first feature timeliness model; When the first effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset suddenly drops after a specific time point, the feature timeliness model of the historical frequency-domain feature set is set as a stepwise decay model, thereby obtaining the second feature timeliness model; When the second effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset decays exponentially over time, the feature timeliness model of the historical semantic feature set is set as an exponential decay model, thereby obtaining the first feature timeliness model; When the second effectiveness-time relationship data shows that the effectiveness of the features in the historical information dataset suddenly drops after a specific time point, the feature timeliness model of the historical semantic feature set is set as a stepwise decay model, thereby obtaining the second feature timeliness model.
5. The information classification and processing method based on big data according to claim 1, characterized in that The construction of the hierarchical clustering framework based on the multi-dimensional time series feature representation data in Step S2 is specifically as follows: Calculate the Pearson correlation coefficient for the multi-dimensional time series feature representation data to obtain a Pearson correlation coefficient matrix, and calculate the mutual information value for the multi-dimensional time series feature representation data to obtain a mutual information matrix. Comprehensively evaluate the feature correlation of the multi-dimensional time series feature representation data according to the Pearson correlation coefficient matrix and the mutual information matrix to obtain a feature correlation heat map; Mark the redundant feature pairs for the feature correlation heat map according to the preset feature correlation threshold to obtain a feature correlation matrix, and the feature correlation matrix represents the intensity of the dependence relationship between the feature dimensions in the multi-dimensional time series feature representation data; Perform feature dimensionality reduction on the multi-dimensional time-series feature representation data according to the feature correlation matrix to obtain the dimensionality-reduced feature space data, and perform outlier detection on the dimensionality-reduced feature space data to obtain the outlier sample label set; Perform anomaly processing on the dimensionality-reduced feature space data according to the outlier sample label set to obtain the purified feature data; Perform distance measurement on the purified feature data to obtain the feature space distance matrix, and construct a neighbor relationship network based on the feature space distance matrix to obtain the sample association graph; Construct an initial hierarchical clustering tree based on the sample association graph and the purified feature data to obtain the initial classification framework data; wherein, the specific construction process of the initial hierarchical clustering tree is as follows: Adopt a bottom-up hierarchical clustering strategy, regard each sample in the purified feature data as an independent cluster, gradually merge the most similar clusters, calculate the distance between samples in the feature space distance matrix and the connection strength in the sample association graph when calculating the similarity between clusters, record the hierarchical structure formed by each merge operation, and finally generate the initial classification framework data representing the internal classification hierarchical relationship of the data.
6. The information classification and processing method based on big data according to claim 1, characterized in that The specific calculation of the dynamic cohesion coefficient for the initial classification framework data in step S2 is as follows: Calculate the cohesion degree of each hierarchical node in the initial classification framework data to obtain the clustering cohesion index; wherein, the clustering cohesion index represents the tightness of each hierarchical clustering. Calculate the separation degree of adjacent clustering categories in the initial classification framework data to obtain the clustering interval index; wherein, the clustering interval index represents the clear boundary degree between clustering categories. Perform a cohesion evaluation on the initial classification framework data according to the clustering cohesion index and the clustering interval index to obtain the initial cohesion coefficient; Identify the time-series change trend of the multi-dimensional time-series feature representation data to obtain the feature distribution dynamic index; Perform a stability evaluation on the feature distribution dynamic index to obtain the distribution stability evaluation result; wherein, the distribution stability evaluation result is either that the feature distribution is stable or that the feature distribution changes violently. When the distribution stability evaluation result is that the feature distribution is stable, linearly adjust the initial cohesion coefficient to obtain the first dynamic cohesion coefficient; When the distribution stability evaluation result is that the feature distribution changes violently, non-linearly adjust the initial cohesion coefficient to obtain the second dynamic cohesion coefficient; Take the first dynamic cohesion coefficient or the second dynamic cohesion coefficient as the dynamic cohesion coefficient, and perform a clustering threshold conversion on the dynamic cohesion coefficient to obtain the clustering tightness parameter.
7. The information classification and processing method based on big data according to claim 1, wherein, The specific information entropy analysis of the dynamic classification basic data in step S3 is as follows: Perform a member distribution statistics on each clustering category in the dynamic classification basic data to obtain the within-class distribution characteristics; Identify the boundary region of adjacent clustering categories in the dynamic classification basic data to obtain the boundary clarity evaluation data; Quantify the classification uncertainty of the dynamic classification basic data according to the within-class distribution characteristics and the boundary clarity evaluation data to obtain the classification information entropy data; Simulate the classification of classification information entropy data under different classification threshold conditions to obtain classification result data sets under different conditions. Identify the threshold-sensitive interval and threshold-stable interval of the classification result data sets under different conditions to obtain threshold interval evaluation data, and generate a classification threshold-performance curve based on the threshold interval evaluation data; Calculate the classification threshold adjustment parameter according to the classification information entropy data and the classification threshold-performance curve to obtain the classification threshold adjustment parameter.
8. The information classification and processing method based on big data according to claim 1, wherein The dynamic classification optimization based on the classification threshold adjustment for the dynamic classification basic data according to the classification threshold adjustment parameter described in step S3 is specifically as follows: Re-evaluate the class attribution of each sample in the dynamic classification basic data according to the classification threshold adjustment parameter to obtain preliminary sample classification evaluation data; Extract the sample feature values from the dynamic classification basic data, and use a preset classification model to calculate the Euclidean distance from each sample feature value to the decision boundary of each category to generate sample-boundary distance data; When the sample-boundary distance data is less than the preset boundary threshold, increase the feature weight coefficient of the corresponding sample feature value according to the preset weight adjustment strategy to obtain the sample classification discrimination condition; Calculate the attribution probability of each sample in each category based on the dynamic classification basic data according to the sample classification discrimination condition to obtain a category attribution probability set; When the difference between the highest category probability value and the second highest category probability value in the category attribution probability set of the sample is less than 0.15, mark the sample as a high-uncertainty sample, and simultaneously retain the category labels corresponding to its highest probability and second highest probability in the preliminary sample classification evaluation data; Calculate a confidence score value for the classification decision of each sample based on the Mahalanobis distance between the sample eigenvalue and the feature center vector of each category, as well as the position of the sample eigenvalue in the within-class distribution. The range of this score value is from 0 to 1, and the score calculation formula is: where Z is the confidence score value, r is the distance from the sample to the category center, R is the category radius, and c is the within-class feature consistency index, and finally generate a threshold-optimized classification result; among them, the threshold-optimized classification result includes the sample identifier, the optimized category attribution, and the corresponding confidence score; Perform post-processing rule optimization on the threshold-optimized classification results to obtain rule-enhanced classification data, and perform consistency verification on the rule-enhanced classification data to obtain a consistency check report; Adjust the rule-enhanced classification data according to the consistency check report to obtain adaptive classification decision data.
9. The information classification and processing method based on big data according to claim 1, characterized in that Step S4 includes the following steps: Step S41: Obtain user interaction behavior data, and perform category hierarchy induction extraction on the adaptive classification decision data to obtain classification system structure data; Step S42: Calculate the multi-dimensional cross-feature correlation degree based on the classification system structure data and the user interaction behavior data to obtain category matching degree scoring data; Step S43: Simulate the trend dynamic evolution of the category matching degree scoring data according to a preset time window to obtain category effectiveness evolution data; Step S44: Mine association rules based on the category effectiveness evolution data and the user interaction behavior data to obtain behavior-classification association rule data; Step S45: Quantify the cognitive pattern difference of the adaptive classification decision data based on the behavior-classification association rule data to obtain classification balance quality index data, and perform consistency comparison based on the classification balance quality index data and the adaptive classification decision data to obtain classification quality evaluation parameters; Step S46: Calculate the confidence decay coefficient of each category according to the classification quality evaluation parameters to obtain a clustering category confidence adjustment matrix.
10. The information classification and processing method based on big data according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Apply the clustering category confidence adjustment matrix to the adaptive classification decision data for confidence recalculation to obtain preliminarily adjusted classification data; Step S52: Re-evaluate the classification boundary based on the preliminarily adjusted classification data to obtain boundary-optimized classification data; Step S53: Perform anomaly detection and correction on the boundary-optimized classification data to obtain adjusted classification result data; Step S54: Obtain the adjusted classification result data of multiple time windows, and perform time window weight assignment on the adjusted classification result data of multiple time windows to obtain time window weight data; Step S55: Perform weighted fusion on the adjusted classification result data of multiple time windows according to the time window weight data to obtain fused classification decision data; Step S56: Perform rule correction on the fused classification decision data to obtain rule-optimized classification data, and perform final consistency verification on the fused classification decision data based on the rule-optimized classification data to obtain final information classification decision data.
Citation Information
Cited By
Storage management method and system based on data encryption
CN120724465A