An intelligent algorithm-driven data analysis system
By calculating conditional entropy and information gain in the data analysis system, screening features and building topological backbone networks, optimizing classification structure and boundaries, and adaptively adjusting classification levels, the problem of inefficient data analysis in the existing technology is solved, and the stability and efficiency of classification are improved.
Patent Information
- Application Number
- CN202510275614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing data analysis technologies have inefficient problems in feature screening, topological optimization and classification boundary adjustment, which affects classification accuracy and efficiency.
By calculating the conditional entropy, information gain and information gain deviation values, filtering one-way dominant features, building a topological backbone network, optimizing classification structure and boundaries, and adaptively adjusting the classification hierarchy.
It improves the stability and computing efficiency of data classification, reduces information redundancy and irrelevant feature interference, optimizes the classification structure and hierarchy division, making it more suitable for the topological hierarchical relationship of the data.
Smart Images

Figure CN119782952B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent data processing, and particularly to a data analysis system driven by intelligent algorithms. Background Art
[0002] The technical field of intelligent data processing includes various methods and systems for collecting, storing, analyzing, and processing data using computer technology. The core contents include data acquisition and preprocessing, data storage and management, data calculation and analysis, data mining and pattern recognition, and data visualization. Data acquisition and preprocessing involve the collection, denoising, standardization, and conversion of raw data to ensure data integrity and consistency. Data storage and management mainly involve database technology, distributed storage technology, and data index optimization to support efficient data access and management. Data calculation and analysis cover a variety of calculation methods for parsing, classifying, and correlating data to extract valuable information. Data mining and pattern recognition involve statistical analysis, pattern matching, and anomaly detection of large amounts of data to discover potential data relationships and patterns. Data visualization intuitively presents the data analysis results in a graphical way for easy user understanding and decision-making. Overall, this technical field is dedicated to improving the efficiency and accuracy of data processing through systematic methods to meet the requirements of different application scenarios.
[0003] Among them, an intelligent algorithm-driven data analysis system refers to a system that uses intelligent algorithms to process and analyze data, mainly involving technical matters such as data acquisition, feature extraction, data conversion, classification, and clustering analysis. Data acquisition obtains data from different sources through sensors, interface devices, or data import methods and performs formatting conversion to meet subsequent processing requirements. Feature extraction uses mathematical calculation methods to extract key information from raw data to reduce redundant data and improve analysis efficiency. Data conversion converts data from one representation form to another through specific calculation methods for easy analysis and modeling. Classification and clustering analysis use data calculation methods to classify data according to specific rules to discover the similarity or difference between data, realizing in-depth analysis and structured processing of data.
[0004] In the process of data feature screening by existing data analysis techniques, most rely on traditional correlation calculation methods and fail to effectively evaluate the impact of features on target classification variables, resulting in some features with high correlation but low contribution entering the classification model and affecting the classification accuracy. In the topological structure optimization link, the construction process mainly relies on static similarity metrics between data points and fails to effectively evaluate the role of data points in global information propagation, resulting in a lack of pertinence in the data classification path and affecting the classification efficiency. In the classification optimization process, the classification boundary adjustment method does not combine the information transmission direction, resulting in the class attribution of some data points being easily affected by local features and making the classification rules unstable. The construction of classification levels mostly relies on static rules and fails to adjust according to the actual hierarchical relationship of the data, making it difficult for the hierarchical division method to adapt to the dynamic changes in data distribution, affecting the rationality of the classification levels, resulting in information redundancy in the data classification process, and reducing the adaptability and accuracy of the classification results. Summary of the Invention
[0005] The purpose of the present invention is to solve the drawbacks existing in the prior art and propose a data analysis system driven by an intelligent algorithm.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A data analysis system driven by an intelligent algorithm includes:
[0007] The feature screening module obtains data feature values, calculates conditional entropy, statistically analyzes the degree of feature information fluctuation, calculates the information gain value and information gain deviation value of features with respect to target classification variables, screens one-way dominant features, and obtains an optimal feature set;
[0008] The topological backbone extraction module calculates the topological adjacency relationship of data points based on the optimal feature set, calculates local connectivity, constructs a minimum spanning tree, calculates information transfer centrality, and obtains a topological backbone network;
[0009] The information transfer optimization module, based on the topological backbone network, determines the classification influence degree, statistically analyzes the change in local information entropy, screens classification reference benchmarks and calculates the information contribution degree of weak connection points, and obtains an optimized data classification structure;
[0010] The classification boundary adjustment module calculates the information transfer direction of data points based on the optimized data classification structure, statistically analyzes the change in classification boundary information entropy, screens core data points and calculates the connection strength, adjusts the classification boundary, and outputs the adjusted classification boundary;
[0011] The adaptive classification level module calculates the local density fluctuation based on the adjusted classification boundary, determines the hierarchical relationship, determines the initial data classification level, calculates the change rate of classification stability, adjusts the hierarchical division strategy, and obtains an adaptive data classification hierarchical structure.
[0012] As a further solution of the present invention, the preferred feature set includes information gain value, information gain deviation value, information gain dominant direction, feature priority ranking, and feature contribution degree threshold; the topological backbone network includes topological adjacency relationship, local connectivity, minimum spanning tree, core connection path, information transfer centrality, and transfer threshold screening result; the data optimization classification structure includes information transfer directivity, local information entropy change analysis record, classification reference benchmark, weak connection point information contribution degree, and data category attribution adjustment record; the adjusted classification boundary includes information propagation trend analysis result, consistency judgment result, classification boundary local information entropy, information propagation core data point, core data point connection strength, and classification boundary optimization result; the adaptive data classification hierarchy structure includes local density fluctuation analysis result, initial classification hierarchy, classification stability change rate, classification entropy value change rate, stability threshold, and data hierarchy division strategy.
[0013] As a further solution of the present invention, the feature screening module includes:
[0014] The feature information calculation sub-module calculates the conditional entropy of each feature based on the data features, obtains the information fluctuation degree of each feature under the condition of given other feature values, calculates the information gain value of each feature for the target classification variable, and obtains the feature information gain matrix.
[0015] The information gain evaluation sub-module calculates the information contribution value of the feature to the target feature based on the feature information gain matrix, using the formula:
[0016] ;
[0017] Calculate the feature for the feature information gain deviation value , screen the features with unidirectional dominant effect, and construct the information gain deviation matrix between features, where represents the information gain value of feature under the condition of given feature , represents the information gain value of feature under the condition of given feature ;
[0018] The preferred feature set construction sub-module ranks the features according to the information gain dominant direction based on the information gain deviation matrix between features, eliminates the features with information contribution degree lower than the information gain deviation threshold, and obtains the preferred feature set.
[0019] As a further solution of the present invention, the topological backbone extraction module includes:
[0020] The topological adjacency calculation sub-module calculates the topological adjacency relationship between data points based on the preferred feature set, calculates the adjacency matrix according to the spatial distance and attribute similarity of data points, and uses the formula:
[0021] ;
[0022] Calculate the data point Local connectivity , output the local connectivity data of the data point, where Represents the relationship strength between the data point And the adjacent data point , Represents the data point And the data point Similarity score between, Represents the set of adjacent points of the data point , Represents the data point Total similarity weight of all adjacent points;
[0023] The minimum spanning tree construction sub-module judges the connection stability of data points in the topological network based on the local connectivity data of the data points, and constructs a minimum spanning tree by sorting the connection weights. The connection weights are determined by the adjacency relationship and local connectivity between data points, and the shortest path tree algorithm is used to construct the data core connection path;
[0024] The information transfer screening sub-module calculates the information transfer centrality of data points in the minimum spanning tree according to the data core connection path. The centrality calculation uses the method of weighted average transfer path, and filters out and eliminates the nodes with information transfer volume lower than the transfer threshold based on the calculated centrality value to obtain the topological backbone network.
[0025] As a further solution of the present invention, the information transfer optimization module includes:
[0026] The influence degree calculation sub-module calculates the connection strength between each data point and other data points based on the topological backbone network, and statistically calculates the local information entropy change value, using the formula:
[0027] ;
[0028] Calculate the information influence degree of the data point , obtain the data point information influence degree data, And And Respectively represent the information state values of the data point And , And Are the local information entropy of the corresponding data points, Represents the neighbor set of data points ;
[0029] Based on the data point information influence degree data, the information dominant point screening sub-module screens the data points with information dominant effect, calculates the average influence degree threshold of the data points, and selects the data points with influence degree higher than the average influence degree threshold as information dominant points to obtain the information dominant point set;
[0030] The optimized classification structure generation sub-module calculates the contribution degree of weak connection points according to the information dominant point set, and judges whether the contribution degree is lower than the information entropy change threshold. If it is lower than the threshold, the target data point is attributed to the adjacent high contribution degree data category to obtain the data optimized classification structure.
[0031] As a further solution of the present invention, the classification boundary adjustment module includes:
[0032] The information propagation trend analysis sub-module obtains the information transfer direction of data points based on the data optimized classification structure, calculates the information propagation path of adjacent data points, statistically analyzes the information flow trend between different categories, judges whether the information propagation direction is consistent with the classification boundary direction, and calculates the information propagation trend deviation value;
[0033] The classification boundary information entropy calculation sub-module calculates the information entropy of the local classification boundary based on the information propagation trend deviation value, analyzes the change of the information entropy in the different regions, judges whether there is a local abnormality in the information entropy change, and uses the formula:
[0034] ;
[0035] Calculate the information entropy gradient value , and output the information entropy gradient data, where represents the information entropy of the th local region, represents the information entropy of the previous region, represents the distance between adjacent regions, represents the information propagation trend deviation value, represents the mean value of all deviation values, represents the total number of local regions, represents the number of calculated deviation values;
[0036] The core data point screening and boundary adjustment sub-module screens the information propagation core data points according to the information entropy gradient data, calculates the connection strength between the core data points and the surrounding categories, analyzes the connection degree between different categories, and adjusts the data classification boundary to obtain the adjusted classification boundary.
[0037] As a further solution of the present invention, the adaptive classification hierarchy module includes:
[0038] Based on the adjusted classification boundary, the topological local density fluctuation calculation sub-module uses the formula:
[0039] ;
[0040] Calculate the average density of each node and calculate the degree of fluctuation:
[0041] ;
[0042] Obtain the local density fluctuation data, where represents the local density of node , represents the set of neighbor nodes of node , represents the connection weight between node and node , represents the local density fluctuation value of node , represents the average density of neighbor nodes;
[0043] Based on the local density fluctuation data, the data-level initial classification sub-module determines whether the data has a hierarchical relationship, determines the hierarchical division method of the data, calculates the node clustering centers and their distribution relationships within each level, and performs clustering analysis to generate an initial classification hierarchical structure;
[0044] According to the initial classification hierarchical structure, the classification stability adjustment sub-module calculates the change rate of classification stability between different levels, calculates the change rate of classification entropy value within the target level, and compares it with the stability threshold. If the change rate exceeds the stability threshold, adjust the data division strategy of the target level to obtain an adaptive data classification hierarchical structure.
[0045] Compared with the prior art, the advantages and positive effects of the present invention are:
[0046] In the present invention, by calculating the conditional entropy of data eigenvalues, analyzing the fluctuation degree of feature information and screening one-way dominant features, the interference of low-contribution features is reduced, the topological adjacency relationship of data points is constructed, the backbone of data classification is optimized, the redundant calculation overhead is reduced, the attribution of data points is adjusted in combination with the change of local information entropy, so that the classification structure is more in line with the data transmission characteristics, the information transmission direction of data points is calculated, the classification boundary is optimized by combining the information entropy fluctuation of the classification boundary and the connection strength of core data points, the local density fluctuation of the topological network is calculated, and the data hierarchy division strategy is adjusted in combination with the change rate of hierarchical stability, so that the classification hierarchy is more in line with the topological hierarchical relationship of the data, the interference of irrelevant features is reduced, the classification structure is optimized, the stability and calculation efficiency of data classification are improved, the data calculation redundancy is reduced, the adjustment of the classification boundary is made more accurate, and the classification hierarchy can be adaptively adjusted. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the system flow chart of the present invention;
[0048] Figure 2 is the flow chart of the feature screening module of the present invention;
[0049] Figure 3 is the flow chart of the topological backbone extraction module of the present invention;
[0050] Figure 4 is the flow chart of the information transmission optimization module of the present invention;
[0051] Figure 5 is the flow chart of the classification boundary adjustment module of the present invention;
[0052] Figure 6 is the flow chart of the adaptive classification hierarchy module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0055] Please refer to Figure 1 , an intelligent algorithm-driven data analysis system includes:
[0056] The feature screening module obtains data feature values, calculates the conditional entropy of each data feature, statistically analyzes the information fluctuation degree of each feature under the condition of given other feature values, calculates the information gain value of the data feature for the target classification variable, records the information contribution value of the target feature to other features, calculates the one-versus-two information gain value and the two-versus-one information gain value for features one and two, calculates the information gain deviation value between the two, screens out the features with unidirectional dominant effects, ranks the features according to the information gain dominant direction, eliminates the features with information contribution degree lower than the information gain deviation threshold, and obtains the preferred feature set;
[0057] The topological backbone extraction module calculates the topological adjacency relationship of data points based on the preferred feature set, calculates the local connectivity of data points, judges the connection stability of data points in the topological network, constructs a minimum spanning tree, determines the core connection path of the data, calculates the information transfer centrality of data points in the minimum spanning tree, screens out and eliminates the nodes with information transfer volume lower than the transfer threshold, and obtains the topological backbone network;
[0058] The information transfer optimization module, based on the topological backbone network, judges the degree of influence of the classification of data points by other data points, statistically analyzes the change of local information entropy of data points, screens out the data points with information dominant effects as the classification reference benchmark, calculates the information contribution degree of weak connection points, and if the contribution degree is lower than the information entropy change threshold, then classifies the target data point into the adjacent data category with high contribution degree, and obtains the optimized data classification structure;
[0059] The classification boundary adjustment module calculates the information transfer direction of data points according to the optimized data classification structure, judges the consistency between the classification boundary and the information propagation trend, statistically analyzes the change of local information entropy of the classification boundary, screens out the core data points of information propagation, calculates the connection strength between the core data points and the surrounding categories, adjusts the data classification boundary, and outputs the adjusted classification boundary;
[0060] The adaptive classification hierarchy module, based on the adjusted classification boundary, calculates the local density fluctuation of the topological network, judges whether the data has a hierarchical relationship, dynamically determines the initial classification hierarchy of the data, calculates the classification stability change rate of different hierarchies, and if the change rate of the classification entropy value within the target hierarchy is greater than the stability threshold, then adjusts the data division strategy of the target hierarchy, and obtains the adaptive data classification hierarchy structure.
[0061] The preferred feature set includes information gain value, information gain deviation value, information gain dominant direction, feature priority ranking, and feature contribution threshold; the topological backbone network includes topological adjacency relationship, local connectivity, minimum spanning tree, core connection path, information transfer centrality, and transfer threshold screening result; the data optimization classification structure includes information transfer directivity, local information entropy change analysis record, classification reference benchmark, weak connection point information contribution degree, and data category attribution adjustment record; the adjusted classification boundary includes information propagation trend analysis result, consistency judgment result, local information entropy of the classification boundary, information propagation core data point, core data point connection strength, and classification boundary optimization result; the adaptive data classification hierarchical structure includes local density fluctuation analysis result, initial classification level, classification stability change rate, classification entropy value change rate, stability threshold, and data level division strategy.
[0062] Please refer to Figure 2 , the feature screening module includes:
[0063] The feature information calculation sub-module calculates the conditional entropy of each feature based on the data features, obtains the information fluctuation degree of each feature under the condition of given other feature values, calculates the information gain value of each feature for the target classification variable, and obtains the feature information gain matrix;
[0064] Based on the original data set, each feature variable is extracted and its conditional entropy is calculated to measure the randomness of each feature. The conditional entropy calculation method is to sum based on the joint probability distribution of different feature values. For example, if a certain feature has possible values, then the conditional entropy calculation is:
[0065] ;
[0066] Among them, represents the number of possible values, that is, the number of categories of the target classification variable , represents the number of possible values of feature , that is, the number of different states of feature , is the target variable, is the case where feature takes value , represents the probability distribution of under the value condition.
[0067] Table 1.1 Example table of conditional entropy calculation:
[0068] ;
[0069] As shown in Table 1.1, the conditional entropy values of each calculated feature measure the randomness of the feature. Further calculate the information gain of each feature:
[0070] ;
[0071] Among them, is the entropy of the target variable , and the calculation formula is:
[0072] ;
[0073] The calculated information gain matrix is as follows:
[0074] Table 1.2 Information gain calculation results:
[0075] ;
[0076] As shown in Table 1.2, the larger the information gain value, the greater the contribution of the feature to the target variable, and finally the feature information gain matrix is obtained.
[0077] The information gain evaluation sub-module calculates the information contribution value of the feature to the target feature based on the feature information gain matrix, using the formula:
[0078] ;
[0079] Calculate the information gain deviation value of feature to feature , screen out the features with one-way dominant effects, and construct the information gain deviation matrix between features. Among them, represents the information gain value of feature under the given condition of feature , and represents the information gain value of feature under the given condition of feature ;
[0080] According to the feature information gain matrix, calculate the information contribution value of feature to feature . The information contribution value is an index to measure the degree of information provided by one feature to other features. In order to measure the information dominant direction between two features, calculate the one-to-two information gain and the two-to-one information gain , and further calculate the information gain deviation.
[0081] Set the information gains of feature and feature as follows:
[0082] ;
[0083] ;
[0084] Substitute into the formula for calculation:
[0085] ;
[0086] Thus, the information gain deviation matrix between features can be obtained.
[0087] Table 1.3 Information gain deviation matrix between features:
[0088] ;
[0089] As shown in Table 1.3, the information gain deviation is used to screen for features with one-way dominant effects. If the deviation value of a certain feature from other features exceeds the set threshold, then this feature is regarded as the dominant feature contributing information.
[0090] The preferred feature set construction sub-module ranks the features according to the information gain dominant direction based on the information gain deviation matrix between features, eliminates the features with information contribution degrees lower than the information gain deviation threshold, and obtains the preferred feature set;
[0091] According to the information gain deviation matrix between features, rank the features according to the information gain dominant direction, and eliminate the features with information contribution degrees lower than the information gain deviation threshold. The setting of the information gain deviation threshold is based on the entropy distribution of the target variable and the average information gain between features. The specific value calculation is as follows:
[0092] The average information gain between features is:
[0093] ;
[0094] The information gain deviation threshold takes 65% of the average information gain between features, that is:
[0095] ;
[0096] The basis for this setting is that when there are large differences in information gain deviations between features, it often leads to an increase in redundant features. Therefore, the threshold is set to 65% of the average information gain, that is, retain the part of the features with higher information gain contributions and discard the features with smaller contributions, making the feature screening more representative.
[0097] According to the calculated threshold of 0.18, screen the preferred feature set from Table 1.3:
[0098] The deviation of feature A - B is 0.20, higher than the threshold, retain feature A;
[0099] The deviation of feature B - C is 0.15, lower than the threshold, eliminate feature B;
[0100] The finally obtained optimal feature set is {A}.
[0101] Please refer to Figure 3 , the topological backbone extraction module includes:
[0102] Based on the optimal feature set, the topological adjacency calculation sub-module calculates the topological adjacency relationship between data points, calculates the adjacency matrix according to the spatial distance and attribute similarity of data points, and uses the formula:
[0103] ;
[0104] Calculate the local connectivity of the data point , and output the local connectivity data of the data point. Among them, represents the relationship strength between the data point and the adjacent data point , represents the similarity score between the data point and the data point , represents the set of adjacent points of the data point , represents the total similarity weight of all adjacent points of the data point ;
[0105] Based on the optimal feature set, first, the feature information of each data point, such as spatial coordinates, attribute similarity, etc., needs to be extracted from the dataset. For example, in a Geographic Information System (GIS), each data point can represent a city, and its feature information includes key indicators such as longitude and latitude, population density, GDP, etc. Then, the topological adjacency relationship between data points is calculated, which requires constructing an adjacency matrix, where each element represents the relationship strength between the data point and the data point , and this strength is determined by the inverse of the spatial distance, that is:
[0106] ;
[0107] Among them, is the Euclidean distance between the data point and . For example, assume there are two data points A(2,3) and B(5,7), and their distance is calculated as follows:
[0108] ;
[0109] Substitute into the weight calculation formula:
[0110] ;
[0111] After calculating the entire data set, the adjacency matrix can be obtained. Next, the local connectivity of data points is calculated through local density estimation, that is:
[0112] ;
[0113] where represents the similarity score between data point and data point , which can be calculated by the cosine similarity between attribute features. For example:
[0114] ;
[0115] Suppose there are data points A and B, and their attribute vectors are , , then:
[0116] ;
[0117] After calculating the adjacency weights and similarity scores of all points, substitute them into the final calculation of the local connectivity of each data point, as shown in Table 2.1:
[0118] Table 2.1 Calculation Table of Data Point Local Connectivity:
[0119] ;
[0120] As shown in Table 2.1, the local connectivity of data points can be obtained through calculation.
[0121] The minimum spanning tree construction sub-module judges the connection stability of data points in the topological network based on the data of data point local connectivity, and constructs a minimum spanning tree through the sorting of connection weights. The connection weights are determined by the adjacency relationship and local connectivity between data points, and the shortest path tree algorithm is used to construct the data core connection path;
[0122] According to the local connectivity of data points, first judge the connection stability of data points in the topological network based on the sorting of local connectivity. For the connection stability in the topological network, it can be calculated by the weight assignment method, that is, if data point is connected to data point , its connection stability can be calculated as:
[0123] ;
[0124] For example, the local connectivities of data points A and B are , , respectively, then:
[0125] ;
[0126] After calculating all possible connections, sort them in descending order of weight values and construct a minimum spanning tree. The minimum spanning tree uses the Kruskal or Prim algorithm to connect data points in order of weight priority and gradually form a data core connection path, as shown in Table 2.2.
[0127] Table 2.2 Core connection path of the spanning tree:
[0128] ;
[0129] As shown in Table 2.2, the data core connection path is calculated.
[0130] The information transfer screening sub-module calculates the information transfer centrality of data points in the minimum spanning tree according to the data core connection path. The centrality calculation uses the method of weighted average transfer path, and filters out and eliminates the nodes with information transfer volume lower than the transfer threshold based on the calculated centrality value to obtain the topological backbone network;
[0131] Based on the data core connection path, calculate the information transfer centrality of data points in the minimum spanning tree. The information transfer centrality can be calculated by the weighted path:
[0132] ;
[0133] Among them, is the transfer centrality of data point , represents all reachable paths of data point . For example, assume that data point A is connected to B, and B is connected to C, and the path weights are as follows:
[0134] ;
[0135] ;
[0136] ;
[0137] Then the transfer centrality of A is:
[0138] ;
[0139] Similarly, calculate the information transfer centrality of all data points, as shown in Table 2.3:
[0140] Table 2.3 Information transfer centrality of data points:
[0141] ;
[0142] The setting basis of the information transfer threshold lies in the average transfer centrality of data points and its standard deviation, ensuring that the selected data points can effectively undertake the information transmission task in the network. When setting the information transfer threshold, first calculate the average value of the transfer centrality of all data points and the standard deviation , and determine the screening threshold according to the data distribution.
[0143] Suppose the transfer centrality distribution of data points is as follows:
[0144] ;
[0145] ;
[0146] ;
[0147] ;
[0148] ;
[0149] Calculate the average value:
[0150] ;
[0151] Calculate the standard deviation:
[0152] ;
[0153] ;
[0154] ;
[0155] The information transfer threshold is set to , that is:
[0156] ;
[0157] According to the calculation results, the information transfer threshold is finally set to 0.35 , and the nodes below this threshold will be excluded. For example, data point A is excluded from the topological network due to . Finally, data points B and C are retained to form the final topological backbone network.
[0158] Please refer to Figure 4 , the information transfer optimization module includes:
[0159] The influence degree calculation sub-module, based on the topological backbone network, calculates the connection strength between each data point and other data points, counts the change value of local information entropy, and uses the formula:
[0160] ;
[0161] Calculate the information influence degree of data points to obtain the data of the information influence degree of data points , and respectively represent the information status values of data points and , and are the local information entropy of the corresponding data points represents the neighbor set of the data point ;
[0162] Based on the topological backbone network, first obtain the data point set , where each data point has a specific location information and an initial information status value , and define the connection relationship between data points , where represents the connection between the data points and . The strength of the connection is determined by the information flow characteristics. If there is information transfer between two data points, then define its connection weight , and its value can be calculated according to the communication frequency, similarity or shared information volume. For each data point , calculate its local information entropy , which can be calculated by the following formula
[0163] ;
[0164] where represents the probability of the data point in the category . This value can be estimated through historical data or the distribution of similar data points. For example, for a data point , if it may belong to the categories , then its probability can be calculated as according to the training data, and then calculate the information entropy
[0165] ;
[0166] After calculating the information entropy of all data points, it is necessary to evaluate the influence degree between data points. Using the influence degree calculation formula, after calculating the value of the data point , the information influence degree of the data point can be obtained. In practical applications, for example, for the analysis of social network information dissemination, if the status information of a certain user is , and the statuses of its neighbor users are respectively , and their information entropies are respectively , then calculate its influence degree:
[0167] ;
[0168] Finally, obtain the information influence degree of data points for subsequent screening processes.
[0169] The information dominant point screening sub-module screens data points with information dominant effects based on the data of the information influence degree of data points, calculates the average influence degree threshold of data points, and selects data points with influence degrees higher than the average influence degree threshold as information dominant points to obtain an information dominant point set;
[0170] Based on the information influence degree of data points, first calculate the average influence degree threshold of data points . Suppose there are five data points, and their influence degrees are respectively , then:
[0171] ;
[0172] Screen out data points with influence degrees higher than as information dominant points. That is, in this example, data points and are selected as information dominant points to obtain an information dominant point set.
[0173] The optimized classification structure generation sub-module calculates the contribution degree of weak connection points according to the information dominant point set, and judges whether the contribution degree is lower than the information entropy change threshold. If it is lower than the threshold, the target data point is classified into the adjacent high contribution degree data category to obtain an optimized data classification structure;
[0174] Call the information dominant point set and calculate the contribution degree of weak connection points :
[0175] ;
[0176] Among them, represents the connection weight between data point and its neighbor , is the influence degree of the neighbor data point. The weight is set based on the information transmission intensity between data point and . This value is calculated from historical data traffic, interaction frequency, and similarity, and usually adopts normalization processing to make the weight value between . The more frequent the information exchange between data points, the closer the weight is to 1. For points with less information exchange, the weight approaches the lower range of 0.1 - 0.2. For example, assume a certain data point and its adjacent data points and if there are 120 interactions with within the past 24 hours, and 80 interactions with , then it can be set as .
[0177] Obtained from the above , then calculate:
[0178] ;
[0179] If the contribution degree is lower than the information entropy change threshold , then this data point belongs to the adjacent high - contribution data category. The threshold is set with reference to the global information entropy distribution. Usually, the mean value of the information entropy of all data points is selected as the reference value, and an adjustable scaling factor is used for fine - tuning, that is:
[0180] ;
[0181] Among them, reflects the strictness of information screening. Usually take to adapt to the discreteness of different data sets. Suppose the mean value of information entropy in a certain data set is 0.65. If is selected, then there is . In this case, , so belongs to the category where the adjacent data point with the highest contribution degree is located, and finally the optimized data classification structure is obtained.
[0182] Please refer to Figure 5 , the classification boundary adjustment module includes:
[0183] The information propagation trend analysis sub - module obtains the information transfer direction of data points based on the optimized data classification structure, calculates the information propagation path of adjacent data points, statistically analyzes the information flow trend between different categories, determines whether the information propagation direction is consistent with the classification boundary direction, and calculates the information propagation trend deviation value;
[0184] Obtain the information transfer direction of data points. Based on historical data and real-time data, analyze the information transfer paths between data points. For example, in urban traffic flow analysis, the real-time traffic flow of each traffic intersection can be obtained, the traffic flow migration between each intersection can be calculated, the information dissemination direction can be statistically analyzed, and an information dissemination relationship matrix between data points can be established for further analysis of the classification boundary, calculate the information transfer paths between adjacent data points, and set the transfer weights between data points. The weight value is set based on the information transfer rate and propagation stability between data points. The flow rate change rate between data points in the past 10 minutes is used as the basic weight calculation item. If the flow rate change rate exceeds the set threshold (such as greater than 50 vehicles / minute), the weight value is increased, otherwise it is decreased. The calculation method is as follows:
[0185] ;
[0186] Among them, represents the transfer weight between data point and . is the data transfer rate from to is the global maximum data transfer rate, is the weight scaling coefficient, set to 2, so that the weight value is distributed in the interval [0, 2]. For example, if the average flow rate change rate between two intersections in the past 10 minutes is 40 vehicles / minute and the global maximum flow rate is 80 vehicles / minute, then
[0187] ;
[0188] Statistically analyze the information flow trends between different categories. By calculating the information flow ratios within and outside each category, determine whether the information dissemination direction is consistent with the classification boundary direction. If the information flow ratio is within a specific range (such as 0.8 - 1.2), it is determined that the information dissemination trend is stable, otherwise it is determined that there may be a problem with the classification boundary. Calculate the information dissemination trend deviation value. The threshold of this deviation value is set to 0.2, which is obtained from the mean calculation of historical data and set after variance analysis of the information dissemination trends in multiple time periods, so that the deviation trend of data points is controllable. If the deviation value of a certain data point is higher than this threshold (such as 0.3), it indicates that there is an abnormality in its classification boundary and further adjustment is required.
[0189] The classification boundary information entropy calculation sub-module calculates the information entropy of the local classification boundary based on the information dissemination trend deviation value, analyzes the change of the information entropy in the difference region, determines whether there is a local abnormality in the information entropy change, and uses the formula:
[0190] ;
[0191] Calculate the information entropy gradient value , and output the information entropy gradient data, where represents the information entropy of the th local region, represents the information entropy of the previous region, represents the distance between adjacent regions, represents the information propagation trend deviation value, represents the mean value of all deviation values, represents the total number of local regions, represents the number of calculated deviation values;
[0192] Based on the information propagation trend deviation value, calculate the information entropy of the local classification boundary. First, set the information entropy calculation range of the data points. For example, in the manufacturing data classification, select the sensor data of adjacent regions as the local range, and calculate the information entropy of each category within the region. If there are three abnormal states (low temperature, high pressure, vibration) in a production region, the information entropy calculation is as follows:
[0193] ;
[0194] where is the occurrence probability of each category. For example, if the proportions of low temperature, high pressure, and vibration are 0.4, 0.3, and 0.3 respectively, then:
[0195]
[0196] Furthermore, analyze the changes in the information entropy of different regions, calculate the local information entropy gradient using difference, and obtain the classification boundary information entropy gradient value through formula operation. For example, if the continuous information entropy of a region is , and the corresponding region spacing is 1 for all, then calculate:
[0197] ;
[0198] The obtained information entropy gradient value is 0.18, and analyze the local information entropy change trend, which can be used to adjust the classification boundary.
[0199] The core data point screening and boundary adjustment sub-module screens the core data points of information propagation according to the information entropy gradient data, calculates the connection strength between the core data points and the surrounding categories, analyzes the connection degree between different categories, adjusts the data classification boundary, and obtains the adjusted classification boundary;
[0200] According to the changing trend of local information entropy, core data points for information dissemination are screened. The screening criteria can be based on the peak region of the local information entropy gradient. For example, when the information entropy change exceeds 0.15, the core data points in this region are identified, and the connection strength between the core data points and surrounding categories is calculated. For example, the distribution of each core data point in different categories is statistically analyzed, and its belonging degree between categories is calculated. For example, if the connection degree of a certain core data point in category A is 0.7 and in category B is 0.3, then its belonging category is category A. If the connection strength is less than 0.6, it is determined as a boundary point and classification adjustment is required. Analyze the connection degree between different categories. The connection strength is calculated using the normalized category density ratio, and its calculation formula is:
[0201] ;
[0202] Among them, represents the connection strength of the data point in a certain category, is the number of neighbors of this data point in this category, is the total number of neighbors. The connection strength threshold is set to 0.6, and this value is obtained from the statistical distribution density of boundary points in the classification dataset, enabling reasonable classification of data points between low-density categories. For example, if the number of neighbors of a certain data point in category A is 15 and the total number of neighbors is 25, then its connection strength is:
[0203] ;
[0204] exactly reaches the classification threshold. If it is lower than 0.6, such as 0.55, then it is determined that this point belongs to the boundary point and the classification needs to be adjusted. If the transfer ratio of core data points in a certain category to other categories exceeds 30%, then boundary adjustment is performed, and finally the adjusted classification boundary is obtained.
[0205] Please refer to Figure 6 , the adaptive classification hierarchy module includes:
[0206] The topological local density fluctuation calculation sub-module, based on the adjusted classification boundary, uses the formula:
[0207] ;
[0208] Calculate the average density of each node and calculate the degree of fluctuation:
[0209] ;
[0210] Obtain the local density fluctuation data. Among them, represents the local density of node , represents the set of neighbor nodes of node , Representative node The connection weight between the node and Representative node The local density fluctuation value of represents the average density of neighbor nodes;
[0211] Based on the adjusted classification boundary, the local density of each node in the topological network needs to be calculated. First, the topological network consists of a series of nodes and their adjacent nodes Each node is connected by an edge The connection weight represents the interaction strength between two nodes. Then, calculate the standard deviation of the density of each node , which is used to measure the degree of local density fluctuation of the node. Assume that the local densities of adjacent nodes are 3.2, 2.8, 4.1, and 3.5 respectively, then their average value is:
[0212] ;
[0213] Substitute into the formula to calculate the degree of fluctuation:
[0214] ;
[0215] When exceeds the set threshold (such as 0.5), it indicates that the local density of this node fluctuates greatly, and it may belong to an abnormal node or the junction of levels in the topological structure, thus affecting the classification hierarchy structure. Calculate the local density fluctuation value.
[0216] The initial classification sub-module of the data level judges whether the data has a hierarchical relationship based on the local density fluctuation data, determines the hierarchical division method of the data, calculates the node clustering centers and their distribution relationships within each level, and conducts clustering analysis to generate the initial classification hierarchy structure;
[0217] Based on the local density fluctuation data, judge whether the data has a hierarchical relationship. When performing hierarchical classification, the threshold method can be used, that is, set the threshold of the local density fluctuation value to judge whether it is necessary to divide levels. For example, set . If the of a certain node, then it is considered to be at the junction of levels and needs to be used as a new level node to guide the subsequent classification strategy. The setting of this threshold is based on the overall density fluctuation range in the network topology. Usually, take the average value of the local density fluctuations of all nodes plus 10% of the standard deviation as this value to ensure the stability of the level division. Assume that the average value of the local density fluctuations of 100 nodes in a certain network is 0.35 and the standard deviation is 0.08, then calculate As follows:
[0218] ;
[0219] When taking the final value, round up to two decimal places, so set .
[0220] In the classification hierarchy division, the classification of nodes is based on topological characteristics. Calculate the clustering centers of nodes within each level and their distribution relationships. Use the mean method to determine the clustering center. For example, assume that a certain level contains 5 nodes with densities of 3.2, 3.5, 3.8, 4.0, and 4.3 respectively. Then the central density of this level can be calculated as follows:
[0221] ;
[0222] To avoid the influence of outliers on the clustering center density, it is necessary to set the weight coefficient to adjust the contribution of each node to the clustering center. This weight coefficient depends on the local density fluctuation value of the node. The calculation method of the weight is:
[0223] ;
[0224] Among them, is the maximum density fluctuation value within the current level. If within a certain level the value range is , then the maximum value , then the weight of a certain node is calculated as follows:
[0225] ;
[0226] Finally, calculate the central density of this level through weighted calculation:
[0227] ;
[0228] According to the density fluctuation value and the clustering center density, further adjust the classification level. For example, if the density of a certain node is more than 15% higher than the clustering center density of the level, then its classification level needs to be adjusted to obtain the initial classification level structure.
[0229] The classification stability adjustment sub-module calculates the change rate of classification stability between different levels based on the initial classification level structure, calculates the change rate of classification entropy value within the target level, and compares it with the stability threshold. If the change rate exceeds the stability threshold, then adjust the data division strategy of the target level to obtain the adaptive data classification level structure;
[0230] Call the initial classification level structure, calculate the change rate of classification stability between different levels, and set the classification entropy value To measure the stability of the hierarchy, the calculation formula is as follows:
[0231] ;
[0232] Among them, represents the proportion of data points in a certain classification. Suppose a hierarchy contains 50 data points, 30 of which belong to classification A and 20 belong to classification B, then the entropy value is calculated as follows:
[0233] ;
[0234] Calculate the change rate of the classification entropy value within the target hierarchy:
[0235] ;
[0236] Suppose the entropy value before adjustment is 0.2923 and the entropy value after adjustment is 0.2500, then the change rate is calculated as follows:
[0237] ;
[0238] The change rate of the entropy value The threshold is set to 0.1. The selection of this threshold is based on the statistical analysis of multiple groups of classification hierarchy data. The setting method is to calculate the change rate of the entropy value of 100 groups of different classification hierarchies and take the value at the 90th percentile as the stability threshold. Suppose the average value of the change rate of the entropy value of 100 groups of data is 0.085 and the standard deviation is 0.02, then the threshold is calculated as follows:
[0239] ;
[0240] When taking the final value, round up to two decimal places, so the stability threshold is set. If the change rate of the entropy value , then the hierarchical structure is unstable and the data partitioning strategy needs to be adjusted.
[0241] If the change rate of the entropy value is greater than the set stability threshold, then the data partitioning strategy of the target hierarchy needs to be adjusted. For example, some data points are redistributed to other hierarchies to optimize the classification stability. The redistribution of this data point is based on its information entropy contribution degree. Set the contribution degree weight The calculation formula is as follows:
[0242] ;
[0243] Suppose the information entropy contribution degree of 30 data points in classification A is calculated as follows:
[0244] ;
[0245] Similarly, calculate the contribution of Category B:
[0246] ;
[0247] If the contribution of a certain category exceeds 0.5, it indicates that the influence weight of this category is large. Then it can be used as the key object for hierarchical stability adjustment. By reducing the number of its data points or adjusting the data point distribution, optimize the hierarchical division strategy, and finally obtain an adaptive data classification hierarchical structure.
[0248] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A data analysis system driven by an intelligent algorithm, characterized in that: The system comprises: The feature screening module obtains data feature values, calculates conditional entropy, counts the fluctuation degree of feature information, calculates the information gain value and information gain deviation value of the feature for the target classification variable, screens unidirectional dominant features, and obtains the optimal feature set; The topological backbone extraction module extracts the characteristic information of each data point based on the preferred feature set, each data point represents a city, and its characteristic information includes latitude and longitude, population density, and GDP key indicators, calculates the topological adjacency relationship of the data points, calculates the local connectivity, constructs a minimum spanning tree, calculates the information transfer centrality, and obtains the topological backbone network; The information transmission optimization module determines the classification influence degree based on the topological backbone network, counts the change of local information entropy, selects the classification reference benchmark and calculates the information contribution of weak connection points, and obtains data to optimize the classification structure; The classification boundary adjustment module calculates the direction of data point information transmission based on the data optimization classification structure, counts the change of classification boundary information entropy, selects core data points and calculates the connection strength, adjusts the classification boundary, and outputs the adjusted classification boundary; The adaptive classification hierarchy module calculates local density fluctuations based on the adjusted classification boundaries, determines hierarchical relationships, determines the initial classification hierarchy of the data, calculates the classification stability change rate, adjusts the hierarchy division strategy, and obtains an adaptive data classification hierarchy structure.
2. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The preferred feature set includes information gain value, information gain deviation value, information gain dominant direction, feature priority ranking, and feature contribution threshold; the topological backbone network includes topological adjacency, local connectivity, minimum spanning tree, core connection path, information transmission centrality, and transmission threshold screening results; the data optimization classification structure includes information transmission directionality, local information entropy change analysis records, classification reference benchmarks, weak connection point information contribution, and data category attribution adjustment records; the adjusted classification boundaries include information propagation trend analysis results, consistency judgment results, local information entropy of classification boundaries, information propagation core data points, core data point connection strength, and classification boundary optimization results; the adaptive data classification hierarchical structure includes local density fluctuation analysis results, initial classification levels, classification stability change rate, classification entropy value change rate, stability threshold, and data hierarchy division strategy.
3. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The feature screening module includes: The feature information calculation submodule calculates the conditional entropy of each feature based on the data features, obtains the information fluctuation degree of each feature under the given other feature values, calculates the information gain value of each feature for the target classification variable, and obtains the feature information gain matrix; The information gain evaluation submodule calculates the information contribution value of the feature to the target feature based on the feature information gain matrix, using the formula: ; Calculate features Features The information gain deviation value , filter out the unidirectional dominant features and construct the information gain deviation matrix between features, where Representative features In Features The information gain value in a given situation is, Representative features In Features The information gain value for a given situation; The preferred feature set construction submodule prioritizes the features according to the information gain dominant direction based on the information gain deviation matrix between the features, eliminates the features whose information contribution is lower than the information gain deviation threshold, and obtains the preferred feature set.
4. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The topology backbone extraction module comprises: The topological adjacency calculation submodule calculates the topological adjacency relationship between data points based on the preferred feature set, and calculates the adjacency matrix according to the spatial distance and attribute similarity of the data points, using the formula: ; Calculate data points The local connectivity , output the local connectivity data of the data point, where Representative data points and adjacent data points The strength of the relationship between Representative data points With data points The similarity score between Representative data points The set of adjacent points of Representative data points The total similarity weight of all neighboring points; The minimum spanning tree construction submodule determines the connection stability of the data points in the topological network based on the local connectivity data of the data points, and constructs a minimum spanning tree by sorting the connection weights. The connection weights are determined by the adjacency relationship and local connectivity between the data points, and the shortest path tree algorithm is used to construct the data core connection path; The information transmission screening submodule calculates the information transmission centrality of the data point in the minimum spanning tree according to the data core connection path. The centrality calculation adopts the weighted average transmission path method. The nodes with information transmission amount lower than the transmission threshold are screened and eliminated according to the calculated centrality value to obtain the topological backbone network.
5. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The information transmission optimization module includes: The influence calculation submodule calculates the connection strength of each data point with other data points based on the topological backbone network, and counts the change value of local information entropy using the formula: ; Calculate data points Information influence , obtain the data point information influence data, and Represents data points and The information status value of and is the local information entropy of the corresponding data point, Represents data points The set of neighbors of ; The information leading point screening submodule screens the information leading data points based on the information influence data of the data points, calculates the average influence threshold of the data points, and selects the data points with influence higher than the average influence threshold as the information leading points, and obtains the information leading point set; The optimized classification structure generation submodule calculates the contribution of the weak connection point according to the information dominant point set, and determines whether the contribution is lower than the information entropy change threshold. If it is lower than the information entropy change threshold, the target data point is classified into the adjacent high-contribution data category to obtain the data optimized classification structure.
6. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The classification boundary adjustment module includes: The information propagation trend analysis submodule obtains the information transmission direction of the data point based on the data optimization classification structure, calculates the information propagation path of adjacent data points, counts the information flow trend between different categories, determines whether the information propagation direction is consistent with the classification boundary direction, and calculates the information propagation trend deviation value; The classification boundary information entropy calculation submodule calculates the information entropy of the local classification boundary based on the information propagation trend deviation value, analyzes the change of information entropy in the difference area, determines whether there is a local anomaly in the information entropy change, and uses the formula: ; Calculate the information entropy gradient value , and output information entropy gradient data, where Representative The information entropy of a local area, represents the information entropy of the previous region, represents the distance between adjacent regions, Represents the deviation value of information dissemination trend, represents the mean of all deviation values, represents the total number of local areas, Represents the number of deviation values calculated; The core data point screening and boundary adjustment submodule screens the information propagation core data points according to the information entropy gradient data, calculates the connection strength between the core data points and the surrounding categories, analyzes the degree of connection between the difference categories, adjusts the data classification boundary, and obtains the adjusted classification boundary.
7. The data analysis system driven by the intelligent algorithm according to claim 1, characterized in that: The adaptive classification hierarchy module comprises: The topological local density fluctuation calculation submodule is based on the adjusted classification boundary and adopts the formula: ; Calculate the average density of each node and calculate the degree of fluctuation: ; Get local density fluctuation data, where Representative Node The local density of Representative Node The set of neighbor nodes of Representative Node With Node The connection weights between Representative Node The local density fluctuation value, Represents the average density of neighbor nodes; The data level initial classification submodule determines whether the data has a hierarchical relationship based on the local density fluctuation data, determines the hierarchical division method of the data, calculates the node cluster centers and their distribution relationships in each level, and performs cluster analysis to generate an initial classification hierarchical structure; The classification stability adjustment submodule calculates the classification stability change rate between different levels and the classification entropy change rate within the target level according to the initial classification hierarchy structure, and compares them with the stability threshold. If the change rate exceeds the stability threshold, the data partitioning strategy of the target level is adjusted to obtain an adaptive data classification hierarchy structure.
Citation Information
Patent Citations
Distributed estimation with adaptive clustering strategy based on element-wise distance over multitask networks.
AU2020103334A4
Lithology intelligent identification method and system based on deep learning
CN118941843A