A multi-dimensional data processing method

By clustering and prioritizing multi-dimensional data, the problem of information loss caused by direct dimensionality reduction of monitoring data from different locations is solved, achieving more accurate dimensionality reduction and monitoring results.

CN120578978BActive Publication Date: 2025-12-02光谷技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511078070.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-12-02
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

In existing technologies, when processing multi-dimensional data, the level of representation of target information by monitoring data from different locations varies. Direct dimensionality reduction analysis may lead to the loss or confusion of important information, resulting in poor monitoring performance.

Method used

By acquiring category sensor data for each location in the monitored target area within a historical time range, cluster analysis is performed to obtain data clusters. The target state range factor, data change sensitivity, and information representation parameters are calculated. Dimensionality reduction is then performed based on retention priorities using the PCA algorithm.

Benefits of technology

It improves the accuracy of dimensionality reduction analysis, retains important information, reduces data loss, and enhances monitoring effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578978B_ABST
    Figure CN120578978B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data dimensionality reduction technology, specifically to a multi-dimensional data processing method. The invention obtains the target state range factor for each dimensional data point based on the morphological characteristics of the data cluster to which it belongs and the number of corresponding dimensional data points of the same category; it obtains the target data change sensitivity of each data cluster by combining the changing trend of the monitoring time-series data corresponding to each dimensional data point; it obtains the target state information representation parameter for each dimensional data point by combining the changing distribution of the monitoring time-series data corresponding to each dimensional data point; it then obtains the retention priority of each dimensional data point; it obtains a dimensional data matrix composed of all monitoring time-series data; and it performs dimensionality reduction processing on the dimensional data matrix to obtain the dimensionality-reduced data. This invention improves the accuracy of dimensionality reduction by obtaining the accurate retention priority of dimensional data points at each location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data dimensionality reduction technology, specifically to a multi-dimensional data processing method. Background Technology

[0002] With the development of sensor technology, the speed and quantity of data generation are growing exponentially, and data structures are becoming increasingly complex. In recent years, the information technology and data acquisition capabilities of the Internet of Things (IoT) have gradually improved, and multi-dimensional data processing has become increasingly important in various fields, helping to extract more accurate and valuable information from the data. However, considering that the amount of data to be monitored during feature analysis is often large and the data analysis efficiency is poor, current technologies use existing dimensionality reduction methods to reduce the dimensionality of data monitored by sensors at multiple locations in the target area. However, because the changes in the target state reflected by multiple types of sensors may be similar during long-term monitoring, the different levels of representation of target information by monitoring data at different locations are not taken into account. Directly performing dimensionality reduction analysis may lead to the loss or confusion of some important information, resulting in poor monitoring results. Summary of the Invention

[0003] To address the technical problem that monitoring data from different locations may represent target information at varying levels, and that direct dimensionality reduction analysis could lead to the loss or obfuscation of certain important information, this invention aims to provide a multi-dimensional data processing method. The specific technical solution adopted is as follows:

[0004] This invention proposes a multi-dimensional data processing method, the method comprising:

[0005] Acquire the monitoring time series data of each type of sensor at each location in the target area within the historical time range to form a dimensional data point;

[0006] Based on the location characteristics of each dimension's data points, all dimension's data points are clustered to obtain multiple data clusters; based on the morphological characteristics of the data cluster to which each dimension's data point belongs and the number of dimension's data points of the same category, the target state range factor for each dimension's data point is obtained.

[0007] Based on the target state range factor corresponding to each category and each dimension data point within each data cluster, and the changing trend of the monitoring time series data corresponding to each dimension data point, the target data change sensitivity of each data cluster is obtained; based on the target data change sensitivity of the data cluster to which each dimension data point belongs, and the change distribution of the monitoring time series data corresponding to each dimension data point, the target state information representation parameters of each dimension data point are obtained.

[0008] Based on the target state information representation parameters and the target state range factor for each dimension data point, the retention priority of each dimension data point is obtained;

[0009] Obtain a dimensional data matrix composed of all monitoring time-series data; perform dimensionality reduction processing according to the retention priority of each dimensional data point and the dimensional data matrix to obtain dimensionality-reduced data of the dimensional data matrix.

[0010] Furthermore, the method for obtaining the target state range factor includes:

[0011] The coordinate region of each data cluster is obtained using the convex hull detection algorithm;

[0012] Based on the number of dimensional data points of the same category in the data cluster where each dimensional data point is located, and the range of the corresponding coordinate region, the target state range factor for each dimensional data point is obtained. The number of dimensional data points of the same category is negatively correlated with the target state range factor, and the range of the coordinate region is positively correlated with the target state range factor.

[0013] Furthermore, the method for obtaining the sensitivity to changes in the target data includes:

[0014] For any category within each data cluster, obtain the fluctuation level of the monitoring time series data corresponding to each dimension data point as the local fluctuation level; obtain the mean of the local fluctuation levels of all dimension data points as the overall fluctuation level.

[0015] Based on the target state range factor of each dimension data point and the difference between the local fluctuation degree and the overall fluctuation degree of each dimension data point, the local sensitivity of each dimension data point is obtained. The target state range factor is negatively correlated with the local sensitivity, and the difference between the local fluctuation degree and the overall fluctuation degree is positively correlated with the local sensitivity.

[0016] The mean of the local sensitivity of all data points in any category is obtained as the category sensitivity; the mean of the category sensitivity within each data cluster is obtained as the target data change sensitivity of each data cluster.

[0017] Furthermore, the method for obtaining the target state information representation parameters includes:

[0018] Based on the change distribution of the monitoring time series data corresponding to each dimension data point, the change information of each dimension data point is obtained;

[0019] The change information is weighted according to the sensitivity of the target data change to the data cluster to which each dimension data point belongs, so as to obtain the target state information representation parameters of each dimension data point.

[0020] Furthermore, the method for obtaining the amount of change information includes:

[0021] Obtain the fitting curve of the monitoring time series data corresponding to each dimension data point;

[0022] The amount of change information is obtained by considering the degree of disorder and the number of monotonic trend intervals of the rate of change for each data point on the corresponding fitted curve. Both the degree of disorder and the number of monotonic trend intervals are positively correlated with the amount of change information.

[0023] Furthermore, the method for obtaining the retention priority includes:

[0024] Calculate the product of the target state information representation parameter and the corresponding target state range factor for each dimension data point, and normalize it to obtain the retention priority of each dimension data point.

[0025] Furthermore, the method for obtaining dimensionality reduction data from the dimensionality data matrix includes:

[0026] During the PCA dimensionality reduction algorithm on the dimensional data matrix, a standardized matrix of the dimensional data matrix is ​​obtained, and the standardized matrix is ​​adjusted according to the retention priority of each dimensional data point to obtain an adjusted matrix;

[0027] Obtain the principal direction matrix of the adjustment matrix; calculate the product of the principal direction matrix and the dimension data matrix to obtain the dimensionality-reduced data of the dimension data matrix.

[0028] Furthermore, the method for obtaining the adjustment matrix includes:

[0029] The monitoring time series data of the corresponding dimension data points in the standardized matrix are weighted according to the retention priority of each dimension data point to obtain the adjustment matrix.

[0030] Furthermore, the method for obtaining the data clusters includes:

[0031] Based on the location characteristics of data points in each dimension, the K-means clustering algorithm is applied to all data points in all dimensions to obtain multiple data clusters.

[0032] Furthermore, the method for obtaining the dimensional data matrix includes:

[0033] Each dimension data point corresponds to a row of monitoring time series data that forms a dimension data matrix.

[0034] The present invention has the following beneficial effects:

[0035] This invention clusters all dimensional data points based on their location characteristics, obtaining multiple data clusters. Clustering dimensional data points in similar locations may identify areas with similar change patterns within the target monitoring region. Based on the morphological characteristics of the data cluster to which each dimensional data point belongs and the number of corresponding dimensional data points of the same category, a target state range factor for each dimensional data point is obtained, quantifying the extent to which each dimensional data point represents a change in the target state within the data cluster. Based on the target state range factor for each category corresponding to each dimensional data point within each data cluster, and the changing trend of the monitoring time-series data corresponding to each dimensional data point, the target data change sensitivity of each data cluster is obtained, identifying which data clusters are more sensitive to subtle changes in the target state, thus more accurately capturing changes in the target state. Based on each dimension... This invention obtains target state information representation parameters for each dimension's data points by analyzing the sensitivity of target data changes within the data cluster and the distribution of changes in the monitoring time series data corresponding to each dimension's data points. This evaluates the ability of each dimension's data points to express the target state during target feature state monitoring, contributing to a comprehensive understanding of target state changes. Based on the target state information representation parameters and target state range factors for each dimension's data points, the retention priority for each dimension's data points is obtained, reflecting the importance of the information contained within the dimension's data points and helping to retain more features that significantly impact the monitored target area. A dimensional data matrix composed of all monitoring time series data is obtained. Dimensionality reduction is then performed based on the retention priority of each dimension's data points and the dimensional data matrix, reducing the number of dimensions in the dataset while preserving as much important information as possible from the original data. This invention improves the accuracy of dimensionality reduction by obtaining accurate retention priorities for dimensional data points at each location. Attached Figure Description

[0036] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 A flowchart illustrating a multi-dimensional data processing method provided in one embodiment of the present invention;

[0038] Figure 2 This is a flowchart illustrating a method for obtaining the sensitivity of target data changes, as provided in an embodiment of the present invention. Detailed Implementation

[0039] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a multi-dimensional data processing method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0041] The following description, in conjunction with the accompanying drawings, details a specific scheme for a multi-dimensional data processing method provided by the present invention.

[0042] Please see Figure 1 The diagram illustrates a flowchart of a multi-dimensional data processing method provided by an embodiment of the present invention, the specific method including:

[0043] Step S1: Obtain the monitoring time series data of the category sensors at each location in the target area within the historical time range, forming a dimensional data point.

[0044] In embodiments of the present invention, when monitoring a target, different types of sensors are typically installed at different locations within the target area. Taking urban area monitoring as an example, it is necessary to monitor the status of urban environment, urban traffic, urban energy, etc. The monitored categories include data such as air quality, noise, temperature, and humidity in the urban environment; data such as road traffic flow, vehicle speed, and vehicle type in urban traffic; and data such as voltage, current, and power of power lines, substations, and charging piles in urban energy. The monitoring time-series data of each type of sensor at each location within the target area are acquired within a historical time range, forming a dimensional data point. Here, one location corresponds to one type of sensor, and within a historical time range, the type of sensor at each location monitors multiple time-series data, forming a dimensional data point.

[0045] It should be noted that, in the embodiments of the present invention, since different categories of data are acquired in the target monitoring area, in order to facilitate subsequent data processing, the range standardization method is used after the data acquisition is completed to eliminate the influence of dimensions, so as to facilitate subsequent comparison and analysis.

[0046] It should be noted that, in one embodiment of the present invention, the historical time range is 3 months and the time interval is 6 hours to acquire monitoring time series data; in other embodiments of the present invention, the size of the historical time range and the time interval can be set according to specific circumstances, and will not be limited or elaborated here.

[0047] Step S2: Based on the location characteristics of each dimension data point, cluster all dimension data points to obtain multiple data clusters; based on the morphological characteristics of the data cluster to which each dimension data point belongs and the number of dimension data points of the same category, obtain the target state range factor for each dimension data point.

[0048] Different locations within a monitored target area may exhibit varying monitoring densities, reflecting different target status ranges. Cluster analysis can group data points based on their location similarity, identifying dimensional data points within regions of similar characteristics and dividing the monitored target area into different ranges. Therefore, based on the location characteristics of each dimensional data point, clustering is performed on all dimensional data points to obtain multiple data clusters.

[0049] Preferably, in one embodiment of the present invention, the method for obtaining data clusters includes:

[0050] Based on the location characteristics of data points in each dimension, the K-means clustering algorithm is applied to all data points in all dimensions to obtain multiple data clusters.

[0051] The K-means clustering algorithm groups multiple clusters of data points into K specified clusters based on their similarity. Each cluster belongs to one and only one cluster whose distance to the center of its dimensional cluster is minimized. It should be noted that, in one embodiment of this invention, when using the K-means clustering algorithm to cluster all dimensional data points, the K value is obtained by using the elbow rule to determine the K value and obtain the corresponding number of dimensional clusters. The specific K-means clustering algorithm and elbow rule are well-known techniques to those skilled in the art and will not be elaborated upon here.

[0052] Within the target monitoring area, the monitoring density of data varies in different areas. The morphological characteristics of data clusters help to understand the spatial distribution of target activities within the target monitoring area. The number of dimensional data points of the corresponding category can reveal the universality and importance of different target states. The more dimensional data points of the same category within a region, the smaller the range of target states represented by a single dimensional data point. Therefore, based on the morphological characteristics of the data cluster to which each dimensional data point belongs and the number of dimensional data points of the same category, the target state range factor for each dimensional data point can be obtained.

[0053] Preferably, in one embodiment of the present invention, the method for obtaining the target state range factor includes:

[0054] The coordinate region of each data cluster is obtained using the convex hull detection algorithm;

[0055] Based on the number of corresponding dimension data points of the same category in the data cluster where each dimension data point is located, and the range of the corresponding coordinate region, the target state range factor for each dimension data point is obtained. The number of corresponding dimension data points of the same category is negatively correlated with the target state range factor, while the range of the coordinate region is positively correlated with the target state range factor.

[0056] Among them, a positive correlation means that the dependent variable increases as the independent variable increases, and the dependent variable decreases as the independent variable decreases. The specific relationship can be a multiplicative relationship, an additive relationship, an exponential function, etc., which is determined by practical application. A negative correlation means that the dependent variable decreases as the independent variable increases, and the dependent variable increases as the independent variable decreases. The relationship can be a subtractive relationship, a division relationship, etc., which is determined by practical application.

[0057] In one embodiment of the present invention, the method for obtaining the target state range factor includes:

[0058] ;

[0059] in, Indicates the first Target state range factor for each dimension of data points; Indicates the first The number of data points of the same category in the data cluster where each dimension data point is located; Indicates the first The coordinate range of the data cluster where each dimension of data point belongs; This represents the normalization function.

[0060] In the formula for the target state range factor, the size of the coordinate region is represented by calculating the area of ​​the coordinate region; the larger the area, the greater the range factor. The larger the coordinate region of the data cluster where the data point in each dimension is located, the more likely it is that the data point in the first dimension will be included in the cluster. The smaller the number of dimensional data points of the same category in the data cluster where the dimensional data point belongs, that is, the fewer dimensional data points of the same category exist in the region, the better. The target state is represented by 3 dimensions of data points, which are relatively scattered. The larger the target state range factor for each dimension of data points, the smaller the range factor; conversely, the smaller the range factor for the third dimension of data points, the smaller the range factor for the fourth dimension of data points. The smaller the coordinate region of the data cluster where the data point in each dimension is located, the better. The larger the number of data points of the same category in the data cluster where a data point of a dimension belongs, the more concentrated the distribution of the data points of that dimension. The smaller the target state range factor for each dimension of data points, the better.

[0061] It should be noted that in other embodiments of the present invention, The smaller the coordinate region range, the larger the number of dimensional data points of the corresponding category, the larger the ratio, and the smaller the target state range factor. This correlation can also be achieved by... function pairs A negative correlation mapping is performed to obtain the correlation that the smaller the coordinate region range, the larger the number of dimensional data points of the corresponding category, the larger the ratio, and the smaller the target state range factor; the specific methods are well known to those skilled in the art and will not be elaborated here.

[0062] Step S3: Based on the target state range factor of each category corresponding to each dimension data point within each data cluster, and the changing trend of the monitoring time series data corresponding to each dimension data point, obtain the target data change sensitivity of each data cluster; based on the target data change sensitivity of the data cluster to which each dimension data point belongs, and the change distribution of the monitoring time series data corresponding to each dimension data point, obtain the target state information representation parameters of each dimension data point.

[0063] Since the actual geographical locations within the monitoring target areas represented by any data cluster are relatively close, the impact on the target status is also relatively similar. A larger target state range factor for each dimension's data point may indicate a locational influence with other dimension's data points, making it less likely to reflect the sensitivity to target data changes. By analyzing the changing trends of the monitoring time-series data corresponding to each dimension's data point, the impact on the target status at that location can be represented. A larger changing trend indicates a greater potential impact and potentially higher sensitivity. Therefore, based on the target state range factor for each dimension's data point within each data cluster, and the changing trends of the monitoring time-series data corresponding to each dimension's data point, the target data change sensitivity of each data cluster is obtained.

[0064] Preferably, in one embodiment of the present invention, the method for obtaining the sensitivity to changes in target data is described in [reference needed]. Figure 2 The diagram illustrates a flowchart of a method for obtaining the sensitivity to changes in target data according to an embodiment of the present invention. The method includes:

[0065] Step S201: For any category within each data cluster, obtain the fluctuation degree of the monitoring time series data corresponding to each dimension data point as the local fluctuation degree; obtain the mean of the local fluctuation degree of all dimension data points as the overall fluctuation degree.

[0066] It should be noted that the degree of fluctuation can be calculated by measuring the variance of the monitoring time series data corresponding to each data point in each dimension. The larger the variance, the greater the degree of fluctuation and the more dispersed the distribution of the monitoring time series data; the smaller the variance, the smaller the degree of fluctuation and the more uniform the distribution of the monitoring time series data.

[0067] To facilitate data comparison, the local fluctuations of all data points across all dimensions are quantified and statistically analyzed by averaging, and the overall fluctuations are compared in subsequent calculations.

[0068] Step S202: Based on the target state range factor of each dimension data point and the difference between the local fluctuation degree and the overall fluctuation degree of each dimension data point, obtain the local sensitivity of each dimension data point. The target state range factor is negatively correlated with the local sensitivity, and the difference between the local fluctuation degree and the overall fluctuation degree is positively correlated with the local sensitivity.

[0069] The greater the difference between the local and overall fluctuation levels of each dimension's data points, the more inconsistent the changes in the local and overall fluctuation levels of the corresponding dimension's data points are, and the greater the possibility of being affected by other factors. The larger the target state range factor of each dimension's data points, the larger the range of representation of the target state by that dimension's data points, the more inconsistent the changes in the target state are affected by, and the less likely it is caused by data sensitivity, and the lower the local sensitivity. Conversely, the smaller the target state range factor of each dimension's data points, the smaller the range of representation of the target state by that dimension's data points, the closer the changes in the target state are affected by, and the more likely it is caused by data sensitivity, and the greater the local sensitivity.

[0070] Among them, a positive correlation means that the dependent variable increases as the independent variable increases, and the dependent variable decreases as the independent variable decreases. The specific relationship can be a multiplicative relationship, an additive relationship, an exponential function, etc., which is determined by practical application. A negative correlation means that the dependent variable decreases as the independent variable increases, and the dependent variable increases as the independent variable decreases. The relationship can be a subtractive relationship, a division relationship, etc., which is determined by practical application.

[0071] Step S203: Obtain the mean of the local sensitivity of all data points in any category as the category sensitivity; obtain the mean of the category sensitivity within each data cluster as the target data change sensitivity of each data cluster.

[0072] In one embodiment of the present invention, the formula for the sensitivity to changes in target data is expressed as:

[0073] ;

[0074] in, Indicates the first Sensitivity to changes in target data for each data cluster; Indicates the first The number of categories within each data cluster; Indicates the first The number of data points in each dimension under each category; Indicates the first Target state range factor for each dimension of data points; Indicates the first The variance of the monitoring time series data corresponding to each dimension data point; This represents the mean variance of the monitoring time series data corresponding to all data points across all dimensions.

[0075] In the formula for sensitivity to changes in target data, Indicates the calculation of the first The difference between the variance of the monitoring data corresponding to each dimension's data point and the mean variance of the monitoring data corresponding to all dimensions' data points, i.e., the difference between the variance of the monitoring data corresponding to each dimension's data point and the mean variance of the monitoring data corresponding to all dimensions' data points. The difference between the local and overall volatility of data points in each dimension; the greater the difference, The less similar the local fluctuations and overall fluctuations of data points in each dimension are, the more inconsistent the changes between data points in one dimension and data points in other dimensions are, the more complex the impact is, and the greater the sensitivity to changes in the target data. Indicates the first The target state range factor is negatively mapped for each dimension of data points; the larger the target state range factor, the better. The smaller; Indicates the first Local sensitivity of data points in each dimension Indicates category sensitivity, This indicates the sensitivity to changes in the target data, i.e., the first... The larger the target state range factor of the data points in each dimension, the greater the... The greater the difference between the local and overall fluctuation levels of data points in each dimension, the greater the local sensitivity, and the greater the sensitivity to changes in the target data of each data cluster.

[0076] Data clusters may represent the state or characteristics of different areas within a monitored target region. By analyzing the sensitivity of each data cluster to changes in target data, we can understand which clusters are more sensitive to subtle changes in the target state, thus capturing these changes more accurately. Monitoring time-series data records the state information of each location within the monitored target region at different points in time. By analyzing the distribution of these changes, we can understand the evolution trend of the target state over time. Combining the sensitivity of data clusters, we can further identify which dimensions of data points have a more significant impact on the target state, potentially providing more information to monitor the target's changing trends. Therefore, based on the sensitivity of each dimension's data point to changes in its data cluster and the distribution of changes in the corresponding monitoring time-series data, we obtain the target state information representation parameters for each dimension's data point.

[0077] Preferably, in one embodiment of the present invention, the method for obtaining the target state information representation parameters includes:

[0078] Based on the change distribution of the monitoring data corresponding to each dimension data point, obtain the change information of each dimension data point;

[0079] The change information is weighted according to the sensitivity of the target data change to the data cluster to which each dimension data point belongs, so as to obtain the target state information representation parameters of each dimension data point.

[0080] It should be noted that the greater the sensitivity of the target data to changes, the greater the degree to which the data is affected and changes, the more information is generated by the changes, the more obvious the uncertainty of the target state, and the larger the target state information representation parameter.

[0081] Preferably, in one embodiment of the present invention, the method for obtaining change information includes:

[0082] Obtain the fitting curve of the monitoring time series data corresponding to each dimension data point; the change process of the fitting curve can reflect the overall change process of the data in different time series.

[0083] The amount of change information is obtained by considering the degree of disorder and the number of monotonic trend intervals of the rate of change for each data point on the corresponding fitted curve. Both the degree of disorder and the number of monotonic trend intervals are positively correlated with the amount of change information.

[0084] It should be noted that the degree of disorder in the rate of change of each data point represents the uncertainty of data change, and the number of monotonic intervals in the fitted curve represents the volatility of data within the monitoring period. The greater the uncertainty of change, the greater the volatility of data, indicating more information about the change, and possibly more uncertainty about the target state.

[0085] It should be noted that, in one embodiment of the present invention, the least squares fitting method is used to fit the monitoring data corresponding to each dimension data point to obtain a fitting curve; the time series is used as the horizontal axis and the corresponding monitoring data is used as the vertical axis; the rate of change of each data point is represented by calculating the derivative of each data point on the corresponding fitting curve, the larger the derivative, the faster the rate of change; the information entropy of the rate of change of each data point is calculated to represent the degree of disorder, the larger the information entropy, the greater the degree of disorder, and the more information it contains.

[0086] It should be noted that a positive correlation means that the dependent variable increases as the independent variable increases, and decreases as the independent variable decreases. The specific relationship can be multiplicative, additive, or an exponential function, determined by the actual application. In one embodiment of the invention, a positive correlation can be constructed by multiplying the degree of disorder and the number of monotonic trend intervals. A higher degree of disorder and a larger number of monotonic trend intervals indicate more change information, and more change information leads to greater sensitivity to changes in the target data and a larger target state information representation parameter. Specific methods are well-known to those skilled in the art and are not limited or elaborated upon here.

[0087] Step S4: Based on the target state information representation parameters and target state range factors of each dimension data point, obtain the retention priority of each dimension data point.

[0088] The target state information representation parameters of each dimension's data points indicate their ability to express the target state during target feature state monitoring. A larger target state information representation parameter indicates a more pronounced representation of the target state by the dimensional data point, requiring more features to be retained for subsequent analysis. Conversely, a larger target state range factor suggests the potential to represent more target state features, necessitating their retention. Therefore, based on the target state information representation parameters and target state range factors of each dimension's data points, the retention priority for each dimension's data points is determined.

[0089] Preferably, in one embodiment of the present invention, the method for obtaining the priority includes:

[0090] Calculate the product of the target state information representation parameter and the corresponding target state range factor for each dimension data point, and normalize it to obtain the retention priority of each dimension data point.

[0091] In one embodiment of the present invention, the formula for retaining priority is expressed as follows:

[0092] ;

[0093] in, Indicates the first Priority for retaining data points in each dimension; Indicates the first Target state range factor for each dimension of data points; Indicates the first Parameters representing the target state information of data points in each dimension; This represents the normalization function.

[0094] In the formula that preserves priority, the first... The larger the target state range factor of the data points in each dimension, the greater the... The larger the target state information representation parameter of each dimension data point, that is, the stronger the ability of the dimension data point to represent the target state and the larger the representation range, the higher the priority of preservation in the dimensionality reduction process, and the more its features need to be preserved.

[0095] It should be noted that in other embodiments of the present invention, a positive correlation can also be constructed by normalizing and adding the parameters, such that the larger the target state information representation parameter, the larger the corresponding target state range factor, and the higher the retention priority. The specific means are well known to those skilled in the art and will not be described in detail here.

[0096] Step S5: Obtain the dimensional data matrix composed of all monitoring time series data; perform dimensionality reduction processing according to the retention priority of each dimensional data point and the dimensional data matrix to obtain the dimensionality-reduced data of the dimensional data matrix.

[0097] Dimensional data matrices allow for a more intuitive observation of changes in the data within the matrix, improving the efficiency, accuracy, and flexibility of data analysis and processing; therefore, a dimensional data matrix composed of all monitored time-series data is obtained.

[0098] Preferably, in one embodiment of the present invention, the method for obtaining the dimensional data matrix includes:

[0099] Each dimension data point corresponds to a row of monitoring time series data that forms a dimension data matrix.

[0100] Prioritization reflects the importance of information contained in dimensional data points. High-priority dimensional data points typically contain more crucial information about the target state and therefore need to be retained during dimensionality reduction. Dimensionality reduction of the dimensional data matrix reduces the number of dimensions in the dataset while preserving as much important information as possible from the original data. This increases the efficiency of the target state evaluation process while minimizing data loss and reducing the amount of data to be analyzed. Therefore, dimensionality reduction is performed based on the retention priority of each dimensional data point and the dimensional data matrix to obtain the dimensionality-reduced data matrix.

[0101] Preferably, in one embodiment of the present invention, the method for obtaining the dimensionality reduction data of the dimensionality data matrix includes:

[0102] In the process of performing PCA dimensionality reduction on a dimensional data matrix, the normalization matrix of the dimensional data matrix is ​​obtained, and the normalization matrix is ​​adjusted according to the retention priority of each dimensional data point to obtain the adjustment matrix;

[0103] Obtain the principal direction matrix of the adjustment matrix; calculate the product of the principal direction matrix and the dimension data matrix to obtain the dimensionality-reduced data of the dimension data matrix.

[0104] It should be noted that, in one embodiment of the present invention, the method for obtaining the principal direction matrix is ​​as follows: during the execution of the PCA dimensionality reduction algorithm, the dimensional data matrix is ​​demeaned to obtain a standardized matrix of the dimensional data matrix. Considering that the degree of influence of the target state on each dimensional data point is inconsistent, an adjustment matrix is ​​obtained based on the retention priority of each dimensional data point. The principal direction matrix is ​​obtained by applying the subsequent original calculation method of the PCA dimensionality reduction algorithm to the adjustment matrix; that is, the covariance matrix of the adjustment matrix is ​​obtained, and the eigenvalues ​​and eigenvectors of the covariance matrix are calculated. The eigenvectors are arranged into a matrix in descending order of their corresponding eigenvalues, and the first k rows are taken to form the principal direction matrix. The first k rows are determined by the cumulative contribution rate of the eigenvalues, which refers to the proportion of the sum of the first k eigenvalues ​​to the total sum of all eigenvalues. In one embodiment of the present invention, the implementer can set the cumulative contribution rate threshold to 85% based on experience, that is, when the cumulative contribution rate is greater than or equal to the cumulative contribution rate threshold, the first k rows of data are obtained. The specific means are well known to those skilled in the art and will not be described in detail here.

[0105] Preferably, in one embodiment of the present invention, the method for obtaining the adjustment matrix includes:

[0106] The monitoring time series data of the corresponding dimension data points in the standardized matrix are weighted according to the retention priority of each dimension data point to obtain the adjustment matrix.

[0107] After obtaining dimensionality-reduced data, the dimensionality of the data can be effectively reduced, which can preserve the main features of the data, reduce computational complexity, reduce the impact of noise and redundant information, and enable accurate and rapid monitoring of target activities and regional characteristics in the target area of ​​m.

[0108] In summary, this invention obtains multiple data clusters of dimensional data points; based on the morphological characteristics of the data cluster to which each dimensional data point belongs and the number of dimensional data points of the same category, it obtains the target state range factor for each dimensional data point; based on the target state range factor for each category corresponding to each dimensional data point within each data cluster, and the changing trend of the monitoring time series data corresponding to each dimensional data point, it obtains the target data change sensitivity of each data cluster; based on the target data change sensitivity of the data cluster to which each dimensional data point belongs, and the changing distribution of the monitoring time series data corresponding to each dimensional data point, it obtains the target state information representation parameters for each dimensional data point; furthermore, it obtains the retention priority of each dimensional data point; it obtains the dimensional data matrix composed of all monitoring time series data; and it performs dimensionality reduction processing based on the retention priority of each dimensional data point and the dimensional data matrix to obtain the dimensionality-reduced data of the dimensional data matrix. This invention improves the accuracy of dimensionality reduction by obtaining the accurate retention priority of dimensional data points at each location.

[0109] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0110] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A multi-dimensional data processing method, characterized in that, The method includes: The monitoring time series data of each type of sensor at each location in the target area within a historical time range are acquired to form a dimensional data point. The target area is an urban area, and the monitored data categories include air quality, noise, temperature, and humidity data in the urban environment; road traffic flow, vehicle speed, and vehicle type data in urban traffic; and voltage, current, and power data of power lines, substations, and charging piles in urban energy. Based on the location characteristics of each dimension's data points, all dimension's data points are clustered to obtain multiple data clusters; based on the morphological characteristics of the data cluster to which each dimension's data point belongs and the number of dimension's data points of the same category, the target state range factor for each dimension's data point is obtained. Based on the target state range factor corresponding to each category and each dimension data point within each data cluster, and the changing trend of the monitoring time series data corresponding to each dimension data point, the target data change sensitivity of each data cluster is obtained; based on the target data change sensitivity of the data cluster to which each dimension data point belongs, and the change distribution of the monitoring time series data corresponding to each dimension data point, the target state information representation parameters of each dimension data point are obtained. Based on the target state information representation parameters and the target state range factor for each dimension data point, the retention priority of each dimension data point is obtained; Obtain a dimensional data matrix composed of all monitoring time series data; perform dimensionality reduction processing according to the retention priority of each dimensional data point and the dimensional data matrix to obtain dimensionality-reduced data of the dimensional data matrix; The method for obtaining the target state range factor includes: The coordinate region of each data cluster is obtained using the convex hull detection algorithm; Based on the number of dimensional data points of the same category in the data cluster where each dimensional data point is located, and the range of the corresponding coordinate region, the target state range factor for each dimensional data point is obtained. The number of dimensional data points of the same category is negatively correlated with the target state range factor, and the range of the coordinate region is positively correlated with the target state range factor.

2. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the sensitivity to changes in the target data includes: For any category within each data cluster, obtain the fluctuation level of the monitoring time series data corresponding to each dimension data point as the local fluctuation level; obtain the mean of the local fluctuation levels of all dimension data points as the overall fluctuation level. Based on the target state range factor of each dimension data point and the difference between the local fluctuation degree and the overall fluctuation degree of each dimension data point, the local sensitivity of each dimension data point is obtained. The target state range factor is negatively correlated with the local sensitivity, and the difference between the local fluctuation degree and the overall fluctuation degree is positively correlated with the local sensitivity. The mean of the local sensitivity of all data points in any category is obtained as the category sensitivity; the mean of the category sensitivity within each data cluster is obtained as the target data change sensitivity of each data cluster.

3. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the target state information representation parameters includes: Based on the change distribution of the monitoring time series data corresponding to each dimension data point, the change information of each dimension data point is obtained; The change information is weighted according to the sensitivity of the target data change to the data cluster to which each dimension data point belongs, so as to obtain the target state information representation parameters of each dimension data point.

4. The multi-dimensional data processing method according to claim 3, characterized in that, The method for obtaining the amount of change information includes: Obtain the fitting curve of the monitoring time series data corresponding to each dimension data point; The amount of change information is obtained by considering the degree of disorder and the number of monotonic trend intervals of the rate of change for each data point on the corresponding fitted curve. Both the degree of disorder and the number of monotonic trend intervals are positively correlated with the amount of change information.

5. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the retention priority includes: Calculate the product of the target state information representation parameter and the corresponding target state range factor for each dimension data point, and normalize it to obtain the retention priority of each dimension data point.

6. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the dimensionality reduction data of the dimensionality data matrix includes: During the PCA dimensionality reduction algorithm on the dimensional data matrix, a standardized matrix of the dimensional data matrix is ​​obtained, and the standardized matrix is ​​adjusted according to the retention priority of each dimensional data point to obtain an adjusted matrix; Obtain the principal direction matrix of the adjustment matrix; calculate the product of the principal direction matrix and the dimension data matrix to obtain the dimensionality-reduced data of the dimension data matrix.

7. The multi-dimensional data processing method according to claim 6, characterized in that, The method for obtaining the adjustment matrix includes: The monitoring time series data of the corresponding dimension data points in the standardized matrix are weighted according to the retention priority of each dimension data point to obtain the adjustment matrix.

8. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the data clusters includes: Based on the location characteristics of data points in each dimension, the K-means clustering algorithm is applied to all data points in all dimensions to obtain multiple data clusters.

9. The multi-dimensional data processing method according to claim 1, characterized in that, The method for obtaining the dimensional data matrix includes: Each dimension data point corresponds to a row of monitoring time series data that forms a dimension data matrix.

Citation Information

Patent Citations

  • Dimension reduction analysis method and system for thermal response data of medium-deep underground heat exchanger

    CN118708953A

  • Massive high-dimensional AIS trajectory data clustering method

    WO2023029461A1