A Machine Learning-Based Sensitive Data Detection Method and System
Through the sensitive data exploration method based on machine learning, the problem of inefficient identification and protection of sensitive data in the prior art is solved, efficient and accurate identification and hierarchical protection of sensitive data are achieved, and data security is enhanced.
Patent Information
- Application Number
- CN202510387379.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art lacks the ability to efficiently identify and protect sensitive data when processing power-related data, resulting in the risk of leakage or insufficient protection of sensitive data during processing.
Using machine learning-based sensitive data exploration method, we collect data from power-related data sources, perform preprocessing, generate multi-dimensional feature vectors, build sensitive data exploration models, and label and hierarchical protection of sensitive data based on model output results.
It significantly improves the efficiency and accuracy of sensitive data identification, dynamically adjusts the probing model to adapt to changes in data distribution, provides a more complete expression of data sensitivity, and enhances data security through encrypted storage and strict access control policies.
Smart Images

Figure CN119884897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and a sensitive data detection method and system based on machine learning. Background Art
[0002] With the rapid development of information technology, the capabilities of data collection, storage, and analysis have been continuously improved. In the power-related fields, data-driven decision-making and management have become important means to improve efficiency and safety. Traditionally, power data mainly exists in the form of regional distribution, including geographical information such as pole tower locations, airport locations, and important station locations. These data are of great value in power network design, operation optimization, and risk assessment, and at the same time, they also face severe challenges in data protection and sensitivity identification.
[0003] However, when dealing with these data, the existing technologies generally lack the ability to efficiently identify and protect sensitive data, and usually rely on manual rules or static classification methods, which are inefficient in the case of complex data dimensions and diverse characteristics. In addition, these traditional methods cannot dynamically adapt to changes in data distribution, nor can they formulate precise protection strategies for different sensitivity levels, resulting in the risk of leakage or insufficient protection of sensitive data during the processing.
[0004] Therefore, it is necessary to develop a new type of sensitive data detection method and system based on machine learning. Summary of the Invention
[0005] This application provides a sensitive data detection method and system based on machine learning to improve the efficiency and accuracy of sensitive data identification.
[0006] This application provides a sensitive data detection method based on machine learning, including:
[0007] Collecting raw data including pole tower location data, airport location data, and important station location data from power-related data sources; preprocessing the raw data, where the preprocessing includes cleaning outliers, unifying data formats, and standardizing geographical coordinates to generate preprocessed structured data;
[0008] Generating a multi-dimensional feature vector based on the preprocessed structured data; combining the multi-dimensional feature vector with its corresponding sensitivity annotation label according to the pre-specified sensitivity annotation rules to generate model training data; where the multi-dimensional feature vector includes location distribution, spatial correlation, and data frequency;
[0009] Based on the model training data, a sensitive data detection model is constructed using a machine learning algorithm; the sensitive data detection model is trained through iterative optimization to obtain model parameters for identifying the sensitivity level of target data; according to the model parameters, a trained sensitive data detection model is obtained;
[0010] The target data set is input into the sensitive data detection model to obtain a model output result; according to the model output result, the sensitive data included in the target data set is labeled to generate a sensitivity classification label; wherein, the target data set uses preprocessed target structured data;
[0011] According to the sensitivity classification label, hierarchical protection is carried out on the detected sensitive data, including encryption storage and access control policies for highly sensitive data.
[0012] Furthermore, generating a multi-dimensional feature vector based on the preprocessed structured data includes:
[0013] The input geographical data is divided into equidistant grid cells according to a preset rule, where each grid cell corresponds to a specified geographical range; the number of geographical data points in each grid cell is calculated, expressed as grid density; the grid density value is normalized to generate a location distribution feature reflecting the local distribution characteristics of the geographical data;
[0014] For each geographical data point, calculate its distance from a set of predefined important facilities; assign weights according to the facility type, and perform weighted averaging on each distance value and the corresponding weight to obtain a spatial association value; normalize the spatial association values of all geographical data points to generate a spatial correlation feature; wherein, the important facilities include power towers, airports, and substations;
[0015] Analyze the time attribute of each geographical data point, and count the number of occurrences within a fixed time period as a frequency value; by setting a time decay weight, perform weighted processing on the frequency value and then normalize it to generate a spatio-temporal frequency feature describing the spatio-temporal activity level;
[0016] Combine the normalized location distribution feature, spatial correlation feature, and spatio-temporal frequency feature of each geographical data point to generate a multi-dimensional feature vector, and output a feature vector data set for model training.
[0017] Furthermore, for each geographical data point, calculate its distance from a set of predefined important facilities; assign weights according to the facility type, and perform weighted averaging on each distance value and the corresponding weight to obtain a spatial association value; normalize the spatial association values of all geographical data points to generate a spatial correlation feature, including:
[0018] For each geographical data point , calculate its important facilities set Each facility distance, generate a distance matrix , among which, Line Column value Indicates geographic data points to The distance to the facility, Calculate according to the following formula (1):
[0019] ;
[0020] in, and Respectively geographic data points and The location coordinates of each facility;
[0021] According to Risk factor of important facilities , a dynamic weight is assigned to each facility according to the following formula (2):
[0022] ;
[0023] in, For the The weight coefficient of each important facility; The number of important facilities;
[0024] According to the following formula (3), the spatial association value of the geographic data point is calculated :
[0025] ;
[0026] in, is the attenuation coefficient, which is used to control the effect of distance on spatial correlation; For the matrix No. The maximum distance value of the column, used for normalized distance; The number of important facilities;
[0027] According to the following formula (4), the spatial correlation value of all geographic data points is Perform normalization:
[0028] ;
[0029] in, Normalized spatial correlation value; For all spatial association values The minimum value in; For all spatial correlation values The maximum value in;
[0030] Generate a spatial correlation feature according to the normalized spatial correlation value.
[0031] The beneficial effects of the technical solution provided by this application include:
[0032] (1) Through the sensitive data exploration model based on machine learning, the present invention uses multi-dimensional feature vectors and sensitivity annotation rules to achieve automatic classification and accurate identification of target data, significantly improving the efficiency and accuracy of sensitive data identification. (2) By iteratively optimizing the training model parameters, the present invention can dynamically adjust the exploration model according to the change of data distribution, ensuring the applicability of the model under various scenarios and data complexity conditions, thus avoiding the limitations of traditional static classification methods. (3) By introducing multi-dimensional features such as position distribution, spatial correlation, and data frequency, the present invention comprehensively models the spatial and temporal correlation characteristics of data, provides a more complete expression of data sensitivity, and solves the problem of insufficient single-dimensional analysis. (4) The present invention classifies and protects data according to sensitivity classification labels, especially adopts encryption storage and strict access control strategies for highly sensitive data, realizes targeted protection of sensitive data, and enhances data security. Brief Description of the Drawings
[0033] Figure 1 is a flowchart of a sensitive data exploration method based on machine learning provided by the first embodiment of this application.
[0034] Figure 2 is a schematic diagram of a sensitive data exploration system based on machine learning provided by the second embodiment of this application. Detailed Embodiments
[0035] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific embodiments disclosed below.
[0036] The first embodiment of this application provides a sensitive data exploration method based on machine learning. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of this application. The following combines Figure 1 to describe in detail a sensitive data exploration method based on machine learning provided by the first embodiment of this application.
[0037] Step S101: Collect raw data including pole position data, airport location data, and important station location data from power-related data sources; preprocess the raw data, where the preprocessing includes cleaning outliers, unifying data formats, and standardizing geographical coordinates to generate preprocessed structured data.
[0038] Step S101 includes two main processes: collection and preprocessing of raw data. The following provides specific descriptions of these two processes.
[0039] First, collect raw data including pole position data, airport location data, and important station location data from power-related data sources. The data sources can be distributed sensor networks, geographic information systems (GIS), satellite remote sensing systems, or manually maintained databases. The collection process can be carried out through programming interface calls, real-time data streams, or batch imports. Taking pole position data as an example, the longitude and latitude coordinate information of poles can be exported through a GIS platform, along with other geographical attributes such as pole types and area numbers where they are located. Airport location data can be obtained from aviation management databases, including the exact location of airports, classification information (such as international airports, local airports), and distribution information of nearby facilities. Important station location data can be extracted from the internal database of power companies, and the data format may be CSV, JSON, or XML files, and the content includes station names, geographical coordinates, and their functional attributes, etc.
[0040] After the collection is completed, enter the preprocessing stage. First, clean the outliers in the raw data. The identification of outliers can adopt statistical methods. For example, conduct range checks on the position coordinates to eliminate points outside reasonable geographical boundaries; eliminate or correct entries with incorrect data types (such as records with longitude and latitude formats not meeting the standards). During the cleaning process, it is necessary to ensure the integrity and accuracy of the data. For example, when some records are missing, reasonable filling can be carried out through interpolation methods or statistical distributions of adjacent area data.
[0041] Secondly, unifying the data format is a key step. In this process, standardize the data from different data sources into a unified file format. For example, convert JSON or XML formats into tabular CSV files, or directly store them as table structures in relational databases. During the process of unifying the data format, ensure that each record has a consistent field structure. For example, each record contains fields such as "data type", "geographical coordinates", "timestamp", etc.
[0042] Then, perform standardization processing on the geographical coordinates.
[0043] Finally, store the preprocessed data as structured data. A relational database (such as MySQL or PostgreSQL) can be used to store each record, or a distributed storage system (such as Hadoop HDFS) can be used to manage large-scale datasets. Each record is represented by standardized fields in the database. For example, the "data type" can be "tower", "airport", or "station", the "geographical coordinates" are represented by latitude and longitude values, and the "additional attributes" store specific functional information, such as the load capacity of a tower or the service range of a station.
[0044] Step S102: Generate a multi-dimensional feature vector based on the preprocessed structured data; combine the multi-dimensional feature vector with its corresponding sensitivity annotation label according to the predefined sensitivity annotation rules to generate model training data; wherein, the multi-dimensional feature vector includes location distribution, spatial correlation, and data frequency.
[0045] In step S102, it is first necessary to extract key information from the preprocessed structured data generated in step S101 and construct a multi-dimensional feature vector for describing data characteristics based on this. The generation process of the multi-dimensional feature vector includes three main aspects: location distribution characteristics, spatial correlation characteristics, and data frequency characteristics.
[0046] The extraction of location distribution characteristics is completed by analyzing the distribution of data points in the geographical space. The geographical locations of the data points are divided into several regions according to preset rules. The division of regions can adopt a fixed grid or a dynamic division method based on geographical characteristics. The distribution density of data points is counted in each region, and the statistical results are standardized to reflect the spatial distribution characteristics of data points in different regions. This process ensures that the generated features can accurately describe the local distribution pattern of data points and maintain consistency with the distribution in the entire spatial range.
[0047] The extraction of spatial correlation characteristics requires analyzing the spatial relationships between data points and between data points and specific facilities (such as towers, airports, and important stations). For each data point, calculate its distance to the target facility, and at the same time assign weights in combination with the type and importance of the target facility to reflect the strength of the spatial correlation. Then, these distance and weight information are combined into a set of feature values that can quantify the spatial correlation. These feature values can describe the relative spatial position relationship between the data point and the facility, thereby capturing potential geographical correlations.
[0048] The extraction of data frequency features is achieved by analyzing the time attributes of data points. The time attributes of data points are divided into several fixed periods. For example, the number of occurrences of each data point within the corresponding period is counted by day, week, or month to construct basic frequency features. On this basis, by introducing time decay weights to adjust the frequency values of data points, the influence of the activity in the recent time on the feature values is made greater, so as to better fit the influence of time changes on data sensitivity in the actual scenario.
[0049] After generating the location distribution features, spatial correlation features, and data frequency features, these features are combined into a multi-dimensional feature vector, where each dimension represents an independent feature value. The generated multi-dimensional feature vector comprehensively reflects the spatial distribution, temporal activity, and correlation with facilities of data points, providing a basis for subsequent sensitivity analysis.
[0050] Next, the generated multi-dimensional feature vector is combined with its corresponding sensitivity annotation label, and model training data is constructed according to the pre-specified annotation rules. The annotation rules can be formulated based on the geographical attributes of data points, the relationship with target facilities, and historical experience knowledge. By combining the multi-dimensional feature vector with the sensitivity label, the generated model training data has rich feature information and clear target annotations, laying a solid foundation for the construction and optimization of subsequent models.
[0051] Furthermore, generating the multi-dimensional feature vector based on the preprocessed structured data includes:
[0052] The input geographical data is divided into equidistant grid cells according to preset rules, where each grid cell corresponds to a specified geographical range; the number of geographical data points within each grid cell is calculated, which is expressed as grid density; the grid density value is normalized to generate location distribution features reflecting the local distribution characteristics of geographical data.
[0053] For each geographical data point, calculate its distance from a set of predefined important facilities; assign weights according to the facility type, and perform weighted averaging on each distance value and the corresponding weight to obtain a spatial correlation value; normalize the spatial correlation values of all geographical data points to generate spatial correlation features; where the important facilities include poles, airports, and substations.
[0054] Analyze the time attributes of each geographical data point, and count the number of occurrences within a fixed time period as a frequency value; by setting time decay weights, perform weighted processing on the frequency value and then normalize it to generate a spatio-temporal frequency feature describing the spatio-temporal activity degree.
[0055] Combine the normalized position distribution features, spatial correlation features, and spatio-temporal frequency features of each geographical data point to generate a multi-dimensional feature vector, and output a feature vector dataset for model training.
[0056] In this embodiment, the process of generating a multi-dimensional feature vector from the preprocessed structured data can be divided into several detailed steps, each with a clear goal and method, to ensure that the core features of the data can be comprehensively and accurately extracted, providing high-quality input for the training of subsequent machine learning models.
[0057] First, according to the input geographical data, the target area is divided according to a preset rule. This division is completed based on an equidistant grid cell method, and each grid cell covers a specified geographical range. The purpose of grid division is to classify geographical data points into each cell for analyzing the distribution characteristics of data points in a local area. After the division is completed, the geographical data points in each grid cell are counted, and the number of data points contained in each cell is recorded, which is expressed as the grid density value. To eliminate the influence of the absolute quantity difference between grid density values on subsequent analysis, all grid density values are normalized to ensure that these values are distributed within a fixed range, for example, standardized to between 0 and 1. This normalization process enables the position distribution features to more intuitively reflect the density of geographical data points in different regions and generates the position distribution features of the local distribution features.
[0058] Next, to further analyze the spatial correlation between geographical data points, the distance between each data point and a set of predefined important facilities needs to be calculated. Important facilities can include power towers, airports, and substations, etc. These facilities are defined as target points with special significance, reflecting the potential association between geographical data points and key infrastructure. For each geographical data point, calculate its straight-line distance to each important facility, and at the same time assign specific weights according to the type of facility. For example, power towers may have a lower weight, while substations have a higher weight. After the calculation is completed, the weighted average of each distance value and the corresponding weight is performed to generate the spatial association value of each geographical data point. To ensure the comparability between these association values, the spatial association values of all data points are normalized to generate the spatial correlation features. This feature can comprehensively describe the relative relationship and potential connection between geographical data points and target facilities.
[0059] Subsequently, in order to capture the activity level of geographical data points in the time dimension, it is necessary to analyze the time attributes of each data point. First, the time attributes are divided into fixed time periods, such as days, weeks, or months as the periodic units, and the number of occurrences of the data points within each time period is counted and expressed as a frequency value. These frequency values reflect the activity of the data points in different time periods. On this basis, by introducing a time decay weight to weight the frequency values, the activity in the recent time period contributes more to the frequency characteristics, while the influence of the far - off time period gradually decreases. This processing method can more realistically reflect the spatio - temporal dynamic characteristics of the data points. Finally, the weighted frequency values are normalized to generate spatio - temporal frequency characteristics that describe the spatio - temporal activity level of the data points.
[0060] After the above - mentioned feature extraction steps are completed, the normalized position distribution feature, spatial correlation feature, and spatio - temporal frequency feature of each geographical data point are combined together to generate a multi - dimensional feature vector. The dimension of each feature vector corresponds to the specific components of these features, thus comprehensively reflecting the local distribution, spatial correlation, and spatio - temporal activity level of the geographical data points. The generated multi - dimensional feature vectors are then output as a dataset for model training, providing high - quality input data for the machine learning algorithms in the subsequent steps.
[0061] Furthermore, for each geographical data point, calculate its distance to a set of predefined important facilities; assign weights according to the facility type, and perform a weighted average of each distance value and the corresponding weight to obtain a spatial correlation value; normalize the spatial correlation values of all geographical data points to generate spatial correlation features, including:
[0062] For each geographical data point , calculate its distance to each facility in the set of important facilities to generate a distance matrix . Among them, is the number of important facilities. The set of important facilities includes facilities such as poles, airports, and substations, and the geographical location of each facility is represented in the form of coordinates. The coordinates of each geographical data point P are , and the coordinates of each facility are .
[0063] Among them, the value in the th row and th column represents the distance from the th geographical data point to the th facility, which is calculated according to the following formula (1):
[0064] ;
[0065] Among them, and are the location coordinates of the th geographical data point and the th facility respectively;
[0066] Next, according to the risk coefficient of the facility, a dynamic weight is assigned to each type of facility. The risk coefficient is a predefined parameter used to measure the importance of the facility. For example, the risk coefficient of a pole tower may be relatively low (such as 0.5), while that of an airport or a substation may be relatively high (such as 0.8 or 1.0), and the specific value can be set according to historical data analysis or business requirements. Formula (2) defines the way of assigning the dynamic weight:
[0067] ;
[0068] Among them, is the weight coefficient of the th important facility; is the number of important facilities;
[0069] According to the following formula (3), the spatial correlation value of the geographical data point is calculated :
[0070] ;
[0071] Among them, is the attenuation coefficient used to control the influence of distance on the spatial correlation; is the maximum distance value of the th column of the matrix for normalizing the distance; is the number of important facilities;
[0072] In formula 2, is the distance attenuation coefficient used to control the influence of distance on the spatial correlation value, and its value range is usually between 0.1 and 1. A larger value indicates that the influence of distance is more significant. For example, when , the contribution of a closer facility to increases significantly. is the maximum distance value of the th column of the matrix used to normalize the distance so that the distances of different facilities are comparable. For example, if the maximum distance of the th column is 50 and , then the normalized distance is 。The weighted summation in the formula ensures that the contribution of each facility is adjusted according to the weight and distance decay, generating the initial spatial association value of the geographical data points.
[0073] According to the following formula (4), the spatial association values of all geographical data points are normalized as follows:
[0074] ;
[0075] where, the spatial association value after normalization; is the minimum value among all spatial association values ; is the maximum value among all spatial association values ;
[0076] According to the spatial association value after the normalization process, a spatial correlation feature is generated.
[0077] Furthermore, the time attribute of each geographical data point is analyzed, and the number of occurrences within a fixed time period is counted as a frequency value; by setting a time decay weight, the frequency value is weighted and then normalized to generate a spatio-temporal frequency feature describing the spatio-temporal activity degree, including:
[0078] For the number of occurrences of a geographical data point within a fixed time period , it is weighted according to the decay weight at the current time point. The selection of the time period can be defined according to the application scenario. For example, the occurrence frequency of each data point is counted in units of days, weeks, or months. If the count is done by day, then represents the number of records of a data point within a certain day. For example, the number of occurrences of a data point P in the past 7 days may be {5, 3, 6, 2, 8, 4, 1} respectively, corresponding to the values of these daily occurrences.
[0079] Next, according to the time interval between the current time point and the time period , the decay weight is calculated through formula (5). The definition of the decay weight is to control the influence of long-term data on the frequency calculation, ensuring that the contribution of recent data is greater and the influence of long-term data gradually weakens. Formula (5) is defined as:
[0080] ;
[0081] where, represents the current time and the time period The time interval; is the attenuation control parameter, used to adjust the weight influence of the forward data; is the adjustable exponential parameter, used to control the non - linear degree of attenuation.
[0082] represents the time interval between the current time point and the time period , usually calculated in days or hours. For example, the current time is November 28, 2024, for the period with a record on November 21, 2024, the time interval is 7 days. The parameter is the attenuation control parameter, controlling the influence intensity of the forward data on the weight. The recommended value range is from 0.1 to 1.0; a larger value makes the weight of the forward data decrease faster. The parameter is the adjustable exponential parameter, used to control the non - linear degree of attenuation. For example, when the weight decays linearly, it shows a square - decreasing effect when and . When , the calculated .
[0083] Within each time period , according to the following formula (6), calculate the weighted frequency value :
[0084] ;
[0085] where, is the total number of time periods; is the weighted cumulative frequency value of the geographical data points over all time periods; according to the following formula (7), generate the normalized spatio - temporal frequency feature :
[0086] ;
[0087] where, is the weighted frequency mean of all time periods; is the standard deviation of the weighted frequency value, used to measure the difference in the activity of the data points relative to the overall distribution.
[0088] Step S103: Based on the model training data, use a machine learning algorithm to construct a sensitive data detection model; through iterative optimization training of the sensitive data detection model, obtain the model parameters for identifying the sensitivity degree of the target data; according to the model parameters, obtain the trained sensitive data detection model.
[0089] In step S103, based on the model training data generated in step S102, a sensitive data detection model is constructed using a machine learning algorithm, and the model is trained by means of iterative optimization to generate model parameters for identifying the sensitivity level of target data.
[0090] First, extract multi-dimensional feature vectors and their corresponding sensitivity annotation labels from the model training data. The feature vectors serve as the input to the model, and the labels serve as the supervision signals to guide the learning process of the model. At the initial stage of training, select an appropriate machine learning algorithm to construct an initial model, such as a deep neural network, a support vector machine, or a random forest model, and set the initial model parameters. Taking a deep neural network as an example, the initial model can consist of several layers of fully connected networks, and the number of nodes in each layer is designed according to the dimension of the feature vectors and the complexity requirements of the model.
[0091] Next, by inputting the training data into the model in batches, calculate the error between the sensitivity label predicted by the model and the true label, and adjust the model parameters according to the error. The error calculation can use common loss functions, such as the cross-entropy loss function, to quantify the difference between the predicted value and the true label. Based on the error feedback, use an optimization algorithm (such as the stochastic gradient descent method or the Adam optimization algorithm) to update the model parameters, so that the prediction result of the model gradually approaches the true label.
[0092] To further improve the generalization ability of the model, regularization techniques, such as L2 regularization and the Dropout method, are introduced during the training process to reduce the risk of overfitting. In addition, monitor the performance change of the model by using a validation set to judge whether the training process tends to converge. When the loss on the validation set no longer decreases significantly, stop the training process and save the current model parameters.
[0093] After completing the model training, evaluate the performance of the model, and use the test set to calculate key indicators such as the sensitivity recognition accuracy, recall rate, and F1 score of the model to verify the sensitivity classification ability of the model for unseen data. If the performance indicators do not meet the expectations, the model can be retrained by adjusting the model structure, optimization algorithm, or hyperparameters until the model meets the requirements.
[0094] Finally, save the trained sensitive data detection model and its parameters. This model can accurately predict the sensitivity label according to the feature vectors of the input data and provide support for the sensitive data detection in the subsequent steps. The entire model construction and training process aims at obtaining the optimal parameters to ensure that the model has stable prediction performance and efficient recognition ability.
[0095] Furthermore, the sensitive data detection model includes a feature extraction layer, a feature fusion layer, and a sensitivity evaluation layer;
[0096] Among them, the feature extraction layer is used to receive the preprocessed target structured data; the feature extraction layer uses parallel spatial feature extraction branches and attribute feature extraction branches to process the received data respectively; among them, the spatial feature extraction branch is implemented using a spatial convolutional network, and is used to extract the position distribution features and spatial correlation features of each data point through multiple layers of convolution, and generate a spatial feature vector; the attribute feature extraction branch is implemented using an attention mechanism, and is used to capture the temporal correlation features and activity levels of the data points, and generate an attribute feature vector.
[0097] The feature fusion layer is used to receive the spatial feature vector and the attribute feature vector provided by the feature extraction layer; the feature fusion layer is implemented using a bidirectional gated recurrent unit network, processes the received data, and generates a unified comprehensive feature vector.
[0098] The sensitivity evaluation layer is used to receive the comprehensive feature vector provided by the feature fusion layer; the sensitivity evaluation layer is used to input the comprehensive feature vector into a multi-layer fully connected network, calculate the sensitivity score of each data point; calibrate the scoring result according to the predefined sensitivity rules; based on the sensitivity score and the result of rule calibration, use a dynamic confidence threshold to determine whether the data point belongs to a specific sensitivity category, and generate a sensitivity classification label.
[0099] In this embodiment, the sensitive data exploration model realizes the efficient processing and sensitivity analysis of the target data through the collaborative action of the feature extraction layer, the feature fusion layer and the sensitivity evaluation layer. Each component of the model has a clear function and implementation method, and the data flow and processing logic between the components are closely connected to ensure the efficient implementation of the overall function.
[0100] First of all, the feature extraction layer receives the preprocessed target structured data, including the geographical location, time attributes and other relevant features of the data points. The feature extraction layer processes the data through parallel spatial feature extraction branches and attribute feature extraction branches. The spatial feature extraction branch is implemented using a spatial convolutional network, which extracts the position distribution features and spatial correlation features of each data point through multiple layers of convolution operations. Specifically, the spatial convolutional network takes the geographical location of the data point as the input, captures the local distribution pattern of the data point through the convolutional kernel with a local receptive field, and at the same time extracts the global association information between the data point and other regions through the receptive field that expands layer by layer. The output spatial feature vector comprehensively reflects the distribution characteristics and correlation of the data point in the spatial dimension.
[0101] The attribute feature extraction branch is implemented using an attention mechanism to capture the temporal correlation features and activity levels of data points. The attention mechanism evaluates the key time points in the time series through dynamic weight allocation, thereby highlighting the importance of recent data and generating an attribute feature vector by combining the time series context information. This feature vector can accurately describe the dynamic performance of data points in the time dimension and provide high-quality input for subsequent feature fusion.
[0102] The feature fusion layer is used to receive the spatial feature vector and the attribute feature vector and perform fusion processing on these two types of features through a bidirectional gated recurrent unit network (GRU). The GRU network has high efficiency and robustness in processing time series data and multi-dimensional input features. Through the bidirectional structure, the network can consider both the forward and backward information of the time series to capture the complex correlations between features. The fused output is a comprehensive feature vector that uniformly contains the spatial distribution features, temporal dynamic features, and other related characteristics of the data points.
[0103] After receiving the comprehensive feature vector, the sensitivity evaluation layer uses a multi-layer fully connected network to calculate the sensitivity score for each data point. The fully connected network maps the comprehensive feature vector to a score value through layer-by-layer weight mapping and activation functions, indicating the likelihood of the data point belonging to a specific sensitivity category. Subsequently, the scoring result is calibrated according to predefined sensitivity rules, such as setting priorities or minimum scoring thresholds for specific categories. Finally, through the judgment of a dynamic confidence threshold, the sensitivity classification label of each data point is determined. The dynamic confidence threshold is adjusted according to the data distribution and application scenario to ensure the accuracy and applicability of the classification result.
[0104] Through the above process, this model realizes a complete data processing flow from feature extraction to classification annotation, can efficiently identify sensitive data and label its sensitivity category in complex multi-dimensional data scenarios, and provides strong support for subsequent data management and protection measures.
[0105] Furthermore, the spatial convolutional network used in the spatial feature extraction branch includes an input layer, a local receptive field convolutional layer, a cross-region convolutional layer, a multi-scale feature fusion layer, and a normalization and output layer;
[0106] Among them, the input layer is used to receive the preprocessed target structured data, organize the data containing the geographical location information and related regional features of each data point into a two-dimensional matrix form; each row of the two-dimensional matrix represents the spatial information of a data point, and each column represents the specific features of the data point, generating an input feature matrix;
[0107] The local receptive field convolutional layer is used to process the input feature matrix. By setting a convolutional kernel with a fixed range, it performs a convolutional operation on the features of each data point and its neighboring data points to extract local spatial distribution characteristics and generate a preliminary local feature map, which is used to describe the neighborhood distribution pattern of the data points and the feature distribution density of the local area.
[0108] The cross-region convolutional layer is used to receive the local feature map and, through a convolutional operation that expands the receptive field range, extract the potential spatial correlation between data points in a large range. Using multiple convolutional kernels of different sizes, it extracts cross-region features at different receptive field scales to generate a global spatial feature map, which reflects the correlation and spatial characteristics between distant data points.
[0109] The multi-scale feature fusion layer is used to receive the local feature map and the global spatial feature map and fuse the two according to a specified weight ratio. The fused comprehensive feature map provides a unified description of the local and global relationships of the data points.
[0110] The normalization and output layer is used to perform batch normalization on the comprehensive feature map to balance the numerical differences between the output features of different convolutional layers and generate a normalized spatial feature vector.
[0111] In this embodiment, the spatial convolutional network of the spatial feature extraction branch realizes the efficient processing and comprehensive feature extraction of geographical data with its hierarchical structure. The entire network consists of an input layer, a local receptive field convolutional layer, a cross-region convolutional layer, a multi-scale feature fusion layer, and a normalization and output layer. Each part is interconnected to ensure that the spatial information of the data points is completely modeled.
[0112] First, the input layer is responsible for receiving the preprocessed target structured data and organizing it in the form of a two-dimensional matrix. Each row corresponds to a geographical data point, representing the spatial information of the data point, such as geographical coordinates, region type, or relevant additional feature information. Each column represents a specific feature dimension, such as the relevance index of a specific facility or other numerical attributes of the data point. The input layer provides a basis for subsequent convolutional operations by converting the original data into an input feature matrix, ensuring that the data is transmitted in a structured form.
[0113] The local receptive field convolutional layer processes the input feature matrix by setting a convolutional kernel with a fixed range to extract local spatial distribution characteristics. In this process, the size of the convolutional kernel determines the neighborhood range covered in each operation. For example, a 3×3 convolutional kernel can capture the spatial distribution pattern of a data point and its closest neighboring data points. Through convolutional operations, the local features of each region are calculated to generate a preliminary local feature map. The local feature map reflects the feature distribution density of data points within a small range, such as the point density in the neighborhood or the correlation between adjacent data points. The local receptive field convolutional layer can capture fine-grained spatial information and is suitable for analyzing feature differences within local regions.
[0114] The cross-region convolutional layer further expands the receptive field range to extract potential global correlations between data points. This layer performs operations based on multiple convolutional kernels of different sizes, such as 5×5 or 7×7 convolutional kernels, which capture medium-range and larger-range spatial relationships respectively. Through these convolutional kernels of different scales, the cross-region convolutional layer can generate a global spatial feature map that reflects the correlations between distant data points. For example, a larger convolutional kernel can identify the relationship between two distant but potentially functionally related data points. The cross-region convolutional layer effectively integrates the global spatial characteristics of the data and provides rich global information for subsequent feature fusion.
[0115] The multi-scale feature fusion layer is responsible for fusing the local feature map and the global spatial feature map according to the specified weight ratio. During the fusion process, through the method of weighted average, the contribution ratios of local features and global features are dynamically adjusted to meet the requirements of specific application scenarios. For example, in some applications that emphasize details, local features may account for a larger proportion, while in scenarios that require considering the overall trend, the weight of global features may be higher. The fused comprehensive feature map retains both the neighborhood distribution information of data points and the global spatial pattern, providing a unified and comprehensive feature description.
[0116] Finally, the normalization and output layer performs batch normalization on the comprehensive feature map to balance the numerical differences between the output features of different convolutional layers. The normalization operation can adjust the distribution of feature values, such as scaling them to a specific range (such as 0 to 1 or a standard normal distribution), to eliminate the impact of scale differences between feature dimensions on subsequent analysis. After normalization, the generated output is a normalized spatial feature vector, with each data point corresponding to a feature vector, which is used to describe its comprehensive characteristics in the spatial dimension.
[0117] Through the collaborative work of the above layers, the spatial convolutional network can efficiently process complex geographical data and comprehensively model the spatial information of data points from both local and global levels, providing high-quality input for subsequent machine learning models.
[0118] Furthermore, the attention mechanism of the attribute feature extraction branch includes the following implementation process:
[0119] Receive the preprocessed target structured data and represent the time attribute as a time series feature vector;
[0120] Calculate the context weight for the feature vector at each time point, and dynamically allocate weights according to time differences and feature similarities;
[0121] Use the context weight to perform weighted aggregation on the time series feature vector to generate an enhanced time feature vector;
[0122] Integrate all enhanced time feature vectors into a global time series feature representation and output it to the feature fusion layer for further processing.
[0123] In this embodiment, the attention mechanism of the attribute feature extraction branch is implemented through a series of clear steps, aiming to extract time series features from the time attributes of the target data, generate high-quality time series feature representations, and provide support for subsequent feature fusion and sensitivity analysis. The entire process is based on the preprocessed target structured data, and realizes the extraction and enhancement of time series features through dynamic weight allocation and feature aggregation.
[0124] First, receive the preprocessed target structured data, extract and represent the time attribute of the data as a time series feature vector. The time attribute usually includes the timestamp information of the data point, such as the time when an event occurs or the time point when a certain geographical data point is recorded. By parsing the timestamp and combining other features of the data point, the relevant information of each time point is encoded as a feature vector. For example, if the features of a data point at different time periods include activity level, the number of associated facilities, and regional change conditions, this information can be represented as a multi-dimensional vector to form a time series feature set.
[0125] Next, calculate the context weight for the feature vector at each time point, which is the core of the attention mechanism. The calculation of the context weight comprehensively considers the time differences and feature similarities between time points. The time difference represents the distance between different time points, measured in hours or days, for example, while the feature similarity is quantified by comparing the values of the time point feature vectors. Through this combination, the attention mechanism can dynamically allocate weights, highlighting the context information more relevant to the current time point while weakening the features in the far future or with low relevance. For example, when two time points are adjacent and have similar features, their context weights are higher, while the weights of time points that are far apart or have large feature differences are lower.
[0126] After calculating the context weights, these weights are used to perform weighted aggregation on the time series feature vectors. The purpose of weighted aggregation is to generate enhanced time feature vectors, such that the features at each time point not only contain their own information but also incorporate the key information in their context. For example, through the aggregation operation, the activity level at a certain time point not only depends on the feature value at that point but is also affected by other associated time points, and this operation makes the generated feature vectors better reflect the global dynamics of the time series.
[0127] Finally, the enhanced feature vectors at all time points are integrated into a global time series feature representation. This integration process is achieved by arranging or concatenating all the feature vectors in the time series, forming a unified time series feature matrix or a set of vectors. The global time series feature representation not only retains the local time information of the data points but also can capture the dynamic patterns and long-term trends in the entire time series. The generated global time series features are used as the output and passed to the feature fusion layer for further processing after being combined with the spatial features.
[0128] Through the above implementation process, the attention mechanism ensures that the extraction of time series features is highly adaptable and flexible. Its dynamic weight allocation mechanism can highlight the feature contributions of key time points according to the time and feature relationships of the actual data, thereby providing accurate time series information support for the sensitive data exploration model.
[0129] Furthermore, the bidirectional gated recurrent unit network used in the feature fusion layer jointly models the correlation and spatial distribution pattern in the time dimension of the spatial feature vector and the attribute feature vector by introducing a dynamic context awareness mechanism, where the dynamic context awareness mechanism dynamically adjusts the weights of the time series features according to the regional importance of the spatial features.
[0130] In this embodiment, the feature fusion layer uses a bidirectional gated recurrent unit network (GRU) combined with a dynamic context awareness mechanism to jointly model the input spatial feature vector and attribute feature vector to capture the influence of spatial features in different regions and the change pattern of time series features in the time dimension, thereby generating more representative comprehensive features.
[0131] First, the spatial feature vector and the attribute feature vector respectively come from the feature extraction layer, and these vectors describe the spatial distribution characteristics of the data points and the dynamic activity level in the time dimension respectively. During the feature fusion process, the bidirectional gated recurrent unit network is used as the core processing unit, which can effectively capture the forward and backward context information in the time series data processing. For example, for a certain time point, the bidirectional GRU not only considers the feature influence of the previous time points but also combines the features of the subsequent time points, making the modeling of the current time point more comprehensive.
[0132] The introduction of the dynamic context awareness mechanism makes the fusion process more flexible and adaptive. Specifically, the dynamic context awareness mechanism dynamically adjusts the weights of the temporal features according to the importance information of each region in the spatial feature vector. The importance of a region can be determined by analyzing the density information or spatial correlation values in the spatial feature vector. For example, if the spatial feature vector of a certain region indicates that it contains multiple highly sensitive facilities (such as a substation or an airport), then a higher weight is assigned to the importance of this region, while regions with lower density or fewer facilities have lower weights. This weight information is introduced into the weighted calculation of the temporal feature vector through the dynamic context awareness mechanism, ensuring that regions with high importance contribute more to the result in the processing of temporal features.
[0133] Taking a specific example, assume that the data points have the following spatial feature weights in three regions: the weight of region A is 0.6, region B is 0.3, and region C is 0.1. At a certain time point, the temporal feature vector contains the activity level, historical trend, and related events at this time point. The dynamic context awareness mechanism weights these temporal features according to the spatial weights. For example, in region A, the activity level feature at this time point will be amplified, while in region C, the impact of this feature on the final comprehensive feature will be correspondingly weakened. This weighted adjustment based on the importance of regions ensures that the model can focus on the regions and time points that are more critical for sensitivity judgment.
[0134] The temporal features processed by the dynamic context awareness mechanism are input into a bidirectional GRU network and participate in the fusion modeling together with the spatial features. The bidirectional GRU filters the input weighted temporal features through its gating units, retains the key features, and at the same time suppresses noise or irrelevant information. Through multiple cycles, the network gradually forms the global comprehensive feature at each time point, which not only contains the dynamic information in the time dimension but also combines the regional characteristics in the spatial dimension.
[0135] Finally, the comprehensive feature vector output by the bidirectional GRU network reflects both the dynamic changes of the data points in the time dimension and captures their distribution patterns in the spatial dimension, providing high-quality input data for subsequent sensitivity assessment. This feature fusion method combined with the dynamic context awareness mechanism can effectively improve the adaptability and recognition ability of the sensitive data exploration model for complex spatio-temporal data.
[0136] Furthermore, the dynamic context awareness mechanism of the bidirectional gated recurrent unit network used in the feature fusion layer is implemented through the following steps:
[0137] Adopt the following formula (8) to calculate the regional importance weight of each spatial region according to the specific region feature value in the spatial feature vector :
[0138] ;
[0139] wherein, represents the total number of spatial regions; represents the normalized weight of the spatial region, used to quantify the importance of this region;
[0140] According to the following formula (9), calculate the dynamic spatio-temporal weight for the attribute features at each time point in the attribute feature vector : :
[0141] ;
[0142] where is the time decay coefficient, is the context relevance degree between the spatial region and the time point ;
[0143] According to the following formula (10), generate the comprehensive feature representation :
[0144] ;
[0145] where represents the total number of time points; is the balance parameter, used to adjust the contribution ratio of the spatial feature and the attribute feature in the comprehensive feature vector; is the spatial feature vector.
[0146] First, according to the regional feature value in the spatial feature vector , calculate the regional importance weight of each spatial region. Formula (8) is defined as follows:
[0147] ;
[0148] where, represents the specific attribute value of the region in the spatial feature vector, such as the number of sensitive facilities in the region or the importance index of the region. represents the total number of spatial regions, is the normalized weight of the region , used to quantify the importance of this region. Through the normalization operation, ensure that the sum of the importance weights of all regions is 1. For example, assume the spatial feature vector , and the feature values corresponding to three regions are respectively If it is 30, then the regional weight is calculated as . This means that Region 2 contributes the most to the comprehensive feature.
[0149] Next, the dynamic spatio-temporal weight of each time point in the attribute feature vector is calculated according to Equation (9) . The equation is defined as follows:
[0150] ;
[0151] where is the spatio-temporal weight of time point , which combines the normalized weight of each spatial region and the context relevance between Region and time point . The context relevance can be defined by the actual application scenario, such as representing the event occurrence frequency or spatial activity of Region at time point . The parameter is the time decay coefficient, which is used to control the influence degree of the time interval on the weight. A larger value will make the influence of the time difference more significant, and the value range is usually from 0.1 to 1.0. For example, assuming that the regional weight , the context relevance , and , then for time point , the dynamic weight is calculated as: ;
[0152] ;
[0153] The calculation result is .
[0154] Finally, based on the dynamic spatio-temporal weight and the attribute feature vector , combined with the spatial feature vector , the comprehensive feature representation is generated. Equation (10) is defined as follows:
[0155] ;
[0156] where represents the total number of time points, is the eigenvalue of time point in the attribute feature vector, which usually includes the dynamic activity degree of time point or other relevant information. The parameter is the balance coefficient, which is used to adjust the weights of spatial features and attribute features in the comprehensive features. The value of can be set according to specific application requirements, usually between 0.5 and 1.0. For example, for the attribute features at three time points , the dynamic spatio-temporal weight , the spatial feature , the balance coefficient , the calculation of the comprehensive feature representation is:
[0157] ;
[0158] The final comprehensive feature representation combines the time point features and the spatial region features, providing a comprehensive and accurate representation of sensitive data.
[0159] Through the above steps, the dynamic context awareness mechanism can flexibly combine spatial and temporal information to generate a highly adaptable comprehensive feature representation, providing strong support for subsequent sensitivity analysis.
[0160] Step S104: Input the target data set into the sensitive data exploration model to obtain the model output result; according to the model output result, label the sensitive data included in the target data set to generate a sensitivity classification label; wherein, the target data set uses the preprocessed target structured data.
[0161] In step S104, input the target data set into the trained sensitive data exploration model. The model generates the corresponding model output result according to the feature vector of the input data, and further labels the sensitive data included in the target data set to generate a sensitivity classification label.
[0162] First, the target data set goes through a preprocessing step to form input features consistent with the structured data format used during model training. These features are represented as multi-dimensional feature vectors, which include feature dimensions such as position distribution, spatial correlation, and data frequency, ensuring a perfect match with the model's input format. The target data set is input into the trained exploration model one by one or in batches, and the model analyzes and infers the feature vectors of each record using the previously learned parameter weights.
[0163] The model inference process is based on the forward propagation mechanism, that is, by passing through the various layers of the model in sequence, the input feature vector is operated on with the model's weights to calculate the probability distribution of each sensitivity category. For each input data, the model outputs a probability distribution vector, where each element represents the probability that the data belongs to the corresponding sensitivity category. For example, the possible classification labels include "highly sensitive", "medium sensitive", and "low sensitive", and the probability vector output by the model indicates the likelihood of the data in these classifications.
[0164] Next, according to the probability distribution output by the model, select the category with the highest probability value as the preliminary sensitivity classification label for this data. To improve the reliability of the classification label, business rules or dynamic confidence thresholds can be combined to further judge some boundary cases. For example, when the probabilities of multiple classifications are close, further weighted adjustment can be made according to specific feature dimensions (such as spatial correlation) to ensure that the classification results better meet the actual application requirements.
[0165] After classification, attach the generated sensitivity classification label to the target dataset and record the classification results of each piece of data in a structured manner. These classification labels can be stored in the database in the form of additional fields, corresponding one by one to the original data, thus forming a complete labeled dataset. The finally generated sensitivity classification label provides a clear basis for subsequent hierarchical protection and data management, and ensures the comprehensive coverage and accurate annotation of the target dataset during the exploration process.
[0166] Step S104, through the inference ability of the model, accurately associates each record in the target dataset with the corresponding sensitivity classification label, laying a foundation for the subsequent processing of sensitive data, and the whole process is very efficient.
[0167] Step S105: According to the sensitivity classification label, perform hierarchical protection on the detected sensitive data, including encryption storage and access control policies for highly sensitive data.
[0168] In step S105, according to the sensitivity classification label generated in step S104, perform hierarchical protection on the detected sensitive data to ensure the security of the data and the controllability of access, especially implementing strict encryption storage and access control policies for highly sensitive data.
[0169] First, divide all data into different protection levels according to the sensitivity classification label, including highly sensitive data, moderately sensitive data, and lowly sensitive data. The protection measures for each category are gradually strengthened according to their sensitivity levels to ensure the security and compliance of sensitive data. For highly sensitive data, strong encryption storage measures are taken, such as using the Advanced Encryption Standard (AES) algorithm to store the data in ciphertext in the database or distributed storage system. The encryption process uses dynamically generated keys, and each access requires key verification, thus reducing the risk of key leakage. The generation and management of keys can be achieved through a secure Key Management System (KMS) to ensure the secure storage and access of keys.
[0170] After the storage is completed, strict access control policies are set for accessing highly sensitive data. Specifically, a multi-factor authentication mechanism is adopted to ensure the credibility of the visitor's identity. For example, verification is carried out through a combination of passwords, fingerprints, or one-time verification codes. At the same time, role-based access control is applied to restrict the access rights of different user roles, and only authorized users are allowed to access the data related to their responsibilities. The allocation of access rights can be dynamically adjusted through the permission management module to adapt to changes in user roles or responsibilities.
[0171] For moderately sensitive data, a lightweight encryption method is adopted, such as a simplified algorithm based on symmetric encryption, to ensure the secure storage of the data. At the same time, simple access control based on IP address or regional restrictions is adopted to reduce unauthorized external access. The access rights of moderately sensitive data are mainly for internal authorized users, and access logs are retained to monitor abnormal access behaviors.
[0172] For low-sensitive data, encryption storage is usually not required, but an access record function needs to be set. The identity, access time, and operation type of each data access are recorded for subsequent auditing and analysis. Although low-sensitive data is not directly involved in high risks, the access record can provide basic security protection and discover potential risk behaviors.
[0173] Finally, by regularly auditing and optimizing the hierarchical protection mechanism, the accuracy of data classification and the effectiveness of protection measures are ensured. The auditing process can evaluate data storage, encryption strength, access behaviors, and permission settings based on automated tools, and dynamically adjust the protection measures according to the update of security policies and the emergence of new threats. Through this hierarchical protection mechanism, the detected sensitive data can be comprehensively and effectively protected, the security of highly sensitive data can be ensured, and at the same time, resource utilization and management efficiency can be optimized.
[0174] In the above embodiments, a method for detecting sensitive data based on machine learning is provided. Correspondingly, the present application also provides a system for detecting sensitive data based on machine learning. Please refer to Figure 2 , which is a schematic diagram of an embodiment of a system for detecting sensitive data based on machine learning according to the present application. Since this embodiment, that is, the second embodiment, is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system embodiments described below are only illustrative.
[0175] A system for detecting sensitive data based on machine learning provided by the second embodiment of the present application includes:
[0176] The acquisition unit 201 is configured to acquire original data including pole tower location data, airport location data, and important station location data from power-related data sources; perform preprocessing on the original data, where the formatting process includes cleaning outliers, unifying data formats, and standardizing geographic coordinates to generate preprocessed structured data;
[0177] The generation unit 202 is configured to generate a multi-dimensional feature vector based on the preprocessed structured data; combine the multi-dimensional feature vector with its corresponding sensitivity annotation label according to a predefined sensitivity annotation rule to generate model training data; where the multi-dimensional feature vector includes location distribution, spatial correlation, and data frequency;
[0178] The construction unit 203 is configured to construct a sensitive data detection model using a machine learning algorithm based on the model training data; obtain model parameters for identifying the sensitivity level of target data by iteratively optimizing and training the sensitive data detection model; obtain a trained sensitive data detection model according to the model parameters;
[0179] The annotation unit 204 is configured to input a target data set into the sensitive data detection model to obtain a model output result; annotate the sensitive data included in the target data set according to the model output result to generate a sensitivity classification label; where the target data set uses preprocessed target structured data;
[0180] The classification unit 205 is configured to perform hierarchical protection on the detected sensitive data according to the sensitivity classification label, including encryption storage and access control policies for highly sensitive data.
[0181] The third embodiment of the present application provides an electronic device, which includes:
[0182] A processor;
[0183] A memory for storing a program, which when read and executed by the processor, executes a method for detecting sensitive data based on machine learning provided in the first embodiment of the present application.
[0184] The fourth embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it executes a method for detecting sensitive data based on machine learning provided in the first embodiment of the present application.
[0185] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.
Claims
1. A sensitive data detection method based on machine learning, characterized in that: include: Collect raw data including tower location data, airport location data, and important station location data from power-related data sources; Preprocessing the raw data, wherein the preprocessing includes cleaning outliers, unifying data formats and standardizing geographic coordinates to generate preprocessed structured data; Based on the preprocessed structured data, a multidimensional feature vector is generated; according to a predetermined sensitivity labeling rule, the multidimensional feature vector is combined with its corresponding sensitivity labeling label to generate model training data; wherein the multidimensional feature vector includes location distribution, spatial correlation, and data frequency; Based on the model training data, a sensitive data detection model is constructed using a machine learning algorithm; the sensitive data detection model is trained by iterative optimization to obtain model parameters for identifying the sensitivity of target data; and a trained sensitive data detection model is obtained according to the model parameters; Input the target data set into the sensitive data detection model to obtain the model output result; according to the model output result, label the sensitive data included in the target data set to generate a sensitivity classification label; wherein the target data set uses the preprocessed target structured data; According to the sensitivity classification labels, the detected sensitive data is protected in different levels, including encryption storage and access control strategies for highly sensitive data; The step of generating a multi-dimensional feature vector based on the pre-processed structured data includes: Divide the input geographic data into equidistant grid cells according to a preset rule, wherein each grid cell corresponds to a specified geographic range; calculate the number of geographic data points in each grid cell, expressed as grid density; normalize the grid density value to generate a location distribution feature reflecting the local distribution feature of the geographic data; For each geographic data point, calculate its distance from a set of predefined important facilities; assign weights according to facility types, perform weighted average of each distance value and the corresponding weight to obtain a spatial correlation value; normalize the spatial correlation values of all geographic data points to generate spatial correlation features; wherein the important facilities include towers, airports and substations; The time attributes of each geographic data point are analyzed, and the number of occurrences within a fixed time period is counted as the frequency value; by setting the time decay weight, the frequency value is weighted and normalized to generate the spatiotemporal frequency characteristics that describe the spatiotemporal activity level; The normalized location distribution features, spatial correlation features, and spatiotemporal frequency features of each geographic data point are combined to generate a multidimensional feature vector, and a feature vector dataset is output for model training.
2. The sensitive data detection method based on machine learning according to claim 1, characterized in that: For each geographic data point, the distance between the point and a set of predefined important facilities is calculated; weights are assigned according to facility types, and each distance value is weighted averaged with the corresponding weight to obtain a spatial association value; The spatial correlation values of all geographic data points are normalized to generate spatial correlation features, including: For each geographic data point , calculate its important facilities set Each facility distance, generate a distance matrix , among which, Line Column value Indicates geographic data points to The distance to the facility, Calculate according to the following formula (1): ; in, and Respectively geographic data points and The location coordinates of each facility; According to Risk factor of important facilities , a dynamic weight is assigned to each facility according to the following formula (2): ; in, For the The weight coefficient of each important facility; The number of important facilities; According to the following formula (3), the spatial association value of the geographic data point is calculated : ; in, is the attenuation coefficient, which is used to control the effect of distance on spatial correlation; For the matrix No. The maximum distance value of the column, used for normalized distance; According to the following formula (4), the spatial correlation value of all geographic data points is Perform normalization: ; in, Normalized spatial correlation value; For all spatial association values The minimum value in ; For all spatial association values The maximum value in ; A spatial correlation feature is generated according to the normalized spatial correlation value.
3. The sensitive data detection method based on machine learning according to claim 1, characterized in that: The time attribute of each geographic data point is analyzed, and the number of occurrences within a fixed time period is counted as a frequency value; by setting a time decay weight, the frequency value is weighted and normalized to generate a spatiotemporal frequency feature describing the spatiotemporal activity, including: For geographic data points at fixed time periods The number of occurrences within , according to the decay weight at the current time point Weighted, where the attenuation weight Determined by the following formula (5): ; in, Indicates the current time and time period time interval; is the attenuation control parameter, which is used to adjust the weight impact of long-term data; It is an adjustable exponential parameter used to control the nonlinear degree of attenuation; In each time period According to the following formula (6), the weighted frequency value is calculated: : ; in, is the total number of time periods; It is the weighted cumulative frequency value of geographic data points in all time periods; According to the following formula (7), the normalized spatiotemporal frequency features are generated: : ; in, is the weighted frequency mean of all time periods; is the standard deviation of the weighted frequency values, which is used to measure the difference in activity of a data point relative to the overall distribution.
4. The sensitive data detection method based on machine learning according to claim 1, characterized in that: The sensitive data detection model includes a feature extraction layer, a feature fusion layer and a sensitivity assessment layer; The feature extraction layer is used to receive the pre-processed target structured data; the feature extraction layer uses a parallel spatial feature extraction branch and an attribute feature extraction branch to process the received data respectively; the spatial feature extraction branch is implemented using a spatial convolutional network, which is used to extract the position distribution characteristics and spatial correlation characteristics of each data point through multi-layer convolution to generate a spatial feature vector; the attribute feature extraction branch is implemented using an attention mechanism to capture the temporal correlation characteristics and activity level of the data point to generate an attribute feature vector; The feature fusion layer is used to receive the spatial feature vector and the attribute feature vector provided by the feature extraction layer; the feature fusion layer is implemented using a bidirectional gated recurrent unit network to process the received data and generate a unified comprehensive feature vector; The sensitivity assessment layer is used to receive the comprehensive feature vector provided by the feature fusion layer; the sensitivity assessment layer is used to input the comprehensive feature vector into a multi-layer fully connected network to calculate the sensitivity score of each data point; the scoring result is calibrated according to predefined sensitivity rules; based on the sensitivity score and rule calibration results, a dynamic confidence threshold is used to determine whether the data point belongs to a specific sensitivity category, and a sensitivity classification label is generated.
5. The sensitive data detection method based on machine learning according to claim 4 is characterized in that: The spatial convolutional network used by the spatial feature extraction branch includes an input layer, a local receptive field convolutional layer, a cross-region convolutional layer, a multi-scale feature fusion layer, a normalization and output layer; The input layer is used to receive the preprocessed target structured data, organize the data containing the geographical location information and related regional features of each data point into a two-dimensional matrix form; each row of the two-dimensional matrix represents the spatial information of a data point, and each column represents the specific features of the data point, thereby generating an input feature matrix; The local receptive field convolution layer is used to process the input feature matrix, and by setting a convolution kernel with a fixed range, performs a convolution operation on the features of each data point and its neighboring data points, extracts the local spatial distribution characteristics, and generates a preliminary local feature map, which is used to describe the neighborhood distribution pattern of the data points and the feature distribution density of the local area; The cross-region convolution layer is used to receive the local feature map and extract the potential spatial correlation between data points in a large range by expanding the convolution operation of the receptive field range; using multiple convolution kernels of different sizes, cross-region features are extracted at different receptive field scales to generate a global spatial feature map to reflect the correlation and spatial characteristics between distant data points; The multi-scale feature fusion layer is used to receive the local feature map and the global spatial feature map, and fuse the two according to a specified weight ratio; the fused comprehensive feature map provides a unified description of the local and global relationship of the data points; The normalization and output layer is used to perform batch normalization on the comprehensive feature map, balance the numerical differences between the output features of different convolutional layers, and generate a normalized spatial feature vector.
6. The sensitive data detection method based on machine learning according to claim 4 is characterized in that: The attention mechanism of the attribute feature extraction branch includes the following implementation process: receiving the preprocessed target structured data, and representing the time attribute as a time series feature vector; Calculate the context weight for the feature vector at each time point and dynamically assign weights based on time differences and feature similarities; Use context weights to perform weighted aggregation on time series feature vectors to generate enhanced time feature vectors; All enhanced temporal feature vectors are integrated into a global temporal feature representation and output to the feature fusion layer for further processing.
7. The sensitive data detection method based on machine learning according to claim 4 is characterized in that: The bidirectional gated recurrent unit network used in the feature fusion layer jointly models the correlation and spatial distribution pattern of the spatial feature vector and the attribute feature vector in the time dimension by introducing a dynamic context-aware mechanism, wherein the dynamic context-aware mechanism dynamically adjusts the weight of the temporal feature according to the regional importance of the spatial feature.
8. The sensitive data detection method based on machine learning according to claim 7 is characterized in that: The dynamic context-aware mechanism of the bidirectional gated recurrent unit network used in the feature fusion layer is implemented by the following steps: Using the following formula (8), according to the spatial eigenvector The specific area characteristic value in Calculate the regional importance weight for each spatial region : ; in, Represents the total number of spatial regions; Represents a spatial region The normalized weight of is used to quantify the importance of the region; According to the following formula (9), the attribute feature of each time point in the attribute feature vector Calculating dynamic spatiotemporal weights : ; in is the time attenuation coefficient, For space area With time point Contextual relevance; According to the following formula (10), a comprehensive feature representation is generated : ; in Indicates the total number of time points; It is a balance parameter used to adjust the contribution ratio of spatial features and attribute features in the comprehensive feature vector; is the spatial feature vector.
9. A sensitive data detection system based on machine learning, characterized in that: include: A collection unit, used to collect raw data including tower location data, airport location data, and important station location data from power-related data sources; Preprocessing the raw data, wherein the preprocessing includes cleaning outliers, unifying data formats and standardizing geographic coordinates to generate preprocessed structured data; A generating unit, configured to generate a multidimensional feature vector based on the preprocessed structured data; and to combine the multidimensional feature vector with its corresponding sensitivity label according to a predetermined sensitivity labeling rule to generate model training data; wherein the multidimensional feature vector includes position distribution, spatial correlation, and data frequency; A construction unit is used to construct a sensitive data detection model based on the model training data using a machine learning algorithm; train the sensitive data detection model through iterative optimization to obtain model parameters for identifying the sensitivity of target data; and obtain a trained sensitive data detection model according to the model parameters; A labeling unit, used for inputting a target data set into the sensitive data detection model to obtain a model output result; labeling the sensitive data included in the target data set according to the model output result to generate a sensitivity classification label; wherein the target data set adopts the preprocessed target structured data; The classification unit is used to perform classified protection on the detected sensitive data according to the sensitivity classification labels, including encryption storage and access control strategies for highly sensitive data; Wherein, the generating unit is specifically used for: Divide the input geographic data into equidistant grid cells according to a preset rule, wherein each grid cell corresponds to a specified geographic range; calculate the number of geographic data points in each grid cell, expressed as grid density; normalize the grid density value to generate a location distribution feature reflecting the local distribution feature of the geographic data; For each geographic data point, calculate its distance from a set of predefined important facilities; assign weights according to facility types, perform weighted average of each distance value and the corresponding weight to obtain a spatial correlation value; normalize the spatial correlation values of all geographic data points to generate spatial correlation features; wherein the important facilities include towers, airports and substations; The time attributes of each geographic data point are analyzed, and the number of occurrences within a fixed time period is counted as the frequency value; by setting the time decay weight, the frequency value is weighted and normalized to generate the spatiotemporal frequency characteristics that describe the spatiotemporal activity level; The normalized location distribution features, spatial correlation features, and spatiotemporal frequency features of each geographic data point are combined to generate a multidimensional feature vector, and a feature vector dataset is output for model training.
Citation Information
Patent Citations
Power sensitive data processing method and device, electronic equipment and storage medium
CN116628584A