Data security access method and system based on distributed structure

By constructing a global transaction database in a distributed system and using entropy analysis and random forest algorithms to identify abnormal behavior and dynamically adjust access permissions, the problem of lagging cross-node threat identification in existing technologies is solved, achieving high-precision security protection for distributed systems.

CN120979790APending Publication Date: 2025-11-18KUNHE DIGITAL TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511306575.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies struggle to identify potential threats across nodes in a timely manner in distributed systems. Static rules and single-dimensional monitoring lead to delays in anomaly detection and the risk of missed detections.

Method used

By acquiring raw data access logs, performing structured processing, and constructing a global transaction database, we analyze the correlation patterns between nodes, use entropy analysis and random forest algorithms to identify abnormal behavior, dynamically adjust access permissions to generate real-time security policies, and restrict access to high-risk nodes.

Benefits of technology

It enables high-precision monitoring and security protection of user behavior in distributed systems, timely identification and blocking of collaborative attack chains, and improves system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979790A_ABST
    Figure CN120979790A_ABST
Patent Text Reader

Abstract

The invention discloses a data security access method based on a distributed structure, and the method comprises the steps: obtaining and processing a data access log, dividing access behavior categories, calculating the access frequency of each node and the distribution of duration according to an access mode classification result, obtaining a node frequency distribution feature, converting the node frequency distribution feature into an object access sequence, and storing the object access sequence in a database; rejecting a pseudo-correlation object combination to obtain a correlation mode; the frequent item set support degree of the relevance mode is calculated, if the support degree exceeds a threshold value, an access behavior topological structure is analyzed, a collaborative attack mode is obtained, an abnormal behavior score is calculated, and if the score is higher than the threshold value, a high-risk node list is generated, abnormal behaviors are judged, and a detection result is obtained; generating a security policy according to the detection result and a preset rule; and if the isolation condition is triggered, distributing a high-risk node list to each node of the distributed system, limiting the access authority of the related node, and obtaining an isolation execution result. According to the method, potential threats in the distributed system can be identified in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data security access technology, and in particular to a data security access method and system based on a distributed structure. Background Technology

[0002] Currently, secure data access has become a crucial component of modern information systems, directly impacting enterprise operational stability, user privacy protection, and national security. In distributed architectures, data is typically stored across multiple nodes, making access control and security management critical for ensuring system security. However, data access patterns in distributed systems are complex and dynamic, with access behaviors adjusting dynamically based on business scenarios, user roles, and time factors, posing significant challenges to traditional secure access management methods.

[0003] In a common technique, static rules are used to define access permissions, and potential abnormal behavior is identified through single-dimensional monitoring metrics (such as single-node access frequency, single-user session count, etc.). Specifically, this type of method typically pre-sets access control lists or role-based permission matrices, and uses a rule engine to determine the legitimacy of access requests. Simultaneously, access logs are collected at the node level, and fixed threshold models (such as frequency thresholds or session over-limit detection) are used to determine if abnormal requests exist. Once an metric is detected to exceed limits, an alarm or blocking mechanism is triggered. However, this method relies on static rules and single-dimensional monitoring, which can lead to limitations when facing dynamically changing access behavior patterns in distributed systems. Static permission rules struggle to adapt to changes in access patterns caused by complex business processes, and fixed threshold detection lacks a global perspective, failing to capture potential correlations between cross-node accesses. This results in the inability to promptly identify potential cross-node threats, and anomaly detection suffers from lag and the risk of missed detections.

[0004] In summary, existing technologies suffer from the problem that potential threats in distributed systems are difficult to identify in a timely manner. Summary of the Invention

[0005] This invention provides a data security access method and system based on a distributed structure to enable timely identification of potential threats in distributed systems.

[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a data security access method based on a distributed structure, comprising: Obtain raw data access logs and perform structured processing on the raw data access logs to obtain an extensible structured access dataset, wherein the raw data access logs include node identifiers, access frequency, and data sensitivity labels; Based on the structured access dataset, access behavior categories are divided to obtain access pattern classification results; Based on the access pattern classification results and the node identifier, the access frequency of each node is calculated, and the distribution of access duration is quantified by entropy analysis to obtain the node frequency distribution characteristics. The node frequency distribution characteristics are transformed into cross-node object access sequences and a global transaction database is constructed. The global transaction database is scanned to remove pseudo-related object combinations and obtain the inter-node association patterns. Calculate the frequent itemset support of the inter-node association pattern. If the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, analyze the access behavior topology between node identifiers to obtain a collaborative attack pattern based on access request type. Based on the cooperative attack mode and the data sensitivity label, an abnormal behavior score is calculated using the random forest algorithm. If the abnormal behavior score is higher than the preset abnormal threshold, a high-risk node list is generated and combined with the preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. Based on the abnormal behavior detection results, the access permissions of the nodes are adjusted according to the preset dynamic adjustment rules, restricting access to the nodes in the high-risk node list and generating a real-time security policy. If the real-time security policy triggers the preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, and the access permissions of the relevant node identifiers are restricted to obtain the isolation execution result.

[0007] In one optional implementation, the step of structuring the original data access logs to obtain a scalable structured access dataset includes: Obtain the raw data access logs, parse the request packets to extract millisecond-level timestamp sequences and global identifiers, and use a distributed message queue to store them to obtain the synchronously collected log dataset; Based on the log dataset, node identifiers and request types are extracted, and hash identifiers are generated. Combined with timestamps, the access frequency per second is calculated to obtain a dataset labeled with request types and frequencies. For the dataset labeled with request type and frequency, regular expressions are used to match sensitive fields to perform data anonymization and verify data integrity, resulting in an anonymized dataset; The de-identified dataset is compressed to obtain the scalable structured access dataset.

[0008] In one optional implementation, the step of classifying access behavior categories based on the structured access dataset to obtain access pattern classification results includes: The user access trajectories in the structured access dataset are segmented, and the number of accesses within each time period is calculated to obtain the access time distribution dataset. The access time distribution dataset is modified by removing outliers and standardizing the data to obtain a standardized access dataset. If the standardized access dataset contains multidimensional features, then dimensionality reduction processing is performed to extract key time and frequency features to obtain the dataset after feature extraction. The dataset after feature extraction is grouped using the k-means clustering algorithm to obtain access pattern classification results. In an optional implementation, the step of calculating the access frequency of each node based on the access pattern classification results and the node identifier, and quantifying the distribution of access duration through entropy analysis to obtain node frequency distribution characteristics includes: Access records are extracted based on the access pattern classification results and the node identifiers, segmented according to a preset time interval, and the number of accesses in each time interval is calculated to obtain a node access frequency dataset. Based on the node access frequency dataset, the access duration entropy value is calculated to obtain the node entropy value dataset. If the entropy value in the node entropy value dataset exceeds the preset entropy value threshold, the entropy value is normalized to obtain a standardized entropy value dataset. Based on the standardized entropy dataset, the node frequency distribution features are divided using the k-means clustering algorithm to obtain the node frequency distribution features.

[0009] In one optional implementation, the step of converting the node frequency distribution characteristics into cross-node object access sequences and constructing a global transaction database, scanning the global transaction database, eliminating pseudo-related object combinations, and obtaining inter-node association patterns includes: Obtain node log data; Obtain node frequency and object access records from the node log data, divide the object access records into time windows, and generate object access sequences between nodes; Based on the object access sequence between nodes, the access sequence within each time window is converted into transaction records to obtain a global transaction database; The global transaction database is scanned using the Apriori algorithm to calculate the frequent itemsets of each object combination. It is then determined whether the lift of the frequent itemsets is greater than a preset lift threshold. If it is greater, the object combination is retained; if it is less, pseudo-related object combinations are removed, and the association pattern between nodes is output.

[0010] In one optional implementation, the calculation of the frequent itemset support of the inter-node association pattern, and if the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, then the access behavior topology between node identifiers is analyzed to obtain a cooperative attack pattern based on access request type, including: Calculate the support of frequent itemsets of the association patterns between nodes. If it exceeds a preset support threshold, use a graph neural network to analyze the topology and request type distribution between nodes in the global transaction database to obtain behavior patterns based on request types. By mining the association rules of the behavioral patterns, and filtering the frequently co-occurring rules according to the preset confidence threshold, a collaborative attack pattern based on access request type is obtained.

[0011] In one optional implementation, the abnormal behavior score is calculated using a random forest algorithm based on the cooperative attack pattern and the data sensitivity label. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack pattern library to determine abnormal behavior, thereby obtaining an abnormal behavior detection result, including: Obtain network logs and extract access behavior data. Combine the data sensitivity tags with the data and use the random forest algorithm to calculate the abnormal threshold of each access behavior to obtain a behavior score. Behavioral pattern analysis is performed based on the behavioral score and the access frequency, and the existence of a collaborative attack is determined by combining the collaborative attack pattern to identify a candidate list of high-risk nodes. If the behavior score in the candidate list of high-risk nodes exceeds the preset abnormal score threshold, then the node risk assessment is performed in conjunction with the data sensitivity label to generate a list of high-risk nodes. By comparing the attack characteristics of the high-risk node list with a pre-established collaborative attack pattern library, it is determined whether there is any abnormal behavior, and an abnormal behavior detection result is generated.

[0012] In one optional implementation, the step of adjusting node access permissions based on the abnormal behavior detection results and a preset dynamic adjustment rule to restrict access to nodes in the high-risk node list and generate a real-time security policy includes: Based on the list of high-risk nodes, obtain the access permission data of the corresponding nodes, dynamically adjust the rules to calculate the permission restriction ratio, and obtain the updated access permission configuration; Based on the updated access permission configuration, the access control rules for high-risk nodes are adjusted to determine the real-time policy content for node access restrictions, thereby obtaining a real-time security policy. In one optional implementation, if the real-time security policy triggers a preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, and access permissions for the relevant node identifiers are restricted to obtain the isolation execution result, including: If the real-time security policy triggers the preset isolation trigger condition, then according to the high-risk node list, the risk value of each node is sorted using the Naive Bayes algorithm to obtain an ordered node list. The ordered list of nodes is distributed to each node in the distributed system, and a consistent hashing algorithm is used to determine the distribution path to obtain the distribution confirmation result. Extract the node identifier from the distribution confirmation result and perform access restrictions to obtain the isolated execution result.

[0013] Secondly, the present invention provides a data security access system based on a distributed architecture, comprising: The log processing module is used to obtain raw data access logs, node identifiers, and data sensitivity tags, and to perform structured processing on the raw data access logs to obtain an extensible structured access dataset. The access pattern classification module is used to classify access behavior categories based on the structured access dataset and obtain access pattern classification results. The node frequency analysis module is used to calculate the access frequency of each node based on the access pattern classification results and the node identifier, and to quantify the distribution of access duration through entropy analysis to obtain the node frequency distribution characteristics. The node relationship analysis module is used to transform the node frequency distribution characteristics into cross-node object access sequences and construct a global transaction database. The global transaction database is scanned to eliminate pseudo-related object combinations and obtain the inter-node association patterns. The coordinated attack pattern analysis module is used to determine whether the support of frequent itemsets of the inter-node association pattern exceeds a preset support threshold, and then analyze the access behavior topology between node identifiers to obtain a coordinated attack pattern based on access request type. An abnormal behavior detection module is used to calculate an abnormal behavior score based on the cooperative attack mode and the data sensitivity label using a random forest algorithm. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. The real-time security policy generation module is used to adjust the access permissions of nodes based on the abnormal behavior detection results and preset dynamic adjustment rules, restrict access to nodes in the high-risk node list, and generate a real-time security policy. The security policy implementation module is used to determine whether the real-time security policy triggers the preset isolation trigger condition, and then distributes the high-risk node list to each node of the distributed system, restricts the access permissions of the relevant node identifiers, and obtains the isolation execution result.

[0014] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the data security access method based on a distributed structure as described above.

[0015] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the data security access method based on a distributed structure as described above.

[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention deploys a log collection module on each distributed node to obtain the millisecond-level timestamps and target object identification information of access requests in real time. After aggregating these data, a unified access behavior sequence across nodes is constructed, which realizes high-precision restoration and reconstruction of user behavior in the distributed system. Then, by combining the set access mode template and abnormal rules for matching analysis, it can accurately locate abnormal behavior of a single node and abnormal path across nodes, thereby realizing accurate monitoring and security protection of access behavior in the distributed system.

[0017] (2) This invention aggregates the access sequences of each node according to timestamps and constructs a collaborative access graph between nodes based on object access similarity and time proximity, thereby identifying suspicious linkage patterns between multiple nodes. Combined with the strategy generation module, the access control rules are dynamically adjusted to isolate and restrict high-risk paths, which can block potential collaborative attack chains in a timely manner, thereby effectively reducing the risk of collaborative attacks and improving system security. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of a data security access method based on a distributed structure provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of a data security access system based on a distributed structure provided in the second embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Reference Figure 1 The first embodiment of the present invention provides a data security access method based on a distributed structure, comprising the following steps: S11, Obtain the raw data access log and perform structured processing on the raw data access log to obtain an extensible structured access dataset, wherein the raw data access log includes node identifiers, access frequency and data sensitivity labels; S12, Based on the structured access dataset, classify access behavior categories to obtain access pattern classification results; S13. Based on the access pattern classification results and the node identifier, calculate the access frequency of each node, quantify the distribution of access duration through entropy analysis, and obtain the node frequency distribution characteristics. S14, the node frequency distribution characteristics are transformed into cross-node object access sequences and a global transaction database is constructed. The global transaction database is scanned to remove pseudo-related object combinations and obtain the inter-node association patterns. S15, calculate the frequent itemset support of the inter-node association pattern. If the frequent itemset support of the inter-node association pattern exceeds the preset support threshold, analyze the access behavior topology between node identifiers to obtain a collaborative attack pattern based on access request type. S16. Based on the cooperative attack mode and the data sensitivity label, the abnormal behavior score is calculated using the random forest algorithm. If the abnormal behavior score is higher than the preset abnormal threshold, a high-risk node list is generated and combined with the preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. S17. Based on the abnormal behavior detection results, adjust the access permissions of the nodes according to the preset dynamic adjustment rules, restrict the access of nodes in the high-risk node list, and generate a real-time security policy. S18, If the real-time security policy triggers the preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, the access permissions of the relevant node identifiers are restricted, and the isolation execution result is obtained.

[0021] In step S11, the raw data access log is obtained and the raw data access log is processed in a structured manner to obtain an extensible structured access dataset. The raw data access log includes node identifiers, access frequency and data sensitivity labels. In one implementation, the step of performing structured processing on the original data access log to obtain a scalable structured access dataset includes: Obtain the raw data access logs, parse the request packets to extract millisecond-level timestamp sequences and global identifiers, and use a distributed message queue to store them to obtain the synchronously collected log dataset; Based on the log dataset, node identifiers and request types are extracted, and hash identifiers are generated. Combined with timestamps, the access frequency per second is calculated to obtain a dataset labeled with request types and frequencies. For the dataset labeled with request type and frequency, regular expressions are used to match sensitive fields to perform data anonymization and verify data integrity, resulting in an anonymized dataset; The de-identified dataset is compressed to obtain the scalable structured access dataset.

[0022] It should be noted that when obtaining real-time logs from distributed nodes, system logs can be captured in real time through a log collection agent deployed on each node. The log collection agent continuously monitors the request log files on the nodes, extracting request packets containing millisecond-level timestamps and global identifiers. For example, a node generates 100 log entries per second, each containing a timestamp of 1625098765432 milliseconds and a global identifier in UUID format, "123e4567-e89b-12d3-a456-426614174000". These logs are stored using a distributed message queue, and topic partitioning is configured to support high throughput, ensuring that logs are collected synchronously in chronological order. The collected log dataset includes node identifiers, access frequency, and data sensitivity tags. When parsing synchronously collected log datasets, node identifiers and request types can be extracted and combined with a hash algorithm to generate unique identifiers. For example, concatenating "Node-001+GET+1625098765432" generates the hash value "5f4dcc3b5aa765d61d8327deb882cf99". Based on timestamps, the frequency of accesses per second can be counted. For example, if there are 80 "GET" requests and 20 "POST" requests in a certain second, a dataset tagged with request types and frequencies can be generated. This approach helps analyze system load and request distribution, improving performance monitoring efficiency.

[0023] In datasets tagged with request types and frequencies, regular expressions can be used to match sensitive fields for data masking. For example, the matching pattern "[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}" replaces IP addresses with "***.***.***.***". Then, the CRC32 verification algorithm is used to confirm data integrity, ensuring that the masked data has not been tampered with. The masked and verified datasets can be compressed using the Gzip compression algorithm to reduce storage space. For example, a 100MB log dataset can be compressed to approximately 20MB.

[0024] For example, in a distributed order processing system of an e-commerce platform, log collection agents deployed on each business node can capture approximately 120 request logs per second. Each log entry contains a node identifier, access frequency, and data sensitivity tag. These logs are stored in a Kafka message queue, partitioned by node, ensuring high throughput and time-order consistency. When parsing the synchronously collected log dataset, the system extracts the node identifier and request type, concatenates them with a timestamp, and generates a unique hash value using the MD5 algorithm. Second-level access frequency statistics based on timestamps show that 90 "POST" requests and 30 "GET" requests occurred within a certain second, generating a structured dataset tagged with request types and frequencies. During the anonymization phase, the system uses regular expressions to match the IPv4 format IP address "203.45.67.89" and replaces it with a mask format "***.***.***.***" to prevent sensitive information leakage. After anonymization, a CRC32 checksum algorithm is used to verify data integrity, ensuring that the data has not been tampered with during processing. The de-identified and verified dataset was compressed using Gzip, reducing its size from the original 150MB to approximately 30MB, and then written to the MySQL database.

[0025] In step S12, based on the structured access dataset, access behavior categories are divided to obtain access pattern classification results; In one implementation, the step of classifying access behavior categories based on the structured access dataset to obtain access pattern classification results includes: The user access trajectories in the structured access dataset are segmented, and the number of accesses within each time period is calculated to obtain the access time distribution dataset. The access time distribution dataset is modified by removing outliers and standardizing the data to obtain a standardized access dataset. If the standardized access dataset contains multidimensional features, then dimensionality reduction processing is performed to extract key time and frequency features to obtain the dataset after feature extraction. The dataset after feature extraction is grouped using the k-means clustering algorithm to obtain the access pattern classification results.

[0026] It should be noted that segmenting user access patterns into time series specifically involves dividing continuous access behavior into a preset one-hour time window and counting the number of accesses within each time period to form an access time distribution dataset. Taking a distributed e-commerce platform as an example, the statistics show that there were 500 accesses between 20:00 and 21:00, but only 50 between 2:00 and 3:00 AM. The segmented statistics reflect the peak and trough of access.

[0027] For the access time distribution dataset, outlier removal and data standardization were performed. Outlier data points with sudden increases in access counts exceeding the mean plus 5 standard deviations were removed. Then, the access counts were normalized to the 0-1 range for unified processing of multi-dimensional features. Specifically, in the 24-hour access data of an e-commerce platform, the mean access count per hour was 200, and the standard deviation was 60. Therefore, the outlier threshold was set to 200 + 5 × 60 = 500 times. If the access count reached 520 times in a certain time period, it was considered outlier and removed. After outlier removal, the access count ranged from 50 to 480. Therefore, the normalized result for the time period with 200 access counts was (200-50) / (480-50)≈0.41.

[0028] Multidimensional feature analysis was performed on the standardized access dataset. Specifically, the standardized access dataset includes access behavior features across multiple dimensions, including access time, access frequency, and page type. Principal Component Analysis (PCA) was used to reduce the dimensionality of the multidimensional features. The covariance matrix of the data was calculated, principal components were extracted, and their contribution rates were calculated. Principal components with a cumulative contribution rate of over 90% were selected as key time and frequency features, forming the dimensionality-reduced feature set.

[0029] The k-means clustering algorithm is used to group the dimensionality-reduced feature set. Through iterative optimization, the data points are divided into a preset k clusters, minimizing the distance between data points within each cluster in the feature space, thus achieving natural clustering of browsing behavior. For example, in an e-commerce platform application, k-means divides user browsing behavior into three categories: high-frequency access group concentrated during peak promotional periods, periodic access group characterized by access at fixed times each day, and low-frequency access group representing users who browse occasionally.

[0030] For example, in user access analysis of a distributed e-commerce platform, access time distribution data showed that the number of visits reached 520 during the 20:00-21:00 time period (exceeding the anomaly threshold of 500 visits), and this abnormal data was removed. After removing the outlier, the range of visits was reduced to 50-480, with 200 visits occurring between 15:00-16:00, and the normalized mapping value was approximately 0.41. The standardized access dataset contains multi-dimensional features such as access time, access frequency, and page type. After dimensionality reduction processing using principal component analysis, key time and frequency features with a cumulative contribution rate of 92% were extracted. Based on the dimensionality-reduced feature set, the k-means clustering algorithm was used to divide user access behavior into three categories: the first category is the high-frequency access group during the promotional peak period of 20:00-22:00; the second category is the periodic access group showing a fixed daily access time; and the third category is the low-frequency access group that browses occasionally. This clustering result helps the platform accurately identify different user behavior patterns, providing data support for subsequent personalized recommendations and system resource allocation.

[0031] In step S13, based on the access pattern classification results and the node identifier, the access frequency of each node is calculated, and the distribution of access duration is quantified by entropy analysis to obtain the node frequency distribution characteristics. In one implementation, the step of calculating the access frequency of each node based on the access pattern classification result and the node identifier, and quantifying the distribution of access duration through entropy analysis to obtain the node frequency distribution characteristics includes: Access records are extracted based on the access pattern classification results and the node identifiers, segmented according to a preset time interval, and the number of accesses in each time interval is calculated to obtain a node access frequency dataset. Based on the node access frequency dataset, the access duration entropy value is calculated to obtain the node entropy value dataset. If the entropy value in the node entropy value dataset exceeds the preset entropy value threshold, the entropy value is normalized to obtain a standardized entropy value dataset. Based on the standardized entropy dataset, the node frequency distribution features are divided using the k-means clustering algorithm to obtain the node frequency distribution features.

[0032] It should be noted that the access records of specific nodes can be accurately extracted from the access pattern classification results based on node identifiers. Specifically, taking user A of an e-commerce platform as an example, the timestamps of each click on a product and each page browsed during a day can be extracted, such as clicking on product A at 8:00 and browsing page B at 8:05, to provide basic data for subsequent analysis.

[0033] The extracted access records are segmented into time windows of 30 minutes to count the number of accesses. Taking user A as an example, he accessed the site 5 times between 8:00 and 8:30 and 3 times between 8:30 and 9:00, forming a node access frequency dataset. This dataset contains the access frequency distribution for 24 windows, providing structured input for entropy analysis.

[0034] Entropy analysis is used to quantify the distribution pattern of node access duration. For example, user A accesses the node 20 times during the promotional period (20:00-21:00), while accessing it at other times is close to zero; a lower entropy value indicates concentrated access. User B, on the other hand, has a higher entropy value if their access frequency is evenly distributed. A node entropy dataset is calculated, such as user A's entropy value of 2.5 and user B's entropy value of 3.8, demonstrating the difference in node access patterns. If the entropy value exceeds a preset threshold of 3.0, it is normalized to the 0-1 range for easier subsequent unified analysis. Taking users A and B as examples, the normalized entropy values ​​are 0.6 and 0.9 respectively, forming a standardized entropy dataset, providing a unified scale for cluster analysis.

[0035] The standardized entropy dataset was grouped using the k-means clustering algorithm. Combined with access frequency characteristics, nodes were divided into a high-frequency concentrated group, a uniformly distributed group, and a low-frequency group. The access frequency characteristics were obtained by aggregating and statistically analyzing the access logs of each node to calculate the average number of visits within a time window. Specifically, user A, whose visits were concentrated during promotional periods, was assigned to the high-frequency concentrated group, while user B, whose visits were more dispersed, was assigned to the uniformly distributed group, reflecting different access behavior patterns.

[0036] For example, in the user access behavior analysis of an e-commerce platform, user A's access records show that they accessed the platform 5 times between 8:00-8:30, 3 times between 8:30-9:00, reached a peak of 18 times between 20:00-20:30, and then dropped sharply to 2 times between 21:00-21:30. Based on 30-minute time windows, a node access frequency dataset containing 48 time periods was formed. Through entropy calculation, user A's access was concentrated during evening promotional periods, with an uneven distribution of access frequency, and an entropy value of 2.3, below the set threshold of 3.0, indicating a significant concentration of access behavior. In contrast, user B's access was more evenly distributed, with access occurring in multiple time periods throughout the day, and an entropy value of 3.7, indicating a more dispersed access behavior. For user B, whose access exceeded the threshold, the entropy value was normalized, resulting in a standardized entropy value of 0.92. User A's standardized entropy value was 0.58. A standardized entropy dataset was constructed to ensure the data was on the same scale for subsequent analysis. Using the k-means clustering algorithm combined with access frequency, user A was categorized into a high-frequency cluster group, indicating that their access was mainly concentrated during promotional periods; user B, on the other hand, was categorized into a uniformly distributed group, reflecting diverse and balanced access behavior. This classification reveals the differences in access patterns among user groups, providing a basis for the platform to accurately push marketing activities and optimize resource allocation.

[0037] In step S14, the node frequency distribution characteristics are transformed into cross-node object access sequences and a global transaction database is constructed. The global transaction database is scanned to remove pseudo-related object combinations and obtain the inter-node correlation patterns. In one implementation, the step of converting the node frequency distribution characteristics into cross-node object access sequences and constructing a global transaction database, scanning the global transaction database, eliminating pseudo-related object combinations, and obtaining the inter-node association patterns includes: Obtain node log data; Obtain node frequency and object access records from the node log data, divide the object access records into time windows, and generate object access sequences between nodes; Based on the object access sequence between nodes, the access sequence within each time window is converted into transaction records to obtain a global transaction database; The global transaction database is scanned using the Apriori algorithm to calculate the frequent itemsets of each object combination. It is then determined whether the lift of the frequent itemsets is greater than a preset lift threshold. If it is greater, the object combination is retained; if it is less, pseudo-related object combinations are removed, and the association pattern between nodes is output.

[0038] It's important to note that generating the object access sequence between nodes specifically involves dividing the access records into fixed-length time windows based on timestamps. A 24-hour day is divided into 24 one-hour windows, and within each window, the node's access sequence to the accessed objects is arranged chronologically. Node A accesses object X 10 times and object Y 5 times within the 8:00-9:00 time window, forming the access sequence {X, Y, X, …}. This method captures the node's access behavior patterns across different time periods.

[0039] Based on the object access sequences between nodes, the access sequences within each time window are transformed into transaction records, resulting in a global transaction database. Specifically, the access sequence of a node within a given time window is converted into a single transaction record, indicating which objects the node accessed during that time window. For example, the sequence {X, Y, X} accessed by node A between 8:00 and 9:00 is transformed into the transaction record {A: (X, Y)}, indicating that node A accessed objects X and Y during that time period. The collection of transaction records for all nodes within each time window constitutes the global transaction database, providing a structured data foundation for subsequent frequent itemset mining.

[0040] The Apriori algorithm is used to scan the global transaction database, calculating frequent itemsets for each object combination. A preset support threshold of 0.2 and a confidence threshold of 0.6 are used. If objects X and Y appear simultaneously in 50% of transactions, the support requirement is met, forming a frequent itemset {X, Y}. Further calculation shows a confidence score P(X|Y) of 0.8, indicating a high probability that X also appears when Y appears. Lift is then calculated using the formula: Lift = Confidence(X|Y) / Support(X), which measures the increase in frequency of X and Y appearing together compared to when they appear independently. If Support(X) = 0.53, then Lift = 0.8 / 0.53 ≈ 1.5, which is higher than the preset threshold of 1.2. Therefore, this frequent itemset is retained as a valid association pattern.

[0041] The validity of association patterns is ensured by eliminating spurious object combinations through a lift threshold. Specifically, although objects X and Z frequently co-occur, their lift is 0.9, which is below the threshold of 1.2, indicating that their co-occurrence may be accidental, and therefore they are eliminated. Ultimately, the output node association patterns, such as {A:(X, Y), B:(Y, Z)}, reflect the actual association relationships when nodes access objects, providing data support for resource optimization, anomaly detection, and other applications.

[0042] For example, in a distributed e-commerce platform, by parsing the access logs of server nodes, node A accesses product X 10 times and product Y 5 times within the 8:00-9:00 time window, generating an access sequence {X, Y, X, …}. This sequence is then converted into a transaction record {A: (X, Y)}, and the transaction records of all nodes in different time windows are aggregated to construct a global transaction database. When scanning this database using the Apriori algorithm, with a support threshold of 0.2 and a confidence threshold of 0.6, it is found that products X and Y appear simultaneously in 50% of the transaction records, satisfying the support requirement and forming a frequent itemset {X, Y}. Further calculation shows that the confidence P(X|Y) is 0.8, indicating that a user is likely to also access X when accessing Y. The lift is calculated to be 1.5, higher than the threshold of 1.2, therefore this frequent itemset is retained as a valid association pattern. Conversely, although products X and Z frequently co-occur, their lift is only 0.9, lower than the threshold, indicating that this co-occurrence may be accidental, and therefore it is discarded. The final output node association pattern, such as {A:(X, Y), B:(Y, Z)}, reveals the association behavior of nodes in product access, providing strong data support for the platform to optimize resource allocation and precise marketing.

[0043] In step S15, the frequent itemset support of the inter-node association pattern is calculated. If the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, the access behavior topology between node identifiers is analyzed to obtain a collaborative attack pattern based on access request type. In one implementation, the frequent itemset support of the inter-node association pattern is calculated. If the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, the access behavior topology between node identifiers is analyzed to obtain a cooperative attack pattern based on access request type, including: Calculate the support of frequent itemsets of the association patterns between nodes. If it exceeds a preset support threshold, use a graph neural network to analyze the topology and request type distribution between nodes in the global transaction database to obtain behavior patterns based on request types. By mining the association rules of the behavioral patterns, and filtering the frequently co-occurring rules according to the preset confidence threshold, a collaborative attack pattern based on access request type is obtained.

[0044] It should be noted that the support of frequent itemsets in the calculation of inter-node association patterns is used. A support exceeding a preset threshold of 0.3 indicates that the combination has a high co-occurrence probability in node access behavior, which is the basis for discovering effective association patterns. Based on these high-frequency itemsets, a graph neural network is used to model and analyze the topological structure of nodes and the distribution of access request types in the transaction database. The graph neural network (GNN) takes node feature matrices and adjacency matrices as inputs. Node feature vectors are constructed by concatenating one-hot encodings of request types with access frequency statistics. Node identifiers constitute the nodes of the graph, and access request types form edges between nodes. The edge weight is a normalized value representing the access frequency of that request type relative to the total number of node visits. The GNN uses a GCN structure, employing a three-layer aggregation network with mean aggregation as the aggregation function, and is trained under supervision using cross-entropy loss. During the inference phase, the trained GNN model generates embedding vectors for target nodes and identifies node behavior patterns based on vector similarity. Vector similarity is calculated using the cosine similarity formula, which involves performing a dot product operation on the embedding vectors of two nodes and dividing by the product of the magnitudes of the two vectors to obtain a similarity score in the range [-1, 1]. The closer the value is to 1, the more similar the node behavior patterns. Specifically, node identifiers constitute the nodes of the graph, access request types form edges between nodes, and edge weights represent request frequencies. Through multi-layer aggregation and transmission, GNN can capture the structural relationships and request behavior similarities between nodes, identify potential behavior patterns based on request types, and effectively reveal the complex relationships of node collaborative access in distributed systems.

[0045] The behavioral patterns obtained through GNN are further analyzed to uncover association rules. A pre-set confidence threshold of 0.7 is used to filter frequently co-occurring access request type rules. These rules reflect the probability of another type of request following a certain type of request initiated by a node, revealing the temporal and behavioral correlations between requests. High-confidence rules frequently appear across multiple nodes, and combined with the tight topological connections between nodes, they can characterize collaborative attack patterns based on access request types, such as collaborative behavior in distributed denial-of-service attacks.

[0046] For example, in a distributed network security monitoring system, the support of frequent itemsets in the association patterns between nodes is first calculated. The support of the frequent itemset {GET_X, POST_Y} in the global transaction database is 0.35, exceeding the preset threshold of 0.3, indicating that this request combination frequently occurs in node access behavior. The system analyzes the access behavior graph constructed from node identifiers, where nodes represent servers, edges represent interactions between nodes based on request types, and edge weights reflect request frequency. Through multi-layer information aggregation, the Generative Neural Network (GNN) captures a tight subgraph composed of nodes A, B, and C, which frequently initiate GET_X and POST_Y requests within multiple time windows, exhibiting highly similar access behavior patterns. Subsequently, association rule mining is performed on the behavior patterns identified by the GNN, with a confidence threshold set at 0.7. The mining results show that the confidence of the rule {GET_X → POST_Y} reaches 0.82, meaning that after a node initiates a GET_X request, there is an 82% probability that it will immediately follow with a POST_Y request. The rule appears multiple times among nodes A, B, and C. Combined with the strong connections formed by their topology, it can be inferred that these nodes are collaboratively launching attack behaviors based on access request types, such as distributed denial-of-service (DDoS) attacks. This analysis process not only confirms the existence of high-frequency request combinations but also achieves accurate detection and early warning of complex collaborative attack patterns through the comprehensive identification of behavioral patterns and rules.

[0047] In step S16, based on the cooperative attack mode and the data sensitivity label, an abnormal behavior score is calculated using the random forest algorithm. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. In one implementation, based on the cooperative attack pattern and the data sensitivity label, an abnormal behavior score is calculated using a random forest algorithm. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack pattern library to determine abnormal behavior, thereby obtaining an abnormal behavior detection result, including: Obtain network logs and extract access behavior data. Combine the data sensitivity tags with the data and use the random forest algorithm to calculate the abnormal threshold of each access behavior to obtain a behavior score. Behavioral pattern analysis is performed based on the behavioral score and the access frequency, and the existence of a collaborative attack is determined by combining the collaborative attack pattern to identify a candidate list of high-risk nodes. If the behavior score in the candidate list of high-risk nodes exceeds the preset abnormal score threshold, then the node risk assessment is performed in conjunction with the data sensitivity label to generate a list of high-risk nodes. By comparing the attack characteristics of the high-risk node list with a pre-established collaborative attack pattern library, it is determined whether there is any abnormal behavior, and an abnormal behavior detection result is generated.

[0048] It should be noted that the Random Forest algorithm calculates behavior scores based on access behavior characteristics, including request frequency, request type distribution, and resource sensitivity. By constructing multiple decision trees, multi-dimensional analysis is performed on each access behavior, outputting anomaly scores. Specifically, when node A's GET request frequency for high-sensitivity resources suddenly increases from 10 times to 100 times within a certain time window, the Random Forest will give a higher score of 0.85 to reflect the risk level of the anomalous behavior.

[0049] Behavioral pattern analysis is performed based on behavior scores and access frequency, combined with collaborative attack patterns to determine the existence of collaborative attacks. Specifically, a rule engine is used to set threshold rules for behavior scores and access frequencies. When a node's behavior score exceeds 0.7 and its access frequency exceeds 10 times per second, the rule engine automatically identifies the node as an abnormal node and triggers an alarm, thereby achieving real-time detection and response to collaborative attacks. If node A and node B both initiate high-frequency GET requests to highly sensitive resources within the same time period, and both scores exceed the preset threshold of 0.7, and the request interval is less than 1 second, it may be determined as a collaborative attack, and both are added to the high-risk node candidate list.

[0050] When generating a list of high-risk nodes, a risk assessment is conducted using data sensitivity tags and a weighted scoring method. High-sensitivity resources are weighted at 1.5, medium-sensitivity resources at 1.0, and low-sensitivity resources at 0.5. The final risk score for a node is calculated by multiplying the behavioral score by the sensitivity weight of the corresponding resource. When this risk score exceeds a preset threshold of 1.05, the node is deemed high-risk. High-sensitivity resources refer to access objects involving core system business, security, privacy, or critical data, including user personal information, financial data, identity authentication interfaces, and core business operation interfaces. These resources are assigned high data sensitivity tags within the system due to their information value and security importance.

[0051] By comparing the attack characteristics of high-risk nodes with a pre-established library of collaborative attack patterns, it is determined whether there are matching abnormal behavior patterns. Specifically, the behavioral characteristics of high-risk nodes are converted into multi-dimensional feature vectors containing request frequency (weight 0.4), access time interval (weight 0.25), resource sensitivity level (weight 0.25), and duration (weight 0.1). Based on the multi-dimensional feature vectors of typical collaborative attack behaviors pre-stored in the collaborative attack pattern library, the cosine similarity between the weighted feature vectors is calculated to quantify the similarity between node behavior and collaborative attack patterns. The cosine similarity is calculated by first multiplying each dimension of the two feature vectors by its weight and then weighting them to obtain a weighted feature vector. Then, the ratio of the dot product of these two weighted vectors to the product of their respective magnitudes is calculated. For example, if the pattern library records attack characteristics of multiple nodes making high-frequency, short-duration requests to high-sensitivity resources, and the behavioral similarity between nodes A and B is higher than 80%, an abnormal behavior detection result indicating a risk of collaborative attack is output.

[0052] For example, a distributed e-commerce platform extracts access behavior data from network logs. Node A initiated 100 GET requests to user personal information resources marked as highly sensitive within a time window from 10:00 to 10:05, far exceeding the normal average access frequency of 10 times. The calculated abnormal behavior score for Node A during this period is 0.85, higher than the preset threshold of 0.7. Simultaneously, Node B also initiated 95 GET requests to the same highly sensitive resource during the same time period, with an abnormal score of 0.82, and the interval between the requests from both nodes is less than 1 second. Combining the behavior score and access frequency, the system determines that nodes A and B may be engaging in a coordinated attack and includes them in a high-risk node candidate list. Subsequently, the system comprehensively assesses the node risk based on data sensitivity tags, confirming that both nodes initiated abnormal access to highly sensitive resources related to core business operations, ultimately generating a high-risk node list. Finally, the system compares the high-risk node list with a pre-established collaborative attack pattern library. Because the behavior of nodes A and B closely matches the high-frequency request characteristics of distributed denial-of-service attacks in the database, the system outputs abnormal behavior detection results, indicating the risk of coordinated attacks and providing support for security response.

[0053] In step S17, based on the abnormal behavior detection results, the access permissions of the nodes are adjusted according to the preset dynamic adjustment rules to restrict access to the nodes in the high-risk node list and generate a real-time security policy. In one implementation, the step of adjusting node access permissions based on the abnormal behavior detection results and a preset dynamic adjustment rule to restrict access to nodes in the high-risk node list and generate a real-time security policy includes: Based on the list of high-risk nodes, obtain the access permission data of the corresponding nodes, dynamically adjust the rules to calculate the permission restriction ratio, and obtain the updated access permission configuration; Based on the updated access permission configuration, the access control rules for high-risk nodes are adjusted to determine the real-time policy content for node access restrictions, thereby obtaining a real-time security policy.

[0054] It should be noted that, combining node anomaly scores and resource sensitivity, the system calculates the access restriction ratio according to preset dynamic adjustment rules. These preset dynamic adjustment rules refer to the system using a linear weighted model to weight and sum the node's anomaly score and resource sensitivity level. Specifically, the rule is: Access Restriction Ratio = Basic Restriction Ratio + Anomaly Score Weight × Anomaly Score + Resource Sensitivity Weight × Resource Sensitivity Level. The basic restriction ratio is the minimum limit for node access (0.3, i.e., at least 30% of access frequency is restricted), and the anomaly score weight and resource sensitivity weight are 0.4 and 0.3, respectively. Specifically, if node C is marked as high-risk and its anomaly score for the highly sensitive resource Z is 0.9 (exceeding the threshold of 0.75), the system will adjust node C's access frequency restriction to 50% of its original permissions, such as allowing a maximum of 100 access requests per hour, thus generating an updated access permission configuration.

[0055] Based on the updated access permission configuration, the system adjusts the access control rules for high-risk nodes in real time to form specific real-time security policies.

[0056] For example, in a distributed e-commerce platform, node C is detected to be behaving abnormally, with an anomaly score of 0.9, far exceeding the preset threshold of 0.75. Furthermore, the resource Z it accesses is a highly sensitive business interface involving user payment information. Based on preset dynamic adjustment rules, the system calculates that node C's access frequency should be limited to 50% of its original permissions, meaning a maximum of 100 access requests per hour. Subsequently, the system obtains node C's current access permission data for resource Z and generates an updated access permission configuration. Based on this configuration, the system automatically adjusts node C's access control rules, limiting its read request frequency to resource Z and prohibiting write operations to prevent abnormal traffic from impacting critical business operations.

[0057] In step S18, if the real-time security policy triggers the preset isolation trigger condition, the high-risk node list is distributed to each node of the distributed system, the access permissions of the relevant node identifiers are restricted, and the isolation execution result is obtained. In one implementation, if the real-time security policy triggers a preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, access permissions for the relevant node identifiers are restricted, and an isolation execution result is obtained, including: If the real-time security policy triggers the preset isolation trigger condition, then according to the high-risk node list, the risk value of each node is sorted using the Naive Bayes algorithm to obtain an ordered node list. The ordered list of nodes is distributed to each node in the distributed system, and a consistent hashing algorithm is used to determine the distribution path to obtain the distribution confirmation result. Extract the node identifier from the distribution confirmation result and perform access restrictions to obtain the isolated execution result.

[0058] It should be noted that if the real-time security policy meets the preset isolation trigger conditions, such as if a node receives more than 100 requests per unit time and has a failure rate higher than 80%, the system will initiate the subsequent isolation response process.

[0059] The system uses the Naive Bayes algorithm to sort the risk values ​​of each node based on the current list of high-risk nodes, generating an ordered list of nodes. Specifically, the system first calculates the risk value of each node based on its multidimensional features (including behavior rating, access frequency, and resource sensitivity). Then, these risk values ​​are used as input features to the Naive Bayes classifier to predict the risk level of each node. Finally, the nodes are sorted according to the prediction results to generate an ordered list of nodes.

[0060] To ensure the efficient and reliable distribution of the ordered node list to all nodes in the distributed system, the system first maps all nodes to a virtual hash ring using a hash function, and evenly distributes multiple virtual nodes on the ring for each node. Then, it calculates a hash value for each piece of data or request to be distributed, and locates the first node with a hash value greater than that value clockwise along the ring, thus distributing the request. In the event of a node failure, the system continues to search for the next node clockwise, ensuring fault tolerance in the distribution. After distribution is complete, the system receives and records the distribution confirmation results from each node, ensuring accurate delivery of the list and providing a basis for subsequent access control restrictions and isolation operations.

[0061] For example, in a distributed system, node D initiates 120 access requests within a unit of time, with 85% of these requests failing, meeting the isolation trigger condition preset by the real-time security policy. The system immediately initiates the isolation process. First, it extracts data for node D and other risky nodes from the high-risk node list, calculates the risk value of each node using the Naive Bayes algorithm, and determines that node D has the highest risk value, thus ranking it first and forming an ordered node list {D, E, F}. Subsequently, the system distributes this ordered node list to all nodes in the system using a consistent hashing algorithm, ensuring that the list is distributed evenly and stably to all relevant nodes. Each node returns a distribution confirmation upon receiving the list, and the system collects the confirmation information to ensure the distribution process is error-free. Finally, based on the distribution confirmation result, the system implements access restriction measures for node D, such as limiting the frequency of node D's access to sensitive resources, or even temporarily blocking its access. This restriction operation is successfully executed, and the system generates an isolation execution result, recording the isolation status of node D, providing a basis for subsequent security audits and risk management.

[0062] In summary, this invention generates a structured dataset by real-time collection of access logs from each node, extracting millisecond-level timestamps and object identifiers, aggregating and calculating access frequencies, and labeling them. It then uses k-means clustering to analyze access patterns, combines entropy analysis to quantify the complexity of node access distribution, transforms the data into cross-node access sequences, and uses the Apriori algorithm to mine frequently co-occurring objects among nodes to determine association patterns. If the support of frequent itemsets exceeds a threshold, a graph neural network is used to construct an access behavior topology, detect collaborative attack patterns, and combine this with a random forest algorithm to calculate abnormal behavior thresholds, generating a list of high-risk nodes. Finally, node access permissions are dynamically adjusted, and high-risk nodes are isolated through a broadcast mechanism to obtain the isolation execution result. This achieves timely identification of potential threats in distributed systems.

[0063] Reference Figure 2 The second embodiment of the present invention provides a data security access device based on a distributed structure, comprising: The log processing module is used to obtain raw data access logs, node identifiers, and data sensitivity tags, and to perform structured processing on the raw data access logs to obtain an extensible structured access dataset. The access pattern classification module is used to classify access behavior categories based on the structured access dataset and obtain access pattern classification results. The node frequency analysis module is used to calculate the access frequency of each node based on the access pattern classification results and the node identifier, and to quantify the distribution of access duration through entropy analysis to obtain the node frequency distribution characteristics. The node relationship analysis module is used to transform the node frequency distribution characteristics into cross-node object access sequences and construct a global transaction database. The global transaction database is scanned to eliminate pseudo-related object combinations and obtain the inter-node association patterns. The coordinated attack pattern analysis module is used to determine whether the support of frequent itemsets of the inter-node association pattern exceeds a preset support threshold, and then analyze the access behavior topology between node identifiers to obtain a coordinated attack pattern based on access request type. An abnormal behavior detection module is used to calculate an abnormal behavior score based on the cooperative attack mode and the data sensitivity label using a random forest algorithm. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. The real-time security policy generation module is used to adjust the access permissions of nodes based on the abnormal behavior detection results and preset dynamic adjustment rules, restrict access to nodes in the high-risk node list, and generate a real-time security policy. The security policy implementation module is used to determine whether the real-time security policy triggers the preset isolation trigger condition, and then distributes the high-risk node list to each node of the distributed system, restricts the access permissions of the relevant node identifiers, and obtains the isolation execution result.

[0064] It should be noted that the data security access device based on a distributed structure provided in this embodiment of the invention is used to execute all the process steps of the data security access method based on a distributed structure in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0065] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a log processing program. When the processor executes the computer program, it implements the steps described in the above embodiments of the data security access method based on a distributed architecture, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module in the above-described device embodiments, such as the log processing module.

[0066] For example, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0067] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0068] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.

[0069] The memory can be used to store the computer programs and modules. The processor implements various functions of the electronic device by running or executing the computer programs and modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0070] If the modules integrated into the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0071] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0072] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A data security access method based on a distributed architecture, characterized in that, include: Obtain raw data access logs and perform structured processing on the raw data access logs to obtain an extensible structured access dataset, wherein the raw data access logs include node identifiers, access frequency, and data sensitivity labels; Based on the structured access dataset, access behavior categories are divided to obtain access pattern classification results; Based on the access pattern classification results and the node identifier, the access frequency of each node is calculated, and the distribution of access duration is quantified by entropy analysis to obtain the node frequency distribution characteristics. The node frequency distribution characteristics are transformed into cross-node object access sequences and a global transaction database is constructed. The global transaction database is scanned to remove pseudo-related object combinations and obtain the inter-node association patterns. Calculate the frequent itemset support of the inter-node association pattern. If the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, analyze the access behavior topology between node identifiers to obtain a collaborative attack pattern based on access request type. Based on the cooperative attack mode and the data sensitivity label, an abnormal behavior score is calculated using the random forest algorithm. If the abnormal behavior score is higher than the preset abnormal threshold, a high-risk node list is generated and combined with the preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. Based on the abnormal behavior detection results, the access permissions of the nodes are adjusted according to the preset dynamic adjustment rules, restricting access to the nodes in the high-risk node list and generating a real-time security policy. If the real-time security policy triggers the preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, and the access permissions of the relevant node identifiers are restricted to obtain the isolation execution result.

2. The data security access method based on a distributed structure according to claim 1, characterized in that, The process of structuring the original data access logs to obtain an extensible structured access dataset includes: Obtain the raw data access logs, parse the request packets to extract millisecond-level timestamp sequences and global identifiers, and use a distributed message queue to store them to obtain the synchronously collected log dataset; Based on the log dataset, node identifiers and request types are extracted, and hash identifiers are generated. Combined with timestamps, the access frequency per second is calculated to obtain a dataset labeled with request types and frequencies. For the dataset labeled with request type and frequency, regular expressions are used to match sensitive fields to perform data anonymization and verify data integrity, resulting in an anonymized dataset; The de-identified dataset is compressed to obtain the scalable structured access dataset.

3. The data security access method based on a distributed structure according to claim 1, characterized in that, The step of classifying access behavior categories based on the structured access dataset to obtain access pattern classification results includes: The user access trajectories in the structured access dataset are segmented, and the number of accesses within each time period is calculated to obtain the access time distribution dataset. The access time distribution dataset is modified by removing outliers and standardizing the data to obtain a standardized access dataset. If the standardized access dataset contains multidimensional features, then dimensionality reduction processing is performed to extract key time and frequency features to obtain the dataset after feature extraction. The dataset after feature extraction is grouped using the k-means clustering algorithm to obtain the access pattern classification results.

4. The data security access method based on a distributed structure according to claim 1, characterized in that, The step involves calculating the access frequency of each node based on the access pattern classification results and the node identifier, quantifying the distribution of access duration through entropy analysis, and obtaining the node frequency distribution characteristics, including: Access records are extracted based on the access pattern classification results and the node identifiers, segmented according to a preset time interval, and the number of accesses in each time interval is calculated to obtain a node access frequency dataset. Based on the node access frequency dataset, the access duration entropy value is calculated to obtain the node entropy value dataset. If the entropy value in the node entropy value dataset exceeds the preset entropy value threshold, the entropy value is normalized to obtain a standardized entropy value dataset. Based on the standardized entropy dataset, the node frequency distribution features are divided using the k-means clustering algorithm to obtain the node frequency distribution features.

5. The data security access method based on a distributed structure according to claim 1, characterized in that, The process of converting the node frequency distribution characteristics into cross-node object access sequences and constructing a global transaction database, scanning the global transaction database, removing pseudo-related object combinations, and obtaining inter-node association patterns includes: Obtain node log data; Obtain node frequency and object access records from the node log data, divide the object access records into time windows, and generate object access sequences between nodes; Based on the object access sequence between nodes, the access sequence within each time window is converted into transaction records to obtain a global transaction database; The global transaction database is scanned using the Apriori algorithm to calculate the frequent itemsets of each object combination. It is then determined whether the lift of the frequent itemsets is greater than a preset lift threshold. If it is greater, the object combination is retained; if it is less, pseudo-related object combinations are removed, and the association pattern between nodes is output.

6. The data security access method based on a distributed structure according to claim 5, characterized in that, The calculation of the frequent itemset support of the inter-node association pattern, if the frequent itemset support of the inter-node association pattern exceeds a preset support threshold, then the access behavior topology between node identifiers is analyzed to obtain a cooperative attack pattern based on access request type, including: Calculate the support of frequent itemsets of the association patterns between nodes. If it exceeds a preset support threshold, use a graph neural network to analyze the topology and request type distribution between nodes in the global transaction database to obtain behavior patterns based on request types. By mining the association rules of the behavioral patterns, and filtering the frequently co-occurring rules according to the preset confidence threshold, a collaborative attack pattern based on access request type is obtained.

7. The data security access method based on a distributed structure according to claim 1, characterized in that, The abnormal behavior score is calculated using a random forest algorithm based on the cooperative attack pattern and the data sensitivity label. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack pattern library to determine abnormal behavior, resulting in an abnormal behavior detection result, including: Obtain network logs and extract access behavior data. Combine the data sensitivity tags with the data and use the random forest algorithm to calculate the abnormal threshold of each access behavior to obtain a behavior score. Behavioral pattern analysis is performed based on the behavioral score and the access frequency, and the existence of a collaborative attack is determined by combining the collaborative attack pattern to identify a candidate list of high-risk nodes. If the behavior score in the candidate list of high-risk nodes exceeds the preset abnormal score threshold, then the node risk assessment is performed in conjunction with the data sensitivity label to generate a list of high-risk nodes. By comparing the attack characteristics of the high-risk node list with a pre-established collaborative attack pattern library, it is determined whether there is any abnormal behavior, and an abnormal behavior detection result is generated.

8. The data security access method based on a distributed structure according to claim 1, characterized in that, The step of adjusting node access permissions based on the abnormal behavior detection results and preset dynamic adjustment rules, restricting access to nodes in the high-risk node list, and generating a real-time security policy includes: Based on the list of high-risk nodes, obtain the access permission data of the corresponding nodes, dynamically adjust the rules to calculate the permission restriction ratio, and obtain the updated access permission configuration; Based on the updated access permission configuration, the access control rules for high-risk nodes are adjusted to determine the real-time policy content for node access restrictions, thereby obtaining a real-time security policy.

9. The data security access method based on a distributed structure according to claim 1, characterized in that, If the real-time security policy triggers a preset isolation trigger condition, the high-risk node list is distributed to each node in the distributed system, access permissions for the relevant node identifiers are restricted, and an isolation execution result is obtained, including: If the real-time security policy triggers the preset isolation trigger condition, then according to the high-risk node list, the risk value of each node is sorted using the Naive Bayes algorithm to obtain an ordered node list. The ordered list of nodes is distributed to each node in the distributed system, and a consistent hashing algorithm is used to determine the distribution path to obtain the distribution confirmation result. Extract the node identifier from the distribution confirmation result and perform access restrictions to obtain the isolated execution result.

10. A data security access system based on a distributed architecture, characterized in that, include: The log processing module is used to obtain raw data access logs, node identifiers and data sensitivity tags, and to perform structured processing on the raw data access logs to obtain an extensible structured access dataset. The access pattern classification module is used to classify access behavior categories based on the structured access dataset and obtain access pattern classification results. The node frequency analysis module is used to calculate the access frequency of each node based on the access pattern classification results and the node identifier, and to quantify the distribution of access duration through entropy analysis to obtain the node frequency distribution characteristics. The node relationship analysis module is used to transform the node frequency distribution characteristics into cross-node object access sequences and construct a global transaction database. The global transaction database is scanned to eliminate pseudo-related object combinations and obtain the inter-node association patterns. The coordinated attack pattern analysis module is used to determine whether the support of frequent itemsets of the inter-node association pattern exceeds a preset support threshold, and then analyze the access behavior topology between node identifiers to obtain a coordinated attack pattern based on access request type. An abnormal behavior detection module is used to calculate an abnormal behavior score based on the cooperative attack mode and the data sensitivity label using a random forest algorithm. If the abnormal behavior score is higher than a preset abnormal threshold, a high-risk node list is generated and combined with a preset cooperative attack mode library to judge the abnormal behavior and obtain the abnormal behavior detection result. The real-time security policy generation module is used to adjust the access permissions of nodes based on the abnormal behavior detection results and preset dynamic adjustment rules, restrict access to nodes in the high-risk node list, and generate a real-time security policy. The security policy implementation module is used to determine whether the real-time security policy triggers the preset isolation trigger condition, and then distributes the high-risk node list to each node of the distributed system, restricts the access permissions of the relevant node identifiers, and obtains the isolation execution result.