Smart city data processing method and platform based on artificial intelligence, and storage medium

By constructing a dynamic data lineage graph and generating minimal access control rules, the problem of delayed response in smart city data processing systems when facing threats and sudden changes in business flow has been solved, achieving precision in data security management and resource optimization, and improving the intelligence and efficiency of the system.

CN121456902APending Publication Date: 2026-02-03CHANGCHUN KESHENGDA INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511600976.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing smart city data processing systems are unable to perceive the instantaneous behavioral patterns and content-derived relationships of data interactions in real time when faced with zero-day attacks, internal threats, or sudden changes in business flow. This results in lagging data security management, inaccurate access control, and serious waste of resources.

Method used

By constructing a dynamic data lineage graph, analyzing the instantaneous behavior patterns and content derivation relationships of real-time data streams, generating minimal access control rules and lifecycle management models, and implementing dynamic access control strategies.

Benefits of technology

It improves the accuracy and agility of data security management, solves the problem of delayed response of traditional strategies, realizes the precise allocation of permissions and dynamic optimization of resources, and enhances the intelligence and efficiency of urban data security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456902A_ABST
    Figure CN121456902A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data security management, and discloses a smart city data processing method and platform based on artificial intelligence, and a storage medium. Comprising the following steps: monitoring real-time data between network entities in real time to obtain a real-time data stream; constructing a dynamic data blood relationship map by analyzing an instantaneous behavior mode and a content derivation relationship of the real-time data stream; constructing a node symbiont by analyzing the data difference degree between different nodes in the dynamic data blood relationship map; calculating an interaction risk of the node symbiont and an external entity, and constructing a minimum access control rule according to the interaction risk; and constructing a life cycle management model, dynamically managing an access control strategy, and managing and controlling real-time data. According to the method, the optimization effect from passive response to active perception is realized on the data stream security management level, the collaborative effect of the authority structure and system resource allocation is improved, and the intelligence and efficiency of the overall data security management are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data security management, more particularly, the present application relates to a smart city data processing method and platform based on artificial intelligence and a storage medium. BACKGROUND

[0002] With the deepening of smart city construction, core scenarios such as traffic control, environmental monitoring, government services, and public security have formed a massive, multi-source, and heterogeneous dynamic data system. These data are the core support for accurate decision-making and efficient operation of smart cities.

[0003] Current smart city data processing has defects in control, making it difficult to protect data security and release data value. For example, existing data flow security management mostly uses static strategies based on IP addresses, port numbers, and protocol types, which can only judge the surface properties of data transmission and cannot deeply understand the context and business intent of data flow. When facing zero-day attacks, covert operations by internal personnel, or traffic surges due to peak demand in normal business, static strategies cannot real-time perceive the instantaneous behavior patterns, content derivative relationships, and dynamic trust states between entities, leading to data processing lag when responding to urban emergency needs, and failing to provide immediate support for decision-making. Furthermore, the current access control strategy lacks detailed evaluation of node interaction risks and does not consider the time decay characteristics of data value. It does not generate "minimum necessary" permission rules based on business scenarios, resulting in permission accumulation and increasing the risk of unauthorized data access. When data value enters the decay period with the life cycle, it may still occupy high-level security resources, causing serious waste of server computing power, storage space, and other resources, and making it difficult to meet the demand for precision, agility, and intelligence in data security management. SUMMARY

[0004] To overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purpose, the present application provides the following technical scheme: a smart city data processing method based on artificial intelligence, comprising:

[0005] Real-time monitoring of real-time data between network entities to obtain real-time data flow;

[0006] Constructing a dynamic data bloodline map by analyzing the instantaneous behavior patterns and content derivative relationships of the real-time data flow;

[0007] Constructing a node symbiont by analyzing the data difference degree between different nodes in the dynamic data bloodline map;

[0008] Calculating the interaction risk between the node symbiont and external entities, and constructing a minimum access control rule based on the interaction risk;

[0009] The time value of the real-time data stream is extracted, and a life cycle management model is constructed according to the decay rate of the time value and a dynamic data bloodline graph, so as to dynamically manage an access control policy;

[0010] A data management strategy that can be loaded in real time is compiled according to the minimum access control rule and the life cycle management model, and the real-time data is managed.

[0011] Preferably, the dynamic data bloodline graph is constructed, including:

[0012] The real-time data stream is parsed into discrete interaction events, and the source and sink endpoints and content fingerprints of each interaction event are marked;

[0013] The semantic correlation strength between continuous interaction events is calculated, and an interaction event pair is constructed according to two interaction events with a semantic correlation strength higher than a preset strength threshold; the semantic correlation strength is calculated based on a hash segment overlap rate and a timing combination probability of an instruction segment of the interaction event;

[0014] The propagation direction of the associated events is determined according to the chronological order of the occurrence of the interaction event pair;

[0015] A derivative link of the interaction event is constructed with the endpoint as a node, the propagation direction of the interaction event pair as a directed edge, and the semantic correlation strength as an edge weight;

[0016] The mutation and behavior deviation of the endpoint behavior in the interaction event are identified, and the trust correlation degree between the endpoints is quantified; the trust correlation degree is obtained based on similarity measurement of the interaction flow data between the source and sink endpoints and the outbound interaction flow data of the source endpoint;

[0017] The derivative link and the trust correlation degree between different endpoints are fused to obtain the dynamic data bloodline graph.

[0018] Preferably, the node symbiotic body is constructed, including:

[0019] The access frequency and path depth of each node in the dynamic data bloodline graph are calculated, and the intrinsic sensitivity of the node is obtained by weighted fusion;

[0020] The betweenness centrality and closeness centrality of the node in the dynamic data bloodline graph are calculated, and the influence of the node is obtained by weighted fusion;

[0021] The propagation speed and range of the influence between different nodes in the node symbiotic body are analyzed, and the key hub node and the ordinary leaf node are identified;

[0022] The comprehensive sensitivity vector of the node is generated according to the intrinsic sensitivity and influence of the key hub node;

[0023] According to the comprehensive sensitivity vector, the nodes are clustered to obtain a plurality of stable node clusters, and a node co-occurrence body is constructed.

[0024] Preferably, the constructing minimizes the access control rules, including:

[0025] An abnormal interaction event affecting the security of the node co-occurrence body is identified, and the sink endpoint, operation type and intrinsic sensitivity of the abnormal interaction event are extracted and combined to obtain abnormal interaction features;

[0026] According to the abnormal interaction features, the real-time data stream is risk assessed to obtain an abnormal risk;

[0027] According to the abnormal risk, an access constraint condition is constructed, which is used as a minimum access control rule between nodes.

[0028] Preferably, the key hub nodes and ordinary leaf nodes are identified, including:

[0029] The out-edge connection weight of all nodes in the dynamic data bloodline graph is extracted, and the initial transition probability of the node is calculated;

[0030] The data flow rate and the number of adjacent nodes of the node are extracted, and the restart preference probability is calculated in combination with the intermediate centrality;

[0031] According to the initial transition probability, the restart preference probability and the intrinsic sensitivity of the node, a virtual walking particle group is initialized, and a plurality of rounds of parallel random walk simulation are performed;

[0032] According to the random walk simulation, the expected residence time and the flow convergence intensity of each node are extracted and calculated, and the flow convergence node is identified;

[0033] The damage degree of the node failure to the connectivity of the dynamic data bloodline graph is analyzed to determine the key hub node.

[0034] Preferably, the abnormal interaction event affecting the security of the node co-occurrence body is identified, including:

[0035] Based on the dynamic data bloodline graph, the interaction events between the key hub nodes are extracted in real time, and the interaction events not in the same node co-occurrence body are identified as cross-node set events;

[0036] The path depth, the number of transit nodes and the boundary jump times of the cross-node set events are extracted and weightedly fused, and the weighted fusion result is used as the confidence of the cross-node set event;

[0037] The confidence decay sequence of the cross-node event is calculated, and the sequence segment whose confidence decay rate exceeds a preset decay threshold is identified, and the cross-node set event corresponding to the sequence segment is marked as an abnormal interaction event.

[0038] Preferably, the determining the key hub node comprises:

[0039] Mark each flow convergence node in the atlas as a failure node in sequence, temporarily remove the failure node from the dynamic data bloodline atlas, and disconnect all edges associated with the failure node;

[0040] After removing each failure node, the following operations are performed:

[0041] Identify the remaining connected subgraphs in the dynamic data bloodline atlas, calculate and compare the size proportion of each connected subgraph, and take the largest size proportion as the connectivity retention rate;

[0042] Calculate the distribution uniformity of the size proportion of all connected subgraphs in the dynamic data bloodline atlas;

[0043] Identify and count the number of isolated node groups as an isolation indicator;

[0044] Weighted calculation is performed on the connectivity retention rate, the distribution uniformity and the reciprocal of the isolation indicator, and the result of the weighted calculation is taken as the connectivity destruction contribution degree of the flow convergence node;

[0045] According to the connectivity destruction contribution degree, the flow convergence nodes are sorted in descending order, and the flow convergence nodes at the top of the sorting are taken as the key hub nodes.

[0046] Preferably, the constructing the life cycle management model comprises:

[0047] Obtaining the heat of the real-time data stream and the derived new data flow, fitting the time value curve, and extracting the value change rate;

[0048] Adding a classification label to the real-time data stream by analyzing the long-term change trend of the time value curve;

[0049] Extracting the interval length of the short-term change feature of the time value curve, and recording it as the feature hiding length under the classification label;

[0050] Calculating the feature expected length according to the feature hiding length under different classification labels, and constructing the life cycle management model.

[0051] The application also provides a smart city data processing platform based on artificial intelligence, which is applied to the smart city data processing method based on artificial intelligence and comprises:

[0052] A data acquisition module for monitoring real-time data between network entities in real time and obtaining real-time data streams;

[0053] A data management and analysis module for constructing a dynamic data bloodline atlas by analyzing the instantaneous behavior mode and content derivative relationship of the real-time data stream;

[0054] a data aggregation analysis module, which constructs a node symbiotic body by analyzing the data difference degree between different nodes in the dynamic data bloodline graph;

[0055] a rule construction module, which calculates the interaction risk of the node symbiotic body and external entities, and constructs a minimum access control rule according to the interaction risk;

[0056] a policy management module, which extracts the time value of real-time data flow, constructs a life cycle management model according to the decay rate of the time value and the dynamic data bloodline graph, and dynamically manages the access control policy;

[0057] a compiling control module, which compiles a real-time loadable data management strategy according to the minimum access control rule and the life cycle management model, and manages the real-time data.

[0058] The application also provides a smart city data storage medium based on artificial intelligence, which is applied to the smart city data processing platform based on artificial intelligence and characterized by storing a computer program.

[0059] The smart city data processing method based on artificial intelligence has the following technical effects and advantages:

[0060] (1) The real-time data flow is parsed into discrete interaction events, and the semantic correlation strength between events and the trust correlation degree between endpoints are calculated in depth, so that a dynamic data bloodline graph reflecting the data instantaneous behavior mode, content derivation relationship and dynamic trust state between entities is constructed, so that the system can understand the context and intention of the data flow; the problem that the traditional static strategy is lagging behind or even invalid when facing zero-day attacks, internal threats or normal business flow mutations is solved, and the optimization effect from passive response to active perception and real-time self-adaptation is achieved, so that the precision and agility of city data security protection are improved, and support is provided for the decision-making of the city in the face of sudden needs.

[0061] (2) By analyzing the interaction risk of the node symbiotic body and external entities, a "minimum" rule granting only necessary permissions is generated, and the precision of permission allocation is realized; at the same time, by fitting the data value decay curve and adding classification labels to the data, a life cycle management model is constructed, so that the access control policy can be updated adaptively with the data value, and adaptive transformation and update of the management and control strategy are realized; the problem of permission accumulation and low-value data occupying high-security resources in the prior art is solved, the synergistic effect of dynamic optimization of permission structure and system resource allocation is improved, and the intelligence and efficiency of city data security management are improved. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 A method flowchart of the intelligent city data processing method based on artificial intelligence.

[0063] Figure 2 A method flowchart of constructing a dynamic data bloodline graph in the intelligent city data processing method based on artificial intelligence.

[0064] Figure 3 A module flowchart of the intelligent city data processing platform based on artificial intelligence. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0066] The present application provides an intelligent city data processing method, platform and storage medium based on artificial intelligence. By analyzing real-time data streams into discrete interactive events, constructing a dynamic data bloodline graph and constructing minimum access control rules, the optimization effect from passive response to active perception is achieved, the collaborative effect of permission structure and system resource allocation is improved, and the intelligence and efficiency of overall data security management are improved.

[0067] Embodiment one, please refer to Figure 1 and Figure 2 In the embodiments of the present application, the intelligent city data processing method based on artificial intelligence is realized by the following steps:

[0068] The data acquisition module monitors the real-time data between network entities in real time to obtain real-time data streams. The network entity refers to any logical unit or physical node in the network that can generate, send, receive or process data. The data stream refers to a sequence of data packets with context association transmitted between two or more network entities to complete a specific business logic.

[0069] The data aggregation analysis module constructs a dynamic data bloodline graph by analyzing the instantaneous behavior patterns and content derivative relationships of real-time data streams.

[0070] The dynamic data bloodline graph includes:

[0071] The real-time data stream is parsed into discrete interaction events, and the source and sink endpoints and content fingerprints of each interaction event are marked; wherein, the interaction event is abstracted from the data stream and is a semantic data exchange action, which is the smallest unit for building data blood relationship; the source and sink endpoints refer to the initiator and receiver of an interaction event, which are respectively marked as source endpoint and sink endpoint; the content fingerprint refers to a digitalized abstract generated by a specific algorithm (such as local sensitivity hashing) on the core content of the data stream load, which can uniquely identify the content and has certain anti-disturbance ability;

[0072] The semantic association strength between continuous interaction events is calculated; wherein, the semantic association degrees between the interaction events of the last complete business cycle are statistically analyzed in terms of probability density, and the maximum semantic association degree is selected as the preset strength threshold in the top three semantic association degree intervals of the probability density; it should be noted that the complete business cycle refers to the time span required for one complete cycle that can represent and cover the core business activities from the beginning to the end, and if an arbitrary and fixed time window (such as the past 1 hour) is used, it may lead to data distortion; the interaction event pair is constructed according to the two interaction events with the semantic association strength higher than the preset strength threshold; the semantic association strength is calculated based on the hash fragment overlap rate and the timing combination probability of the instruction fragments; specifically including:

[0073] The content fingerprints of the continuous interaction events are sliced into fixed length, and the fixed length slices are mapped by local sensitivity hashing to generate their corresponding hash fragment sequences; wherein, the length of the slice is obtained based on the historical interaction event content load big data training; local sensitivity hashing is a technique for fast approximate nearest neighbor search, which will not be described in detail;

[0074] The two hash fragment sequences are compared to calculate the fragment sequence overlap rate; wherein, the fragment overlap rate refers to the proportion of the number of identical fingerprints in the two hash fragment sequences to the total number of fingerprints;

[0075] The operation instructions in the interaction events are combined in time sequence to obtain an operation instruction sequence;

[0076] Adjacent operation instructions in the operation instruction sequence are identified, and an operation instruction pair is constructed according to the two adjacent operation instructions;

[0077] The occurrence frequency of the operation instruction pair and the timing combination probability of the preceding operation instruction to the subsequent operation instruction in each operation instruction pair are extracted; the timing combination probability is calculated, specifically including: in the interaction event, the total frequency of the operation instruction pair with a specific operation instruction as the starting point is counted, and in the total frequency, the specific frequency of the target subsequent operation instruction immediately following the preceding operation instruction is separated out; the occurrence frequency of the target subsequent operation instruction is divided by the total frequency with the preceding operation instruction as the starting point to obtain the timing combination probability;

[0078] calculate the instruction continuity of the operation instruction sequence according to the frequency of occurrence and the time sequence combination probability of each operation instruction pair; wherein the instruction continuity can take the average of the frequency of occurrence and the time sequence combination probability;

[0079] weighting calculation of the segment coincidence rate and the instruction continuity, and taking the result of the weighting calculation as the semantic association strength between the interaction events; wherein the segment coincidence rate and the instruction continuity in the last complete business cycle are weighted calculated by using different weight combinations, the semantic association strength is calculated by using the standard deviation, and the weight corresponding to the minimum standard deviation is taken as the weighting weight of the segment coincidence rate and the instruction continuity;

[0080] determine the propagation direction of the associated events according to the time sequence of the occurrence of the interaction event pairs;

[0081] construct the derived link of the interaction events by taking the end points as nodes, the propagation direction of the interaction event pairs as directed edges, and the semantic association strength as edge weights;

[0082] identify the mutation and behavior deviation degree of the end point behavior in the interaction events, and quantify the trust association degree between the end points; the trust association degree is obtained based on the similarity measurement of the interaction flow data between the source and sink end points and the outbound interaction flow data of the source end point; specifically including:

[0083] The historical interaction frequency, data throughput, operation type distribution, and delay distribution of the interaction response between the acquisition endpoints are obtained as the statistical quantities of the behavior characteristics. The historical interaction frequency refers to the number of interaction events from one source endpoint to one or more destination endpoints in a unit of time. A probe can be deployed on a key node (such as a gateway or proxy server) through which all interaction events (such as network connections, API calls, and database query requests) pass to capture all the interaction events. The data throughput is the data volume transmitted in the interaction events, which is usually divided into uplink throughput (the amount of data sent by the source endpoint) and downlink throughput (the amount of data received by the source endpoint), and measures the scale of the interaction. The network probe analyzes the payload length of each data packet on the basis of capturing the interaction events. The operation type distribution refers to the proportion of various operation commands initiated by one source endpoint to the destination endpoint in a unit of time. The captured real-time data stream is subjected to deep packet analysis to identify the application layer protocol (such as HTTP, SQL, and Redis protocol). The key operation instructions (such as HTTP: extraction of request methods (GET, POST, etc.), SQL: extraction of SQL command types (SELECT, INSERT, etc.), and custom API: extraction of API endpoints or function names) are extracted from the analyzed protocol. The number of occurrences of each operation type is counted in a time window to obtain the operation type distribution data. The interaction response delay refers to the time taken from the request sent by one source endpoint to the response received from the destination endpoint. The delay distribution refers to the statistical distribution (such as the average, median, and quantile) of all the delay values in a period of time.

[0084] The statistical quantities of different types in each interaction event subsequence are combined to obtain the behavior characteristic vector of the source endpoint to each destination endpoint. The behavior characteristic vector is exemplarily illustrated as follows. The existing frequency [12.5, 2.1, 1.8] has data values representing the mean 12.5 qps, the standard deviation 2.1, and the peak-to-mean ratio 1.8. The throughput [0.8, 4.2, 0.19] has data values representing the uplink 0.8 MB / s, the downlink 4.2 MB / s, and the uplink-to-downlink ratio 0.19. The operation type distribution [0.88, 0.09, 0.03, 0.45] has data values representing [SELECT proportion 88%, READ proportion 9%, OTHER proportion 3%, and entropy value 0.45]. The delay distribution [45, 120, 15] has data values representing the P50 delay 45 ms, the P95 delay 120 ms, and the delay standard deviation 15 ms. The behavior characteristic vector is [12.5, 2.1, 1.8, 0.8, 4.2, 0.19, 0.88, 0.09, 0.03, 0.45, 45, 120, 15].

[0085] The outbound interaction flow of the source endpoint is acquired in real time, a real-time behavior feature vector is constructed, difference measurement is performed on the real-time behavior feature vector and the behavior feature vector, and a behavior deviation degree of the source endpoint is obtained; wherein the difference measurement adopts vector space distance calculation of the real-time behavior feature vector and the behavior feature vector;

[0086] The deviation degrees are time-series combined to obtain a deviation degree sequence, and mutation features in the deviation degree sequence are identified;

[0087] In the embodiment of the application, the mutation features in the deviation degree sequence are identified, and specifically include:

[0088] A preset window length is set, a deviation degree sub-sequence is obtained by using a sliding window, morphological clustering is performed on the deviation degree sub-sequence, deviation degree sub-sequences with similar waveforms (such as peak position, fluctuation amplitude and trend direction) are classified into the same category, and a period baseline is obtained by sequentially performing weighted average calculation on each deviation degree value of all deviation degree sub-sequences in the same category; wherein the weight is determined based on the appearance frequency and the clarity of the deviation degree sub-sequence; for example, three deviation degree sub-sequence categories are obtained according to morphological clustering, which are categories A, B and C, respectively including 12, 5 and 3 deviation degree sub-sequences, and the appearance frequencies are 0.6, 0.25 and 0.15, respectively; the clarity is obtained by measuring the waveform sharpness and intra-class consistency; it should be noted that the morphological clustering can adopt a K-Shape clustering algorithm;

[0089] For example, assume that the values of the three representative deviation sub-sequences in category A are {0.1, 0.15, 0.25, 0.2, 0.18, 0.3, 0.35, 0.25, 0.15}, {0.12, 0.18, 0.28, 0.23, 0.2, 0.32, 0.38, 0.27, 0.17} and {0.08, 0.13, 0.22, 0.18, 0.16, 0.28, 0.33, 0.23, 0.13}; where the sharpness calculates the significance of the peak-valley difference of each segment, for deviation sub-sequence A1, sharpness = (0.35-0.1) / average value ≈ 0.83; for deviation sub-sequence A2, sharpness = (0.38-0.12) / average value ≈ 0.85; for deviation sub-sequence A3, sharpness = (0.33-0.08) / average value ≈ 0.81; the intra-class consistency calculates the correlation coefficient of each segment and the mean value of the category, for deviation sub-sequence A1, the correlation coefficient with the mean value of category A = 0.98; for deviation sub-sequence A2, the correlation coefficient with the mean value of category A = 0.99; for deviation sub-sequence A3, the correlation coefficient with the mean value of category A = 0.97; for each deviation sub-sequence, its clarity is calculated by summing the correlation coefficients of sharpness; the clarity of A1, A2 and A3 is normalized to 0.32, 0.38 and 0.30 respectively; then the weights of deviation sub-sequences A1, A2 and A3 are obtained by multiplying the frequency of occurrence by the clarity and normalizing the multiplication result;

[0090] Each deviation sub-sequence is taken in turn as a temporary sliding window, starting from the starting position on the periodic baseline, and sliding point by point backward at a time unit (such as 1 second) as a step, at each sliding position, the local cross-correlation value of the deviation sub-sequence and the current overlapping part of the baseline is calculated, after traversing all possible sliding positions, a position is found at which the local cross-correlation value calculated at the position reaches a global maximum value, the global maximum value is recorded as the best phase alignment position; the dynamic time warping distance between the deviation sub-sequence and the aligned periodic baseline is calculated, and the distance is normalized to obtain the waveform fit score at the time point;

[0091] The waveform fit score is combined in time sequence to obtain a waveform fit score sequence, and a first-order difference calculation is performed on the waveform fit score sequence to obtain a fit rate sequence; the score change rate whose absolute value exceeds a preset score change rate threshold in the change rate sequence is identified as a mutation feature; wherein the probability density of the waveform fit score change rate data is statistically analyzed through the historical deviation sub-sequence of the last complete business period, and the maximum change rate is selected as the preset score change rate threshold in the waveform fit score change rate interval where the probability density is in the top three;

[0092] Fusion of deviation degree and mutation characteristics, get the trust correlation degree of the source endpoint to the destination endpoint; wherein, through the deviation degree and mutation characteristic data of the source endpoint in the last complete business cycle, different weights are used for weighted calculation, the standard deviation of the obtained trust correlation degree is calculated, and the minimum value of the standard deviation corresponds to a set of weights as the weighted weight of the deviation degree and the mutation characteristic;

[0093] Fusion of derived link and trust correlation degree between different endpoints, get dynamic data blood atlas; Specifically, the source and destination endpoints of each interaction event in the interaction event pair are obtained, the semantic correlation strength of the interaction event pair is taken as the correlation strength coefficient of each source and destination endpoint combination in the two source and destination endpoint combinations; For different interaction event pairs, multiple correlation strength coefficients of the same source and destination endpoint combination are extracted and mean value calculation is performed, and the calculation result is taken as the expected correlation coefficient, combined with the trust correlation degree between the corresponding source and destination endpoints, the behavior vector between the corresponding endpoints is constructed, the Euclidean distance between the behavior vectors is taken as the similarity measure, the behavior vectors are clustered based on density, and multiple stable behavior vector clusters are obtained; wherein, in the behavior vector, the Euclidean distance between each two behavior vectors is less than a preset Euclidean distance threshold; wherein, by obtaining the historical behavior vectors between different source and destination endpoints in the last complete business cycle, the probability density of the historical Euclidean distance between the historical behavior vectors is calculated, and in the Euclidean distance interval where the probability density is in the top three inverses, the minimum Euclidean distance is selected as the preset Euclidean distance threshold;

[0094] According to the endpoints corresponding to each behavior vector in the behavior vector cluster, a general endpoint set is constructed; the average Euclidean distance in each behavior vector cluster is calculated, and if the average Euclidean distance is less than the behavior vector cluster of the preset stability threshold, the corresponding general endpoint set is re-marked as a stable endpoint set; wherein, according to the behavior vector of the last complete business cycle, the corresponding historical vector cluster is constructed, the average Euclidean distance in different behavior vector clusters is statistically calculated, and in the average Euclidean distance interval where the probability density is in the top three inverses, the minimum value of the average Euclidean distance is taken as the preset stability threshold;

[0095] Determine whether the source and destination endpoints of the data flow belong to two different stable endpoint sets, if they belong to two different stable endpoint sets, a strong connection edge is established between the two stable endpoint sets; if one belongs to a stable endpoint set and the other belongs to a general endpoint set, a medium-strength connection edge is established; all connection relationships are taken as edges, and endpoint sets are taken as nodes to construct a dynamic data blood atlas;

[0096] In view of the defect that the static security policy in the existing data flow security management strategy cannot adapt to the dynamic changes of real-time data flow, by parsing the real-time data flow into discrete interactive events and deeply calculating the semantic correlation strength between events and the trust correlation degree between endpoints, a dynamic data bloodline graph reflecting the data instantaneous behavior mode, content derivative relationship and dynamic trust state between entities is constructed, so that the system can understand the context and intention of the data flow; the problem that the traditional static strategy lags behind or even fails to respond when facing zero-day attacks, internal threats or normal business flow mutations is solved, the optimization effect from passive response to active perception and real-time self-adaptation is achieved, and the precision and agility of security protection are improved;

[0097] The data aggregation analysis module constructs a node symbiotic body by analyzing the data difference degree between different nodes in the dynamic data bloodline graph;

[0098] The construction of the node symbiotic body includes:

[0099] The access frequency and path depth of each node in the dynamic data bloodline graph are calculated, and the intrinsic sensitivity of the node is obtained by weighted fusion; wherein the access frequency refers to the frequency of data nodes being requested, read or referenced by other entities in the network within a unit time (such as 1 second), which is calculated by ratio of the number of basic operations (such as API call, database query) in the system background log to the unit time length; the path depth is calculated by traversing the data derivative link in reverse from the current node, recording the longest hop number required to reach all data source nodes (i.e. no more upstream nodes), simulating a "data change" or "data pollution" starting from the current node, propagating along the data flow direction, calculating the number of nodes it passes through and the number of hops required to reach the stable state; the number of nodes passed through and the number of hops required to reach the stable state are summed to obtain the path depth; the weight of weighted fusion is based on the statistical correlation strength of the access frequency and the path depth with the security events that have occurred, and the risk preference configuration under the current business environment; the significance of intrinsic sensitivity is to change the cornerstone of security management from the static identity of data to the dynamic value and risk performance in the dynamic data bloodline graph;

[0100] The number of hops required to reach the stable state is exemplarily illustrated, for example, starting from a certain node, marking it as the 0th hop, initializing a propagation entropy value H0=1, representing complete influence; for each node i of the ith hop, the out-edge weight distribution p i is calculated, and the entropy decay factor ss i of the node i is calculated by the formula ss i =1-entropy(p i ); wherein g x is the out-edge weight of the xth out-edge connected node in the node i;

[0101] entropy(p i ) is a Shannon entropy calculation function for measuring the core index of uncertainty of a random variable; the influence speed receiving value of node i is calculated by the formula G i = G i-1 × ss i × (1-δ); where δ is the system basic attenuation rate; G i is the influence speed receiving value; the weighted average influence speed V i of the i-th jump is calculated by the formula V i =∑(η i × G i ) / i ; in the formula, η i is the out-edge connection weight of node i; when the weighted average influence speed V i is less than the preset convergence judgment threshold, it is judged that i is the number of jumps required to reach the stable state; wherein, by statistically analyzing the probability density of the change of influence in the propagation data of the last complete business cycle, the lowest influence speed is selected as the preset convergence judgment threshold in the influence speed interval where the probability density is in the top three inverses;

[0102] The influence of the node in the dynamic data bloodline graph is calculated by the betweenness centrality and the closeness centrality, and the influence is obtained by weighted fusion; wherein, the betweenness centrality is used to measure the frequency of a specified node as an intermediate node on the shortest path between all other node pairs, by statistically analyzing the total number of shortest paths between all node pairs, marking the shortest path including the specified node as a specified path, counting the number of specified paths, and calculating the betweenness centrality by ratio of the number of specified paths to the total number of shortest paths; the closeness centrality is used to measure the comprehensive distance from a specified node to all other nodes, by summing the shortest paths from the specified node to all other nodes, and performing an inverse operation on the sum to obtain the closeness centrality; by statistically analyzing the betweenness centrality and the closeness centrality of the node in the dynamic data bloodline graph in the last complete business cycle, the influence is calculated by weighted calculation with different weights, and the standard deviation of the influence is calculated; the group of weights corresponding to the minimum standard deviation is used as the corresponding weighted weight;

[0103] By analyzing the propagation speed and range of the influence between different nodes in the node symbiotic body, key hub nodes and ordinary leaf nodes are identified;

[0104] Identifying key hub nodes and ordinary leaf nodes includes:

[0105] Extract the out-edge connection weight of all nodes in the dynamic data bloodline graph, and calculate the initial transition probability of the node; wherein, the out-edge connection refers to the directed connection from one node to other nodes; specifically, traverse all nodes in the dynamic data bloodline graph, and mark the nodes as target nodes in turn; after marking each target node, traverse all other nodes connected to the target node and pointing to the target node, which are called out-edge connection nodes of the target node, combine all out-edge connection nodes to obtain an out-edge connection list;

[0106] Obtain the direction attribute and establishment timestamp of the out-edge connection of the target node; set a sliding time window; wherein, the time window covers the most recent complete business cycle;

[0107] Count the total number of interaction events triggered by each out-edge connection within the time window; accumulate the total size of the data packets transmitted by each out-edge connection within the time window; according to the time interval of each interaction event and the current time interval, according to the decay coefficient of the time interval; wherein, the calculation formula of the decay coefficient is:

[0108] SJ = exp(-λ×ΔT); in the formula, SJ is the decay coefficient; λ is an attenuation rate parameter for controlling the decay speed of the weight with time, which is set according to business requirements;

[0109] ΔT is the time interval; exp() is the exponential function with natural number e as the base;

[0110] Multiply the number of each interaction event by the corresponding decay coefficient to obtain the weighted event quantity; multiply each data stream size by the corresponding decay coefficient to obtain the weighted data stream size; sum the weighted interaction event quantity and data stream size to obtain the original connection strength;

[0111] For each target node, sum the original connection total strength of all out-edges thereof; calculate the ratio of the original connection strength of each out-edge to the original connection total strength to obtain the initial transition probability;

[0112] Extract the data stream turnover rate and the number of adjacent nodes of the node, and calculate the restart preference probability combined with the intermediate centrality; specifically, within the same time window as above, the data stream turnover rate of the node is calculated; the number of out-edge connection nodes is calculated as the number of adjacent nodes; the data stream turnover rate, the number of adjacent nodes and the intermediate centrality are weighted and fused to obtain the node comprehensive scale; wherein, the weight is based on historical walk simulation to set different weight combinations, the standard deviation of the node comprehensive scale calculated by different weight combinations is calculated, and the weight corresponding to the minimum standard deviation is taken as the corresponding weighted weight; the node comprehensive scale is taken as the base of the probability distribution, and the softmax function is applied to normalize the node comprehensive scale of all nodes, and the normalization result is taken as the restart preference probability of the corresponding node;

[0113] The virtual walking particle swarm is initialized based on the initial transition probability, restart preference probability, and intrinsic sensitivity of the nodes, and multiple rounds of parallel random walk simulation are performed; specifically, this includes: obtaining the total number of nodes and edges in the dynamic data lineage graph, according to the formula:

[0114] The basic cover density BCD is calculated; where N is the total number of nodes; d is the edge density, based on the formula.

[0115] d = 2M / [N(N-1)] is calculated, where M is the total number of edges; C v The coefficient of variation is calculated by dividing the standard deviation of node degree by the mean of node degree; where node degree refers to the number of edges directly connected to a given node; this is expressed by the formula...

[0116] The initial total number of particles P is calculated; where k1 is the first scaling factor, usually taken as 2.5-3.5; the operator in the formula is an up-rounding function;

[0117] The intrinsic sensitivity of the node is normalized, and the result of the normalization is used as the initial number of particles for the corresponding node. The initial number of particles P is multiplied by the initial number allocation weight to obtain the initial number of particles for the node.

[0118] Based on the target node and the outgoing edge connection node, a corresponding walking path is generated; for the same target node, the original transition probability of its outgoing edge connection node is normalized, and the normalization result is used as the walking direction probability of the particle in the walking path; the initial number of particles of the target node is multiplied by the particle walking direction probability to obtain the number of particles walking in the walking path.

[0119] The influence of a node is normalized using the formula: E initial The initial energy value E is calculated as E0 + k2 × ZX. initial In the formula, ZX represents the normalized influence; k2 is the second scaling factor.

[0120] Through formula E cost =E base ×(1+w1×D n +w2×C n Calculate the estimated energy consumption E for the target node to jump to the outgoing edge connection node. cost In the formula, E cost E represents the estimated energy consumption required for the current node to migrate to candidate node n. base The basic energy constant is obtained through training based on historical node jump data; D n C is the topology depth coefficient of the node; nw1 represents the complexity coefficient of the node; w2 and w2 represent the topology depth coefficients D. n The complexity coefficient C of the nodes n The weights, w1 and w2, are obtained by training based on historical node jump big data;

[0121] For the topology depth coefficient D n The topology depth coefficient D is obtained by enumerating node pairs consisting of two nodes, counting the paths formed between each pair and the number of nodes contained in each path, and using this count as the node hop count of the path. For all paths between different node pairs, the path with the fewest nodes is selected and recorded as the global shortest path. The node pairs formed by the target node and its outgoing edge are marked as target node pairs. Among the paths of the target node pairs, the path with the fewest hops is marked as the target shortest path. The topology depth coefficient D is obtained by calculating the ratio of the hop count of the target shortest path to that of the global shortest path. n For the complexity coefficient C n Obtain the in-degree and out-degree of each node and sum them to get the total degree. Normalize the total degree of all nodes and use the result as the complexity coefficient C. n ;

[0122] Based on the initial energy value E of the particle initial Compared with the estimated energy consumption E cost The remaining energy value of a particle jumping from the target node to the outgoing edge connection node is calculated. If the remaining energy value is lower than the preset safe energy threshold, the outgoing edge connection node is determined to be an unreachable node; otherwise, the outgoing edge connection node is determined to be a reachable node. In this process, by performing a random walk simulation on the previous complete business cycle, the probability density statistics of the particle's historical remaining energy value are performed. Among the remaining energy value intervals with the probability density in the bottom three, the minimum remaining energy value is selected as the preset convergence judgment threshold.

[0123] The ratio of the remaining energy value to the initial energy value is calculated, and the result is used as the particle's active jump probability.

[0124] Obtain the list of reachable nodes for each particle. Normalize the remaining energy of particles in the list after they jump to different reachable nodes. Use the normalized energy as the jump weight for the current node to jump to the corresponding reachable node. Normalize the influence of each node using the formula: E initial The initial energy value E is calculated as E0 + k2 × ZX. initial In the formula, ZX represents the normalized influence; k2 is the second scaling factor.

[0125] Through formula E cost =E base ×(1+w1×D n +w2×Cn )calculating the estimated energy consumption E of the target node jumping to the out-edge connection node cost ; in the formula, E cost is the estimated energy consumption required for the current node to transfer to the candidate node n; E base is a basic energy constant, which is obtained by training based on historical node jump data; D n is a topological depth coefficient of the node; C n is a complexity coefficient of the node; w1 and w2 are weights of the topological depth coefficient D n and the complexity coefficient C n of the node respectively, and w1 and w2 are obtained by training based on historical node jump big data;

[0126] For the topological depth coefficient D n , by enumerating node pairs composed of two nodes, the number of nodes contained in the path formed between each node pair is counted as the node hop number of the path; for all paths between different node pairs, the path with the least number of nodes is selected and recorded as the global shortest path; the node pair formed by the target node and its out-edge connection node is marked as the target node pair, and in the path of the target node pair, the path with the least number of hops is marked as the target shortest path, and the topological depth coefficient D n is obtained by ratio calculation of the number of hops of the target shortest path and the global shortest path; for the complexity coefficient C n , the in-degree and out-degree of each node are obtained and summed to obtain the total degree, and the total degrees of all nodes are normalized to obtain the complexity coefficient C n ;

[0127] The walk path containing reachable nodes is marked as a reachable path, and the walk path containing unreachable nodes is marked as an unreachable path;

[0128] For the reachable path, the jumpable weight of the reachable node, the initial transfer probability and the active jump probability of the particle are multiplied, and the result obtained by multiplication is normalized to obtain the first jump success rate of the particle jumping to the corresponding reachable node in the reachable path, and the particle jumps according to the first jump success rate;

[0129] For the unreachable path, whether to jump is decided according to the active jump probability of the particle; when it is decided to jump, the restart preference probability of all reachable nodes is taken as the second jump success rate of jumping to the corresponding reachable node, and the particle jumps according to the second jump success rate;

[0130] According to random walk simulation, the expected residence time and the traffic convergence intensity of each node are extracted and calculated to identify the traffic convergence node; specifically, in the random walk simulation process, the total residence time of all particles visiting each node in the dynamic data bloodline graph and the total number of times of visiting the corresponding node are extracted, and the total residence time and the total number of times are calculated by ratio to obtain the expected residence time of the node;

[0131] In the random walk simulation process, the total number of jumps from other nodes to the specified node and the total number of jumps from the specified node to other nodes are counted, and are denoted as the total number of jumps into the specified node and the number of jumps out of the node, respectively; for the particles visiting the specified node, the number of different particles is counted as the total number of visiting particles of the specified node, and the total number of visiting particles is calculated by ratio with the total number of particles used in the simulation to obtain the concentration of the specified node; the traffic convergence intensity LQ is calculated by the formula LQ = a x log((TR+e) / (TC+e))+b x FJ; in the formula, TR is the total number of jumps into the specified node; TC is the total number of jumps out of the specified node; FJ is the concentration; e is a minimum value to prevent division by zero; a and b are corresponding weight coefficients, and the expected residence time and the traffic convergence intensity data obtained from historical random walk simulation are used to calculate the traffic convergence intensity by weighting with different weights, and the standard deviation of the obtained traffic convergence intensity is calculated, and the weight corresponding to the minimum standard deviation is taken as the corresponding weight coefficient;

[0132] The mean value of the expected residence time and the traffic convergence intensity is calculated, and the result of the mean value calculation is taken as the traffic convergence degree of the node, and the nodes are sorted in descending order according to the traffic convergence degree, and the nodes with high sorting positions are marked as traffic convergence nodes; for example, the top 3 nodes in the sorting can be marked as traffic convergence nodes;

[0133] By analyzing the damage degree of node failure to the connectivity of the dynamic data bloodline graph, the key hub node is determined;

[0134] Determining the key hub node includes:

[0135] Each traffic convergence node in the graph is marked as a failed node in turn, and the failed node is temporarily removed from the dynamic data bloodline graph, and all edges associated with the failed node are disconnected;

[0136] After removing the failed node each time, the following operations are performed:

[0137] Identify the remaining connected subgraphs in the dynamic data bloodline graph, calculate and compare the size proportion of each connected subgraph, and take the maximum size proportion as the connectivity retention rate; wherein, the subgraph refers to a subgraph composed of multiple nodes, and in the subgraph, any two nodes can reach each other through at least one path;

[0138] calculate the distribution uniformity of the size proportion of all connected subgraphs in the dynamic data bloodline graph; wherein, the distribution uniformity is used to measure the uniformity of the size of each connected subgraph after the node failure; the size proportion standard deviation of different connected subgraphs is calculated, and the size proportion standard deviation is taken as the distribution uniformity;

[0139] identify and count the number of isolated node groups as an isolation index; wherein, the isolated node group refers to a connected subgraph with very small node size formed after the node failure, and the very small node size can be that the number of nodes in the isolated node group is less than a preset node number threshold; the node numbers of the connected subgraphs in the previous complete business cycle are sorted in descending order, and the node number at 95% is taken as the preset convergence determination threshold; for example, node A is connected to node B, node B is connected to node C, node C is connected to node D, and node D is connected to node E, wherein D is a flow convergence node, the preset node number threshold is 2, and after D is removed, E has a node size of 1, which is less than the preset node number threshold, and E becomes an isolated node group;

[0140] the connectedness retention rate, the distribution uniformity and the reciprocal of the isolation index are weighted calculated, and the result of the weighted calculation is taken as the connectedness destruction contribution degree of the flow convergence node; wherein, the weight of the weighted calculation is obtained based on the importance evaluation of the connectedness destruction, and the importance evaluation can be realized in the form of expert scoring;

[0141] the flow convergence nodes are sorted in descending order according to the connectedness destruction contribution degree, and the flow convergence nodes at the front of the sorting are taken as the key hub nodes; specifically, the remaining flow convergence nodes are taken as ordinary leaf nodes; wherein, the flow convergence nodes at the front can be the flow convergence nodes at the top 3 positions;

[0142] generate a comprehensive sensitivity vector of the node according to the intrinsic sensitivity and influence of the key hub node;

[0143] cluster the nodes according to the comprehensive sensitivity vector to obtain a plurality of stable node clusters, and construct a node symbiotic body; specifically, a node is selected and recorded as a selected node, the Euclidean distance from the selected node to all other nodes is calculated, and a distance sorting sequence is generated; the distance of the Kth nearest neighbor node in the distance sorting sequence is selected as the reference distance, wherein K is a neighborhood scale parameter adapted according to the size of the data set, and is set as the integer part of the square root of the total number of nodes; the personalized neighborhood range of the target node is determined with z times of the reference distance as the neighborhood radius, wherein z is an expansion coefficient set according to the sparsity of data distribution;

[0144] The selected node calculates the comprehensive sensitivity vector similarity with other nodes within its personalized neighborhood range; the relative difference degrees of intrinsic sensitivity and influence between the comprehensive sensitivity vectors are calculated respectively, and the calculation formula is: CYD = γ × (a-b); in the formula, a and b are the normalized values of two same-dimension data respectively; γ is the specificity weight of the dimension, which is determined based on the clustering full risk assessment of historical comprehensive sensitivity vectors; the two difference degrees are weighted and summed, and the weighted sum result is taken as the comprehensive sensitivity vector similarity;

[0145] The Euclidean distance of the selected node to other nodes is multiplied by the vector similarity of the corresponding comprehensive sensitivity vector to obtain the association strength of the selected node and other nodes; the association strengths of the selected node and all nodes within the personalized neighborhood are weighted and accumulated to obtain the coarse-grained density value; wherein, different weights are given according to the topological role difference of the nodes within the personalized neighborhood, and the association strength contribution of the key hub node is higher than that of the ordinary leaf node; the weighted accumulation sum is normalized by the neighborhood volume, and the normalized value is taken as the distribution density of the node;

[0146] The local maximum points of the distribution density are identified as initial cluster center candidate points; the vector distance and density ratio between adjacent candidate cluster center candidate points are calculated, and the separation degree index is obtained by weighted calculation; when the separation degree index exceeds the preset separation degree threshold, the pair of cluster centers is confirmed to be retained; wherein, the vector distance and density ratio between the candidate points in the last complete business cycle are weighted by different weights, the standard deviation of the obtained separation degree index is calculated, and the group of weights corresponding to the minimum standard deviation is taken as the weighted weight of the vector distance and density ratio; the probability density of the separation degree index between the candidate points in the last complete business cycle is statistically analyzed, and the minimum separation degree index is selected as the preset separation degree threshold in the interval of the separation degree index whose probability density is in the top three inverses;

[0147] All candidate points are iteratively processed, and the number of remaining cluster centers separated from each other is generated according to the nodes within different cluster centers;

[0148] The rule construction module calculates the interaction risk of the node symbiont and the external entity, and constructs the minimum access control rule according to the interaction risk;

[0149] The minimum access control rule is constructed, including:

[0150] An abnormal interaction event affecting the security of the node symbiont is identified, and the sink point, operation type and intrinsic sensitivity of the abnormal interaction event are extracted and combined to obtain abnormal interaction features;

[0151] The abnormal interaction event affecting the security of the node symbiont is identified, including:

[0152] Based on the dynamic data bloodline graph, real-time interaction events between key hub nodes are extracted, and interaction events not in the same node symbiotic body are identified as cross-node set events; wherein the interaction events not in the same node symbiotic body refer to interaction events in which only the source endpoint of the interaction event is in the corresponding node symbiotic body;

[0153] The path depth, the number of transit nodes and the number of boundary jumps of the cross-node set event are extracted and weightedly fused, and the weightedly fused result is taken as the confidence of the cross-node set event; wherein when the path length abnormally increases, it will gradually jump from a compromised node to other internal nodes, resulting in an access path much longer than the normal business path; the number of transit nodes is the number of intermediate nodes, the more the transit nodes, the more complex the path, which means that the data is exposed to more system components during transmission, and the potential attack surface is larger; the number of boundary jumps is the number of times the interaction event data flow passes through the node community during transmission; when there is continuous and multiple boundary jumps in a path, it does not belong to normal business flow, and there may be an attack chain with the risk of data leakage; the significance of weighted fusion is to construct a risk portrait of the interaction event and form a confidence indicating the comprehensive risk; wherein the weight is based on the correlation strength of each parameter in the historical abnormal interaction data and the security threat, and the contribution of each parameter to the overall risk is determined by statistical analysis;

[0154] The confidence decay sequence of the cross-node event is calculated, and the sequence segment with a confidence decay rate exceeding a preset decay threshold is identified as an abnormal interaction event; specifically including: constructing a confidence decay curve according to the confidence decay sequence, and determining an abnormal interaction event when the instantaneous slope of the confidence decay curve exceeds the preset decay threshold; by statistically analyzing the slope of the confidence decay curve during the historical confidence fluctuation period in the last complete business cycle, the maximum slope is selected as the preset decay threshold in the top three slope intervals of the probability density;

[0155] According to the abnormal interaction characteristics, the real-time data stream is risk assessed to obtain an abnormal risk; specifically including: statistics of the total number of historical interaction events pointing to the sink endpoint in the last complete business cycle, statistics of the number of abnormal interaction events, recorded as the first abnormal number, and ratio calculation of the first abnormal number and the total number of historical interaction events to obtain the first risk index; statistics of the number of abnormal interaction events caused by the operation type in the last complete business cycle, recorded as the second abnormal number, and ratio calculation of the second abnormal number and the total number of historical interaction events to obtain the second risk index; weightedly fusing the first risk index, the second risk index and the intrinsic sensitivity to obtain the abnormal risk; wherein the weight of the fusion is based on the statistical correlation strength of the two risk indicators and the intrinsic sensitivity with the occurred security events respectively;

[0156] The access constraint condition is constructed according to the abnormal risk, and is used as a minimum access control rule between nodes; specifically, historical abnormal risk data is obtained and normalized, and the abnormal risk is mapped to the interval (0, 1); a plurality of constraint actions are predefined, and the abnormal risk is mapped to the corresponding constraint action through a dynamic mapping function; for example, if the abnormal risk is in the low risk interval, such as [0.1, 0.3), it is mapped to an observation type action, for example: detailed logs are recorded for post-audit and analysis; if the abnormal risk is in the low-to-medium risk interval, such as [0.3, 0.6), it is mapped to a light constraint type action, such as limiting the request rate of the session; if the abnormal risk is in the medium-to-high risk interval, such as [0.6, 0.8), it is mapped to a strong intervention type action, such as reducing the current session authority and limiting the access range; if the abnormal risk is in the high risk interval, such as [0.8, 0.1), it is mapped to a blocking type action, such as immediately terminating the current session; each constraint type action is used as a minimum access control rule of the node symbiont and the external entity;

[0157] The policy management module extracts the time value of the real-time data stream, constructs a life cycle management model according to the decay rate of the time value and the dynamic data blood atlas, and dynamically manages the access control policy;

[0158] The life cycle management model is constructed, including:

[0159] The heat of the real-time data stream and the derived new data flow are obtained, the time value curve is fitted, and the value change rate is extracted;

[0160] The heat of the real-time data stream and the derived new data flow are obtained, the time value curve is fitted, and the value change rate is extracted; wherein the heat of the data stream includes the access frequency and the reference frequency of the data stream, the access frequency of the data stream refers to the number of times that an external entity (user, application program, system process) initiates and successfully establishes a session to read or consume the content of the data stream, the target address (such as URL, API endpoint) of the request, HTTP method (GET / POST) and response status code are extracted through port mirroring at the node (such as core switch, API gateway) using a deep packet analysis engine, and the access times of known data streams are aggregated and counted according to the data stream identifier (such as URI, topic name), the total access times of all data streams accessed by the external entity in the last complete business cycle are counted, the access times of the known data stream are compared with the total access times, and the calculation result is used as the access frequency of the known data stream;

[0161] The data stream is cited frequency refers to the content, identifier or conclusion of the data stream is cited as the component or decision basis of other newly generated data stream, file, report or code, by generating a unique content fingerprint (such as SHA-256 hash value) for each data stream, when a new data stream is generated, the system will scan its content to detect whether it includes the content fingerprint of the known data stream, if it includes, it is recorded as the known data stream is cited once, the generation frequency of the new data stream in the last complete business cycle is calculated, and the generation frequency of the known data stream is calculated as the citation frequency of the corresponding known data stream; the access frequency and the citation frequency are summed up to obtain the heat;

[0162] The derived new data flow refers to the number of new data streams generated after processing, calculation, aggregation or conversion with the data stream as the main and direct input, which has independent identification;

[0163] The heat and the derived new data flow are multiplied to obtain the time value, the time value of the same real-time data stream at different time points is calculated, and the time value curve is generated;

[0164] By analyzing the long-term trend of the time value curve, a classification label is added to the real-time data stream; the time scale of the long-term trend is determined based on the trend decomposition of the time value curve and the time distribution of the big data and data stream; the long-term trend of the value change rate can be realized by trend decomposition based on AI time series model, such as ARIMA model, which will not be described in detail; the time value curve with a sustained downward trend is added with a high decay label; the time value curve with a stable and no obvious slope trend is added with a stable label; the time value curve with a sustained upward trend is added with a value-added label;

[0165] The interval length of the short-term change feature of the time value curve is extracted, which is recorded as the feature hiding length under the classification label; specifically, for high decay type data stream, exponential decay model is used to segment fit its value change curve, and the interval length of the value falling to the preset safety value threshold is marked as invalid hidden length; for stable type data stream, segmented constant model is used to describe its value, and the length of time value change rate greater than the preset value rate threshold is marked as fluctuation hidden length; for value-added type data stream, the value accumulation is segmented and fitted by using the logic growth model, and the interval length of the value growth rate lower than the preset growth rate threshold is marked as mature hidden length; wherein, the probability density of the time value of the data stream in the last complete business cycle is statistically analyzed, and the minimum time value is selected as the preset safety value threshold in the time value interval where the probability density is in the top three;

[0166] According to the feature hiding time length under different classification labels, the feature expected time length is calculated, and a life cycle management model is constructed; specifically, the real-time data stream of the most recent complete business cycle is sampled as sample data stream, and the feature hiding time length is determined for each sample data stream; for different sample data streams, the mean value of the feature hiding time length of the same kind is calculated to obtain the feature expected time length; in the data stream of each node symbiont and external entity, the minimum access control rule is obtained, and for each data stream, based on the long-term change trend of the time value, the corresponding classification label is added, the feature expected time length under the classification label is taken as the life cycle of the minimum access control rule, and the corresponding data stream is controlled; that is, when it is identified which label the data stream belongs to, the corresponding feature expected time length is added to the minimum access control rule as the life cycle, and only the data stream is controlled; when the life cycle ends, the long-term trend of the data stream is re-identified and the life cycle is updated;

[0167] The compilation control module compiles the data control strategy that can be loaded in real time according to the minimum access control rule and the life cycle management model, and controls the real-time data;

[0168] In view of the defects of the existing access control strategy, such as lack of detailed risk assessment and time dimension management, leading to non-detailed strategy and resource waste, by analyzing the interaction risk of the node symbiont and the external entity, the "minimum" rule is generated, which only grants necessary permissions, and the accurate delivery of permissions is realized; at the same time, by fitting the data value decay curve and adding classification labels to the data, a life cycle management model is constructed, so that the access control strategy can be updated adaptively with the data value, realizing the adaptive transformation and update of the control strategy; solve the common problems of permission accumulation and low-value data occupying high-security resources in the prior art, achieve the synergistic effect of dynamically optimizing the permission structure and system resource allocation while ensuring security, and improve the intelligence and efficiency of the overall data security management.

[0169] Embodiment two, please refer to Figure 3 The embodiment two of the present application provides a smart city data processing platform based on artificial intelligence, which is applied to the smart city data processing method based on artificial intelligence provided in embodiment one, and comprises:

[0170] The data acquisition module monitors the real-time data between network entities in real time to obtain real-time data streams;

[0171] The data management analysis module constructs a dynamic data blood relationship map by analyzing the instantaneous behavior mode and content derivative relationship of the real-time data stream;

[0172] The data aggregation analysis module constructs a node symbiont by analyzing the data difference degree between different nodes in the dynamic data blood relationship map;

[0173] A rule construction module is configured to calculate an interaction risk of the node symbiont and an external entity, and to construct a minimum access control rule according to the interaction risk.

[0174] A policy management module is configured to extract a time value of a real-time data stream, to construct a life cycle management model according to a decay rate of the time value and a dynamic data bloodline graph, and to dynamically manage an access control policy.

[0175] A compiling control module is configured to compile a real-time loadable data management strategy according to the minimum access control rule and the life cycle management model, and to manage the real-time data.

[0176] In the third embodiment, the smart city data storage medium based on artificial intelligence is applied to the smart city data processing platform based on artificial intelligence in the first embodiment, and stores a computer program. The computer program can be run by a processor of a device where the storage medium is located, so as to realize the smart city data processing method based on artificial intelligence in the first embodiment.

[0177] The above merely provides preferred embodiments of the present application but is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can make modifications to the technical solutions recorded in the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall fall within the protection scope of the present application.

[0178] It should be noted that, in this document, the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or device including the element.

[0179] The formulas in the present specification are dimensionless numerical calculations, the formulas are obtained by software simulation of a large amount of data to obtain a formula of the most recent real situation, and the preset parameters and threshold values in the formula are set by a person skilled in the art according to the actual situation.

[0180] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, replacements, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.

Claims

1. A smart city data processing method based on artificial intelligence, characterized in that, include: Real-time monitoring of data between network entities to obtain real-time data streams; By analyzing the instantaneous behavioral patterns and content derivation relationships of real-time data streams, a dynamic data lineage map is constructed. By analyzing the degree of data differences between different nodes in the dynamic data lineage map, a node symbiosis is constructed. Calculate the interaction risk between the node symbiont and external entities, and construct a minimum access control rule based on the interaction risk; Extract the timeliness value of real-time data streams, and construct a lifecycle management model based on the decay rate of timeliness value and dynamic data lineage graph to dynamically manage access control policies; Data management strategies that can be loaded in real time are compiled based on minimal access control rules and lifecycle management models to manage real-time data.

2. The smart city data processing method based on artificial intelligence according to claim 1, characterized in that, The construction of the dynamic data lineage map includes: The real-time data stream is parsed into discrete interactive events, and the source and destination endpoints and content fingerprint of each interactive event are labeled. Calculate the semantic association strength between consecutive interactive events, and construct interactive event pairs based on two interactive events with semantic association strength higher than a preset strength threshold; the semantic association strength is calculated based on the hash fragment overlap rate of interactive events and the temporal combination probability of instruction fragments. Determine the propagation direction of related events based on the chronological order of their occurrence. Using endpoints as nodes, the propagation direction of interactive event pairs as directed edges, and the semantic association strength as edge weights, a derivative link of interactive events is constructed. Identify abrupt changes and deviations in endpoint behavior during interactive events, and quantify the trust correlation between endpoints; the trust correlation is based on a similarity measurement of the interaction flow data between source and destination endpoints and the outbound interaction flow data of the source endpoint; By integrating the trust relationships between derived links and different endpoints, a dynamic data lineage map is obtained.

3. The smart city data processing method based on artificial intelligence according to claim 2, characterized in that, The constructed node symbiotic includes: The frequency of visits and path depth of each node in the dynamic data lineage graph are calculated, and the intrinsic sensitivity of the node is obtained by weighted fusion. The influence of a node is obtained by weighting and fusing the betweenness centrality and proximity centrality of the node in the dynamic data lineage graph. By analyzing the propagation speed and range of influence between different nodes within a node symbiosis, key hub nodes and ordinary leaf nodes can be identified. Generate a comprehensive sensitivity vector for nodes based on their inherent sensitivity and influence. Clustering nodes based on the comprehensive sensitivity vector yields multiple stable node clusters, thus constructing a node symbiosis.

4. The smart city data processing method based on artificial intelligence according to claim 3, characterized in that, The construction of the minimum access control rules includes: Identify abnormal interaction events that affect the security of node symbionts, extract and combine the destination endpoints, operation types and intrinsic sensitivities of abnormal interaction events to obtain abnormal interaction features; Based on the characteristics of abnormal interactions, a risk assessment is performed on the real-time data stream to obtain the abnormal risk. Access constraints are constructed based on abnormal risks and used as minimal access control rules between nodes.

5. The smart city data processing method based on artificial intelligence according to claim 3, characterized in that, The identification of key hub nodes and ordinary leaf nodes includes: Extract the outgoing edge connection weights of all nodes in the dynamic data pedigree graph and calculate the initial transition probability of the nodes; Extract the data flow rate and the number of adjacent nodes of a node, and calculate the restart preference probability by combining the degree center number; The virtual walking particle swarm is initialized based on the initial transition probability, restart preference probability, and intrinsic sensitivity of the nodes, and multiple rounds of parallel random walk simulation are performed. Based on random walk simulation, the expected dwell time and traffic convergence intensity of each node are extracted and calculated to identify traffic convergence nodes; By analyzing the extent to which node failures disrupt the connectivity of dynamic data lineage graphs, key hub nodes are identified.

6. The smart city data processing method based on artificial intelligence according to claim 4, characterized in that, The identification of abnormal interaction events affecting the security of the node symbiote includes: Based on dynamic data lineage graphs, interaction events between key hub nodes are extracted in real time, and interaction events that are not in the same node symbiotic body are identified and recorded as cross-node set events. Extract the path depth, number of transit nodes, and number of boundary jumps of cross-node set events and fuse them in a weighted manner. Use the weighted fusion result as the confidence level of cross-node set events. Calculate the confidence decay sequence of cross-node events, identify sequence segments whose confidence decay rate exceeds a preset decay threshold, and mark the cross-node set events corresponding to the sequence segments as abnormal interaction events.

7. The smart city data processing method based on artificial intelligence according to claim 5, characterized in that, The determination of key hub nodes includes: Each traffic convergence node in the graph is marked as a failed node in turn. The failed nodes are temporarily removed from the dynamic data lineage graph, and all edges associated with the failed nodes are disconnected. After each failed node is removed, perform the following operations: Identify the remaining connected subgraphs in the dynamic data lineage graph, calculate and compare the size proportion of each connected subgraph, and take the largest size proportion as the connectivity preservation rate; Calculate the uniformity of the distribution of the size proportion of all connected subgraphs in a dynamic data lineage graph; Identify and count the number of isolated node clusters as an isolation indicator; The connectivity retention rate, distribution uniformity, and the reciprocal of the isolation index are weighted and calculated, and the result of the weighted calculation is used as the connectivity disruption contribution of the traffic aggregation node. Traffic aggregation nodes are sorted in descending order based on their contribution to connectivity disruption, and the top-ranked traffic aggregation nodes are designated as key hub nodes.

8. The smart city data processing method based on artificial intelligence according to claim 1, characterized in that, The construction of the lifecycle management model includes: To obtain the popularity of real-time data streams and the new data traffic generated therefrom, fit the timeliness value curve, and extract the rate of value change; By analyzing the long-term trend of the timeliness value curve, classification tags are added to the real-time data stream; The interval between the short-term changes in the time-sensitive value curve is extracted and denoted as the feature hiding time under the classification label. Calculate the expected duration of features based on the feature hiding duration under different classification labels, and construct a lifecycle management model.

9. An AI-based smart city data processing platform, applied to the AI-based smart city data processing method described in any one of claims 1 to 8, characterized in that, include: The data acquisition module monitors real-time data between network entities and acquires real-time data streams. The data management and analysis module constructs a dynamic data lineage graph by analyzing the instantaneous behavioral patterns and content derivation relationships of real-time data streams. The data aggregation and analysis module constructs a node symbiosis by analyzing the degree of data differences between different nodes in the dynamic data lineage graph. The rule building module calculates the interaction risk between the node symbiont and external entities, and builds minimal access control rules based on the interaction risk; The policy management module extracts the timeliness value of real-time data streams, and constructs a lifecycle management model based on the decay rate of timeliness value and dynamic data lineage graph to dynamically manage access control policies. The compilation control module compiles real-time data management strategies based on minimal access control rules and lifecycle management models to manage real-time data.

10. A smart city data storage medium based on artificial intelligence, applied to the smart city data processing method based on artificial intelligence as described in claims 1 to 8, characterized in that, It stores computer programs that can be run by the processor of the device where the storage medium is located, in order to implement a smart city data processing method based on artificial intelligence.

Citation Information

Cited By

  • Vehicle electronic tag intelligent management method based on big data processing

    CN122090623A

  • Intelligent Management Method for Vehicle Electronic Tags Based on Big Data Processing

    CN122090623B