Data Security Risk Early Warning Method and System Based on Big Data Analysis

By constructing a data security risk feature map and a distributed risk feature learning network, multi-source heterogeneous data from mobile communication networks are collected and analyzed in real time. This solves the problem of difficulty in identifying and warning of data security risks in existing technologies, and enables accurate risk warning and rapid response for mobile communication networks, thereby improving network security and reliability.

CN120935575BActive Publication Date: 2026-05-05CHINA MOBILE COMM GRP TIBET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE COMM GRP TIBET CO LTD
Filing Date
2025-08-01
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for integrating and analyzing multi-source heterogeneous data in mobile communication networks, making it difficult to comprehensively identify data security risks, cope with complex security threats such as identity theft and cyberattacks, and lack targeted analysis and early warning mechanisms.

Method used

A data security risk feature map is constructed. Heterogeneous data is collected in real time through multi-source data access interfaces, and distributed cleaning and standardization processing is performed to generate a multi-source heterogeneous data fusion stream. A distributed risk feature learning network is used for feature mapping and correlation enhancement processing to generate real-time risk feature vectors. Big data correlation analysis is performed to determine the risk diffusion level and key impact nodes, and security risk warning instructions are generated.

Benefits of technology

It enables precise risk warnings for mobile communication networks, improves the timeliness and accuracy of risk identification, reduces the impact of data security risks on the network, and ensures the safe and stable operation of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935575B_ABST
    Figure CN120935575B_ABST
Patent Text Reader

Abstract

This invention provides a data security risk early warning method and system based on big data analysis. First, a data security risk feature map is constructed. Heterogeneous data sets from mobile communication networks are collected in real time through multi-source data access interfaces. After distributed cleaning and standardization, a multi-source heterogeneous data fusion stream is generated. This stream is then input into a pre-defined distributed risk feature learning network. Based on the data security risk feature map, feature mapping and association enhancement processing are performed to obtain real-time risk feature vectors. Big data association analysis is then conducted on these real-time risk feature vectors to uncover the transmission dependencies of risk features, generating a risk propagation path weight set. Finally, the risk diffusion level and key influencing nodes are determined, and a security risk early warning command containing risk propagation path identifiers is generated and pushed to the mobile communication security management platform. This achieves accurate early warning and rapid response to data security risks, ensuring the safe and stable operation of the mobile communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and more specifically, to a data security risk early warning method and system based on big data analysis. Background Technology

[0002] With the rapid development of mobile communication networks, data volume is exploding, and data security issues are becoming increasingly prominent. Traditional data security measures mainly focus on handling single data sources or specific security events. For example, user authentication typically uses simple username and password verification methods, which are insufficient to cope with increasingly complex identity theft and fraud methods; for access behavior, only basic information such as access time and location is considered, failing to comprehensively analyze abnormal access patterns; in signaling interaction, traditional methods only perform basic format checks on signaling, making it difficult to detect potential security threats hidden in signaling interactions; during data transmission, data security is mainly ensured by encryption technology, but there is a lack of effective monitoring of abnormal traffic and security risks in transmission paths; and network slicing, as a key technology in mobile communication networks, also lacks targeted analysis and early warning mechanisms for its security risks.

[0003] Furthermore, existing technologies lack effective methods for fusion and processing when analyzing multi-source heterogeneous data, making it difficult to uncover potential relationships between different data sources and to comprehensively and accurately identify data security risks. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a data security risk early warning method based on big data analysis, the method comprising:

[0005] Construct a data security risk feature map, which includes a network of relationships between user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features;

[0006] Heterogeneous data sets in mobile communication networks are collected in real time through multi-source data access interfaces. The heterogeneous data sets are then subjected to distributed cleaning and standardization processing to generate a multi-source heterogeneous data fusion stream.

[0007] The multi-source heterogeneous data fusion stream is input into a preset distributed risk feature learning network, and feature mapping and association enhancement processing are performed based on the data security risk feature map to obtain a real-time risk feature vector.

[0008] Big data correlation analysis is performed on the real-time risk feature vectors to uncover the transmission dependencies between risk features and generate a set of risk propagation path weights.

[0009] Based on the risk propagation path weight set, the risk diffusion level and key impact nodes are determined, a security risk warning instruction containing the risk propagation path identifier is generated, and the security risk warning instruction is pushed to the mobile communication security management platform.

[0010] In another aspect, embodiments of the present invention also provide a data security risk early warning system based on big data analysis, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-described method.

[0011] Based on the above, this invention constructs a data security risk feature map that includes a relational network of user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features. It collects heterogeneous datasets from mobile communication networks in real time, merges them for distributed cleaning and standardization, and generates a multi-source heterogeneous data fusion stream. This multi-source heterogeneous data fusion stream is input into a preset distributed risk feature learning network. Based on the data security risk feature map, feature mapping and association enhancement processing are performed to obtain real-time risk feature vectors. This accurately captures the dynamic changes of data security risks, improving the timeliness and accuracy of risk identification. Big data association analysis is performed on the real-time risk feature vectors to uncover the transmission dependencies between risk features and generate a risk propagation path weight set. This helps to deeply understand the propagation mechanism of data security risks and provides a scientific basis for taking effective prevention and control measures. Based on the risk propagation path weight set, the risk diffusion level and key impact nodes are determined, and a security risk warning instruction containing risk diffusion path identifiers is generated and pushed to the mobile communication security management platform. This achieves accurate early warning and rapid response to data security risks, effectively reducing the impact of data security risks on mobile communication networks, ensuring the safe and stable operation of mobile communication networks, and improving the security and reliability of the entire mobile communication system. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the execution flow of the data security risk early warning method based on big data analysis provided in the embodiments of the present invention.

[0013] Figure 2 This is a schematic diagram of exemplary hardware and software components of a data security risk early warning system based on big data analysis provided in an embodiment of the present invention. Detailed Implementation

[0014] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1This is a flowchart illustrating a data security risk warning method based on big data analysis provided in one embodiment of the present invention. The following is a detailed description of this data security risk warning method based on big data analysis.

[0015] Step S110: Construct a data security risk feature map, which includes a network of relationships between user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features.

[0016] In the scenario of early warning of data security risks in mobile communication networks, the construction of the aforementioned relational network is to clearly present the interactions and influence paths between various features. In this embodiment, it is first necessary to comprehensively share user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features, clarifying their respective connotations and possible relationships.

[0017] For example, changes in user identity characteristics may affect access behavior characteristics, and anomalies in access behavior characteristics may in turn trigger anomalies in signaling interaction characteristics, thereby affecting data transmission characteristics and network slicing characteristics.

[0018] Step S111: Obtain the core feature dimensions of the data security risk domain of mobile communication networks and determine the basic feature set. The basic feature set includes user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features. User identity features include user identifier type, authentication credential status, and permission level label. Access behavior features include access time distribution, access location sequence, and access device fingerprint. Signaling interaction features include signaling type sequence, signaling field variation, and signaling node hopping. Data transmission features include transmission protocol type, encryption status identifier, and transmission traffic fluctuation. Network slicing features include slice identifier information, resource allocation parameters, and isolation policy configuration.

[0019] In this embodiment, obtaining the core feature dimensions requires in-depth analysis of the operating mechanism of the mobile communication network and data security risks. Throughout the entire process from user access to network to data transmission, there may be various risk factors, each with different corresponding feature dimensions.

[0020] Regarding user identity characteristics, user identifier types can be identifiers based on mobile phone numbers, temporary identifiers based on network allocation, etc.; authentication credential status includes credential validity, credential expiration, credential forgery, etc.; permission level labels are divided into different levels according to user business needs and network management policies, such as ordinary user permissions, administrator permissions, etc.

[0021] Among the access behavior characteristics, access time distribution refers to the regularity of user access to the network during different time periods, such as specific time periods on weekdays and different time periods on weekends; access location sequence is a sequence of geographical location information of users when accessing the network in chronological order, which can be obtained through base station positioning and other methods; access device fingerprint is the unique identification information of the device, such as the characteristics formed by the combination of the device's hardware model, operating system version, network card address, etc.

[0022] In signaling interaction features, signaling type sequence is the sequence of various signaling types generated in order during communication, such as call setup signaling and location update signaling; signaling field variation refers to abnormal changes in certain fields in signaling, such as abnormal field length or field content that does not conform to protocol specifications; signaling node jump is the change in the path of signaling transmission between network nodes, such as the order of nodes that the signaling passes through under normal circumstances not matching the actual situation.

[0023] Among the characteristics of data transmission, the transmission protocol type includes Transmission Control Protocol, User Datagram Protocol, etc.; the encryption status identifier is divided into unencrypted, encrypted but the encryption algorithm is insecure, encrypted and the encryption algorithm is secure, etc.; transmission traffic fluctuation refers to the change in the amount of data transmitted per unit time, such as a sudden surge or drop in traffic.

[0024] In terms of network slicing characteristics, slice identification information is a unique identifier used to distinguish different network slices; resource allocation parameters include the bandwidth, computing resources, storage resources, etc. allocated to the slice; isolation policy configuration is a strategy to ensure that data and resources between different slices do not interfere with each other, such as specific configuration methods such as logical isolation and physical isolation.

[0025] Step S112: Extract risk feature combinations and causal relationships between features from historical data security incident cases, and establish a risk feature association rule base. The risk feature association rule base includes feature co-occurrence frequency, feature propagation delay, and feature impact weight.

[0026] In this embodiment, extracting relevant information from historical data security incident cases requires collecting a large number of past mobile communication network data security incidents. These incident cases should cover different types of security risks, such as data breaches, cyberattacks, and unauthorized access.

[0027] For each incident case, it is necessary to analyze the combination of risk characteristics involved, that is, which user identity characteristics, access behavior characteristics, etc. appear simultaneously or sequentially. For example, in a data breach incident, there may be a combination of characteristics such as abnormal user authentication credential status, unfamiliar fingerprints of access devices, and abnormal fluctuations in transmission traffic.

[0028] Next, the causal relationships between these features are analyzed. For example, does an abnormal user authentication credential status lead to an unfamiliar access device fingerprint, or does an unfamiliar access device fingerprint cause abnormal fluctuations in transmission traffic? Through the analysis of multiple cases, common causal relationship patterns between features are summarized.

[0029] When establishing a risk feature association rule base, feature co-occurrence frequency refers to the ratio of the number of times two or more features appear simultaneously in multiple event cases to the total number of cases; feature propagation delay refers to the time interval between the appearance of one feature and the subsequent appearance of another related feature; feature influence weight is a parameter that measures the degree of influence of one feature on another feature, and the greater the influence, the higher the weight.

[0030] Step S113: Construct a feature node layer based on the basic feature set. Each feature node corresponds to a specific feature in the basic feature set. Assign a unique feature identifier and feature attribute description to each feature node. The feature attribute description includes the feature data type, feature acquisition source, and feature security level.

[0031] In this embodiment, a feature node is created for each specific feature based on the previously determined set of basic features. For example, the specific feature of user identity type in the user identity feature corresponds to a feature node; the access time distribution in the access behavior feature also corresponds to a feature node.

[0032] The unique identifier assigned to each feature node can be a string consisting of letters and numbers, ensuring that it is not repeated throughout the entire graph. For example, the feature identifier for a user identifier type could be "UAT001", and the feature identifier for access time distribution could be "ATD001", etc.

[0033] In the description of feature attributes, the data type of the feature is determined according to the nature of the feature. For example, the data type of the user identifier type can be a string, the data type of the access time distribution can be a time series, and the data type of the transmission traffic fluctuation can be a numeric type, etc.

[0034] The source of feature collection is clearly defined, such as where the feature data is obtained. For example, user identification type can be collected from the user subscription information database, access time distribution can be collected from the access records of the base station, and transmission traffic fluctuation can be collected from the core network user plane interface, etc.

[0035] Security levels are determined by the sensitivity of the feature data. For example, features related to user identity have a higher security level, while features related to access time distribution have a relatively lower security level. The classification of security levels can be referenced in relevant cybersecurity standards and specifications.

[0036] Through the above operations, a feature node layer is constructed. Each node has a clear feature identifier and attribute description, so that each feature has clear individual characteristics in the graph.

[0037] Step S114: Construct a feature association edge layer based on the risk feature association rule base. Each feature association edge connects two feature nodes that have an association relationship. Configure an association strength parameter and a transmission direction identifier for each feature association edge. The association strength parameter is calculated based on the feature co-occurrence frequency and influence weight. The transmission direction identifier is used to indicate the direction of influence between features.

[0038] In this embodiment, based on information in the risk feature association rule base, it is determined which feature nodes are related. For example, if the feature node "abnormal user identifier type" is related to the feature node "unfamiliar access device fingerprint," then a feature association edge is used to connect them.

[0039] The association strength parameter is calculated by combining feature co-occurrence frequency and feature influence weight. Specifically, the feature co-occurrence frequency and feature influence weight are first normalized to ensure they are within the same dimension. Then, they are weighted according to a set ratio to obtain the association strength parameter. For example, if the normalized value of the feature co-occurrence frequency is A, the normalized value of the feature influence weight is B, and the weight ratios are 0.6 and 0.4, then the association strength parameter is equal to A multiplied by 0.6 plus B multiplied by 0.4.

[0040] The direction of influence is indicated by specific symbols or markings, such as an arrow pointing from the source node to the affected node, indicating the direction of influence between features. For example, if an abnormal user authentication credential status points to an unfamiliar fingerprint on the access device, it means that the former may affect the latter.

[0041] By configuring associated edges and related parameters for each pair of associated nodes, a feature-related edge layer is constructed, which clearly presents the degree of association and the direction of influence between feature nodes.

[0042] Step S115: Perform topological integration of the feature node layer and the feature association edge layer to generate a data security risk feature map containing node attributes, edge attributes, and topological structure. The topological structure is used to describe the direct association relationship and indirect association path between feature nodes. The data security risk feature map supports querying the association strength parameters of associated features and association paths based on feature nodes.

[0043] In this embodiment, topology integration organically combines the feature node layer and the feature association edge layer to form a complete network structure. During the integration process, it is necessary to ensure that each feature association edge accurately connects the corresponding two feature nodes, and that the attributes of the nodes and the attributes of the edges correspond correctly.

[0044] Topological structures visually represent the direct relationships between characteristic nodes through the arrangement of nodes and the connection of edges. For example, two nodes directly connected by an edge are directly related. Indirect relationships, on the other hand, are paths formed by connecting multiple intermediate nodes and edges. For instance, if node A is connected to node C through node B, then there is an indirect relationship path from A to C.

[0045] Once the data security risk feature map is built, it will have a query function. When a feature node is entered, all feature nodes directly and indirectly associated with that node, as well as the association strength parameters on these association paths, can be retrieved. For example, querying the node "abnormal user authentication credential status" will yield the associated unfamiliar access device fingerprint nodes, their association strength parameters, and also the nodes with abnormal transmission traffic fluctuations associated with these unfamiliar access device fingerprint nodes, along with their corresponding association strength parameters.

[0046] Step S120: Collect heterogeneous data sets in the mobile communication network in real time through the multi-source data access interface, perform distributed cleaning and standardization processing on the heterogeneous data sets, and generate a multi-source heterogeneous data fusion stream.

[0047] In mobile communication network data security risk early warning scenarios, real-time and comprehensive data collection is fundamental to risk analysis. Due to the diverse sources of data in the network, and the different data formats and types, data needs to be collected through multi-source data access interfaces and processed to generate a fused stream for subsequent analysis.

[0048] Step S121: Construct a multi-source data access interface. The multi-source data access interface includes a radio access network data interface, a core network control plane interface, a core network user plane interface, a network slice management interface, and a user equipment interaction interface. The radio access network data interface is used to collect access data from the base station side. The core network control plane interface is used to collect signaling data related to mobility management and session management. The core network user plane interface is used to collect user data transmission records. The network slice management interface is used to collect slice resource allocation and status monitoring data. The user equipment interaction interface is used to collect user equipment operation behavior data.

[0049] In this embodiment, constructing a multi-source data access interface requires designing corresponding interface specifications and communication protocols based on different data sources and data types.

[0050] The wireless access network data interface establishes a connection with each base station and collects access data from the base station side according to the set format and frequency. This data includes the time of user access to the base station, location information, and preliminary identification of the access device.

[0051] The core network control plane interface is connected to the control part of the core network and is specifically used to collect mobility management-related signaling, such as signaling during user location updates and base station handover, as well as session management-related signaling, such as signaling for establishing and releasing data connections.

[0052] The core network user plane interface focuses on collecting user data transmission records, including the amount of data transmitted, the source and destination addresses, and the protocols used for transmission.

[0053] The network slice management interface connects with the network slice management system to collect information on the resources allocated to different slices, such as the bandwidth and computing resource usage of each slice, as well as slice operation status monitoring data, such as whether it is running normally and whether congestion has occurred.

[0054] The user device interaction interface communicates with relevant applications or system modules on the user device to collect user device operation behavior data, such as the time when the user opens the application, data upload and download operations, and device settings changes.

[0055] Step S122: Configure the collection parameters of each data access interface, set the data collection frequency, data field filtering rules and data format conversion protocol, so that the collected data contains the original data fields of user identity characteristics, access behavior characteristics, signaling interaction characteristics, data transmission characteristics and network slicing characteristics.

[0056] In this embodiment, the acquisition parameters are configured to ensure that the acquired data meets the requirements for subsequent processing and analysis.

[0057] The data collection frequency is set according to the update speed and importance of different data. For example, user device operation behavior data can be collected at a higher frequency because it changes more frequently; while network slice resource allocation data can be collected at a lower frequency because it changes relatively slowly.

[0058] Data field filtering rules are used to select the necessary raw data fields and exclude irrelevant or redundant fields. For example, when collecting access data from the base station side, only fields related to user identity characteristics and access behavior characteristics, such as user identifier, access time, and access location, are retained, while some device internal status fields that are irrelevant to risk analysis are filtered out.

[0059] Data format conversion protocols are used to convert raw data from different sources into a unified intermediate format to facilitate subsequent processing. For example, they can convert binary data collected by the radio access network data interface into text format, or convert specific signaling formats collected by the core network control plane interface into a general structured format.

[0060] Step S123: Start the multi-source data access interface to collect heterogeneous data sets in the mobile communication network in real time. The heterogeneous data sets include structured data, semi-structured data and unstructured data. The structured data includes user subscription information stored in database tables, the semi-structured data includes signaling messages in JSON format, and the unstructured data includes user plane data in binary format.

[0061] In this embodiment, after the multi-source data access interface is activated, each interface begins to operate according to the configured parameters. The radio access network data interface continuously collects access data from the base station, the core network control plane interface acquires signaling data in real time, and other interfaces also collect data synchronously.

[0062] The collected heterogeneous data sets are diverse. Structured data, such as user subscription information, is stored in database tables with clearly defined fields and row structures, containing basic user information, subscription service types, permission levels, etc.

[0063] Semi-structured data, such as signaling messages in JSON format, has a certain structure, but it is not as strict as a database table. It includes the signaling type, sending and receiving nodes, timestamps, key-value pairs of each field, etc.

[0064] Unstructured data, such as user plane data in binary format, has no fixed structure and mainly consists of the actual content transmitted by the user, such as voice data, video data, and file data. These data require corresponding parsing methods to understand their content.

[0065] Step S124: Distribute the heterogeneous data set to the distributed data processing cluster, clean the heterogeneous data set using a partitioned parallel processing method, standardize the cleaned heterogeneous data, and generate standardized data records. The standardized data records include feature identifier fields, data value fields, timestamp fields, and data source fields.

[0066] In this embodiment, since heterogeneous datasets typically contain large amounts of data, using a distributed data processing cluster can improve efficiency. Data is distributed to multiple nodes in the cluster, with each node processing a portion of the data, achieving partitioned parallel processing.

[0067] Data cleaning primarily involves removing noise, missing values, and outliers from the data. For example, missing fields in user subscription information are filled by querying historical records or using default values; malformed data in signaling messages is corrected or removed; and transmission traffic data that clearly does not fall within the normal range is treated as outliers.

[0068] Standardization processing converts data of different formats and types into a unified format. In the generated standardized data records, the feature identifier field corresponds to the feature node identifier in the data security risk feature map, such as the feature identifier corresponding to the user identifier; the data value field is the specific value of the feature; the timestamp field records the time when the data was generated; and the data source field indicates which interface the data was collected from.

[0069] For example, a standardized data record about access time has the feature identifier field "ATD001", the data value field is the specific access time, the timestamp field is the timestamp corresponding to that time, and the data source field is "wireless access network data interface".

[0070] Step S125: Based on the timestamp field, perform time-series alignment of standardized data records from different sources, and aggregate the time-series aligned standardized data records according to the preset time window length to generate a multi-source heterogeneous data fusion stream containing multi-source features. Each data unit in the multi-source heterogeneous data fusion stream contains user identity feature data, access behavior feature data, signaling interaction feature data, data transmission feature data, and network slice feature data within the same time window.

[0071] In this embodiment, time-series alignment links data from different sources but belonging to the same time range. Based on the timestamp field in the standardized data records, data with similar times are grouped together, ensuring that subsequent analysis can be performed based on the same time dimension.

[0072] The preset time window length can be set according to actual needs, such as one second or one minute. The standardized data records, aligned to the time sequence, are aggregated according to the time window, that is, all data within each time window is integrated together.

[0073] In the generated multi-source heterogeneous data fusion stream, each data unit corresponds to a time window. For example, within a one-minute time window, the data unit contains user identity feature data, such as user identifier and authentication status; access behavior feature data, such as access time distribution and access location sequence; signaling interaction feature data, such as signaling type sequence and signaling field changes; data transmission feature data, such as transmission protocol and traffic fluctuations; and network slicing feature data, such as slice identifier and resource allocation status.

[0074] Step S130: Input the multi-source heterogeneous data fusion stream into the preset distributed risk feature learning network, perform feature mapping and association enhancement processing based on the data security risk feature map, and obtain the real-time risk feature vector.

[0075] In mobile communication network data security risk early warning, distributed risk feature learning networks can extract and enhance risk features from fused multi-source data to form real-time risk feature vectors, providing more representative feature representations for subsequent risk analysis.

[0076] Step S131: Initialize the distributed risk feature learning network. The distributed risk feature learning network includes a feature input layer, a graph mapping layer, an association enhancement layer, and a feature output layer. The feature input layer is used to receive multi-source heterogeneous data fusion streams. The graph mapping layer is used to realize the mapping of data features to risk feature graph nodes. The association enhancement layer is used to enhance the association relationship between features. The feature output layer is used to output real-time risk feature vectors.

[0077] In this embodiment, initializing the distributed risk feature learning network requires setting the parameters of each layer. The input dimension of the feature input layer matches the number of features in each data unit of the multi-source heterogeneous data fusion stream, ensuring that the input data can be completely received and carried. The graph mapping layer pre-sets a correspondence table between the feature node identifiers of the data security risk feature graph and the input data features to facilitate rapid mapping. The association reinforcement layer loads the association edge information in the risk feature graph, including association strength parameters and propagation direction identifiers. The output dimension of the feature output layer is determined based on the number of feature nodes in the risk feature graph and subsequent analysis requirements, and the initial weights and bias values ​​of the fully connected network are set.

[0078] During initialization, it is also necessary to set the training status of the network, such as the initial value of the learning rate and the number of iterations, to ensure that the network can perform feature processing normally after receiving input data.

[0079] Step S132: Perform feature parsing on the data units in the multi-source heterogeneous data fusion stream through the feature input layer to extract the original feature values ​​corresponding to user identity feature data, access behavior feature data, signaling interaction feature data, data transmission feature data, and network slice feature data.

[0080] In this embodiment, after receiving the multi-source heterogeneous data fusion stream, the feature input layer parses it according to the structure of the data units. Each data unit contains multiple feature data within the same time window, and the feature input layer identifies the type of these feature data one by one.

[0081] For example, for a certain data unit, the feature input layer first locates the user identity feature data part and extracts the corresponding original feature values ​​such as user identifier type, authentication credential status, and permission level label from it; then it parses the access behavior feature data part and extracts the original feature values ​​of access time distribution, access location sequence, and access device fingerprint; in the same way, it extracts the corresponding original feature values ​​from the signaling interaction feature data, data transmission feature data, and network slicing feature data parts respectively.

[0082] The extracted raw feature values ​​are temporarily stored in the buffer of the feature input layer, awaiting transmission to the graph mapping layer for further processing. During the extraction process, the feature input layer performs preliminary validation of the data format to ensure that the extracted raw feature values ​​meet the input requirements of the graph mapping layer.

[0083] Step S133: Input the extracted original feature values ​​into the graph mapping layer. Based on the feature node identifier of the data security risk feature graph, the graph mapping layer maps the original feature values ​​to the corresponding feature nodes to generate the initial node feature vector. The dimension of the initial node feature vector is consistent with the number of feature nodes in the risk feature graph, and each dimension corresponds to the original feature value of a feature node.

[0084] In this embodiment, after receiving the original feature values, the graph mapping layer maps them according to the feature node identifiers of the data security risk feature graph. The graph mapping layer stores the correspondence between the original features and the feature node identifiers. For example, the original feature value of the user identifier type corresponds to the feature node identifier "UAT001", and the original feature value of the access time distribution corresponds to the feature node identifier "ATD001", etc.

[0085] For each original feature value, the graph mapping layer finds its corresponding feature node and assigns the original feature value to the corresponding dimension of that node in the initial node feature vector. If a feature node does not have a corresponding original feature value in the current data unit, the value of that dimension will be set to a default value, such as zero or null, depending on the type of feature.

[0086] For example, assuming there are 100 feature nodes in the risk feature map, the initial node feature vector is a 100-dimensional vector. When the original feature value of the user identifier type in the user identity feature is extracted, it will be mapped to the 5th dimension (assuming "UAT001" corresponds to the 5th dimension). Then, the 5th dimension value of the initial node feature vector is the original feature value, and the other dimensions that do not have corresponding original feature values ​​are default values.

[0087] Through the above mapping process, an initial node feature vector is generated, which converts the original feature values ​​of the multi-source heterogeneous system into vector representations corresponding to the nodes of the risk feature map.

[0088] Step S134: Input the initial node feature vector into the association enhancement layer. The association enhancement layer calls the association edge information in the risk feature map. Based on the association strength parameter and the propagation direction identifier, the initial node feature vector is subjected to neighborhood feature aggregation processing. The neighborhood feature aggregation processing includes, for each feature node, aggregating the feature value of its directly associated node with the product of the association strength parameter to generate neighborhood aggregated features.

[0089] In this embodiment, after receiving the initial node feature vector, the association strengthening layer first retrieves the association edge information related to each feature node from the data security risk feature map. This information includes the direct associated nodes of each feature node, the association strength parameter, and the propagation direction identifier.

[0090] For each feature node in the initial node feature vector, the association strengthening layer filters out the directly associated nodes pointed to by the association edges originating from that node, based on the propagation direction identifier. Then, it calculates the product of the feature value of each directly associated node (i.e., the value of the corresponding dimension in the initial node feature vector) and the association strength parameter of that association edge.

[0091] Next, these products are aggregated, either by summation, to obtain the neighborhood aggregated feature of the feature node. For example, if feature node A has directly associated nodes B and C with association strength parameters S1 and S2 respectively, and the feature value of B in the initial node feature vector is Vb and the feature value of C is Vc, then the neighborhood aggregated feature of A is Vb multiplied by S1 plus Vc multiplied by S2.

[0092] For feature nodes without directly related nodes, their neighborhood aggregation feature is zero or a default value. Through the above neighborhood feature aggregation process, each feature node obtains the comprehensive influence feature of its neighboring nodes.

[0093] Step S135: The initial node feature vector is weighted and fused with the neighborhood aggregation feature to generate a reinforced association feature vector. The weight of the weighted fusion is determined based on the importance score of the feature node in the risk feature map. The importance score is obtained by calculating the degree centrality of the feature node.

[0094] In this embodiment, degree centrality is an indicator for measuring the importance of a feature node. It is determined by calculating the number of associated edges (i.e., the number of direct associations with other nodes) of the feature node in the risk feature graph. The more associated edges, the higher the degree centrality, and the higher the importance score of the node.

[0095] When calculating importance scores, the degree centrality value of each feature node is first calculated, and then these values ​​are normalized to make them fall within the same value range to obtain the importance score of each node. The sum of the importance scores of all nodes can be 1, or it can be set to other fixed values ​​according to the actual situation.

[0096] During weighted fusion, for each feature node, the feature value of that node in the initial node feature vector is multiplied by its corresponding importance score, and then the neighborhood aggregation feature of that node is multiplied by (1 minus the importance score). The two results are then added together to obtain the value of that node in the enhanced association feature vector.

[0097] For example, if the importance score of feature node D is P, the feature value of D in the initial node feature vector is Vd, and the neighborhood aggregation feature is Ad, then the dimension value of D in the reinforcement association feature vector is Vd multiplied by P plus Ad multiplied by (1 minus P).

[0098] Through the above weighted fusion, the original feature information of the feature node itself is preserved, while the association influence of its neighboring nodes is incorporated, so that the generated enhanced association feature vector can more comprehensively reflect the association relationship between features.

[0099] Step S136: Input the enhanced correlation feature vector into the feature output layer, and use the feature output layer to perform a nonlinear transformation on the enhanced correlation feature vector through a fully connected network. Then, perform dimension normalization on the transformed feature vector to generate a real-time risk feature vector with a unified dimension representation. Each dimension of the real-time risk feature vector corresponds to the enhanced feature value of a feature node in the risk feature map.

[0100] In this embodiment, the fully connected network of the feature output layer contains multiple neurons, each of which is connected to all dimensions of the reinforced associated feature vector. During the nonlinear transformation, each dimension value of the reinforced associated feature vector is multiplied by the weight matrix of the fully connected network, plus a bias term, and then nonlinearly processed by an activation function, such as using the ReLU function or the Sigmoid function, to obtain the transformed feature vector.

[0101] Nonlinear transformations can capture complex nonlinear relationships between features, enhancing their expressive power. The transformed feature vectors may differ in dimensionality from the enhanced correlation feature vectors, but their number of dimensions is fixed to allow for subsequent unified processing.

[0102] Next, the transformed feature vector is normalized. The normalization process can be done by mapping the value of each dimension to the range of 0 to 1. The specific calculation method is as follows: for each dimension, subtract the minimum value of all dimensions from the feature value of that dimension, and then divide by the difference between the maximum value and the minimum value of all dimensions (if the maximum value and the minimum value are equal, then all values ​​of that dimension are set to 0.5).

[0103] The generated real-time risk feature vector has each dimension corresponding to a feature node in the risk feature map. The dimension value is the enhanced feature value of that node after a series of processing steps. For example, the dimension corresponding to the feature node "UAT001" has a comprehensive feature value after mapping, correlation enhancement, nonlinear transformation and normalization, which can reflect the risk characteristics of that node within the current time window.

[0104] Step S140: Perform big data correlation analysis on the real-time risk feature vectors to explore the transmission dependency relationship between risk features and generate a set of risk propagation path weights.

[0105] In mobile communication network data security risk early warning, the real-time risk feature vector contains the enhanced feature values ​​of each feature node. By performing big data correlation analysis on it, the transmission path and dependency relationship between risk features can be found, and the weight of each path can be determined, providing a basis for subsequent risk level determination.

[0106] Step S141: Based on the preset risk feature threshold, filter out the abnormal risk features that exceed the threshold from the real-time risk feature vector to obtain an abnormal feature set, which includes the abnormal feature identifier and the corresponding feature value.

[0107] In this embodiment, the preset risk characteristic thresholds are determined based on historical data security incident analysis and network security policies. Different characteristic nodes may have different thresholds due to their different roles in risk propagation. These thresholds can be stored in a threshold configuration table, with each characteristic node identifier corresponding to a specific threshold.

[0108] The feature value of each dimension in the real-time risk feature vector (i.e., the enhanced feature value of each feature node) is compared with the risk feature threshold corresponding to that node. If the enhanced feature value of a feature node is greater than its corresponding threshold, then the feature is considered an abnormal risk feature.

[0109] For example, the risk feature threshold of the feature node "UAT001" (user identification type) is T1, and the feature value of this dimension in the real-time risk feature vector is V1. If V1 is greater than T1, then "UAT001" is judged as an abnormal risk feature.

[0110] Collect all identified risk characteristics and their corresponding feature values ​​to form an anomaly feature set. This set can be stored as a list, with each element containing an anomaly feature identifier and its corresponding feature value, facilitating subsequent risk propagation path discovery using it as a starting point.

[0111] Step S142: Using the set of abnormal features as the starting feature node, based on the topology of the data security risk feature map, mine the directly related feature nodes and indirectly related feature nodes of the starting feature node to generate a set of feature association paths. The set of feature association paths contains multiple path sequences, and each path sequence is composed of the starting feature node and its related feature nodes arranged in the order of transmission.

[0112] In this embodiment, the topological structure of the data security risk feature map provides a structured display of the relationships between feature nodes, based on which the associated nodes of the starting feature node can be gradually mined.

[0113] First, select each abnormal feature identifier from the abnormal feature set as the starting feature node. Then, based on the associated edge information recorded in the topology, find the feature nodes directly connected to the starting feature node, i.e., the directly associated feature nodes.

[0114] Next, starting with the directly related feature nodes, we search for their directly related nodes again based on the topology. These nodes are the indirectly related feature nodes of the starting feature nodes (related through an intermediate node). We expand step by step in this way until we reach the preset maximum association depth or there are no new related nodes.

[0115] During the mining process, the node sequence needs to be recorded according to the transmission order to form a path sequence. For example, if the direct associated node of the starting feature node A is B, and the direct associated node of B is C, then the path sequence can be [A, B, C], representing the associated path from A through B to C.

[0116] All mined path sequences are collected, duplicate paths are removed, and a set of feature-related paths is formed. Each path sequence clearly shows the propagation path from the starting feature node to other related nodes.

[0117] Step S1421: Initialize the breadth-first search queue, add each abnormal feature identifier in the abnormal feature set as the starting node to the search queue, set the current path depth of each starting node to 1, and the current path sequence is a single-node sequence containing only the starting node.

[0118] In this embodiment, a breadth-first search queue is used to mine associated nodes in hierarchical order. During initialization, an empty queue is first created, and then each anomaly feature identifier in the anomaly feature set is traversed.

[0119] For each anomaly identifier, it is added to the queue as a starting node. Simultaneously, the current path depth is recorded for each starting node; since the path currently contains only that starting node, the path depth is 1. The current path sequence also only contains that starting node, such as [starting node A], [starting node B], etc.

[0120] The queue can be stored as a list, with each element containing the starting node identifier, the current path sequence, and the current path depth, for processing in subsequent search operations.

[0121] Step S1422: Remove the head node from the search queue and query the adjacent nodes in the data security risk feature map that have a direct relationship with the head node. The adjacent nodes are determined based on the propagation direction of the associated edge and only include the nodes pointed to by the associated edge originating from the current node.

[0122] In this embodiment, the head node retrieved from the search queue includes its identifier, current path sequence, and path depth. Based on the node's identifier, the associated edge information is searched in the data security risk feature graph.

[0123] Based on the propagation direction of the associated edges, the nodes pointed to by the associated edges originating from the current node are selected. These nodes are the adjacent nodes of the current node. For example, if the current node is D, and the associated edge information shows that the associated edges originating from D point to E and F, then E and F are the adjacent nodes of D.

[0124] During the query process, it is necessary to accurately match the node identifier and the propagation direction of the associated edges to ensure the accuracy of adjacent nodes and avoid including nodes with reverse associations.

[0125] Step S1423: For each adjacent node, determine whether the node has already appeared in the current path sequence. If it has not appeared, generate a new path sequence. The new path sequence adds the adjacent node to the end of the current path sequence, and the new path depth is the current path depth plus 1.

[0126] In this embodiment, for each neighboring node found, it is necessary to check whether it already exists in the current path sequence. This is to avoid circular paths and ensure the validity of the path sequence.

[0127] If an adjacent node is not in the current path sequence, the adjacent node is added to the end of the current path sequence to form a new path sequence. Simultaneously, since a node has been added to the path, the new path depth is the current path depth plus 1.

[0128] For example, if the current path sequence is [A, B], the path depth is 2, and the adjacent node is C which is not in [A, B], then the new path sequence is [A, B, C], and the new path depth is 3.

[0129] Step S1424: Determine whether the new path depth exceeds the preset maximum path depth threshold. If it does not exceed the threshold, add the new path sequence to the search queue and record the starting node, node sequence, and path depth of the path sequence.

[0130] In this embodiment, the preset maximum path depth threshold is used to control the scope of path mining, preventing excessively long paths from causing excessive computation or rendering the paths meaningless. This maximum path depth threshold can be set according to the complexity of the data security risk feature map and the actual application requirements.

[0131] After a new path sequence is generated, its path depth is compared with the maximum path depth threshold. If the new path depth is less than or equal to the maximum path depth threshold, the new path sequence, its starting node, path depth, and other information are added to the search queue, awaiting further mining of its adjacent nodes.

[0132] If a new path depth exceeds the maximum path depth threshold, it will no longer be added to the search queue, and the mining process for that path sequence will end there.

[0133] Step S1425: Repeat the steps of retrieving the head node from the search queue, querying adjacent nodes, and generating a new path sequence until the search queue is empty or the path depth of all path sequences reaches the maximum path depth threshold.

[0134] In this embodiment, the repetitive process is the core of breadth-first search. It gradually expands the path sequence by continuously retrieving nodes from the queue for processing. Each time the head node is retrieved from the search queue, steps S1422 to S1424 are followed to generate a new path sequence and determine whether to add it to the search queue. When there are no more nodes in the search queue, it means that all mineable path sequences have been generated; or when the path depth of all path sequences reaches the maximum path depth threshold, continuing to mine will not generate any new valid paths, and the mining process stops at this point.

[0135] Step S1426: Collect all generated path sequences, remove duplicate path sequences, and obtain a set of feature-related paths containing the starting feature node and its directly and indirectly related feature nodes. Each path sequence in the set of feature-related paths has a unique path identifier, which is generated by combining the starting node identifier, path depth, and node sequence hash value.

[0136] In this embodiment, all path sequences generated during the mining process are collected. These sequences may be duplicated, meaning that different mining processes may generate the exact same node sequence.

[0137] To ensure the uniqueness and validity of the feature association path set, the collected path sequences need to be deduplicated. This can be done by comparing the node order and node identifiers within the path sequences; only one identical sequence is retained.

[0138] After deduplication, a unique path identifier is generated for each path sequence. This can be achieved by combining the starting node identifier, the path depth, and the hash value of the node sequence. The hash value is a unique string obtained by hashing the node sequence, ensuring that different node sequences will have different path identifiers even if the starting node and path depth are the same.

[0139] Step S143: For each path sequence in the feature association path set, extract the association strength parameter between adjacent feature nodes in the path, and calculate the path propagation probability based on the association strength parameter. The path propagation probability is equal to the product of the association strength parameters of all adjacent nodes in the path.

[0140] In this embodiment, each path sequence in the feature association path set consists of multiple feature nodes arranged in the transmission order. There is an association edge between two adjacent nodes, and the association edge is configured with an association strength parameter.

[0141] For each path sequence, query the association strength parameter between adjacent feature nodes from the data security risk feature map. For example, in the path sequence [A, B, C], the association strength parameter between A and B is S1, and the association strength parameter between B and C is S2.

[0142] The path propagation probability is calculated by multiplying the association strength parameters between all adjacent nodes in the path. For the path sequence [A, B, C], its path propagation probability is the product of S1 and S2. In this embodiment, since the association strength parameters have been normalized in the previous stage to ensure that they are within the same dimension range, the result of the multiplication still has reasonable physical meaning and can reflect the overall propagation probability of the path.

[0143] During the calculation, it is necessary to extract the association strength parameters of adjacent nodes in each path sequence one by one to ensure that no parameters of any associated edge are missed. For example, for a path sequence [D, E, F, G] containing four nodes, it is necessary to extract the association strength parameter S3 between D and E, the association strength parameter S4 between E and F, and the association strength parameter S5 between F and G, and then calculate the path propagation probability as S3 multiplied by S4 multiplied by S5.

[0144] By using the above method, each path sequence is assigned a path propagation probability. The magnitude of this path propagation probability can, to a certain extent, reflect the likelihood of risk characteristics being propagated along that path.

[0145] Step S144: Collect data on the actual impact of risk propagation paths in historical data security incidents, and construct a path impact assessment model. The path impact assessment model takes the path propagation probability, path length, and number of abnormal features contained in the path as inputs, and outputs the path impact weights.

[0146] In this embodiment, the data collected regarding historical data security incidents needs to cover events of different types and severity. This data can be obtained from network security incident logs, records from security management platforms, and other sources.

[0147] For each historical data security incident, it is necessary to extract data on the actual impact of its risk propagation path. The actual impact can be described through multiple dimensions, such as the duration of network service interruption caused by the incident, the range of affected users, the scale of data breach, and the estimated economic losses.

[0148] Before constructing a path impact assessment model, the collected data needs to be preprocessed. Relevant parameters of the risk propagation path, such as path transmission probability, path length (i.e., the number of feature nodes in the path), and the number of abnormal features in the path, should be organized. Simultaneously, the actual impact data should be quantified and converted into a numerical form suitable for model training.

[0149] The path impact assessment model is constructed using a machine learning model, such as a gradient boosting decision tree model. The input layer of this gradient boosting decision tree model receives three parameters: path propagation probability, path length, and the number of anomalous features contained in the path. These input parameters require feature scaling to suit the model's training requirements.

[0150] The model's output layer outputs path influence weights, which are parameters that comprehensively reflect the degree of influence of a path on risk diffusion. During model training, input parameters from historical data are used as features of training samples, and the corresponding quantified values ​​of actual influence are used as sample labels. By adjusting the model's internal parameters, the model can learn the mapping relationship between input parameters and output path influence weights.

[0151] For example, for a certain historical risk propagation path, its path transmission probability is 0.6, the path length is 3, the number of abnormal features it contains is 2, and the corresponding actual impact quantification value is 0.7. Then, during training, these three parameters are used as inputs, and 0.7 is used as the expected output to train the model.

[0152] By training with a large amount of historical data, the path impact assessment model can accurately output reasonable path impact weights based on the input path parameters.

[0153] Step S145: Input each path sequence and its corresponding path propagation probability from the feature-associated path set into the path influence assessment model to calculate the path influence weight of each path sequence.

[0154] In this embodiment, after obtaining the path propagation probability of each path sequence in the feature-associated path set, it is also necessary to determine the path length and the number of abnormal features contained in each path sequence.

[0155] The path length can be obtained by counting the number of feature nodes in each path sequence. For example, the path length of the path sequence [A, B, C] is 3. The number of anomalous features contained in the path is the number of feature nodes in the path sequence that belong to the set of anomalous features. For example, if A and C are anomalous features in the path sequence [A, B, C], then the number of anomalous features contained in the path is 2.

[0156] The path propagation probability, path length, and number of abnormal features contained in each path sequence are organized according to the format required by the model and then input into the pre-constructed path impact assessment model.

[0157] The model processes these input parameters through an internal computational mechanism. For example, a gradient boosting decision tree model processes the input features layer by layer through a combination of multiple decision trees, and finally outputs a path influence weight value.

[0158] For example, a path sequence has a path propagation probability of 0.5, a path length of 4, and contains 1 outlier. When these parameters are input into the model, the path influence weight output by the model might be 0.3. Another path sequence has a path propagation probability of 0.8, a path length of 2, and contains 3 outliers. When input into the model, the path influence weight might be 0.9.

[0159] Using the above method, a corresponding path influence weight is calculated for each path sequence in the feature association path set. The larger the weight value, the greater the influence of the path in the risk propagation process.

[0160] Step S146: Sort the set of feature-related paths based on path influence weights, and select path sequences with path influence weights greater than a preset weight threshold as key risk propagation paths.

[0161] In this embodiment, the preset weight threshold is determined based on the experience of handling historical data security incidents and risk control requirements. Its function is to distinguish the paths that have a significant impact on the spread of risk.

[0162] First, all path sequences in the feature association path set are sorted in descending order of their path influence weights. After sorting, the path influence weight of each path sequence is compared with a preset weight threshold.

[0163] For example, with a preset weight threshold of 0.5, path sequences with a path influence weight greater than 0.5, such as 0.6, 0.7, and 0.8, are selected and identified as key risk propagation paths; while path sequences with a path influence weight less than 0.5, such as 0.4 and 0.3, are not considered as key risk propagation paths.

[0164] By employing the above screening methods, we can focus on those paths that have a significant impact on risk spread, reducing the workload of subsequent analysis and processing, and making risk warnings more targeted. The key risk propagation paths identified will serve as an important basis for subsequently determining the risk spread level and key impact nodes.

[0165] Step S147: For each key risk propagation path, integrate the feature node identifiers, path transmission probabilities, and path influence weights in the path sequence to generate a risk propagation path weight set containing path identifiers, feature node sequences, transmission probability sequences, and influence weights.

[0166] In this embodiment, for each key risk propagation path, its relevant information needs to be integrated. The path identifier is a unique identifier assigned to each path sequence when generating the feature-associated path set, such as "P001", "P002", etc.

[0167] A feature node sequence is a sequence of feature nodes in a path arranged in the order of propagation, such as [A, B, C, D]. A propagation probability sequence is a sequence of association strength parameters between adjacent feature nodes in the path, arranged in order, corresponding to the feature node sequence. For example, for the feature node sequence [A, B, C, D], the propagation probability sequence is [S1, S2, S3], where S1 is the association strength parameter between A and B, S2 is the association strength parameter between B and C, and S3 is the association strength parameter between C and D.

[0168] The impact weight is the path impact weight of the key risk propagation path calculated through the path impact assessment model.

[0169] This information is integrated to form a risk propagation path weight set. For example, for a critical risk propagation path identified as "P001", its characteristic node sequence is [A, B, C], its transmission probability sequence is [S1, S2], and its influence weight is 0.7. Then, the record of this path in the risk propagation path weight set contains this information.

[0170] The risk propagation path weight set fully records the key information of all critical risk propagation paths.

[0171] Step S150: Determine the risk diffusion level and key impact nodes based on the risk propagation path weight set, generate a security risk warning instruction containing the risk propagation path identifier, and push the security risk warning instruction to the mobile communication security management platform.

[0172] In mobile communication network data security risk early warning, the severity and key nodes of the current risk can be assessed based on the information in the risk propagation path weight set, and then corresponding early warning instructions can be generated so that the security management platform can take timely measures to deal with it.

[0173] Step S151: Analyze the risk propagation path weight set and extract the path influence weight, path transmission probability and feature node sequence of each key risk propagation path.

[0174] In this embodiment, parsing the risk propagation path weight set requires extracting relevant information for each key risk propagation path one by one according to the data structure of the set.

[0175] For each key risk propagation path, firstly, its path influence weight is extracted, which reflects the degree of influence of the path on risk diffusion; then, the path transmission probability is extracted, which reflects the possibility of risk characteristics being transmitted along the path; finally, the feature node sequence is extracted, which shows the transmission order of risk characteristics along the path and the feature nodes involved.

[0176] For example, analyzing a key risk propagation path reveals that its path influence weight is 0.8, its path propagation probability is 0.7, and its characteristic node sequence is [abnormal user authentication credentials, unfamiliar access device fingerprint, abnormal transmission traffic fluctuations].

[0177] By using the above analytical method, the information in the risk propagation path weight set is decomposed into specific parameters that can be directly used for calculation and analysis, thus preparing for the subsequent calculation of the comprehensive risk diffusion index and the analysis of node importance.

[0178] Step S152: Calculate the risk diffusion comprehensive index. The risk diffusion comprehensive index is the weighted sum of the path influence weights and path transmission probabilities of all key risk propagation paths. The weights are determined according to the path length, and the shorter the path length, the greater the weight.

[0179] In this embodiment, calculating the comprehensive risk diffusion index requires considering the impact of all key risk propagation paths. First, a weighted weight is determined for each key risk propagation path, based on the path length.

[0180] Path length refers to the number of feature nodes in a path. The shorter the path length, the faster the risk feature may be transmitted along the path, and the more direct its impact on risk diffusion may be. Therefore, it is given a larger weighting. Conversely, the longer the path length, the smaller the weighting.

[0181] For example, a critical risk transmission path with a path length of 2 can have a weighted weight of 0.4; a path with a path length of 3 can have a weighted weight of 0.3; a path with a path length of 4 can have a weighted weight of 0.2, and so on. The specific weighted weight values ​​can be determined based on the actual situation and historical data, and the sum of the weighted weights of all paths is 1.

[0182] Next, for each key risk propagation path, calculate the product of its path influence weight and path transmission probability, and then multiply this product by the weighted weight of the path to obtain the path's contribution value in the composite index.

[0183] Finally, the contribution values ​​of all key risk propagation paths are summed to obtain the comprehensive risk diffusion index. For example, if there are three key risk propagation paths, the first path has a path influence weight of 0.8, a path transmission probability of 0.7, and a weighted average weight of 0.4, with a contribution value of 0.8 multiplied by 0.7 multiplied by 0.4; the second path has a path influence weight of 0.6, a path transmission probability of 0.6, and a weighted average weight of 0.3, with a contribution value of 0.6 multiplied by 0.6 multiplied by 0.3; and the third path has a path influence weight of 0.5, a path transmission probability of 0.5, and a weighted average weight of 0.3, with a contribution value of 0.5 multiplied by 0.5 multiplied by 0.3. Summing these three contribution values ​​yields the comprehensive risk diffusion index.

[0184] The comprehensive risk diffusion index can comprehensively reflect the overall degree of risk diffusion brought about by all current key risk propagation paths, and is an important basis for determining the risk diffusion level.

[0185] Step S153: Based on the risk diffusion comprehensive index, query the preset risk level classification rules to determine the risk diffusion level. The risk level classification rules include multiple index ranges and corresponding risk levels.

[0186] In this embodiment, the preset risk level classification rule is formulated based on the security management strategy of mobile communication networks and the experience in handling historical risk events. This risk level classification rule divides the comprehensive risk diffusion index into multiple consecutive intervals, each interval corresponding to a specific risk level.

[0187] For example, a risk diffusion index between 0 and 0.3 corresponds to a low risk level; between 0.3 and 0.6 corresponds to a medium risk level; and between 0.6 and 1.0 corresponds to a high risk level. Different risk levels correspond to different response measures and handling priorities.

[0188] Based on the calculated comprehensive risk diffusion index, a query is performed within the preset risk level classification rules to find the interval to which the index belongs, and then the corresponding risk diffusion level is determined.

[0189] For example, the calculated risk diffusion composite index is 0.7. After consulting the risk level classification rules, it is found that the index is in the range of 0.6 to 1.0. Therefore, the current risk diffusion level is determined to be high.

[0190] Once the risk diffusion level is determined, it can intuitively reflect the severity of the current risk.

[0191] Step S1531: Obtain the preset risk level classification rules. The risk level classification rules are formulated by the mobile communication security management department based on the impact assessment results of historical data security incidents. They include multiple continuous intervals of the comprehensive risk diffusion index and the risk level name, risk handling priority, and response time limit requirements for each interval.

[0192] In this embodiment, the preset risk level classification rules can be obtained by accessing a database or configuration file storing these rules. These risk level classification rules are formulated by the mobile communication security management department and are authoritative and practical.

[0193] The risk level classification rules include multiple consecutive intervals covering the possible range of values ​​for the comprehensive risk diffusion index. Each interval can be classified as low, medium, high, or severe.

[0194] Risk handling priorities correspond to risk level names; the higher the level, the higher the handling priority. For example, low risk level corresponds to low handling priority, medium risk level corresponds to medium handling priority, high risk level corresponds to high handling priority, and severe risk level corresponds to the highest handling priority.

[0195] The response time limit requirements specify the timeframe within which the safety management department must respond and take initial measures after receiving an early warning instruction. For example, the response time limit is 24 hours for low-risk levels, 12 hours for medium-risk levels, 4 hours for high-risk levels, and 1 hour for severe-risk levels.

[0196] Step S1532: Extract the specific value of the comprehensive risk diffusion index, compare the specific value with the index range in the risk level classification rules, and determine the target range to which the specific value belongs.

[0197] In this embodiment, after extracting the specific value of the comprehensive risk diffusion index, it is compared with the index range in the risk level classification rules in ascending order.

[0198] For example, the specific value of the comprehensive risk diffusion index is 0.5, and the index range in the risk level classification rules is 0-0.3, 0.3-0.6, and 0.6-1.0. Comparing 0.5 with these ranges, 0.5 is greater than 0.3 and less than 0.6, therefore its target range is determined to be 0.3-0.6.

[0199] During the comparison process, attention should be paid to the handling of boundary values ​​within the intervals. For example, when a specific value equals the upper limit of the interval, it should be classified into a higher-level interval. For instance, if the specific value is 0.6, it should be classified into the interval 0.6-1.0.

[0200] Step S1533: Query the risk level name, risk handling priority and response time limit requirements corresponding to the target interval, and determine the risk level name as the current risk diffusion level.

[0201] In this embodiment, after determining the target range, the risk level name, risk handling priority, and response time limit requirements corresponding to the target range are queried according to the risk level classification rules.

[0202] For example, if the target range is 0.3-0.6, the corresponding risk level name obtained after querying the rules is medium, the risk handling priority is medium, and the response time limit requirement is within 12 hours.

[0203] The retrieved risk level name is designated as the current risk spread level, which is currently medium. Simultaneously, the corresponding risk handling priority and response time limit are recorded. This information will be included in the security risk warning instruction to guide the security management department's handling work.

[0204] Step S1534: Generate risk level description information, which includes the risk diffusion level name, risk handling priority, response time limit requirements, and a description of the typical impact scenarios corresponding to the risk level.

[0205] In this embodiment, risk level description information is generated to help the staff of the security management platform to more intuitively understand the meaning of the current risk level and the response requirements.

[0206] The risk diffusion level name, risk handling priority, and response time limit requirements in the risk level description information are consistent with the information found in the risk level classification rules.

[0207] The typical impact scenarios corresponding to this risk level are summarized based on historical data security incidents, describing the typical impacts that may occur under this risk level. For example, the typical impact scenario for the medium risk level might be: network services for some users in certain areas experience brief interruptions, data transmission for a small number of users is affected, which may cause some user complaints, but will not pose a serious threat to the overall operation of the network.

[0208] Step S1535: Add the risk level description information to the basic information field of the security risk warning instruction to help the administrators of the mobile communication security management platform understand the meaning of the risk level and the handling requirements.

[0209] In this embodiment, the basic information field of the security risk warning instruction is an important component of the instruction, containing core basic information related to risk warning.

[0210] Adding the generated risk level description information to this field allows administrators of the mobile communication security management platform to directly obtain a detailed explanation of the risk level when viewing warning instructions.

[0211] For example, the basic information field includes not only the warning timestamp and risk spread level, but also a risk level description: "The current risk spread level is medium, the risk handling priority is medium, and the response time limit is within 12 hours. The typical impact scenario is that network services for some users in some areas will be briefly interrupted, and data transmission for a small number of users will be affected."

[0212] As a result, managers can quickly understand the current risk situation and the requirements for handling it, thereby improving response efficiency.

[0213] Step S154: Perform node importance analysis on the feature node sequence of each key risk propagation path, and calculate the intermediary centrality of each feature node. The intermediary centrality is the frequency with which the node appears in all key risk propagation paths. The higher the frequency, the greater the intermediary centrality.

[0214] In this embodiment, the importance of each key risk propagation path is analyzed for the characteristic node sequence, aiming to identify the key characteristic nodes that play a crucial role in the risk propagation process.

[0215] Calculating the betweenness centrality of each feature node requires counting the number of times that node appears in all key risk propagation paths. First, iterate through the sequence of feature nodes in all key risk propagation paths and record the frequency of each feature node.

[0216] For example, suppose there are three key risk propagation paths with characteristic node sequences of [A, B, C], [B, D, E], and [A, B, E], respectively. Then the frequency of characteristic node A is 2, the frequency of characteristic node B is 3, the frequency of characteristic node C is 1, the frequency of characteristic node D is 1, and the frequency of characteristic node E is 2.

[0217] Next, the total number of all key risk propagation paths is calculated; in this example, the total number is 3. Then, the frequency of each feature node is divided by the total number to obtain the betweenness centrality of each feature node. For example, the betweenness centrality of feature node A is 2 divided by 3, the betweenness centrality of feature node B is 3 divided by 3, the betweenness centrality of feature node C is 1 divided by 3, the betweenness centrality of feature node D is 1 divided by 3, and the betweenness centrality of feature node E is 2 divided by 3.

[0218] The above calculation method can quantify the importance of each feature node in the risk propagation path. The greater the intermediary centrality, the more significant the bridging and hub role the node plays in the risk propagation process, and the more likely it is to be a key node in risk propagation.

[0219] For example, step S1541: Collect the feature node sequences of all key risk propagation paths in the risk propagation path weight set, and construct a node-path association matrix. The rows of the node-path association matrix represent feature nodes, and the columns represent key risk propagation paths. The matrix element value of 1 indicates that the feature node appears in the corresponding key risk propagation path, and 0 indicates that it does not appear.

[0220] In this embodiment, after collecting the feature node sequences of all key risk propagation paths, a node-path association matrix needs to be constructed. First, all the feature nodes that have appeared are listed as rows of the matrix, and each key risk propagation path is listed as a column of the matrix.

[0221] For example, given feature nodes A, B, C, D, and E, and three key risk propagation paths P1, P2, and P3, with feature node sequences [A, B, C], [B, D, E], and [A, B, E] respectively, the constructed node-path association matrix will have rows A, B, C, D, and E, and columns P1, P2, and P3 respectively.

[0222] For feature node A, it appears in P1 and P3, so the element value in columns P1 and P3 is 1, and the element value in column P2 is 0; feature node B appears in P1, P2, and P3, so the element value in all three columns is 1; feature node C appears only in P1, so the element value in column P1 is 1, and the element value in other columns is 0; feature node D appears only in P2, so the element value in column P2 is 1, and the element value in other columns is 0; feature node E appears in P2 and P3, so the element value in columns P2 and P3 is 1, and the element value in column P1 is 0.

[0223] The node-path association matrix above clearly shows the association between each feature node and each key risk propagation path.

[0224] Step S1542: Calculate the sum of the element values ​​of the corresponding row in the node-path association matrix for each feature node to obtain the path occurrence frequency of the feature node. The path occurrence frequency is equal to the total number of times the feature node appears in the key risk propagation path.

[0225] In this embodiment, after constructing the node-path association matrix, for each feature node, the values ​​of all elements in the corresponding row are added together, and the sum is the frequency of the path occurrence of that feature node.

[0226] Taking the example in step S1541, the row element values ​​corresponding to feature node A are 1, 0, and 1, and the sum of the element values ​​is 2, meaning the path occurrence frequency is 2; the row element values ​​corresponding to feature node B are 1, 1, and 1, and the sum of the element values ​​is 3, meaning the path occurrence frequency is 3; the row element values ​​corresponding to feature node C are 1, 0, and 0, and the sum of the element values ​​is 1, meaning the path occurrence frequency is 1; the row element values ​​corresponding to feature node D are 0, 1, and 0, and the sum of the element values ​​is 1, meaning the path occurrence frequency is 1; the row element values ​​corresponding to feature node E are 0, 1, and 1, and the sum of the element values ​​is 2, meaning the path occurrence frequency is 2.

[0227] The above statistical methods can accurately determine the total number of times each feature node appears in the key risk propagation path.

[0228] Step S1543: Calculate the total number of all critical risk propagation paths, divide the frequency of occurrence of the path of the feature node by the total number, and obtain the betweenness centrality of the feature node. The value of betweenness centrality ranges from 0 to 1. The larger the value, the higher the criticality of the node in risk propagation.

[0229] In this embodiment, the total number of all key risk propagation paths is first determined. This number can be obtained by counting the number of key risk propagation paths in the risk propagation path weight set.

[0230] For example, if there are 5 key risk propagation paths in the risk propagation path weight set, then the total number is 5. For each feature node, the betweenness centrality can be obtained by dividing the frequency of its path occurrence by this total number.

[0231] Suppose a feature node appears 3 times in a path, and the total number of paths is 5. Then the betweenness centrality of this feature node is 3 divided by 5. Since the frequency of a path occurrence is at most equal to the total number of paths, the maximum value of betweenness centrality is 1, which is obtained when the feature node appears in all key risk propagation paths; the minimum value is 0, which is obtained when the feature node does not appear in any key risk propagation path.

[0232] The above calculations can standardize the frequency of path occurrence of different characteristic nodes, resulting in a unified range of indicators, which facilitates the comparison of the criticality of different characteristic nodes in risk propagation.

[0233] Step S1544: Sort the feature nodes in descending order of their intermediary centrality to generate a node importance ranking list.

[0234] In this embodiment, after obtaining the betweenness centrality of each feature node, the feature nodes are sorted from largest to smallest according to the betweenness centrality value.

[0235] For example, if there are feature nodes M, N, O, and P, and their betweenness centralities are 0.8, 0.6, 0.9, and 0.5 respectively, then the sorted list of node importance is O (0.9), M (0.8), N (0.6), and P (0.5).

[0236] Generating the above-mentioned list of node importance rankings can intuitively show the order of importance of each feature node, enabling staff to quickly understand which feature nodes are more critical in risk communication.

[0237] Step S1545: Based on the preset centrality threshold, select feature nodes with a middleness centrality greater than the threshold from the node importance ranking list as key influential nodes. The centrality threshold is determined based on the average middleness centrality of key nodes in historical data security events.

[0238] In this embodiment, the preset centrality threshold needs to be determined in conjunction with historical data security events. By analyzing the betweenness centrality of feature nodes identified as key nodes in historical data security events, the average value is calculated and used as a reference for the centrality threshold.

[0239] For example, in historical data security incidents, if the average betweenness centrality of key nodes is 0.6, then the centrality threshold can be set to 0.6. Then, feature nodes with a betweenness centrality greater than 0.6 are selected from the node importance ranking list as key influencing nodes.

[0240] Assuming the intermediary centralities of the feature nodes in the node importance ranking list are 0.9, 0.8, 0.7, 0.5, and 0.4 respectively, and the centrality threshold is 0.6, then the feature nodes with intermediary centralities of 0.9, 0.8, and 0.7 will be selected as key influencing nodes.

[0241] The key impact nodes selected through the above methods can accurately reflect the nodes that play a crucial role in the spread of risk in the current risk scenario.

[0242] Step S155: Select feature nodes with a middleness centrality greater than a preset centrality threshold as key influencing nodes, and extract the feature identifiers and corresponding feature values ​​of the key influencing nodes.

[0243] In this embodiment, after identifying the key impact nodes, it is necessary to extract the feature identifiers and corresponding feature values ​​of these nodes. The feature identifier is a unique identifier assigned to each feature node when constructing the data security risk feature map, such as "UAT001" and "ATD001" in the previous example.

[0244] The eigenvalues ​​are the eigenvalues ​​of key influencing nodes in the real-time risk eigenvector, and these eigenvalues ​​reflect the specific state of the node within the current time window.

[0245] For example, the key impact node is the abnormal fingerprint node of the access device, whose feature identifier is "ADF002". The corresponding feature value may be the matching degree value between the device fingerprint and the fingerprint of the user's commonly used devices.

[0246] Step S156: Generate the basic information fields of the security risk warning instruction. The basic information fields include the warning timestamp, risk spread level, list of key impact nodes, and comprehensive risk spread index.

[0247] In this embodiment, the warning timestamp is the time when the security risk warning instruction is generated, accurate to the millisecond level, so as to accurately record the moment when the warning occurs.

[0248] The risk diffusion level is determined based on the comprehensive risk diffusion index, such as "high risk", "medium risk" and "low risk" as determined in the previous steps.

[0249] The list of key impact nodes contains the feature identifiers and corresponding feature values ​​of all selected key impact nodes, arranged in descending order of betweenness centrality.

[0250] The risk diffusion composite index is a value obtained by weighting and summing the path influence weights and path transmission probabilities of all key risk propagation paths, and reflects the overall degree of current risk diffusion.

[0251] This information is integrated into the basic information fields to form the core content framework of the security risk warning instruction, providing the recipient with overall information about the risk.

[0252] Step S157: Generate a path description field for each key risk propagation path. The path description field includes path identifier, feature node sequence, propagation probability sequence, and path impact weight.

[0253] In this embodiment, the path identifier is a unique identifier assigned to each critical risk propagation path, such as "P001" or "P002", to distinguish different paths.

[0254] The feature node sequence is a sequence of feature nodes included in the critical risk propagation path arranged in the order of transmission, such as [A, B, C, D].

[0255] A transmission probability sequence is a sequence of transmission probabilities between adjacent feature nodes in a path, arranged in order. It corresponds to the feature node sequence. For example, the transmission probability sequence corresponding to the feature node sequence [A, B, C, D] is [S1, S2, S3], where S1 is the transmission probability from A to B, S2 is the transmission probability from B to C, and S3 is the transmission probability from C to D.

[0256] The path impact weight is the impact weight value of the path calculated by the path impact assessment model, which reflects the degree of impact of the path in risk propagation.

[0257] The above-mentioned path description field is generated for each key risk propagation path, which can show the specific details of each path and help the recipient understand the propagation path and characteristics of the risk.

[0258] Step S158: Integrate the basic information field and the path description field, and encapsulate them into a security risk warning instruction according to the preset instruction format. The instruction format includes an instruction header, an instruction body, and an instruction tail. The instruction header includes an instruction type identifier and a version number, the instruction body includes the basic information field and the path description field, and the instruction tail includes a checksum.

[0259] In this embodiment, when integrating the basic information field and the path description field, they need to be organized according to a preset instruction format. The instruction type identifier in the instruction header is used to indicate that the instruction is a security risk warning instruction, such as "SRW001"; the version number is used to identify the version of the instruction format, which is convenient for the receiver to parse, such as "V1.0".

[0260] The instruction body is the core of the instruction, containing basic information fields and path description fields for all key risk propagation paths. Each field is clearly separated by delimiters such as commas and semicolons to ensure correct parsing by the receiver. The checksum at the end of the instruction is a value calculated from the contents of the instruction header and body. Upon receiving the instruction, the receiver calculates the checksum using the same algorithm and compares it with the checksum at the end of the instruction to verify whether the instruction has been erroneous or tampered with during transmission. This encapsulation method ensures the integrity, accuracy, and parsability of security risk warning instructions, facilitating correct reception and processing by the mobile communication security management platform.

[0261] Step S159: Push the security risk warning instruction to the mobile communication security management platform.

[0262] In this embodiment, pushing security risk warning commands can be achieved through various communication methods, such as communication connections based on Transmission Control Protocol (TCP) and message queues. Before pushing, encrypted transmission can be used, such as encrypting the command using Secure Sockets Layer (SSL) protocol, to prevent the command from being stolen or tampered with during transmission. During the push process, the time and status of the push need to be recorded, such as push success or push failure. If the push fails, a retry mechanism is required to ensure that the command is ultimately delivered to the mobile communication security management platform. After receiving the security risk warning command, the mobile communication security management platform will process and respond accordingly based on the information in the command, such as notifying relevant personnel and taking risk prevention and control measures, thereby achieving timely warning and handling of mobile communication network data security risks.

[0263] Figure 2 The illustration shows exemplary hardware and software components of a data security risk early warning system 100 based on big data analytics, which can implement the ideas of this application, according to some embodiments of this application. For example, processor 120 can be used in the data security risk early warning system 100 based on big data analytics and to perform the functions in this application.

[0264] For example, a data security risk early warning system 100 based on big data analytics may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the data security risk early warning system 100 based on big data analytics may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The data security risk early warning system 100 based on big data analytics also includes an I / O interface 150 between the computer and other input / output devices.

[0265] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the data security risk warning method based on big data analysis as described above is implemented.

[0266] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A data security risk early warning method based on big data analysis, characterized in that, The method includes: Construct a data security risk feature map, which includes a network of relationships between user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features; Heterogeneous data sets in mobile communication networks are collected in real time through multi-source data access interfaces. The heterogeneous data sets are then subjected to distributed cleaning and standardization processing to generate a multi-source heterogeneous data fusion stream. The multi-source heterogeneous data fusion stream is input into a preset distributed risk feature learning network, and feature mapping and association enhancement processing are performed based on the data security risk feature map to obtain a real-time risk feature vector. Big data correlation analysis is performed on the real-time risk feature vectors to uncover the transmission dependencies between risk features and generate a set of risk propagation path weights. Based on the risk propagation path weight set, the risk diffusion level and key impact nodes are determined, a security risk warning instruction containing the risk propagation path identifier is generated, and the security risk warning instruction is pushed to the mobile communication security management platform. The process of inputting the multi-source heterogeneous data fusion stream into a preset distributed risk feature learning network, and performing feature mapping and association enhancement processing based on the data security risk feature map to obtain a real-time risk feature vector includes: Initialize a distributed risk feature learning network, which includes a feature input layer, a graph mapping layer, an association enhancement layer, and a feature output layer. The feature input layer performs feature parsing on the data units in the multi-source heterogeneous data fusion stream to extract the original feature values ​​corresponding to user identity feature data, access behavior feature data, signaling interaction feature data, data transmission feature data, and network slice feature data. The extracted raw feature values ​​are input into the graph mapping layer. Based on the feature node identifiers of the data security risk feature graph, the graph mapping layer maps the raw feature values ​​to the corresponding feature nodes to generate an initial node feature vector. The dimension of the initial node feature vector is consistent with the number of feature nodes in the risk feature graph, and each dimension corresponds to the raw feature value of a feature node. The initial node feature vector is input into the association enhancement layer. The association enhancement layer calls the association edge information in the risk feature map. Based on the association strength parameter and the propagation direction identifier, the initial node feature vector is subjected to neighborhood feature aggregation processing. The neighborhood feature aggregation processing includes, for each feature node, aggregating the feature value of its directly associated node and the product of the association strength parameter to generate neighborhood aggregated features. The initial node feature vector is weighted and fused with the neighborhood aggregation feature to generate a reinforced association feature vector. The weights of the weighted fusion are determined based on the importance score of the feature node in the risk feature map. The importance score is obtained by calculating the degree centrality of the feature node. The enhanced correlation feature vector is input into the feature output layer. The feature output layer performs a nonlinear transformation on the enhanced correlation feature vector through a fully connected network. The transformed feature vector is then normalized in dimension to generate a real-time risk feature vector with a unified dimension representation. Each dimension of the real-time risk feature vector corresponds to the enhanced feature value of a feature node in the risk feature map.

2. The data security risk early warning method based on big data analysis according to claim 1, characterized in that, The construction of the data security risk feature map includes: To identify the core characteristic dimensions of data security risks in mobile communication networks, a basic feature set is determined. This basic feature set includes user identity features, access behavior features, signaling interaction features, data transmission features, and network slicing features. The user identity features include user identifier type, authentication credential status, and permission level label. The access behavior features include access time distribution, access location sequence, and access device fingerprint. The signaling interaction features include signaling type sequence, signaling field variation, and signaling node hopping. The data transmission features include transmission protocol type, encryption status identifier, and transmission traffic fluctuation. The network slicing features include slice identification information, resource allocation parameters, and isolation policy configuration. Extract risk feature combinations and causal relationships between features from historical data security incident cases, and establish a risk feature association rule base, which includes feature co-occurrence frequency, feature propagation delay, and feature impact weight; A feature node layer is constructed based on the basic feature set. Each feature node corresponds to a specific feature in the basic feature set. A unique feature identifier and feature attribute description are assigned to each feature node. The feature attribute description includes the feature data type, feature acquisition source, and feature security level. A feature association edge layer is constructed based on the risk feature association rule base. Each feature association edge connects two feature nodes that have an association relationship. Each feature association edge is configured with an association strength parameter and a transmission direction identifier. The association strength parameter is calculated based on the feature co-occurrence frequency and influence weight. The transmission direction identifier is used to indicate the influence direction between features. The feature node layer and the feature association edge layer are topologically integrated to generate a data security risk feature map that includes node attributes, edge attributes, and topological structure. The topological structure is used to describe the direct association relationships and indirect association paths between feature nodes. The data security risk feature map supports querying association strength parameters of associated features and association paths based on feature nodes.

3. The data security risk early warning method based on big data analysis according to claim 1, characterized in that, The process of collecting heterogeneous data sets from mobile communication networks in real time through multi-source data access interfaces, performing distributed cleaning and standardization on the heterogeneous data sets, and generating a multi-source heterogeneous data fusion stream includes: A multi-source data access interface is constructed, which includes a radio access network data interface, a core network control plane interface, a core network user plane interface, a network slice management interface, and a user equipment interaction interface. The radio access network data interface is used to collect base station-side access data, the core network control plane interface is used to collect mobility management and session management related signaling data, the core network user plane interface is used to collect user data transmission records, and the network slice management interface is used to collect slice resource allocation and status monitoring data. The user equipment interaction interface is used to collect user equipment operation behavior data. Configure the collection parameters of each data access interface, set the data collection frequency, data field filtering rules and data format conversion protocol, so that the collected data contains raw data fields of user identity characteristics, access behavior characteristics, signaling interaction characteristics, data transmission characteristics and network slicing characteristics; The multi-source data access interface is activated to collect heterogeneous data sets in the mobile communication network in real time. The heterogeneous data sets include structured data, semi-structured data, and unstructured data. The structured data includes user subscription information stored in a database table. The semi-structured data includes signaling messages in JSON format. The unstructured data includes user plane data in binary format. The heterogeneous data set is distributed to a distributed data processing cluster, and the heterogeneous data set is cleaned using a partitioned parallel processing method. The cleaned heterogeneous data is then standardized to generate standardized data records. The standardized data records include feature identifier fields, data value fields, timestamp fields, and data source fields. Based on the timestamp field, standardized data records from different sources are time-aligned. The time-aligned standardized data records are aggregated according to a preset time window length to generate a multi-source heterogeneous data fusion stream containing multi-source features. Each data unit in the multi-source heterogeneous data fusion stream contains user identity feature data, access behavior feature data, signaling interaction feature data, data transmission feature data, and network slice feature data within the same time window.

4. The data security risk early warning method based on big data analysis according to claim 1, characterized in that, The step of performing big data correlation analysis on the real-time risk feature vectors to mine the transmission dependencies between risk features and generate a risk propagation path weight set includes: Based on a preset risk feature threshold, abnormal risk features that exceed the threshold are filtered out from the real-time risk feature vector to obtain an abnormal feature set, which includes an abnormal feature identifier and corresponding feature value. Using the set of abnormal features as the starting feature node, based on the topology of the data security risk feature map, the directly associated feature nodes and indirectly associated feature nodes of the starting feature node are mined to generate a set of feature association paths. The set of feature association paths contains multiple path sequences, and each path sequence is composed of the starting feature node and its associated feature nodes arranged in the order of transmission. For each path sequence in the feature association path set, the association strength parameter between adjacent feature nodes in the path is extracted, and the path propagation probability is calculated based on the association strength parameter. The path propagation probability is equal to the product of the association strength parameters of all adjacent nodes in the path. Collect data on the actual impact of risk propagation paths in historical data security incidents, and construct a path impact assessment model. The path impact assessment model takes the path propagation probability, path length, and number of abnormal features contained in the path as inputs, and outputs the path impact weights. Each path sequence and its corresponding path propagation probability in the feature-associated path set are input into the path influence assessment model to calculate the path influence weight of each path sequence. The feature-associated path set is sorted based on the path influence weight, and path sequences with path influence weight greater than a preset weight threshold are selected as key risk propagation paths. For each key risk propagation path, the characteristic node identifiers, path transmission probabilities, and path impact weights in the path sequence are integrated to generate a risk propagation path weight set containing path identifiers, characteristic node sequences, transmission probability sequences, and impact weights.

5. The data security risk early warning method based on big data analysis according to claim 4, characterized in that, The step of using the abnormal feature set as the starting feature node, and based on the topological structure of the data security risk feature map, mining the directly and indirectly related feature nodes of the starting feature node to generate a feature association path set includes: Initialize a breadth-first search queue, add each anomaly feature identifier in the anomaly feature set as a starting node to the search queue, set the current path depth of each starting node to 1, and the current path sequence is a single-node sequence containing only the starting node; Take the head node from the search queue and query the adjacent nodes in the data security risk feature map that have a direct relationship with the head node. The adjacent nodes are determined based on the propagation direction of the associated edge and only include the nodes pointed to by the associated edge starting from the current node. For each adjacent node, determine whether the node has already appeared in the current path sequence. If it has not appeared, generate a new path sequence. The new path sequence is the current path sequence with the adjacent node added to the end. The depth of the new path is the current path depth plus 1. Determine whether the new path depth exceeds the preset maximum path depth threshold. If it does not exceed the threshold, add the new path sequence to the search queue and record the starting node, node sequence, and path depth of the path sequence. Repeat the steps of retrieving the head node from the search queue, querying adjacent nodes, and generating a new path sequence until the search queue is empty or the path depth of all path sequences reaches the maximum path depth threshold. Collect all generated path sequences, remove duplicate path sequences, and obtain a set of feature-related paths containing the starting feature node and its directly and indirectly related feature nodes. Each path sequence in the set of feature-related paths has a unique path identifier, which is generated by combining the starting node identifier, path depth, and node sequence hash value.

6. The data security risk early warning method based on big data analysis according to claim 5, characterized in that, The collection of historical data on the actual impact of risk propagation paths in security incidents, and the construction of a path impact assessment model, includes: Collect historical data security incident records. The historical data security incident records include the time of the incident, the sequence of risk characteristic nodes involved, the scope of the incident's impact, and the duration of the incident handling. The scope of the incident's impact is described by the number of affected users, the number of affected business types, and the amount of data leakage. The risk propagation path is extracted from the historical data security event records. The risk propagation path is a sequence of risk feature nodes involved in the event arranged in chronological order. For each risk propagation path, path feature parameters are calculated. The path feature parameters include path length, number of abnormal features in the path, average correlation strength parameter, and maximum correlation strength parameter. The path length is the number of feature nodes in the path, the average correlation strength parameter is the average value of all correlation edge strength parameters in the path, and the maximum correlation strength parameter is the maximum value of the correlation edge strength parameters in the path. The impact of the event is quantified into an impact score, which is obtained by weighted summation of the number of affected users, the number of affected business types, and the amount of data leakage. The weights are determined according to the importance of each indicator in the data security event. Construct a training dataset, which includes sample inputs and sample labels. The sample inputs are path feature parameters, and the sample labels are the corresponding influence scores. Initialize the path impact assessment model. The path impact assessment model adopts a gradient boosting decision tree structure, which contains multiple decision tree weak classifiers. Each decision tree weak classifier takes the path feature parameters as input and outputs the preliminary impact weights. The path impact assessment model is trained using the training dataset. The model parameters are adjusted by minimizing the mean squared error between the sample labels and the model output. The training is iterated until the mean squared error of the path impact assessment model on the validation set is lower than a preset threshold. The trained path impact assessment model is then saved and used to calculate the path impact weights for new risk propagation paths.

7. The data security risk early warning method based on big data analysis according to claim 1, characterized in that, The step of determining the risk diffusion level and key impact nodes based on the risk propagation path weight set, and generating a security risk warning instruction containing risk propagation path identifiers, includes: The risk propagation path weight set is analyzed to extract the path influence weight, path transmission probability, and feature node sequence of each key risk propagation path; Calculate the risk diffusion comprehensive index, which is the weighted sum of the path influence weights and path transmission probabilities of all key risk propagation paths. The weights are determined based on the path length, with shorter path lengths having larger weights. Based on the risk diffusion comprehensive index, a preset risk level classification rule is queried to determine the risk diffusion level. The risk level classification rule includes multiple index ranges and corresponding risk levels. For each key risk propagation path, a node importance analysis is performed on the sequence of characteristic nodes, and the betweenness centrality of each characteristic node is calculated. The betweenness centrality is the frequency with which the node appears in all key risk propagation paths. The higher the frequency, the greater the betweenness centrality. Feature nodes with a betweenness centrality greater than a preset centrality threshold are selected as key influencing nodes, and their feature identifiers and corresponding feature values ​​are extracted. The basic information fields for generating security risk warning instructions include warning timestamp, risk spread level, list of key impact nodes, and comprehensive risk spread index. A path description field is generated for each key risk propagation path. The path description field includes a path identifier, a sequence of feature nodes, a sequence of propagation probability, and a path impact weight. The basic information field and the path description field are integrated and encapsulated into a security risk warning instruction according to a preset instruction format. The instruction format includes an instruction header, an instruction body, and an instruction tail. The instruction header includes an instruction type identifier and a version number. The instruction body includes a basic information field and a path description field. The instruction tail includes a checksum. The security risk warning instruction is pushed to the mobile communication security management platform.

8. The data security risk early warning method based on big data analysis according to claim 7, characterized in that, The step of determining the risk diffusion level based on the preset risk level classification rules using the comprehensive risk diffusion index includes: Obtain preset risk level classification rules, which are formulated by the mobile communication security management department based on the impact assessment results of historical data security incidents. These rules include multiple continuous intervals of the comprehensive risk diffusion index and the risk level name, risk handling priority, and response time limit requirements corresponding to each interval. Extract the specific value of the comprehensive risk diffusion index, compare the specific value with the index range in the risk level classification rules, and determine the target range to which the specific value belongs. Query the risk level name, risk handling priority, and response time limit requirements corresponding to the target range, and determine the risk level name as the current risk diffusion level; Generate risk level description information, which includes the risk diffusion level name, risk handling priority, response time limit requirements, and a description of the typical impact scenarios corresponding to the risk level. The risk level description information is added to the basic information field of the security risk warning instruction to help the administrators of the mobile communication security management platform understand the meaning of the risk level and the handling requirements.

9. A data security risk early warning system based on big data analysis, characterized in that, The system includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the data security risk early warning method based on big data analysis as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Power supply service risk identification system based on big data analysis

    CN119809319A

  • Underground engineering geological safety dynamic risk assessment method based on multi-source data fusion

    CN120373874A