Smart community data security dynamic analysis method and system

By using topological analysis neural networks in smart communities for dynamic feature extraction and circular adjustment, the problem of poor adaptability of traditional data security analysis methods in complex network environments is solved, and higher analysis accuracy and risk identification capabilities are achieved.

CN119996214AInactive Publication Date: 2025-05-13GUANGZHOU LONGNENG CITY OPERATION & MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510128666.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional data security analysis methods are difficult to cope with the complex and changing network environment and evolving attack methods in smart communities, and there are problems such as high false alarm rate, low false alarm rate and poor adaptability.

Method used

Provide a dynamic analysis method for smart community data security, by obtaining the target topology device and its corresponding example communication subtopology structure from the example community communication topology, using topology analysis neural network for feature extraction and cyclic adjustment, and dynamically analyze data security risks.

Benefits of technology

It improves the accuracy and adaptability of data security analysis, reduces the false alarm rate and missed alarm rate, and can more effectively identify the data security risks of topological devices in the smart community.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996214A_ABST
    Figure CN119996214A_ABST
Patent Text Reader

Abstract

The invention provides a smart community data security dynamic analysis method and system, and the method comprises the steps: carrying out the feature extraction of an example communication sub-topological structure based on a to-be-adjusted topological analysis neural network, obtaining a frame coding feature representation, and carrying out the feature extraction of a target topological device, and obtaining a device coding feature; pooling the frame coding feature representation to obtain an example sub-topology feature vector; according to the feature fitting degree of the example topology feature vector and the equipment coding feature in the high-dimensional space, obtaining an environmental deviation metric value, and determining a target deviation cost value based on the environmental deviation metric value; determining a feature similarity cost value according to a spatial similarity coefficient between the example topology feature vector and the equipment coding feature; and determining a target cost value based on the target deviation cost value and the feature similarity cost value, and performing cyclic adjustment on the topology analysis neural network to be adjusted based on the target cost value, so that the feature extraction and risk identification capabilities of the neural network are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and system for dynamic analysis of data security in a smart community. Background Art

[0002] With the rapid development of smart community construction, the number of various topological devices and data interaction in the community has increased dramatically, which has brought unprecedented challenges to community data security. Traditional data security analysis methods often rely on static rule matching and feature detection, which is difficult to cope with complex and changing network environments and evolving attack methods. Especially in smart communities, due to the wide variety of devices, complex communication protocols, and frequent data interactions, traditional analysis methods use simple rule-based judgment methods, which often have problems such as high false alarm rate, low false alarm rate, and poor adaptability, making it difficult to meet the data security needs of smart communities. Summary of the invention

[0003] In view of this, the present application provides a smart community data security dynamic analysis method and system.

[0004] The technical solution of this application is implemented as follows:

[0005] On the one hand, the present application provides a method for dynamic analysis of data security in a smart community, the method comprising: obtaining a target topology device and an example communication sub-topology structure corresponding to the target topology device from an example community communication topology structure; based on a topology analysis neural network to be calibrated, extracting features from the example communication sub-topology structure to obtain a framework coding feature representation, and extracting features from the target topology device to obtain a device coding feature; pooling the framework coding feature representation to obtain an example sub-topology feature vector; obtaining an environmental deviation measurement value based on the feature fit between the example sub-topology feature vector and the device coding feature in a high-dimensional space, and determining a target deviation cost value based on the environmental deviation measurement value; determining a feature similarity cost value based on the spatial similarity coefficient between the example sub-topology feature vector and the device coding feature; determining a target cost value based on the target deviation cost value and the feature similarity cost value, and cyclically calibrating the topology analysis neural network to be calibrated based on the target cost value, and obtaining a target topology analysis neural network when the topology analysis neural network to be calibrated converges, and the target topology analysis neural network is used to perform data security detection on the topology devices of the smart community.

[0006] On the other hand, the present application provides a computer system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.

[0007] The smart community data security dynamic analysis method and system provided by the present application obtains the target topology device and the example communication sub-topology structure corresponding to the target topology device from the example community communication topology structure to form a sample for network training, and extracts features from the example communication sub-topology structure and the target topology device to obtain the framework coding feature representation and the device coding feature. The framework coding feature representation is merged to obtain the example sub-topology feature vector, and the example sub-topology feature vector is considered to be the environmental data of the target topology device. In this way, based on the feature fit between the example sub-topology feature vector and the device coding feature in the high-dimensional space, the environmental deviation metric is obtained, and the environmental deviation metric can evaluate the degree of interaction between the target topology device and the example communication sub-topology structure, that is, the risk coefficient. The present application also determines the feature similarity cost value based on the spatial similarity coefficient between the example sub-topology feature vector and the device coding feature in the feature space, and the feature similarity cost value can effectively balance the semantic field during feature extraction. Based on this, the target cost value generated by the environmental deviation metric value and the feature similarity cost value performs label-free cyclic adjustment on the topological analysis neural network to be adjusted, so that the neural network's ability to extract features and identify risks is enhanced.

[0008] As mentioned above, this application is based on unlabeled training, which can save a large number of labeled samples, prevent the problem of neural network accuracy distortion due to insufficient sample base, and further improve the accuracy of community data security analysis.

[0009] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0011] Figure 1 A schematic diagram of the implementation process of a smart community data security dynamic analysis method provided in an embodiment of the present application.

[0012] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0014] The embodiment of the present application provides a method for dynamic analysis of data security in a smart community, which can be executed by a processor of a computer system, wherein the computer system can refer to a device with data processing capabilities such as a server, a laptop computer, and a desktop computer.

[0015] Figure 1 A schematic diagram of the implementation process of a smart community data security dynamic analysis method provided in an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0016] Step S110: acquiring a target topology device and an example communication sub-topology structure corresponding to the target topology device from the example community communication topology structure.

[0017] In an embodiment of the present application, the example community communication topology is a data graph, in which each topological point represents a topological device, and the connecting lines represent the communication relationship between the devices. The example community communication topology is a basic sample for training, which is used to simulate the actual communication environment in the smart community. The target topology device is a specific topological point in the graph structure, that is, the topological device that needs to be analyzed. The example communication sub-topology is a sub-graph centered on the target topological device, including its neighboring devices and their connection relationships, which represents the local communication environment in which the target device is located.

[0018] In step S110, the computer system first traverses the example community communication topology structure, that is, visits it in sequence to identify all topological devices. This process can be implemented by a general graph traversal algorithm. For example, assuming that the example community communication topology structure contains 100 topological devices, the computer system will start from any device and gradually visit all reachable devices until the entire graph is completely traversed.

[0019] Next, the computer system screens the traversed topology devices to determine which devices can be used as target topology devices. The screening criteria can be determined based on actual needs, such as selecting devices with the largest communication volume, devices closest to the network center, or devices with specific functions. Taking the largest communication volume as an example, the computer system can calculate the degree of each device (i.e., the number of devices directly connected to it), and then select the device with the largest degree as the target topology device.

[0020] After determining the target topological device, the computer system constructs an example communication sub-topology structure corresponding to the device. This process can include two steps: first, starting from the target topological device, access to its neighboring topological devices; second, stop accessing when the access number threshold is reached, and the accessed devices and their connection relationships form an example communication sub-topology structure. The access number threshold can be set according to actual needs, for example, the number of devices directly connected to the target device, the number of devices in the two-layer neighborhood, or dynamically adjusted according to the communication strength between devices. Taking the graph as an example, assuming that the target topological device is topological point A, the access number threshold is set to 2. The computer system will take topological point A as the starting point and first access its directly connected devices (i.e., neighbors of topological point A), such as topological points B, C, and D. Then, for each neighbor topological point (such as topological point B), the computer system will further access its neighbors (i.e., the second-layer neighbors of topological point A), such as topological points E and F. When the number of accessed devices reaches or exceeds the access number threshold, the computer system will stop accessing and organize the accessed devices (topological points A, B, C, D, E and F) and the connection relationships between them into an example communication sub-topology structure.

[0021] When constructing the example communication sub-topology, the computer system can suppress the features of the target topology device to avoid excessive influence in the subsequent feature extraction and neural network training. The specific method of suppression can be to set the feature value of the target device to zero, replace it with a specific value, or perform normalization. For example, in the feature extraction stage, if the feature vector of the target device is represented as [1,2,3], the computer system can replace it with [0,0,0] to ensure that the subsequent analysis focuses on the environment of the target device rather than the device itself.

[0022] Through step S110, the computer system extracts the target topology device and its corresponding example communication sub-topology structure from the example community communication topology structure. These structures will be input into the topology analysis neural network as training samples for subsequent feature extraction, neural network adjustment and data security detection.

[0023] Step S120: Based on the topology analysis neural network to be adjusted, feature extraction is performed on the example communication sub-topology structure to obtain a framework coding feature representation, and feature extraction is performed on the target topology device to obtain a device coding feature.

[0024] In the embodiment of the present application, step S120 uses the topology analysis neural network to be adjusted to extract features from the example communication sub-topology structure and the target topology device to obtain the framework coding feature representation and device coding feature. This step is the basis for subsequent neural network adjustment, data security detection, and environmental deviation and feature similarity evaluation.

[0025] The topology analysis neural network is a neural network specially designed to process graph structure data. It can extract useful feature information from the graph structure. The frame encoding feature representation refers to the feature representation obtained after feature extraction of the example communication sub-topology structure. It reflects the context characteristics of the target topology device and can be expressed as a matrix. The device encoding feature is the feature representation obtained after feature extraction of the target topology device itself. It reflects the inherent attribute characteristics of the device and can be expressed as a vector.

[0026] In step S120, the computer system first uses a topology analysis neural network to extract features from the example communication sub-topology structure. This process may include: the computer system first pre-processes the example communication sub-topology structure and converts it into a format that the neural network can process. This may include representing the graph structure in the form of an adjacency matrix or a degree matrix for subsequent matrix operations. Design and implement a neural network module specifically for feature extraction. This module can be a convolutional neural network (CNN), a graph convolutional network (GCN), or other neural network variants suitable for graph structure data. The function of this module is to extract useful feature information from the graph structure and convert it into a frame encoding feature representation. The computer system inputs the pre-processed example communication sub-topology structure into the feature extraction network, extracts the feature information through the forward propagation process of the network, and encodes it into a frame encoding feature representation. This process may involve calculations of multiple network layers, including operations such as convolution, pooling, and activation functions.

[0027] Taking the graph convolutional network (GCN) as an example, assuming that the example communicator topology has been represented as the adjacency matrix A and the degree matrix D, the feature extraction network is designed as a two-layer GCN. The calculation formula of the first layer of GCN can be expressed as:

[0028]

[0029] Where X is the input feature matrix (which can be the identity matrix or a randomly initialized feature matrix at the beginning), W (0) is the weight matrix of the first layer of GCN, σ is the activation function (such as ReLU), It is the result of normalizing the adjacency matrix A. Through the calculation of this layer, the hidden layer feature matrix H can be obtained. (1) .

[0030] The calculation formula of the second layer GCN can be expressed as:

[0031]

[0032] Among them, W (1) is the weight matrix of the second layer of GCN. Through the calculation of this layer, the output feature matrix H can be obtained(2) , which is the frame encoding feature representation.

[0033] While obtaining the framework coding feature representation, the computer system extracts features from the target topology device to obtain the device coding features. Because the target topology device is just a topological point, not a graph structure, the computer system can directly extract feature information from the attributes of the target device and encode it into a vector. For example, if the target device has attributes such as IP address, MAC address, device type, etc., the computer system can convert these attributes into numerical features and combine them into a vector as the device coding feature.

[0034] In the feature extraction process, the computer system can normalize or reduce the dimension of the features according to actual needs to improve the quality of the features and the training efficiency of the neural network. For example, the maximum and minimum normalization method can be used to scale the features, or the principal component analysis (PCA) method can be used to reduce the dimension of the features.

[0035] Step S130: pooling the frame encoding feature representation to obtain an example sub-topology feature vector.

[0036] In an embodiment of the present application, the framework encoding feature representation is a feature representation obtained after feature extraction of the example communication sub-topology structure by a topological analysis neural network in step S120. It can be a matrix containing contextual feature information of the target topology device and its surrounding environment. Pooling operation is a dimensionality reduction method commonly used in neural networks. It reduces the size of the feature map by aggregating statistics of local areas of the feature map (or feature matrix) while retaining important feature information. The example sub-topology feature vector is the output result of the pooling operation. It is a vector used to represent the overall characteristics of the example communication sub-topology structure.

[0037] In step S130, the computer system first performs a pooling operation on the frame encoding feature representation. This process may include: first determining the size and step size of the pooling window. The pooling window is a small area used to slide on the feature map, which determines the feature range covered by each pooling operation. The step size is the step size of the pooling window moving on the feature map, which determines the degree of overlap between two adjacent pooling operations. Select the type of pooling operation, such as Max Pooling and Average Pooling. Max pooling selects the maximum value as the output within the pooling window, which can retain the significant features in the feature map; average pooling calculates the average value of all values ​​within the pooling window as the output, which can smooth the feature map and reduce noise. According to the selected pooling window size, step size and pooling type, a sliding window operation is performed on the frame encoding feature representation, and the feature values ​​in each window are aggregated and statistically analyzed to obtain a pooled feature map. This process may involve multiple iterations until the size of the feature map is reduced to a predetermined target size.

[0038] It reduces the number of features by reducing the dimension, thus reducing the complexity of subsequent calculations. Secondly, the pooling operation is translation invariant, that is, for a small translation in the feature map, the pooled result remains unchanged, which helps improve the robustness of the neural network. Finally, the pooling operation can retain the significant features in the feature map while reducing noise and redundant information, which helps improve the performance of the neural network.

[0039] Through the pooling operation in step S130, the computer system converts the frame encoding feature representation into an example sub-topology feature vector. This vector, as the basic input data for subsequent analysis, is of great significance for steps such as evaluating the degree of interaction between the target topology device and the example communication sub-topology structure, determining the environmental deviation metric, and performing feature similarity evaluation.

[0040] Step S140: obtaining an environmental deviation measurement value according to the feature fit between the example sub-topology feature vector and the device coding feature in the high-dimensional space, and determining a target deviation cost value according to the environmental deviation measurement value.

[0041] In an embodiment of the present application, step S140 quantifies the degree of deviation between the device topology device and the context environment by calculating the feature fit between the example sub-topology feature vector and the device coding feature in the high-dimensional space, that is, the environmental deviation metric, and further determines the target deviation cost value based on the metric value. The example sub-topology feature vector is extracted from the framework coding feature representation through a pooling operation in step S130, and it represents the overall characteristics of the example communication sub-topology structure. The device coding feature is obtained by feature extraction of the target topology device in step S120, which reflects the inherent attribute characteristics of the device. Feature fit refers to the degree of proximity between two feature vectors in a high-dimensional space, which can be measured by calculating the distance or similarity between them. The environmental deviation metric is calculated based on the feature fit and is used to indicate the degree of deviation between the device topology device and the context environment. The target deviation cost value is further determined based on the environmental deviation metric value, and is used to evaluate the performance of the model during the neural network calibration process.

[0042] In step S140, the computer system first calculates the feature fit between the example sub-topology feature vector and the device coding feature. This process can be implemented by feature distance calculation, such as cosine distance, Euclidean distance, Manhattan distance, etc., which is not specifically limited.

[0043] Cosine distance evaluates the proximity of two vectors by calculating the cosine of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are; the closer it is to -1, the more opposite the two vectors are; and the closer it is to 0, the two vectors are orthogonal. Euclidean distance calculates the straight-line distance between two vectors in multidimensional space. The smaller the distance, the closer the two vectors are; the larger the distance, the farther the two vectors are. Manhattan distance is similar to Euclidean distance, but it calculates the sum of the absolute differences between two vectors in each dimension in multidimensional space. Manhattan distance is more robust to sparse features in high-dimensional data.

[0044] Taking the cosine distance as an example, assuming that the example sub-topology feature vector is v = [v1, v2, ..., v n ], the device encoding feature vector is u=[u1,u2,...,u n ], then the cosine distance between them can be calculated by the following formula:

[0045]

[0046] Among them, v·u represents the dot product of two vectors, and ∥v∥ and ∥u∥ represent the modulus lengths of the two vectors respectively.

[0047] After obtaining the feature fit, the computer system further calculates the environment deviation metric. The environment deviation metric can be defined as 1 minus the feature fit (for cosine distance), or the inverse of the feature fit (for Euclidean distance or Manhattan distance), to ensure that a larger deviation value indicates a higher degree of deviation between the device topology and the context environment.

[0048] Finally, the computer system determines the target deviation cost value based on the environmental deviation metric value. The target deviation cost value can be proportional to the environmental deviation metric value, that is, the larger the deviation, the higher the cost value. This can be achieved through a simple linear transformation or a nonlinear transformation, depending on the requirements of the neural network tuning process.

[0049] Through the calculation in step S140, the computer system quantifies the degree of interaction between the example sub-topology feature vector and the device coding feature as an environmental deviation metric value, and further determines the target deviation cost value. These metrics will serve as an important basis for subsequent neural network adjustment and data security detection, helping to improve the accuracy and efficiency of the entire analysis method.

[0050] Step S150: determining a feature similarity cost value according to a spatial similarity coefficient between the example sub-topology feature vector and the device coding feature.

[0051] In the embodiment of the present application, step S150 determines the feature similarity cost by calculating the spatial similarity coefficient between the two. In addition to the cosine distance mentioned above, other measurement methods can be used to calculate the spatial similarity coefficient, such as the Pearson correlation coefficient, the Spearman rank correlation coefficient or the mutual information.

[0052] The example sub-topology feature vector is extracted from the framework encoding feature representation through a pooling operation in step S130, and it represents the overall characteristics of the example communication sub-topology structure. The device encoding feature is obtained by extracting features from the target topology device in step S120, and it reflects the inherent attribute characteristics of the device. The spatial similarity coefficient is a measure of the similarity between two feature vectors, and the feature similarity cost value is calculated based on the spatial similarity coefficient and is used to evaluate the performance of the model during the neural network tuning process.

[0053] In step S150, the computer system uses the Pearson correlation coefficient as a method for calculating the spatial similarity coefficient. The Pearson correlation coefficient is a statistic that measures the degree of linear correlation between two variables, and its value is between -1 and 1. When the correlation coefficient is 1, it means that the two variables are completely positively correlated; when the correlation coefficient is -1, it means that the two variables are completely negatively correlated; when the correlation coefficient is 0, it means that there is no linear correlation between the two variables. The calculation formula of the Pearson correlation coefficient is as follows:

[0054]

[0055] Where Cov(X,Y) is the covariance of variables X and Y, σ X and σ Y are the standard deviations of variables X and Y respectively.

[0056] In practical applications, the computer system first performs standardization processing on the example sub-topology feature vector and the device coding feature to eliminate the influence of different feature dimensions. The standardization processing may include subtracting the mean and dividing by the standard deviation, so that the processed feature vector has a mean of 0 and a standard deviation of 1. Then, the computer system calculates the Pearson correlation coefficient between the two standardized feature vectors as the spatial similarity coefficient between them.

[0057] Assume that the example subtopology feature vector is v = [v1, v2, ..., v n ], the device encoding feature vector is u=[u1,u2,...,u n ], their vectors after normalization are v'=[v'1,v'2,...,v' n ] and u'=[u'1,u'2,...,u' n ]. Then the Pearson correlation coefficient between them can be calculated by the following formula:

[0058]

[0059] in, and are the means of vectors v' and u' respectively.

[0060] After obtaining the spatial similarity coefficient, the computer system further calculates the feature similarity cost. The feature similarity cost can be proportional to the spatial similarity coefficient, that is, the higher the similarity, the lower the cost; the lower the similarity, the higher the cost. This can be achieved through simple linear transformation or nonlinear transformation, depending on the needs of the neural network tuning process. For example, the feature similarity cost can be calculated using the following formula:

[0061] FeatureSimilarityCost=1-|ρ v ' ,u '|;

[0062] Among them, |ρ v',u' | is the absolute value of the spatial similarity coefficient, which is used to ensure that the cost value is always non-negative.

[0063] Through the calculation in step S150, the computer system quantifies the similarity between the example sub-topology feature vector and the device coding feature as a feature similarity cost value. This metric value will serve as an important basis for subsequent neural network calibration and data security detection, helping to improve the accuracy and efficiency of the entire analysis method.

[0064] Step S160: Determine a target cost value based on the target deviation cost value and the feature similarity cost value, and perform cyclic adjustment on the topology analysis neural network to be adjusted based on the target cost value. When the topology analysis neural network to be adjusted converges, obtain a target topology analysis neural network, and the target topology analysis neural network is used to perform data security detection on the topology devices of the smart community.

[0065] In the embodiment of the present application, step S160 combines the target deviation cost value and the feature similarity cost value to determine a comprehensive target cost value, and based on this cost value, the topology analysis neural network to be adjusted is cyclically adjusted. This process is intended to enable the neural network to more accurately identify the data security risks of topological devices in the smart community through iterative optimization.

[0066] The target deviation cost value is calculated in step S140, which reflects the degree of deviation between the target topology device and the context environment. The feature similarity cost value is calculated in step S150, which measures the similarity between the example sub-topology feature vector and the device coding feature. The target cost value is obtained by combining these two cost values ​​and is used to evaluate the current performance of the neural network. Cyclic tuning refers to the process of continuously adjusting the parameters of the neural network through multiple iterations to optimize its performance. Convergence refers to the state in which the performance of the neural network gradually stabilizes and no longer improves significantly during the tuning process.

[0067] In step S160, the computer system first determines the target cost value based on the target deviation cost value and the feature similarity cost value. This process may include: the computer system assigns different weights to the target deviation cost value and the feature similarity cost value. The allocation of weights can be adjusted according to actual needs to balance the importance of deviation and similarity in the neural network calibration process. For example, if you pay more attention to the deviation between the device and its context environment, you can assign a higher weight to the target deviation cost value; if you pay more attention to the similarity between the features, you can assign a higher weight to the feature similarity cost value. After assigning the weights, the computer system performs a weighted summation of the target deviation cost value and the feature similarity cost value to obtain the target cost value.

[0068] In some embodiments, the computer system can dynamically adjust the weights according to the rounds of tuning. For example, in the early stages of tuning, the feature similarity cost value can be given a higher weight to help the neural network learn the similarities between features more quickly; in the later stages of tuning, the weight of the target deviation cost value can be gradually increased to more accurately adjust the deviation recognition capability of the neural network.

[0069] After determining the target cost value, the computer system will cyclically adjust the topological analysis neural network to be adjusted based on this cost value. This process may include: the computer system inputs the training sample (including the example communication sub-topology structure and the target topology device) into the neural network, and obtains the output value through forward propagation calculation. According to the output value and the true value (or expected output value) of the neural network, the computer system calculates the loss function value. The loss function value reflects the gap between the current performance and the expected performance of the neural network. The computer system passes the loss function value backward layer by layer through the back propagation algorithm to calculate the gradient of each layer of neurons. According to the calculated gradient, the computer system updates the parameters of the neural network to reduce the loss function value. This process may involve the gradient descent algorithm or its variants. The computer system repeats the four steps of forward propagation, loss calculation, back propagation and parameter update until the neural network converges or reaches a preset number of iterations.

[0070] During the cyclic tuning process, the performance of the neural network will gradually improve and the target cost value will gradually decrease. When the neural network converges, that is, when its performance no longer improves significantly after multiple iterations, the computer system will stop tuning and use the neural network at this time as the target topological analysis neural network. This neural network will be used to perform data security detection on the topological devices of the smart community, and evaluate the security risks of the device by identifying the deviations between the device and its context and the similarities between the features.

[0071] It should be noted that during the cyclic tuning process, the computer system can use a variety of optimization strategies to improve the tuning efficiency and neural network performance. For example, the learning rate decay strategy can be used to prevent the neural network from falling into the local optimal solution; the early stopping strategy can be used to prevent overfitting; and regularization methods such as batch normalization and Dropout can be used to improve the generalization ability of the neural network.

[0072] Through the cyclic adjustment process of step S160, the computer system combines the target deviation cost value and the feature similarity cost value into a target cost value, and gradually improves the performance of the neural network through iterative optimization. The target topology analysis neural network finally obtained will be able to more accurately identify the data security risks of topological devices in the smart community, providing a strong guarantee for the data security of the smart community.

[0073] In one implementation, the example communication sub-topology structure includes a safe example communication sub-topology structure and a risky example communication sub-topology structure; the step S110 of acquiring a target topology device and an example communication sub-topology structure corresponding to the target topology device from the example community communication topology structure includes:

[0074] Step S111: sequentially access each topology device in the example community communication topology structure, and screen each accessed topology device to obtain a target topology device;

[0075] Step S112: starting from the target topology device in the example community communication topology structure, access is made to adjacent topology devices, and the access is terminated when the access number threshold is reached, and a secure example communication sub-topology structure corresponding to the target topology device is obtained according to the accessed topology devices;

[0076] Step S113: determining a risky example communication sub-topology structure corresponding to the target topology device according to the safe example communication sub-topology structures of other target topology devices different from the target topology device.

[0077] Step S111 is the process of screening the target topology device. The computer system first traverses the example community communication topology structure to access each topology device therein. This step can be implemented through a graph traversal algorithm. During the traversal process, the computer system screens each accessed topology device to determine which devices can be used as target topology devices. The screening criteria can be determined according to actual needs. For example, you can select devices with the largest communication volume, devices closest to the network center, devices with specific functions, or devices with specific risk characteristics. Taking the largest communication volume as an example, the computer system can calculate the degree of each device (that is, the number of devices directly connected to it), and then select the device with the largest degree as the target topology device. Assuming that the example community communication topology structure contains 100 topology devices, the computer system finally determines device A as the target topology device through traversal and screening.

[0078] Next, step S112 is the process of constructing a secure example communication sub-topology. The computer system takes the target topology device (such as device A) determined in step S111 as the starting point and accesses its neighboring topology devices. The number of neighboring devices accessed is determined by the access number threshold, which can be set according to actual needs. For example, the number of devices directly connected to the target device, the number of devices in the two-layer neighborhood, or the number of devices can be dynamically adjusted according to the communication strength between devices. During the access process, the computer system will record the accessed devices and the connection relationship between them to construct a secure example communication sub-topology corresponding to the target topology device. Taking device A as an example, assuming that the access number threshold is set to 2, the computer system will access the direct neighbors of device A (such as devices B, C and D) and the neighbors of these neighbors (such as devices E and F), and these devices and their connection relationships will form a secure example communication sub-topology. In this sub-topology, device A is the central topology point, and its connection relationship with other devices reflects the local communication environment in which device A is located.

[0079] Finally, step S113 is the process of constructing a risk example communication sub-topology structure. Unlike step S112, step S113 does not construct a sub-topology structure directly from the target topology device, but is constructed based on the secure example communication sub-topology structure of other target topology devices. Specifically, the computer system first selects other target topology devices (such as devices G, H, etc.) that are different from the target topology device (such as device A), and these devices have been screened out in step S111. Then, for these other target topology devices, the computer system constructs their respective secure example communication sub-topology structures according to the method of step S112. Finally, the computer system regards these secure example communication sub-topology structures as risk examples and compares them with the secure example communication sub-topology structure of the target topology device (device A) to determine the risk example communication sub-topology structure corresponding to device A. This process can be understood as that by comparing the secure communication environments of different target topology devices, the computer system can identify the security environment and possible risk environment similar to the target topology device, thereby providing a reference for subsequent data security detection.

[0080] In practical applications, for example, when processing a large community communication topology, the computer system uses an efficient graph traversal algorithm and data storage structure to ensure the efficiency of traversal and screening. At the same time, when constructing an example communication sub-topology, the computer system also needs to consider how to deal with redundant information and noise data in the sub-topology to improve the accuracy of subsequent feature extraction and neural network calibration. In addition, the implementation of step S110 may also be affected by some factors, such as the selection criteria of the target topology device, the setting of the access number threshold, the construction method of the example communication sub-topology, etc. The selection and setting of these factors will directly affect the effects of subsequent steps and the performance of the entire analysis method. Therefore, when implementing step S110, the computer system performs detailed parameter adjustments and experimental verifications according to actual needs to ensure that high-quality training samples can be constructed.

[0081] Through the implementation of step S110, the computer system extracts the target topology device and its corresponding example communication sub-topology structure (including security examples and risk examples) from the example community communication topology structure. These structures will be input into the subsequent neural network as training samples for feature extraction and adjustment, providing strong support for data security detection in smart communities.

[0082] In one implementation, the step S112, starting from the target topology device in the example community communication topology structure, accesses the adjacent topology devices, and terminates when the access number threshold is reached, and obtains the secure example communication sub-topology structure corresponding to the target topology device according to the accessed topology device, including:

[0083] Step S1121: Starting from the target topology device in the example community communication topology structure, access is made to adjacent topology devices, and the access is terminated when the access number threshold is reached to obtain a sub-topology structure, and the features of the target topology device in the obtained sub-topology structure are suppressed to obtain a safe example communication sub-topology structure corresponding to the target topology device.

[0084] Step S1121 starts from the target topology device, accesses the adjacent topology devices, builds a sub-topology structure including the target device and its adjacent devices, and suppresses the characteristics of the target device to eliminate its influence on the overall characteristics of the sub-topology structure. This process is intended to simulate the communication environment of the target device in an ideal security state, providing a benchmark for subsequent data security detection and neural network adjustment.

[0085] In step S1121, the computer system first takes the target topological device (assuming it is device A) as the starting point and accesses its neighboring topological devices. The number of neighboring devices accessed is determined by the access number threshold, which can be set according to actual needs. For example, the number of devices directly connected to the target device, the number of devices in the two-layer neighborhood, or dynamically adjusted according to the communication strength between devices can be selected. Assuming that the access number threshold is set to 2, the computer system will access the direct neighbors of device A (such as devices B, C, and D) and the neighbors of these neighbors (such as devices E and F), and record these devices and their connection relationships to form a preliminary sub-topology structure.

[0086] However, this preliminary sub-topology structure contains the characteristic information of the target device A, which may interfere with the subsequent analysis of the secure communication environment. Therefore, in step S1121, the computer system suppresses the characteristics of the target device A. The specific method of the suppression process can be determined according to actual needs, but the purpose can be to reduce or eliminate the impact of the target device characteristics on the overall characteristics of the sub-topology structure. For example, a zeroing process can be used, that is, each element in the characteristic vector of the target device A is set to 0, so that its contribution to the sub-topology structure is 0. Assume that the characteristic vector of device A is [a1, a2, ..., an], after the zeroing process, its characteristic vector becomes [0, 0, ..., 0].

[0087] The sub-topology after suppression is the required secure example communication sub-topology. This sub-topology only includes the neighboring devices of target device A and their connection relationships, and excludes the influence of the target device's own characteristics, so it can more accurately reflect the communication environment of the target device in an ideal security state.

[0088] In practical applications, the implementation of step S1121 may involve complex graph structures and data processing techniques. For example, when processing a large community communication topology, the computer system uses an efficient graph traversal algorithm and data storage structure to ensure the efficiency of access and recording. At the same time, when suppressing the target device features, the computer system also needs to consider how to retain other important feature information of the sub-topology to ensure the accuracy of subsequent analysis.

[0089] In addition, the implementation of step S1121 may also be affected by some factors, such as the setting of the access number threshold, the method of suppressing the characteristics of the target device, etc. The selection and setting of these factors will directly affect the construction quality of the secure example communication sub-topology structure and the effect of subsequent analysis. Therefore, when implementing step S1121, the computer system performs detailed parameter adjustment and experimental verification according to actual needs to ensure that a high-quality secure example communication sub-topology structure can be constructed.

[0090] Through the implementation of step S1121, the computer system extracts the secure example communication sub-topology structure corresponding to the target topology device from the example community communication topology structure. This structure will serve as an important basis for subsequent feature extraction, neural network adjustment and data security detection, and will help improve the accuracy and efficiency of the entire analysis method.

[0091] In one implementation, the step S120 of extracting features from the example communication sub-topology structure to obtain a framework coding feature representation includes:

[0092] Step S121: obtaining an example topology device degree table indicating the number of connections of each topology device in the example communication sub-topology structure;

[0093] Step S122: performing weight balancing on the connection relationships between topology devices in the example connection relationship table of the example communication sub-topology structure based on the example topology device degree table, to obtain an example connection density table for indicating the connection density between topology devices in the example communication sub-topology structure;

[0094] Step S123: merging the example connection density table with the example source feature table of the example communication sub-topology structure to obtain a framework coding feature representation.

[0095] The degree table, also known as the degree matrix, is a matrix used to indicate the number of connections of each topological device in the example communication sub-topology. In a computer system, step S121 may involve traversing the example communication sub-topology, counting the number of neighbors of each device, and organizing this information into the form of a degree matrix. For example, assuming that the example communication sub-topology contains 5 topological devices (labeled as A, B, C, D, E), its degree matrix may be as follows:

[0096] In this matrix, each row represents a device, each column also represents a device, and the elements in the matrix represent the degree (i.e. the number of neighbors) of the corresponding device. For example, the degree of device A is 2, which means it has two neighbors (assuming they are devices B and C).

[0097] Next, step S122 is a process of weight balancing the example connection relationship table (adjacency matrix) of the example communication sub-topology structure based on the example topology device degree table. The adjacency matrix is ​​a matrix used to represent the connection relationship between topological devices, where an element of 1 indicates that two devices are connected, and 0 indicates that they are not connected. The process of weight balancing completes the purpose of normalization and adjusts the element values ​​in the adjacency matrix to reflect the closeness of the connection relationship between devices. In a computer system, this step may involve dividing the element values ​​in the adjacency matrix by the degree of the corresponding device to obtain an example connection density table (association matrix). Continuing with the above example, assuming that the adjacency matrix of the example communication sub-topology structure is:

[0098]

[0099] After weight balancing, the resulting correlation matrix may be as follows:

[0100]

[0101] In this matrix, the value of each element represents the closeness of the connection relationship between the corresponding devices, that is, the connection density. For example, the connection density between device A and device B is 1 / 2, which means that among all the neighbors of device A, device B accounts for half of the proportion.

[0102] Finally, step S123 is the process of merging the example connection density table with the example source feature table (initial attribute matrix) of the example communication sub-topology. The source feature table is a matrix used to represent the inherent attribute characteristics of topological devices, in which each row represents a device and each column represents an attribute feature. In a computer system, this step may involve some form of merging the association matrix and the source feature table to obtain a framework encoding feature representation containing structural features and attribute features. The merging method can be determined according to actual needs, but the purpose can be to effectively combine structural features and attribute features so that subsequent neural networks can learn useful information from them. For example, assuming that the source feature table of the example communication sub-topology is:

[0103]

[0104] Among them, each row represents two attribute characteristics of a device (for example, the computing power and storage capacity of the device). In order to merge the association matrix and the source feature table, the computer system can use matrix multiplication, weighted average or other merging strategies. Assuming that matrix multiplication is used as the merging method, the specific calculation process involves performing a dot product operation on each row of the association matrix and each column of the source feature table, and organizing the results into a new matrix. This new matrix is ​​the framework encoding feature representation, which contains the structural characteristics and attribute feature information of the example communication sub-topology structure, and provides important input data for the subsequent feature extraction and adjustment of the neural network.

[0105] In one implementation, the step S123, merging the example connection density table with the example source feature table of the example communication sub-topology structure to obtain a framework coding feature representation, includes:

[0106] Step S1231: Based on the connection density between topological devices in the example connection density table, the units in the example source feature table are cyclically merged, and the merging rounds are stopped when the merging rounds meet the preset rounds, and the merging result obtained by the last cyclic merging is determined as the framework coding feature representation.

[0107] The goal of step S1231 is to combine the units in the example source feature table (i.e., the attribute features of the device) with the connection density information in the example connection density table by means of loop merging to form a framework encoding feature representation that can reflect the relationship between topological devices and their attribute features. This process aims to make full use of the structural information and attribute information in the example communication sub-topology structure to provide richer and more effective input data for subsequent neural network data security detection.

[0108] In step S1231, the computer system first uses an example connection density table as a basis, which reflects the connection density between topological devices. For example, assuming that the example connection density table is: The element values ​​in the matrix represent the closeness of the connection relationship between the corresponding devices (for example, obtained through normalization). Next, the computer system will perform a cyclic merging of the cells in the example source feature number table based on the association matrix.

[0109] The example source feature table is a matrix containing the attribute features of topological devices, each row of which represents a device and each column represents an attribute feature. For example, suppose the example source feature table is: In this matrix, each row represents two attribute characteristics of a device (for example, the computing power and storage capacity of the device).

[0110] During the cyclic merging process, the computer system will iteratively update the cells in the source feature number table according to a certain merging strategy (such as weighted average, maximum value selection, etc.). Each round of merging will merge the attribute features of adjacent devices according to the connection density information in the association matrix to form a new feature representation. For example, in the first round of merging, the computer system can perform a weighted average of their attribute features based on the connection density (0.5) between device 1 and device 2 to obtain a new feature vector. Then, in the second round of merging, the computer system can further merge based on the updated feature vector and the connection density (0.7) between device 3.

[0111] This process will continue until the number of merge rounds meets the preset number of rounds. The preset number of rounds is a parameter set according to actual needs, which determines the number of loop merges and the complexity of the final frame encoding feature representation. In practical applications, the selection of the preset number of rounds needs to take into account factors such as the limitation of computing resources, the dimension of feature representation, and the processing power of the subsequent neural network.

[0112] After the loop merging is completed, the computer system determines the merged result obtained from the last loop merging as the frame encoding feature representation. This feature representation is a matrix containing the relationship between topological devices and their attribute characteristics, which will be used as part of the subsequent neural network input data to support tasks such as data security detection.

[0113] The merging strategy in step S1231 can be flexibly selected according to actual needs. For example, in addition to weighted averaging, maximum selection, minimum selection, summation or other merging methods can also be used. In addition, in order to further improve the richness and effectiveness of feature representation, the computer system may also consider introducing technical means such as nonlinear transformation, feature selection or dimensionality reduction in the merging process. Through the implementation of step S1231, the computer system effectively merges the example connection density table with the example source feature table to construct a framework encoding feature representation containing structural features and attribute features. This feature representation not only retains the connection relationship information between topological devices, but also incorporates the attribute feature information of the device, providing more comprehensive and effective input data support for subsequent neural network data security detection.

[0114] In one implementation, the example communication sub-topology structure includes a safe example communication sub-topology structure and a risky example communication sub-topology structure; in step S140, obtaining an environmental deviation metric value according to a feature fit between the example sub-topology feature vector and the device coding feature in a high-dimensional space includes:

[0115] Step S141: performing a scalar product solution based on the device coding feature and the mirror vector of the example sub-topology feature vector of the secure example communication sub-topology structure to obtain a first scalar product;

[0116] Step S142: estimating a first deviation factor according to the first quantitative product;

[0117] Step S143: performing a scalar product solution based on the device coding feature and the mirror vector of the example sub-topology feature vector of the risk example communication sub-topology structure to obtain a second scalar product;

[0118] Step S144: estimating a second deviation factor according to the second quantitative product;

[0119] Step S145: Determine an environmental deviation metric value based on the first deviation factor and the second deviation factor.

[0120] Step S141 involves solving the scalar product between the device coding feature and the mirror vector of the example sub-topology feature vector of the secure example communication sub-topology structure. In this step, the computer system first obtains the device coding feature vector of the target topology device, which can be a high-dimensional vector containing the inherent attribute characteristics of the device. At the same time, the computer system also obtains the example sub-topology feature vector of the secure example communication sub-topology structure, which is obtained by performing feature extraction and pooling operations on the secure example communication sub-topology structure, and it reflects the contextual environment characteristics of the target topology device in an ideal security state. In order to calculate the feature fit between the two vectors, the computer system mirrors the example sub-topology feature vector, that is, inverts all its elements (or performs other forms of transformation to keep the vector dimension unchanged) to obtain a mirror vector. Then, the scalar product between the device coding feature vector and the mirror vector is calculated, that is, the sum of the products of the corresponding elements of the two vectors. This process can be expressed as a mathematical formula:

[0121]

[0122] Where d represents the device encoding feature vector, The mirror vector of the example subtopology feature vector representing the secure example communication subtopology, n is the dimension of the vector, d i and v safe,i are the element values ​​of the two vectors in the i-th dimension.

[0123] Next, step S142 estimates the first scalar product to obtain a first deviation factor. The purpose of this step is to convert the scalar product result into a metric that can reflect the degree of deviation between the device coding feature and the secure example communication sub-topology. Since the size of the scalar product is proportional to the cosine value of the angle between the two vectors, the deviation factor can be obtained by performing some form of transformation on the scalar product (such as negation, scaling, or application of a nonlinear function). For example, the first deviation factor can be calculated using the following formula:

[0124]

[0125] Where f is a nonlinear function used to map the scalar product result to the range of the deviation factor (for example, [0,1] or [-1,1]). The inversion here is to make the deviation factor inversely proportional to the size of the scalar product, that is, the smaller the scalar product (indicating that the two vectors are less similar), the larger the deviation factor.

[0126] Similarly, steps S143 and S144 involve solving the dot product between the device encoding feature and the mirror vector of the example sub-topology feature vector of the risk example communication sub-topology structure and estimating the deviation factor. This process is similar to steps S141 and S142, but the example sub-topology feature vector of the risk example communication sub-topology structure is used. By calculating the dot product between the device encoding feature vector and the risk example mirror vector and applying the same transformation function, the second deviation factor can be obtained:

[0127]

[0128] In step S145, the computer system determines the environmental deviation metric based on the first deviation factor and the second deviation factor. This process may involve some form of comprehensive processing of the two deviation factors to obtain a single metric that can fully reflect the degree of deviation between the device topology and the context environment. The comprehensive processing method can be determined according to actual needs, but it may be necessary to consider the relative importance of the two deviation factors and the interaction between them. For example, the environmental deviation metric can be calculated by weighted summation, and the size of the weight reflects the importance of the two deviation factors in the calculation of the environmental deviation metric.

[0129] In the case where the target topology device corresponds to multiple safety example communication sub-topologies and multiple risk example communication sub-topologies, the computer system calculates the scalar product and the deviation factor for each safety example and risk example respectively, and averages the deviation factors of all safety examples (or performs other forms of aggregation processing) to obtain the final first deviation factor and second deviation factor. Then, the environmental deviation metric is calculated based on the two comprehensive deviation factors.

[0130] In one implementation, the method further includes:

[0131] Step S1301: Based on the topological analysis neural network to be adjusted, feature extraction is performed again on the example communication sub-topology structure to obtain a mixed feature matrix;

[0132] Step S1302: performing feature inversion reconstruction on the mixed feature matrix to obtain an example inversion feature matrix;

[0133] Step S1303: determining a topology device deviation cost value according to the difference between the inversion feature of the target topology device in the example inversion feature matrix and the source feature of the target topology device.

[0134] Based on the above steps S1301 to S1303, in step S140, determining the target deviation cost value according to the environmental deviation measurement value includes:

[0135] Step S146: determining a target deviation cost value based on the environment deviation metric value and the topology device deviation cost value.

[0136] First, step S1301 involves re-feature extraction of the example communication sub-topology structure to obtain a mixed feature matrix. This process is different from the feature extraction in step S120. It not only focuses on the attribute features or structural features of the topological device, but attempts to combine these two features to form a more comprehensive feature representation. In order to achieve this goal, the computer system uses the topological analysis neural network to be adjusted to perform feature extraction on the example communication sub-topology structure again. During this feature extraction process, the neural network not only considers the connection relationship between the topological devices (i.e., graph structure information), but also considers the attribute characteristics of each device (such as computing power, storage capacity, etc.). Through the nonlinear transformation and feature fusion mechanism of the neural network, the computer system obtains a mixed feature matrix, each row of which represents a topological device, and each column represents a mixed feature (i.e., a feature that combines graph structure information and attribute information).

[0137] For example, suppose the example communication sub-topology contains three topological devices A, B, and C, whose attribute characteristics are computing power, storage capacity, and communication rate. After feature extraction by the neural network, the computer system may obtain a mixed feature matrix as follows:

[0138]

[0139] Among them, each row represents a mixed feature vector of a device, and each column represents a mixed feature.

[0140] Next, step S1302 performs feature inversion reconstruction on the mixed feature matrix to obtain an example inversion feature matrix. Feature inversion reconstruction is an inverse process that attempts to recover the original feature information from the mixed feature matrix. In a computer system, this process can be implemented by a decoder network, which is the inverse process of the feature extraction network (i.e., the encoder network). Through the nonlinear transformation and feature mapping mechanism of the decoder network, the computer system converts the mixed feature vector into an example inversion feature vector, and then obtains an example inversion feature matrix. Each row of this matrix represents an inversion feature vector of a topological device, which should be as close as possible to the original attribute feature vector (i.e., the source feature vector), but there may also be certain deviations due to information loss in the feature extraction process or the influence of nonlinear transformation.

[0141] For example, for device A, its original attribute feature vector is [0.5, 0.3, 0.7]. After feature inversion and reconstruction, the inverted feature vector obtained may be [0.48, 0.32, 0.69]. There is a certain deviation between this inverted feature vector and the original feature vector, which reflects the information loss and nonlinear transformation in the process of feature extraction and inversion reconstruction.

[0142] Then, step S1303 determines the topological device deviation cost value based on the difference between the inversion feature and the source feature of the target topological device in the example inversion feature matrix. The purpose of this step is to quantify the deviation of the target topological device in the feature space, that is, the degree of inconsistency between the inverted feature vector and the original feature vector. In a computer system, this can be achieved by calculating metrics such as the Euclidean distance, Manhattan distance, or cosine distance between two vectors. The obtained topological device deviation cost value reflects the stability or consistency of the target topological device during feature extraction and inversion reconstruction.

[0143] For example, for device A, its inverted eigenvector is [0.48, 0.32, 0.69], and its original eigenvector is [0.5, 0.3, 0.7]. By calculating the Euclidean distance between these two vectors, the computer system can obtain the topological device deviation cost value as: The smaller the cost value is, the closer the inverted feature vector is to the original feature vector, and the more stable the feature extraction and inversion reconstruction process is.

[0144] Finally, step S146 determines the target deviation cost value based on the environmental deviation metric value and the topological device deviation cost value. This step combines the environmental deviation metric value calculated in step S140 with the topological device deviation cost value obtained in step S1303 to form a more comprehensive deviation metric. In a computer system, this can be achieved by weighted summation or nonlinear combination. The target deviation cost value obtained not only reflects the degree of environmental deviation between the target topological device and the example communication sub-topology structure, but also takes into account the deviation of the target topological device in the feature space.

[0145] For example, assuming that the environmental deviation metric value is 0.2 (indicating that there is a certain degree of deviation between the target topology device and the example communication sub-topology structure), and the topology device deviation cost value is 0.03 (indicating that the deviation of the target topology device in the feature space is small), the computer system can obtain the target deviation cost value by weighted summation.

[0146] Through the implementation of steps S1301 to S1303 and step S146, the computer system can more comprehensively evaluate the degree of interaction between the target topology device and the example communication sub-topology structure and the deviation of the target topology device in the feature space. This information is of great significance for the adjustment of neural networks and data security detection. It helps to improve the generalization ability and detection accuracy of neural networks and provide more reliable protection for the data security of smart communities.

[0147] In one implementation, the example communication sub-topology structure includes a safe example communication sub-topology structure and a risky example communication sub-topology structure; the step S150, determining a feature similarity cost value according to a spatial similarity coefficient between the example sub-topology feature vector and the device coding feature, includes:

[0148] Step S151: determining a security space similarity coefficient between the example sub-topology feature vector of the secure example communication sub-topology structure and the device coding feature;

[0149] Step S152: determining a risk space similarity coefficient between the example sub-topology feature vector of the risk example communication sub-topology structure and the device coding feature;

[0150] Step S153: Determine a feature similarity cost based on the measurement difference between the safety space similarity coefficient and the risk space similarity coefficient.

[0151] First, step S151 involves determining the security space similarity coefficient between the example sub-topology feature vector of the secure example communication sub-topology structure and the device coding feature. In a computer system, this step can be implemented by calculating the similarity of two vectors in space. There are many ways to measure similarity, including but not limited to cosine distance, Pearson correlation coefficient, Spearman rank correlation coefficient, etc. Here, cosine distance is used as an example for explanation. The cosine distance evaluates the similarity between two vectors by calculating the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are, the closer it is to -1, the more opposite it is, and close to 0, it means orthogonal.

[0152] Next, step S152 determines the risk space similarity coefficient between the example sub-topology feature vector of the risk example communication sub-topology structure and the device encoding feature. This process is similar to step S151, but the example sub-topology feature vector of the risk example communication sub-topology structure is used. The risk space similarity coefficient reflects the feature similarity of the target topology device in the risk communication environment.

[0153] Finally, step S153 determines the feature similarity cost value based on the metric difference (i.e., comparison error) between the similarity coefficient of the safe space and the similarity coefficient of the risk space. In a computer system, this step can be implemented by calculating a function of the difference or ratio of the two similarity coefficients. The size of the metric difference reflects the degree of inconsistency of the feature similarity of the target topology device in the safe communication environment and the risk communication environment. The higher the feature similarity cost value, the greater the degree of inconsistency, and the target topology device may face a higher data security risk.

[0154] The calculation method of the feature similarity cost value can be set according to actual needs. For example, the following formula can be used for calculation:

[0155] FeatureSimilarityCost=g(|cos(θ)-cos(θ')|);

[0156] Where g is a nonlinear function used to map the metric difference to the range of feature similarity cost values ​​(e.g., [0, 1] or [-1, 1]). The absolute value is taken here to ensure that the metric difference is always non-negative. The choice of nonlinear function g can be adjusted according to the actual application scenario to reflect the degree of influence of the metric difference on the feature similarity cost value.

[0157] In the case where the target topology device corresponds to multiple safety example communication sub-topology structures and multiple risk example communication sub-topology structures, the implementation of steps S151 to S153 will be more complicated. In this case, the computer system may need to calculate the spatial similarity coefficient for each safety example and risk example respectively, and average the similarity coefficients of all safety examples (or other forms of aggregation processing) to obtain the final safety space similarity coefficient and risk space similarity coefficient. Then, the feature similarity cost is calculated based on these two comprehensive similarity coefficients.

[0158] Through the implementation of step S150, the computer system can accurately evaluate the feature similarity between the target topology device and the example communication sub-topology structure, and quantify the inconsistency of this similarity in a secure communication environment and a risky communication environment. This metric is of great significance for subsequent tasks such as neural network calibration, data security detection, and risk assessment. It helps to improve the generalization ability and detection accuracy of neural networks, and provide more reliable protection for the data security of smart communities. At the same time, by introducing the concepts of secure examples and risk examples, step S150 can also more comprehensively consider the performance of devices in different communication environments, providing richer information support for data security analysis.

[0159] In one implementation, the step S160, determining the target cost value based on the target deviation cost value and the feature similarity cost value, includes:

[0160] Step S161: obtaining a first adjustment weight indicating the importance of the target deviation cost value during adjustment, wherein the first adjustment weight increases with an increase in the number of adjustment rounds;

[0161] Step S162: obtaining a second adjustment weight indicating the importance of the feature similarity cost value during adjustment, wherein the second adjustment weight decreases with an increase in the number of adjustment rounds;

[0162] Step S163: According to the first adjustment weight and the second adjustment weight, the target deviation cost value and the feature similarity cost value are combined to obtain a target cost value.

[0163] The step S161 of obtaining a first adjustment weight indicating the importance of the target deviation cost value during adjustment includes:

[0164] Step S1611: Obtain the maximum set adjustment rounds and the real-time completed rounds;

[0165] Step S1612: Divide the real-time completed rounds by the maximum set adjustment rounds to obtain the first adjustment weight.

[0166] Step S161 involves obtaining a first adjustment weight indicating the importance of the target deviation cost value in the adjustment. The purpose of this step is to dynamically adjust the weight of the target deviation cost value in the target cost value calculation according to the change of the adjustment round. In a computer system, this process can be achieved by calculating the ratio between the real-time completed rounds and the maximum set adjustment rounds.

[0167] Step S1611 first obtains the maximum set calibration rounds and the real-time completed rounds. The maximum set calibration rounds are the total number of preset neural network calibration rounds, which determines the termination condition of the neural network calibration. The real-time completed rounds are the number of rounds that have been completed during the current calibration process. These two parameters change dynamically. As the calibration process proceeds, the real-time completed rounds will gradually increase until the maximum set calibration rounds are reached.

[0168] Step S1612 divides the real-time completed rounds by the maximum set adjustment rounds to obtain the first adjustment weight. This weight is a value between 0 and 1, which reflects the proportion of the current adjustment round in the total adjustment rounds. As the adjustment rounds increase, the real-time completed rounds will gradually approach the maximum set adjustment rounds, so the first adjustment weight will gradually increase. This means that in the early stage of adjustment, the target deviation cost value has a smaller weight in the target cost value calculation, and in the later stage of adjustment, its weight will gradually increase, thereby emphasizing the importance of the target deviation cost value in neural network adjustment.

[0169] For example, assuming that the maximum set adjustment rounds are 100 rounds, and the current real-time completed rounds are 30 rounds, then the first adjustment weight is 0.3.

[0170] Next, step S162 involves obtaining a second adjustment weight indicating the importance of the feature similarity cost value during the adjustment. In contrast to the first adjustment weight, the second adjustment weight gradually decreases with the increase of the adjustment round. This is because in the early stage of adjustment, the feature similarity cost value is of great significance in helping the neural network to quickly learn the intrinsic structure and feature relationship of the data, while in the later stage of adjustment, as the neural network's ability to adapt to the data gradually increases, the importance of the feature similarity cost value is relatively weakened.

[0171] In a computer system, the second adjustment weight may be calculated by subtracting the first adjustment weight from 1. Continuing with the above example, the second adjustment weight is 1-0.3=0.7.

[0172] Finally, step S163 combines the target deviation cost value and the feature similarity cost value according to the first adjustment weight and the second adjustment weight to obtain the target cost value. The purpose of this step is to combine the cost values ​​of two different dimensions to form a single cost value that can fully reflect the current performance of the neural network. In a computer system, the merging process can be implemented by weighted summation.

[0173] Through the implementation of step S160, the computer system can dynamically adjust the weights of the target deviation cost value and the feature similarity cost value in the target cost value calculation, thereby ensuring that the neural network can gradually optimize its performance during the calibration process. In the early stage of calibration, the feature similarity cost value has a higher weight, which helps the neural network to quickly learn the intrinsic structure and feature relationship of the data; and in the later stage of calibration, the target deviation cost value has a higher weight, which helps the neural network to more accurately identify the degree of deviation between the target topological device and the context environment. This dynamic weight adjustment mechanism enables the neural network to adjust parameters more efficiently and improve the accuracy and efficiency of data security detection.

[0174] In one implementation, the method further includes an application phase of the neural network, comprising the following steps:

[0175] Step S170: acquiring a topology device to be analyzed and a target sub-topology structure corresponding to the topology device to be analyzed from the communication topology structure of the community to be analyzed;

[0176] Step S180: extracting features of the target sub-topology structure based on the target topology analysis neural network to obtain a target framework coding feature representation, and extracting features of the topology device to be analyzed to obtain a target device coding feature;

[0177] Step S190: pooling the target framework coding feature representation to obtain a target sub-topology feature vector, and determining an environmental deviation metric value according to a feature fit between the target sub-topology feature vector and the target device coding feature in a high-dimensional space;

[0178] Step S200: Based on the environmental deviation metric value, determine whether the topology device to be analyzed is a risky topology device.

[0179] Step S170 involves obtaining the topology device to be analyzed and the corresponding target sub-topology from the communication topology of the community to be analyzed. In a computer system, this step can be implemented by traversing the communication topology of the community to be analyzed. The traversal process can use a graph traversal algorithm to ensure that all devices in the topology can be accessed. During the traversal process, the computer system will identify the topology device to be analyzed (i.e., the device whose security needs to be analyzed) and other devices directly or indirectly connected to it. These devices and their connection relationships together constitute the target sub-topology.

[0180] For example, suppose that the community communication topology to be analyzed contains multiple topological devices, one of which is selected as the topological device to be analyzed (labeled as device D). By traversing the topological structure, the computer system can identify the devices directly connected to device D (such as devices E, F, and G) and the neighbors of these devices (such as devices H and I), thereby forming a target sub-topological structure. This sub-topological structure reflects the local communication environment of device D and is the basis for subsequent feature extraction and deviation measurement.

[0181] Next, step S180 involves using the trained target topology analysis neural network to extract features from the target sub-topology structure and the topology device to be analyzed. For the target sub-topology structure, the neural network will treat it as a graph structure, and extract features through a graph convolutional network (GCN) or other neural network models suitable for graph data to obtain a target framework encoding feature representation. This feature representation is a high-dimensional vector that contains the structural information and attribute information of the target sub-topology structure, which reflects the overall characteristics of the communication environment in which device D is located.

[0182] At the same time, for the topological device to be analyzed (device D), the neural network will directly extract its attribute features to obtain the target device encoding features. This feature representation is a high-dimensional vector that contains the inherent attribute information of device D, such as computing power, storage capacity, communication rate, etc.

[0183] In the feature extraction process, the neural network will use the weights and bias parameters learned during its training process to perform nonlinear transformation and feature fusion on the input data, thereby extracting useful feature information. This feature information will be used for subsequent deviation measurement and risk judgment.

[0184] Then, step S190 involves performing a pooling operation on the target frame encoding feature representation to obtain a target sub-topology feature vector, and determining the environmental deviation metric based on the feature fit between the vector and the target device encoding feature in a high-dimensional space. Pooling is a dimensionality reduction method that reduces the size of a feature map and retains important feature information by performing local region aggregation statistics on a feature map (or feature vector). In a computer system, pooling can be implemented by methods such as maximum pooling or average pooling.

[0185] For the target frame encoding feature representation (assuming it is a matrix), the computer system will perform a pooling operation on it to obtain a one-dimensional target sub-topology feature vector. This vector contains the overall feature information of the target sub-topology structure, and its dimension is low, which is convenient for subsequent calculations.

[0186] Then, the computer system calculates the feature fit between the target sub-topology feature vector and the target device encoding feature in the high-dimensional space. Feature fit can be achieved by calculating the cosine distance, Euclidean distance or other similarity metrics between the two vectors. In the computer system, the cosine distance can be used to evaluate the similarity between two vectors because the cosine distance can reflect the proximity of the vectors in direction without being affected by the vector modulus.

[0187] The calculation result of the feature fit reflects the degree of deviation between the topological device to be analyzed (device D) and the communication environment it is in. If the feature fit is high (i.e., the cosine distance is close to 1), it means that the features between device D and its communication environment are relatively consistent, and there may be a low security risk; if the feature fit is low (i.e., the cosine distance is close to -1 or 0), it means that there is a large deviation between the features of device D and its communication environment, and there may be a high security risk.

[0188] Finally, step S200 determines whether the topology device to be analyzed is a risk topology device based on the environmental deviation metric value. In a computer system, this step can be implemented by setting a threshold. If the environmental deviation metric value is greater than the threshold, the topology device to be analyzed is judged to be a risk topology device; otherwise, it is judged to be a non-risk topology device.

[0189] The threshold setting can be adjusted according to actual needs. For example, in scenarios with high security requirements, the threshold can be set lower to more strictly screen risky devices; in scenarios with relatively low security requirements, the threshold can be set higher to reduce the false alarm rate.

[0190] Through the implementation of the neural network application phase, the computer system can efficiently use the trained target topology analysis neural network to perform security analysis on the devices in the communication topology structure of the community to be analyzed. This process not only takes into account the attribute characteristics of the device itself, but also the overall characteristics of the communication environment in which the device is located and the interaction between the device and the environment, thereby improving the accuracy and efficiency of data security detection. At the same time, by dynamically adjusting the threshold of the environmental deviation metric and comprehensively considering the combined impact of multiple factors, the computer system can more flexibly respond to data security challenges in different scenarios.

[0191] In one implementation, the target sub-topology structure includes a first target sub-topology structure covering the topology device to be analyzed and a second target sub-topology structure not covering the topology device to be analyzed; in step S190, determining the environmental deviation metric value according to the feature fit between the target sub-topology feature vector and the target device encoding feature in the high-dimensional space includes:

[0192] Step S191: obtaining a first similarity coefficient between a target sub-topology feature vector of the first target sub-topology structure and the target device coding feature in a high-dimensional space;

[0193] Step S192: obtaining a second similarity coefficient between the target sub-topology feature vector of the second target sub-topology structure and the target device coding feature in a high-dimensional space;

[0194] Step S193: determining an environment deviation metric value target environment deviation metric value according to the difference between the first similarity coefficient and the second similarity coefficient.

[0195] Step S191 involves obtaining a first similarity coefficient between a target sub-topology feature vector of a first target sub-topology structure and a target device encoding feature in a high-dimensional space. The first target sub-topology structure is a sub-topology structure covering the topology device to be analyzed, which reflects the communication relationship between the topology device to be analyzed and its directly or indirectly connected devices. In a computer system, the target sub-topology feature vector of the first target sub-topology structure is obtained by inputting the first target sub-topology structure into a trained target topology analysis neural network for feature extraction and pooling operations. This feature vector contains the overall feature information of the first target sub-topology structure and is the basis for the subsequent calculation of the similarity coefficient.

[0196] The target device coding feature is obtained by directly extracting the attribute features of the topological device to be analyzed, which reflects the inherent attribute information of the topological device to be analyzed. In a computer system, the target device coding feature can be a high-dimensional vector that contains multiple attribute features such as the computing power, storage capacity, and communication rate of the device.

[0197] The calculation method of the first similarity coefficient can refer to the determination process of the first deviation factor above. Specifically, the computer system can obtain the first similarity coefficient by calculating the cosine distance, Euclidean distance or other similarity measurement indicators between the target sub-topology feature vector and the target device coding feature.

[0198] Next, step S192 involves obtaining a second similarity coefficient between the target sub-topology feature vector of the second target sub-topology structure and the target device encoding feature in the high-dimensional space. The second target sub-topology structure is a sub-topology structure that does not cover the topology device to be analyzed, and it reflects the communication relationship between other devices outside the topology device to be analyzed. Similar to the first target sub-topology structure, the target sub-topology feature vector of the second target sub-topology structure is also obtained by inputting it into the target topology analysis neural network for feature extraction and pooling operations.

[0199] The calculation method of the second similarity coefficient can refer to the determination process of the second deviation factor above. Unlike the first similarity coefficient, the second similarity coefficient reflects the degree of feature consistency between the topology device to be analyzed and its communication environment (within the scope covered by the second target sub-topology structure). Since the second target sub-topology structure does not cover the topology device to be analyzed, the second similarity coefficient may be low, indicating that there are large feature differences between the topology device to be analyzed and the devices covered by the second target sub-topology structure.

[0200] Finally, step S193 determines the environmental deviation metric value based on the difference between the first similarity coefficient and the second similarity coefficient. In a computer system, this step can be implemented by calculating a function of the difference or ratio of the two similarity coefficients. The size of the difference reflects the inconsistency of the degree of feature consistency of the topology device to be analyzed in different communication environments. If the first similarity coefficient is much higher than the second similarity coefficient, it means that there is a high degree of feature consistency between the topology device to be analyzed and the device directly or indirectly connected to it, and there is a large feature difference between the topology device to be analyzed and the devices covered by the second target sub-topology structure, which may mean that the topology device to be analyzed is in a relatively independent or special communication environment, and there is a high security risk.

[0201] It should be noted that when calculating the first similarity coefficient and the second similarity coefficient, the computer system may need to adopt different feature extraction and pooling strategies to adapt to the characteristics of different sub-topological structures. In addition, when determining the environmental deviation metric, the comprehensive influence of other factors may also be considered, such as the size of the sub-topological structure, the connection density between devices, etc.

[0202] By implementing step S190, the computer system can more comprehensively evaluate the degree of deviation between the topological device to be analyzed and the communication environment it is in. This process not only takes into account the attribute characteristics of the topological device to be analyzed itself, but also takes into account the inconsistency of the degree of consistency of its characteristics with different communication environments, thereby providing a more accurate and reliable basis for data security detection.

[0203] In one implementation, the method further includes:

[0204] Step S201: extracting features of the target sub-topological structure again based on the target topological analysis neural network to obtain a target mixed feature matrix;

[0205] Step S202: performing feature inversion reconstruction on the target mixed feature matrix to obtain a target inversion feature matrix;

[0206] Step S203: determining a topology device risk coefficient according to a difference between the target inversion feature matrix and the source feature of the topology device in the target sub-topology structure;

[0207] Based on the above steps S201 to S203, the step S200, based on the environmental deviation metric value, determines whether the topology device to be analyzed is a risk topology device, including:

[0208] Step S210: Based on the environmental deviation metric value and the topology device risk coefficient, determine whether the topology device to be analyzed is a risky topology device.

[0209] In the embodiment of the present application, steps S201 to S203 and sub-step S210 of step S200 constitute a further refinement of the security assessment of the topological device to be analyzed. These steps not only take into account the degree of deviation between the topological device to be analyzed and its communication environment (i.e., the environmental deviation metric), but also introduce a new indicator, the topological device risk coefficient, to more comprehensively evaluate the security of the device. Below, the principles and implementation methods of these steps will be explained in depth and technology, and detailed explanations will be given with examples.

[0210] First, step S201 involves re-extracting features of the target sub-topology structure based on the target topology analysis neural network to obtain a target mixed feature matrix. This process is similar to the feature extraction in step S180, but the goal is clearer: to extract a comprehensive feature representation that can simultaneously reflect the structural features and attribute features of the target sub-topology structure. In a computer system, this goal can be achieved by inputting the target sub-topology structure into a trained target topology analysis neural network and utilizing the multi-layer nonlinear transformation and feature fusion mechanism of the neural network.

[0211] The target hybrid feature matrix is ​​a high-dimensional matrix that contains the overall feature information of the target sub-topology. Each row represents a topological device, and each column represents a hybrid feature (i.e., a feature that combines structural features and attribute features). For example, assuming that the target sub-topology contains three topological devices A, B, and C, and each device has two attribute features (such as computing power and storage capacity), the target hybrid feature matrix may be as follows:

[0212] Among them, m ij represents the j-th hybrid feature value of device i, and n is the total number of hybrid features.

[0213] Next, step S202 performs feature inversion reconstruction on the target mixed feature matrix to obtain the target inversion feature matrix. Feature inversion reconstruction is an inverse process that attempts to recover the original feature information from the mixed feature matrix. In a computer system, this process can be implemented by a decoder network, which is the inverse process of the feature extraction network (i.e., the encoder network). Through the nonlinear transformation and feature mapping mechanism of the decoder network, the computer system converts the mixed feature vector into an inversion feature vector, thereby obtaining the target inversion feature matrix.

[0214] Each row of the target inversion feature matrix represents the inversion feature vector of a topological device, which should be as close as possible to the source feature vector (i.e., the original attribute vector) of the device. However, due to the information loss and nonlinear transformation during feature extraction and inversion reconstruction, there may be a certain deviation between the inversion feature vector and the source feature vector. This deviation reflects the stability or consistency of the topological device in the feature space and is an important basis for assessing device risk.

[0215] Then, step S203 determines the topological device risk coefficient based on the difference between the target inversion feature matrix and the source features of the topological device in the target sub-topological structure. In a computer system, this step can be achieved by calculating metrics such as the Euclidean distance, Manhattan distance, or cosine distance between the inversion feature vector and the source feature vector. The obtained topological device risk coefficient reflects the degree of deviation or inconsistency of the device in the feature space. The higher the risk coefficient, the worse the stability of the device during feature extraction and inversion reconstruction, and there may be a higher security risk.

[0216] For example, for device A, its source feature vector is [s A1 ,s A2 ], the inversion eigenvector is [r A1 ,r A2 By calculating the Euclidean distance between these two vectors, the risk factor of device A can be obtained: Similarly, the risk factors of device B and device C can be calculated.

[0217] Finally, step S210 determines whether the topological device to be analyzed is a risky topological device based on the environmental deviation metric and the topological device risk coefficient. In a computer system, this step can be implemented by setting a threshold. If the sum of the environmental deviation metric and the topological device risk coefficient is greater than the threshold, the topological device to be analyzed is judged to be a risky topological device; otherwise, it is judged to be a non-risky topological device.

[0218] For example, suppose the environmental deviation metric is E, the topology device risk factor is R, and the threshold is T. For the topology device to be analyzed (assuming it is device A), if the following conditions are met: E+R A >T, then device A is judged to be a risky topology device. This judgment process comprehensively considers the degree of deviation between the device and its communication environment and the stability of the device in the feature space, providing a more comprehensive and reliable basis for data security detection.

[0219] Through the implementation of steps S201 to S203 and step S210, the computer system can more accurately evaluate the security of the topological device to be analyzed. This process not only takes into account the interactive relationship between the device and its communication environment, but also introduces a new indicator, the topological device risk coefficient, to evaluate the stability of the device in the feature space. This comprehensive evaluation method helps to improve the accuracy and efficiency of data security detection and provide more reliable protection for the data security of smart communities.

[0220] In one implementation, the step S210, based on the environmental deviation metric and the topology device risk coefficient, determines whether the topology device to be analyzed is a risky topology device, includes:

[0221] Step S211: determining a maximum environment deviation metric value and a minimum environment deviation metric value from the environment deviation metric values ​​of each topology device in the communication topology structure of the community to be analyzed;

[0222] Step S212: Based on the maximum environmental deviation metric value and the minimum environmental deviation metric value, weight balancing is performed on the environmental deviation metric values ​​of the topology devices to be analyzed to obtain an environmental risk balance value;

[0223] Step S213: determining a maximum topology device risk coefficient and a minimum topology device risk coefficient from the topology device risk coefficients of each topology device in the communication topology structure of the community to be analyzed;

[0224] Step S214: Based on the maximum topology device risk coefficient and the minimum topology device risk coefficient, weight balancing is performed on the topology device risk coefficients of the topology devices to be analyzed to obtain a topology device risk balancing value;

[0225] Step S215: The environmental risk balance value and the topology device risk balance value are combined to obtain a risk coefficient of the topology device to be analyzed, and when the risk coefficient is greater than a risk threshold, the topology device to be analyzed is determined as a risky topology device.

[0226] Step S211 involves determining the maximum environmental deviation metric value and the minimum environmental deviation metric value from the environmental deviation metric value of each topology device in the community communication topology structure to be analyzed. The environmental deviation metric value reflects the degree of deviation between the topology device and the communication environment in which it is located, and is one of the important indicators for evaluating the security of the device. In a computer system, this step can be achieved by sorting the environmental deviation metric values ​​of all topology devices. After sorting, the value at the beginning of the sequence is the maximum environmental deviation metric value, and the value at the end of the sequence is the minimum environmental deviation metric value.

[0227] For example, suppose the community communication topology to be analyzed contains five topological devices (labeled as devices A to E), and their environmental deviation metrics are 0.3, 0.6, 0.4, 0.8, and 0.5, respectively. By sorting, it can be obtained that the maximum environmental deviation metric is 0.8 (corresponding to device D) and the minimum environmental deviation metric is 0.3 (corresponding to device A).

[0228] Next, step S212 performs weight balancing on the environmental deviation metric values ​​of the topology devices to be analyzed based on the maximum environmental deviation metric value and the minimum environmental deviation metric value to obtain an environmental risk balance value. The purpose of weight balancing is to normalize the environmental deviation metric value to a unified scale for comparison and merging with other indicators. This calculation process converts the environmental deviation metric value of device C from the original scale to the [0,1] interval, which is convenient for subsequent comparison and merging with other indicators.

[0229] Then, steps S213 and S214 involve determining the maximum and minimum topological device risk coefficients, and weighting the topological device risk coefficients of the topological devices to be analyzed to obtain the topological device risk balance value. The topological device risk coefficient reflects the stability or consistency of the device in the feature space and is another important indicator for evaluating device security. Similar to steps S211 and S212, steps S213 and S214 are also implemented through sorting and normalization operations.

[0230] For example, assume that the topology device risk coefficients of the five topology devices are 0.2, 0.5, 0.3, 0.7, and 0.4, respectively. By sorting, it can be obtained that the maximum topology device risk coefficient is 0.7 (corresponding to device D), and the minimum topology device risk coefficient is 0.2 (corresponding to device A). Then, for the topology device to be analyzed (assuming it is still device C), its topology device risk balance value can be calculated by the following formula:

[0231] Device Risk BalancedValue C =(0.3-0.2) / (0.7-0.2)=0.2;

[0232] This calculation process also converts the topological device risk coefficient of device C from the original scale to the [0,1] interval.

[0233] Finally, step S215 combines the environmental risk balance value and the topological device risk balance value to obtain the final risk coefficient of the topological device to be analyzed. In a computer system, this step can be implemented by weighted summation.

[0234] After obtaining the final risk coefficient, the computer system compares it with the preset risk threshold. If the risk coefficient is greater than the risk threshold, the topology device to be analyzed is judged as a risk topology device; otherwise, it is judged as a non-risk topology device. For example, assuming the risk threshold is 0.5, for device C, its final risk coefficient is:

[0235] FinalRiskCoefficient C =0.5×0.2+0.5×0.2=0.2;

[0236] Since 0.2 is less than 0.5, device C is determined to be a non-risk topology device.

[0237] Through the implementation of step S210, the computer system can comprehensively consider the two indicators of environmental deviation measurement value and topological device risk coefficient to evaluate the security of the topological device to be analyzed. This process not only considers the interactive relationship between the device and its communication environment, but also considers the stability of the device in the feature space, providing a more comprehensive and reliable basis for data security detection. At the same time, through weight balancing and merging operations, the computer system can unify indicators of different scales to the same scale for comparison and judgment, thereby improving the accuracy and efficiency of the evaluation.

[0238] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application is shown in FIG. Figure 2As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0239] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data). This can be achieved through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0240] When the processor 1001 executes the program, the steps of any of the above-mentioned smart community data security dynamic analysis methods are implemented. The processor 1001 generally controls the overall operation of the computer system 1000.

[0241] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the smart community data security dynamic analysis method as described in any of the above embodiments.

[0242] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for dynamic analysis of data security in a smart community, characterized in that: The method comprises: Acquire a target topology device and an example communication sub-topology structure corresponding to the target topology device from an example community communication topology structure; Based on the topology analysis neural network to be adjusted, feature extraction is performed on the example communication sub-topology structure to obtain a framework coding feature representation, and feature extraction is performed on the target topology device to obtain a device coding feature; Pooling the frame encoding feature representation to obtain an example sub-topology feature vector; Obtaining an environmental deviation measurement value according to a feature fit between the example sub-topology feature vector and the device coding feature in a high-dimensional space, and determining a target deviation cost value according to the environmental deviation measurement value; Determining a feature similarity cost value according to a spatial similarity coefficient between the example sub-topology feature vector and the device coding feature; A target cost value is determined based on the target deviation cost value and the feature similarity cost value, and the topology analysis neural network to be calibrated is cyclically calibrated based on the target cost value. When the topology analysis neural network to be calibrated converges, a target topology analysis neural network is obtained. The target topology analysis neural network is used to perform data security detection on the topology devices of the smart community.

2. The method according to claim 1, characterized in that The example communication sub-topology structure includes a safe example communication sub-topology structure and a risk example communication sub-topology structure; the step of obtaining a target topology device and an example communication sub-topology structure corresponding to the target topology device from the example community communication topology structure includes: Visit each topology device in the example community communication topology structure in turn, and filter each visited topology device to obtain a target topology device; Starting from the target topology device in the example community communication topology structure, access is made to adjacent topology devices, and the access is terminated when the access number threshold is reached, and a secure example communication sub-topology structure corresponding to the target topology device is obtained according to the accessed topology devices; Determine a risky example communication sub-topology structure corresponding to the target topology device according to the safe example communication sub-topology structures of other target topology devices different from the target topology device; The extracting features of the example communication sub-topology structure to obtain a frame coding feature representation includes: Obtaining an example topology device degree table indicating the number of connections of each topology device in the example communication sub-topology structure; Based on the example topology device degree table, weight balancing is performed on the connection relationships between topology devices in the example connection relationship table of the example communication sub-topology structure, to obtain an example connection density table for indicating the connection density between each topology device in the example communication sub-topology structure; The example connection density table is combined with the example source feature table of the example communication sub-topology structure to obtain a framework encoding feature representation.

3. The method according to claim 2, characterized in that The method of accessing adjacent topology devices from a target topology device in the example community communication topology structure as a starting point, terminating when a threshold of the number of accesses is reached, and obtaining a secure example communication sub-topology structure corresponding to the target topology device according to the accessed topology device, includes: Accessing adjacent topology devices from a target topology device in the example community communication topology structure as a starting point, and terminating when a threshold of the number of accesses is reached to obtain a sub-topology structure, suppressing the features of the target topology device in the obtained sub-topology structure, and obtaining a secure example communication sub-topology structure corresponding to the target topology device; The step of merging the example connection density table with the example source feature table of the example communication sub-topology structure to obtain a framework encoding feature representation includes: Based on the connection density between topological devices in the example connection density table, the units in the example source feature table are cyclically merged, and the merging rounds are stopped when the merging rounds meet the preset rounds, and the merging result obtained by the last cyclic merging is determined as the framework coding feature representation.

4. The method according to claim 1, characterized in that: The example communication sub-topology structure includes a safe example communication sub-topology structure and a risk example communication sub-topology structure; the obtaining of the environmental deviation metric value according to the feature fit between the example sub-topology feature vector and the device coding feature in the high-dimensional space includes: Solving a scalar product based on the device coding feature and a mirror vector of an example sub-topology feature vector of the secure example communication sub-topology structure to obtain a first scalar product; Estimating a first deviation factor according to the first quantitative product; Solving the scalar product based on the device coding feature and the mirror vector of the example sub-topology feature vector of the risk example communication sub-topology structure to obtain a second scalar product; Estimating a second deviation factor according to the second quantitative product; Based on the first deviation factor and the second deviation factor, an environmental deviation metric value is determined.

5. The method according to claim 1, characterized in that The method further comprises: Based on the topological analysis neural network to be adjusted, feature extraction is performed again on the example communication sub-topology structure to obtain a mixed feature matrix; Performing feature inversion reconstruction on the mixed feature matrix to obtain an example inversion feature matrix; Determine a topology device deviation cost value according to a difference between an inversion feature of the target topology device in the example inversion feature matrix and a source feature of the target topology device; Determining the target deviation cost value according to the environmental deviation measurement value includes: A target deviation cost value is determined based on the environment deviation metric value and the topology device deviation cost value.

6. The method according to claim 1, characterized in that The example communication sub-topology structure includes a safe example communication sub-topology structure and a risk example communication sub-topology structure; the determining of the feature similarity cost value according to the spatial similarity coefficient between the example sub-topology feature vector and the device coding feature includes: Determining a security space similarity coefficient between an example sub-topology feature vector of the secure example communication sub-topology structure and the device encoding feature; Determining a risk space similarity coefficient between an example sub-topology feature vector of the risk example communication sub-topology structure and the device coding feature; Determining a feature similarity cost based on a metric difference between the safety space similarity coefficient and the risk space similarity coefficient; The determining of the target cost value based on the target deviation cost value and the feature similarity cost value includes: Obtaining a first adjustment weight indicating the importance of the target deviation cost value during adjustment, wherein the first adjustment weight increases with an increase in the number of adjustment rounds; Obtaining a second adjustment weight indicating the importance of the feature similarity cost value during adjustment, wherein the second adjustment weight decreases with an increase in the number of adjustment rounds; According to the first adjustment weight and the second adjustment weight, the target deviation cost value and the feature similarity cost value are combined to obtain a target cost value; The obtaining of a first adjustment weight indicating the importance of the target deviation cost value during adjustment includes: Get the maximum set adjustment rounds and the real-time completed rounds; The real-time completed rounds are divided by the maximum set adjustment rounds to obtain the first adjustment weight.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Acquire a topology device to be analyzed and a target sub-topology structure corresponding to the topology device to be analyzed from the communication topology structure of the community to be analyzed; Based on the target topology analysis neural network, feature extraction is performed on the target sub-topology structure to obtain a target framework coding feature representation, and feature extraction is performed on the topology device to be analyzed to obtain a target device coding feature; Pooling the target framework coding feature representation to obtain a target sub-topology feature vector, and determining an environmental deviation metric value according to a feature fit between the target sub-topology feature vector and the target device coding feature in a high-dimensional space; Based on the environmental deviation metric value, it is determined whether the topology device to be analyzed is a risky topology device.

8. The method according to claim 7, characterized in that The target sub-topology structure includes a first target sub-topology structure covering the topology device to be analyzed and a second target sub-topology structure not covering the topology device to be analyzed; and determining the environmental deviation metric value according to the feature fit between the target sub-topology feature vector and the target device encoding feature in the high-dimensional space includes: Obtaining a first similarity coefficient between a target sub-topology feature vector of the first target sub-topology structure and the target device encoding feature in a high-dimensional space; Obtaining a second similarity coefficient between a target sub-topology feature vector of the second target sub-topology structure and the target device encoding feature in a high-dimensional space; Determining an environmental deviation metric value target environmental deviation metric value according to a difference between the first similarity coefficient and the second similarity coefficient; The method further comprises: Based on the target topology analysis neural network, feature extraction is performed again on the target sub-topology structure to obtain a target mixed feature matrix; Performing feature inversion reconstruction on the target mixed feature matrix to obtain a target inversion feature matrix; Determining a topological device risk coefficient according to a difference between the target inversion characteristic matrix and a source characteristic of a topological device in the target sub-topological structure; The determining, based on the environmental deviation metric value, whether the topology device to be analyzed is a risky topology device includes: Based on the environmental deviation metric value and the topology device risk coefficient, it is determined whether the topology device to be analyzed is a risky topology device.

9. The method according to claim 8, characterized in that The determining, based on the environmental deviation metric value and the topology device risk coefficient, whether the topology device to be analyzed is a risky topology device includes: Determine a maximum environment deviation metric value and a minimum environment deviation metric value from the environment deviation metric value of each topology device in the communication topology structure of the community to be analyzed; Based on the maximum environmental deviation metric value and the minimum environmental deviation metric value, weight balancing is performed on the environmental deviation metric values ​​of the topology devices to be analyzed to obtain an environmental risk balance value; Determine a maximum topology device risk coefficient and a minimum topology device risk coefficient from the topology device risk coefficient of each topology device in the communication topology structure of the community to be analyzed; Based on the maximum topology device risk coefficient and the minimum topology device risk coefficient, weight balancing is performed on the topology device risk coefficients of the topology devices to be analyzed to obtain a topology device risk balancing value; The environmental risk balance value and the topology device risk balance value are combined to obtain a risk coefficient of the topology device to be analyzed, and when the risk coefficient is greater than a risk threshold, the topology device to be analyzed is determined as a risky topology device.

10. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 9 are implemented.