Equipment terminal identification method for RIS phased array network deployment
By monitoring and analyzing the initial feature set and abnormal attack behavior of the terminal in the RIS phased array network, and dynamically adjusting the identification strategy using local outliers and machine learning models, the problem of identification response lag in the dynamic attacks of the RIS phased array network is solved, and the recognition accuracy of potential secondary attacks and the accuracy of terminal identification is improved.
Patent Information
- Application Number
- CN202510580248.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When existing RIS phased array network deployment terminals encounter dynamic attacks, due to the lack of accurate identification of potential secondary anomaly attacks during the authentication response lag period, the identification of the degree of initial attacks in the terminal is biased.
By monitoring the target identification terminals in the RIS phased array system that are subject to initial network attacks, collecting the initial feature set and monitoring abnormal attack behaviors within the response lag time period, building a local outlier method to screen out effective attack behaviors, using machine learning models for mixed clustering calculations, setting a dynamic threshold strategy for identification and determination, ensuring that influential attack behaviors are re-identified.
It enhances the perception of dynamic attacks, avoids blind spots in the system's response lag, improves the recognition accuracy of potential secondary attacks, and ensures accurate identification of terminals.
Smart Images

Figure CN120455069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of terminal detection, and in particular to a device terminal identification method for RIS phased array network deployment. Background Art
[0002] Traditional terminal authentication methods rely primarily on static information, such as MAC addresses or user credentials. However, these methods are easily forged or misused, making them difficult to accurately identify access terminals in high-security scenarios. To enhance physical-layer perception of terminal behavior, recent research has introduced sensing devices such as Reconfigurable Intelligent Surface (RIS) phased arrays to improve wireless propagation environments and enhance the spatial granularity of signal feature recognition. RIS, with its controllable reflection and directional transmission properties, can enhance channel state perception in physical space, contributing to improved interpretability and traceability of terminal behavior.
[0003] In the deployment scenario of RIS phased arrays, the existing device authentication mechanism for WAPI-converged networks accessing RIS phased arrays primarily identifies and manages terminals based on static identity information and single-point behavioral data, without establishing a time series model for the changing behavior of terminals within the network. In practice, malicious terminals often do not maintain a persistent abnormal state, but instead conceal their attack intent through intermittent triggering, phased penetration, or delayed execution. When existing systems identify potential time-series anomalies, they often require a long response lag to determine the specific location of the abnormal data point before generating a specific identification assessment result. This response lag may allow attackers ample time to conduct a secondary, in-depth attack, resulting in biased identification results prior to the secondary, in-depth attack. This results in low accuracy in identifying previous attack results when screening digital terminal devices vulnerable to complex attacks. Summary of the Invention
[0004] The present invention provides a device terminal identification method for RIS phased array network deployment, which solves the problem that when existing RIS phased array network deployment terminals encounter dynamic attacks, they lack accurate identification of potential secondary abnormal attacks during the identification response lag period, resulting in deviation in the identification of the initial attack severity of the terminal.
[0005] The present invention is achieved through the following technical solutions: A device terminal identification method for RIS phased array network deployment, the method comprising: Step S1: monitoring a target identification terminal in the RIS phased array system that is subjected to an initial network attack, collecting an initial feature set in the target identification terminal, and monitoring abnormal attack behaviors suffered by the target identification terminal within a response lag period for identifying the initial network attack; Step S2: Construct a local outlier method to screen out effective attack behaviors from abnormal attack behaviors according to their effectiveness, extract the initial feature set of the target identification terminal corresponding to the effective attack behavior when it is effective, and mark it as a composite feature set; Step S3: Construct a machine learning model, use the machine learning model to perform mixed clustering calculations on the initial feature set and the composite feature set according to the effective attack behavior, and annotate the final generated result as a correlation scalar representing the target identification terminal feature; Step S4: A dynamic threshold policy is set to identify and determine the correlation scalar. The valid attack behavior corresponding to the correlation scalar that does not comply with the dynamic threshold policy is determined to have no valid association with the initial network attack, and the target identification terminal is determined to not need to be re-identified. The valid attack behavior corresponding to the correlation scalar that complies with the dynamic threshold policy is determined to have a valid association with the initial network attack, and the target identification terminal is determined to need to be re-identified.
[0006] Because identifying the specific location of abnormal data points often requires a long response lag time to evaluate and produce a specific identification assessment result, this response lag time may leave attackers ample time to launch a secondary, in-depth attack, resulting in identification bias in the identification results before the secondary, in-depth attack. This results in a low accuracy rate for the identification of previous attack results when screening digital terminal devices vulnerable to complex attacks. Based on this, the present invention provides a device terminal identification method for RIS phased array network deployment, which solves the problem that existing WAPI converged network devices, when encountering dynamic attacks, lack accurate identification of potential secondary abnormal attacks during the identification response lag time period, resulting in a bias in the identification of the terminal's initial attack severity.
[0007] Furthermore, the local outlier method includes a LOF outlier factor method; and the process of screening out effective attack behaviors using the local outlier method includes: Normalize all abnormal data points caused by abnormal attack behaviors and set each abnormal data point as an independent sample point. Set the number of neighbors used in local density calculation, calculate the neighbor set corresponding to the number of neighbors for each sample point, and use it to calculate the local density of the sample point. Calculate the LOF value of each sample point based on the local density. A first density threshold is set for the LOF value. When the LOF value of the sample point is greater than the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point is an effective attack behavior and has an impact on the initial network attack; when the LOF value of the sample point is less than or equal to the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point has no effective impact on the initial network attack.
[0008] Furthermore, the initial feature set is constructed based on the collection of RIS feature signals of the target identification terminal under the RIS phased array reflection path; the RIS feature signals are annotated by the spatial channel variation features captured by the RIS enhancement path.
[0009] Furthermore, the calculation process of the LOF value is set as follows: Let the LOF value be the ratio of the average local density of the neighboring samples to the local density of the target sample point; Among them, the preset number of neighbors is represented as f, and the distance value of the neighbor sample with the ordinal number f from the target sample point from near to far in the neighbor sample is d, and the set of sample points in the neighbor sample whose distance to the target sample point is less than or equal to d is represented as the neighbor point set. The average local density of the neighbor sample is the average local density of all neighbor point sets of the target sample point.
[0010] Furthermore, a second density threshold is set for the LOF value, and the first density threshold is greater than the second density threshold; when the LOF value of the sample point is less than or equal to the second density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has no effect on the initial network attack; when the LOF value of the sample point is between the second density threshold and the first density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has a potential impact on the target identification terminal, and the abnormal attack behavior corresponding to the current sample point is marked as a potential attack behavior, and defensive repairs are performed on the abnormal data points caused by the potential attack behavior.
[0011] Furthermore, the process of using the machine learning model to perform hybrid clustering calculations includes: Step A1: The initial feature set and the composite feature set are weighted and merged according to similar data to form a multidimensional feature matrix containing multiple multidimensional data related to effective attack behaviors, and the data points in the multidimensional feature matrix are normalized; Step A2: Preset the range of cluster numbers and the number of iterations, use the K-means clustering method to calculate the cluster centers, randomly generate a set of initial centroids, assign each data point to the nearest centroid, and then update the centroid of each cluster along the number of iterations; Step A3: After the iteration is completed, the cluster with the largest distance between cluster centroids is marked as an abnormal behavior cluster with potential effective abnormal behavior, and the data samples corresponding to the abnormal behavior cluster are marked as correlation scalars. The correlation scalars are denormalized and then output.
[0012] Furthermore, the updating calculation process of the centroid includes: Let the center of mass be μ, let x i represents the i-th data point, the ω(x i ) represents the data point x i The weight of the cluster is expressed as k, and C k Denotes the set of all data points in the kth cluster, let |C k | represents the total number of data points, Then the centroid μ of the kth cluster k The calculation formula is expressed as: .
[0013] Furthermore, the denormalization process uses Min-Max denormalization to process each data point, and the denormalized feature set is compared with the historical behavior data of the target identification terminal along the time series points for re-evaluation.
[0014] Furthermore, the dynamic threshold strategy includes: setting a correlation threshold range for the correlation scalar, and setting a sliding window mechanism to monitor the historical judgment error in real time; The definition of the sliding window mechanism is as follows: a target identification rate indicating that a result requires re-identification is preset, and an initial time window of fixed length is preset; the initial time window records the determination results of each target identification terminal, including whether re-identification is required and whether re-identification is not required; each time the sliding window is updated, the actual identification rate indicating that a result requires re-identification within the initial time window is counted; When the actual recognition rate is higher than the target recognition rate, the range of the correlation threshold is narrowed; when the actual recognition rate is lower than the target recognition rate, the range of the correlation threshold is expanded.
[0015] Furthermore, an identification frequency reference value is set for the length of the initial time window; when the number of identifications per unit time reaches the identification frequency reference value, the initial time window is updated.
[0016] Furthermore, when the number of identifications per unit time reaches a reference value of identification frequency, the target identification rate is dynamically adjusted to minimize the total number of misjudgments.
[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: The behavior of terminals deployed in RIS phased array networks after being attacked is usually intermittent or periodic. The present invention continues to monitor potential abnormal attack behaviors during the initial attack identification response lag period, capturing the dynamic changes of target identification terminals in real time, enhancing the perception of dynamic attacks and avoiding the system's blind spots after the initial attack. By introducing the local outlier detection method, "effective attacks" directly related to the attack behavior are further screened out after the initial attack, and irrelevant or noise data are excluded, avoiding deviations caused by irrelevant abnormal interference, ensuring that only influential attack behaviors are re-identified, and improving the recognition accuracy of potential secondary attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings: Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with examples and drawings. The exemplary embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention. Example
[0020] like Figure 1 As shown, this embodiment is a device terminal identification method for RIS phased array network deployment, the method comprising: Step S1: monitoring a target identification terminal in the RIS phased array system that is subjected to an initial network attack, collecting an initial feature set in the target identification terminal, and monitoring abnormal attack behaviors suffered by the target identification terminal within a response lag period for identifying the initial network attack; Step S2: Construct a local outlier method to screen out effective attack behaviors from abnormal attack behaviors according to their effectiveness, extract the initial feature set of the target identification terminal corresponding to the effective attack behavior when it is effective, and mark it as a composite feature set; Step S3: Construct a machine learning model, use the machine learning model to perform mixed clustering calculations on the initial feature set and the composite feature set according to the effective attack behavior, and annotate the final generated result as a correlation scalar representing the target identification terminal feature; Step S4: A dynamic threshold policy is set to identify and determine the correlation scalar. The valid attack behavior corresponding to the correlation scalar that does not comply with the dynamic threshold policy is determined to have no valid association with the initial network attack, and the target identification terminal is determined to not need to be re-identified. The valid attack behavior corresponding to the correlation scalar that complies with the dynamic threshold policy is determined to have a valid association with the initial network attack, and the target identification terminal is determined to need to be re-identified.
[0021] The RIS is an intelligent reflective surface composed of a large number of controllable metasurface units. A phased array traditionally refers to an antenna system that directionally controls the beam direction by adjusting the phase difference between the transmitted or received signals of multiple antennas. In specific implementations, the RIS phased array can refer to a RIS device as a passive phased array or an intelligent reflective phased array. This system, based on reconfigurable smart surface technology, simulates and extends the beam steering capabilities of a traditional phased array by intelligently controlling the reflective properties of a large number of units. A target identification terminal in the RIS phased array system refers to a terminal device selected by the system as requiring identity authentication or behavior analysis in a wireless network integrated with the RIS reflectarray system. The terminal's communication behavior and channel characteristics are controlled by the RIS system and used as identification. An initial network attack refers to the first or early network attack behavior a terminal has just encountered, including but not limited to unauthorized access, man-in-the-middle attacks, attempts to tamper with communication protocols, and server authentication fraud. These attacks are typically exploratory and tentative, and the attacker may not immediately reveal their full attack intent, but rather gradually infiltrate. At this stage, the system has detected or suspected an attack, but is still in a response lag period, requiring further confirmation and evaluation. The initial feature set represents the state information of the target authentication terminal at the initial time of the test. It refers to the terminal's various features or attribute data recorded by the system immediately when the terminal first experiences the initial network attack. In practical applications, this initial feature set can include data such as network layer information, encryption and authentication information, communication behavior characteristics, and security event logs. The authentication response lag period refers to the processing time required by the system after detecting an attack to analyze the attack source, attack methods, and verify abnormal behavior. In actual systems, this process is typically not instantaneous and may take tens of milliseconds, several seconds, or even longer. This period is called the "response lag time." Traditional systems often suspend monitoring after the initial attack is detected, awaiting authentication results before deciding on subsequent actions. This allows attackers to further infiltrate or hide during this lag period. During this response lag period, the terminal may continue to experience new abnormal behavior. During this lag period, the system needs to monitor in real time to observe whether new abnormal attack behavior has emerged, rather than waiting for the initial authentication to be fully completed before re-analyzing.
[0022] The local outlier method is used to screen out attacks that actually impact system security and eliminate those that are irrelevant or have minimal impact. All abnormal attacks are evaluated based on their actual threat to system security, retaining only those that are truly effective and threatening. This completes the initial screening of abnormal attacks. The local outlier method identifies points in a dataset that are significantly different from other data points (i.e., outliers). In attack behavior detection, these outliers typically represent potential security threats. Local outliers refer not only to globally abnormal data, but also to data that is abnormal relative to its surroundings or neighboring data. In other words, a data point may not be conspicuous in the entire dataset, but its difference from its neighboring data may be significant enough to be considered an anomaly. Abnormal attack behavior refers to terminal behavior that is inconsistent with normal network behavior, typically manifesting as sudden traffic surges, irregular communication patterns, or unauthorized access, and is a manifestation of part or all of an attack. Effective attack behavior refers to attack behavior that can actually impact network security and poses a substantial threat to the target system, network, or terminal. The composite feature set refers to a collection of annotated feature data sets extracted at the time when effective attacks take effect. It is a dynamic extension of the original feature set, which includes the actual changes of the terminal under the influence of the attack. As a specific implementation method, for example, assuming that the initial characteristics of the target terminal are normal signal strength, stable data packet frequency, normal CPU usage and stable connection, this is the initial feature set; later, an anomaly is detected during the lag time, and a pre-penetration probing attack behavior occurs, such as the terminal starts to send a large number of small packets with irregular intervals, and the CPU usage suddenly increases. After local outlier analysis, this is determined to be an effective attack behavior. At this time, the terminal's signal changes, packet frequency changes, and CPU load anomalies and other status data will be extracted and marked as a composite feature set. The parameter category of the composite feature set is consistent with that of the initial feature set.
[0023] Clustering is a common unsupervised machine learning method that groups samples in a dataset based on certain similarities or proximity. In this embodiment, the goal of clustering is to classify terminal behavioral features based on valid attack behaviors. Clustering can identify which behavioral patterns remain consistent between the initial attack and subsequent attacks, and which may be triggered by new attacks. During the clustering process, the model comprehensively processes data from both feature sets and analyzes the similarities and correlations between them to identify potential attack behaviors or emerging abnormal patterns. The correlation scalar is a numerical value that indicates the similarity or strength of the relationship between one of the current parameter characteristics of the target identification terminal, such as behavioral pattern or attack manifestation, and the initial network attack characteristics. The specific parameter being measured is not necessary here; the correlation scalar serves only to measure whether valid attack behaviors have a substantial impact on the initial network attack identification results. If the correlation scalar value is high, it indicates that the currently monitored characteristics are highly correlated with previously recorded attack characteristics, which may indicate that the terminal is experiencing further attacks and requires re-identification. If the value is low, it indicates that the current attack behavior is weakly correlated with the initial attack signature, possibly indicating that the attack has ceased or been mitigated, and therefore, the terminal does not need to be re-authenticated. The correlation scalar can be considered the output of the machine learning model. It is a comprehensive assessment of the target terminal's status based on the calculations from the clustering algorithm, feature extraction, and learning process. The dynamic threshold strategy automatically adjusts a threshold used to determine whether an attack is effective based on changes in the attack situation. This threshold is not fixed; it adjusts dynamically based on changes in certain factors. If the correlation scalar value is lower than the dynamic threshold, it indicates that the subsequent attack behavior is weakly correlated with the initial attack, indicating that the attack may have ended or that no further threats exist. The attack behavior is then determined to have no valid correlation with the initial attack, and the target terminal does not need to be re-authenticated. The system deems the terminal to be unaware of any further threats and can continue to trust the terminal. If the value of the correlation scalar is higher than the dynamic threshold, it indicates that there is a strong correlation between the subsequent attack behavior and the initial attack, which may indicate that the attack has not been completely completed or the attacker is conducting more complex infiltration. In this case, the attack behavior is determined to be effectively correlated with the initial attack. The system considers that the security of the terminal is threatened and needs to be re-authenticated. This may trigger stricter security measures or a review of the identity and behavior of the terminal.
[0024] Furthermore, as a feasible implementation method, the initial feature set is constructed based on the collection of RIS feature signals of the target identification terminal under the RIS phased array reflection path; the RIS feature signals are annotated with spatial channel variation features captured by the RIS enhancement path.
[0025] The RIS phased array reflection path refers to the path a signal takes from the transmitter, through reflection from the RIS system, to the receiving terminal. By controlling the configuration of the RIS system, signal propagation characteristics, such as reflection angle, attenuation, and delay, can be modified. This variation in the RIS signal's reflection path provides each target terminal with a unique signal signature, containing information about the target terminal's specific location and environment. Traditional signature collection methods typically rely on terminal identity information, such as MAC address, IP address, or behavior logs. These features rely more heavily on the data link and network layers. By using the signal signatures along the RIS reflection path, more detailed and dynamic terminal behavior information can be obtained from the physical layer. These signal signatures are more spatially diverse and less susceptible to spoofing or tampering, thereby enhancing accurate identification of the target terminal. The RIS signature signal refers to a specific signal generated and controlled by the RIS reflectarray, encompassing characteristics such as reflection angle, signal strength, propagation path, and arrival time. Unlike traditional signals, the RIS signature signal possesses unique spatial properties that reflect the influence of the RIS array on the signal. The RIS enhanced path refers to the signal path formed by reflection from the RIS system. These paths may exhibit different propagation characteristics than traditional signal transmission paths. In practical applications, the RIS system analyzes the signal data and spatial channel variation characteristics collected by the system, classifying and labeling them based on the temporal and spatial variations of these signal characteristics. This implementation improves the accuracy of terminal identification through physical layer signal characteristics, enabling the system to more efficiently identify abnormal behavior in complex network environments.
[0026] Example 2 In this embodiment, the local outlier method includes the LOF outlier factor method; the process of screening out effective attack behaviors using the local outlier method includes: Normalize all abnormal data points caused by abnormal attack behaviors and set each abnormal data point as an independent sample point. Set the number of neighbors used in local density calculation, calculate the neighbor set corresponding to the number of neighbors for each sample point, and use it to calculate the local density of the sample point. Calculate the LOF value of each sample point based on the local density. A first density threshold is set for the LOF value. When the LOF value of the sample point is greater than the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point is an effective attack behavior and has an impact on the initial network attack; when the LOF value of the sample point is less than or equal to the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point has no effective impact on the initial network attack.
[0027] The LOF outlier factor method measures the degree of abnormality of each point by comparing its local density with the local density of its neighbors. Before outlier detection, all data points related to abnormal attack behaviors must first be normalized. This step eliminates the influence of data with different dimensions or scales, ensuring that all data points are compared under the same standard. Treating each abnormal data point as an independent sample point facilitates the calculation of its local density and relative abnormality. The number of neighbors considered when calculating the local density, namely the neighborhood range of each data point, is set. This neighborhood range influences the density calculation results; it determines how data points within the local range affect the density calculation of the target data point. For each sample point, the density of its neighbors is calculated, that is, the distribution of the neighbors is calculated. If the density within a point's neighborhood is high, it indicates that the point is likely within the normal range; conversely, if the neighborhood density is low, the point is likely an outlier. The first density threshold is used to determine whether an abnormal data point represents a valid attack behavior. If the LOF value of a data point is greater than the first density threshold, the behavior of that point is anomalous and significant within its neighborhood. Therefore, it is considered to have contributed to the initial attack and is considered a valid attack. If the LOF value of a data point is less than or equal to the threshold, it indicates that the data point is similar to other points in the neighborhood and may be noise or non-aggressive behavior. Therefore, the anomalous behavior is considered to have had no significant impact on the initial attack. Through the above process, the solution can effectively distinguish which anomalous behaviors still affect the attack outcome after the initial attack, and thus re-identify the terminal.
[0028] Furthermore, as a feasible implementation method, the calculation process of the LOF value is set as follows: Let the LOF value be the ratio of the average local density of the neighboring samples to the local density of the target sample point; Among them, the preset number of neighbors is represented as f, and the distance value of the neighbor sample with the ordinal number f from the target sample point from near to far in the neighbor sample is d, and the set of sample points in the neighbor sample whose distance to the target sample point is less than or equal to d is represented as the neighbor point set. The average local density of the neighbor sample is the average local density of all neighbor point sets of the target sample point.
[0029] The ratio of the LOF values reflects whether the target point is in an area with a significantly lower density than its neighbors, thereby determining whether the point is an outlier. The preset number of neighbors f defines the number of neighbors of the target sample point that needs to be considered when calculating the LOF value. A neighbor point refers to a sample point that is closest to the target point in spatial distance. Set the neighbor sample distance of f to d, which represents the distance between the target sample point and the fth neighbor point. This distance value is usually used to determine the size of the neighbor sample set. The target sample point and the sample points with a distance less than or equal to d from it form a neighbor point set. This set contains the target sample point and its f closest sample points, and is the basis for calculating the local density of the target sample point. Calculate the average local density of the neighbor samples of the target sample point, that is, calculate the local density of all neighbor point sets of the target sample point and obtain the average value of the set. The local density is usually the inverse of the distance between the target point and its neighbors, or a density value weighted by the distance.
[0030] Furthermore, as a feasible implementation method, a second density threshold is set for the LOF value, and the first density threshold is greater than the second density threshold; when the LOF value of the sample point is less than or equal to the second density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has no effect on the initial network attack; when the LOF value of the sample point is between the second density threshold and the first density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has a potential impact on the target identification terminal, and the abnormal attack behavior corresponding to the current sample point is marked as a potential attack behavior, and defensive repairs are performed on the abnormal data points caused by the potential attack behavior.
[0031] The second density threshold is used to further refine the identification of abnormal attack behavior. The first density threshold is larger than the second density threshold and is used to determine whether an abnormal attack behavior has effectively impacted the initial network attack. The second density threshold provides a criterion for identifying potential attack behavior. This threshold helps identify attack behaviors that have not yet posed a direct threat but may subsequently develop into a threat. If the LOF value of a sample point is less than or equal to the second density threshold, this indicates that the abnormal behavior at that data point has no impact on the initial network attack and may even be noise or non-malicious behavior. In this case, the system determines that the abnormal attack behavior corresponding to this sample point has no impact on the initial attack and therefore does not perform further processing or identification. If the LOF value of a sample point is between the second density threshold and the first density threshold, the abnormal behavior corresponding to this point is identified as potential attack behavior. Potential attack behavior refers to behavior that may not pose an immediate threat but may develop into a real attack in the future or be persistent malicious behavior. The system will mark this as a potential attack behavior and implement defensive remediation measures. Such remediation measures can include encrypted communications, enhanced authentication, and more frequent monitoring, thereby proactively preventing potential attack risks. For abnormal data points identified as potential attack behaviors, the system no longer waits for them to escalate into serious attacks before taking action. Instead, it implements proactive, defensive remediation, enabling the system to respond to potential attacks in their early stages and prevent them from escalating. This helps mitigate future attack risks for intermittent attacks or gradual penetration attacks, such as those that develop gradually. In specific implementations, since the LOF value is defined as the ratio of the average local density of neighboring samples to the local density of the target sample point, the second density threshold can be set to 1.1 or 1.2. When the LOF value is less than 1, it means that the local density of the point is higher than that of its neighbors, indicating that the point is a very normal and stable point, possibly even the center or core sample of a cluster. Setting the second density threshold slightly above 1 creates redundancy, allowing for slight density variations. Sample points with LOF values greater than 1.1 or 1.2 are considered likely to be potential attack behaviors. This effectively screens out potential threats and avoids overly strict or lenient judgments. The first density threshold may be set according to an empirical rule, for example, to a value between 1.5 and 2.0.
[0032] Example 3 In this embodiment, the process of performing hybrid clustering calculation using a machine learning model includes: Step A1: The initial feature set and the composite feature set are weighted and merged according to similar data to form a multidimensional feature matrix containing multiple multidimensional data related to effective attack behaviors, and the data points in the multidimensional feature matrix are normalized; Step A2: Preset the range of cluster numbers and the number of iterations, use the K-means clustering method to calculate the cluster centers, randomly generate a set of initial centroids, assign each data point to the nearest centroid, and then update the centroid of each cluster along the number of iterations; Step A3: After the iteration is completed, the cluster with the largest distance between cluster centroids is marked as an abnormal behavior cluster with potential effective abnormal behavior, and the data samples corresponding to the abnormal behavior cluster are marked as correlation scalars. The correlation scalars are denormalized and then output.
[0033] The data points in the initial feature set and the composite feature set are weighted and merged to form a multidimensional feature matrix. This matrix contains multidimensional data related to valid attack behaviors. Different feature sets may have different importance. For example, the initial feature set may represent basic information about network access, while the composite feature set may represent features caused by attack behaviors. By merging these data points, the multidimensional features related to attack behaviors can be better captured. To eliminate the influence of different feature scales, the data in the multidimensional feature matrix needs to be normalized. The cluster number range refers to the number of clusters you want the data to be divided into before performing K-means clustering. The core goal of the K-means clustering method is to divide data points into a specified number of clusters, each with a center, called the cluster centroid. The goal of clustering is to make data points within a cluster as similar as possible and to minimize the differences between clusters. When starting K-means clustering, several data points are randomly selected as the initial cluster centroids. These initial centroids are randomly generated, so the positions of the initial cluster centroids may vary between each K-means clustering run.
[0034] Furthermore, after the initial centroids are selected, K-means clustering calculates the distance between each data point and each centroid. Each data point is assigned to the cluster with the centroid closest to it. This process is called cluster assignment. After the data points are assigned, the centroids are recalculated: the centroid of each cluster needs to be recalculated. During recalculation, the feature values of all data points in the cluster are averaged to obtain a new centroid. This process continues, that is, after each centroid update, the data points are reallocated according to the new centroid. This process is iterated repeatedly until the number of iterations is exhausted. The cluster centroid distance refers to the distance between the centroids of different clusters. Clusters with large centroid distances usually mean that the behavior of these clusters is very different. In K-means clustering, the distance between the centroids of different clusters is usually calculated to determine whether they belong to completely different categories. In this embodiment, if the centroid of a cluster is far away from the centroid of other clusters and the distance reaches the maximum, it means that the behavior of this cluster is very different from that of other clusters, and this cluster is identified as an abnormal behavior cluster associated with effective attack behavior. Data samples within a cluster are labeled with a correlation scalar. During the clustering process, the size of the correlation scalar indicates whether the sample point is related to potential attack behavior. The larger the correlation, the more likely the point is to be part of an attack.
[0035] Among them, as a feasible implementation method, the denormalization process uses Min-Max denormalization to process each data point, and compares the denormalized feature set with the historical behavior data of the target identification terminal along the time series points for re-evaluation.
[0036] The Min-Max denormalization is a common method in data preprocessing, which restores the normalized data to its original range by the following formula: norm ×(max(x)-min(x))+min(x), where x represents the original data, max(x) and min(x) represent the maximum and minimum values of the original data, and x norm Represents the normalized value, which returns the denormalized data to its original scale. Time series data is arranged in chronological order and records the historical behavior of the target identification terminal. Time series analysis can be used to identify regularities and changing trends in terminal behavior. Comparing denormalized data with historical behavior data aims to assess the consistency and differences between current and historical behavior. By comparing denormalized data with historical behavior data, the system can reassess the behavior of the target terminal and issue early warnings if any significant abnormal behavior is observed.
[0037] Furthermore, as a feasible implementation method, the center of mass update calculation process includes: Let the center of mass be μ, let x irepresents the i-th data point, the ω(x i ) represents the data point x i The weight of the cluster is expressed as k, and C k Denotes the set of all data points in the kth cluster, let |C k | represents the total number of data points, Then the centroid μ of the kth cluster k The calculation formula is expressed as: .
[0038] The center of mass μ k represents the center of all data points in cluster k. It is the weighted average of all data points in the cluster and is used to redistribute data points according to their distance from the cluster center. The goal of centroid calculation is to make the data points in the cluster as close to the center of the cluster as possible and minimize the square error within the cluster. The weight of each data point ω(x i ) reflects the weight of the data point in the cluster centroid calculation. Data points with high weights have a greater impact on the cluster centroid position, while data points with low weights have a smaller impact on the centroid. k | represents the total number of elements, 1 / |C k | represents a normalization factor, which is used to normalize the weighted sum. xi ∙ω(x i )∙x i represents the weighted summation of all data points in cluster k; the ω(x i )⋅x i The summation process accumulates the weighted contributions of all data points within the cluster. The calculated centroid is a weighted average, ensuring that different data points have different influences on the cluster's location when calculating the cluster's center.
[0039] Example 4 The dynamic threshold strategy includes: setting a correlation threshold range for the correlation scalar, and setting a sliding window mechanism to monitor the historical judgment error in real time; The definition of the sliding window mechanism is as follows: a target identification rate indicating that a result requires re-identification is preset, and an initial time window of fixed length is preset; the initial time window records the determination results of each target identification terminal, including whether re-identification is required and whether re-identification is not required; each time the sliding window is updated, the actual identification rate indicating that a result requires re-identification within the initial time window is counted; When the actual recognition rate is higher than the target recognition rate, the range of the correlation threshold is narrowed; when the actual recognition rate is lower than the target recognition rate, the range of the correlation threshold is expanded.
[0040] The correlation threshold range is used to determine whether there is a valid correlation between the effective attack behavior and the initial network attack. The target identification rate represents the expected accuracy or frequency of re-identification. This target value reflects the target ratio that the system should re-identify under ideal circumstances. The initial time window is a time window of fixed length, which is used to record the judgment results of each target identification terminal, including situations where re-identification is not required and situations where re-identification is required. The initial time window is used to collect the identification results within a certain period of time for subsequent judgment and adjustment. Whenever the time window slides, the actual identification rate that needs to be re-identified is represented in the statistical window. The actual identification rate refers to the ratio of the number of terminals that actually need to be re-identified to the total number of judgments during the current window period.
[0041] If the number of terminals actually requiring re-identification within the window is too high, exceeding the target identification rate, this indicates that the system may be over-identifying. To reduce misjudgments and over-identification, the system will narrow the correlation threshold range, increasing its sensitivity to the correlation scalar and reducing unnecessary re-identification. If the number of terminals actually requiring re-identification within the window is lower than the target identification rate, this indicates that the system's identification accuracy may be insufficient, preventing it from promptly identifying terminals requiring re-identification. To improve the identification rate, the system will expand the correlation threshold range, reducing the requirement for the correlation scalar, enabling more sensitive identification of terminals requiring re-identification. Using a sliding window mechanism, the system can evaluate and adjust the threshold range in real time, allowing the dynamic threshold strategy to adapt to actual identification situations. This allows the system to adapt to changes and effectively respond to diverse attack scenarios.
[0042] Furthermore, as a feasible implementation method, an identification frequency benchmark value is set for the length of the initial time window; when the number of identifications per unit time reaches the identification frequency benchmark value, the initial time window is updated; when the number of identifications per unit time reaches the identification frequency benchmark value, the target identification rate is dynamically adjusted to minimize the total number of false positives.
[0043] The identification frequency benchmark value is a preset value, which represents the minimum threshold number of target identifications per unit time. This benchmark value is used to determine when the initial time window needs to be updated and the system needs to be adjusted. The benchmark value ensures that the system monitors and evaluates the identification of the target terminal frequently enough. When the number of identifications per unit time reaches the preset benchmark value, it means that the system has enough data to evaluate the current identification accuracy. At this time, the initial time window will be updated, that is, the content of the window will slide to the latest data, so that the system can make new evaluations and adjustments based on the current identification results. The identification frequency benchmark value ensures the system's response speed to changes, avoids long delays, and ensures that the window evaluation is always up to date, so that it can better adapt to dynamic changes in the network.
[0044] The target identification rate represents the expected frequency of re-identification, and the system uses this value to set the expected identification result. When the number of identifications per unit time reaches a baseline value, the system adjusts the target identification rate based on the actual identification situation, such as false positives and missed detections. The purpose of adjusting the target identification rate is to minimize the total number of false positives. If the actual re-identification rate is too high, it may indicate that the threshold is too loose, resulting in an increase in false positives. In this case, the target identification rate will be reduced to reduce the frequency of re-identification. If the actual re-identification rate is too low, it may indicate that the threshold is too strict, resulting in an increase in missed detections. In this case, the target identification rate will be increased to increase the frequency of re-identification. The sliding window mechanism and dynamic adjustment strategy enable the system to update strategies based on real-time identification data, reduce false positives and missed detections, and improve the security of the entire network.
[0045] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A device terminal identification method for RIS phased array network deployment, characterized in that: The method includes: Step S1: monitoring a target identification terminal in the RIS phased array system that is subjected to an initial network attack, collecting an initial feature set in the target identification terminal, and monitoring abnormal attack behaviors suffered by the target identification terminal within a response lag period for identifying the initial network attack; Step S2: Construct a local outlier method to screen out effective attack behaviors from abnormal attack behaviors according to their effectiveness, extract the initial feature set of the target identification terminal corresponding to the effective attack behavior when it is effective, and mark it as a composite feature set; Step S3: Construct a machine learning model, use the machine learning model to perform mixed clustering calculations on the initial feature set and the composite feature set according to the effective attack behavior, and annotate the final generated result as a correlation scalar representing the target identification terminal feature; Step S4: A dynamic threshold policy is set to identify and determine the correlation scalar. The valid attack behavior corresponding to the correlation scalar that does not comply with the dynamic threshold policy is determined to have no valid association with the initial network attack, and the target identification terminal is determined to not need to be re-identified. The valid attack behavior corresponding to the correlation scalar that complies with the dynamic threshold policy is determined to have a valid association with the initial network attack, and the target identification terminal is determined to need to be re-identified.
2. The device terminal identification method for RIS phased array network deployment according to claim 1 is characterized in that: The local outlier method includes the LOF outlier factor method; The process of screening out effective attack behaviors using the local outlier method includes: Normalize all abnormal data points caused by abnormal attack behaviors and set each abnormal data point as an independent sample point; Set the number of neighbors used in local density calculation, calculate the neighbor set corresponding to the number of neighbors for each sample point and use it to calculate the local density of the sample point, and calculate the LOF value of each sample point based on the local density; A first density threshold is set for the LOF value. When the LOF value of the sample point is greater than the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point is an effective attack behavior and has an impact on the initial network attack; when the LOF value of the sample point is less than or equal to the first density threshold, it is judged that the abnormal attack behavior corresponding to the current sample point has no effective impact on the initial network attack.
3. The device terminal identification method for RIS phased array network deployment according to claim 2 is characterized in that: The calculation process of the LOF value is set as: Let the LOF value be the ratio of the average local density of the neighboring samples to the local density of the target sample point; Among them, the preset number of neighbors is represented as f, and the distance value of the neighbor sample with the ordinal number f from the target sample point from near to far in the neighbor sample is d, and the set of sample points in the neighbor sample whose distance to the target sample point is less than or equal to d is represented as the neighbor point set. The average local density of the neighbor sample is the average local density of all neighbor point sets of the target sample point.
4. The device terminal identification method for RIS phased array network deployment according to claim 3 is characterized in that: A second density threshold is also set for the LOF value, and the first density threshold is greater than the second density threshold; when the LOF value of the sample point is less than or equal to the second density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has no effect on the initial network attack; when the LOF value of the sample point is between the second density threshold and the first density threshold, it is determined that the abnormal attack behavior corresponding to the current sample point has a potential impact on the target identification terminal, and the abnormal attack behavior corresponding to the current sample point is marked as a potential attack behavior, and defensive repairs are performed on the abnormal data points caused by the potential attack behavior.
5. The device terminal identification method for RIS phased array network deployment according to claim 1 is characterized in that: The initial feature set is constructed based on the collection of RIS feature signals of the target identification terminal under the RIS phased array reflection path; the RIS feature signals are annotated with spatial channel variation features captured by the RIS enhancement path.
6. The device terminal identification method for RIS phased array network deployment according to claim 1 is characterized in that: The process of performing hybrid clustering calculations using machine learning models includes: Step A1: The initial feature set and the composite feature set are weighted and merged according to similar data to form a multidimensional feature matrix containing multiple multidimensional data related to effective attack behaviors, and the data points in the multidimensional feature matrix are normalized; Step A2: Preset the range of cluster numbers and the number of iterations, use the K-means clustering method to calculate the cluster centers, randomly generate a set of initial centroids, assign each data point to the nearest centroid, and then update the centroid of each cluster along the number of iterations; Step A3: After the iteration is completed, the cluster with the largest distance between cluster centroids is marked as an abnormal behavior cluster with potential effective abnormal behavior, and the data samples corresponding to the abnormal behavior cluster are marked as correlation scalars. The correlation scalars are denormalized and then output.
7. The device terminal identification method for RIS phased array network deployment according to claim 6 is characterized in that: The updating calculation process of the centroid includes: Let the center of mass be μ, let x i represents the i-th data point, the ω(x i ) represents the data point x i The weight of the cluster is expressed as k, and C k Denotes the set of all data points in the kth cluster, let |C k | represents the total number of data points, Then the centroid μ of the kth cluster k The calculation formula is expressed as: .
8. The device terminal identification method for RIS phased array network deployment according to claim 6 is characterized in that: The denormalization process uses Min-Max denormalization to process each data point, and compares the denormalized feature set with the historical behavior data of the target identification terminal along the time series points for re-evaluation.
9. The device terminal identification method for RIS phased array network deployment according to claim 1 is characterized in that: The dynamic threshold strategy includes: setting a correlation threshold range for the correlation scalar, and setting a sliding window mechanism to monitor the historical judgment error in real time; The definition of the sliding window mechanism is as follows: a target identification rate indicating that a result requires re-identification is preset, and an initial time window of fixed length is preset; the initial time window records the determination results of each target identification terminal, including whether re-identification is required and whether re-identification is not required; each time the sliding window is updated, the actual identification rate indicating that a result requires re-identification within the initial time window is counted; When the actual recognition rate is higher than the target recognition rate, the range of the correlation threshold is narrowed; when the actual recognition rate is lower than the target recognition rate, the range of the correlation threshold is expanded.
10. The device terminal identification method for RIS phased array network deployment according to claim 9 is characterized in that: A reference value for identification frequency is set for the length of the initial time window; when the number of identifications per unit time reaches the reference value for identification frequency, the initial time window is updated; When the number of identifications per unit time reaches a reference value of the identification frequency, the target identification rate is dynamically adjusted to minimize the total number of misjudgments.
Citation Information
Patent Citations
Network traffic anomaly detection method and device
CN109067725A
A low-speed denial of service attack detection method based on an SNN-LOF algorithm
CN109726553A
Network attack determination method and device, storage medium and electronic device
CN118353683A
AU2021102263A4