A cloud-based collaborative intrusion anomaly behavior detection method and system
By receiving and analyzing data packets from edge computing nodes and combining them with internal operating status information to perform context-aware judgment, the problem of high false alarm rate in the industrial internet has been solved, enabling accurate differentiation of data flow anomalies and improving the accuracy and security of detection.
Patent Information
- Application Number
- CN202610638742.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-14
AI Technical Summary
In the industrial internet environment, existing technologies struggle to effectively distinguish between data flow anomalies caused by internal factors of edge computing nodes and external attacks, leading to misjudgments by cloud-based detection systems, increasing false alarm rates, and hindering effective defense responses.
By receiving data packets from edge computing nodes, analyzing their internal operating status information, and combining this with context-aware judgment logic, the system distinguishes between interpretable anomalies and external attacks. This includes adaptive processing of business data and metadata analysis, and adjusting the judgment logic based on the internal operating status information.
It effectively reduced the false alarm rate, improved the accuracy and response efficiency of intrusion detection, avoided resource waste and defense failure, and enhanced security protection capabilities in the industrial internet environment.
Smart Images

Figure CN122394936A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intrusion anomaly detection technology, and more specifically, to a cloud-based collaborative intrusion anomaly detection method and system. Background Technology
[0002] In the modern Industrial Internet environment, smart manufacturing equipment, edge computing nodes, and cloud platforms engage in continuous and high-frequency real-time data interaction. To ensure the security and data integrity of the entire industrial control system, operators typically implement a series of stringent security strategies, such as a periodic rotation mechanism for the authentication tokens used for communication between edge computing nodes and the cloud platform. However, in actual industrial deployments, the challenge of device heterogeneity often exists. For example, a large smart factory may deploy edge computing nodes from different vendors, different batches, and even running different operating system versions.
[0003] When these edge nodes fail to authenticate due to expired tokens, they continuously attempt to authenticate with the cloud platform. This persistent authentication failure is detected by the industrial network's security policy engine. If the system identifies this persistent authentication failure as a potential device malfunction or abnormal activity, When the memory buffer of an edge node approaches its overflow threshold, the task scheduler of its internal operating system may be forced to sacrifice the real-time and ordered nature of packet transmission in order to prevent system crashes and ensure the continuous operation of core services (such as local device control). This can cause significant, irregular timestamp delays and jitter in the packaging and transmission of normal device status data that should be sent in strict order and with consecutive sequence numbers. Summary of the Invention
[0004] This application discloses a cloud-based collaborative method and system for detecting intrusion anomalies, aiming to address the challenges faced by existing technologies in identifying complex attack patterns in industrial internet environments. These challenges include a high false alarm rate, and the fact that data flow anomalies caused by internal factors of edge computing nodes are similar to external attacks, leading to misjudgments by cloud detection systems and hindering effective defense responses.
[0005] The technical solution of this application is as follows: In a first aspect, this application discloses a cloud-based collaborative method for detecting intrusion anomalies, comprising: receiving a data packet from an edge computing node, the data packet containing business data and embedded internal operating status information, the internal operating status information indicating the memory usage, authentication connection status, business data processing method, and network traffic management status of the edge computing node, wherein the business data is business data that has undergone adaptive processing when the edge computing node detects a specific operating status; the adaptive processing includes compression processing or aggregation processing; The metadata of the business data stream in the data packet is analyzed to identify whether there are any anomalies. When an anomaly is identified, data stream metadata anomaly is generated. Based on internal operating status information, context-aware judgment is performed on data stream element information anomalies. Context-aware judgment includes: adjusting the judgment logic for data stream element information anomalies based on internal operating status information, and classifying data stream element information anomalies as interpretable anomalies or external attacks; interpretable anomalies refer to anomalies caused by internal predicaments of edge computing nodes.
[0006] Secondly, this application also discloses a cloud-based collaborative intrusion anomaly detection system, which includes: The data packet receiving module is used to receive data packets from the edge computing node. The data packets contain business data and embedded internal operating status information. The internal operating status information indicates the memory usage, authentication connection status, business data processing method, and network traffic management status of the edge computing node. The business data is the business data after adaptive processing when the edge computing node detects a specific operating status. Adaptive processing includes compression processing or aggregation processing. The metadata analysis module is used to analyze the metadata of the business data stream in the data packet to identify whether there are any anomalies. When an anomaly is identified, a data stream metadata anomaly is generated. The context-aware judgment module is used to perform context-aware judgment on data stream element information anomalies based on internal operating status information. The context-aware judgment includes: adjusting the judgment logic for data stream element information anomalies based on internal operating status information, and classifying data stream element information anomalies as interpretable anomalies or external attacks; interpretable anomalies refer to anomalies caused by internal predicaments of edge computing nodes.
[0007] Beneficial effects This application provides a cloud-based collaborative intrusion anomaly detection method. By receiving data packets containing business data and internal operational status information, and performing anomaly identification on the metadata of the business data stream, combined with context-aware judgment based on the internal operational status information, it can effectively distinguish between internal interpretability anomalies caused by internal predicaments of edge computing nodes, such as memory usage, authentication connection status, business data processing methods, and network traffic management status, and genuine external attacks. This method solves the problem in existing technologies where data stream anomalies caused by internal factors of edge computing nodes resemble external attacks, leading to misjudgments by cloud-based detection systems. By adjusting the judgment logic based on internal operational status information, this application can avoid misjudging internal predicaments as external attacks, thereby reducing the false alarm rate, improving the accuracy and reliability of detection, enabling the system to take targeted defensive measures, avoiding resource waste and defense failures caused by misjudgments, and effectively improving security protection capabilities in the industrial internet environment. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of a cloud-based collaborative intrusion anomaly detection method provided in this application.
[0009] Figure 2 This is a schematic diagram of a cloud-based collaborative intrusion anomaly detection system provided in this application. Detailed Implementation
[0010] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0011] Reference Figure 1 The diagram illustrates an embodiment of a cloud-based collaborative intrusion anomaly detection method according to the present invention, which may specifically include the following steps: S101, Receive a data packet from an edge computing node. The data packet contains service data and embedded internal operating status information. The internal operating status information indicates the memory usage, authentication connection status, service data processing method, and network traffic management status of the edge computing node. The service data is service data that has undergone adaptive processing when the edge computing node detects a specific operating status. The adaptive processing includes compression processing or aggregation processing. S102, Analyze the metadata of the service data stream in the data packet to identify whether there is an anomaly. When an anomaly is identified, generate a data stream metadata anomaly. S103, based on the internal operating status information, perform context-aware judgment on the abnormality of the data stream element information. The context-aware judgment includes: adjusting the judgment logic for the abnormality of the data stream element information based on the internal operating status information, and classifying the abnormality of the data stream element information as an interpretable abnormality or an external attack; the interpretable abnormality refers to an abnormality caused by the internal predicament of the edge computing node.
[0012] This application, by receiving data packets containing internal operating status information and combining this information to perform context-aware judgment on data stream meta-information anomalies, can effectively distinguish between interpretable anomalies caused by internal predicaments of edge computing nodes and genuine external attacks, thereby significantly reducing false alarm rates and improving the accuracy and response efficiency of intrusion detection.
[0013] To better understand the technical solution proposed in this application, some key terms involved will be explained first.
[0014] "Edge computing nodes" refer to computing devices deployed at the edge of the network, close to the data source. They have certain computing, storage and networking functions and can perform preliminary processing and analysis of data.
[0015] A "data packet" is a basic unit of data transmitted in a network, which encapsulates business data and control information.
[0016] "Business data" refers to data generated by intelligent manufacturing equipment connected to edge computing nodes, reflecting core information such as equipment operating status and production process.
[0017] "Internal operating status information" refers to the operating status data of the edge computing node itself, such as memory usage, authentication connection status, business data processing methods, and network traffic management status. This information is crucial for understanding the underlying causes of abnormal data packets.
[0018] "Adaptive processing" refers to the preprocessing of business data by edge computing nodes under specific operating conditions, which aims to optimize data transmission efficiency or alleviate resource pressure, such as compression or aggregation processing.
[0019] "Data stream metadata" refers to information that describes the characteristics of a data stream, such as the timestamp, sequence number, source / destination address, and protocol type of data packets. By analyzing this metadata, the normality of the data stream can be preliminarily determined.
[0020] "Abnormal data stream metadata" refers to features that do not conform to the normal data stream pattern, identified through analysis of data stream metadata.
[0021] "Context-aware judgment" refers to combining the internal operating status information of edge computing nodes to conduct a deeper analysis and judgment of abnormal data stream metadata in order to distinguish the root cause of the abnormality.
[0022] "Explainability anomalies" refer to anomalies caused by internal factors of the edge computing node (such as resource constraints, misconfiguration, etc.), rather than by external attacks.
[0023] "External attack" refers to an action initiated by a malicious third party with the aim of compromising system security, stealing data, or interfering with normal operation.
[0024] The cloud-based collaborative intrusion anomaly detection method proposed in this application is based on the ability to effectively distinguish between interpretable anomalies caused by internal predicaments and external attacks through context-aware judgment.
[0025] Specifically, when receiving data packets from edge computing nodes, the nodes will adaptively process the business data when they detect specific operational states, such as memory usage reaching a preset threshold or network bandwidth being limited. This adaptive processing can include compression, which involves lossless or lossy compression of the original business data to reduce the size of the data packets, thereby reducing network transmission load and the memory footprint of the edge computing node. For example, when an edge computing node detects that memory usage exceeds 80%, it can activate a data compression module to compress the collected sensor data using the LZ77 algorithm to reduce the payload of each data packet. Another adaptive processing method is aggregation, which combines multiple small data packets into a larger one to reduce the number of data packets, thereby reducing network protocol overhead and processing frequency. For example, when an edge computing node receives a large number of small data updates from the same sensor within a short period, it can aggregate them into a single data packet containing data from multiple time points for reporting. This adaptively processed business data, along with internal operational status information indicating the edge computing node's memory usage, authentication connection status, business data processing method, and network traffic management status, is encapsulated in the data packet and sent to the cloud platform.
[0026] To analyze the metadata of business data streams within data packets to identify anomalies, the cloud platform first extracts the metadata of the business data streams upon receiving the packets. This metadata includes, but is not limited to, the packet's timestamp, sequence number, source IP address, destination IP address, protocol type, and packet size. By performing pattern matching, statistical analysis, or anomaly detection based on machine learning models on this metadata, anomalies that deviate from normal patterns can be identified in the data stream. For example, if an irregular, large jump in the packet's timestamp is detected, or if the packet's sequence number is discontinuous, it will be initially identified as an anomaly in the data stream's metadata. At this point, the system will generate a flag indicating an anomaly in the data stream's metadata, suggesting a potential problem.
[0027] In terms of context-aware judgment of data stream metadata anomalies based on internal operational status information, the system does not immediately classify anomalies as external attacks upon detection. Instead, it combines the internal operational status information embedded in the data packets to perform context-aware judgment. The core of context-aware judgment lies in adjusting the judgment logic for data stream metadata anomalies based on internal operational status information. For example, if internal operational status information indicates that the memory utilization of the edge computing node has reached 95%, and its business data processing method is compression, then the previously detected anomalies such as timestamp jitter or discontinuous sequence numbers in the data packets are likely not caused by external attacks, but rather by the edge computing node being forced to perform data compression and task scheduling adjustments under extreme memory pressure to maintain basic operation. In this case, the judgment logic will be adjusted to classify the data stream metadata anomaly as an interpretable anomaly. Conversely, if internal operational status information shows that the edge computing node is running normally, has sufficient memory, and has no special business data processing methods, but data stream metadata anomalies are still detected, then the judgment logic will tend to classify it as an external attack. In this way, this application can effectively distinguish between interpretability anomalies caused by internal dilemmas of edge computing nodes and genuine external attacks, thus avoiding false alarms.
[0028] The cloud-based collaborative intrusion anomaly detection method proposed in this application effectively solves the problems of high false alarm rates and difficulty in distinguishing between internal predicaments and external attacks faced by traditional intrusion detection systems in industrial internet environments by introducing internal operational status information of edge computing nodes for context-aware judgment. Traditional methods often rely solely on the analysis of data flow metadata. When the data flow of an edge computing node becomes abnormal due to its own resource limitations or policy adjustments, it is easily misjudged as an external attack. For example, when an edge computing node is short of memory, it may be forced to compress or aggregate data, resulting in irregular changes in the timing or sequence number of data packets, which is easily identified as an anomaly in traditional detection models. However, this application receives internal operational status information embedded in the data packets, such as memory usage, authentication connection status, business data processing methods, and network traffic management status, and uses this as contextual information to make a deeper judgment on the anomalies in data flow metadata.
[0029] For example, when an anomaly in data stream metadata is detected (such as packet timing jitter or sequence number jumps), if the internal operating status information simultaneously indicates that the edge computing node is operating under high memory load and its business data is undergoing compression processing, the judgment logic of this application will consider this anomaly to be due to adaptive measures taken by the edge computing node due to limited internal resources, thus classifying it as an interpretable anomaly. This contrasts sharply with traditional methods that directly classify it as an external attack. Through this context-aware judgment, this application can avoid misjudging normal (albeit abnormal) behaviors of edge computing nodes due to internal difficulties as malicious attacks, thereby significantly reducing the false alarm rate and improving detection accuracy. Furthermore, this method can help the cloud platform more accurately understand the true operating status of edge computing nodes, providing a more reliable basis for subsequent fault diagnosis and resource optimization, avoiding unnecessary defensive measures due to misjudgment, and thus improving the resilience and efficiency of the entire industrial internet system.
[0030] This application further proposes that the above-mentioned steps for context-aware judgment of data stream element information anomalies based on internal operating status information include: Collect external network behavior information related to edge computing nodes; the external network behavior information includes the edge computing node's authentication attempt records at the network layer, the number of authentication failures, the traffic management logs imposed on it by network security policies, and the connection status and data transmission statistics recorded by network devices; The internal operating status information is correlated and compared with the external network behavior information to evaluate the consistency between the internal operating status information and the external network behavior information, and a consistency evaluation result is obtained. Based on the internal operating status information and the consistency assessment results, the judgment logic for the abnormality of the data stream element information is adjusted, and the abnormality of the data stream element information is classified as interpretability anomaly or external attack.
[0031] Specifically, external network behavior information refers to network activity data observed from the perspective of an edge computing node, aiming to provide an external reference frame independent of the edge computing node's internal state. This information can include network-layer authentication attempt logs of the edge computing node, such as logs of attempts to connect to other network services or devices; the number of authentication failures, which can indicate brute-force attempts or configuration errors; traffic management logs imposed by network security policies, such as traffic restrictions or blocking events recorded by firewalls or intrusion prevention systems; and connection status and data transmission statistics recorded by network devices, such as port connection status, packet loss rate, and bandwidth usage recorded by routers or switches. This external information can reflect the behavioral patterns of the edge computing node and the network environment it faces at the network level.
[0032] The purpose of correlating and comparing internal operational status information with external network behavior information is to assess whether there are any contradictions or inconsistencies between the two. For example, if internal operational status information shows high memory usage and slow business data processing, while external network behavior information shows normal network traffic and no abnormal authentication attempts, this may indicate that the anomaly mainly stems from internal problems. Conversely, if the internal status is normal but external network behavior information shows a large number of authentication failures or abnormal traffic, it may point to an external attack. The consistency assessment result can be a quantitative indicator, such as a consistency score, or a qualitative judgment, such as "highly consistent," "partially consistent," or "inconsistent."
[0033] In practical applications, the judgment logic for data flow metadata anomalies is adjusted based on internal operational status information and consistency assessment results to make anomaly classification more accurate. For example, when internal operational status information indicates resource constraints, and the consistency assessment results show a high degree of consistency between the internal state and external network behavior (i.e., no obvious external anomalies), the judgment logic may tend to classify the data flow metadata anomaly as an interpretability anomaly. Conversely, if there is a significant inconsistency between internal status information and external network behavior information—for example, the internal state is normal but there are numerous authentication failures externally—the judgment logic may be more inclined to classify it as an external attack. This adjustment can be achieved by modifying judgment thresholds, weights, or enabling / disabling specific judgment rules.
[0034] Through the above technical solution, this application can more comprehensively and accurately perform context-aware judgment on anomalies in data stream metadata. By introducing external network behavior information and correlating and comparing it with internal operational status information, the judgment bias that may result from relying solely on internal status information can be effectively compensated for. This enables the system to more accurately distinguish between interpretable anomalies caused by internal predicaments of edge computing nodes and external attacks caused by malicious external behavior, thereby significantly reducing false positive and false negative rates. In addition, this method improves the robustness of intrusion anomaly detection. Even when the internal status information of edge computing nodes may be tampered with or not entirely reliable, cross-validation can be provided through external network behavior information to ensure the accuracy of the judgment.
[0035] In some preferred embodiments, a specific example is given below. Suppose that an edge computing node, during a certain period, shows that its internal operating status information indicates a consistently high memory usage rate and a significant decrease in business data processing speed, resulting in abnormal data stream metadata. Based solely on internal operating status information, the system might initially determine that this is an interpretable anomaly caused by internal resource constraints. However, the solution in this application further collects external network behavior information related to the edge computing node. For example, the system discovers that during the same period, the edge computing node's network security policy logs show a large number of authentication attempts from unknown IP addresses, with a sharp increase in the number of authentication failures, while network device connection status records show abnormal port scanning behavior.
[0036] At this point, the system compares and correlates internal operational status information (high memory usage, slow business processing) with external network behavior information (numerous authentication failures, port scanning). Through evaluation, the system discovers a significant inconsistency between internal anomalies and external network attack behavior. Specifically, internal resource constraints may be an independent issue, while the large number of authentication failures and port scans clearly point to external attacks. Based on this consistency assessment, the system adjusts its logic for judging data stream metadata anomalies. For example, the rule weights in the judgment logic are adjusted, increasing the weight of external attack characteristics. Ultimately, despite internal resource constraints, due to the strong attack characteristics of external network behavior, the data stream metadata anomaly is classified as an external attack, rather than a simple interpretable anomaly. This allows for timely triggering of appropriate security response mechanisms, effectively addressing potential intrusion threats.
[0037] This application further proposes the following steps for adjusting the judgment logic for data flow element information anomalies based on internal operating status information and consistency assessment results, and classifying data flow element information anomalies as interpretable anomalies or external attacks: External network behavior information from network devices from different vendors is processed in a unified manner to obtain unified external network behavior information. The unified processing includes converting logs of different formats into a standard format. Based on the time synchronization protocol and accuracy of the network devices for each external network behavior information, the timestamp of the unified external network behavior information is calibrated to obtain calibrated external network behavior information. Based on the source device type of the calibrated external network behavior information and its time-calibrated time deviation range, the weight of the consistency assessment between the internal operating status information and the calibrated external network behavior information is adjusted to obtain a weighted consistency assessment result. Based on the weighted consistency assessment results and internal operating status information, determine the type of anomaly in the data stream metadata.
[0038] Specifically, standardizing external network behavior information from network devices of different vendors refers to converting external network behavior information collected from different network devices (such as routers, switches, firewalls, etc.), which may have different log formats or data structures, into a unified, standardized data format through predefined transformation rules or parsers. The aim is to eliminate data heterogeneity and provide a consistent foundation for subsequent data processing and analysis. For example, Syslog, SNMP Trap, or proprietary vendor log formats can be uniformly converted into standard formats such as JSON or XML.
[0039] Specifically, the timestamps of the unified external network behavior information are calibrated based on the time synchronization protocols and accuracy of the network devices used to obtain the external network behavior information. This can be understood as a fine-tuning of the timestamps of these external network behavior information after data unification, considering that different network devices may use different time synchronization mechanisms (e.g., NTP, PTP) or have different clock accuracies. The purpose is to ensure that all external network behavior information has a high degree of consistency in the time dimension, thereby enabling accurate time-series correlation and comparison with internal operating status information. For example, the timestamps can be forward or backward corrected based on the NTP synchronization status and clock drift rate of the devices.
[0040] In practical applications, adjusting the weight of the consistency assessment between the internal operating status information and the calibrated external network behavior information based on the source device type and the time-series deviation range after time calibration means that when assessing the consistency between internal operating status information and external network behavior information, all external information is no longer simply assigned the same weight. Instead, its importance in the assessment is dynamically adjusted based on the reliability, accuracy, and relevance of the external information to the internal state. For example, logs from core network devices with high time synchronization accuracy can have a higher weight, while logs from edge devices with large time-series deviations can have a lower weight. The aim is to improve the accuracy and robustness of the consistency assessment and reduce misjudgments caused by unreliable external information.
[0041] Therefore, determining the type of data flow element information anomaly based on the weighted consistency assessment results and internal operating status information means that after obtaining the consistency assessment results after unification, calibration, and weighting, the previously identified data flow element information anomalies are more accurately classified in conjunction with the internal operating status information of the edge computing node itself. The purpose is to distinguish whether the anomaly is caused by "interpretable anomalies" such as resource bottlenecks or configuration errors within the edge computing node, or by "external attacks" caused by malicious external attacks.
[0042] Through the above technical solution, this application effectively addresses the problem of insufficient accuracy in context-aware judgment caused by inconsistent data formats, inaccurate time synchronization, and differences in data reliability in environments with multi-source heterogeneous external network behavior information. Compared to basic solutions, this application significantly improves the accuracy and robustness of the correlation and comparison between internal operational status information and external network behavior information through unified processing, timestamp calibration, and weighted consistency evaluation. This results in more accurate classification of anomalies in data stream metadata, avoiding misjudging internal predicaments as external attacks or vice versa, and greatly improving the reliability of intrusion anomaly detection and the effectiveness of decision-making.
[0043] In some preferred embodiments, suppose that in a cloud-edge collaborative system, the data stream metadata of an edge computing node is identified as abnormal. In this case, context-aware judgment is required.
[0044] First, the system receives external network behavior information from network devices (such as routers, firewalls, and switches) from different vendors (e.g., Huawei, Cisco, H3C). This information may exist in Syslog, SNMP Trap, or vendor-defined binary formats. For effective analysis, this heterogeneous external network behavior information is processed in a unified manner. For example, a pre-defined parser converts all logs into a standardized JSON format, ensuring that all key fields (such as source IP, destination IP, port, event type, and timestamp) have consistent naming and data types.
[0045] Secondly, after unification, the system will detect potential discrepancies in timestamps across different network devices. For example, a firewall might synchronize time via an NTP server with millisecond-level precision, while another switch might synchronize only via its local clock with second-level precision and some clock drift. In this case, the system will calibrate the timestamps of the unified external network behavior information based on the NTP synchronization status, clock drift rate, and time synchronization protocol of each device, ensuring that the timestamps of all events are aligned to a unified system time base. For example, all timestamps may be calibrated to nanosecond-level precision.
[0046] Furthermore, when assessing the consistency between internal operational status information and calibrated external network behavior information, the system dynamically adjusts the evaluation weights based on the source device type of the external information and its time-calibrated timing deviation range. For example, traffic logs from the core router with high time synchronization accuracy and small timing deviation will have a weight of 0.8; while connection logs from an older edge switch with lower time synchronization accuracy and larger timing deviation may have a weight of 0.4. This weighting mechanism can more accurately reflect the reliability of different external information sources.
[0047] Finally, the system combines the weighted consistency assessment results with the internal operating status information of the edge computing nodes (e.g., high memory usage, numerous authentication connection failures) to determine the type of data flow metadata anomaly. For example, if the consistency assessment results show normal external network traffic, but the internal operating status information indicates persistently high memory usage, and the weighted assessment results favor internal factors, then this data flow metadata anomaly may be classified as an "interpretability anomaly" (e.g., resource exhaustion). Conversely, if the consistency assessment results show a large number of abnormal authentication attempts from the outside, and the weighted assessment results favor external factors, then this anomaly may be classified as an "external attack." In this way, performance degradation caused by internal resource constraints can be avoided from being misjudged as an external attack, or a real external attack can be misjudged as an internal problem, thereby improving the accuracy of detection and the targeting of the response.
[0048] This application further proposes a step for determining the type of anomaly in the data stream metadata based on the weighted consistency assessment results and internal operating status information, including: Collect operational behavior data of edge computing nodes; The cloud platform receives the operational behavior data and identifies the current operating mode of the edge computing node based on the operational behavior data. The operating mode includes normal operation, resource-constrained operation, or authentication-abnormal operation. The current operating mode of the edge computing node is identified, and the parameters in the judgment logic of the abnormal data stream metadata are adjusted to obtain the adjusted judgment logic. The parameters include judgment thresholds or rule weights. Based on the adjusted judgment logic, combined with the weighted consistency assessment results and the internal operating status information, the abnormal data stream metadata is classified.
[0049] Specifically, collecting operational behavior data from edge computing nodes refers to gathering real-time or historical data related to their current operational status and performance. This data can include CPU utilization, memory usage, disk I / O, network bandwidth usage, process lists, system logs, error reports, and performance metrics for specific applications. This operational behavior data can comprehensively reflect the workload, resource health, and potential internal problems of edge computing nodes.
[0050] The cloud platform receives the operational behavior data and identifies the current operating mode of the edge computing nodes based on this data. The operating mode includes normal operation, resource-constrained operation, or authentication-abnormal operation. The cloud platform, acting as a centralized processing and analysis center, is responsible for receiving operational behavior data reported by each edge computing node. Upon receiving the data, the cloud platform uses pre-trained models or rule-based analysis methods to perform in-depth analysis of this data to identify the current operating mode of the edge computing nodes. For example, when memory usage consistently exceeds a certain threshold or CPU utilization remains high for an extended period, it may be identified as a resource-constrained operating mode; when a large number of authentication failure logs or connection attempts are rejected, it may be identified as an authentication-abnormal operating mode; and when all indicators fluctuate within the normal range, it is identified as a normal operating mode.
[0051] In practical applications, after identifying the current operating mode of the edge computing node, the parameters in the judgment logic for data flow metadata anomalies are adjusted to obtain the adjusted judgment logic. These parameters include judgment thresholds or rule weights. Once the current operating mode of the edge computing node is identified, the cloud platform dynamically adjusts the logic parameters used to judge data flow metadata anomalies based on this operating mode. For example, in a resource-constrained operating mode, certain anomaly judgment thresholds related to resource consumption may be appropriately relaxed to avoid misjudging normal performance degradation as an attack; while in an authentication-related anomaly operating mode, the weights of authentication-related anomaly rules may be increased to more sensitively detect potential authentication attacks. This dynamic adjustment ensures that the judgment logic can adapt to the actual operating environment of the edge computing node.
[0052] Furthermore, based on the adjusted judgment logic, combined with the weighted consistency assessment results and the internal operating status information, the data flow metadata anomalies are categorized. After the judgment logic parameters are adjusted, the system will use this new, more adaptive judgment logic to comprehensively consider the previously obtained weighted consistency assessment results and the internal operating status information reported by the edge computing nodes to finally categorize the data flow metadata anomalies. This categorization will more accurately identify whether the anomaly is an interpretable anomaly (such as performance degradation caused by resource constraints) or an external attack (such as malicious authentication attempts).
[0053] In some preferred embodiments, a specific example is given below. Suppose an edge computing node is deployed in a smart factory environment, responsible for processing production line data in real time.
[0054] First, the system will continuously collect operational behavior data of the edge computing node, such as CPU utilization, memory usage, network I / O, and system logs.
[0055] When factory production peaks, the CPU and memory utilization of this edge computing node may remain above 90% for extended periods, while network I / O also increases significantly. Upon receiving this operational data, the cloud platform analyzes it and identifies that the edge computing node is currently operating in a "resource-constrained" mode.
[0056] At this point, if the metadata analysis module detects an increase in packet processing latency and generates an anomaly in the data stream metadata, without this solution, this latency might be directly classified as a potential external attack. However, in this solution, because the "resource-constrained operation" mode is identified, the cloud platform dynamically adjusts the parameters in the judgment logic. For example, it appropriately relaxes the judgment threshold for data processing latency or reduces the weight of anomaly rules related to resource consumption.
[0057] Based on the adjusted judgment logic and combined with the previously obtained weighted consistency assessment results (e.g., external network behavior information did not show any abnormal authentication attempts or traffic patterns), the system ultimately classified the data processing delay anomaly as an "interpretability anomaly," rather than an external attack. This indicates that the anomaly is a normal phenomenon caused by resource constraints within the edge computing node, rather than a malicious intrusion.
[0058] In this way, this application avoids false alarms during peak load periods on edge nodes, allowing operations and maintenance personnel to focus more on addressing genuine security threats, thus improving the system's usability and efficiency. In some of the embodiments described above in this application, although it is proposed to adjust the parameters in the judgment logic for abnormal data flow metadata based on the current operating mode of the edge computing node, in actual implementation, adjusting parameters solely based on broad operating modes (such as normal operation, resource-constrained operation, or authentication-abnormal operation) may not be sufficient to capture subtle changes in the internal operating state of the edge computing node. This coarse-grained pattern recognition may lead to insufficient precision in the parameter adjustment of the judgment logic, thereby affecting the accuracy and adaptability of intrusion anomaly detection. To address this, this application further proposes to perform multi-dimensional analysis of the internal operating state information reported by the edge computing node to identify fine-grained operating states, and based on this, select matching judgment thresholds or rule weights from a preset parameter configuration library, thereby more accurately adjusting the parameters in the judgment logic.
[0059] The steps described above for identifying the current operating mode of the edge computing node, adjusting the parameters in the judgment logic for abnormal data stream metadata, and obtaining the adjusted judgment logic include: The internal operating status information reported by the edge computing nodes is analyzed from multiple dimensions to identify the fine-grained operating status of the edge computing nodes. The fine-grained operating status includes the classification of memory usage, the classification of authentication failure reasons, and the values of network interface traffic patterns and task scheduling latency. The multi-dimensional analysis refers to classifying memory usage, refining the classification of authentication failure reasons, or combining specific values of network interface traffic patterns and task scheduling latency to identify the fine-grained operating status of the edge computing nodes. Based on the identified fine-grained operating state, a judgment threshold or rule weight matching the fine-grained operating state is selected from a preset parameter configuration library. Based on the matching judgment threshold or rule weight, the parameters in the judgment logic of abnormal data stream metadata are adjusted to obtain the adjusted judgment logic.
[0060] Specifically, a multi-dimensional analysis is performed on the internal operating status information reported by the edge computing nodes to identify the current fine-grained operating status of the edge computing nodes. This step aims to delve into the details of the internal operating status of the edge computing nodes. Multi-dimensional analysis can be understood as a more refined division of a single operating status indicator or a comprehensive consideration of multiple related indicators. Its purpose is to obtain a more specific and accurate status description than "normal operation" or "resource-constrained operation". The fine-grained operating status includes memory usage grading, authentication failure reason classification, network interface traffic patterns, and task scheduling latency values. Specifically, memory usage grading can divide memory usage into multiple ranges, such as "low (0-25%)", "medium (25-50%)", "high (50-75%)", and "very high (75-100%)", rather than simply judging whether it is "resource-constrained". The classification of authentication failure reasons can be refined into specific reasons such as "incorrect password", "expired certificate", and "insufficient permissions". The numerical values of network interface traffic patterns and task scheduling latency provide quantitative indicators, such as whether the traffic is bursty or under sustained high load, and whether the task scheduling latency is in milliseconds or seconds. These detailed classifications and values can more accurately reflect the actual situation faced by edge computing nodes. The multi-dimensional analysis refers to classifying memory usage, refining the classification of authentication failure reasons, or combining specific values of network interface traffic patterns and task scheduling latency to identify the fine-grained operating state of the edge computing node. Specific implementation methods for multi-dimensional analysis may include, but are not limited to: applying clustering algorithms to classify memory usage data; performing text analysis or pattern matching on authentication failure logs for refined classification; or obtaining feature values of network interface traffic patterns and average and maximum values of task scheduling latency through real-time monitoring and statistics. Through these methods, the operating state of the edge computing node can be comprehensively characterized from multiple perspectives.
[0061] Further, based on the identified fine-grained operating state, a judgment threshold or rule weight matching the fine-grained operating state is selected from a preset parameter configuration library. The parameter configuration library is a pre-established knowledge base that stores the optimal or recommended judgment thresholds and rule weights corresponding to different fine-grained operating states. For example, when a fine-grained state of "extremely high memory usage and significantly increased task scheduling latency" is identified, the system can select a set of more lenient anomaly judgment thresholds from the library to avoid misjudging normal behavior caused by resource constraints as an attack. Conversely, in a state of "low memory usage and normal network traffic patterns," a stricter threshold may be selected. Thus, based on the matched judgment threshold or rule weight, the parameters in the judgment logic for the abnormal data flow metadata are adjusted to obtain the adjusted judgment logic. This step applies the specific parameters selected from the parameter configuration library to the anomaly judgment logic. These parameters can be numerical thresholds used to determine whether the data flow metadata is abnormal (e.g., how much a certain indicator exceeds is considered abnormal), or the importance weights of different rules in the comprehensive judgment. In this way, the judgment logic can dynamically and accurately adapt to the real-time fine-grained status of the edge computing nodes.
[0062] In some preferred embodiments, a specific example is given below. Suppose that at a certain moment, the internal operating status information reported by an edge computing node shows that its memory utilization is 85%, the authentication connection status shows a large number of authentication failure attempts from unknown sources, and the network interface traffic pattern shows an abnormal burst of growth, while the task scheduling latency also increases slightly.
[0063] First, the system performs multi-dimensional analysis of this internal operational status information. Specifically, a memory utilization rate of 85% is identified as a "high memory utilization" level; a large number of failed authentication attempts from unknown sources are categorized as "authentication failure reason: unknown source brute-force attack attempt"; and sudden increases in network interface traffic patterns and task scheduling latency are identified as specific numerical characteristics. Combining this information, the system identifies that the edge computing node is currently in a fine-grained operational state of "unknown source authentication attack under high memory pressure and abnormal network traffic".
[0064] Next, the system queries a pre-defined parameter configuration library based on this identified fine-grained operational state. This library pre-stores optimized parameter sets for different fine-grained states. For example, for the state of "unknown source authentication attack under high memory pressure and abnormal network traffic," the library may be configured with a specific set of judgment thresholds and rule weights: for example, appropriately relaxing the anomaly judgment threshold related to memory to avoid misjudging performance degradation caused by memory pressure as an attack; at the same time, significantly increasing the rule weights related to authentication failure and abnormal network traffic to enhance sensitivity to potential attack behaviors.
[0065] Finally, based on the matching threshold and rule weights selected from the parameter configuration library, the parameters in the data flow metadata anomaly judgment logic are dynamically adjusted. Therefore, when subsequent data flow metadata is analyzed, the judgment logic can more accurately identify whether the anomaly is caused by an external attack or by performance fluctuations due to internal challenges in edge computing nodes (such as high memory pressure), thus making a more accurate classification judgment.
[0066] This application further proposes steps for multi-dimensional analysis of the internal operating status information reported by the aforementioned edge computing nodes to identify the current fine-grained operating status of the edge computing nodes, including: The internal operating status information reported by the edge computing node is structured and parsed to obtain structured internal operating status information. The structured parsing refers to extracting memory usage, authentication failure reasons, network interface traffic patterns, and task scheduling delay, and converting the memory usage, authentication failure reasons, network interface traffic patterns, and task scheduling delay into a unified data structure. Based on a preset set of verification rules, the structured internal operating status information is subjected to consistency checks and integrity verification to obtain the results of the consistency checks and integrity verification. Based on the results of the consistency check and integrity verification, the internal operating status information is quality-marked or data-processed to obtain internal operating status information that has been quality-marked or data-processed. Based on the internal operating status information that has been marked with quality tags or processed by data, the fine-grained operating status of the edge computing node is identified.
[0067] Specifically, structured parsing of the internal operational status information reported by edge computing nodes refers to processing raw, potentially diverse internal operational status information—such as log files, sensor data, or system indicator reports—through a parser to extract key technical features, including memory usage, authentication failure reasons, network interface traffic patterns, and task scheduling latency. These extracted technical features are then converted into a unified data structure, such as JSON objects, XML documents, or database records, with the aim of providing standardized input for subsequent data processing and analysis. This structured processing eliminates format differences caused by different data sources or device models, ensuring data consistency.
[0068] The consistency and integrity verification of the structured internal operating status information, based on a predefined set of verification rules, refers to evaluating data quality using a series of predefined rules. Consistency checks can include data range verification (e.g., memory usage should be between 0% and 100%), data type verification (e.g., authentication failure reason should be string type), or logical consistency verification (e.g., if the authentication connection status is "connected," the number of authentication failures should be zero). Integrity verification aims to detect whether there are missing critical fields or records in the data; for example, if memory usage data for a certain timestamp is missing, it will be marked. The verification rule set can be customized according to the type of edge computing node, operating system, or application scenario, with the aim of ensuring data accuracy, reliability, and availability.
[0069] In practical applications, based on the results of the aforementioned consistency checks and integrity verifications, internal operational status information undergoes quality labeling or data processing. Quality labeling involves attaching metadata tags to data to indicate its quality level or existing problems, such as "incomplete data," "abnormal data," or "high-confidence data." Data processing can include various operations. For example, missing data can be processed using interpolation, mean imputation, or record deletion; inconsistent data can be corrected according to preset correction logic; and outliers can be smoothed or isolated. The aim is to improve the overall quality of the data, making it more suitable for subsequent fine-grained operational status identification.
[0070] In some preferred embodiments, a specific example is given below. Suppose an edge computing node reports internal operational status information including system logs, performance metrics, and network events. First, this raw information is parsed in a structured manner. For example, text lines in the system logs are parsed into structured records containing timestamps, event types, memory usage, and authentication attempt results; performance metric data is extracted to include values such as CPU utilization, memory usage, and task scheduling latency; and network events are parsed to include fields such as source IP, destination IP, port, and authentication failure reason. All this information is then converted into a unified JSON format.
[0071] Furthermore, based on a predefined set of validation rules, consistency checks and integrity verifications are performed on these structured JSON data. For example, validation rules might include: the memory usage field must have a value between 0 and 100; the authentication failure reason field must be a value from a predefined list (e.g., "incorrect password," "user does not exist"); all records must contain a valid timestamp; and critical fields (such as memory usage) cannot be empty. If a record's memory usage exceeds 100%, or its authentication failure reason is not in the predefined list, that record is marked as inconsistent. If a record lacks a task scheduling delay field, it is marked as incomplete.
[0072] Subsequently, based on the results of consistency checks and integrity verifications, the internal operational status information is marked with quality indicators or processed. For example, for records with memory usage exceeding 100%, the memory usage can be corrected to 100% and marked as "Data Corrected"; for records where the authentication failure reason does not match the preset list, they can be classified as "Unknown Authentication Failure" and marked as "Low Confidence"; for records missing the task scheduling delay field, the previous valid value can be used to fill it in, or it can be directly marked as "Incomplete Data, Interpolated".
[0073] Ultimately, based on this internal operational status information, which has been quality-marked or processed, the system can accurately identify the fine-grained operational status of the edge computing nodes. For example, if the processed data shows that memory usage consistently exceeds 90% and task scheduling latency increases significantly, it can be identified as a "memory bottleneck state in resource-constrained operation mode"; if authentication failures frequently result in "incorrect password" and the authentication connection status remains "not connected," it can be identified as a "credential problem state in authentication anomaly operation mode." This precise, fine-grained operational status information will be used to adjust the parameters in the judgment logic for data stream metadata anomalies, thereby achieving more accurate detection of intrusion anomalies.
[0074] This application further proposes the following steps for performing consistency checks and integrity verifications on the structured internal operating state information based on a preset set of verification rules: Receive novel device identification information or internal dilemma mode characteristics from the edge computing node. The novel device identification information indicates the model, manufacturer, or firmware version of the edge computing node, and the internal dilemma mode characteristics indicate an abnormal behavior pattern of the edge computing node that is not covered by the existing rule set. Based on the novel device identification information, the device operating parameters, data reporting specifications, and known internal distress manifestations corresponding to the novel device identification information are obtained from a preset device characteristic database; the internal distress manifestations include the data compression method of a specific model of device when memory is insufficient. Based on the internal dilemma pattern characteristics, analyze the differences between the internal dilemma pattern characteristics and the existing verification rule set to identify the coverage blind spots of the existing rule set; Based on the device operating parameters, data reporting specifications, known internal distress manifestations, and analysis results of coverage blind spots, generate or adjust the verification rules in the verification rule set. The verification rules include data range verification rules, timing consistency verification rules, or integrity check rules. Based on the verification rules in the generated or adjusted verification rule set, the structured internal operating status information is subjected to consistency checks and integrity verification.
[0075] Specifically, upon receiving novel device identification information or internal predicament pattern characteristics from edge computing nodes, the system can recognize these new or unknown behavioral patterns. Novel device identification information can be understood as the unique identification information of the edge computing device, such as its model, manufacturer, or firmware version. This information is crucial for understanding the device's inherent behavior and data characteristics. Internal predicament pattern characteristics refer to abnormal behavioral patterns exhibited by edge computing nodes under specific internal conditions that have not yet been recognized or covered by existing verification rule sets, such as data transmission delay patterns due to resource constraints or specific error log sequences.
[0076] Furthermore, upon receiving identification information for a new type of device, the system queries a pre-defined device characteristic database. This database stores the operating parameters, data reporting specifications, and behavior patterns of various known devices under specific internal predicaments. For example, for a particular device model, when its memory is insufficient, it may employ a specific data compression method to process business data; this behavior is an inherent characteristic of the device, not an external attack. By acquiring this information, the system can provide the foundational data for generating or adjusting subsequent verification rules.
[0077] Furthermore, when an internal dilemma pattern characteristic is received, the system performs a difference analysis between that characteristic and the existing set of validation rules. This analysis aims to identify blind spots in the coverage of the existing rule set when handling such novel abnormal behavior patterns, i.e., scenarios that the existing rules cannot accurately judge or interpret.
[0078] Therefore, based on the acquired equipment operating parameters, data reporting specifications, known internal predicament manifestations, and analysis results of coverage blind spots, the system can dynamically generate or adjust the verification rules in the verification rule set. These verification rules may include data range verification rules (e.g., the reasonable value range of a parameter), timing consistency verification rules (e.g., whether the order and time interval of events meet expectations), or integrity check rules (e.g., whether data packets contain all required fields). Finally, using these generated or adjusted verification rules in the verification rule set, consistency checks and integrity verifications are performed on the structured internal operating status information, thereby ensuring the accuracy and adaptability of the verification.
[0079] In some preferred embodiments, it is assumed that a novel IoT edge gateway device is deployed in the network. This gateway device has a unique firmware version and data reporting protocol, and when memory utilization reaches 90%, it initiates a special log compression mechanism, resulting in a difference between the log format it reports and the standard format expected by existing rule sets.
[0080] At this point, the detection method of this application receives the identification information of the novel device, such as its firmware version. The system queries the device feature database based on the firmware version to obtain the log compression method of this model device under high memory load. Simultaneously, if the system detects an abnormal log format reported by the gateway device, and this abnormal pattern is not covered by the existing rule set, it will identify it as an internal dilemma mode feature.
[0081] Based on device characteristics (log compression method) obtained from the database and analysis of blind spots in the existing rule set, the system dynamically generates a new verification rule. For example, when a device with a specific firmware version is detected and its memory usage exceeds 85%, its logs are allowed to use a specific compression format. This newly generated verification rule is then added to the verification rule set.
[0082] Therefore, when the new gateway device reports compressed logs when memory is low, the system will be able to accurately identify, based on the updated verification rule set, that this is a normal behavior caused by an internal predicament rather than an external attack, thus avoiding false alarms and ensuring accurate verification of the internal operating status information of the new device.
[0083] Specifically, the steps described above, which involve obtaining the equipment operating parameters, data reporting specifications, and known internal distress manifestations corresponding to the new equipment identification information from a pre-set equipment characteristic database, can be further refined into the following sub-steps.
[0084] The steps described above, which involve retrieving the equipment operating parameters, data reporting specifications, and known internal distress manifestations corresponding to the new equipment identification information from a pre-set equipment characteristic database, include: Receive identification information for new equipment; The new equipment identification information is subjected to format verification and standardization processing to obtain the processed new equipment identification information; Based on the processed new equipment identification information, a query is performed in the equipment characteristic database to obtain the corresponding equipment operating parameters, data reporting specifications, and known internal distress manifestations.
[0085] Receiving novel device identification information refers to the system receiving unique identification information about a novel device from an edge computing node or other relevant source. This novel device identification information may include the device's model, manufacturer, firmware version, etc., and is used to uniquely identify an edge computing device.
[0086] Furthermore, the new equipment identification information undergoes format validation and standardization to obtain processed new equipment identification information, aiming to ensure that the received identification information conforms to preset data formats and specifications. Format validation may include checking character sets, length, and the validity of specific fields. Standardization may involve unifying identification information from different sources or formats into a standard format, for example, mapping the naming rules of equipment models from different manufacturers, or removing unnecessary special characters to facilitate subsequent database queries and matching.
[0087] Therefore, based on the processed new device identification information, a query is performed in the device characteristic database to obtain the corresponding device operating parameters, data reporting specifications, and known internal distress manifestations. The device characteristic database is a pre-built knowledge base that stores detailed information on various known edge devices. By using the standardized new device identification information as the query key, the system can accurately retrieve from the database the device's specific operating parameters (e.g., CPU frequency, memory size, sensor type), data reporting specifications (e.g., data transmission protocol, data format, reporting frequency), and known behavioral patterns or characteristics that the device may exhibit under specific internal distress conditions. For example, for a specific model of device, its data compression method may exhibit specific manifestations when memory is insufficient; all this information can be obtained from the database.
[0088] In some of the embodiments described above in this application, format verification and standardization of novel device identification information are proposed to obtain corresponding device operating parameters, data reporting specifications, and known internal dilemma manifestations. However, in practical applications, novel device identification information may originate from multiple device manufacturers, and its data format, character encoding, or transmission method may differ, and it may even contain encrypted fields. If these complexities are not adequately addressed before format verification and standardization, it may lead to information parsing failure or inaccurate parsing, thereby affecting the query results of the subsequent device characteristic database, reducing the accuracy of the generation or adjustment of the verification rule set, and ultimately affecting the reliability of intrusion anomaly detection.
[0089] In this regard, this application further proposes the following steps for performing format verification and standardization processing on the new equipment identification information to obtain the processed new equipment identification information: Receive identification information for new equipment; The character encoding of the novel device identification information is detected to identify the character encoding type used in the novel device identification information; Based on the identified character encoding type, the novel device identification information is decoded to obtain the decoded novel device identification information; Encrypted fields are identified in the decoded new device identification information to identify the encrypted fields present in the decoded new device identification information; The identified encrypted fields are decrypted to obtain the decrypted new device identification information; The decrypted new device identification information is then subjected to format verification and standardization processing.
[0090] Specifically, receiving new device identification information refers to obtaining raw device identification data from edge computing nodes or other data sources. This information may exist as a string, binary stream, or other forms. Character encoding detection of the new device identification information can be understood as determining the character encoding type used by the information, such as UTF-8, GBK, or ASCII, by analyzing the byte sequence of the data stream or using a preset encoding recognition algorithm (e.g., based on BOM header, character frequency statistics, or machine learning models). Its purpose is to ensure the correctness of subsequent decoding. In practical applications, decoding the new device identification information based on the identified character encoding type means converting the data in the original encoding format into a unified character encoding format within the system (e.g., Unicode) to eliminate data garbled characters or parsing errors caused by encoding inconsistencies. Further, identifying encrypted fields in the decoded new device identification information specifically refers to detecting the presence of encrypted regions or parameters in the data through pattern matching, metadata analysis, or a preset rule base. Its purpose is to identify the parts that need decryption. Therefore, decrypting the identified encrypted fields means using a preset key, decryption algorithm, or security protocol to restore the encrypted fields to the original plaintext data. The purpose is to obtain complete and readable device identification information. Finally, the decrypted new device identification information undergoes format verification and standardization processing. This involves checking the completeness and legality of the information according to a preset device identification specification or data model, and converting it into a unified and standardized data format (such as JSON, XML, or key-value pairs) to facilitate efficient querying and matching in the device characteristic database.
[0091] In some preferred embodiments, assume the cloud platform receives a new device identification information, the original form of which is a Base64 encoded JSON string with some fields encrypted by AES, for example: “eyJkZXZpY2VfaWQiOiAiMTIzNDUiLCAiZW5jcnlwdGVkX2RhdGEiOiAiS0pISkxzZGtqZmRzYSJ9”. First, the system receives this new device identification information. Next, it performs character encoding detection on the information, identifying that it uses UTF-8 encoding. Then, based on the identified UTF-8 encoding, it decodes the information to obtain the decoded JSON string: “{"device_id": "12345", "encrypted_data": "KJHJLsdkjfdsa"}”. Then, the system identifies the “encrypted_data” field as an encrypted field. Using a pre-defined AES key, the "encrypted_data" field is decrypted, assuming the result is "firmware_v1.2". The new device identification information then becomes: "{"device_id":"12345", "encrypted_data": "firmware_v1.2"}". Finally, the decrypted new device identification information undergoes format validation and standardization to ensure it conforms to the pre-defined device identification JSON format specification, and the field names are standardized. For example, if the specification requires camelCase naming, it might be converted to "{"deviceId": "12345", "encryptedData": "firmware_v1.2"}". After this series of processes, the new device identification information is reliably parsed and standardized, enabling accurate queries in the device characteristic database to obtain detailed operating parameters and predicament behaviors of the device, thereby supporting more precise detection of intrusion anomalies.
[0092] Secondly, referring to Figure 2 This application further proposes a cloud-based collaborative intrusion anomaly detection system, which includes: The data packet receiving module 201 is used to receive data packets from the edge computing node. The data packets contain service data and embedded internal operating status information. The internal operating status information indicates the memory usage, authentication connection status, service data processing method, and network traffic management status of the edge computing node. The service data is the service data after adaptive processing when the edge computing node detects a specific operating status. Adaptive processing includes compression processing or aggregation processing. The metadata analysis module 202 is used to analyze the metadata of the business data stream in the data packet to identify whether there is an anomaly. When an anomaly is identified, a data stream metadata anomaly is generated. The context-aware judgment module 203 is used to perform context-aware judgment on data stream element information anomalies based on internal operating status information. The context-aware judgment includes: adjusting the judgment logic for data stream element information anomalies based on internal operating status information, and classifying data stream element information anomalies as interpretable anomalies or external attacks; interpretable anomalies refer to anomalies caused by internal predicaments of edge computing nodes.
[0093] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A cloud-based collaborative method for detecting abnormal intrusion behavior, characterized in that, include: The system receives data packets from an edge computing node. These data packets contain service data and embedded internal operating status information. The internal operating status information indicates the edge computing node's memory usage, authentication connection status, service data processing method, and network traffic management status. The service data is adaptively processed when the edge computing node detects a specific operating status. The adaptive processing includes compression or aggregation. The metadata of the business data stream in the data packet is analyzed to identify whether there are any anomalies. When an anomaly is identified, an anomaly is generated in the data stream metadata. Based on the internal operating status information, a context-aware judgment is made on the abnormality of the data stream element information. The context-aware judgment includes: adjusting the judgment logic for the abnormality of the data stream element information based on the internal operating status information, and classifying the abnormality of the data stream element information as an interpretable anomaly or an external attack; the interpretable anomaly refers to an anomaly caused by the internal predicament of the edge computing node.
2. The cloud-based collaborative intrusion anomaly detection method according to claim 1, characterized in that, The step of performing context-aware judgment on the abnormality of the data stream element information based on the internal operating status information includes: Collect external network behavior information related to edge computing nodes. The external network behavior information includes the edge computing node's authentication attempt records at the network layer, the number of authentication failures, the traffic management logs imposed on it by network security policies, and the connection status and data transmission statistics recorded by network devices. The internal operating status information is correlated and compared with the external network behavior information to evaluate the consistency between the internal operating status information and the external network behavior information, and a consistency evaluation result is obtained. Based on the internal operating status information and the consistency assessment results, the judgment logic for the abnormality of the data stream element information is adjusted, and the abnormality of the data stream element information is classified as interpretability anomaly or external attack.
3. The cloud-based collaborative intrusion anomaly detection method according to claim 2, characterized in that, The step of adjusting the judgment logic for the abnormal data stream metadata based on the internal operating status information and the consistency assessment result, and classifying the abnormal data stream metadata as interpretable anomalies or external attacks, includes: External network behavior information from network devices from different vendors is processed in a unified manner to obtain unified external network behavior information. The unified processing includes converting logs of different formats into a standard format. Based on the time synchronization protocol and accuracy of the network devices for each external network behavior information, the timestamp of the unified external network behavior information is calibrated to obtain calibrated external network behavior information. Based on the source device type of the calibrated external network behavior information and its time-calibrated time deviation range, the weight of the consistency assessment between the internal operating status information and the calibrated external network behavior information is adjusted to obtain a weighted consistency assessment result. Based on the weighted consistency assessment results and internal operating status information, determine the type of anomaly in the data stream metadata.
4. The cloud-based collaborative intrusion anomaly detection method according to claim 3, characterized in that, The step of determining the type of anomaly in the data stream metadata based on the weighted consistency assessment result and internal operating status information includes: Collect operational behavior data of edge computing nodes; The cloud platform receives the operational behavior data and identifies the current operating mode of the edge computing node based on the operational behavior data. The operating mode includes normal operation, resource-constrained operation, or authentication-abnormal operation. The current operating mode of the edge computing node is identified, and the parameters in the judgment logic of the abnormal data stream metadata are adjusted to obtain the adjusted judgment logic. The parameters include judgment thresholds or rule weights. Based on the adjusted judgment logic, combined with the weighted consistency assessment results and the internal operating status information, the abnormal data stream metadata is classified.
5. The cloud-based collaborative intrusion anomaly detection method according to claim 4, characterized in that, The steps of identifying the current operating mode of the edge computing node and adjusting the parameters in the judgment logic for abnormal data stream metadata to obtain the adjusted judgment logic include: The internal operating status information reported by the edge computing nodes is analyzed from multiple dimensions to identify the fine-grained operating status of the edge computing nodes. The fine-grained operating status includes the classification of memory usage, the classification of authentication failure reasons, and the values of network interface traffic patterns and task scheduling latency. The multi-dimensional analysis refers to classifying memory usage, refining the classification of authentication failure reasons, or combining the values of network interface traffic patterns and task scheduling latency to identify the fine-grained operating status of the edge computing nodes. Based on the identified fine-grained operating state, a judgment threshold or rule weight matching the fine-grained operating state is selected from a preset parameter configuration library. Based on the matching judgment threshold or rule weight, the parameters in the judgment logic of abnormal data stream metadata are adjusted to obtain the adjusted judgment logic.
6. The cloud-based collaborative intrusion anomaly detection method according to claim 5, characterized in that, The step of performing multi-dimensional analysis on the internal operating status information reported by the edge computing node to identify the current fine-grained operating status of the edge computing node includes: The internal operating status information reported by the edge computing node is structured and parsed to obtain structured internal operating status information. The structured parsing refers to extracting memory usage, authentication failure reasons, network interface traffic patterns, and task scheduling delay, and converting the memory usage, authentication failure reasons, network interface traffic patterns, and task scheduling delay into a unified data structure. Based on a preset set of verification rules, the structured internal operating status information is subjected to consistency checks and integrity verification to obtain the results of the consistency checks and integrity verification. Based on the results of the consistency check and integrity verification, the internal operating status information is quality-marked or data-processed to obtain internal operating status information that has been quality-marked or data-processed. Based on the internal operating status information that has been marked with quality tags or processed by data, the fine-grained operating status of the edge computing node is identified.
7. The cloud-based collaborative intrusion anomaly detection method according to claim 6, characterized in that, The steps of performing consistency checks and integrity verifications on the structured internal operating status information based on a preset set of verification rules include: Receive novel device identification information or internal dilemma mode characteristics from the edge computing node. The novel device identification information indicates the model, manufacturer, or firmware version of the edge computing node, and the internal dilemma mode characteristics indicate an abnormal behavior pattern of the edge computing node that is not covered by the existing rule set. Based on the novel device identification information, the device operating parameters, data reporting specifications, and known internal distress manifestations corresponding to the novel device identification information are obtained from a preset device characteristic database; the internal distress manifestations include the data compression method of a specific model of device when memory is insufficient. Based on the internal dilemma pattern characteristics, analyze the differences between the internal dilemma pattern characteristics and the existing verification rule set to identify the coverage blind spots of the existing rule set; Based on the device operating parameters, data reporting specifications, known internal distress manifestations, and analysis results of coverage blind spots, generate or adjust the verification rules in the verification rule set. The verification rules include data range verification rules, timing consistency verification rules, or integrity check rules. Based on the verification rules in the generated or adjusted verification rule set, the structured internal operating status information is subjected to consistency checks and integrity verification.
8. The intrusion anomaly detection method based on cloud collaboration according to claim 7, characterized in that, The step of obtaining the equipment operating parameters, data reporting specifications, and known internal distress manifestations corresponding to the novel equipment identification information from a preset equipment characteristic database based on the novel equipment identification information includes: Receive identification information for new equipment; The new equipment identification information is subjected to format verification and standardization processing to obtain the processed new equipment identification information; Based on the processed new equipment identification information, a query is performed in the equipment characteristic database to obtain the corresponding equipment operating parameters, data reporting specifications, and known internal distress manifestations.
9. The cloud-based collaborative intrusion anomaly detection method according to claim 8, characterized in that, The step of performing format verification and standardization on the novel device identification information to obtain the processed novel device identification information includes: Receive identification information for new equipment; The character encoding of the novel device identification information is detected to identify the character encoding type used in the novel device identification information; Based on the identified character encoding type, the novel device identification information is decoded to obtain the decoded novel device identification information; Encrypted fields are identified in the decoded new device identification information to identify the encrypted fields present in the decoded new device identification information; The identified encrypted fields are decrypted to obtain the decrypted new device identification information; The decrypted new device identification information is then subjected to format verification and standardization processing.
10. A cloud-based collaborative intrusion anomaly detection system, characterized in that, The system includes: A data packet receiving module is used to receive data packets from an edge computing node. The data packets contain business data and embedded internal operating status information. The internal operating status information indicates the memory usage, authentication connection status, business data processing method, and network traffic management status of the edge computing node. The business data is adaptively processed when the edge computing node detects a specific operating status. The adaptive processing includes compression or aggregation. The metadata analysis module is used to analyze the metadata of the business data stream in the data packet to identify whether there is an anomaly. When an anomaly is identified, a data stream metadata anomaly is generated. The context-aware judgment module is used to perform context-aware judgment on the abnormality of the data stream element information based on the internal operating status information. The context-aware judgment includes: adjusting the judgment logic for the abnormality of the data stream element information based on the internal operating status information, and classifying the abnormality of the data stream element information as an interpretable abnormality or an external attack; the interpretable abnormality refers to an abnormality caused by the internal predicament of the edge computing node.