Real-time data exchange anomaly detection method and device based on deep learning

By employing a deep learning-based real-time data exchange anomaly detection method, which utilizes entropy-aware data encapsulation and hidden flow modulation techniques, the method addresses the shortcomings in accuracy and real-time performance of existing methods, achieving more efficient anomaly detection and data exchange security.

CN120956531AActive Publication Date: 2025-11-14JIANGSU YIJIESI INFORMATION TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511470460.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing methods for detecting anomalies in real-time data exchange are insufficient in terms of accuracy and real-time performance, making it difficult to adapt to complex and ever-changing network environments and data exchange scenarios, leading to missed detections and false detections.

Method used

A deep learning-based real-time data exchange anomaly detection method is adopted. This method generates pre-exchange data at the source end and encapsulates it with entropy-aware data. It also deploys a hidden flow protocol for hidden flow modulation and combines it with dual-branch detection technology to achieve real-time monitoring and anomaly detection of the data exchange process.

Benefits of technology

It improves the accuracy and real-time performance of anomaly detection and enhances the security and reliability of real-time data exchange.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956531A_ABST
    Figure CN120956531A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time data exchange anomaly detection method and device based on deep learning, and relates to the related field of data communication, and the method comprises the steps: an information source end generates pre-exchange data, and a trigger end encoder carries out entropy sensing data packaging according to a data gene, and determines exchange packaging data; the method comprises the following steps: deploying an implicit flow protocol in a switched network, activating along with data generation of an information source end, executing implicit flow modulation based on network background flow on switched encapsulation data, and generating implicit flow switched data; and performing exchange transmission from the information source end to the information sink end on the hidden stream exchange data, synchronously driving a data detection component, and performing double-branch verification along with an exchange process until data exchange is completed, the double-branch verification including first conformal anomaly verification and second anomaly access autophagy. The technical problem that existing real-time data exchange anomaly detection is insufficient in accuracy and real-time performance is solved, and the technical effects that the accuracy and the real-time performance of anomaly detection are improved, and then the safety and the reliability of real-time data exchange are improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data communication, and in particular to a method and apparatus for detecting anomalies in real-time data exchange based on deep learning. Background Technology

[0002] In today's era of rapid digital information development, the accuracy and security of real-time data exchange play a crucial role in the normal operation and decision-making of many key areas. Anomalies in data exchange can lead to economic losses and privacy breaches. Currently, the main method for detecting anomalies in real-time data exchange is the use of traditional rule-based matching and simple statistical analysis techniques. These techniques monitor and compare various indicators during the data exchange process by pre-setting a series of fixed rules and thresholds to determine if anomalies exist. However, these traditional methods, relying primarily on manually preset rules and fixed thresholds, are ill-suited to the complex and ever-changing network environment and data exchange scenarios. With the dynamic changes in network traffic, the emergence of new attack methods, and the increasing complexity of data structures, traditional methods cannot capture various potential anomaly patterns in a timely and accurate manner, resulting in poor accuracy and real-time performance, and a high likelihood of missed or false detections.

[0003] Currently, real-time data exchange anomaly detection suffers from insufficient accuracy and real-time performance. Summary of the Invention

[0004] This application provides a method and apparatus for real-time data exchange anomaly detection based on deep learning. It employs a method where the source generates pre-exchange data, the end encoder encapsulates the data based on its genetic entropy to obtain encapsulated exchange data, a hidden flow protocol is deployed in the exchange network and activated upon data generation, and the encapsulated data is modulated to generate hidden flow exchange data. This hidden flow exchange data is transmitted from the source to the destination, synchronously driving the data detection component. Dual-branch verification is performed throughout the exchange process until the data exchange is completed. These techniques address the shortcomings in accuracy and real-time performance of existing real-time data exchange anomaly detection methods, thereby improving the accuracy and real-time performance of anomaly detection and ultimately enhancing the security and reliability of real-time data exchange.

[0005] This application provides a real-time data exchange anomaly detection method based on deep learning, comprising: generating pre-exchange data at the source end, triggering an encoder at the trigger end, encapsulating entropy-aware data according to data genes, and determining exchange encapsulated data; deploying a hidden flow protocol in the exchange network, activating it as the data is generated at the source end, performing hidden flow modulation on the exchange encapsulated data based on network background traffic, and generating hidden flow exchange data; performing exchange transmission from the source end to the destination end on the hidden flow exchange data, synchronously driving a data detection component, performing dual-branch verification as the exchange process proceeds, until the data exchange is completed, wherein the dual-branch verification includes a first conformal anomaly detection based on operational morphology and a second anomaly access autophagy based on data genes and entropy values.

[0006] In a possible implementation, entropy-aware data encapsulation is performed based on data genes, and the following processing is carried out: defining gene tags, encoding the pre-exchange data into genes, defining data genes, wherein the gene tags at least include behavioral patterns and value metabolism rates; defining data entropy values ​​based on the information value and randomness of the pre-exchange data, wherein the data entropy values ​​include the entropy values ​​of each smallest data unit; and encapsulating the pre-exchange data based on the data genes and the data entropy values.

[0007] In a possible implementation, the pre-exchange data is encapsulated based on the data gene and the data entropy value, and the following processing is performed: the pre-exchange data is divided into blocks according to the data entropy value to determine the exchange data blocks; each exchange data block is encapsulated using the average entropy value as the first encapsulation code and the block data gene of the data exchange block as the second encapsulation code to generate the exchange encapsulated data.

[0008] In a possible implementation, the exchanged encapsulated data is subjected to hidden flow modulation based on network background traffic, and the following processes are performed: obtaining the exchanged network thread from the source end to the destination end; for the exchanged network thread, determining network statistical characteristics by monitoring background traffic, wherein the network statistical characteristics include at least packet size distribution, packet interval time, and packet sending rate; and triggering the hidden flow protocol to modulate the exchanged encapsulated data according to the network statistical characteristics.

[0009] In a possible implementation, the exchange encapsulation data is modulated and the following processing is performed: based on the network statistical characteristics, a first generator is assisted to generate a pseudo-random data packet sequence, wherein the pseudo-random data packet sequence contains multiple data packets consistent with the network statistical characteristics; based on the pseudo-random data packet sequence, the exchange encapsulation data is block-modulated to determine the hidden flow exchange data.

[0010] In a possible implementation, a synchronously driven data detection component performs dual-branch verification during the exchange process until the data exchange is completed, and performs the following processing: The first and second detection branches are run in parallel to build the data detection component, which is then embedded in the peripheral interface of the interactive network; as the source of the hidden flow exchange data is sent, the data detection component is retrieved and directional distillation training is performed to determine a short-term effective data detection card, wherein the data gene and data entropy value of the pre-exchanged data are used as the guide for directional distillation training; the data detection card is then transmitted along with the hidden flow exchange data.

[0011] In a possible implementation, the first detection branch is constructed and the following processing is performed: for the data operation set, invariant features are extracted for each data operation, and an abstract syntax tree is constructed; based on the abstract syntax tree, a high-dimensional morphological space is constructed; based on the high-dimensional morphological space, the first detection branch is constructed through sample deep learning.

[0012] In a possible implementation, the second detection branch is constructed, and the following processing is performed: using the data gene set, data autophagy rules under abnormal access are defined; using the data entropy level, autophagy relaxation is defined; based on the data autophagy rules and the autophagy relaxation, the second detection branch is constructed through sample deep learning.

[0013] In a possible implementation, after data exchange is completed, the following processing is performed: when data exchange is completed, the data detection card outputs an anomaly detection chain of the data exchange process, and at the same time, the data detection card is invalidated; the anomaly detection chain is interpreted to locate the anomaly pattern and the anomaly network location, and stored in the exchange feature library; according to a preset period, the exchange feature library is called to execute data exchange constraints based on the exchange network.

[0014] This application also provides a real-time data exchange anomaly detection device based on deep learning, comprising: an exchange encapsulation data determination module, used to generate pre-exchange data at the source end, trigger an encoder at the trigger end, perform entropy-aware data encapsulation based on data genes, and determine the exchange encapsulation data; a hidden flow modulation module, used to deploy a hidden flow protocol in the exchange network, activate it as the data is generated at the source end, perform hidden flow modulation on the exchange encapsulation data based on network background traffic, and generate hidden flow exchange data; and a dual-branch verification module, used to perform exchange transmission from the source end to the destination end on the hidden flow exchange data, synchronously drive a data detection component, and perform dual-branch verification as the exchange process progresses until the data exchange is completed, wherein the dual-branch verification includes a first conformal anomaly verification based on the operation mode and a second abnormal access autophagy based on data genes and entropy values.

[0015] The proposed real-time data exchange anomaly detection method and apparatus based on deep learning, as described in this application, firstly generates pre-exchange data at the source end, triggering an encoder to perform entropy-aware data encapsulation based on data genes, determining the exchange encapsulated data. Next, a hidden flow protocol is deployed in the exchange network and activated upon data generation at the source end. Hidden flow modulation based on network background traffic is performed on the exchange encapsulated data to generate hidden flow exchange data. Finally, the hidden flow exchange data is exchanged from the source end to the destination end, synchronously driving a data detection component to perform dual-branch detection throughout the exchange process until data exchange is complete. The dual-branch detection includes a first conformal anomaly detection based on operational morphology and a second anomaly access autophagy based on data genes and entropy values. This achieves the technical effect of improving the accuracy and real-time performance of anomaly detection, thereby enhancing the security and reliability of real-time data exchange. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the apparatus according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0017] Figure 1 This is a flowchart illustrating the real-time data exchange anomaly detection method based on deep learning provided in an embodiment of this application.

[0018] Figure 2 This is a schematic diagram of the structure of a real-time data exchange anomaly detection device based on deep learning provided in an embodiment of this application.

[0019] Explanation of reference numerals in the attached diagram: 10 for exchanging encapsulated data determination module, 20 for concealed current modulation module, and 30 for dual-branch verification module. Detailed Implementation

[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below.

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or apparatuses. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.

[0023] This application provides a real-time data exchange anomaly detection method based on deep learning, such as... Figure 1 As shown, the method includes:

[0024] In step S100, the source end generates pre-exchange data, triggers the encoder at the end, performs entropy-sensing data encapsulation based on the data genes, and determines the exchange encapsulation data.

[0025] Specifically, the source end is the originator of the data, that is, the party that generates the data to be exchanged. For example, in a financial trading system, this could be the client initiating the transaction. The pre-exchange data is the raw data generated by the source end and prepared for exchange, without any encapsulation or processing. The end encoder is an encoding device or algorithm module located at the source end, used to encapsulate and process the pre-exchange data.

[0026] First, data feature extraction is performed on the pre-exchange data to extract key information that uniquely identifies each data feature as its data genes. For example, for text data, keywords and thematic terms can be extracted using methods such as word frequency statistics and semantic analysis; for image data, edge features, texture features, and color distribution features can be extracted as data genes. Then, information entropy theory is applied to calculate the entropy value of each part of the pre-exchange data. Information entropy measures the uncertainty of data; a higher entropy value indicates greater uncertainty. By calculating the entropy values ​​of different data segments, the distribution and complexity of the data can be understood. For example, text data containing a large number of repeated characters has a relatively low entropy value; a randomly generated character sequence has a higher entropy value.

[0027] Based on the calculated entropy value and data genes, the end encoder encapsulates the pre-exchange data. For example, a layered encapsulation method is used, where the data genes are used as the identification information of the inner encapsulation, and the entropy-related parameters are used as the control information of the outer encapsulation, thus wrapping the pre-exchange data within them to form exchange-encapsulated data.

[0028] In one possible implementation, entropy-aware data encapsulation is performed based on data genes. Step S100 further includes step S110, defining gene tags, encoding the pre-exchange data genetically, and defining data genes. The gene tags at least include behavioral patterns and value metabolism rates. Specifically, behavioral patterns describe the operational characteristics of pre-exchange data in a specific business scenario or system environment. For example, in order data from an e-commerce platform, behavioral patterns can include the operational processes and patterns of different stages such as order creation, payment, delivery, and receipt. By defining behavioral patterns, data can be classified and identified from the perspective of operational processes. Value metabolism rate reflects the change in the value of data during the business process. Taking financial transaction data as an example, a transaction has high value when the transaction occurs, but its value gradually decreases over time and as the business progresses. Value metabolism rate is used to quantify the rate of this value change. For example, a time function can be set to describe the degree of data value decay over time. By defining the value metabolism rate, the importance and timeliness of the data can be understood.

[0029] Based on the defined genetic tags, the pre-exchange data is encoded. For example, binary encoding can be used, representing behavioral patterns and value metabolism rates with different binary bits. Assuming there are four different behavioral patterns, they can be represented by two binary bits; and value metabolism rates are divided into five levels, which can be represented by three binary bits. Combining these two encodings forms the genetic code for the pre-exchange data. This encoding result is the data gene, which uniquely identifies the characteristics of the pre-exchange data.

[0030] Step S120: Define a data entropy value based on the information value and randomness of the pre-exchanged data, wherein the data entropy value includes the entropy value of each smallest data unit. Specifically, information value measures the importance of the pre-exchanged data to business decisions or system operation. For example, in medical data, patient diagnoses and treatment plans have high information value, while some routine monitoring data, such as regular temperature records, have relatively low information value. The information value of the pre-exchanged data can be determined through expert evaluation, data analysis models, etc. For example, experts in relevant fields can be invited to score different types of data, or machine learning models can be used to predict information value based on historical data.

[0031] Randomness reflects the uncertainty and disorder of the pre-exchanged data. Information entropy theory is used to calculate the randomness of the pre-exchanged data. The pre-exchanged data is divided into multiple smallest data units, such as characters in text data or pixels in image data. For each smallest data unit, its probability of occurrence is calculated, and then the entropy value of each smallest data unit is calculated according to the information entropy formula. These entropy values ​​together constitute the data entropy value. The information value can be quantified into a weighting coefficient, and the calculated entropy value is adjusted according to this weighting coefficient to reflect the relative importance of the data in the overall data exchange and processing.

[0032] Step S130: Encapsulate the pre-exchange data according to the data genes and the data entropy value. Specifically, the encapsulation priority and method are determined based on data genes, behavioral patterns, value metabolism rates, and other information. For example, for data with complex behavioral patterns and high value metabolism rates, stricter encryption and protection measures are used for encapsulation to ensure data security and integrity. Simultaneously, combined with the data entropy value, specific encoding and compression algorithms are used for encapsulation of data with high entropy values ​​to optimize data storage and transmission efficiency. Following the established encapsulation strategy, the data genes and data entropy values ​​are integrated and encapsulated with the pre-exchange data as metadata. For example, the data genes and data entropy values ​​can be added to the header or footer of the pre-exchange data in a specific format to form a complete encapsulated exchange data. The encapsulated data has clear identification and characteristic information for transmission and anomaly detection in the network.

[0033] In one possible implementation, the pre-exchange data is encapsulated based on the data gene and the data entropy value. Step S130 further includes step S131, which involves dividing the pre-exchange data into blocks based on the data entropy value to determine the exchange data blocks. Specifically, the data entropy value reflects the information value and randomness of the pre-exchange data. Data with high entropy values ​​has higher randomness and complexity, containing more detailed information or noise; data with low entropy values ​​is relatively more ordered and has stronger regularity. Based on this characteristic, the pre-exchange data is divided into blocks according to the data entropy value. Specifically, a sliding window algorithm can be used for block division. First, a fixed-size window is determined, and the window slides from the starting position of the pre-exchange data. At each window position, the entropy value of the data within the window is calculated. If the entropy value of the data within the window is within a preset threshold range, the data within that window is considered as an exchange data block; if the entropy value exceeds the range, the window size or sliding step size is adjusted, and the entropy value is recalculated and judged until a suitable block boundary is found. Another method is to divide the data into blocks based on a clustering algorithm, treating each smallest data unit of the pre-exchange data as a data point and clustering it according to its entropy value. For example, the K-Means clustering algorithm is used to group data points with similar entropy values ​​into one class, with each class corresponding to a swapped data block.

[0034] Step S132: Using the average entropy value as the first encapsulation code and the block data element of the data exchange block as the second encapsulation code, each exchanged data block is encapsulated to generate the exchanged encapsulated data. Specifically, for each exchanged data block, the average entropy value of all its smallest data units is calculated. The average entropy value comprehensively reflects the overall randomness and information value level of the data block. For example, in a data block containing multiple text paragraphs, the entropy value of each character or word is calculated, and then the average value is calculated. This average value is used as the first encapsulation code to identify the characteristics of the data block in terms of randomness and information value. During the data exchange process, the receiver can quickly understand the overall characteristics of the data block based on the first encapsulation code. If the first encapsulation code shows that the average entropy value of the data block is high, it indicates that the data block may contain more important information or a complex structure, and the receiver can prioritize processing and parsing it; conversely, if the average entropy value is low, the receiver can process it according to the conventional process, thereby improving the efficiency and targeting of data processing.

[0035] The block data gene is specific gene information extracted for each exchanged data block, based on the gene encoding of the pre-exchange data in step S110. The block data gene serves as the second encapsulation code, uniquely identifying the source and characteristics of each exchanged data block. During data exchange, the receiver can identify the type and attributes of the data block using the second encapsulation code, ensuring correct data parsing and integration. For example, in data exchange scenarios involving multiple data sources, data from different sources may have similar entropy characteristics, but the block data gene can clearly distinguish them, avoiding data confusion and incorrect processing.

[0036] The first encapsulation code (mean entropy value) and the second encapsulation code (block data gene) are integrated with the exchanged data block according to a certain format to generate exchanged encapsulated data. For example, a fixed-length field can be added to the header of the exchanged data block to store the first and second encapsulation codes. The first and second encapsulation codes can be arranged in sequence or combined using a specific encoding method to ensure data integrity and readability. The generated exchanged encapsulated data has a clear structure and identification information, facilitating transmission and processing over the network. During transmission, even if some data is lost or damaged, the receiver can perform preliminary verification and recovery of the data using the encapsulation code, improving the reliability and security of data transmission.

[0037] Step S200: By deploying a hidden flow protocol in the switching network and activating it as data is generated at the source end, hidden flow modulation based on network background traffic is performed on the switched encapsulated data to generate hidden flow switched data.

[0038] Specifically, a switching network is a network system composed of multiple network nodes, such as routers and switches, used for data exchange and transmission. Hidden flow protocols are protocols used to hide the data exchange process within the network. Through specific modulation and transmission mechanisms, they make the exchanged data difficult for third parties to detect. Hidden flow protocol software modules are installed and configured on each node of the switching network to ensure the protocol functions correctly within the network. For example, the hidden flow protocol code can be integrated into the operating system of network devices, and relevant parameters such as modulation methods and activation conditions can be set through configuration files.

[0039] The hidden flow protocol is activated through a trigger mechanism based on data generation at the source. After the source generates pre-switched data and the encoder at the trigger end encapsulates it, a specific activation signal is sent to the switching network. Upon receiving this signal, nodes in the switching network check whether their locally configured hidden flow protocol meets the activation conditions. If so, the protocol is activated. For example, the activation signal can be a specific data packet containing identifier information for protocol activation; network nodes trigger protocol activation by parsing this data packet.

[0040] After activation, the hidden flow protocol modulates the encapsulated data based on the characteristics of the network background traffic. Network background traffic refers to all data traffic normally transmitted within the network, excluding the exchanged data that needs to be hidden. Examples include traffic generated by web browsing, file downloads, and video playback. Adaptive modulation technology is employed to adjust modulation parameters according to real-time changes in network background traffic, making them similar to the background traffic in terms of frequency, amplitude, and phase, thus blending into the background traffic and achieving the goal of data concealment. For example, if the network background traffic is predominantly high-frequency signals, the hidden flow protocol can use high-frequency modulation to modulate the encapsulated data; if the background traffic varies significantly, the protocol can dynamically adjust the modulation frequency and amplitude to ensure that the modulated data is difficult to detect. After hidden flow modulation processing, the resulting hidden flow exchanged data can be transmitted under the cover of network background traffic, making it difficult for external detection.

[0041] In one possible implementation, the exchanged encapsulated data undergoes hidden flow modulation based on network background traffic. Step S200 further includes step S210, obtaining the exchange network thread from the source to the destination. Specifically, the exchange network thread refers to the data transmission path from the source (data sender) to the destination (data receiver) and the associated set of network resources. This thread includes the nodes, links, and related network configuration information that the data passes through during transmission in the network. For example, in a large enterprise's wide area network, the source might be a server located in the Shanghai branch, and the destination might be a server located in the Beijing headquarters. The exchange network thread would then include a series of routers, switches, and the physical links between these two servers.

[0042] Network topology discovery tools, such as Simple Network Management Protocol (SMMP), can be used to collect configuration information of network devices and identify the connections between adjacent devices using link-layer discovery protocols, thereby constructing a complete network topology and determining the switching network threads. Alternatively, route tracing techniques can be employed, using the `traceroute` or `tracert` commands to send probe packets with different TTL (Time To Live) values ​​to obtain information about the routing nodes that data packets pass through from the source to the destination, thus determining the transmission path. In a software-defined network environment, the controller has a global network view that can directly obtain the switching network threads from the source to the destination and dynamically adjust the path based on network conditions.

[0043] Step S220: For the switching network thread, network statistical characteristics are determined by monitoring background traffic. These characteristics include at least packet size distribution, packet interval time, and packet transmission rate. Specifically, network packet capture tools, such as Wireshark, are used to capture data packets on the network interface and record detailed information, including header information and payload content. Background traffic data is obtained by deploying Wireshark at key nodes of the switching network thread for real-time packet capture. Protocols such as NetFlow or sFlow are used to collect network traffic statistics and send them to a traffic analysis server. The server can summarize and analyze this information to provide detailed statistical data on background traffic. Specifically, packet size distribution is obtained by statistically analyzing the proportion of data packets of different sizes in the background traffic during the monitoring period. For example, data packets are divided into several ranges based on size, such as 0-64 bytes, 65-1500 bytes, and above 1501 bytes, and the proportion of data packets in each range to the total number of data packets is calculated. Packet interval time is obtained by measuring the time interval between two adjacent data packets arriving at the receiving end and calculating the average interval time, minimum interval time, maximum interval time, and the distribution of interval times. The packet sending rate is obtained by calculating the number of data packets and the amount of data sent per unit time. For example, the packet sending rate can be obtained by counting the number of data packets collected over a period of time and dividing by the time interval.

[0044] Step S230: Based on the network statistical characteristics, the hidden flow protocol is triggered to modulate the exchanged encapsulated data. Specifically, the hidden flow protocol adjusts the transmission parameters of the exchanged encapsulated data using adaptive modulation technology based on the statistical characteristics of the background traffic. This includes: segmenting the exchanged encapsulated data into data packets with sizes similar to those commonly found in the background traffic, based on packet size distribution characteristics. For example, if most data packets in the background traffic are between 65 and 1500 bytes in size, the exchanged encapsulated data is segmented into data packets within this range for transmission. Adjusting the transmission interval of the exchanged encapsulated data packets according to the packet interval distribution pattern of the background traffic. For example, if the packet interval of the background traffic follows a certain probability distribution, the transmission interval of the exchanged encapsulated data packets is made to follow the same distribution. Matching the transmission rate of the exchanged encapsulated data with the transmission rate of the background traffic can be achieved by adjusting parameters such as the size of the transmission buffer and the size of the transmission window, making the transmission rate similar to that of the background traffic to avoid detection.

[0045] In one possible implementation, the exchanged encapsulated data is modulated, and step S230 further includes step S231, whereby, based on the network statistical characteristics, the first generator is assisted in generating a pseudo-random data packet sequence, wherein the pseudo-random data packet sequence contains multiple data packets consistent with the network statistical characteristics. Specifically, the first generator can employ a pseudo-random number generation algorithm such as a linear congruential generator, adjusting the algorithm parameters according to the network statistical characteristics to generate the pseudo-random number sequence through a linear equation. For example, suppose the packet size distribution of the background traffic monitored in step S220 is: 30% small packets (0-64 bytes), 60% medium packets (65-1500 bytes), and 10% large packets (over 1501 bytes). When generating the pseudo-random data packet sequence, the first generator will randomly generate data packets of different sizes according to this proportion. For example, when generating 100 data packets, approximately 30 small packets (0-64 bytes), 60 medium packets (65-1500 bytes), and 10 large packets (over 1501 bytes) will be generated. For example, if the average packet interval of the monitored background traffic is 50 milliseconds, and its distribution follows a normal distribution with a standard deviation of 10 milliseconds, the first generator, when generating a pseudo-random data packet sequence, will randomly generate the transmission interval of each data packet with a mean of 50 milliseconds and a standard deviation of 10 milliseconds. This ensures that the generated pseudo-random data packet sequence maintains consistency with the background traffic in terms of packet interval. For instance, if the background traffic transmits packets at a rate of 1000 packets per second, the first generator will control the speed at which it generates the pseudo-random data packet sequence to ensure that the number of data packets generated per unit time is similar to the transmission rate of the background traffic.

[0046] Step S232: Based on the pseudo-random data packet sequence, the exchange encapsulation data is block-modulated to determine the hidden flow exchange data. Specifically, the exchange encapsulation data to be transmitted is block-based according to certain rules. For example, the exchange data can be divided into several data blocks of matching size based on the size range of data packets in the pseudo-random data packet sequence. The block-based exchange encapsulation data is then filled into the data packets generated in the pseudo-random data packet sequence. During the filling process, it must be ensured that the filled data packets still conform to network statistical characteristics. For example, if the data packets in the pseudo-random data packet sequence are generated according to a certain packet interval time, then after filling the data blocks, the transmission interval time of these data packets must remain unchanged. In order to correctly identify and reassemble the hidden flow exchange data at the receiving end, some identification information can be added to the filled data packets. For example, a specific identification field can be added to the header of the data packet to identify that the data packet belongs to the hidden flow exchange data and the position information of the data packet in the original exchange data. At the receiving end, the data packets filled with exchange encapsulation data are reassembled according to the identification information in the data packets. First, the data blocks in the data packets are arranged in the correct order according to the position information in the identification information. Then, the arranged data blocks are pieced together to reconstruct the original exchanged encapsulated data. To ensure the integrity of the hidden stream exchanged data, a checksum mechanism can be added during transmission. For example, a checksum can be added to each data packet. After receiving the data packet, the receiving end calculates the checksum and compares it with the checksum sent by the sending end. If they match, it means that the data packet was not corrupted during transmission; if they do not match, retransmission or other error handling operations are required.

[0047] Step S300: Perform source-to-destination exchange transmission on the hidden flow exchange data, synchronously drive the data detection component, and perform dual-branch verification as the exchange process progresses until the data exchange is completed. The dual-branch verification includes a first conformal anomaly verification based on the operation mode and a second anomalous access autophagy based on data genes and entropy values.

[0048] Specifically, a reliable network transmission protocol, such as TCP, is employed to ensure that the hidden stream exchange data is transmitted accurately from the source to the destination. During transmission, the data is segmented, encapsulated, and numbered to ensure correct reception and reassembly at the destination. Data detection components are deployed at both the source and destination to detect any anomalies during transmission, and software programming is used to synchronize the data detection components with the data transmission process. When the hidden stream exchange data transmission begins, the data detection components are triggered to initiate the detection process. For example, a callback function is set in the data transmission program to automatically invoke the data detection components for detection when data is sent or received.

[0049] Two different detection methods are employed simultaneously to monitor the data transmission process, thereby improving the accuracy and reliability of anomaly detection. The first method, conformal anomaly detection, analyzes the operational behavior characteristics during data transmission and compares them with a normal operational model to detect any abnormal operations. Specifically, a normal operational model is established by analyzing the operational patterns of data transmission, such as data read / write operations, access frequency, and operation sequence. During detection, the actual operational behavior is compared with the normal model; if the operational behavior deviates from the normal model, it is determined to be an anomaly. For example, in a database system, the normal operational pattern involves data query, insertion, and update operations performed in a certain order, with access frequency within a certain range. If a sudden, excessive, and frequent access to a specific data table is detected, and the operation sequence is disordered, an anomaly may exist.

[0050] The second type of abnormal access autophagy utilizes the data's genetic information and entropy values ​​to monitor whether the data access behavior conforms to its characteristics and detects the existence of abnormal access. Specifically, it uses extracted genetic information and calculated entropy values ​​to monitor the access behavior of data during transmission. If the access behavior of data is found to be inconsistent with the characteristics reflected by the genetic information and entropy values, it is judged as abnormal. For example, if the genetic information of a certain data indicates that it should only be accessed within a specific time period, and the entropy value shows that its data distribution is relatively stable, but during the detection process it is found that the data is frequently accessed outside of the specific time period, and the data distribution changes significantly, then abnormal access may exist.

[0051] In one possible implementation, a synchronously driven data detection component performs dual-branch verification during the exchange process until the data exchange is completed. Step S300 further includes step S310, where the first and second detection branches are run in parallel to build the data detection component, which is then embedded in the peripheral interface of the interactive network. Specifically, two independent detection branches are constructed simultaneously, namely the first and second detection branches, each employing different detection methods to detect the data transmission process. Based on this, a comprehensive data detection component is built, integrating the functions of the two detection branches to achieve comprehensive detection of anomalies during data transmission. The constructed data detection component is embedded in the peripheral interface of the interactive network, which is a critical node for data transmission. By deploying the data detection component here, it is ensured that data can be detected in real time and effectively during data transmission, promptly identifying potential anomalies and guaranteeing the security and reliability of data exchange.

[0052] Step S320: As the hidden flow exchange data is sent from the source, the data detection component is retrieved and directed distillation training is performed to determine a short-term effective data detection card. The data genes and data entropy values ​​of the pre-exchanged data serve as the training guide for directed distillation. Specifically, when the hidden flow exchange data begins to be sent from the source, the data detection component deployed on the peripheral interface of the interactive network is immediately retrieved, and the directed distillation training process is initiated. Directed distillation training is a special training method used to quickly generate an effective detection model based on specific targets and features. During directed distillation training, the data genes and data entropy values ​​of the pre-exchanged data serve as the training guide. Data genes uniquely identify the features of the pre-exchanged data, and data entropy reflects the information value and randomness of the data. By using these two key pieces of information as the training guide, the trained data detection card can be more specifically adapted to the characteristics of the current data, improving the accuracy and efficiency of detection. After directed distillation training, a short-term effective data detection card is generated. The short-term validity indicates that the detection card is generated specifically for the characteristics of the currently transmitted hidden stream exchange data. It can detect whether there are any abnormalities in the data during transmission in real time and accurately. As the data transmission progresses, the data detection card can be adjusted or regenerated in a timely manner according to changes in data characteristics.

[0053] Step S330: The data detection card is transmitted along with the hidden flow exchange data. Specifically, the data detection card, trained and determined through directional distillation, is transmitted together with the hidden flow exchange data. This allows the data detection card to monitor the data in real time throughout the entire data exchange process. Regardless of which node the data is transmitted to, the data detection card remains active, promptly detecting any anomalies that may occur during transmission.

[0054] By integrating data detection cards with data transmission, end-to-end monitoring and management of the data exchange process is achieved. During data transmission, the data detection cards can monitor data operation behaviors and access status in real time based on the real-time status and characteristics of the data, combined with information such as data genes and entropy values ​​used in previous training. Once an anomaly is detected, corresponding measures can be taken promptly, such as issuing alarms or interrupting transmission, to ensure the security and reliability of the data exchange process. Furthermore, this detection and management method is fast and lightweight, allowing for flexible adjustments based on the characteristics of the current data without significantly impacting data transmission efficiency.

[0055] In one possible implementation, the first detection branch is constructed, and step S310 further includes step S311, which involves extracting invariant features for each data operation in the data operation set and constructing an abstract syntax tree. Specifically, during data transmission, a series of data operations are involved, constituting a data operation set. For example, in a file transfer system, data operations include reading files, writing files, and modifying file permissions; in a database operation scenario, these include data querying, inserting, updating, and deleting operations. To accurately analyze and detect data operations, it is necessary to analyze each data operation in the data operation set and extract its invariant features. Invariant features refer to characteristics that remain relatively stable during data operations and do not easily change with the operating environment and specific parameters. For example, for a file reading operation, its invariant features may include the type of object being operated on and the basic purpose of the operation; for a database query operation, invariant features may include the data table structure involved in the query and the basic logical relationship of the query. Based on the extracted invariant features, an abstract syntax tree is constructed. An abstract syntax tree is a data structure that represents the syntactic structure of a program or operation in a tree-like form. In data manipulation analysis, the root node of the abstract syntax tree represents the entire set of data operations, and its child nodes represent specific data operations. Each data operation node can also have its own child nodes, used to represent detailed attributes and parameters of the operation. For example, for a set of data operations that includes file read and write operations, the root node of the abstract syntax tree could be "file operation set," with child nodes for "read operation" and "write operation." The "read operation" node has child nodes representing attributes such as the file path to be read and the reading method, and the "write operation" node also has similar child nodes representing attributes such as the target path to be written and the content to be written.

[0056] Step S312: Construct a high-dimensional morphological space based on the abstract syntax tree. Specifically, the abstract syntax tree is mapped to a higher-dimensional space, i.e., a high-dimensional morphological space is constructed. The construction of the high-dimensional morphological space is based on the attributes and feature information of each node in the abstract syntax tree, with each node attribute in the abstract syntax tree serving as a dimension in the high-dimensional space. For example, if the operation types involved in the abstract syntax tree include read, write, query, and update, then the operation type can be considered as a dimension; if the operation object is a file, the file type, size, etc., can serve as other dimensions; operation parameters such as the number of bytes read and the query filtering conditions also correspond to different dimensions. In this way, each data operation in the abstract syntax tree is mapped to a point in the high-dimensional morphological space. Different data operations, due to their different attributes, will occupy different positions in the high-dimensional space. Normal operations, due to their certain regularity and stability, will exhibit a relatively concentrated distribution in the high-dimensional space; while abnormal operations, due to differences in attributes from normal operations, will occupy relatively off-center positions in the high-dimensional space. Constructing a high-dimensional morphological space provides a suitable input data representation for anomaly detection using deep learning algorithms, thereby improving the accuracy and effectiveness of detection.

[0057] Step S313: Based on the high-dimensional morphological space, a first detection branch is constructed through sample deep learning. Specifically, after the high-dimensional morphological space is constructed, deep learning is performed using sample data to construct the first detection branch, thereby detecting abnormal data operations. First, a large number of normal and abnormal data operation samples are collected, and these samples are mapped into the high-dimensional morphological space according to the method described above to obtain corresponding sample points. These sample data contain various types of data operation situations, where normal samples are used to train the model to learn the feature patterns of normal operations, and abnormal samples are used to help the model distinguish between normal and abnormal operations. Then, a deep learning algorithm is selected, such as a deep neural network, convolutional neural network, or recurrent neural network, depending on the characteristics of the data operation and the needs of anomaly detection. For example, if the data operation has obvious spatial structure features, a convolutional neural network is a better choice; if the data operation involves time series information, a recurrent neural network and its variants are more suitable. The sample data in the high-dimensional morphological space is input into the selected deep learning model for training. During the training process, the model continuously adjusts its parameters to learn the distribution characteristics of normal data operation samples in the high-dimensional space, as well as the differences between abnormal data operation samples and normal samples. Through extensive iterative training, the model gradually becomes able to accurately distinguish between normal and abnormal data operations. After training, this deep learning model is used as the core detection module of the first detection branch. During actual data transmission, when a new data operation occurs, it is mapped to a high-dimensional morphological space using the same method to obtain the corresponding point, which is then input into the trained deep learning model. The model judges the data operation based on its learned knowledge and outputs a detection result indicating whether it is an abnormal operation. In this way, the first detection branch can effectively detect abnormal operations during data transmission by leveraging the powerful feature learning and pattern recognition capabilities of deep learning algorithms, thus ensuring the security of data exchange.

[0058] In one possible implementation, the second detection branch is constructed, and step S310 further includes step S314, defining data self-evolution rules under abnormal access using a data gene set. Specifically, the data gene set is a high-level abstraction and generalization of the essential characteristics and inherent attributes of data, containing key information about the data at each stage of generation, storage, processing, and transmission, determining the essential characteristics of the data. For example, in medical data, the data gene set includes the types of sensitive patient information, the source of the data, and the data usage guidelines. In the case of abnormal data access, the data self-evolution rule refers to a self-protection and self-destruction mechanism inherent in the data itself. When abnormal access is detected, such as unauthorized access or abnormal reading caused by malicious attacks, the rules defined in the data gene set are triggered. These rules include under what abnormal access conditions data self-evolution is initiated, and the specific methods of self-evolution, such as encrypting the data, partially deleting the data, or completely destroying the data. For example, if the data gene set defines that frequent and large-scale data requests from unauthorized IP addresses are considered abnormal access, then the data self-evolution rule is triggered to encrypt the data involved in the request, preventing further data theft.

[0059] Step S315 defines autophagy relaxation based on data entropy level. Specifically, data entropy is an indicator that measures the degree of disorder or uncertainty in data. Data entropy level is the result of classifying and dividing data entropy values; different data entropy levels represent different degrees of disorder and uncertainty. Autophagy relaxation refers to the flexibility and leniency set according to the data entropy level during the execution of data autophagy rules. Different data entropy levels correspond to different autophagy relaxations; that is, for data with different degrees of disorder and uncertainty, the execution strength and method of data autophagy rules differ when facing abnormal access. For example, for data with a high entropy level, due to its inherent high uncertainty, a certain degree of autophagy relaxation is allowed in the event of abnormal access. That is, after triggering the data autophagy rule, a certain buffer time is given or a relatively mild autophagy method is adopted, such as data isolation rather than direct destruction. For data with a low entropy level, due to its relative order and importance, the autophagy relaxation is lower. Once abnormal access is detected, the data autophagy rules will be executed quickly and strictly to protect data security to the greatest extent.

[0060] Step S316: Based on the data autophagy rules and the autophagy relaxation, a second detection branch is constructed through sample deep learning. Specifically, sample deep learning utilizes a large amount of sample data for training and learning, enabling the model to automatically extract features and patterns from the data and make decisions and predictions based on these features and patterns. In this process, the defined data autophagy rules and autophagy relaxation are used as learning objectives and constraints, and input into the deep learning model for training. The sample data includes normal access data, various types of abnormal access data, and data access situations at different data entropy levels. Through sample deep learning, the model can learn the complex relationship between data autophagy rules, autophagy relaxation, and data access behavior, thereby constructing the second detection branch. The second detection branch is a supplement and enhancement to the original data detection system, capable of monitoring data access in real time based on the learned rules and relaxation. When data access behavior is detected, the second detection branch combines the data autophagy rules to determine whether it is abnormal access, and decides whether to trigger data autophagy and how to perform autophagy based on the data entropy level and autophagy relaxation. For example, when an access request is detected, the second detection branch first determines whether the request is an abnormal access based on the data autophagy rules. If so, it then decides whether to immediately encrypt the data or directly isolate the data based on the current data entropy level and the corresponding autophagy relaxation, thereby achieving dynamic protection of data security.

[0061] In one possible implementation, after data exchange is completed, the method further includes: when data exchange is completed, the data detection card outputs an anomaly detection chain of the data exchange process, and at the same time, the data detection card invalidates; the anomaly detection chain is interpreted to locate the anomaly pattern and the anomaly network location, and stored in the exchange feature library; according to a preset period, the exchange feature library is called to execute data exchange constraints based on the exchange network.

[0062] Specifically, the data detection card monitors all stages of data exchange in real time, including data transmission, processing, and storage, to detect any anomalies. Once the data exchange is complete, the data detection card outputs an anomaly detection chain, which records all detected anomalies and their related information. After outputting the anomaly detection chain, the data detection card performs invalidation; that is, after completing its monitoring and recording tasks, the data detection card no longer remains continuously active in the current data exchange process but enters a relatively idle or standby state, waiting for the next data exchange task to trigger its reactivation. This invalidation optimizes system resource allocation and prevents the data detection card from unnecessarily consuming excessive resources.

[0063] By using decoding algorithms and tools, the raw information in the anomaly detection chain is transformed into more readable and understandable anomaly patterns and anomaly network location information. Anomaly patterns include specific data tampering methods and abnormal data transmission frequencies; anomaly network locations refer to the specific network nodes or devices that generate abnormal behavior, such as a server or terminal device. For example, decoding algorithms can identify frequently occurring abnormal data access patterns, determine which node or device in the network these abnormal accesses originate from, and the data type targeted by the abnormal access. The decoded anomaly patterns and anomaly network location information are stored in an exchange feature database, a database specifically designed to store various characteristic information during data exchange, recording anomalies and other relevant characteristic information from historical data exchanges. By storing this information in the exchange feature database, references and bases can be provided for subsequent data exchanges, enabling better identification and handling of anomalies.

[0064] A preset period is defined, which can be set according to actual needs, such as daily, weekly, or monthly. At the end of each preset period, information from the exchange feature database is automatically retrieved. Based on the abnormal patterns and abnormal network location information stored in the exchange feature database, data exchange constraints based on the exchange network are executed, including data access restrictions on specific network nodes or devices, and transmission restrictions on specific data types. For example, if the exchange feature database records that a network node has repeatedly exhibited abnormal data access behavior in the past, that node can be more strictly monitored and restricted during subsequent data exchange processes to prevent it from engaging in abnormal operations again, thereby ensuring the security and stability of data exchange.

[0065] This application's embodiments employ techniques such as generating pre-exchange data at the source end, obtaining exchanged encapsulated data through data gene entropy perception and encapsulation at the end encoder, deploying a hidden flow protocol in the exchange network, activating it upon data generation, modulating the encapsulated data to generate hidden flow exchange data, transmitting the hidden flow exchange data from the source end to the destination end, synchronously driving the data detection component, and performing dual-branch verification along with the exchange process until data exchange is completed. These techniques solve the technical problems of insufficient accuracy and real-time performance in existing real-time data exchange anomaly detection, achieving the technical effect of improving the accuracy and real-time performance of anomaly detection, thereby improving the security and reliability of real-time data exchange.

[0066] In the above text, refer to Figure 1 A method for detecting anomalies in real-time data exchange based on deep learning, according to embodiments of the present invention, is described in detail. Next, reference will be made to... Figure 2 A deep learning-based real-time data exchange anomaly detection device according to an embodiment of the present invention is described.

[0067] The deep learning-based real-time data exchange anomaly detection device according to embodiments of the present invention addresses the technical problems of insufficient accuracy and real-time performance in existing real-time data exchange anomaly detection methods, thereby improving the accuracy and real-time performance of anomaly detection and ultimately enhancing the security and reliability of real-time data exchange. The deep learning-based real-time data exchange anomaly detection device includes: an exchange encapsulation data determination module 10, a hidden current modulation module 20, and a dual-branch detection module 30.

[0068] The exchange encapsulation data determination module 10 is used to generate pre-exchange data at the source end, trigger the encoder at the end, and encapsulate entropy-aware data according to the data gene to determine the exchange encapsulation data; the hidden flow modulation module 20 is used to deploy a hidden flow protocol in the exchange network, which is activated as the data is generated at the source end, and performs hidden flow modulation on the exchange encapsulation data based on the network background traffic to generate hidden flow exchange data; the dual-branch verification module 30 is used to perform exchange transmission from the source end to the destination end on the hidden flow exchange data, synchronously drive the data detection component, and perform dual-branch verification as the exchange process proceeds until the data exchange is completed, wherein the dual-branch verification includes a first conformal anomaly verification based on the operation mode and a second abnormal access autophagy based on the data gene and entropy value.

[0069] The detailed description of the specific configuration of the exchange-encapsulated data determination module 10 is explained below: As mentioned above, the exchange-encapsulated data determination module 10 further includes: a gene encoding unit for defining gene tags, encoding the pre-exchange data, and defining data genes, wherein the gene tags at least include behavioral patterns and value metabolism rates; a data entropy value definition unit for defining data entropy values ​​based on the information value and randomness of the pre-exchange data, wherein the data entropy values ​​include the entropy values ​​of each smallest data unit; and a data encapsulation unit for encapsulating the pre-exchange data based on the data genes and the data entropy values.

[0070] Specifically, the pre-exchange data is encapsulated based on the data gene and the data entropy value. The data encapsulation unit may further include: a data segmentation subunit for segmenting the pre-exchange data into blocks based on the data entropy value to determine the exchange data blocks; and a data block encapsulation subunit for encapsulating each exchange data block using the average entropy value as the first encapsulation code and the block data gene of the data exchange block as the second encapsulation code to generate the exchange encapsulated data.

[0071] The specific configuration of the hidden flow modulation module 20 is described in detail below: As mentioned above, the hidden flow modulation module 20 performs hidden flow modulation based on network background traffic on the exchanged encapsulated data. The hidden flow modulation module 20 may further include: a switching network thread acquisition unit for acquiring the switching network threads from the source end to the destination end; a network statistical feature determination unit for determining network statistical features for the switching network threads by monitoring background traffic, wherein the network statistical features include at least packet size distribution, packet interval time, and packet sending rate; and a modulation unit for triggering the hidden flow protocol to modulate the exchanged encapsulated data according to the network statistical features.

[0072] The modulation unit for modulating the exchange encapsulation data may further include: a pseudo-random data packet sequence generation subunit for assisting the first generator in generating a pseudo-random data packet sequence based on the network statistical characteristics, wherein the pseudo-random data packet sequence contains multiple data packets consistent with the network statistical characteristics; and a block modulation subunit for performing block modulation on the exchange encapsulation data based on the pseudo-random data packet sequence to determine the hidden flow exchange data.

[0073] The specific configuration of the dual-branch verification module 30 is described in detail below: As mentioned above, the synchronously driven data detection component performs dual-branch verification along with the exchange process until the data exchange is completed. The dual-branch verification module 30 may further include: a data detection component building unit for paralleling the first detection branch and the second detection branch, building a data detection component, and embedding the data detection component into the peripheral interface of the interactive network; a directional distillation training unit for retrieving the data detection component and performing directional distillation training as the source end of the hidden flow exchange data is sent, determining the short-term effective data detection card, wherein the data gene and data entropy value of the pre-exchanged data are used as the guide for directional distillation training; and a transmission unit for transmitting the data detection card along with the hidden flow exchange data.

[0074] The data detection component building unit, which builds the first detection branch, may further include: an abstract syntax tree construction subunit for extracting invariant features for each data operation set and building an abstract syntax tree; a high-dimensional morphological space construction subunit for building a high-dimensional morphological space based on the abstract syntax tree; and a first detection branch construction subunit for building the first detection branch based on the high-dimensional morphological space through sample deep learning.

[0075] The data detection component building unit for building the second detection branch may further include: a data autophagy rule definition subunit for defining data autophagy rules under abnormal access based on the data gene set; an autophagy relaxation definition subunit for defining autophagy relaxation based on the data entropy level; and a second detection branch building subunit for building the second detection branch through sample deep learning based on the data autophagy rules and the autophagy relaxation.

[0076] After data exchange is completed, the device may further include: an anomaly detection chain output module, which outputs an anomaly detection chain of the data exchange process when the data exchange is completed, and simultaneously, invalidates the data detection card; an anomaly location module, which interprets the anomaly detection chain, locates the anomaly pattern and the anomaly network location, and stores it in the exchange feature library; and a data exchange constraint module, which calls the exchange feature library according to a preset period and executes data exchange constraints based on the exchange network.

[0077] The deep learning-based real-time data exchange anomaly detection device provided in this embodiment of the invention can execute the deep learning-based real-time data exchange anomaly detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0078] Although this application makes various references to certain modules in the apparatus according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not intended to limit the scope of protection of this invention.

[0079] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A real-time data exchange anomaly detection method based on deep learning, characterized in that, The method includes: The source generates pre-exchange data, triggers the encoder at the end, performs entropy-sensing data encapsulation based on the data genes, and determines the encapsulated data to be exchanged. By deploying a hidden flow protocol in the switching network and activating it when data is generated at the source end, hidden flow modulation based on network background traffic is performed on the switched encapsulated data to generate hidden flow switched data. The hidden flow exchange data is transmitted from the source end to the destination end, and the data detection component is driven synchronously. Dual-branch verification is performed as the exchange process progresses until the data exchange is completed. The dual-branch verification includes a first conformal anomaly verification based on the operation mode and a second anomalous access autophagy based on data genes and entropy values.

2. The method as described in claim 1, characterized in that, Entropy-aware data encapsulation based on data genes includes: Define gene tags, encode the pre-exchange data into genes, and define data genes, wherein the gene tags include at least behavioral patterns and value metabolism rates; Based on the information value and randomness of the pre-exchanged data, a data entropy value is defined, wherein the data entropy value includes the entropy value of each smallest data unit; The pre-exchange data is encapsulated based on the data gene and the data entropy value.

3. The method as described in claim 2, characterized in that, Based on the data gene and the data entropy value, the pre-exchange data is encapsulated, including: Based on the data entropy value, the pre-exchange data is divided into blocks to determine the exchange data blocks; The average entropy value is used as the first encapsulation code, and the block data gene of the data exchange block is used as the second encapsulation code to encapsulate each exchange data block, thereby generating the exchange encapsulated data.

4. The method as described in claim 1, characterized in that, Performing hidden-flow modulation based on network background traffic on the exchanged encapsulated data includes: Obtain the network thread for the exchange between the source and the destination; For the aforementioned switching network thread, network statistical characteristics are determined by monitoring background traffic, wherein the network statistical characteristics include at least packet size distribution, packet interval time, and packet sending rate; Based on the network statistical characteristics, the hidden flow protocol is triggered to modulate the exchanged encapsulated data.

5. The method as described in claim 4, characterized in that, Modulating the exchanged encapsulated data includes: Based on the network statistical characteristics, the first generator is assisted in generating a pseudo-random data packet sequence, wherein the pseudo-random data packet sequence contains multiple data packets consistent with the network statistical characteristics; Based on the pseudo-random data packet sequence, the exchange encapsulation data is modulated in blocks to determine the hidden flow exchange data.

6. The method as described in claim 1, characterized in that, The synchronously driven data detection component performs dual-branch verification during the exchange process until the data exchange is completed, including: The first and second detection branches are parallelized to build a data detection component, which is then embedded in the peripheral interface of the interactive network. As the source of the hidden flow exchange data is sent, the data detection component is retrieved and directional distillation training is performed to determine a short-term effective data detection card, wherein the data gene and data entropy value of the pre-exchanged data are used as the guide for directional distillation training. The data detection card is transmitted along with the hidden stream exchange data.

7. The method as described in claim 6, characterized in that, The first detection branch includes: For each data operation set, invariant features are extracted for each data operation, and an abstract syntax tree is constructed. Based on the abstract syntax tree, construct a high-dimensional morphological space; Based on the high-dimensional morphological space, a first detection branch is constructed through sample deep learning.

8. The method as described in claim 6, characterized in that, The second detection branch includes: Define data autophagy rules under abnormal access using data gene sets; Autophagy relaxation is defined by the level of data entropy values; Based on the data autophagy rules and the autophagy relaxation, a second detection branch is constructed through sample deep learning.

9. The method as described in claim 6, characterized in that, After the data exchange is completed, the following will be included: Once the data exchange is complete, the data detection card outputs an anomaly detection chain for the data exchange process, and simultaneously, the data detection card resolves invalidity. The anomaly detection chain is interpreted to locate the anomaly pattern and anomaly network position, and stored in the exchange feature library; According to a preset period, the exchange feature library is invoked to execute data exchange constraints based on the exchange network.

10. A real-time data exchange anomaly detection device based on deep learning, characterized in that, The apparatus is used to implement the real-time data exchange anomaly detection method based on deep learning as described in any one of claims 1-9, the apparatus comprising: The data exchange and encapsulation determination module is used to generate pre-exchange data at the source end, trigger the encoder at the trigger end, perform entropy-sensing data encapsulation based on the data genes, and determine the exchange and encapsulation data. The hidden flow modulation module is used to perform hidden flow modulation on the exchanged encapsulated data based on the network background traffic by deploying a hidden flow protocol in the switching network and activating it when the data is generated at the source end, thereby generating hidden flow exchanged data. The dual-branch verification module is used to perform source-to-destination exchange transmission of the hidden flow exchange data, synchronously drive the data detection component, and perform dual-branch verification as the exchange process proceeds until the data exchange is completed. The dual-branch verification includes a first conformal anomaly verification based on the operation mode and a second anomalous access autophagy based on data genes and entropy values.

Citation Information

Patent Citations

  • Real-time data exchange anomaly detection method and device based on artificial intelligence

    CN118972280A

  • Intelligent file cabinet data exchange management and control method

    CN119066697A

  • Network anomaly monitoring method and system of switch

    CN119071052A

  • Nursing data sharing method and system based on Internet of Things

    CN120196604A

  • Civil aviation logistics data exchange method

    CN120223612A