Message feature extraction method and related equipment

By acquiring and parsing the original message data of network traffic and physical layer signal waveforms, building a cross attention weight matrix and generating target message characteristics, the problem of insufficient feature representation capabilities in the face of encrypted traffic and protocol obfuscation attacks is solved, and the ability to efficiently identify complex network attacks is achieved.

CN120165982AActive Publication Date: 2025-06-17BYZORO NETWORK LTD +1

Patent Information

Application Number
CN202510629653.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-17
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

When existing network traffic feature extraction technology faces encrypted traffic, private protocol nesting and protocol obfuscation attacks, the feature representation capability is limited, making it difficult to effectively identify hidden network abnormal behaviors.

Method used

By obtaining the original message data of the target network traffic and the physical layer signal waveform, analyzing the message data based on the preset protocol hierarchical rules, determining the multi-layer protocol field set and frequency domain feature vectors, constructing a cross attention weight matrix, and generating the target message features.

Benefits of technology

It improves the feature resolution capability of encrypted traffic, enhances the robustness of feature to protocol obfuscation and signal interference, and gives features endogenous security defense capabilities through anti-playback encoding and dynamic nonlinear transformation mechanisms, and can accurately identify complex network attack behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165982A_ABST
    Figure CN120165982A_ABST
Patent Text Reader

Abstract

The invention discloses a message feature extraction method and related equipment, and relates to the technical field of feature extraction, and the method comprises the steps: obtaining original message data of target network flow and a corresponding physical layer signal waveform; analyzing the original message data based on a preset protocol layering rule, and determining a multi-layer protocol field set; determining a protocol layer feature vector based on the multi-layer protocol field set; determining a frequency domain feature vector based on the physical layer signal waveform; constructing a cross attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; and generating a target message feature based on the cross attention weight matrix. According to the method, through dual-mode deep fusion of protocol semantics and physical signals, the analysis capability of the features on encrypted traffic and protocol confusion attacks is enhanced, meanwhile, the feature robustness is improved through dynamic weight distribution and an anti-replay mechanism, and endogenous security features with high discrimination power are provided for complex network threat detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of feature extraction, and particularly to a method for extracting packet features and related devices. Background Art

[0002] Current network traffic feature extraction technologies mainly rely on static parsing of protocol layer fields, such as field extraction based on fixed protocol templates or shallow statistical feature analysis. Traditional methods have limited feature representation capabilities when faced with encrypted traffic, private protocol nesting, and protocol obfuscation attacks, and it is difficult to effectively identify hidden network abnormal behaviors. Therefore, there is an urgent need for a method for extracting packet features to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further described in detail in the Detailed Implementation section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a method for extracting packet features, including: Obtain the original packet data of the target network traffic and the corresponding physical layer signal waveform; Parse the original packet data based on a preset protocol layering rule to determine a multi-layer protocol field set; Determine a protocol layer feature vector based on the multi-layer protocol field set; Determine a frequency domain feature vector based on the physical layer signal waveform; Construct a cross-attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; Generate target packet features based on the cross-attention weight matrix.

[0005] In some embodiments, parsing the original packet data based on a preset protocol layering rule to determine a multi-layer protocol field set includes: Determine a protocol identification segment based on the first preset length of bytes at the head of the original packet data; Determine a protocol parsing template based on the matching result between the protocol identification segment and the protocol feature library; Perform secondary parsing on the unmatched protocol fields in the original packet data based on a sliding window mechanism to determine the position markers of variable-length fields; Generate a protocol tree containing a multi-layer protocol structure based on the protocol parsing template and the position markers of variable-length fields; Determine a multi-layer protocol field set based on the hierarchical structure of the protocol tree.

[0006] In some embodiments, determining a protocol layer feature vector based on a multi-layer protocol field set includes: Determining a first semantic encoding based on the mapping relationship between the transport layer port number and the application layer protocol type in the multi-layer protocol field set; Determining a second semantic encoding based on the temporal correlation coefficient between the network layer TTL value and the data link layer frame interval in the multi-layer protocol field set; Determining a protocol layer feature vector based on the normalized concatenation result of the first semantic encoding and the second semantic encoding.

[0007] In some embodiments, determining a frequency domain feature vector based on the physical layer signal waveform includes: Dividing the frequency bands of the physical layer signal waveform based on the wavelet packet decomposition algorithm to determine the energy entropy of each frequency band; Determining a signal stability coefficient based on the decay rate of the autocorrelation function of the signal amplitude between adjacent packets; Determining a frequency domain feature vector based on the weighted calculation result of the energy entropy of each frequency band and the signal stability coefficient.

[0008] In some embodiments, constructing a cross-attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector includes: Determining a protocol timing transition probability based on the mapping relationship between the transport layer port number and the application layer protocol type in the protocol layer feature vector; Determining a frequency domain mutation intensity based on the distribution characteristics of the energy entropy of each frequency band in the frequency domain feature vector; Determining a cross-modal association matrix based on the protocol timing transition probability and the frequency domain mutation intensity; Performing temporal modeling on the cross-modal association matrix based on a gated recurrent unit to determine a feature dependence relationship; Determining a cross-attention weight matrix based on weight allocation for the feature dependence relationship using a self-attention mechanism.

[0009] In some embodiments, generating a target packet feature based on the cross-attention weight matrix includes: Determining a bimodal feature by weighted fusion of the protocol layer feature vector and the frequency domain feature vector based on the cross-attention weight matrix; Generating a pseudo-hash sequence based on the byte sampling result of the encrypted payload in the original packet data; Determining an anti-replay encoding based on the exclusive-or confusion operation between the pseudo-hash sequence and the bimodal feature; Determining a target packet feature based on non-linearly transforming the anti-replay encoding using a dynamic S-box substitution algorithm.

[0010] In some embodiments, it further includes: Based on the synchronization relationship between the timestamp and the physical layer signal waveform of the target packet feature, embed a watermark identifier to generate a target packet feature that includes anti-replay coding and the watermark identifier.

[0011] In a second aspect, the present application proposes a packet feature extraction device, including: A packet data acquisition unit for acquiring the original packet data of the target network traffic and the corresponding physical layer signal waveform; A protocol field determination unit for parsing the original packet data based on a preset protocol layering rule to determine a multi-layer protocol field set; A protocol feature generation unit for determining a protocol layer feature vector based on the multi-layer protocol field set; A frequency domain feature generation unit for determining a frequency domain feature vector based on the physical layer signal waveform; A weight matrix construction unit for constructing a cross-attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; A packet feature extraction unit for generating a target packet feature based on the cross-attention weight matrix.

[0012] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, it implements the steps of the packet feature extraction method according to any one of the first aspects above.

[0013] In a fourth aspect, the present application proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the packet feature extraction method according to any one of the first aspects.

[0014] In summary, the present application realizes the dual-modal deep feature fusion of protocol content and physical signals by fusing the semantic association features of protocol layer fields and the frequency domain analysis results of physical layer signals, constructing a dynamic cross-attention weight matrix. The present application improves the feature parsing ability of encrypted traffic, effectively enhances the robustness of features against protocol obfuscation and signal interference, and at the same time endows the features with an endogenous security defense ability through anti-replay coding and a dynamic non-linear transformation mechanism, and can accurately identify complex network attack behaviors, providing reliable technical support for high-concealment threat detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit this specification. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1Schematic flowchart of a method for extracting packet features provided by an embodiment of the present application; Figure 2 Schematic structural diagram of a device for extracting packet features provided by an embodiment of the present application; Figure 3 Structural diagram of an electronic device for extracting packet features provided by an embodiment of the present application. Detailed implementation manners

[0016] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0017] Please refer to Figure 1 , which is a schematic flowchart of a method for extracting packet features provided by an embodiment of the present application, and specifically may include: S110. Obtain the original packet data of the target network traffic and the corresponding physical layer signal waveform; Exemplarily, obtaining the original packet data of the target network traffic and the corresponding physical layer signal waveform is the basic data collection stage of packet feature extraction. The original packet data refers to the actual data packets in network transmission, including structured information such as protocol headers and payload contents, and is used to parse the semantic features of protocol fields; while the physical layer signal waveform is the electrical or optical signal manifestation when data is transmitted in the physical medium, reflecting physical characteristics such as the timing fluctuation and frequency distribution of the signal. The combination of the two can make up for the limitations of a single data dimension. The protocol layer data provides semantic information at the logical level, and the physical layer signal implies underlying features such as channel state and transmission stability, providing complementary data support for subsequent cross-modal feature fusion.

[0018] This step synergistically acquires dual-modal data of protocol content and physical signals, breaking through the limitation of traditional feature extraction that only relies on protocol fields. The original packet data covers the complete protocol stack information from the application layer to the physical layer, while the physical layer signal waveform captures hardware layer features such as signal strength and frequency offset. This dual-modal data acquisition not only enhances the diversity of feature sources but also lays a foundation for subsequent dynamic modeling of the interaction relationship between protocol and physical features, thereby improving the adaptability of feature expression to complex network environments.

[0019] S120. Parse the original packet data based on the preset protocol layering rules to determine the multi-layer protocol field set. Exemplarily, parsing the original packet data based on the preset protocol layering rules aims to extract the multi-layer protocol field set by structurally decomposing the protocol layers of network traffic. The protocol layering rules define the protocol stack parsing logic from the physical layer to the application layer. By identifying the packet header features and protocol interaction patterns, the key field information of different protocol layers is peeled off layer by layer to form a hierarchical protocol field set. This process overcomes the problem of insufficient adaptability of traditional static parsing methods to protocol nesting and dynamic changes, providing a structured data basis for subsequent feature vector generation.

[0020] Through the hierarchical parsing mechanism, it is possible to effectively handle the field ambiguity in private protocols or encrypted traffic. The preset rules combined with the dynamic matching strategy, based on identifying the known protocol identification segments, use the sliding window mechanism to make a secondary inference on unknown or variable-length fields, thereby constructing a complete protocol tree structure. This hierarchical parsing method not only enhances the accuracy of protocol field extraction but also provides hierarchical data support for cross-protocol layer semantic association analysis, ensuring that the feature extraction process adapts to the changing requirements of complex network environments.

[0021] S130. Determine the protocol layer feature vector based on the multi-layer protocol field set. Exemplarily, the construction of the protocol layer feature vector aims to extract highly discriminative feature expressions from the semantic association and temporal dependence relationships of multi-layer protocol fields. By analyzing the mapping relationship between the transport layer and the application layer, the semantic logic features of protocol interaction are extracted; at the same time, combined with the temporal dynamic parameters of the network layer and the data link layer, the context relevance of protocol behavior is captured. This cross-layer feature fusion mechanism breaks through the limitation of traditional single-layer field parsing and can more comprehensively depict the collaborative operation mode of the protocol stack, providing a semantically rich vector representation for subsequent feature fusion.

[0022] This step enhances the adaptability of features to encrypted traffic and protocol obfuscation attacks by dynamically modeling the multi-dimensional associations between protocol fields. Compared with the static field extraction method, the protocol layer feature vector not only contains the static attributes of the protocol content but also embeds the temporal dynamic characteristics of protocol interactions, thereby enhancing the robustness of feature representation against encrypted traffic and protocol obfuscation attacks and laying a foundation for the accurate recognition of complex network behaviors.

[0023] S140. Determine the frequency-domain feature vector based on the physical layer signal waveform; Exemplarily, the physical layer signal waveform contains the underlying characteristics of network traffic during transmission over the physical medium, such as the frequency distribution, energy intensity, and temporal fluctuation pattern of the signal. Through frequency-domain analysis, the time-domain signal can be mapped to the frequency dimension to extract the energy distribution characteristics and signal stability indicators of different frequency bands. The construction of the frequency-domain feature vector aims to quantify the energy concentration in the frequency-domain space, the interaction relationship between frequency bands, and the frequency-domain distortion characteristics caused by channel interference during transmission, thereby providing a complementary physical-dimensional representation for the protocol layer features and enhancing the adaptability of the features to hardware layer noise and transmission environment changes.

[0024] This step captures the dynamic behavior rules of the signal in the transmission medium through frequency-domain statistical features. For example, abnormal high-frequency band energy may indicate that the signal is subject to sudden interference, while the stability of the low-frequency band energy entropy is closely related to the channel quality. The frequency-domain feature vector can not only represent the instantaneous frequency-domain state of a single packet signal but also reveal the overall transmission trend of the signal through cross-packet correlation analysis, providing the underlying signal characteristics support for the subsequent deep fusion of protocol and physical features and enhancing the comprehensive representation ability of the features for complex network environments.

[0025] S150. Construct a cross-attention weight matrix based on the protocol layer feature vector and the frequency-domain feature vector; Exemplarily, the goal of constructing the cross-attention weight matrix is to dynamically quantify the association strength between the protocol layer features and the physical layer frequency-domain features and achieve deep interaction of the bimodal features. By analyzing the protocol interaction timing rules in the protocol layer feature vector (such as the dynamic mapping relationship between the transport layer port and the application protocol) and the sudden change characteristics of the signal frequency band energy in the frequency-domain feature vector, the internal coupling relationship between protocol behaviors and physical signal fluctuations is revealed. This matrix adaptively adjusts the contribution weights of different feature dimensions through cross-modal association modeling, providing dynamic guidance for subsequent feature fusion to ensure that the feature expression can accurately reflect the multi-dimensional characteristics of network traffic.

[0026] The construction mechanism of the cross-attention weight matrix enhances the context awareness ability of feature fusion. The semantic logic carried by the protocol layer features and the dynamic changes of the physical signals contained in the frequency domain features are jointly analyzed through temporal modeling and self-attention mechanism. This dynamic weight allocation strategy can reflect the synergy strength between the protocol layer and physical layer features, provide dynamic guidance for the weighted fusion of bimodal features, and thus improve the global consistency and environmental adaptability of feature representation.

[0027] S160. Generate target packet features based on the cross-attention weight matrix.

[0028] Exemplarily, generating target packet features based on the cross-attention weight matrix is the core step to optimize feature representation through dynamic fusion of protocol layer and physical layer features. The cross-attention weight matrix quantifies the correlation strength between protocol semantic features and physical signal characteristics, guides the weighted fusion process of bimodal features, and enables the protocol interaction logic and signal transmission dynamics to be jointly characterized. This step performs context-aware fusion of protocol layer feature vectors and frequency domain feature vectors through matrix operations to generate a comprehensive feature vector with both semantic richness and physical robustness, providing a highly discriminative feature input for subsequent network behavior analysis.

[0029] The generation of target packet features not only completes the integration of multi-dimensional features, but also endows the features with endogenous security through mechanisms such as anti-replay coding. Under the dynamic regulation of cross-attention weights, the semantic relevance in protocol layer features and the frequency domain stability of physical layer features are adaptively strengthened, effectively suppressing noise interference and attack confusion, ensuring that the generated features can accurately depict the essential behavior patterns of network traffic, and providing reliable feature support for complex threat detection.

[0030] In summary, the embodiments of the present application construct a dynamic cross-attention weight matrix by obtaining the original packet data of the target network traffic and its physical layer signal waveform, and combining the semantic features of the protocol layer and the frequency domain characteristics of the physical layer, realizing the dual-modal deep feature fusion of protocol content and physical signals. The embodiments of the present application break through the limitation of traditional feature extraction relying on a single protocol field. By synchronously collecting the structured information of the protocol layer and the physical layer signal waveform, a complementary input space is constructed from the dual dimensions of logical semantics and underlying signal characteristics. Based on the hierarchical parsing rules and dynamic matching strategies, a multi-layer protocol field set is extracted, and combined with the semantic mapping relationship between the transport layer and the application layer, and the timing correlation between the network layer and the data link layer, a protocol layer feature vector with cross-layer semantic association is generated; at the same time, the frequency domain dynamic characteristics of the physical layer signal are quantified by using the frequency band energy entropy and the signal stability coefficient to form a frequency domain feature vector. By dynamically modeling the correlation between the protocol timing transition probability and the frequency domain mutation intensity through the cross-attention weight matrix, the fusion weights of the dual-modal features are adaptively allocated, enhancing the global consistency and environmental adaptability of the feature expression. The generated comprehensive feature vector not only has the semantic richness of the protocol interaction logic, but also integrates the robustness of the physical signal transmission, improving the accuracy of encrypted traffic recognition and protocol obfuscation attack detection. In addition, through the anti-replay coding mechanism and dynamic non-linear transformation, the feature is endowed with an endogenous security defense ability, effectively suppressing replay attacks and noise interference, ensuring the accurate representation of the feature for complex network threats, and providing reliable technical support for the detection of highly concealed abnormal behaviors.

[0031] Obtaining the original packet data of the target network traffic and the corresponding physical layer signal waveform is the core step of data acquisition in the packet feature extraction method. The original packet data is captured in real time by a capture device (such as a network card, a probe, or a mirrored port) deployed on the network node, and the content of the data packets transmitted in the network traffic is completely recorded, including the protocol header fields, the payload data, and the metadata information. The original packet data covers the full protocol stack information from the physical layer to the application layer, and serves as the input for the protocol layer parsing, which is used to extract the semantic association features of the multi-layer protocol fields subsequently, such as the mapping relationship between the transport layer port number and the application layer protocol type, and the timing change rule of the network layer TTL value, etc., providing a structured data basis for the construction of the protocol layer feature vector.

[0032] The physical layer signal waveform is synchronously collected by a radio frequency sensor or a baseband chip, capturing the electrical or optical signal waveform when the data is transmitted in the physical medium (such as a cable, an optical fiber, or a wireless channel), specifically including the signal amplitude, frequency, phase, and timing fluctuation characteristics. The synchronous collection of the two is realized through a hardware clock synchronization mechanism, ensuring that the timestamps of the protocol content and the physical signal are strictly aligned, providing a data basis with consistent timing for cross-modal feature fusion.

[0033] The collaborative acquisition of bimodal data is achieved through the data association mechanism between the protocol layer and the physical layer. By binding the timestamp of the original packet data with the sampling moment of the physical layer signal waveform, a one-to-one correspondence of cross-layer data is established. This association mechanism ensures that the field information parsed by the protocol layer can be strictly synchronized with the signal characteristics of the physical layer in terms of time sequence, providing a data alignment basis for the construction of the subsequent cross-attention weight matrix. For example, a certain packet shows a communication behavior with a specific port number at the transport layer, and its corresponding physical layer signal waveform may exhibit a sudden increase in energy in the high-frequency band. After the two are associated through the timestamp, the dynamic coupling relationship between protocol interactions and signal physical characteristics can be revealed. This bimodal data acquisition not only breaks through the limitation of the traditional method that solely relies on protocol fields, but also supplements the noise, interference, and channel state information in the hardware transmission environment through the physical layer signal, providing multi-dimensional data input for dynamically modeling the internal relationship between protocol semantics and physical signals, and enhancing the adaptability of feature extraction to complex network environments.

[0034] In some instances, the original packet data is parsed based on the preset protocol layering rules to determine the multi-layer protocol field set, including: Based on the first preset length of bytes at the head of the original packet data, the protocol identification segment is determined; Based on the matching result between the protocol identification segment and the protocol feature library, the protocol parsing template is determined; Based on the sliding window mechanism, the un-matched protocol fields in the original packet data are re-parsed to determine the position markers of the variable-length fields; Based on the protocol parsing template and the position markers of the variable-length fields, a protocol tree containing a multi-layer protocol structure is generated; Based on the hierarchical structure of the protocol tree, the multi-layer protocol field set is determined.

[0035] Exemplarily, during the protocol layering parsing process, first, the first preset length of bytes is extracted from the head of the original packet data as the protocol identification segment. The first preset length is pre-set according to the length range of typical protocol identification fields. For example, the first 14 bytes of the Ethernet frame head or the first 20 bytes of the IP packet head. By intercepting the fixed-length byte stream at the head, the basic identification information of the current protocol layer can be quickly identified. This step utilizes the fixed-position characteristic of the protocol identification segment to provide the initial data input for subsequent protocol type matching, ensuring the efficiency and accuracy of the parsing process. For example, when parsing an Ethernet frame, the first 14 bytes contain the destination MAC address, source MAC address, and Ethernet type fields. By extracting this part of the bytes, the protocol family or protocol stack level to which the packet belongs can be initially identified, providing the basic identification information for subsequent protocol matching.

[0036] The protocol feature library stores the identification features of known protocol types and the corresponding parsing rules. The protocol identification segment extracted in Step 1 is matched with the identification features in the protocol feature library. If the match is successful, the corresponding protocol parsing template is loaded. The protocol parsing template defines the field structure, field length, and parsing order of the current protocol layer. For example, the parsing rules for the request line, header fields, and payload body of the HTTP protocol. This step ensures that the fields of known protocols can be accurately extracted through a templated parsing mechanism and provides structured guidance for subsequent hierarchical parsing. If no known protocol type is matched, it is marked as an unknown protocol and the subsequent secondary parsing mechanism is triggered. By dynamically loading predefined parsing templates, efficient parsing of standardized protocols is ensured, while providing extended processing capabilities for private protocols or encrypted traffic.

[0037] For protocol fields or variable-length fields where the protocol identification segment does not match (such as protocols with an unfixed payload length), a sliding window mechanism is used for secondary parsing. The sliding window slides byte by byte in the message data with a preset step size (such as 1 byte), and dynamically infers the field boundaries in combination with the context information. By calculating the statistical characteristics of the data within the window (such as field type identifiers, length field values, or checksum matching degrees), the start and end positions of variable-length fields are identified, and position markers are generated. For example, when parsing the IP option field, the option type and length fields are detected through the sliding window to determine the exact range of each option. The above effectively solves the problem of insufficient adaptability of traditional fixed-length parsing to variable-length protocol fields and improves the parsing accuracy in complex protocol nesting scenarios.

[0038] The protocol parsing template is combined with the variable-length field position markers to hierarchically parse the message data and construct a protocol tree. The nodes of the protocol tree correspond to the field information of each protocol layer, and the hierarchical relationship is determined by the protocol stack structure. For example, when parsing the TCP / IP protocol stack, the root node is the Ethernet frame, and the child nodes are the IP packet, TCP packet, and application layer data in sequence. Each node contains the field name, field value, and position range information of the protocol layer, and at the same time records the dynamic parsing results of variable-length fields. For example, after parsing the IP layer, according to the protocol type field in the IP header (such as the protocol number 6 for TCP), the TCP parsing template is called to continue parsing the transport layer fields. By structurally modeling the protocol hierarchical relationship, a complete protocol stack representation is formed, providing hierarchical data support for the extraction of multi-layer protocol fields.

[0039] Traverse the nodes of each layer of the protocol tree, extract the key field information of each layer of the protocol, and form a multi-layer protocol field set. The field set includes the identification field, control field, and payload data of each protocol layer. For example, the source IP address of the network layer, the destination port number of the transport layer, the HTTP method of the application layer, etc. Through the hierarchical association of the protocol tree, ensure that the field set completely reflects the context relationship of protocol interaction. For example, when parsing HTTPS traffic, the field set covers the TLS handshake protocol version, the list of cipher suites, and the certificate chain information, while retaining its association with the underlying TCP sequence number and IP fragmentation offset. Finally, output a structured and hierarchical protocol field set, providing an accurate input for the construction of protocol layer feature vectors, and ensuring the semantic integrity and cross-layer correlation of protocol layer features.

[0040] In the embodiments of the present application, through the preset protocol layering rules, combined with the protocol feature library matching and sliding window dynamic parsing mechanism, the structured layering parsing of the original message data is realized. From the extraction of the header identification segment to the positioning of variable-length fields, a multi-layer protocol field set is finally generated, ensuring that the protocol parsing process takes into account the efficient matching of standard protocols and the dynamic adaptation ability of non-standard protocols, and providing an accurate and complete protocol layer data basis for subsequent feature extraction.

[0041] In some instances, based on the multi-layer protocol field set, determine the protocol layer feature vector, including: Based on the mapping relationship between the transport layer port number and the application layer protocol type in the multi-layer protocol field set, determine the first semantic encoding; Based on the temporal correlation coefficient between the network layer TTL value and the data link layer frame interval in the multi-layer protocol field set, determine the second semantic encoding; Based on the normalized splicing result of the first semantic encoding and the second semantic encoding, determine the protocol layer feature vector.

[0042] Exemplarily, based on the mapping relationship between the transport layer port number and the application layer protocol type in the multi-layer protocol field set, the first semantic encoding is determined. Specifically, through a pre-set protocol mapping rule library, a static or dynamic association is established between the transport layer port number (such as TCP / UDP port number) and the application layer protocol type (such as HTTP, DNS, FTP, etc.). For example, port number 80 is mapped to the HTTP protocol type, and port number 53 is mapped to the DNS protocol type. Such mapping relationships are predefined through a protocol feature library. For non-standard ports or encrypted traffic, a sliding window matching algorithm is used to deeply analyze the application layer payload content, infer the actual protocol type, and update the mapping relationship library. The first semantic encoding is achieved by mapping the combination of the port number and the protocol type into a numerical vector of a fixed length. For example, one-hot encoding or embedding encoding techniques are used to convert the discrete protocol type identifiers into a continuous vector space representation. This encoding not only represents the explicit correspondence between the port number and the application protocol but also enhances the adaptability to non-standard protocols through a dynamic inference mechanism, ensuring the integrity and scalability of the semantic encoding.

[0043] Based on the temporal correlation coefficient between the Time To Live (TTL) value in the network layer and the frame interval in the data link layer in the multi-layer protocol field set, the second semantic encoding is determined. First, the temporal sequence of the network layer TTL field and the temporal sequence of the data link layer frame interval are extracted, and the two are aligned through timestamps. By calculating the Pearson correlation coefficient between the TTL change and the frame interval fluctuation within a sliding time window, the dynamic association strength between the two is quantified. For example, the hop-by-hop decreasing rule of the TTL value and the sudden increase in the data link layer frame interval may reflect network congestion or routing path changes. After the correlation coefficient is normalized, it is mapped into a numerical vector of a fixed dimension to form the second semantic encoding. The second semantic encoding captures the co-variation pattern of the behaviors of the network layer and the data link layer by embedding the temporal dynamic characteristics of protocol interactions, providing a context-aware temporal representation for the protocol layer feature vector.

[0044] The first semantic encoding and the second semantic encoding are respectively normalized to eliminate the dimension difference. Specifically, the Z-score normalization method is adopted to adjust each encoding dimension to a distribution with a mean of 0 and a variance of 1. Subsequently, the normalized first semantic encoding and the second semantic encoding are concatenated in the dimension order to form a high-dimensional composite vector. For example, if the first encoding dimension is 128 dimensions and the second encoding dimension is 64 dimensions, a 192-dimensional protocol layer feature vector is generated after concatenation. This vector realizes multi-dimensional feature representation across protocol layers by fusing the semantic mapping relationship between the transport layer and the application layer and the temporal correlation between the network layer and the data link layer. The finally output protocol layer feature vector has both the static semantic logic and the dynamic temporal characteristics of protocol interaction, providing hierarchical and structured input data for the construction of the subsequent cross-attention weight matrix and ensuring the robustness of feature expression against complex protocol nesting and encrypted traffic.

[0045] In some instances, based on the physical layer signal waveform, a frequency domain feature vector is determined, including: Based on the wavelet packet decomposition algorithm, the physical layer signal waveform is divided into frequency bands to determine the energy entropy of each frequency band; Based on the decay rate of the autocorrelation function of the signal amplitude between adjacent packets, the signal stability coefficient is determined; Based on the weighted calculation result of the energy entropy of each frequency band and the signal stability coefficient, the frequency domain feature vector is determined.

[0046] Exemplarily, based on the wavelet packet decomposition algorithm, the physical layer signal waveform is divided into frequency bands to determine the energy entropy of each frequency band. Specifically, a preset wavelet basis function (such as Daubechies wavelet) and the decomposition level are used to perform multi-level decomposition on the physical layer signal waveform, and the signal is gradually divided into sub-frequency bands ( is the decomposition level). For example, after performing 3-level wavelet packet decomposition, the original signal is divided into 8 frequency bands, and each frequency band corresponds to a signal component in a specific frequency range. The wavelet packet coefficients of each sub-frequency band are normalized to obtain the energy probability distribution of each frequency band, and the energy entropy of each frequency band is quantified by the Shannon entropy formula, expressed as: Among them, is the energy entropy of the th frequency band, characterizing the distribution complexity of the signal energy in this frequency band; is the energy proportion of the th wavelet packet coefficient in the th frequency band; is the total number of coefficients. The energy entropy characterizes the distribution complexity of the signal energy in each frequency band. A sudden increase in the entropy value in the high-frequency band may reflect signal interference or distortion, while a stable entropy value in the low-frequency band indicates the smoothness of the channel transmission.

[0047] The signal stability coefficient is determined based on the decay rate of the autocorrelation function of the signal amplitude between adjacent messages. Specifically, the physical layer signal amplitude sequence corresponding to the adjacent messages is extracted, and its autocorrelation function is calculated. ,in, is the time delay parameter. By fitting the decay curve of the autocorrelation function with time delay (such as using an exponential decay model ), calculate the decay rate parameter Specifically, the logarithm of the autocorrelation function is taken and linear regression is performed. The absolute value of the slope is the decay rate, and the signal stability coefficient is defined as ,in, A small constant (such as 1e-6) to prevent division by zero. The larger the value, the more drastic the signal amplitude fluctuation. The smaller the value, the smoother the signal transmission. This coefficient quantifies the temporal consistency of the signal amplitude between adjacent messages and provides a dynamic stability indicator for the frequency domain characteristics.

[0048] The energy entropy of each frequency band Signal stability factor Perform linear fusion according to preset weights to generate frequency domain feature vectors , assuming the total number of frequency bands is , The weight distribution strategy is expressed as: Frequency domain eigenvector It is composed of the weighted energy entropy and stability coefficient of each frequency band, expressed as: The energy entropy of each frequency band is assigned a preset weight, and the signal stability coefficient is independently assigned a fixed weight. The energy entropy of each frequency band is fused with the stability coefficient through a linear weighted formula to generate a comprehensive frequency domain feature vector. The energy entropy of 8 frequency bands and 1 stability coefficient are weighted and spliced ​​into a 9-dimensional feature vector. Through adaptive weight allocation, frequency bands with complex energy distribution and signal stability indicators jointly characterize the physical layer signal characteristics. The final output frequency domain feature vector contains both the static statistical characteristics of the frequency domain energy distribution and the embedded dynamic stability information of the signal transmission, providing a robust physical layer representation for subsequent cross-modal feature fusion.

[0049] In some instances, a cross-attention weight matrix is ​​constructed based on the protocol layer feature vector and the frequency domain feature vector, including: Determine the protocol timing transition probability based on the mapping relationship between the transport layer port number and the application layer protocol type in the protocol layer feature vector; Based on the distribution characteristics of energy entropy of each frequency band in the frequency domain feature vector, the frequency domain mutation intensity is determined; Determine the cross-modal correlation matrix based on the protocol timing transition probability and frequency domain mutation intensity; Based on the gated recurrent unit, the cross-modal correlation matrix is ​​temporally modeled to determine the feature dependencies; Based on the attention mechanism, the feature dependencies are weighted and the cross-attention weight matrix is ​​determined.

[0050] Exemplarily, the protocol timing transition probability is used to quantify the dynamic association strength between the transport layer port number and the application layer protocol type. The mapping relationship between the transport layer port number and the application layer protocol type is extracted from the protocol layer feature vector, and the relationship is established through a predefined protocol mapping rule base or a real-time parsing mechanism. For example, port number 80 usually corresponds to the HTTP protocol, and port number 443 corresponds to the HTTPS protocol. For non-standard ports or encrypted traffic, the actual protocol type is dynamically inferred by analyzing the application layer payload content, and the mapping relationship is updated. The calculation of the protocol timing transition probability is based on the co-occurrence frequency of the port number and the application protocol type in the sliding time window, which is specifically achieved by counting the proportion of the number of co-occurrences of the port number and the application protocol type in a specific time period to the total number of occurrences. This probability value reflects the dynamic distribution characteristics of the application protocol type under a specific port number, and characterizes the timing law of protocol interaction.

[0051] The frequency domain mutation intensity is used to characterize the abnormal fluctuation degree of the physical layer signal in the frequency domain space. From the energy entropy distribution of each frequency band generated by wavelet packet decomposition, the standard deviation or range of the energy entropy in the sliding time window is calculated to identify the mutation phenomenon of the energy distribution. For example, when the standard deviation of the energy entropy of a certain frequency band exceeds the preset threshold, it is determined that there is a significant fluctuation in the frequency band. The calculation of the frequency domain mutation intensity combines the ratio of the change amplitude of the energy entropy of each frequency band to its maximum value to quantify the stability of the signal in the frequency domain dimension. The larger the mutation intensity value, the more drastic the energy fluctuation of the physical signal in this frequency band, which may be caused by channel interference or signal distortion.

[0052] The cross-modal correlation matrix constructs a dynamic correlation model of protocol layer and physical layer features by integrating the protocol timing transition probability and the frequency domain mutation intensity. Specifically, the protocol transition probability matrix is ​​expanded into a one-dimensional vector and the outer product operation is performed with the frequency domain mutation intensity vector to generate the initial correlation matrix. Subsequently, the matrix elements are weighted and modified, where linear weights are used to strengthen significant correlation areas and nonlinear weights are used to capture complex interaction patterns. For example, areas where the protocol interaction pattern is strongly correlated with high-frequency band mutations will be given higher weights.

[0053] The final output correlation matrix comprehensively characterizes the coordinated variation relationship between protocol behavior and physical signal fluctuations, providing input for subsequent timing modeling.

[0054] The gated recurrent unit (GRU) is used to capture the temporal dependencies of the cross-modal association matrix. The cross-modal association matrix is input into the GRU network in a time series, and the fusion ratio of historical information and the current input is dynamically controlled through the update gate and the reset gate. The update gate determines how much of the historical hidden state to retain, and the reset gate controls the degree of correction of the current input to the historical state. The hidden state of the GRU gradually encodes the dynamic evolution law of the cross-modal association relationship, forming a feature vector reflecting the temporal dependence relationship. Through this process, the co-variation pattern between the protocol layer features and the physical layer features is effectively modeled, providing temporal context information for weight allocation.

[0055] The self-attention mechanism generates a cross-attention weight matrix by dynamically allocating weights for feature dependencies. Taking the feature dependency vector output by the GRU as input, the query matrix, key matrix, and value matrix are calculated respectively. The relevance scores at different time steps are calculated through scaled dot-product attention and normalized to form an attention weight matrix. This matrix dynamically adjusts the fusion weights between the protocol layer features and the physical layer features, strengthens the contribution of key feature dimensions, and suppresses noise interference. The finally generated cross-attention weight matrix accurately quantifies the association strength of the bimodal features, guides the deep fusion of the protocol content and the physical signal features, and ensures that the generated target message features have both the integrity of semantic logic and the robustness of physical transmission.

[0056] In the embodiments of this application, the temporal transfer law is extracted from the protocol layer features in sequence, the frequency domain mutation intensity is quantified from the physical layer signals, a cross-modal association matrix is constructed through tensor product operations, and the GRU and self-attention mechanisms are used to achieve temporal modeling and dynamic weight allocation. The mapping relationship between the port number and protocol type in the protocol layer features and the energy entropy fluctuation characteristics of the physical layer signals form a dynamic representation of cross-modal feature interaction through the correlation analysis of temporal transfer probability and mutation intensity. The GRU network captures the temporal co-variation of protocol and physical features, and the self-attention mechanism further optimizes the feature fusion weights. The finally generated cross-attention weight matrix provides dynamic guidance for the efficient fusion of multi-dimensional features, ensuring the strong adaptability of feature expression to complex network environments.

[0057] In some instances, based on the cross-attention weight matrix, target message features are generated, including: Weighted fusion of the protocol layer feature vector and the frequency domain feature vector based on the cross-attention weight matrix to determine bimodal features; Generating a pseudo-hash sequence based on the byte sampling results of the encrypted payload in the original message data; Determining the anti-replay encoding based on the exclusive-or confusion operation between the pseudo-hash sequence and the bimodal features; Determining the target message features based on the non-linear transformation of the anti-replay encoding using the dynamic S-box substitution algorithm.

[0058] Exemplarily, the cross-attention weight matrix dynamically represents the correlation strength between the protocol layer features and the physical layer frequency domain features, providing weight guidance for the fusion of bimodal features. Specifically, the protocol layer feature vector is multiplied by the cross-attention weight matrix through matrix multiplication to obtain the weighted vector of the protocol layer features; at the same time, the frequency domain feature vector is subjected to a similar operation with the transpose of the cross-attention weight matrix to obtain the weighted vector of the frequency domain features. The weighted protocol layer features and frequency domain features are fused by element-wise addition or concatenation to generate bimodal features. For example, the dimension of the protocol layer feature vector is 192, and the frequency domain feature vector is 9. After weighted fusion, a 201-dimensional bimodal feature vector is generated. The elements in the cross-attention weight matrix represent the correlation strength between the protocol layer features and the frequency domain features in different dimensions. The high-weight dimensions indicate the collaborative region of the key features. For example, if the mapping relationship between the transport layer port number and the application protocol type in the protocol layer features has a strong correlation with the high-frequency band energy entropy mutation in the frequency domain features, the corresponding weight value will increase significantly, ensuring that the fused features have both the integrity of logical semantics and the stability of physical transmission.

[0059] The byte sampling of the encrypted payload aims to extract its key structural features while avoiding direct exposure of the encrypted content. Specifically, first, according to the preset position rule, equal-interval sampling is performed within the range from the X-th byte to the Y-th byte after the starting offset of the encrypted payload. Among them, the values of X and Y are predefined according to the characteristics of the encryption protocol or actual requirements. For example, X = 16 and Y = 80, covering the payload header and key parameter areas. The sampling interval is dynamically adjusted according to the total length of the payload to ensure complete coverage of the key structural area. For example, if the payload length is 200 bytes, the sampling interval can be set to 5% of the total length (i.e., sampling once every 10 bytes), so as to capture the distribution characteristics of the payload at limited sampling points.

[0060] After sampling, the obtained byte sequence is input into a preset hash function (such as SHA-256) to generate a pseudo-hash sequence. The reason for choosing SHA-256 is its strong collision resistance, fixed output length (32 bytes), and wide applicability in security scenarios. The sampled byte sequence needs to be preprocessed according to the requirements of the hash function. For example, sequences with insufficient length are padded (PKCS#7 standard) to ensure that the input conforms to the data block size of the hash operation. Through the hash operation, the original byte sequence is compressed into a fixed-length pseudo-hash sequence. This sequence not only represents the local structural features of the payload but also cannot reverse-derive the original content due to the one-way nature of the hash function, thus avoiding the leakage of sensitive information.

[0061] The generation process of the pseudo-hash sequence includes calculating the sampling interval in real time according to the payload length to ensure that key areas (such as encryption algorithm parameters and initialization vectors) are covered; performing normalization processing on the sampled bytes (such as byte alignment and padding) to meet the input requirements of the hash function; inputting the processed data into the SHA-256 algorithm to generate a 32-byte pseudo-hash sequence; and performing secondary encoding (such as Base64) on the hash result for efficient processing of subsequent confusion operations. Through the above mechanism, the pseudo-hash sequence not only retains the key structural information of the encrypted payload but also achieves anti-tampering protection through the irreversible characteristic of the hash function. This sequence provides a stable digest input for subsequent XOR confusion operations, ensuring that the generated anti-replay encoding is unique and dynamic, thereby effectively resisting replay attacks and feature forgery.

[0062] The core of the XOR confusion operation is to generate a dynamic anti-replay encoding by fusing the differential information of the bimodal features and the pseudo-hash sequence. Specifically, first, the bimodal feature vector (such as a 201-dimensional floating-point vector) and the pseudo-hash sequence (such as a 32-byte binary sequence) are aligned. Since bimodal features are usually represented in floating-point form, they need to be converted into a binary bit sequence to match the format of the pseudo-hash sequence. For example, each floating-point number is converted into a 32-bit binary representation according to the IEEE 754 standard, and the 201-dimensional feature vector will be extended to a 201×32 = 6432-bit binary sequence. If the length of the pseudo-hash sequence (such as 256 bits) is less than the length of the feature vector binary sequence, the pseudo-hash sequence is circularly extended, that is, the pseudo-hash sequence is repeatedly concatenated until its length is the same as that of the feature vector binary sequence. For example, the 256-bit pseudo-hash sequence is circularly extended to 6432 bits to ensure that the dimensions of the two are strictly matched.

[0063] After the data alignment is completed, a bitwise exclusive OR (XOR) operation is performed on the two. The XOR operation rule is that if the corresponding bit values are the same, 0 is output, and if they are different, 1 is output. For example, if a certain bit of the feature binary sequence is 1 and the corresponding bit of the pseudo-hash sequence is 0, the XOR result is 1. Through the bitwise operation, the semantic information of the bimodal features and the local structural features of the pseudo-hash sequence are deeply fused to generate the binary sequence of the anti-replay encoding. This process breaks the predictability of the features through the confusion mechanism, making the encoding generated by the same message at different transmission times or under different channel conditions have significant differences, thereby resisting attacks launched by attackers by replaying the same features.

[0064] To ensure the irreversibility of the confusion result and its anti - analysis ability, post - processing of the generated binary sequence is required. Specifically, the XOR result is converted into a byte sequence according to a preset grouping length (such as 8 bits per group). If there are padding bits, the padding bits are removed. For example, a 6432 - bit XOR result is converted into an 804 - byte replay - resistant encoding. This encoding not only contains the fusion information of protocol - layer and physical - layer features but also embeds the digest characteristics of the pseudo - hash sequence, making it difficult for attackers to restore the original features through reverse analysis or forge valid encodings.

[0065] The finally generated replay - resistant encoding implements a security protection mechanism, including that the sampling position of the pseudo - hash sequence and the hash operation result change dynamically with the message content to ensure the timeliness of the encoding; through cyclic expansion and type conversion, the dimensional differences between the dual - mode features and the pseudo - hash sequence are resolved to ensure the integrity of the operation; the XOR operation destroys the statistical rules of the feature sequence, making it impossible for attackers to infer future encodings from historical data. This step endows the target message features with an endogenous security attribute through multi - level data transformation and confusion, provides an anti - tampering input basis for subsequent dynamic S - box replacement, and comprehensively enhances the defense ability of features against replay attacks.

[0066] The dynamic S - box replacement algorithm performs irreversible confusion processing on the replay - resistant encoding by introducing a key - driven non - linear permutation rule, and finally generates target message features with endogenous security attributes. Specifically, the generation of the dynamic S - box is based on preset dynamic parameters (such as timestamps or random seeds) to ensure that the replacement rule is unique each time the features are generated. For example, using the hash value of the current timestamp as a seed, a 16×16 S - box matrix is dynamically generated through a pseudo - random number generation algorithm (such as the AES - CTR mode). Each element (from 0x00 to 0xFF) of this matrix is randomly permuted, and the generated S - box is only applicable to the current message, thus avoiding the risk of reverse analysis of static S - boxes.

[0067] Each byte of the replay - resistant encoding (such as 0xAB) is split into the high 4 bits (0xA) and the low 4 bits (0xB), which are used as the row index and column index of the S - box respectively. For example, in the dynamically generated S - box matrix, the row index 0xA (decimal 10) corresponds to the 10th row, and the column index 0xB (decimal 11) corresponds to the 11th column. Looking up the table, the replaced byte value (such as 0xCD) is obtained. Through this operation, each byte of the replay - resistant encoding is non - linearly mapped to a new value, and the mapping rule changes dynamically with the S - box. For example, if the same replay - resistant encoding is generated at different times, due to different S - boxes, its replacement results will be significantly different, thus destroying the attacker's ability to predict the feature rules.

[0068] To enhance the irreversibility of features, multiple rounds of operations need to be performed on the replaced byte sequence. Each round includes S-box substitution and bit shifting. First-round substitution: Byte substitution is completed based on the dynamic S-box; Circular left shift: The 8 bits of each byte are circularly shifted left by 3 bits to disrupt the bit order; Second-round substitution: Replace again using a new dynamic S-box (derived from the result of the previous round); Final shift: A block-wise circular right shift operation is performed on the overall byte sequence. Multiple rounds of operations completely obscure the correlation between the feature and the original anti-replay encoding through cascaded non-linear transformations and bit perturbations. For example, after 3 rounds of substitution and shifting, even if the attacker obtains the final feature, they cannot reverse-derive and restore the original encoded content.

[0069] In the final feature generation stage, the current timestamp (such as Unix timestamp) is embedded into the target message feature according to a preset rule. Specifically, the timestamp is converted into a 4-byte binary format and XORed with the replaced byte sequence to ensure that the timestamp is scattered and embedded in the feature. At the same time, the synchronization of the timestamp is guaranteed by the hardware clock to avoid watermark conflicts caused by clock offsets. The feature with the embedded watermark not only has temporal uniqueness but can also be used to detect replay attacks. The finally generated target message feature, through the above process, combines protocol semantics, physical signal characteristics, and dynamic security attributes, providing a robust and forgery-proof feature input for high-concealment threat detection.

[0070] In some instances, it also includes: Based on the synchronization relationship between the timestamp of the target message feature and the physical layer signal waveform, a watermark identifier is embedded to generate a target message feature containing anti-replay encoding and the watermark identifier.

[0071] Exemplarily, during the process of generating the target message feature, first, through the hardware clock synchronization mechanism, the timestamp of the target message feature is strictly aligned with the sampling moment of the physical layer signal waveform to ensure that the two are completely consistent in timing. The timestamp is generated by a high-precision clock source (such as GPS synchronized clock or Network Time Protocol NTP), with an accuracy reaching the microsecond level, to avoid synchronization errors caused by clock offsets. The synchronization data of the physical layer signal waveform is obtained through the hardware timestamp marking function of the baseband chip or radio frequency sensor, ensuring that its corresponding relationship with the protocol layer timestamp is accurate to a single sampling point. Subsequently, the timestamp is converted into a binary sequence of a fixed length (such as 4-byte Unix timestamp), and a watermark identifier is generated through a preset hash function (such as SHA-256). This identifier contains the uniqueness information of the timestamp and the synchronization check code of the physical layer signal waveform. The embedding of the watermark identifier is achieved through exclusive OR operation: perform an exclusive OR operation bit by bit on the specified positions (such as 4 bytes reserved at the head, middle, and tail) of the anti-replay encoding with the watermark identifier to generate the target message feature containing the watermark. For example, if the anti-replay encoding is 804 bytes and the watermark identifier is 12 bytes (3 groups of 4-byte timestamps), then perform exclusive OR operations on the 0-3 bytes, 400-403 bytes, and 800-803 bytes of the encoding respectively to ensure that the watermark is dispersed and embedded without affecting the overall structure of the feature. The generated target message feature realizes the uniqueness verification of the transmission moment through the strong binding relationship between the timestamp and the physical layer signal. If an attacker attempts to replay a historical feature, the timestamp verification of the watermark identifier will fail due to being out of sync with the current physical layer signal waveform, thus effectively resisting replay attacks. This mechanism ensures that the watermark is tamper-proof and has dynamic uniqueness through hardware-level synchronization, hash digest generation, and decentralized embedding strategies, providing end-to-end security protection for the message feature.

[0072] Please refer to Figure 2 , which is a schematic structural diagram of a message feature extraction device provided by an embodiment of the present application, including: A message data acquisition unit 21, configured to acquire the original message data of the target network traffic and the corresponding physical layer signal waveform; A protocol field determination unit 22, configured to parse the original message data based on a preset protocol layering rule to determine a multi-layer protocol field set; A protocol feature generation unit 23, configured to determine a protocol layer feature vector based on the multi-layer protocol field set; A frequency domain feature generation unit 24, configured to determine a frequency domain feature vector based on the physical layer signal waveform; A weight matrix construction unit 25, configured to construct a cross-attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; A message feature extraction unit 26, configured to generate a target message feature based on the cross-attention weight matrix.

[0073] Please refer toFigure 3 , Embodiment of the present application further provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for extracting message features.

[0074] Since the electronic device introduced in this embodiment is the device adopted for implementing a message feature extraction device in an embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope protected by the present application.

[0075] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0076] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0077] Those skilled in the art should understand that the embodiments of the present application can provide a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program codes.

[0078] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more of the processes Figure 1 and / or boxes Figure 1 specified in one or more of the processes and / or boxes.

[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 and / or boxes Figure 1 specified in one or more of the boxes.

[0081] An embodiment of the present application also provides a computer program product, which includes computer software instructions that, when running on a processing device, cause the processing device to execute Figure 1 the process of a message feature extraction method in the corresponding embodiment.

[0082] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatuses. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.

[0083] Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working processes of the systems, apparatuses, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0084] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0085] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware and / or software functional units.

[0087] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device to execute all or part of the steps of the methods in various embodiments of the present application.

[0088] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.

[0089] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of this specification.

[0090] Obviously, those skilled in the art can make various modifications and deformations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, then this specification is also intended to include these modifications and deformations.

Claims

1. A message feature extraction method, characterized in that: include: Obtain the original message data of the target network traffic and the corresponding physical layer signal waveform; Parsing the original message data based on a preset protocol layering rule to determine a multi-layer protocol field set; Determining a protocol layer feature vector based on the multi-layer protocol field set; Based on the physical layer signal waveform, determining a frequency domain feature vector; Constructing a cross attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; Based on the cross-attention weight matrix, target message features are generated.

2. The method according to claim 1, characterized in that The parsing of the original message data based on the preset protocol layering rule to determine the multi-layer protocol field set includes: Determine a protocol identification segment based on a first preset length byte of the header of the original message data; Determine a protocol parsing template based on the matching result between the protocol identification segment and the protocol feature library; Performing secondary parsing on the unmatched protocol fields in the original message data based on a sliding window mechanism to determine the position marker of the variable-length field; Generate a protocol tree including a multi-layer protocol structure based on the protocol parsing template and the position mark of the variable-length field; Based on the hierarchical structure of the protocol tree, a multi-layer protocol field set is determined.

3. The method according to claim 1, characterized in that The determining of the protocol layer feature vector based on the multi-layer protocol field set includes: Determine a first semantic code based on a mapping relationship between a transport layer port number and an application layer protocol type in the multi-layer protocol field set; Determine a second semantic code based on a timing correlation coefficient between a network layer TTL value and a data link layer frame interval in the multi-layer protocol field set; A protocol layer feature vector is determined based on a normalized concatenation result of the first semantic code and the second semantic code.

4. The method according to claim 1, characterized in that: The determining of the frequency domain feature vector based on the physical layer signal waveform includes: Divide the physical layer signal waveform into frequency bands based on a wavelet packet decomposition algorithm to determine the energy entropy of each frequency band; Determine the signal stability coefficient based on the decay rate of the autocorrelation function of the signal amplitude between adjacent messages; Based on the weighted calculation result of the energy entropy of each frequency band and the signal stability coefficient, a frequency domain feature vector is determined.

5. The method according to claim 1, characterized in that The constructing a cross attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector includes: Determining the protocol timing transition probability based on the mapping relationship between the transport layer port number and the application layer protocol type in the protocol layer feature vector; Determining the frequency domain mutation intensity based on the distribution characteristics of the energy entropy of each frequency band in the frequency domain feature vector; Determining a cross-modal correlation matrix based on the protocol timing transition probability and the frequency domain mutation intensity; Performing temporal modeling on the cross-modal association matrix based on a gated recurrent unit to determine feature dependencies; Based on the self-attention mechanism, weights are assigned to the feature dependencies to determine a cross-attention weight matrix.

6. The method according to claim 1, characterized in that The generating target message features based on the cross attention weight matrix includes: Performing weighted fusion on the protocol layer feature vector and the frequency domain feature vector based on the cross attention weight matrix to determine a bimodal feature; Generate a pseudo hash sequence based on the byte sampling result of the encrypted payload in the original message data; Determining an anti-replay code based on an XOR confusion operation of the pseudo-hash sequence and the bimodal feature; The anti-replay code is nonlinearly transformed based on a dynamic S-box replacement algorithm to determine the target message characteristics.

7. The method according to claim 6, characterized in that Also includes: Based on the synchronization relationship between the timestamp of the target message feature and the physical layer signal waveform, a watermark identifier is embedded to generate the target message feature including the anti-replay code and the watermark identifier.

8. A message feature extraction device, characterized in that: include: A message data acquisition unit, used to acquire the original message data of the target network traffic and the corresponding physical layer signal waveform; A protocol field determination unit, which parses the original message data based on a preset protocol layering rule to determine a multi-layer protocol field set; A protocol feature generating unit, which determines a protocol layer feature vector based on the multi-layer protocol field set; A frequency domain feature generating unit, which determines a frequency domain feature vector based on the physical layer signal waveform; A weight matrix construction unit, which constructs a cross-attention weight matrix based on the protocol layer feature vector and the frequency domain feature vector; The message feature extraction unit generates target message features based on the cross-attention weight matrix.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the message feature extraction method as described in any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the message feature extraction method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Vehicle-mounted CAN bus intrusion detection system based on cross-protocol layer fusion

    CN117857145A

  • Abnormal network flow detection system and method based on multi-modal fusion features

    CN118353690A

  • Bilinear attention mechanism Internet of Things attack and defense method and system

    CN118523958A

  • Chain information collection and vulnerability checking method and related product

    CN118764326A

  • Protocol interface test method, device, computer device and storage medium

    WO2020119430A1

Cited By

  • Internet of Things protocol analysis method and device based on multi-mode AI and medium

    CN120512486A

  • IoT protocol analysis method, equipment and media based on multimodal AI

    CN120512486B

  • Internet of Things communication method and device based on Internet of Things protocol analysis, equipment and medium

    CN120658773A

  • Encrypted traffic data leakage traceability method and system based on spatio-temporal characteristics

    CN120915604A

  • Cross-protocol identification data intelligent analysis middleware method and system

    CN121125869A