Abnormal traffic content identification method and device based on large model

Through the large-model-based anomaly traffic content recognition method, combined with byte entropy, target information and multi-dimensional data analysis, the problem that a single-dimensional analysis method in the existing technology is difficult to identify complex network attack patterns, significantly improving the accuracy of anomaly traffic recognition.

CN120090875APending Publication Date: 2025-06-03BEIJING FULE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510554309.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing abnormal traffic detection scheme relies on a single-dimensional analysis framework, and it is difficult to fully capture the subtle differences between normal traffic and abnormal behavior, resulting in inaccurate identification results.

Method used

The abnormal traffic content recognition method based on the big model is adopted. By obtaining the byte entropy and target information in the traffic data packet, combining preset information tables and multi-dimensional data analysis, we accurately distinguish encrypted and non-encrypted data, and using the big model to identify the traffic data packets.

Benefits of technology

It significantly improves the accuracy of abnormal traffic content recognition, avoids the limitations of single-dimensional detection, and can more accurately identify complex and changeable network attack patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090875A_ABST
    Figure CN120090875A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal traffic content identification method and device based on a large model, and relates to the technical field of large models. The method comprises the following steps: acquiring a first traffic data packet; obtaining a plurality of bytes from the first traffic data packet, and performing information entropy calculation on the plurality of bytes to obtain a byte entropy; when the byte entropy is greater than the preset byte entropy, obtaining target information from the first traffic data packet; when the target information exists in the preset information table, determining that the first traffic data packet is an encrypted data packet; calling a first analysis mode according to the encrypted data packet, and processing the first traffic data packet based on the first analysis mode to obtain a plurality of first data; and inputting the plurality of first data into a preset large model for processing to obtain an abnormal result, and determining that the first traffic data packet is abnormal traffic content according to the abnormal result. By implementing the technical scheme provided by the invention, the limitation of a single-dimension analysis mode in facing a complex and changeable network attack mode is effectively solved, and the recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of large models, and specifically relates to a method and device for identifying abnormal traffic content based on large models. Background Art

[0002] Under the macro background of the accelerating global informatization process, network security has leaped to become a core strategic element for maintaining social stability and promoting the safe development of the economy. With the in-depth application and cross-border integration of frontier technologies such as cloud computing and the Internet of Things, the data volume in the cyber space shows an unprecedented exponential expansion trend. This structural transformation not only significantly enhances the diversity, dynamics and interactive complexity of network behaviors, but also gives rise to the evolution of network attack means towards highly concealed, diverse in form and intelligent in strategy. To effectively address this increasingly severe network security challenge, abnormal traffic detection technology, as the key technical support for building an active defense system, is undergoing continuous technological iteration and innovative breakthroughs. Its core goal is to achieve accurate identification and forward-looking early warning of potential security threats from massive, high-dimensional and dynamically evolving network data streams.

[0003] However, current mainstream abnormal traffic detection solutions still generally rely on a single-dimensional analysis framework, such as quantitative analysis based on traffic statistical features and feature matching of fixed patterns. When facing complex and changeable network attack patterns, a single detection method often fails to comprehensively capture the subtle differences between normal traffic and abnormal behaviors, resulting in inaccurate identification results.

[0004] Therefore, there is an urgent need for a method and device for identifying abnormal traffic content based on large models that can solve the above technical problems. Summary of the Invention

[0005] The present application provides a method and device for identifying abnormal traffic content based on large models. This method accurately distinguishes encrypted and unencrypted data by comprehensively capturing network behaviors and combining byte entropy and target information, and then uses a large model to identify traffic data packets, effectively solving the limitations of the single-dimensional analysis method when facing complex and changeable network attack patterns and improving the accuracy of identification.

[0006] In a first aspect, the present application provides a method for identifying abnormal traffic content based on a large model. The method includes: obtaining a first traffic data packet; obtaining multiple bytes from the first traffic data packet, calculating the information entropy of the multiple bytes to obtain the byte entropy; determining whether the byte entropy is less than or equal to a preset byte entropy; when the byte entropy is greater than the preset byte entropy, obtaining target information from the first traffic data packet, where the target information includes duplicate information, padding character information, and file naming information; determining whether the target information exists in a preset information table, where the preset information table is information corresponding to encrypted data; when the target information exists in the preset information table, determining that the first traffic data packet is an encrypted data packet; invoking a first analysis method based on the encrypted data packet, and processing the first traffic data packet based on the first analysis method to obtain multiple first data, where the multiple first data includes communication frequency, connection time, destination IP address, source IP address, and device port number; inputting the multiple first data into a preset large model for processing to obtain an abnormal result, and determining that the first traffic data packet is abnormal traffic content based on the abnormal result.

[0007] By adopting the above technical solution, the byte entropy in the first traffic data packet is calculated to identify highly random encrypted traffic as a preliminary screening index. Then, the calculated byte entropy is compared with the preset byte entropy. When the byte entropy is greater than the preset byte entropy, the target information in the first traffic data packet is extracted, and then the extracted target information is matched with the preset information table to accurately identify the encrypted data packet. Combining multiple encryption features to determine that the first traffic data packet is an encrypted data packet reduces misjudgment. Then, based on the encrypted data packet, the first analysis method is determined, and the first traffic data packet is processed using the first analysis method to obtain first data in multiple dimensions. Based on multi-dimensional data analysis, the traffic behavior pattern can be captured. The multi-dimensional first data is input into a preset large model for processing to obtain an abnormal result, and based on the abnormal result, it is determined that the first traffic data packet is abnormal traffic content. Through comprehensive judgment using multiple indicators and selecting different extraction methods according to whether the traffic data packet is encrypted, the limitation of single-dimensional detection is avoided, and the accuracy of identifying abnormal traffic content is significantly improved.

[0008] Optionally, after determining whether the byte entropy is less than or equal to the preset byte entropy, the method further includes: when the byte entropy is less than or equal to the preset byte entropy, obtaining a second traffic data packet, where the second traffic data packet is the data packet after the first traffic data packet; obtaining a first length from the first traffic data packet and a second length from the second traffic data packet; calculating the first length and the second length to obtain the data packet length entropy; determining whether the data packet length entropy is less than or equal to a preset length entropy; when the data packet length entropy is greater than the preset length entropy, obtaining target information from the first traffic data packet.

[0009] By adopting the above technical solution, when the byte entropy is less than or equal to the preset byte entropy, by calculating the packet length entropy of the first traffic packet and the second traffic packet, the abnormal change of the packet length is identified. When the packet length entropy is greater than the preset length entropy, it indicates that the packet length distribution is abnormal, and there may be malicious traffic or attack behavior. The attacker may randomize the packet length to confuse the traffic. The packet length entropy can capture this abnormality. By combining the byte entropy and the packet length entropy, legal encrypted traffic and malicious encrypted traffic can be more accurately distinguished, thereby significantly improving the accuracy of abnormal traffic detection.

[0010] Optionally, after determining whether the packet length entropy is less than or equal to the preset length entropy, the method further includes: when the packet length entropy is less than or equal to the preset length entropy, obtaining a third traffic packet, where the third traffic packet is a packet received after the second traffic packet; obtaining a first time interval, where the first time interval is the time interval between the first arrival time and the second arrival time, the first arrival time is the time when the first traffic packet is obtained, and the second arrival time is the time when the second traffic packet is obtained; obtaining a second time interval, where the second time interval is the time interval between the second arrival time and the third arrival time, and the third arrival time is the time when the third traffic packet is obtained; calculating the first time interval and the second time interval to obtain a time interval entropy; determining whether the time interval entropy is less than or equal to the preset time entropy; when the time interval entropy is greater than the preset time entropy, obtaining target information from the first traffic packet.

[0011] By adopting the above technical solution, when the packet length entropy is less than or equal to the preset length entropy, by calculating the time interval entropy values between the first traffic packet and the second traffic packet, and between the second traffic packet and the third traffic packet, the abnormal change of the packet arrival time can be identified. When the time interval entropy is greater than the preset time entropy, it indicates that the packet arrival time distribution is abnormal, and there may be malicious traffic or attack behavior. Pulse attacks will cause abnormal packet arrival time intervals, and the time interval entropy calculation can effectively identify such attacks, thereby providing data support for subsequent distinguishing normal and abnormal traffic.

[0012] Optionally, after determining whether the target information exists in the preset information table, the method further includes: when the target information does not exist in the preset information table, determining that the first traffic packet is an unencrypted packet; retrieving a second analysis method according to the unencrypted packet, and parsing the first traffic packet based on the second analysis method to obtain second data, where the second data includes text information, image information, and link information; determining whether there is abnormal data in the second data; when there is abnormal data in the second data, determining that the first traffic packet is abnormal traffic content, and performing an abnormal mark on the first traffic packet according to the abnormal traffic content.

[0013] By adopting the above technical solution, when the target information does not exist in the preset information table, the data packet is determined to be an unencrypted data packet, and the second analysis method is invoked for parsing. By parsing the content of the unencrypted data packet, potential abnormal data such as malicious text, malicious images, or malicious links are identified, which complements the detection of encrypted data packets and comprehensively covers the abnormal detection in network traffic. Since attackers may hide malicious content through obfuscation techniques, in-depth parsing can reveal this hidden information. By combining text, image, and link information, multi-dimensional comprehensive judgment is performed to improve the accuracy of abnormal detection. Different analysis strategies are adopted for encrypted and unencrypted traffic to improve the flexibility of detection.

[0014] Optionally, after determining that the first traffic data packet is an encrypted data packet when the target information exists in the preset information table, the method further includes: analyzing the first traffic data packet to obtain an encryption protocol; determining whether the encryption protocol exists in the preset encryption protocol table; when the encryption protocol does not exist in the preset encryption protocol, determining to use the first analysis method to process the first traffic data packet.

[0015] By adopting the above technical solution, when the encryption protocol used by the data packet is not in the preset encryption protocol table, it can be identified that this is an unknown or unauthorized encryption protocol. When the identified encryption protocol does not exist in the preset encryption protocol table, the first analysis method is used to process the data packet. Through the first analysis method, more detailed inspection of the unauthorized encrypted traffic is performed to improve the accuracy of abnormal detection. Combining with the encryption protocol whitelist mechanism, false alarms caused by misjudging normal encrypted traffic as abnormal are reduced.

[0016] Optionally, after determining whether the encryption protocol exists in the preset encryption protocol table, the method further includes: when the encryption protocol exists in the preset encryption protocol table, decrypting the first traffic data packet according to the encryption protocol to obtain the decrypted traffic data packet; parsing the decrypted traffic data packet to obtain third data, where the third data includes text data and image data; determining whether there are abnormal keywords in the text data and whether the target similarity is less than or equal to the preset similarity, where the target similarity is the similarity between the image corresponding to the image data and the suspicious image; when there are no abnormal keywords in the text data and the target similarity is less than or equal to the preset similarity, determining that the first traffic data packet is normal traffic content.

[0017] By adopting the above technical solution, when there is an identified encryption protocol in the preset encryption protocol table, the first traffic data packet is decrypted according to the protocol, and the encrypted traffic data packet is converted into plaintext, which is convenient for subsequent content analysis. Through decryption, the actual content of the first traffic data packet can be obtained for more in-depth security detection. Ensure that all authorized encrypted traffic can be decrypted and analyzed, prevent malicious content in the encrypted traffic from being ignored, identify potential malicious text content by checking whether there are abnormal keywords in the text data, identify potential malicious image content by calculating the similarity between the image data and the suspicious image, and combine the detection results in the two dimensions of text and image for comprehensive judgment to improve the accuracy of anomaly detection.

[0018] Optionally, after inputting multiple first data into a preset large model for processing to obtain an anomaly result and determining that the first traffic data packet is abnormal traffic content according to the anomaly result, the method further includes: troubleshooting the network device according to the first traffic data packet to determine the target device, where the target device is the device in the network device that has received the first traffic data packet; disconnecting the network connection between the target device and the sending device so that the target device stops running the relevant program, and the sending device is the device that sends the first traffic data packet.

[0019] By adopting the above technical solution, after detecting abnormal traffic, quickly locate the affected device, reduce the time window for attack spread, only isolate the affected device, avoid misoperation of normal devices, and reduce the impact on the overall network operation. Cut off the lateral movement path of the attacker in the network to prevent the attack from spreading to other devices. After disconnecting the network connection, the target device cannot communicate with the attacker, and the malicious program cannot continue to execute or receive instructions, preventing the malicious program from continuing to steal or leak sensitive data on the target device.

[0020] In a second aspect of the present application, an abnormal traffic content recognition device based on a large model is provided. The device includes an acquisition unit, a processing unit, and a determination unit; the acquisition unit acquires a first traffic data packet; obtains multiple bytes from the first traffic data packet, calculates the information entropy of the multiple bytes to obtain the byte entropy; the processing unit determines whether the byte entropy is less than or equal to a preset byte entropy; when the byte entropy is greater than the preset byte entropy, obtains target information from the first traffic data packet, and the target information includes duplicate information, padding character information, and file naming information; determines whether the target information exists in a preset information table, and the preset information table is information composed of corresponding encrypted data; when the target information exists in the preset information table, determines that the first traffic data packet is an encrypted data packet; invokes a first analysis method based on the encrypted data packet, and processes the first traffic data packet based on the first analysis method to obtain multiple first data, and the multiple first data includes communication frequency, connection time, destination IP address, source IP address, and device port number; the determination unit inputs the multiple first data into a preset large model for processing to obtain an abnormal result, and determines that the first traffic data packet is abnormal traffic content according to the abnormal result.

[0021] Optionally, the acquisition unit is configured to, when the byte entropy is less than or equal to the preset byte entropy, acquire a second traffic data packet, and the second traffic data packet is a data packet after the first traffic data packet; obtain a first length from the first traffic data packet, and obtain a second length from the second traffic data packet; the processing unit is configured to calculate the first length and the second length to obtain a data packet length entropy; determine whether the data packet length entropy is less than or equal to a preset length entropy; when the data packet length entropy is greater than the preset length entropy, obtain target information from the first traffic data packet.

[0022] Optionally, the processing unit is configured to, when the data packet length entropy is less than or equal to the preset length entropy, acquire a third traffic data packet, and the third traffic data packet is a data packet received after the second traffic data packet; the acquisition unit is configured to obtain a first time interval, and the first time interval is the time interval between a first arrival time and a second arrival time, the first arrival time is the time when the first traffic data packet is acquired, and the second arrival time is the time when the second traffic data packet is acquired; obtain a second time interval, and the second time interval is the time interval between the second arrival time and a third arrival time, and the third arrival time is the time when the third traffic data packet is acquired; the processing unit is configured to calculate the first time interval and the second time interval to obtain a time interval entropy; determine whether the time interval entropy is less than or equal to a preset time entropy; when the time interval entropy is greater than the preset time entropy, obtain target information from the first traffic data packet.

[0023] Optionally, the processing unit is configured to determine that the first traffic data packet is an unencrypted data packet when the target information does not exist in the preset information table; retrieve a second analysis method according to the unencrypted data packet, and parse the first traffic data packet based on the second analysis method to obtain second data, where the second data includes text information, image information, and link information; determine whether there is abnormal data in the second data; the determination unit is configured to determine that the first traffic data packet is abnormal traffic content when there is abnormal data in the second data, and perform an abnormal mark on the first traffic data packet according to the abnormal traffic content.

[0024] Optionally, the processing unit is configured to analyze the first traffic data packet to obtain an encryption protocol; determine whether the encryption protocol exists in the preset encryption protocol table; when the encryption protocol does not exist in the preset encryption protocol, determine to use the first analysis method to process the first traffic data packet.

[0025] Optionally, the processing unit is configured to decrypt the first traffic data packet according to the encryption protocol to obtain a decrypted traffic data packet when the encryption protocol exists in the preset encryption protocol table; parse the decrypted traffic data packet to obtain third data, where the third data includes text data and image data; determine whether there is an abnormal keyword in the text data and whether the target similarity is less than or equal to a preset similarity, where the target similarity is the similarity between the image corresponding to the image data and the suspicious image; the determination unit is configured to determine that the first traffic data packet is normal traffic content when there is no abnormal keyword in the text data and the target similarity is less than or equal to the preset similarity.

[0026] Optionally, the processing unit is configured to check the network device according to the first traffic data packet to determine a target device, where the target device is a device in the network device that has received the first traffic data packet; cut off the network connection between the target device and the sending device so that the target device stops running relevant programs, and the sending device is the device that sends the first traffic data packet.

[0027] In a third aspect of the present application, an electronic device is provided, where the electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory, so that an electronic device executes the method according to any one of the above in the present application.

[0028] In a fourth aspect of the present application, a computer-readable storage medium is provided, where the computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of the above in the present application is executed.

[0029] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Calculate the byte entropy in the first traffic data packet to identify encrypted traffic with high randomness as a preliminary screening indicator. Then compare the calculated byte entropy with the preset byte entropy. When the byte entropy is greater than the preset byte entropy, extract the target information in the first traffic data packet, and then match the extracted target information with the preset information table to accurately identify the encrypted data packet. Combine multiple encryption features to determine that the first traffic data packet is an encrypted data packet, reducing misjudgment. Then determine the first analysis method based on the encrypted data packet, and use the first analysis method to process the first traffic data packet to obtain first data in multiple dimensions. Based on multi-dimensional data analysis, the traffic behavior pattern can be captured. Input the multi-dimensional first data into the preset large model for processing to obtain an abnormal result. Determine that the first traffic data packet is abnormal traffic content according to the abnormal result. Make a comprehensive judgment through multiple indicators, and select different extraction methods according to whether the traffic data packet is encrypted to avoid the limitations of single-dimensional detection and significantly improve the accuracy of identifying abnormal traffic content.

[0030] 2. When the byte entropy is less than or equal to the preset byte entropy, calculate the packet length entropy of the first traffic data packet and the second traffic data packet to identify abnormal changes in the packet length. When the packet length entropy is greater than the preset length entropy, it indicates that the packet length distribution is abnormal, and there may be malicious traffic or attack behavior. The attacker may randomize the packet length to confuse the traffic, and the packet length entropy can capture this abnormality. Combining the byte entropy and the packet length entropy can more accurately distinguish between legitimate encrypted traffic and malicious encrypted traffic, thereby significantly improving the accuracy of abnormal traffic detection. Description of the Drawings

[0031] Figure 1 is a schematic flowchart of a method for identifying abnormal traffic content based on a large model provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a device for identifying abnormal traffic content based on a large model provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device disclosed by an embodiment of the present application.

[0032] Description of the reference numerals: 201, acquisition unit; 202, processing unit; 203, determination unit; 300, electronic device; 301, processor; 302, memory; 303, user interface; 304, network interface; 305, communication bus. Detailed Embodiment

[0033] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments.

[0034] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.

[0035] In the description of the embodiments of this application, the meaning of the term "a plurality of" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0036] Under the macro background of the accelerating global informatization process, network security has leaped to become a core strategic element for maintaining social stability and promoting economic security development. With the in-depth application and cross-border integration of frontier technologies such as cloud computing and the Internet of Things, the data volume in the cyber space shows an unprecedented exponential expansion trend. This structural transformation not only significantly enhances the diversity, dynamics and interaction complexity of network behaviors, but also gives rise to the evolution of network attack means towards the directions of high concealment, diverse forms and intelligent strategies. To effectively address this increasingly severe network security challenge, as a key technical support for building an active defense system, the abnormal traffic detection technology is undergoing continuous technical iteration and innovative breakthroughs. Its core goal is to accurately identify potential security threats and provide forward-looking warnings from the massive, high-dimensional and dynamically evolving network data streams.

[0037] However, the current mainstream abnormal traffic detection solutions still generally rely on a single-dimensional analysis framework, such as quantitative analysis based on traffic statistical features and feature matching of fixed patterns. When facing complex and changeable network attack patterns, a single detection method often fails to comprehensively capture the subtle differences between normal traffic and abnormal behaviors, resulting in inaccurate identification results.

[0038] Therefore, how to solve the limitation problem of the single-dimensional analysis method when facing complex and changeable network attack patterns. An abnormal traffic content recognition method based on a large model provided by an embodiment of the present application is applied to a server. The server of the present application can be a platform that provides abnormal traffic recognition services for the network security of an enterprise. Figure 1 It is a schematic flowchart of an abnormal traffic content recognition method based on a large model provided by an embodiment of the present application. Refer to Figure 1 , and this method includes the following steps S101 - step S108.

[0039] S101: Obtain a first traffic data packet.

[0040] In the above S101, due to the development of network technology, when employees of an enterprise use electronic devices for work, they need to focus on the network interfaces of each device, that is, analyze the traffic data packets received or transmitted in real time on the network interfaces to ensure that the received or transmitted traffic data packets are normal traffic, rather than content carrying malicious network attacks, resulting in the situation where the device is infected or sensitive data is leaked. Accurately identifying malicious network attacks is a key technology to ensure the security of the internal data network of an enterprise. However, the current method for identifying abnormal traffic relies too much on feature matching of fixed patterns. However, the current network attack patterns are constantly changing, and the feature matching of fixed patterns cannot comprehensively identify between normal traffic and abnormal behaviors, resulting in inaccurate identification results. The present application provides an abnormal traffic content recognition method based on a large model to solve the problem of inaccurate current identification results. Next, it will be described in detail how the present application identifies traffic data packets.

[0041] First, the present application takes the identification of traffic data packets received or sent by an enterprise as an example for description. It is possible to first check the device information that can receive or send data in the enterprise, and then determine the corresponding network interface for each device. At this time, the network interface refers to a network card. Capture traffic data packets in real time from the network interface of any device, or read them from an offline traffic log (such as a PCAP file). It is also possible to use a commercial traffic analysis tool to obtain traffic data packets. At this time, the obtained traffic data packets refer to the first traffic data packets obtained for the first time.

[0042] S102: Obtain multiple bytes from the first traffic data packet, calculate the information entropy of the multiple bytes, and obtain the byte entropy.

[0043] In the above S102, after obtaining the first traffic data packet, each byte in the first traffic data packet is traversed to count the number of occurrences of each byte value (0 - 255). The total number of bytes in the first traffic data packet is determined, and then the probability of each byte value is calculated in turn. At this time, the probability is the number of occurrences of the byte value divided by the total number of bytes in the first traffic data packet. Then, the probability of each byte value is substituted into the information entropy formula to calculate the byte entropy. The specific calculation formula is as follows: ;H byte represents the byte entropy value, x i represents the byte value corresponding to the i-th byte, P(x i ) represents the probability of occurrence of the byte value corresponding to the i-th byte. The byte entropy represents the degree of chaos or uncertainty of the data stream. For example, the first traffic data packet contains the following byte sequence [01, 02, 01, 03, 02, 01]. First, calculate the frequency of each byte occurrence. At this time, 01 appears 3 times, 02 appears 2 times, and 03 appears 1 time. Then calculate the probability of each byte appearing in the first traffic data packet. P(01) = 3 / 6 = 0.5, P(02) = 2 / 6 = 0.3, P(03) = 1 / 6 = 0.167. Then, input the obtained probabilities into the byte entropy for calculation, H byte is equal to 1.459 bits.

[0044] S103: Determine whether the byte entropy is less than or equal to the preset byte entropy.

[0045] In the above S103, after calculating the byte entropy for the first traffic data packet and obtaining the byte entropy, compare the byte entropy with the preset byte entropy. The preset byte entropy is the basis for determining whether the first traffic data packet is encrypted data or non-encrypted data, and the preset byte entropy can be set according to experience. When the byte entropy is high, the first traffic data packet may be encrypted data. When the byte entropy is low, it may be non-encrypted data.

[0046] S104: When the byte entropy is greater than the preset byte entropy, obtain target information from the first traffic data packet. The target information includes duplicate information, padding character information, and file naming information.

[0047] In the above S104, when the byte entropy is greater than the preset byte entropy, it is determined that the first traffic data packet may be encrypted data. Although encrypting the data disrupts the pattern of the original data, resulting in a high entropy value, compressing the data can also remove redundancy in the data and may lead to a high entropy value. Therefore, directly defaulting the first traffic data packet with a byte entropy greater than the preset byte entropy as an encrypted data packet may result in misjudgment and ignoring compressed files. Thus, it is necessary to search for target information in the first traffic data packet. At this time, the target information refers to repeated byte sequences (such as padding characters, key fragments), padding character information, and file naming information. The padding character information is to search for padding characters, and the file naming information is to search for the naming method of the file name.

[0048] In addition, when the byte entropy is less than or equal to the preset byte entropy, obtain a second traffic data packet, where the second traffic data packet is the data packet after the first traffic data packet; obtain a first length from the first traffic data packet and a second length from the second traffic data packet; calculate the data packet length entropy for the first length and the second length; determine whether the data packet length entropy is less than or equal to the preset length entropy; when the data packet length entropy is greater than the preset length entropy, obtain the target information from the first traffic data packet. Specifically, the byte entropy of the first traffic data packet has been calculated above and it is determined that the byte entropy is less than or equal to the preset byte entropy. To more accurately determine whether the first traffic data packet is encrypted data, in addition to the byte entropy, the data packet length entropy and the time interval entropy can also be calculated, and then determine whether the first traffic data packet is encrypted data. Then continue to capture the next data packet from the network interface, or two consecutive data packets can also be read from the relevant file. Then obtain the first length from the first traffic data packet and the second length from the second traffic data packet. The data packet length usually refers to the number of bytes of the entire data packet (including the header and the payload). Then calculate the first length and the second length. Multiple traffic data packets can be obtained from the network interface first, and then count the number of occurrences of different data packet lengths. Then calculate the probability of each data packet length. At this time, the probability refers to the number of occurrences of this data packet length divided by the total number of data packets. Then substitute the probability of each data packet length into the information entropy formula to calculate the data packet length entropy. The specific calculation formula is as follows: ; where m is the number of different data packet lengths, H length represents the data packet length entropy value, l i represents the length of the i-th data packet, P(l i)(i) represents the probability of the occurrence of the length of the i-th data packet. The data packet length entropy is used to measure the randomness of the data packet length. After obtaining the data packet length entropy, the data packet length entropy is compared with a preset length entropy, which is the basis for determining whether the first traffic data packet is encrypted data or non-encrypted data, and the preset length entropy can be set according to experience. When the data packet length entropy is greater than the preset length entropy, it is determined that the first traffic data packet is not non-encrypted data, and it is also necessary to further obtain target information from the first traffic data packet, extract duplicate information, padding characters, and file naming information, so as to judge whether the first traffic data packet is compressed data or encrypted data.

[0049] Further, when the packet length entropy is less than or equal to a preset length entropy, obtain a third traffic packet, where the third traffic packet is a packet received after the second traffic packet; obtain a first time interval, where the first time interval is the time interval between the first arrival time and the second arrival time, the first arrival time is the time when the first traffic packet is obtained, and the second arrival time is the time when the second traffic packet is obtained; obtain a second time interval, where the second time interval is the time interval between the second arrival time and the third arrival time, and the third arrival time is the time when the third traffic packet is obtained; calculate the first time interval and the second time interval to obtain a time interval entropy; determine whether the time interval entropy is less than or equal to a preset time entropy; when the time interval entropy is greater than the preset time entropy, obtain target information from the first traffic packet. Specifically, after calculating and determining the packet length entropy, it is confirmed that the packet length entropy is less than or equal to the preset length entropy. After capturing the second traffic packet, continue to capture the next packet (i.e., the third traffic packet). The capture method can be real-time traffic monitoring based on a network interface or sequential reading from an existing traffic dataset. If the traffic data is stored in a PCAP file, read the third traffic packet sequentially. When obtaining traffic packets, it is necessary to ensure that the third traffic packet is received after the second traffic packet to maintain the timing of the packets. When capturing each packet, record its timestamp. First, obtain the timestamp corresponding to the first traffic packet, i.e., the first arrival time, then obtain the timestamp corresponding to the second traffic packet, i.e., the second arrival time, and then calculate the first arrival time and the second arrival time to obtain a first time interval. At this time, the first time interval is the time difference between the first arrival time and the second arrival time. Then obtain the timestamp corresponding to the third traffic packet, i.e., the third arrival time, and then calculate the second arrival time and the third arrival time to obtain a second time interval. At this time, the second time interval is the time difference between the second arrival time and the third arrival time. That is, after obtaining multiple traffic packets, it is necessary to sequentially obtain the time interval differences between two traffic packets. Then calculate the first time interval and the second time interval to obtain a time interval entropy, and the time interval entropy is used to measure the randomness of the time interval. Collect the time intervals of multiple consecutive packets (such as the first time interval, the second time interval, etc.). Count the occurrence frequency of each time interval and calculate its probability distribution. Probability refers to the number of times the time interval appears divided by the total number of time intervals, and then substitute the probability of each time interval into the information entropy formula to calculate the time interval entropy. The specific calculation formula is as follows: ; where k is the number of different time intervals, Hinterval represents the packet length entropy value, ti represents the i-th time interval, and P(ti) represents the probability of the occurrence times of the i-th time interval. Then, compare the calculated time interval entropy with the preset time entropy. The preset time entropy is used as the basis for determining whether the first traffic packet is encrypted data or non-encrypted data, and the preset time entropy can be set according to experience. If the time interval entropy is less than or equal to the preset time entropy, it is considered that the randomness of the time interval is within the normal range, that is, it is confirmed that the first traffic packet is non-encrypted data, and then use the analysis method of non-encrypted data to identify abnormal traffic for the first traffic packet. If the time interval entropy is greater than the preset time entropy, it is considered that the randomness of the time interval is abnormal, and it is possible that the first traffic packet is encrypted or compressed data. To further determine whether the first traffic packet is encrypted data or compressed data, information extraction can be performed on the first traffic packet, the second traffic packet, and the third traffic packet in sequence to check whether there are repeated byte sequences in the traffic packets. For example, consecutive identical bytes may indicate a padding or encryption pattern. Check whether there are common padding characters in the payload. Padding characters may be used to align the packet length or hide the real content. Try to decode the payload content to find possible file names or file extensions. For example, look for strings such as ".exe", ".pdf", etc. Perform byte-by-byte analysis on the payload of the first traffic packet.

[0050] S105: Determine whether there is target information in the preset information table, and the preset information table is the information composed of corresponding encrypted data.

[0051] In the above S105, after obtaining the target information from the traffic packet, a feature table common in encrypted data, that is, the preset information table, is constructed in advance. Then, input the target information into the preset information table for query to determine whether there is target information in the preset information table. The preset information table can be gradually updated based on the update of encrypted data to make the encrypted features stored in the preset information table match the actual application.

[0052] S106: When there is target information in the preset information table, determine that the first traffic packet is an encrypted packet.

[0053] In the above S106, when the byte entropy of the first traffic packet is high and the target information exists in the preset information table, it can be determined that the first traffic packet is an encrypted packet. Then, different analysis methods are selected for different traffic packets to solve the limitation of relying on a fixed recognition method currently and improve the accuracy of traffic packet recognition.

[0054] In addition, after determining that the first traffic data packet is an encrypted data packet, it is also necessary to identify the encryption method of the encrypted data packet, effectively identify the unknown encryption method, and adopt corresponding processing measures to analyze the first traffic data packet to obtain the encryption protocol; determine whether there is an encryption protocol in the preset encryption protocol table; when there is no encryption protocol in the preset encryption protocol, determine to use the first analysis method to process the first traffic data packet. Specifically, use a network protocol parsing tool (such as a custom protocol parser) to parse the first traffic data packet layer by layer, and first extract the protocol information of each layer of the data packet, including the transport layer protocol (such as TCP, UDP), and the application layer protocol (such as HTTP, TLS, SSH, FTP, etc.). Then check the source port and destination port of the data packet, and the port numbers of common encryption protocols TLS / SSL: 443 (HTTPS), 993 (IMAPS), 995 (POP3S). Extract the payload content of the data packet and check whether there is a characteristic field or pattern of a specific protocol in the payload. If the destination port of the first traffic data packet is 443, it is preliminarily determined to be a TLS protocol. Use a regular expression or pattern matching algorithm to find the protocol identifier in the payload. For complex protocols (such as TLS), use the state machine model to parse the handshake process and confirm the protocol type. First, analyze the first traffic data packet to determine the encryption protocol used by the first traffic data packet (such as TLS, SSH, custom encryption protocol, etc.). Predefine an encryption protocol table containing known encryption protocols and their characteristics, that is, a preset encryption protocol table. At this time, the preset encryption protocol table can also be set to the encryption protocol or method used within the enterprise. Then accurately match the encryption protocol name obtained by analysis with the protocol name in the preset encryption protocol table. If the match is successful, it is considered that the encryption protocol exists in the preset encryption protocol table. If the match fails, it is considered that the encryption protocol does not exist in the preset encryption protocol table. Confirm that the method for encrypting the first traffic data packet is an unknown protocol. Using an unknown protocol to encrypt the traffic data packet may be malicious traffic, so it is necessary to call the first analysis method according to the encrypted data packet to process the first traffic data packet.

[0055] The step of decrypting the encrypted data and then analyzing the decrypted data in this application is not necessary. This application can also directly process the encrypted data to determine whether the encrypted data is an abnormal traffic packet. By directly adopting the first analysis method to process the first traffic data packet, the first traffic data packet can be extracted through communication mode and correlation analysis, and information such as communication frequency, connection time, destination IP address, source IP address, and device port number can be extracted from the first traffic data packet. Use a network protocol parsing library (such as Scapy) to extract the data packet header information. First, observe the characteristics of the first traffic data packet, such as communication frequency, data packet size distribution, connection time, etc. By identifying these characteristics, abnormal behaviors can be identified. Then, combined with information such as source IP address, destination IP address, port number, etc., analyze the context of the data packet. Identify the communication with known malicious IPs or domain names to obtain multiple first data.

[0056] Further, when there is an encryption protocol in the preset encryption protocol table, decrypt the first traffic data packet according to the encryption protocol to obtain the decrypted traffic data packet; parse the decrypted traffic data packet to obtain the third data, where the third data includes text data and image data; determine whether there are abnormal keywords in the text data and whether the target similarity is less than or equal to the preset similarity, and the target similarity is the similarity between the image corresponding to the image data and the suspicious image; when there are no abnormal keywords in the text data and the target similarity is less than or equal to the preset similarity, determine that the first traffic data packet is normal traffic content. Specifically, find a matching encryption protocol in the preset encryption protocol table. Use the decryption algorithm and key corresponding to the matching encryption protocol to decrypt the first traffic data packet. Find an entry in the preset encryption protocol table that matches the identified encryption protocol. The matching item should include the protocol name, version, and information required for decryption (such as encryption algorithm, key, etc.). Select the corresponding decryption algorithm (such as AES, RSA, etc.) according to the matching encryption protocol. Ensure that the key required for decryption is securely stored and correctly used during the decryption process. Use an encryption library (such as OpenSSL) or a custom decryption module for decryption. The decrypted data packet should be restored to plaintext form for subsequent analysis. Successfully decrypt the first traffic data packet to obtain the decrypted traffic data packet, providing a basis for subsequent analysis. Parse the protocol and extract the content of the decrypted traffic data packet. Classify the extracted data into text data and image data. According to the encryption protocol type of the first traffic data packet, use the corresponding parser to extract data, and extract text content, image data, etc. from the first traffic data packet payload. Extract the text information in the data packet, such as HTML text, email body, etc. Identify and extract the image file or image data in the data packet. Use an image processing library (such as PIL, OpenCV) to process the image data. Successfully extract the text data and image data from the decrypted traffic data packet, providing a data basis for subsequent anomaly detection. First, maintain a list containing abnormal keywords, such as malicious code features, SQL injection keywords, etc. Then, check whether there are abnormal keywords in the text data. Extract features from the image data (such as SIFT, SURF, HOG features). Calculate the similarity between images using cosine similarity, Euclidean distance, or perceptual hashing algorithm. Maintain a library containing suspicious images for comparison with the extracted images. Compare the calculated target similarity with the preset similarity to determine whether the images are similar. Accurately determine whether there are abnormal keywords in the text data and the similarity between the image data and the suspicious image, providing a basis for subsequent traffic content judgment. At this time, the suspicious image refers to an image containing malicious content or features. Make a comprehensive judgment based on the abnormal keyword detection result and the image similarity comparison result.There are no abnormal keywords in the text data, and the target similarity is less than or equal to the preset similarity. If it is determined that the normal traffic condition is met, the first traffic data packet is marked as normal traffic content. Accurately identify and mark normal traffic content, reduce false alarms, and improve the accuracy of traffic analysis.

[0057] S107: Retrieve the first analysis method according to the encrypted data packet, and process the first traffic data packet based on the first analysis method to obtain multiple first data.

[0058] In the above S107, after determining that the first traffic data packet is an encrypted data packet and the first traffic data packet is encrypted by an unknown encryption protocol, according to the characteristics of the encrypted data packet, select the first analysis method to extract the first traffic data packet. Since the plaintext content cannot be directly obtained from the first traffic data packet, but the first traffic data packet can be extracted through the communication mode and correlation analysis. Information such as communication frequency, connection time, destination IP address, source IP address, and device port number can be extracted from the first traffic data packet. The network protocol parsing library (such as Scapy) can be used to extract the data packet header information. First, observe the characteristics of the first traffic data packet such as communication frequency, data packet size distribution, and connection time. By identifying these characteristics, abnormal behaviors can be identified. Then, combined with information such as source IP address, destination IP address, and port number, analyze the context of the data packet. Identify the communication with known malicious IPs or domain names to obtain multiple first data for subsequent anomaly detection.

[0059] S108: Input the multiple first data into a preset large model for processing to obtain an anomaly result, and determine that the first traffic data packet is abnormal traffic content according to the anomaly result.

[0060] In the above S108, after processing the first traffic data packet using the first analysis method and obtaining multiple first data, the extracted first data is input into a preset large model, which is an anomaly detection model based on machine learning (such as random forest, deep learning model). Historical traffic data can be collected first, and the historical traffic data includes but is not limited to network packet headers, DNS query logs, SSL / TLS handshake details, and port scan records. New image data sources such as Web page screenshots, malicious email attachment pictures, etc. are added, and then the historical data is preprocessed. The image and text data can be processed respectively based on computer vision technology and natural language processing technology, transformed into high-dimensional feature vectors, and the input data for the large model is obtained. An initial large model integrating a large language model and a large image model is pre-constructed, and the obtained input data is input into the initial large model for training. A large number of hyperparameters are involved in the training process, such as learning rate, batch size, etc. The selection of these hyperparameters is crucial for the training effect. Therefore, hyperparameter tuning is required, and the effects of different hyperparameter combinations are tested through a large number of experiments. When the output result of the model training is within the actual range of the actual result, it is determined that the initial large model at the end of the training is used as the preset large model. Then, multiple first data are input into the preset large model for processing to obtain an anomaly result. According to the output result of the preset large model, it is judged whether the first traffic data packet is abnormal traffic. After determining that the first traffic data packet is abnormal traffic content, corresponding measures need to be taken.

[0061] In addition, when determining that the first traffic data packet is abnormal traffic, effectively investigate and cut off the connection of the device receiving the first traffic data packet, improve the security and controllability of the network, and prevent potential malicious behaviors or abnormal activities from harming the network. Investigate the network devices based on the first traffic data packet to determine the target device, where the target device is the device in the network devices that has received the first traffic data packet; cut off the network connection between the target device and the sending device so that the target device stops running relevant programs, and the sending device is the device that sends the first traffic data packet. Specifically, analyze the first traffic data packet to determine the network path it flows through and the devices involved. Based on the information in the traffic data packet, identify the network devices that have received the packet. The header information of the packet, such as the IP header, MAC address, VLAN tag, etc., can be parsed to determine the source address, destination address, and the devices passed through in the middle of the packet. Combine with the network topology diagram to determine the transmission path of the packet in the network. The logs of the network devices can also be checked to find the records of the devices that have received the first traffic data packet. Compare the identified devices with the network device list to confirm the target device. Extract the source IP address or MAC address from the first traffic data packet to determine the sending device. Configure an access control list (ACL) or firewall rules on network devices (such as switches, routers) to block the communication between the target device and the sending device. In some cases, it may be necessary to physically disconnect the network cable between the target device and the sending device. In a virtualized environment, the connection can be cut off through a virtual switch or security group policy. Ensure that only authorized personnel can perform the connection cut-off operation. Successfully cut off the network connection between the target device and the sending device to prevent potential malicious communication or abnormal behavior. Determine the programs related to the first traffic data packet running on the target device. Take measures to stop the running of relevant programs on the target device. First, identify the running programs on the target device and use system commands to terminate relevant processes. Before terminating the process, confirm the identity and relevance of the process again to avoid misoperation. Ensure that terminating the program will not affect the normal operation of the target device, and perform backup or recovery operations if necessary. Successfully stop the programs related to the first traffic data packet on the target device to prevent potential malicious behaviors or abnormal activities from continuing to occur.

[0062] Further, when the target information does not exist in the preset information table, determine that the first traffic data packet is an unencrypted data packet; retrieve the second analysis method according to the unencrypted data packet, and parse the first traffic data packet based on the second analysis method to obtain second data, where the second data includes text information, image information, and link information; determine whether there is abnormal data in the second data; when there is abnormal data in the second data, determine that the first traffic data packet is abnormal traffic content, and perform an abnormal mark on the first traffic data packet according to the abnormal traffic content. Specifically, search for the target information in the preset information table. The preset information table is a database or list containing known encryption features or identifiers. For example, encrypted traffic may contain specific encryption protocol identifiers, encryption algorithm identifiers, or key exchange information. Use database query or list search operations to check whether the target information exists in the preset information table. If it does not exist, it is considered that the first traffic data packet does not contain encryption features and is determined to be an unencrypted data packet. Accurately distinguishing encrypted traffic and unencrypted traffic provides a basis for subsequent analysis. According to the characteristics of the unencrypted data packet, select a suitable second analysis method. The second analysis method is usually for unencrypted traffic and may include text parsing, image recognition, link extraction, etc. For example, for HTTP traffic, an HTTP protocol parser can be used to extract text information, image links, and URLs. Use the second analysis method to parse the first traffic data packet and extract the second data. Parse the text content in the data packet, such as the HTML text in the HTTP response body. Identify the image data in the data packet, which may require decoding or format conversion. Extract the URL or file path information in the data packet. Extract the second data from the unencrypted data packet to provide a basis for subsequent anomaly detection. Define the characteristics of abnormal data in advance, match the second data with the abnormal data, and determine whether there is abnormal data. Abnormal data includes abnormal keywords, malicious image features, malicious URLs, phishing website links, and abnormal domain names, etc. Compare the text information, image information, and link information with the abnormal data in turn, and then accurately identify whether there is abnormal data in the second data, providing a basis for determining abnormal traffic content. If there is one or more abnormal data in the second data, it is considered that the first traffic data packet contains abnormal traffic content. When there is abnormal data in the second data, determine that the first traffic data packet is abnormal traffic content, and mark the first traffic data packet as abnormal traffic content. The abnormal status of the first traffic data packet can be recorded in the traffic log or database, which can be done by adding tags, modifying the attributes of the data packet, or generating an alert. Select a suitable marking method to perform an abnormal mark on the first traffic data packet and generate a real-time alert to notify the security administrator. Ensure that abnormal traffic content is effectively marked and recorded for subsequent security analysis and response.

[0063] Through the above method, calculate the byte entropy in the first traffic data packet, identify encrypted traffic with high randomness as a preliminary screening indicator, then compare the calculated byte entropy with the preset byte entropy. When the byte entropy is greater than the preset byte entropy, extract the target information in the first traffic data packet, and then match the extracted target information with the preset information table to accurately identify the encrypted data packet. Combine multiple encryption features to determine that the first traffic data packet is an encrypted data packet, reducing misjudgment. Then, determine the first analysis method based on the encrypted data packet, and use the first analysis method to process the first traffic data packet to obtain first data in multiple dimensions. Based on multi-dimensional data analysis, the traffic behavior pattern can be captured. Input the multi-dimensional first data into the preset large model for processing to obtain an abnormal result. Determine that the first traffic data packet is abnormal traffic content according to the abnormal result. Make a comprehensive judgment through multiple indicators, and select different extraction methods according to whether the traffic data packet is encrypted, avoiding the limitations of single-dimensional detection and significantly improving the accuracy of abnormal traffic content recognition.

[0064] An embodiment of the present application further provides an abnormal traffic content recognition device based on a large model. Figure 2 It is a structural schematic diagram of an abnormal traffic content recognition device based on a large model provided by an embodiment of the present application. Refer to Figure 2 The device includes an acquisition unit 201, a processing unit 202, and a determination unit 203.

[0065] The acquisition unit 201 acquires the first traffic data packet; acquires multiple bytes from the first traffic data packet, and calculates the information entropy of the multiple bytes to obtain the byte entropy.

[0066] The processing unit 202 determines whether the byte entropy is less than or equal to the preset byte entropy; when the byte entropy is greater than the preset byte entropy, acquires the target information from the first traffic data packet, and the target information includes duplicate information, padding character information, and file naming information; determines whether the target information exists in the preset information table, and the preset information table is the information composed of corresponding encrypted data; when the target information exists in the preset information table, determines that the first traffic data packet is an encrypted data packet; retrieves the first analysis method based on the encrypted data packet, and processes the first traffic data packet based on the first analysis method to obtain multiple first data, and the multiple first data includes communication frequency, connection time, destination IP address, source IP address, and device port number.

[0067] The determination unit 203 inputs the multiple first data into the preset large model for processing to obtain an abnormal result, and determines that the first traffic data packet is abnormal traffic content according to the abnormal result.

[0068] In a possible implementation, the obtaining unit 2011 is configured to obtain a second traffic data packet when the byte entropy is less than or equal to a preset byte entropy, where the second traffic data packet is a data packet after the first traffic data packet; obtain a first length from the first traffic data packet, and obtain a second length from the second traffic data packet; the processing unit 202 is configured to calculate the first length and the second length to obtain a data packet length entropy; determine whether the data packet length entropy is less than or equal to a preset length entropy; when the data packet length entropy is greater than the preset length entropy, obtain target information from the first traffic data packet.

[0069] In a possible implementation, the processing unit 202 is configured to obtain a third traffic data packet when the data packet length entropy is less than or equal to a preset length entropy, where the third traffic data packet is a data packet received after the second traffic data packet; the obtaining unit 201 is configured to obtain a first time interval, where the first time interval is the time interval between a first arrival time and a second arrival time, the first arrival time is the time when the first traffic data packet is obtained, and the second arrival time is the time when the second traffic data packet is obtained; obtain a second time interval, where the second time interval is the time interval between the second arrival time and a third arrival time, and the third arrival time is the time when the third traffic data packet is obtained; the processing unit 202 is configured to calculate the first time interval and the second time interval to obtain a time interval entropy; determine whether the time interval entropy is less than or equal to a preset time entropy; when the time interval entropy is greater than the preset time entropy, obtain target information from the first traffic data packet.

[0070] In a possible implementation, when the target information does not exist in the preset information table, the processing unit 202 is configured to determine that the first traffic data packet is an unencrypted data packet; retrieve a second analysis method according to the unencrypted data packet, and parse the first traffic data packet based on the second analysis method to obtain second data, where the second data includes text information, image information, and link information; determine whether there is abnormal data in the second data; when there is abnormal data in the second data, the determining unit 203 is configured to determine that the first traffic data packet is abnormal traffic content, and perform an abnormal mark on the first traffic data packet according to the abnormal traffic content.

[0071] In a possible implementation, the processing unit 202 is configured to analyze the first traffic data packet to obtain an encryption protocol; determine whether the encryption protocol exists in a preset encryption protocol table; when the encryption protocol does not exist in the preset encryption protocol, determine to use a first analysis method to process the first traffic data packet.

[0072] In a possible implementation, the processing unit 202 is configured to, when there is an encryption protocol in the preset encryption protocol table, decrypt the first traffic data packet according to the encryption protocol to obtain a decrypted traffic data packet; parse the decrypted traffic data packet to obtain third data, where the third data includes text data and image data; determine whether there are abnormal keywords in the text data and whether the target similarity is less than or equal to a preset similarity, where the target similarity is the similarity between the image corresponding to the image data and a suspicious image; the determination unit 203 is configured to, when there are no abnormal keywords in the text data and the target similarity is less than or equal to the preset similarity, determine that the first traffic data packet is normal traffic content.

[0073] In a possible implementation, the processing unit 202 is configured to troubleshoot the network device according to the first traffic data packet to determine a target device, where the target device is the device in the network device that has received the first traffic data packet; cut off the network connection between the target device and the sending device so that the target device stops running relevant programs, and the sending device is the device that sends the first traffic data packet.

[0074] It should be noted that: when the device provided in the above embodiment implements its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0075] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 302, and at least one communication bus 305.

[0076] Among them, the communication bus 305 is used to realize the connection and communication between these components.

[0077] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0078] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0079] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 302, and by invoking the data stored in the memory 302, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application requests, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0080] Among them, the memory 302 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 302 includes a non-transitory computer-readable storage medium. The memory 302 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 302 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 302 may also be at least one storage device located far from the aforementioned processor 301.

[0081] As Figure 3 shown, the memory 302, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for identifying abnormal traffic content based on a large model.

[0082] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 302 for identifying abnormal traffic content based on the large model. When executed by one or more processors, the electronic device performs one or more of the methods described in the above embodiments.

[0083] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0084] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0085] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0086] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0088] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0089] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the embodiments, those skilled in the art will readily think of other implementation manners of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure.

Claims

1. A method for identifying abnormal traffic content based on a large model, characterized in that: The method comprises: Obtaining a first traffic data packet; Acquire multiple bytes from the first traffic data packet, and perform information entropy calculation on the multiple bytes to obtain byte entropy; Determining whether the byte entropy is less than or equal to a preset byte entropy; When the byte entropy is greater than the preset byte entropy, obtaining target information from the first traffic data packet, the target information including repetition information, padding character information, and file naming information; Determine whether the target information exists in a preset information table, where the preset information table is information corresponding to the encrypted data; When the target information exists in the preset information table, determining that the first traffic data packet is an encrypted data packet; Retrieving a first analysis method according to the encrypted data packet, processing the first traffic data packet based on the first analysis method to obtain a plurality of first data, wherein the plurality of first data include a communication frequency, a connection time, a destination IP address, a source IP address, and a device port number; Input a plurality of the first data into a preset large model for processing to obtain abnormal results, and determine that the first traffic data packet is abnormal traffic content according to the abnormal results.

2. The method according to claim 1, characterized in that After determining whether the byte entropy is less than or equal to a preset byte entropy, the method further includes: When the byte entropy is less than or equal to the preset byte entropy, obtaining a second traffic data packet, where the second traffic data packet is a data packet after the first traffic data packet; Obtaining a first length from the first traffic data packet, and obtaining a second length from the second traffic data packet; Calculating the first length and the second length to obtain data packet length entropy; Determining whether the data packet length entropy is less than or equal to a preset length entropy; When the data packet length entropy is greater than the preset length entropy, the target information is obtained from the first traffic data packet.

3. The method according to claim 2, characterized in that After determining whether the data packet length entropy is less than or equal to a preset length entropy, the method further includes: When the data packet length entropy is less than or equal to the preset length entropy, obtaining a third traffic data packet, where the third traffic data packet is a data packet received after the second traffic data packet; Obtain a first time interval, where the first time interval is a time interval between a first arrival time and a second arrival time, where the first arrival time is a time for obtaining the first traffic data packet, and the second arrival time is a time for obtaining the second traffic data packet; Obtain a second time interval, where the second time interval is a time interval between the second arrival time and a third arrival time, and the third arrival time is a time for obtaining the third traffic data packet; Calculating the first time interval and the second time interval to obtain time interval entropy; Determining whether the time interval entropy is less than or equal to a preset time entropy; When the time interval entropy is greater than the preset time entropy, the target information is obtained from the first traffic data packet.

4. The method according to claim 1, characterized in that After determining whether the target information exists in the preset information table, the method further includes: When the target information does not exist in the preset information table, determining that the first traffic data packet is a non-encrypted data packet; Retrieving a second analysis method according to the non-encrypted data packet, parsing the first traffic data packet based on the second analysis method to obtain second data, where the second data includes text information, image information, and link information; Determining whether there is abnormal data in the second data; When the abnormal data exists in the second data, it is determined that the first traffic data packet is the abnormal traffic content, and the first traffic data packet is marked as abnormal according to the abnormal traffic content.

5. The method according to claim 1, characterized in that When the target information exists in the preset information table, after determining that the first traffic data packet is an encrypted data packet, the method further includes: Analyzing the first traffic data packet to obtain an encryption protocol; Determine whether the encryption protocol exists in the preset encryption protocol table; When the encryption protocol does not exist in the preset encryption protocols, it is determined to use the first analysis method to process the first traffic data packet.

6. The method according to claim 5, characterized in that After determining whether the encryption protocol exists in the preset encryption protocol table, the method further includes: When the encryption protocol exists in the preset encryption protocol table, decrypting the first traffic data packet according to the encryption protocol to obtain a decrypted traffic data packet; Parsing the decrypted traffic data packet to obtain third data, wherein the third data includes text data and image data; Determine whether there are abnormal keywords in the text data, and whether the target similarity is less than or equal to a preset similarity, where the target similarity is the similarity between the image corresponding to the image data and the suspicious image; When the abnormal keyword does not exist in the text data and the target similarity is less than or equal to the preset similarity, it is determined that the first traffic data packet is normal traffic content.

7. The method according to claim 1, characterized in that After inputting the plurality of first data into a preset large model for processing to obtain an abnormal result, and determining that the first traffic data packet is abnormal traffic content according to the abnormal result, the method further includes: Check the network devices according to the first traffic data packet to determine the target device, where the target device is a device in the network devices that has received the first traffic data packet; The network connection between the target device and the sending device is cut off so that the target device stops running related programs. The sending device is the device that sends the first traffic data packet.

8. A device for identifying abnormal traffic content based on a large model, characterized in that: The device comprises an acquisition unit (201), a processing unit (202) and a determination unit (203); The acquisition unit (201) acquires a first traffic data packet; acquires a plurality of bytes from the first traffic data packet, and performs information entropy calculation on the plurality of bytes to obtain byte entropy; The processing unit (202) determines whether the byte entropy is less than or equal to a preset byte entropy; when the byte entropy is greater than the preset byte entropy, obtains target information from the first traffic data packet, the target information including repetition information, fill character information and file naming information; determines whether the target information exists in a preset information table, the preset information table is information corresponding to the encrypted data; when the target information exists in the preset information table, determines that the first traffic data packet is an encrypted data packet; calls a first analysis method according to the encrypted data packet, processes the first traffic data packet based on the first analysis method, and obtains a plurality of first data, the plurality of first data including a communication frequency, a connection time, a destination IP address, a source IP address and a device port number; The determination unit (203) inputs the plurality of the first data into a preset large model for processing to obtain an abnormal result, and determines that the first traffic data packet is an abnormal traffic content according to the abnormal result.

9. An electronic device, characterized in that: The electronic device (300) comprises a processor (301), a memory (302), a user interface (303) and a network interface (304), wherein the memory (302) is used to store instructions, the user interface (303) and the network interface (304) are used to communicate with other devices, and the processor (301) is used to execute the instructions stored in the memory (302) so that the electronic device (300) executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Computer data management method and device based on big data and storage medium

    CN120434060A