Virus detection method, device, equipment and product
By extracting and analyzing the feature information of the user terminal data packets and matching them with the preset feature information in the labeled dataset, the problem of low virus detection efficiency in the prior art is solved, and fast and effective virus detection is achieved, reducing detection costs.
Patent Information
- Application Number
- CN202510104453.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
Smart Images

Figure CN119945780A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to virus detection methods, devices, equipment and products. Background Art
[0002] At present, in cloud security networks that use blockchain to transmit data, network viruses usually invade computer systems through various channels. A common way of transmission is to induce users to download seemingly harmless files that carry malicious network virus programs through the operating behaviors of the control end and the controlled end. In this case, there is a big difference between the data packets of users during normal browsing and the data packets carrying malicious network virus programs. Network viruses can use the characteristics of self-replication to cause normal data packets to be lost or redundant. Therefore, how to quickly identify the security of user data packets under the cloud security platform is crucial. In related technologies, virus detection is achieved through the detection program by installing a detection program in the user terminal. Since viruses are highly variable and complex, and the characteristics and behavior patterns of viruses are constantly changing. If the detection program cannot be updated or upgraded in time to adapt to the characteristics of new viruses, the efficiency of virus detection will be reduced. Summary of the invention
[0003] The main purpose of this application is to provide a virus detection method, device, equipment and product, aiming to solve the technical problem of low virus detection efficiency.
[0004] To achieve the above objectives, the present application proposes a virus detection method, comprising: Obtaining a data packet on a user terminal and extracting characteristic information of the data packet, wherein the characteristic information includes characteristic information of each blockchain node through which the data packet passes; Acquire preset characteristic information matched by the characteristic information from the marked data set, wherein the marked data set includes a normal data packet and an abnormal data packet, and the normal data packet and the abnormal data packet are marked with corresponding preset characteristic information; Determine the virus detection result of the data packet based on the preset feature information.
[0005] In one embodiment, the virus detection method further includes: Obtain preset feature information corresponding to each blockchain node that a single data packet passes through; Determine the similarity of preset feature information of every two adjacent blockchain nodes; According to the similarity of the preset feature information of every two adjacent blockchain nodes, the mean similarity of a single data packet to all blockchain nodes is determined; Determine the marking result of a single data packet according to the similarity threshold and similarity mean corresponding to each two adjacent blockchain nodes of the single data packet; According to the labeling results of each data packet, a labeled data set is constructed.
[0006] In one embodiment, obtaining preset characteristic information corresponding to each blockchain node passed by a single data packet includes: Obtain the latency, jitter, packet length, packet transmission throughput, and transmission protocol level of each blockchain node that a single data packet passes through; According to the delay time, jitter time, data packet length, data packet transmission throughput and transmission protocol level of each blockchain node, the preset characteristic information corresponding to each blockchain node is obtained.
[0007] In one embodiment, determining the marking result of a single data packet according to the similarity threshold and the similarity mean corresponding to each two adjacent blockchain nodes of the single data packet includes: If the similarity thresholds corresponding to the two adjacent blockchain nodes of a single data packet are both less than or equal to the average similarity value of the single data packet, the marking result of the single data packet is determined to be a normal data packet; Alternatively, if the similarity thresholds corresponding to two adjacent blockchain nodes of a single data packet are both greater than the average similarity of the single data packet, the marking result of the single data packet is determined to be an abnormal data packet.
[0008] In one embodiment, obtaining a data packet on a user terminal and extracting characteristic information of the data packet includes: Obtain the data packet on the user terminal, and perform word segmentation on the data in the data packet to obtain the data code and data content; Extract features from data coding and data content to obtain text feature vectors and content feature vectors; The text feature vectors and content feature vectors are sorted in chronological order to obtain time series data, and the time series data are processed using a long short-term memory network model to obtain a time series feature vector; A fused feature vector is obtained according to the first weights between text feature vectors and text feature vectors, the second weights between content feature vectors and content feature vectors, and the third weights between timing feature vectors and timing feature vectors, wherein the feature information of the data packet includes the fused feature vector.
[0009] In one embodiment, obtaining preset feature information matching feature information from the labeled data set includes: Inputting the characteristic information of the data packet into the virus detection model, the virus detection model obtains preset characteristic information matching the characteristic information from the labeled data set, and determines the timing detection result of the data packet according to the preset characteristic information; wherein the virus detection model is obtained by training the initial recurrent neural network with normal data packet samples and abnormal data packet samples, and the normal data packet samples and the abnormal data packet samples are both marked with corresponding preset characteristic information and corresponding timing detection results; According to the preset characteristic information, the virus detection result of the data packet is determined to include: According to the timing detection result, the virus detection result of the data packet is determined.
[0010] In one embodiment, determining the virus detection result of the data packet according to the timing detection result includes: The time series detection result is input into the Sigmoid activation function, and the relationship between the time series detection result and the probability threshold is calculated to obtain the detection result; According to the detection results, the time series detection results are reconstructed through the autoencoder for error or the similarity with any virus type is calculated; If the similarity exceeds a similarity threshold or the reconstructed error exceeds an error threshold, the data packet is determined to be an abnormal data packet, and the virus type of the abnormal data packet is obtained.
[0011] In addition, to achieve the above objectives, the present application also proposes a virus detection device, comprising: An acquisition and extraction module is used to acquire a data packet on a user terminal and extract characteristic information of the data packet, wherein the characteristic information includes characteristic information of each blockchain node through which the data packet passes; and to acquire preset characteristic information matching the characteristic information from a marked data set, wherein the marked data set includes a normal data packet and an abnormal data packet, and the normal data packet and the abnormal data packet are marked with corresponding preset characteristic information; The virus detection module is used to determine the virus detection result of the data packet according to preset characteristic information.
[0012] In addition, to achieve the above objectives, the present application also proposes a virus detection device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the virus detection method as described above.
[0013] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the virus detection method as described above are implemented.
[0014] The present application obtains data packets on the user terminal, extracts characteristic information of the data packets, obtains preset characteristic information that matches the characteristic information from the marked data set, and then determines the virus detection result of the data packet based on the preset characteristic information. Since the marked data set includes normal data packets and abnormal data packets, by classifying and identifying normal data packets and abnormal data packets, it can be determined whether the data packets on the user terminal are normal or abnormal based on the existing marked data packets, so that when similar abnormal data is encountered again, the virus can be quickly detected without re-detection based on the virus detection program, which can improve the virus detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] Figure 1 A schematic diagram of a process of a virus detection method in some embodiments of the present application; Figure 2 Another schematic diagram of the process of virus detection method in some embodiments of the present application; Figure 3 This is a detailed flow chart of step S10 of the virus detection method in some embodiments of the present application; Figure 4 This is a schematic diagram of the module structure of a virus detection device in some embodiments of the present application; Figure 5 This is a schematic diagram of the structure of a virus detection device in some embodiments of the present application.
[0018] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0020] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0021] At present, cloud security technology mainly refers to a new form of network computing generated under the influence of modern technologies such as network communication technology and big data processing technology. Cloud security technology has powerful computing, storage and processing capabilities, and is being fully applied in various fields in my country. In cloud security networks that use blockchain to transmit data, network viruses usually invade computer systems through various channels. A common way of transmission is to induce users to download seemingly harmless files that carry malicious network virus programs through the operation behavior of the control end and the controlled end. In this case, there is a big difference between the data packets of users during normal browsing and the data packets carrying malicious network virus programs. Network viruses can use the characteristics of self-replication to cause normal data packets to be lost or redundant. Therefore, how to quickly identify the security of user data packets under the cloud security platform is crucial. In related technologies, virus detection is achieved by installing a detection program in the user terminal. Since viruses are highly variable and complex, and the characteristics and behavior patterns of viruses are constantly changing. If the detection program cannot be updated or upgraded in time to adapt to the characteristics of new viruses, the efficiency of virus detection will be reduced.
[0022] In view of the above problems, the present application proposes a virus detection method, the main technical scheme of which includes: obtaining the data packet on the user terminal, extracting the characteristic information of the data packet, obtaining the preset characteristic information matching the characteristic information from the marked data set, and then determining the virus detection result of the data packet according to the preset characteristic information. Since the marked data set includes normal data packets and abnormal data packets, by classifying and identifying the normal data packets and the abnormal data packets, it can be determined whether the data packet on the user terminal is normal or abnormal based on the existing marked data packets, so that the virus can be quickly detected when encountering similar abnormal data again, without installing a detection program on the user terminal, and re-detecting based on the virus detection program, which can improve the efficiency of virus detection and reduce the cost of virus detection.
[0023] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a virus detection device capable of realizing the above functions, etc. The following takes the virus detection device as an example to illustrate this embodiment and the following embodiments.
[0024] It should be noted that the present application provides an analysis of data packets on user terminals based on blockchain and cloud security to improve the security and reliability of user access. Specifically, by obtaining user feedback data in historical cloud logs, a labeled data set is established, data packets in the labeled data set are captured, characteristic information of data packets are extracted, and normal data packets and abnormal data packets are marked respectively; the labeled data set is preprocessed; at the same time, a virus detection model is constructed, the virus detection model is trained using the labeled data set in the training set, and the accuracy of the virus detection model is verified using the test set; finally, the data packets on the user terminal are obtained in real time, characteristic information of the data packets is extracted, and it is determined whether the current data packet is abnormal. If the current data packet is abnormal, the user is given an early warning maintenance reminder to improve the security and reliability of user access.
[0025] It should be noted that the virus detection method of the present application is applicable to private chains, public chains and consortium chains.
[0026] Private chain: Since private chains are usually controlled by a single organization or enterprise, they have high requirements for data privacy and security. This application analyzes data packets on user terminals and establishes a virus detection model to identify abnormal behavior, which is suitable for private chain environments. In this environment, data access can be more effectively monitored and controlled to ensure the security of internal data.
[0027] Public chain: Public chain is open to everyone, so it is particularly important to implement data packet analysis on user terminals on public chain. This application can be used to protect users on public chain from malicious attacks and improve the security and reliability of the entire network.
[0028] Alliance chain: Alliance chain is a blockchain jointly maintained by multiple organizations, which shares data among members. The solution of this application is also applicable to alliance chain, because it can help alliance members monitor data access behavior, so that authorized users can access specific data, and can also detect and handle abnormal behavior in a timely manner.
[0029] It should be noted that the node configuration of the blockchain includes hardware resource configuration, software resource configuration, and network configuration. Hardware resource configuration includes: CPU (Central Processing Unit), memory, storage, and network bandwidth. For the CPU, its processing power needs to match the tasks performed by the node, such as verifying transactions, executing smart contracts, etc. For memory, sufficient memory is used to process transactions and block data. For storage, sufficient storage space is required to save a complete copy of the blockchain. For network bandwidth, high-bandwidth network connections are required to ensure that nodes can synchronize data quickly. Software configuration includes blockchain client, operating system, and security configuration. For blockchain client, select appropriate client software and customize it as needed. For operating system, select a stable and secure operating system. For security, configure firewalls, encrypted communications, and other security measures. Network configuration includes node discovery and port mapping. For node discovery, configure the node so that it can discover other nodes in the network. For port mapping, ensure that the node can receive connections from other nodes.
[0030] For example, for the node configuration of the blockchain, in a small network, the CPU can be of medium performance, the memory can be 4-8GB, and the storage can be at least 256GB SSD (Solid State Drive) to accommodate fast read and write requirements. In medium to large networks, the CPU can be a high-performance multi-core processor, the memory can be more than 16GB, and the storage can be at least 1TB SSD or higher to store the complete blockchain data. In an ultra-large network, the CPU can be a server-level multi-core processor, the memory can be 64GB or higher, and the storage can be multiple TB-level SSDs or RAID (Redundant Arrays of Independent Disks) configuration.
[0031] It should be noted that for the number of nodes in a blockchain, if the blockchain adopts a public chain, the more nodes the blockchain has, the more robust the network will be. However, considering performance and resource constraints, it may be necessary to set certain thresholds, such as requiring nodes to have certain computing power and storage space. If the blockchain adopts a private chain / consortium chain, the number of nodes can be determined based on the required fault tolerance, such as Byzantine fault tolerance. For example, for Byzantine fault tolerance, at least more than 2 / 3 of the honest nodes in the network are required. The number of nodes can also be determined based on the expected transaction volume and processing speed.
[0032] It should be noted that the virus detection system architecture of this application includes: The data acquisition module is used to obtain user feedback data from historical cloud logs. Access the cloud service provider's log management system through the API (Application Programming Interface) interface or log collection tools. Use log parsing tools to filter out log entries related to user feedback based on specific keywords or log levels. You can write scripts or use data processing tools to extract specific user feedback data from the filtered logs. It can be set as a scheduled task such as hourly, daily, or real-time streaming processing based on system requirements and resource conditions. It usually includes log data from a recent period of time, such as the last month or three months, depending on storage capacity and analysis requirements.
[0033] The marking module includes a feature extraction unit and a marking unit. The feature extraction unit is used to capture the data packets in the marked data set and extract the feature information of the data packets; the marking unit is used to mark normal data packets and abnormal data packets. In the specific implementation, the features that may be extracted include: source IP (Internet Protocol) address, destination IP address, port number, protocol type, data packet size, data packet transmission time, request type such as HTTP (Hypertext Transfer Protocol) request method, user agent information, session duration, traffic pattern such as burst traffic, periodic traffic. Feature selection criteria and rules include: features should be related to user behavior and security, features should be able to effectively distinguish normal and abnormal behaviors, features should remain stable within different time windows, and features should be easy to understand in order to analyze the causes of abnormalities.
[0034] A preprocessing module is used to split the labeled dataset into training and testing sets.
[0035] The prediction model building and analysis module includes a prediction model building unit and an analysis unit.
[0036] The prediction model is used to train the model using the user data packets of the training set; the analysis unit tests the accuracy of the model using the user data packets of the test set.
[0037] The monitoring module includes a feature real-time monitoring unit and a discrimination unit. In the specific implementation, a network sniffer or a packet capture library is used to capture packets in real time. An efficient packet processing algorithm is implemented, such as multi-threading or asynchronous I / O. In-memory data structures such as hash tables and trees are used to quickly retrieve and update packet features. The packet processing time is reduced by optimizing algorithms and hardware resources. The accuracy of feature extraction and classification is improved by continuously training and optimizing prediction models. Among them, the feature real-time monitoring unit is used to obtain the user's data packets on the terminal network in real time and extract the user's data packet features; the discrimination unit is used to determine whether the current user data packet is abnormal; The early warning maintenance module is used to provide early warning maintenance reminders to users when abnormalities are detected in user data packets. In actual implementation, the early warning mechanism includes the identification unit detecting abnormal data packets, triggering early warnings according to preset rules, notifying users through emails, text messages, application push, etc., and recording abnormal events and notification details for subsequent analysis. Among them, the appropriate notification method can be selected according to user preferences and urgency.
[0038] Based on this, some embodiments of the present application provide a virus detection method, referring to Figure 1 , Figure 1 The flowchart of some embodiments of the virus detection method of the present application is shown in FIG. In the embodiments of the present application, the virus detection method includes steps S10 to S30: Step S10, obtaining a data packet on the user terminal and extracting characteristic information of the data packet, wherein the characteristic information includes characteristic information of each blockchain node through which the data packet passes.
[0039] You can use a network packet capture tool to capture packets on a network interface; or you can develop a custom capture program, write a program in a programming language, directly access the network interface, or use the API provided by the operating system to capture packets. It should be noted that capturing packets requires appropriate permissions and may need to be run in administrator mode. The captured packets may contain sensitive information, so user authorization is required.
[0040] After capturing the data packets, useful data is extracted from the captured data packets. The structure and content of the data packets can be parsed according to the network protocol used by the data packets, such as HTTP. The required data, such as the header information, request body, response body, etc. of the HTTP request, is extracted from the parsed data packets. The extracted data is cleaned and formatted for word segmentation, including removing irrelevant information such as unnecessary fields in the HTTP header and advertising content in the response body; text encoding conversion to ensure that the data is UTF-8 (a variable-length character encoding method) or other encoding formats suitable for word segmentation; data formatting organizes text data into a format suitable for word segmentation, such as removing HTML (Hypertext Markup Language, the standard language for building web pages) tags and extracting plain text.
[0041] The preprocessed text data is divided into words or phrases to obtain data encoding such as word identifiers and data content, i.e. the segmentation result. Specifically, segmentation algorithms such as rule-based segmentation, statistical-based segmentation such as hidden Markov models, conditional random fields, and deep learning-based segmentation can be used. Select the segmentation tool that provides ready-made segmentation algorithms and interfaces. Assign a unique identifier or code to each segmentation result for subsequent processing.
[0042] Output the data encoding and data content after word segmentation for subsequent processing. The word segmentation results can be stored in a relational database or a non-relational database. Alternatively, the word segmentation results can be written into a text file or other format file. Alternatively, the word segmentation results can be passed to a virus detection model for training or prediction.
[0043] After obtaining the data coding and data content output, feature extraction can be performed based on the data coding and data content to obtain a text feature vector and a content feature vector, and the text feature vector and the content feature vector are used as feature information of the data packet.
[0044] Step S20, obtaining preset characteristic information that matches the characteristic information from the marked data set, wherein the marked data set includes normal data packets and abnormal data packets, and the normal data packets and the abnormal data packets are marked with corresponding preset characteristic information.
[0045] Obtain user feedback data from historical cloud logs, capture data packets in the user feedback data, extract feature information of data packets, and mark normal data packets and abnormal data packets respectively, and obtain a marked data set based on the marked normal data packets and abnormal data packets. Among them, historical cloud logs are user feedback data of different users on the blockchain.
[0046] Step S30, determining the virus detection result of the data packet according to the preset characteristic information.
[0047] A preset text feature vector matched by a text feature vector and a preset content feature vector matched by a content feature vector can be obtained from the marked data set, and then the virus detection result of the data packet is determined according to the preset text feature vector and the preset content feature vector.
[0048] In the embodiment of the present application, by acquiring a data packet on a user terminal, extracting characteristic information of the data packet, acquiring preset characteristic information that matches the characteristic information from a marked data set, and then determining the virus detection result of the data packet according to the preset characteristic information. Since the marked data set includes normal data packets and abnormal data packets, by classifying and identifying normal data packets and abnormal data packets, it is possible to determine whether the data packet on the user terminal is normal or abnormal according to the existing marked data packets, so that when similar abnormal data is encountered again, the virus can be quickly detected, without installing a detection program on the user terminal, and re-detection is performed based on the virus detection program, which can improve the efficiency of virus detection and reduce the cost of virus detection.
[0049] Reference Figure 2 In some embodiments of the present application, before step S20 or before step S10, the virus detection method further includes: Step S01, obtaining preset feature information corresponding to each blockchain node passed by a single data packet.
[0050] Each blockchain node has corresponding preset characteristic information, which at least includes delay time, jitter time, data packet length, data packet transmission throughput rate and transmission protocol level.
[0051] In a feasible implementation, the delay time, jitter time, packet length, data packet transmission throughput rate, and transmission protocol level of each blockchain node that a single data packet passes through can be obtained. According to the delay time, jitter time, packet length, data packet transmission throughput rate, and transmission protocol level of each blockchain node, the preset characteristic information corresponding to each blockchain node is obtained. For example, the preset characteristic information of the i-th blockchain node is recorded as , ; The next blockchain node of the i-th blockchain node is the j-th node, and the preset feature information of the j-th blockchain node is recorded as , .
[0052] Latency refers to the time it takes for a data packet to be sent from one blockchain node to another. Latency includes one-way latency and round-trip latency. One-way latency refers to the time it takes for a data packet to be transmitted from a source node to a destination node. For example, it records the time it takes for a data packet to be sent from the source node. , record the time when the data packet is received at the destination node , the one-way delay is Round trip delay refers to the time it takes for a data packet to travel between the source node and the destination node. For example, if you send a data packet and record the time it takes to send it, ; When a data packet is received by the destination node and immediately returned, the receiving time returned to the source node is recorded The round trip delay is .
[0053] Jitter time refers to the variation of packet delay, which reflects the stability of network transmission. Jitter time includes average jitter and maximum jitter. Jitter can be calculated by the following steps: measuring the delay of multiple packets in continuous transmission to obtain a series of delay times: G1, G2, G3, ..., G n . Calculate the difference between these delay times: G1=|G2-G1|; G2=|G3-G2|; G3=|G4-G3|; ...; G n-1 =|G n -G n-1 |. The average or maximum of these differences is calculated to represent the jitter: Average jitter: ; Maximum jitter: .
[0054] The packet length is usually one of the fields in the packet header and can be read directly.
[0055] The data packet transmission throughput refers to the amount of data successfully transmitted per unit time.
[0056] The transport protocol level is usually a categorical variable indicating the type of protocol used by the packet and does not need to be calculated.
[0057] Step S02, determining the similarity of preset feature information of every two adjacent blockchain nodes.
[0058] The similarity can be cosine similarity, Euclidean distance, or Manhattan distance, etc. The following formula can be used to calculate the similarity of the preset feature information of each two adjacent blockchain nodes: .
[0059] in, Indicates the similarity between the preset feature information of the i-th blockchain node and the preset feature information of the j-th blockchain node. Indicates the preset feature information of the i-th blockchain node. Since there are multiple preset feature information of the i-th blockchain node, superscripts 1 to 5 can be used to distinguish them. Indicates the preset feature information of the j-th blockchain node. Since there are multiple preset feature information of the j-th blockchain node, superscripts 1 to 5 can be used to distinguish them.
[0060] Step S03, determining the mean similarity of a single data packet to all blockchain nodes based on the similarity of preset feature information of every two adjacent blockchain nodes.
[0061] The following formula can be used to calculate the mean similarity of a single data packet to all blockchain nodes: .
[0062] in, Represents the mean similarity of a single data packet to all blockchain nodes.
[0063] Step S04, determining the marking result of the single data packet according to the similarity threshold and the similarity mean corresponding to each two adjacent blockchain nodes of the single data packet.
[0064] It should be noted that the similarity thresholds corresponding to two adjacent blockchain nodes of a single data packet may be the same or different.
[0065] It should be noted that the marking results of the data packet include normal or abnormal.
[0066] In a feasible implementation, if the similarity thresholds corresponding to each two adjacent blockchain nodes of a single data packet are less than or equal to the average similarity of the single data packet, the marking result of the single data packet is determined to be a normal data packet; if the similarity thresholds corresponding to each two adjacent blockchain nodes of a single data packet are greater than the average similarity of the single data packet, the marking result of the single data packet is determined to be an abnormal data packet.
[0067] Set the similarity threshold of any data packet between two adjacent blockchain nodes, denoted as ; When the average similarity of a single data packet in n blockchain nodes exceeds the similarity threshold of any data packet in two adjacent blockchain nodes When , a single data packet is marked as a normal data packet; when the average similarity of a single data packet in n blockchain nodes does not exceed the similarity threshold of any data packet in two adjacent blockchain nodes When a packet is detected, a single packet is marked as an abnormal packet.
[0068] In another feasible implementation manner, the similarity threshold may be determined in any of the following ways: 1. Calculate the mean and standard deviation of the similarity of all normal data packets, and then set the threshold to the mean plus a certain multiple, such as 1 or 2 times the standard deviation.
[0069] 2. Use quartiles to set the similarity threshold, which can usually be set to an IQR (interquartile range) that is a certain multiple higher than the third quartile (Q3).
[0070] 3. Use cross-validation to determine the similarity threshold. Divide the dataset into a training set and a validation set. Select the best similarity threshold by adjusting and observing the performance on the validation set.
[0071] 4. By plotting the receiver operating characteristic curve and calculating the area under the curve, select the similarity threshold that best balances the true positive rate and the false positive rate.
[0072] Step S05: construct a labeled data set according to the labeling results of each data packet.
[0073] The labeling results of each data packet are combined to obtain a labeled data set.
[0074] In an embodiment of the present application, preset feature information corresponding to each blockchain node passed by a single data packet is obtained, the similarity of the preset feature information of each two adjacent blockchain nodes is determined, and the similarity mean of the single data packet for all blockchain nodes is determined according to the similarity of the preset feature information of each two adjacent blockchain nodes. The marking result of the single data packet is determined according to the similarity threshold and the similarity mean corresponding to each two adjacent blockchain nodes of the single data packet. Finally, a marked data set is constructed according to the marking result of each data packet. The marked data set is constructed in the above manner. Since the marked data set includes normal data packets and abnormal data packets, by classifying and identifying normal data packets and abnormal data packets, it is possible to determine whether the data packet on the user terminal is normal or abnormal based on the existing marked data packets, so that when similar abnormal data is encountered again, the virus can be quickly detected, and there is no need to install a detection program on the user terminal. Re-detection is performed based on the virus detection program, which can improve the virus detection efficiency and reduce the virus detection cost.
[0075] Reference Figure 3 In some embodiments of the present application, obtaining a data packet on a user terminal and extracting characteristic information of the data packet includes: Step S11, obtaining a data packet on the user terminal, and performing word segmentation processing on the data in the data packet to obtain data coding and data content.
[0076] Jieba word segmenter (jieba word segmenter is a very important open source tool in the field of Chinese natural language processing) can be used to segment the data in the data packet to obtain the data encoding and data content.
[0077] Step S12, extracting features from the data encoding and data content to obtain a text feature vector and a content feature vector.
[0078] In a feasible implementation manner, the data encoding and the data content are subjected to de-meaning and normalization processing to obtain a text feature vector and a content feature vector.
[0079] In another feasible implementation, for the extraction of text feature vectors, the bag-of-words model can be used to count the number of occurrences of each word in the text to form a sparse matrix. Alternatively, TF-IDF (term frequency–inverse document frequency, a commonly used weighting technique for information retrieval and data mining) is used to consider the frequency of words in the document and the inverse document frequency in the entire corpus to measure the importance of the word. Alternatively, word embedding methods such as Word2Vec (a tool for generating word vectors) are used to map words to a high-dimensional vector space to capture the semantic relationship between words. For the extraction of content feature vectors, the sentence structure of the text can be parsed to extract syntactic features such as subject, predicate, object, and dependency relationships. Alternatively, the meaning and context of the text can be understood to extract semantic features such as entities, relationships, and events.
[0080] Step S13, sorting the text feature vectors and the content feature vectors in chronological order to obtain time series data, and processing the time series data using a long short-term memory network model to obtain a time series feature vector.
[0081] The long short-term memory network is a special recurrent neural network that can capture long-term dependencies in sequence data. Using the long short-term memory network model to process time series data and obtain time series feature vectors includes: data preparation, long short-term memory network model construction, long short-term memory network model training, feature extraction, etc.
[0082] Data preparation: It is used to ensure that time series data can be processed by LSTM (Long Short-Term Memory). This includes removing irrelevant data and processing missing values. Time series data is converted into a format suitable for LSTM processing, usually including scaling the data to a specific range such as 0-1. Time series data is divided into multiple fixed-length sequences as input to the LSTM model.
[0083] Long short-term memory network model construction: used to build an LSTM model suitable for processing time series data. The LSTM model includes an input layer for defining the shape of the input data, usually the sequence length and the number of features, an LSTM layer for adding LSTM layers, configuring the number of units, i.e. the memory capacity and other parameters of the LSTM unit, and an output layer for defining the output layer according to task requirements. For feature extraction tasks, the output layer may be a fully connected layer, which is used to convert the output of the LSTM into a fixed-length feature vector.
[0084] Long short-term memory network model training is used to train the LSTM model to learn the characteristics of time series data. Among them, select a suitable loss function, such as mean square error for regression tasks and cross entropy loss for classification tasks. Select an optimization algorithm such as Adam (a first-order gradient-based optimization algorithm) to minimize the loss function. Use the training data to iteratively train the LSTM model and update the model weights through the back-propagation algorithm.
[0085] Feature extraction is used to extract time series feature vectors from the trained LSTM model. This includes inputting time series data into the trained LSTM model to obtain the output of the LSTM layer. As needed, the output of the last time step of the LSTM layer can be selected as the time series feature vector, or the output of all time steps can be selected and aggregated in some form, such as averaging, maximum pooling, etc.
[0086] Step S14, obtaining a fused feature vector according to the first weights between the text feature vectors and the text feature vectors, the second weights between the content feature vectors and the content feature vectors, and the third weights between the timing feature vectors and the timing feature vectors, wherein the feature information of the data packet includes the fused feature vector.
[0087] The text feature vector, content feature vector and time series feature vector are cascaded, and the softmax function (a normalized exponential function used to convert a set of arbitrary real numbers into real numbers representing probability distribution) is used to calculate the weight of each feature vector to obtain a fused feature vector.
[0088] In the embodiment of the present application, by determining a fused feature vector based on a text feature vector, a content feature vector, and a time series feature vector, the extracted features are enriched, thereby improving the accuracy of subsequent virus detection.
[0089] In some embodiments of the present application, obtaining preset feature information matching feature information from the labeled data set includes: Step S21, inputting the characteristic information of the data packet into the virus detection model, the virus detection model obtains preset characteristic information that matches the characteristic information from the labeled data set, and determines the timing detection result of the data packet based on the preset characteristic information; wherein the virus detection model is obtained by training the initial recurrent neural network with normal data packet samples and abnormal data packet samples, and both the normal data packet samples and the abnormal data packet samples are marked with corresponding preset characteristic information and corresponding timing detection results.
[0090] In a feasible implementation, a fused feature vector obtained by fusing the text feature vector, the content feature vector and the timing feature vector is input into a virus detection model. The virus detection model obtains a preset fused feature vector that matches the fused feature vector from the labeled data set, and determines the timing detection result of the data packet based on the preset fused feature vector.
[0091] In another feasible implementation, the text feature vector and the content feature vector are input into the virus detection model, and the virus detection model obtains the preset text feature vector matched by the text feature vector and the preset content feature vector matched by the content feature vector from the marked data set, and then determines the virus detection result of the data packet according to the preset text feature vector and the preset content feature vector.
[0092] The virus detection model can be obtained by using 70% of the labeled data set in the training set to train the initial recurrent neural network and using 30% of the labeled data set to test the initial recurrent neural network. The virus detection model includes an input layer, a hidden layer, and an output layer.
[0093] The number of nodes in the above input layer matches the number of features. For example, if there are 20 features, the input layer has 20 nodes. The number of hidden layers mentioned above usually starts with 1-3 hidden layers and is adjusted according to the model performance. The number of nodes in each layer usually decreases between the input layer and the output layer, such as 64, 32, 16, etc. Common activation functions include nonlinear activation functions, hyperbolic tangent functions or Sigmoid functions. The sigmoid function is used for the output of hidden layer neurons, with a value range of (0,1). It can map a real number to the interval of (0,1) and can be used for binary classification. For the binary classification problem, the output layer usually has 1 node. For the binary classification problem, the Sigmoid function is usually used because it can output probability values between 0 and 1. In the specific implementation, a learning rate decay strategy can be used, such as learning rate decay or adaptive learning rate algorithm. It is determined according to the model performance and training time, usually between tens to hundreds of times. L1, L2 regularization or Dropout (a regularization technique commonly used in model training) is used to prevent overfitting. Confusion matrix is used to evaluate the performance of the model in classification tasks, including true positive rate, false positive rate, precision and recall. The ROC (Receiver Operating Characteristic) curve shows the true positive rate and false positive rate under different thresholds. The closer the AUC (Area Under the Curve) value is to 1, the better the model performance. K-fold cross validation is used to evaluate the generalization ability of the model. The performance indicator F1 score, the harmonic mean of precision and recall, is used to evaluate the accuracy of the model. Loss is used to evaluate the accuracy of the model's predicted probability. Grid search or random search is used to find the optimal hyperparameter combination. The stability of the model performance is observed through multiple training and testing.
[0094] According to the preset characteristic information, the virus detection result of the data packet is determined to include: Step S31, determining the virus detection result of the data packet according to the timing detection result.
[0095] In a feasible implementation, the time series detection result is input into the Sigmoid activation function, and the relationship between the time series detection result and the probability threshold is calculated to obtain the detection result. According to the detection result, the time series detection result is error reconstructed through the autoencoder or the similarity with any virus type is calculated. If the similarity exceeds the similarity threshold or the reconstructed error exceeds the error threshold, the data packet is determined to be an abnormal data packet, and the virus type of the abnormal data packet is obtained. Among them, the probability threshold, the error threshold and the similarity threshold can be set according to the actual situation.
[0096] In another feasible implementation, according to the detection result, reconstructing the error of the time series detection result through an autoencoder or calculating the similarity with any virus type includes: Preprocess the time series detection results to ensure that they are suitable for input into the autoencoder. Specifically, organize the time series detection results into a format suitable for machine learning, such as a two-dimensional array, where each row represents a detection sample and each column represents a time step or feature. Standardize or normalize the detection results to ensure that the data is on the same scale. If the purpose is to calculate the similarity with the virus type, it is necessary to prepare label data, that is, the time series feature vector of the known virus type.
[0097] Next, build an autoencoder model suitable for processing time series detection results. Specifically, the input layer of the autoencoder model defines the shape of the input data to match the sorted time series detection results; build an encoding layer to compress the input data into a low-dimensional representation, i.e., encoding; build a decoding layer to restore the encoding to an approximate representation of the original data, i.e., reconstruction. Select an appropriate loss function, such as mean square error, to measure the reconstruction error. Select an optimization algorithm to minimize the loss function.
[0098] Next, the autoencoder is trained to learn the potential features of the time series detection results. The training process uses a set of training data, which can be normal time series data, to iteratively train the autoencoder and update the model weights through the back propagation algorithm. The performance of the autoencoder is evaluated on the validation set, such as measured by the reconstruction error.
[0099] Next, the trained autoencoder is used to reconstruct the error or calculate the similarity with the virus type. Error reconstruction includes inputting the time series detection results into the trained autoencoder to obtain the reconstruction results. The error between the original detection results and the reconstruction results is calculated, such as using MSE (mean-square error). Samples with large errors may indicate abnormal or potential virus activities. Similarity calculation includes inputting the time series detection results into the encoding layer of the autoencoder to obtain a low-dimensional representation. Calculate the similarity between the encoding and the known virus type encoding, such as using cosine similarity, Euclidean distance, etc. Based on the similarity, determine which virus type the detection result is most similar to.
[0100] In the embodiment of the present application, a virus detection model is constructed. Since the virus detection model adopts a recurrent neural network, the recurrent neural network incorporates the temporal correlation between data into the network calculation compared with most neural networks such as multi-layer perceptrons and BP (back propagation) neural networks. The recurrent neural network connects the nodes between the hidden layers so that neurons at different times can have information from the previous time, and can also continue to pass the information in the neurons at this time downward. In this way, the input of the hidden layer includes not only the input data at the current time, but also the output data of the hidden layer at the previous time. Therefore, the recurrent neural network can process data with time series characteristics very well. According to the time series characteristic data, viruses with propagation properties can be judged. For viruses that do not have propagation properties, monitoring and judgment can also be achieved through the comparison of time series characteristics.
[0101] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the virus detection method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0102] Based on the same inventive concept, this application also provides a virus detection device, please refer to Figure 4 , the virus detection device comprises: The acquisition and extraction module 10 is used to acquire the data packet on the user terminal and extract the characteristic information of the data packet, wherein the characteristic information includes the characteristic information of each blockchain node through which the data packet passes; and acquire the preset characteristic information matching the characteristic information from the marked data set, wherein the marked data set includes normal data packets and abnormal data packets, and the normal data packets and the abnormal data packets are marked with corresponding preset characteristic information.
[0103] The virus detection module 20 is used to determine the virus detection result of the data packet according to preset characteristic information.
[0104] Optionally, the virus detection device also includes a module for constructing a marked data set, which is used to obtain preset feature information corresponding to each blockchain node passed by a single data packet; determine the similarity of the preset feature information of every two adjacent blockchain nodes; determine the mean similarity of the single data packet for all blockchain nodes based on the similarity of the preset feature information of every two adjacent blockchain nodes; determine the marking result of the single data packet based on the similarity threshold and the similarity mean corresponding to each two adjacent blockchain nodes of the single data packet; and construct a marked data set based on the marking results of each data packet.
[0105] Optionally, the construction module of the labeled data set is also used to obtain the delay time, jitter time, data packet length, data packet transmission throughput rate and transmission protocol level of each blockchain node passed by a single data packet; according to the delay time, jitter time, data packet length, data packet transmission throughput rate and transmission protocol level of each blockchain node, the preset feature information corresponding to each blockchain node is obtained.
[0106] Optionally, the construction module of the marked data set is also used to determine that the marking result of the single data packet is a normal data packet if the similarity thresholds corresponding to each two adjacent blockchain nodes of the single data packet are less than or equal to the similarity mean of the single data packet; or, if the similarity thresholds corresponding to each two adjacent blockchain nodes of the single data packet are greater than the similarity mean of the single data packet, determine that the marking result of the single data packet is an abnormal data packet.
[0107] Optionally, the acquisition and extraction module 10 is also used to acquire data packets on the user terminal, and perform word segmentation on the data in the data packets to obtain data codes and data content; perform feature extraction on the data codes and data content to obtain text feature vectors and content feature vectors; sort the text feature vectors and content feature vectors in chronological order to obtain time series data, and use a long short-term memory network model to process the time series data to obtain a time series feature vector; obtain a fused feature vector based on a first weight between text feature vectors and text feature vectors, a second weight between content feature vectors and content feature vectors, and a third weight between time series feature vectors and time series feature vectors, wherein the feature information of the data packet includes the fused feature vector.
[0108] Optionally, the acquisition and extraction module 10 is also used to input the characteristic information of the data packet into the virus detection model, and the virus detection model obtains preset characteristic information that matches the characteristic information from the labeled data set, and determines the timing detection result of the data packet based on the preset characteristic information; wherein the virus detection model is obtained by training the initial recurrent neural network using normal data packet samples and abnormal data packet samples, and the normal data packet samples and the abnormal data packet samples are both marked with corresponding preset characteristic information and corresponding timing detection results.
[0109] Optionally, the virus detection module 20 is further configured to determine a virus detection result of the data packet according to the timing detection result.
[0110] Optionally, the virus detection module 20 is also used to input the timing detection result into a Sigmoid activation function, calculate the relationship between the timing detection result and the probability threshold, and obtain the detection result; according to the detection result, the timing detection result is error reconstructed through an autoencoder or the similarity with any virus type is calculated; if the similarity exceeds the similarity threshold or the reconstructed error exceeds the error threshold, the data packet is determined to be an abnormal data packet, and the virus type of the abnormal data packet is obtained.
[0111] The virus detection device provided by the present application adopts the virus detection method in the above embodiment, which can solve the technical problem of low virus detection efficiency. Compared with the prior art, the beneficial effects of the virus detection device provided by the present application are the same as the beneficial effects of the virus detection method provided by the above embodiment, and other technical features in the virus detection device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0112] Based on the same inventive concept, the present application provides a virus detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the virus detection method in the above-mentioned embodiment.
[0113] like Figure 5As shown, the virus detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the virus detection device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the virus detection device to communicate wirelessly or wired with other devices to exchange data. Although the virus detection device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0114] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0115] The virus detection device provided by the present application adopts the virus detection method in the above embodiment, which can solve the technical problem of low virus detection efficiency. Compared with the prior art, the beneficial effects of the virus detection device provided by the present application are the same as the beneficial effects of the virus detection method provided by the above embodiment, and other technical features in the virus detection device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0116] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0117] Based on the same inventive concept, the present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned virus detection method when executed by a processor.
[0118] The computer program product provided by this application can solve the technical problem of low virus detection efficiency. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the virus detection method provided by the above embodiment, which will not be repeated here.
[0119] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A virus detection method, characterized in that: The virus detection method comprises: Obtain a data packet on a user terminal, and extract characteristic information of the data packet, wherein the characteristic information includes characteristic information of each blockchain node through which the data packet passes; Acquire preset characteristic information matched by the characteristic information from the marked data set, wherein the marked data set includes a normal data packet and an abnormal data packet, and the normal data packet and the abnormal data packet are marked with corresponding preset characteristic information; The virus detection result of the data packet is determined according to the preset characteristic information.
2. The virus detection method according to claim 1, characterized in that: The virus detection method further comprises: Obtain preset feature information corresponding to each blockchain node that a single data packet passes through; Determine the similarity of preset feature information of every two adjacent blockchain nodes; Determine, according to the similarity of the preset characteristic information of each two adjacent blockchain nodes, a mean similarity of the single data packet to all blockchain nodes; Determine the marking result of the single data packet according to the similarity threshold corresponding to each two adjacent blockchain nodes of the single data packet and the similarity mean; The labeled data set is constructed according to the labeling results of each of the data packets.
3. The virus detection method according to claim 2, characterized in that: The obtaining of preset characteristic information corresponding to each blockchain node through which a single data packet passes includes: Obtain the delay time, jitter time, data packet length, data packet transmission throughput rate, and transmission protocol level of each blockchain node passed by the single data packet; According to the delay time, the jitter time, the data packet length, the transmission throughput rate of the data packet and the transmission protocol level of each blockchain node, the preset characteristic information corresponding to each blockchain node is obtained.
4. The virus detection method according to claim 2, characterized in that: Determining the marking result of the single data packet according to the similarity threshold corresponding to each two adjacent blockchain nodes of the single data packet and the similarity mean includes: If the similarity thresholds corresponding to the two adjacent blockchain nodes of the single data packet are both less than or equal to the average similarity value of the single data packet, it is determined that the marking result of the single data packet is a normal data packet; Alternatively, if the similarity thresholds corresponding to two adjacent blockchain nodes of the single data packet are both greater than the average similarity of the single data packet, it is determined that the marking result of the single data packet is an abnormal data packet.
5. The virus detection method according to claim 1, characterized in that: The obtaining of a data packet on a user terminal and extracting characteristic information of the data packet includes: Obtaining a data packet on a user terminal, and performing word segmentation processing on the data in the data packet to obtain data coding and data content; Extracting features of the data encoding and the data content to obtain a text feature vector and a content feature vector; The text feature vector and the content feature vector are sorted in chronological order to obtain time series data, and the time series data is processed using a long short-term memory network model to obtain a time series feature vector; A fused feature vector is obtained based on a first weight between the text feature vector and the text feature vector, a second weight between the content feature vector and the content feature vector, and a third weight between the timing feature vector and the timing feature vector, wherein the feature information of the data packet includes the fused feature vector.
6. The virus detection method according to claim 1 or 5, characterized in that: The obtaining, from the labeled data set, preset feature information matched by the feature information comprises: Inputting the characteristic information of the data packet into a virus detection model, the virus detection model obtains preset characteristic information matching the characteristic information from the labeled data set, and determines the timing detection result of the data packet according to the preset characteristic information; wherein the virus detection model is obtained by training an initial recurrent neural network using normal data packet samples and abnormal data packet samples, and the normal data packet samples and the abnormal data packet samples are both marked with corresponding preset characteristic information and corresponding timing detection results; Determining the virus detection result of the data packet according to the preset characteristic information includes: A virus detection result of the data packet is determined according to the timing detection result.
7. The virus detection method according to claim 6, characterized in that: Determining the virus detection result of the data packet according to the timing detection result includes: The time series detection result is input into the Sigmoid activation function, and the relationship between the time series detection result and the probability threshold is calculated to obtain the detection result; According to the detection result, the time series detection result is error reconstructed through an autoencoder or similarity with any virus type is calculated; If the similarity exceeds a similarity threshold or the reconstructed error exceeds an error threshold, the data packet is determined to be an abnormal data packet, and the virus type of the abnormal data packet is obtained.
8. A virus detection device, characterized in that: The virus detection device comprises: An acquisition and extraction module, used to acquire a data packet on a user terminal and extract characteristic information of the data packet, wherein the characteristic information includes characteristic information of each blockchain node through which the data packet passes; and to acquire preset characteristic information matched by the characteristic information from a marked data set, wherein the marked data set includes a normal data packet and an abnormal data packet, and the normal data packet and the abnormal data packet are marked with corresponding preset characteristic information; The virus detection module is used to determine the virus detection result of the data packet according to the preset characteristic information.
9. A virus detection device, characterized in that: The virus detection device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the virus detection method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the virus detection method according to any one of claims 1 to 7 are implemented.