Network traffic anomaly determination method, apparatus, electronic device, and readable storage medium
By acquiring the associated prediction data of the Internet of Things system, determining the rank of the prediction vector matrix, and evaluating the data complexity by combining Fourier transform, the number of multi-head attention mechanisms in the deformer is adjusted, solving the problem of selecting the number of self-attention mechanism heads in the deformer model, and achieving accurate network traffic anomaly detection.
Patent Information
- Application Number
- CN202311126233.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-09-01
AI Technical Summary
Existing technologies struggle to accurately detect traffic anomalies in IoT networks, especially given the unresolved issue of head selection for the self-attention mechanism of deformer models under the influence of multiple factors.
By acquiring the associated prediction data of the Internet of Things system, the rank of the prediction vector matrix is determined after word embedding processing. The number of multi-head attention mechanisms of the deformer is determined based on the rank. The data complexity is evaluated by combining Fourier transform, and the number of multi-head attention mechanisms of the deformer is adjusted to train the deformer to perform anomaly detection.
It achieves accurate anomaly detection in network traffic, improves model training efficiency and detection accuracy, reduces trial and error, and is suitable for both complex and simple data scenarios.
Smart Images

Figure CN117176417B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer and Internet technology, and in particular to a method and apparatus for determining network traffic anomalies, an electronic device, and a computer-readable storage medium. Background Technology
[0002] This section is intended to provide background or context for the embodiments of this application as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.
[0003] The rapid development of the Internet of Things (IoT) has brought many potential dangers to mobile communication networks. Network attacks can originate from any IoT connection, making traffic anomaly detection in the IoT increasingly important.
[0004] Therefore, the technical problem to be solved by this application is how to accurately detect whether there are abnormalities in network traffic. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, electronic device, and computer-readable storage medium for determining abnormal network traffic, which can accurately detect abnormal traffic.
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0007] This application provides a method for determining network traffic anomalies, comprising: acquiring IoT traffic association prediction data from an IoT system; performing word embedding processing on the IoT traffic association prediction data to obtain a prediction vector matrix of IoT association data; determining the rank of the prediction vector matrix; determining the number of multi-head attention mechanisms of a deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix; setting the parameters of the deformer based on the number of multi-head attention mechanisms; and training the deformer using the IoT traffic association prediction data so as to determine whether network traffic is abnormal using the trained deformer.
[0008] In some embodiments, determining the number of multi-head attention mechanisms in the deformer based on the rank of the prediction vector matrix includes: acquiring first training data of IoT traffic and network anomaly labels corresponding to the first training data of IoT traffic; performing word embedding processing on the first training data of IoT traffic to obtain a first training vector matrix; training the deformer using the first training vector matrix and the network anomaly labels to obtain a first number of training multi-head attention mechanisms determined for the deformer based on the prediction vector matrix; predicting the rank of the first training vector matrix using a target network model to obtain a first number of predicted multi-head attention mechanisms; training the target network model using the first number of training multi-head attention mechanisms and the first number of predicted multi-head attention mechanisms; and processing the rank of the prediction vector matrix using the trained target network model to determine the number of multi-head attention mechanisms in the deformer.
[0009] In some embodiments, the IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device feature data from the IoT system; wherein, the prediction vector matrix of the IoT association data includes: a sensor prediction vector matrix, an operational prediction vector matrix, a maintenance record prediction vector matrix, an environmental condition prediction vector matrix, and a device feature prediction vector matrix; wherein, performing word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of the IoT association data includes: performing word embedding processing on the sensor data, the operational data, the maintenance record data, the environmental condition data, and the device feature data respectively to obtain... The sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix; wherein, determining the rank of the prediction vector matrix includes: determining the ranks of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix respectively; and determining the number of multi-head attention mechanisms of the deformer based on the ranks of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix.
[0010] In some embodiments, determining the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix includes: performing a Fourier transform on the prediction vector matrix to obtain a Fourier spectrum matrix; determining the spectral mean based on the Fourier spectrum matrix; and determining the number of multi-head attention mechanisms of the deformer based on the spectral mean and the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to both the rank of the prediction vector matrix and the spectral mean.
[0011] In some embodiments, determining the number of multi-head attention mechanisms of the deformer based on the mean of the spectrum and the rank of the prediction vector matrix includes: acquiring second training data of IoT traffic and network anomaly labels corresponding to the second training data of IoT association; performing word embedding processing on the second training data of IoT association to obtain a second training vector matrix; training the deformer using the second training vector matrix and the network anomaly labels to obtain a second number of training multi-head attention mechanisms determined for the deformer based on the prediction vector matrix; performing prediction processing on the rank of the second training vector matrix and the mean of the spectrum using a target network model to obtain a second number of predicted multi-head attention mechanisms; training the target network model using the second number of training multi-head attention mechanisms and the second number of predicted multi-head attention mechanisms; and processing the rank of the prediction vector matrix and the mean of the spectrum using the trained target network model to determine the number of multi-head attention mechanisms of the deformer.
[0012] In some embodiments, the IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device feature data from the IoT system; wherein, the prediction vector matrix of the IoT association data includes: a sensor prediction vector matrix, an operational prediction vector matrix, a maintenance record prediction vector matrix, an environmental condition prediction vector matrix, and a device feature prediction vector matrix; wherein, the Fourier spectrum matrix includes: a sensor word spectrum matrix, an operational spectrum matrix, a maintenance record spectrum matrix, an environmental condition spectrum matrix, and a device feature spectrum matrix; wherein, performing word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of the IoT association data includes: performing word embedding processing on the sensor data, the operational data, the maintenance record data, the environmental condition data, and the device feature data respectively to obtain the sensor prediction vector matrix, the operational prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix; wherein, determining the rank of the prediction vector matrix includes: determining the rank of the sensor prediction vector matrix, the operational prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix. The rank; wherein, performing Fourier transform processing on the prediction vector matrix to obtain the Fourier spectrum matrix includes: performing Fourier transform processing on the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix to obtain the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment feature spectrum matrix, respectively; wherein, determining the spectral mean based on the Fourier spectrum matrix includes: determining the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, and the environmental condition spectrum matrix, respectively. The mean spectral values of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the device feature spectrum matrix; wherein, determining the number of multi-head attention mechanisms of the deformer based on the mean spectral values and the rank of the prediction vector matrix includes: determining the number of multi-head attention mechanisms of the deformer based on the mean spectral values of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the device feature spectrum matrix, and the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix.
[0013] In some embodiments, determining the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix includes: obtaining a numerical fitting relationship between the number of multi-head attention mechanisms of the deformer and the rank of the matrix; and performing interpolation processing on the numerical fitting relationship to determine the number of multi-head attention mechanisms corresponding to the rank of the prediction vector matrix as the number of multi-head attention mechanisms of the deformer.
[0014] This application provides a network traffic anomaly determination device, including: an association prediction data acquisition module, a word embedding module, a rank determination module, an attention mechanism quantity determination module, a parameter setting module, and a training module.
[0015] The system includes the following modules: a correlation prediction data acquisition module for acquiring IoT traffic correlation prediction data from an IoT system; a word embedding module for performing word embedding processing on the IoT traffic correlation prediction data to obtain a prediction vector matrix of the IoT correlation data; a rank determination module for determining the rank of the prediction vector matrix; a number of attention mechanisms determination module for determining the number of multi-head attention mechanisms in the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms in the deformer is proportional to the rank of the prediction vector matrix; a parameter setting module for setting the parameters of the deformer based on the number of multi-head attention mechanisms; and a training module for training the deformer using the IoT traffic correlation prediction data to determine whether network traffic is abnormal.
[0016] This application provides an electronic device comprising: a memory and a processor; the memory is used to store computer program instructions; the processor invokes the computer program instructions stored in the memory to implement the network traffic anomaly determination method described above.
[0017] This application provides a computer-readable storage medium storing computer program instructions to implement the network traffic anomaly determination method as described in any of the above claims.
[0018] This application provides a computer program product or computer program that includes computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and the processor executes the computer program instructions to implement the aforementioned method for determining abnormal network traffic.
[0019] The network traffic anomaly determination method, apparatus, electronic device, and computer-readable storage medium provided in this application embodiment can set the parameters of the deformer by using the rank of the prediction vector matrix corresponding to the IoT traffic association prediction data, so that the deformer can accurately detect whether the predicted network traffic is abnormal.
[0020] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0022] Figure 1 A schematic diagram of a scenario that can be applied to the network traffic anomaly determination method or device in the embodiments of this application is shown.
[0023] Figure 2 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0024] Figure 3 This is a schematic diagram illustrating word vector weight matrix training according to an exemplary embodiment.
[0025] Figure 4 This is a schematic diagram illustrating a skip-gram model according to an exemplary embodiment.
[0026] Figure 5 This is a schematic diagram of word encoding according to an exemplary embodiment.
[0027] Figure 6 This is a schematic diagram illustrating deformer parameter guidance according to an exemplary embodiment.
[0028] Figure 7 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0029] Figure 8 This is a schematic diagram of the structure corresponding to a network model training method according to an exemplary embodiment.
[0030] Figure 9 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0031] Figure 10This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0032] Figure 11 This is a schematic diagram illustrating a network model training method according to an exemplary embodiment.
[0033] Figure 12 This is a flowchart illustrating a network model training method according to an exemplary embodiment.
[0034] Figure 13 This is a schematic diagram illustrating a network model training method according to an exemplary embodiment.
[0035] Figure 14 This is a flowchart illustrating a network model training method according to an exemplary embodiment.
[0036] Figure 15 This is an architectural diagram of a transformer model according to an exemplary embodiment.
[0037] Figure 16 This is a block diagram illustrating a network traffic anomaly determination device according to an exemplary embodiment.
[0038] Figure 17 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of this application is shown. Detailed Implementation
[0039] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this application will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0040] Those skilled in the art will understand that embodiments of this application can be a system, apparatus, device, method, or computer program product. Therefore, this application can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0041] The features, structures, or characteristics described in this application can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of the embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced with one or more specific details omitted, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0042] The accompanying drawings are merely illustrative of this application; the same reference numerals denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0043] The flowchart shown in the accompanying drawings is merely illustrative and does not necessarily include all content and steps, nor does it require execution in the described order. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0044] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences; the terms "contains," "includes," and "has" are used to indicate an open-ended inclusion and mean that additional elements / components / etc. may exist besides the listed elements / components / etc.
[0045] To better understand the above-mentioned objectives, features and advantages of this application, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0046] With the widespread adoption of mobile devices and the rapid development of mobile network technology, network security issues are becoming increasingly serious. Anomaly traffic detection technology has emerged in this context, and it holds significant importance in the field of mobile communications, especially in industrial sectors with diverse connection types. It can effectively identify and prevent potential network attacks, providing security for mobile communication networks. Applications of anomaly traffic detection include, but are not limited to, the following: 1. Intrusion Detection and Prevention: Detecting potential network attacks, such as Distributed Denial of Service (DDoS) attacks, botnets, and malware propagation. 2. Data Breach Prevention: Detecting abnormal behaviors that may lead to data breaches, such as stealing user information or attacking internal systems. 3. Load Balancing and Optimization: Helping network operators detect and resolve traffic congestion problems. 4. Network Behavior Analysis: Conducting in-depth analysis of user behavior to identify malicious and fraudulent activities.
[0047] In related technologies, many classic neural networks have been applied to anomaly traffic detection: Convolutional Neural Networks (CNNs) can automatically learn local features of network traffic, while Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs) can capture temporal dependencies in network traffic sequence data; both can effectively detect anomaly traffic. Furthermore, unsupervised learning methods such as autoencoders (AEs) and variational autoencoders (VAEs) have also achieved breakthroughs in anomaly traffic detection. These methods can learn normal patterns of network traffic even with unlabeled data and detect anomaly traffic that deviates significantly from the normal pattern. Researchers have then proposed many improvement methods to enhance the performance of anomaly traffic detection. For example, using Generative Adversarial Networks (GANs) to generate more training samples improves the model's generalization ability; using techniques such as multi-task learning and transfer learning to transfer knowledge learned in certain scenarios to other scenarios; and employing ensemble learning and multi-model fusion methods to combine the predictions of multiple models, improving the accuracy and robustness of detection.
[0048] In some embodiments, a transformer with a self-attention mechanism can be used to identify anomalous traffic. However, in practice, the challenge of choosing the number of heads for the transformer's self-attention mechanism is a key research topic, as the performance of the transformer model often depends on the head selection for the attention mechanism. However, how to choose the most suitable head number is not clear, as it is influenced by various factors such as model complexity, data size, and computational resources.
[0049] In some embodiments, the more heads and parameters a transformer has, the more complex relationships it can fit. Simply put, a single head in a self-attention mechanism can view the data from one perspective, while multiple heads can view the data from multiple perspectives. In other words, the more complex the data, the more heads the self-attention mechanism needs.
[0050] To better understand data complexity and thus better train the network model, enabling the trained model to detect traffic anomalies, this application provides a method to predict the number of multi-head attention mechanisms in a Transformer in advance. This application uses the rank sum and / or Fourier transform of the matrix to evaluate data complexity, infers the possible size of the data space, and adjusts the number of multi-head attention mechanisms in the Transformer based on this prediction.
[0051] The exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0052] Figure 1 A schematic diagram of a scenario that can be applied to the network traffic anomaly determination method or device in the embodiments of this application is shown.
[0053] Please refer to Figure 1 The diagram illustrates an implementation environment provided by an exemplary embodiment of this application.
[0054] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0055] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, desktop computers, wearable devices, virtual reality devices, smart home devices, etc.
[0056] Server 105 can be a server that provides various services, such as a backend management server that supports the devices operated by users using terminal devices 101, 102, and 103. The backend management server can analyze and process received requests and other data, and feed the processing results back to the terminal devices.
[0057] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This application does not impose any restrictions on this.
[0058] Server 105 may, for example, acquire IoT traffic association prediction data from an IoT system; server 105 may, for example, perform word embedding processing on the IoT traffic association prediction data to obtain a prediction vector matrix of IoT association data; server 105 may, for example, determine the rank of the prediction vector matrix; server 105 may, for example, determine the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix; server 105 may, for example, set the parameters of the deformer based on the number of multi-head attention mechanisms; server 105 may, for example, train the deformer using the IoT traffic association prediction data so that the trained deformer can be used to determine whether network traffic is abnormal.
[0059] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Server 105 can be a single physical server or a combination of multiple servers. Depending on actual needs, it can have any number of terminal devices, networks, and servers.
[0060] Under the above system architecture, this application provides a method for determining network traffic anomalies, which can be executed by any electronic device with computing capabilities.
[0061] Figure 2 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment. The method provided in this application embodiment can be executed by any electronic device with computing power; for example, the method can be implemented by the aforementioned... Figure 1 The execution can be performed by a server or terminal device in the embodiments, or it can be performed by both a server and a terminal device. In the following embodiments, the server is used as the execution subject for illustration, but this application is not limited to this.
[0062] Reference Figure 2 The network traffic anomaly determination method provided in this application embodiment may include the following steps.
[0063] Step S202: Obtain IoT traffic association prediction data from the IoT system.
[0064] The aforementioned IoT traffic-related data can refer to data related to the identification of IoT traffic anomalies. Examples include sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data from the IoT system.
[0065] The following section will explain the various IoT traffic correlation prediction data.
[0066] Sensor data: Sensors can collect real-time data on various equipment parameters, such as temperature, pressure, vibration, current, and voltage. This data can be used to monitor the operating status and performance indicators of the equipment, as well as detect any abnormalities.
[0067] Operational data: This data includes equipment operating time, work cycle, speed, rotational speed, etc. Operational data can provide basic information about the equipment's operation, providing a basis for prediction and analysis.
[0068] Maintenance records: Maintenance records include the equipment's maintenance history, maintenance activities, and repair records. This data can be used to analyze the equipment's maintenance needs and maintenance effectiveness in order to optimize maintenance strategies.
[0069] Environmental conditions: Environmental condition data includes environmental parameters of the environment in which the equipment is located, such as temperature, humidity, and air pressure. Environmental conditions have a certain impact on the operating status and performance of the equipment; therefore, monitoring and recording environmental data is also important for maintenance decisions.
[0070] Equipment characteristic data: This data includes equipment specifications, model, production date, parts information, etc. Equipment characteristic data can be used to build benchmark models of the equipment and conduct comparative analyses to determine the equipment's health status.
[0071] Step S204: Perform word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of IoT association data.
[0072] In this embodiment, the text-based network traffic association prediction data first needs to be preprocessed to allow the deep learning model, the transformer, to better understand the data. One of the key steps in preprocessing is word embedding. Word embedding is a technique that maps words or phrases in text to a continuous vector space, capturing the semantic and grammatical relationships between words. In this application, a word embedding algorithm called Word2Vec can be used.
[0073] Word2Vec is a popular algorithm for learning word embeddings. It's a technique for representing text data, converting each word into a relatively low-dimensional continuous vector that captures the semantic and syntactic relationships between words. Specifically, it uses an embedding space where semantically similar words are close together. For example, apples and pears are both fruits, so their word embeddings will be similar, while semantically unrelated words, such as apples and bricks, will have larger numerical differences. Its implementation involves training a neural network, then using the hidden layers of this network to calculate the probability distribution of the input words, selecting the words with the highest probabilities for output. Since this process doesn't require manual labeling, it's an unsupervised learning method. We'll first introduce the training method for the weight matrix, as illustrated in the diagram below. Figure 3 As shown. Among them, Figure 3 This is a schematic diagram illustrating word vector weight matrix training according to an exemplary embodiment.
[0074] In some embodiments, by feeding IoT traffic association prediction data into... Figure 3 After multiple iterations, the input and output ends of the neural network can be trained to produce an embedding matrix with rich weight information.
[0075] Furthermore, depending on the training method and output, the Word2Vec algorithm has two basic forms: the Skip-Gram model and the Continuous Bag of Words (CBOW) model. In some embodiments, the Skip-Gram model can be used, where each input word is used to predict the words around it. That is, given a word, it predicts the context, as illustrated in the diagram below. Figure 4 As shown. Figure 4 This is a schematic diagram illustrating a skip-gram model according to an exemplary embodiment.
[0076] Here, w(t) represents the current input word, w(t-2) represents the second word before this word, w(t-1) represents the word before it, w(t+1) represents the word after it, and so on. The number of words to predict each time is determined by the window size, which is 2 (representing the prediction of 2 words before and after). The window size can be set according to your needs.
[0077] After determining the training method, the weight matrix (i.e., the trained neural network) calculates the probability of the input word and predicts the word with the highest probability in its context. This prediction is then used as the output, and the difference is calculated by comparing it with the actual context. The training is then performed using the backpropagation algorithm until the weight information is fully trained. At this point, the weight matrix completes the conversion bridge between text data and numerical vectors in our patented technology. Any new IoT traffic data can be used to calculate its word vector representation.
[0078] In some embodiments, positional encoding is also introduced during word embedding processing. Positional encoding can be interpreted as follows: The transformer is a deep learning model based on a self-attention mechanism. It eliminates the traditional RNN architecture used in seq2seq networks for text processing, and therefore lacks the ability to understand the positional relationships between words. However, when processing text data, understanding the positional information of words in a sentence is crucial for capturing grammatical structure and semantic relationships. To introduce positional information into the transformer, this application requires the use of positional encoding. The role of positional encoding is to generate a unique vector representation for each position, so as to preserve word order information in subsequent self-attention calculations. The dimension of the positional encoding vector is the same as the dimension of the word embedding vector, so they can be directly added. The main significance of positional encoding in the transformer model is to introduce the positional information of words in the sentence. After being added to the word embedding vector, it forms a new vector containing positional information, which will be fed into subsequent neural network layers for processing (see details for reference). Figure 5 (Example shown).
[0079] Step S206: Determine the rank of the prediction vector matrix.
[0080] The rank of a matrix is defined as the maximum number of linearly independent rows or columns in the matrix. Linearly independent rows or columns are those that cannot be represented by a linear combination of other rows or columns. This means that the rank describes the amount of independent information in the matrix.
[0081] Physically, rank can be interpreted as the degrees of freedom or dimension of the transformation described by a matrix. More specifically, if the rank of a matrix is r, then the linear transformation described by this matrix will operate in an r-dimensional space. This means that the transformation can map at most one r-dimensional vector space to another r-dimensional vector space.
[0082] For example, consider a 2x2 matrix describing a rotational transformation on a two-dimensional plane. If the matrix has a rank of 2, it means the transformation can completely preserve all information on the two-dimensional plane, including length, angle, and shape. However, if the matrix has a rank of 1, it means the transformation can only map all points on the two-dimensional plane to a straight line, losing one dimension of the plane.
[0083] In summary, the rank of a matrix provides information about the dimension of the matrix describing a linear transformation. Physically, it can be interpreted as the degrees of freedom of the transformation or the dimension of the action space.
[0084] In summary, the rank of a matrix can describe the complexity of the original data, such as the complexity of IoT traffic correlation prediction data. Generally speaking, the more complex the IoT traffic correlation prediction data and the lower the data correlation, the larger the rank of the prediction vector matrix; conversely, the simpler the IoT traffic correlation prediction data and the higher the data correlation, the smaller the rank of the prediction vector matrix.
[0085] Step S208: Determine the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix.
[0086] like Figure 6 As shown, the number of multi-head attention mechanisms in the deformer can be determined based on the rank of the prediction vector matrix.
[0087] Generally speaking, the more complex the IoT traffic association prediction data and the lower the data correlation, the larger the rank of the prediction vector matrix, and the more attention mechanisms the deformer trained on the IoT traffic association prediction data will have. Conversely, the simpler the IoT traffic association prediction data and the higher the data correlation, the smaller the rank of the prediction vector matrix, and the fewer attention mechanisms the deformer trained on the IoT traffic association prediction data will have.
[0088] In some embodiments, the rank of the prediction vector matrix can be processed by a trained network model to determine the number of multi-head attention mechanisms in the deformer.
[0089] In some embodiments, the number of multi-head attention mechanisms in the deformer can also be determined by a numerical fitting relationship between the number of multi-head attention mechanisms and the rank of the matrix. Specifically, this may include the following steps: obtaining a numerical fitting relationship between the number of multi-head attention mechanisms in the deformer and the rank of the matrix; and performing interpolation on the numerical fitting relationship to determine the number of multi-head attention mechanisms corresponding to the rank of the prediction vector matrix, which is then used as the number of multi-head attention mechanisms in the deformer. The numerical fitting relationship between the number of multi-head attention mechanisms in the deformer and the rank of the matrix can be obtained by fitting prior data, which will not be elaborated further in this embodiment.
[0090] Step S210: Set the parameters of the deformer according to the number of multi-head attention mechanisms.
[0091] Step S212: Train the deformer using IoT traffic association prediction data so that the trained deformer can determine whether network traffic is abnormal.
[0092] In some embodiments, Fourier transform can be used to calculate the embedding matrix of IoT traffic association prediction data (such as sensor data, operational data, maintenance records, environmental conditions, and device characteristic data words), and its spectrogram can be used to determine the size of the transformer multi-head attention mechanism. Subsequently, this set multi-head attention mechanism is used to perform anomaly detection and classification of IoT traffic based on a large model neural network built on the transformer.
[0093] In some embodiments, the rank of the prediction vector matrix and the spectrogram of the prediction vector matrix can be combined to determine the number of multi-head attention mechanisms in the deformer. Specific details can be found in the following embodiments, which will not be elaborated upon here.
[0094] The methods described above can effectively evaluate data on traditional machine learning models (linear regression, decision trees, gradient boosting trees, etc.), and the simpler the model, the stronger the reference value of the data evaluation. This makes them more suitable for industry application scenarios where small data and small models are more commonly used. These methods can determine the quality of the data, thus allowing for an early assessment of the difficulty of model construction and training. This enables trade-offs to be made based on project benefits, avoiding a large amount of trial and error. Furthermore, these methods can assist in predicting the network perception experience of industry customers by classifying the spectrum of the data, providing assistance for network fault diagnosis and intelligent operation and maintenance of industry networks.
[0095] In summary, the above methods propose a novel approach to constructing large transformer models by incorporating data feature extraction into the construction of transformer models; transforming the network traffic classification and detection problem into a text processing problem and connecting it with the most popular transformer architecture through word embedding; and using Fourier transform and matrix rank to guide the construction of large transformer models to evaluate the number of multi-head attention mechanisms required for the problem.
[0096] Figure 7 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0097] refer to Figure 7 The method for determining network traffic anomalies in the above network model may include the following steps.
[0098] Step S702: Obtain the first training data of IoT traffic and the network anomaly label corresponding to the first training data of IoT traffic.
[0099] The aforementioned first training data for IoT traffic may also include sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data, etc., and this application does not impose any restrictions on this.
[0100] The aforementioned network anomaly labels can be used to indicate whether the network traffic corresponding to the first training data of IoT traffic is abnormal, and this application does not limit the specific representation format.
[0101] Step S704: Perform word embedding processing on the first training data of IoT traffic to obtain the first training vector matrix.
[0102] Step S706: The deformer is trained using the first training vector matrix and network anomaly labels to obtain the first number of multi-head attention mechanisms for training the deformer based on the prediction vector matrix.
[0103] In some embodiments, the deformer can be parameter-tuned during training to determine the number of first training multi-head attention mechanisms corresponding to the deformer when training the deformer using the prediction vector matrix. The number of first training multi-head attention mechanisms may be the number of optimal multi-head attention mechanisms corresponding to the deformer under the prediction vector matrix.
[0104] Step S708: The rank of the first training vector matrix is predicted by the target network model to obtain the first number of predicted multi-head attention mechanisms.
[0105] In some embodiments, the rank of the first training vector matrix can be input into the target network model to predict the first number of predicted multi-head attention mechanisms.
[0106] like Figure 8 As shown, the rank (e.g., X, Y, ...) of the first training vector matrix can be input into the target network model (e.g., ReLU) to predict the number of first prediction multi-head attention mechanisms.
[0107] Step S710: Train the target network model using the first number of trained multi-head attention mechanisms and the first number of predicted multi-head attention mechanisms.
[0108] In some embodiments, the first number of trained multi-head attention mechanisms is the optimal number of multi-head attention mechanisms corresponding to the deformer under the first training vector matrix, and the first number of predicted multi-head attention mechanisms is the number of multi-head attention mechanisms of the deformer determined by the target network model for the first training vector matrix.
[0109] In some embodiments, the difference between the first predicted number of multi-head attention mechanisms and the first trained number of multi-head attention mechanisms can be determined to determine a loss function value, and then the target network model can be trained using the loss function value.
[0110] Step S712: The rank of the prediction vector matrix is processed by the trained target network model to determine the number of multi-head attention mechanisms in the deformer.
[0111] In some embodiments, the rank of the prediction vector matrix can be processed by the trained target network model to accurately determine the number of multi-head attention mechanisms in the deformer.
[0112] In some embodiments, IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data in the IoT system; wherein, the prediction vector matrix of IoT association data includes: sensor prediction vector matrix, operational prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and device characteristic prediction vector matrix.
[0113] Figure 9 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0114] refer to Figure 9 The above-mentioned method for determining abnormal network traffic may include the following steps.
[0115] Step S902: Obtain sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data from the Internet of Things system.
[0116] Step S904: Perform word embedding processing on sensor data, operation data, maintenance record data, environmental condition data, and equipment feature data respectively to obtain sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and equipment feature prediction vector matrix.
[0117] Step S906: Determine the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix, respectively.
[0118] Step S908: Determine the number of multi-head attention mechanisms of the deformer based on the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
[0119] In some embodiments, the rank summation, mean, or median of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix can be calculated, and then the obtained numbers can be input into the target network model to determine the number of multi-head attention mechanisms in the deformer.
[0120] In some embodiments, the target network model described above can be a ReLU activation function capable of extracting nonlinear relationships.
[0121] In some embodiments, the above process may be specifically described as follows:
[0122] Define a function f(x) = ReLU(aX + bT + cZ + dH + eI) and use the ReLU activation function for the output (in order to capture nonlinear information) to determine the number of multi-head attention mechanisms of the deformer, where X is sensor data, Y is operational data, Z is maintenance record, H is environmental conditions, and I is equipment characteristic data.
[0123] Specifically, after obtaining the above five types of text data, the following steps can be taken.
[0124] 1. Use the word2vec algorithm introduced earlier to perform word embeddings to obtain different word embedding matrices.
[0125] 2. Different ranks are obtained by embedding different words into the matrix.
[0126] 3. Record the average value of each of the five ranks and record it as a parameter.
[0127] 4. Using these five ranks as input to the proposed ReLU function, the ReLU function calculates the result, obtaining y^*. Here, X corresponds to the rank of the first type of text, Y corresponds to the rank of the second type of text, and so on.
[0128] 5. Through experiments, we tested the optimal parameters that these data should possess: the number of multi-head attention points, y^.
[0129] 6. Subtract y^* from y^~ to obtain the difference, and use this difference for backpropagation to update the parameters a, b, c, d, e, etc. of the ReLU function.
[0130] 7. Repeat the above steps until the difference between the calculation result of the ReLU function and the actual optimal transformer multi-head number is minimized or a certain requirement is met.
[0131] Step S910: Set the parameters of the deformer according to the number of multi-head attention mechanisms.
[0132] Step S912: Train the deformer using IoT traffic association prediction data so that the trained deformer can determine whether network traffic is abnormal.
[0133] Step S914: Train the deformer using IoT traffic association prediction data so that the trained deformer can determine whether network traffic is abnormal.
[0134] Figure 10 This is a flowchart illustrating a method for determining network traffic anomalies according to an exemplary embodiment.
[0135] refer to Figure 10 The above-mentioned method for determining abnormal network traffic may include the following steps.
[0136] Step S1002: Obtain IoT traffic association prediction data from the IoT system.
[0137] Step S1004: Perform word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of IoT association data.
[0138] Step S1006: Determine the rank of the prediction vector matrix.
[0139] Step S1008: Perform Fourier transform on the prediction vector matrix to obtain the Fourier spectrum matrix.
[0140] Introduction to Fourier Transform: The Fourier Transform is a mathematical transform used to convert time-domain signals into frequency-domain signals. It is an important tool in signal processing, and its mathematical formula is as follows:
[0141]
[0142] Where f(t) represents the time-domain signal, F(u) represents the frequency-domain signal, u is the frequency, i is the imaginary unit, and t is the time. The Fourier transform obtains the spectral distribution of the signal, that is, the intensity distribution of the signal at each frequency, by analyzing the amplitude and phase of each frequency component.
[0143] Step S1010: Determine the mean of the spectrum based on the Fourier spectrum matrix.
[0144] Step S1012: Determine the number of multi-head attention mechanisms of the deformer based on the mean of the spectrum and the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to both the rank of the prediction vector matrix and the mean of the spectrum.
[0145] The trained target network function is used to process the spectral mean and rank of the prediction vector matrix to determine the number of multi-head attention mechanisms in the deformer.
[0146] Step S1014: Set the parameters of the deformer according to the number of multi-head attention mechanisms.
[0147] Step S1016: Train the deformer using IoT traffic association prediction data so that the trained deformer can determine whether network traffic is abnormal.
[0148] The above embodiments not only use the rank of the word embedding matrix corresponding to sensor data, operational data, maintenance records, environmental conditions, and equipment feature data to measure the data complexity of these data, but also use Fourier transform to calculate the word embedding matrix of these data, and determine the size of the transformer multi-head attention mechanism number through its rank and spectrogram. Subsequently, this pre-defined multi-head attention mechanism number is used to build a large-scale neural network based on the transformer for IoT traffic detection and classification.
[0149] The above embodiments assess data quality by using the rank and Fourier transform of the matrix to anticipate model data requirements and provide a reference for whether further data collection is necessary.
[0150] Figure 12 This is a flowchart illustrating a network model training method according to an exemplary embodiment.
[0151] Step S1202: Obtain the network anomaly labels corresponding to the second training data of IoT traffic and the second training data of IoT association.
[0152] Step S1204: Perform word embedding processing on the IoT-related second training data to obtain the second training vector matrix.
[0153] The aforementioned second training data for IoT traffic may also include sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data, etc., and this application does not impose any restrictions on this.
[0154] Step S1206: The deformer is trained using the second training vector matrix and network anomaly labels to obtain the number of second training multi-head attention mechanisms determined for the deformer based on the prediction vector matrix.
[0155] Step S1208: The rank and spectral mean of the second training vector matrix are predicted using the target network model to obtain the number of second predicted multi-head attention mechanisms.
[0156] Step S1210: Train the target network model by using the second number of trained multi-head attention mechanisms and the second number of predicted multi-head attention mechanisms.
[0157] The target network model described above can be a ReLU activation function model, and the training process described above can be specifically as follows: Figure 13 The process is shown.
[0158] Define a function f(x) = ReLU(aX + bT + cZ + dH + eI), where X is sensor data, Y is operational data, Z is maintenance records, H is environmental conditions, and I is equipment characteristic data. The ReLU activation function is used for the output (to capture nonlinear information). After obtaining five types of text data, the following steps can be performed:
[0159] 1. Use the word2vec algorithm introduced earlier to perform word embeddings to obtain different word embedding matrices.
[0160] 2. Perform Fourier transform and rank calculation on different word embedding matrices to obtain different spectrograms and ranks.
[0161] 3. Record the average frequency of the rank and spectrogram corresponding to the 5 types of text as parameters.
[0162] 4. Using the average frequency corresponding to the rank sum and spectrogram of these 5 types of text as input to the proposed ReLU function, the ReLU function calculates the result, obtaining y^* (e.g., Figure 12 As shown, the parameters corresponding to the rank sum spectrograms of the above five types of text can be processed by ReLU to calculate the multi-head attention mechanism number y^*) of the deformer. Here, X corresponds to the sum of the rank sum and average frequency of the first type of text, Y corresponds to the sum of the rank sum and average frequency of the second type of text, and so on.
[0163] 5. Through experiments, we tested the optimal parameters that these data should possess: the number of multi-head attention points, y^.
[0164] 6. Subtract y^* from y^~ to obtain the difference, and use this difference for backpropagation to update the parameters a, b, c, d, e, etc. of the ReLU function.
[0165] 7. Repeat the above steps until the difference between the calculation result of the ReLU function and the actual optimal transformer multi-head number is minimized or a certain requirement is met.
[0166] Step S1212: The rank and spectral mean of the prediction vector matrix are processed by the trained target network model to determine the number of multi-head attention mechanisms in the deformer.
[0167] In some embodiments, IoT traffic association prediction data may include sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data from the IoT system.
[0168] In some embodiments, the prediction vector matrix for IoT-related data may include: a sensor prediction vector matrix, an operation prediction vector matrix, a maintenance record prediction vector matrix, an environmental condition prediction vector matrix, and a device feature prediction vector matrix.
[0169] In some embodiments, the Fourier spectrum matrix may include: a sensor word spectrum matrix, an operational spectrum matrix, a maintenance record spectrum matrix, an environmental condition spectrum matrix, and a device characteristic spectrum matrix.
[0170] Figure 14 This is a flowchart illustrating a network model training method according to an exemplary embodiment.
[0171] refer to Figure 14 The above-mentioned network model training method may include the following steps.
[0172] Step S1402: Obtain sensor data, operation data, maintenance record data, environmental condition data, and device characteristic data from the Internet of Things system.
[0173] Step S1404: Perform word embedding processing on sensor data, operation data, maintenance record data, environmental condition data, and equipment feature data respectively to obtain sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and equipment feature prediction vector matrix.
[0174] Step S1406: Determine the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
[0175] Step S1408: Perform Fourier transform processing on the sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and equipment feature prediction vector matrix to obtain the sensor word spectrum matrix, operation spectrum matrix, maintenance record spectrum matrix, environmental condition spectrum matrix, and equipment feature spectrum matrix, respectively.
[0176] Step S1410: Determine the mean spectral values of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment characteristic spectrum matrix, respectively.
[0177] Step S1412: Determine the number of multi-head attention mechanisms of the deformer based on the mean of the spectrum of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment feature spectrum matrix, as well as the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
[0178] Step S1414: Set the parameters of the deformer according to the number of multi-head attention mechanisms.
[0179] Step S1416: Train the deformer using IoT traffic association prediction data so that the trained deformer can determine whether network traffic is abnormal.
[0180] The above embodiments can simultaneously learn the local features and long-term dependencies of abnormal traffic. This application incorporates the rank of the matrix and the information from the Fourier transform into the transformer construction process, integrating the characteristics of the traffic data itself—the spectral information of the data and the rank information of the matrix—into the training and construction of the transformer multi-head attention mechanism.
[0181] Figure 15 This is an architectural diagram of a transformer model according to an exemplary embodiment.
[0182] In this application, the self-attention mechanism in the transformer is the core of the model, responsible for capturing long-distance dependencies in the input sequence. The self-attention mechanism generates a weighted representation by calculating the relationship between each word and other words in the input sequence, which is then used for subsequent network layer processing. Generally, transformers use multiple self-attention mechanisms (e.g., eight), therefore this patent employs a multi-head attention mechanism.
[0183] The self-attention mechanism is derived from three weight matrices: Q (Query), K (Key), and V (Value), representing the query, key, and value, respectively. Specifically, the input sequence is first transformed into Q, K, and V vectors through a linear layer. Then, the dot product of Q and K is calculated to measure the contribution of each word in the input sequence to the current word. This dot product result is further normalized using a softmax function to obtain the final attention weights.
[0184] Feedforward neural networks typically consist of two linear layers (fully connected layers) and an activation function. The activation function used can capture the non-linear characteristics of the input data (such as ReLU). The Transformer also uses skip connections (also known as residual connections) and layer normalization to optimize network performance.
[0185] Skip connections directly add the input to the output of a feedforward neural network, thus achieving a "skip" transfer of the original input. This structure helps alleviate the vanishing gradient problem, enabling the model to be trained more effectively at deeper levels.
[0186] In addition, by normalizing the output of each layer, this application can ensure smoother information transfer between different layers in the network and avoid gradient explosion or vanishing problems.
[0187] It should be particularly noted that the steps in each embodiment of the above-described method for determining network traffic anomalies can be overlapped, substituted, added, or deleted. Therefore, these reasonable permutations and combinations should also fall within the protection scope of this application, and the protection scope of this application should not be limited to the embodiments.
[0188] Based on the same inventive concept, this application also provides a network traffic anomaly determination device, as shown in the following embodiment. Since the principle of this device embodiment in solving the problem is similar to that of the above method embodiment, the implementation of this device embodiment can refer to the implementation of the above method embodiment, and repeated details will not be described again.
[0189] Figure 16 This is a block diagram illustrating a network traffic anomaly determination apparatus according to an exemplary embodiment. (Refer to...) Figure 16 The network traffic anomaly determination device 1600 provided in this application embodiment may include: an association prediction data acquisition module 1601, a word embedding module 1602, a rank determination module 1603, an attention mechanism quantity determination module 1604, a parameter setting module 1605, and a training module 1606.
[0190] The module 1601 for acquiring IoT traffic association prediction data in the IoT system is used to acquire IoT traffic association prediction data. The word embedding module 1602 is used to perform word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of the IoT association data. The rank determination module 1603 is used to determine the rank of the prediction vector matrix. The attention mechanism number determination module 1604 is used to determine the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix. The parameter setting module 1605 is used to set the parameters of the deformer based on the number of multi-head attention mechanisms. The training module 1606 is used to train the deformer using IoT traffic association prediction data so that the trained deformer can be used to judge whether the network traffic is abnormal.
[0191] It should be noted that the aforementioned association prediction data acquisition module 1601, word embedding module 1602, rank determination module 1603, attention mechanism quantity determination module 1604, parameter setting module 1605, and training module 1606 correspond to S202 to S212 in the method embodiment. The examples and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content claimed in the above method embodiment. It should also be noted that these modules, as part of the apparatus, can be executed in a computer system such as a set of computer-executable instructions.
[0192] In some embodiments, the attention mechanism number determination module 1604 may include: a first training data acquisition submodule, a first training vector matrix acquisition submodule, a first training multi-head attention mechanism number determination submodule, a first prediction multi-head attention mechanism number determination submodule, a first training submodule, and a first rank determination submodule.
[0193] The first training data acquisition submodule can be used to acquire the first training data of IoT traffic and the network anomaly labels corresponding to the first training data of IoT traffic; the first training vector matrix acquisition submodule can be used to perform word embedding processing on the first training data of IoT traffic to obtain the first training vector matrix; the first training multi-head attention mechanism quantity determination submodule can be used to train the deformer through the first training vector matrix and the network anomaly labels to obtain the first training multi-head attention mechanism quantity determined for the deformer based on the prediction vector matrix; the first predicted multi-head attention mechanism quantity determination submodule can be used to predict the rank of the first training vector matrix through the target network model to obtain the first predicted multi-head attention mechanism quantity; the first training submodule can be used to train the target network model through the first training multi-head attention mechanism quantity and the first predicted multi-head attention mechanism quantity; the first rank determination submodule can be used to process the rank of the prediction vector matrix through the trained target network model to determine the multi-head attention mechanism quantity of the deformer.
[0194] In some embodiments, IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data in the IoT system; wherein, the prediction vector matrix of the IoT association data includes: sensor prediction vector matrix, operational prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and device characteristic prediction vector matrix;
[0195] In some embodiments, the word embedding module 1602 may include a device feature prediction vector matrix determination submodule.
[0196] The device feature prediction vector matrix determination submodule can be used to perform word embedding processing on sensor data, operation data, maintenance record data, environmental condition data and device feature data respectively to obtain sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction matrix, environmental condition prediction matrix and device feature prediction vector matrix.
[0197] In some embodiments, the rank determination module 1603 may include: a rank determination submodule for the device feature prediction vector matrix and a multi-head attention mechanism quantity determination submodule.
[0198] The rank determination submodule for the equipment feature prediction vector matrix can be used to determine the ranks of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix, respectively. The multi-head attention mechanism number determination submodule can be used to determine the number of multi-head attention mechanisms of the deformer based on the ranks of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
[0199] In some embodiments, the attention mechanism quantity determination module 1604 may include: a Fourier spectrum matrix determination submodule, a spectrum mean determination submodule, and a deformer parameter setting submodule.
[0200] The Fourier spectrum matrix determination submodule can be used to perform Fourier transform processing on the prediction vector matrix to obtain the Fourier spectrum matrix; the spectrum mean determination submodule can be used to determine the spectrum mean based on the Fourier spectrum matrix; the deformer parameter setting submodule can be used to determine the number of multi-head attention mechanisms of the deformer based on the spectrum mean and the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to both the rank of the prediction vector matrix and the spectrum mean.
[0201] In some embodiments, the deformer parameter setting submodule may include: a second training data acquisition unit, a second training vector matrix determination unit, a training parameter acquisition unit, a second prediction multi-head attention mechanism quantity determination unit, a second training unit, and a prediction unit.
[0202] The second training data acquisition unit can be used to acquire network anomaly labels corresponding to the second training data of IoT traffic and the second training data of IoT association; the second training vector matrix determination unit can be used to perform word embedding processing on the second training data of IoT association to obtain the second training vector matrix; the training parameter acquisition unit can be used to train the deformer through the second training vector matrix and the network anomaly labels to obtain the number of second training multi-head attention mechanisms determined for the deformer based on the prediction vector matrix; the second prediction multi-head attention mechanism number determination unit can be used to predict the rank and spectral mean of the second training vector matrix through the target network model to obtain the number of second prediction multi-head attention mechanisms; the second training unit can be used to train the target network model through the number of second training multi-head attention mechanisms and the number of second prediction multi-head attention mechanisms; the prediction unit can be used to process the rank and spectral mean of the prediction vector matrix through the trained target network model to determine the number of multi-head attention mechanisms of the deformer.
[0203] In some embodiments, the IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device feature data in the IoT system; wherein, the prediction vector matrix of the IoT association data includes: sensor prediction vector matrix, operational prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and device feature prediction vector matrix; wherein, the Fourier spectrum matrix includes: sensor word spectrum matrix, operational spectrum matrix, maintenance record spectrum matrix, environmental condition spectrum matrix, and device feature spectrum matrix; wherein, the word embedding module 1602 may include: a sensor prediction vector matrix acquisition submodule.
[0204] The sensor prediction vector matrix acquisition submodule can be used to perform word embedding processing on sensor data, operation data, maintenance record data, environmental condition data and equipment feature data respectively to obtain sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix and equipment feature prediction vector matrix.
[0205] The rank determination module 1603 may include: a rank determination submodule that runs the prediction vector matrix.
[0206] The rank determination submodule for the operation prediction vector matrix can be used to determine the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
[0207] The Fourier spectrum matrix determination submodule may include: a sensor word spectrum matrix determination unit.
[0208] The sensor word spectrum matrix determination unit can be used to perform Fourier transform processing on the sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix and equipment feature prediction vector matrix to obtain the sensor word spectrum matrix, operation spectrum matrix, maintenance record spectrum matrix, environmental condition spectrum matrix and equipment feature spectrum matrix, respectively.
[0209] The spectrum mean determination submodule may include: mean determination unit.
[0210] The mean value determination unit can be used to determine the mean values of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment characteristic spectrum matrix, respectively.
[0211] The deformer parameter setting submodule may include a multi-head number determination unit.
[0212] The multi-head number determination unit can be used to determine the number of multi-head attention mechanisms in the deformer based on the spectral mean of the sensor word spectrum matrix, operation spectrum matrix, maintenance record spectrum matrix, environmental condition spectrum matrix, and equipment feature spectrum matrix, as well as the rank of the sensor prediction vector matrix, operation prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and equipment feature prediction vector matrix.
[0213] In some embodiments, the attention mechanism quantity determination module 1604 may include: a fitting relationship acquisition submodule and an interpolation submodule.
[0214] The fitting relationship acquisition submodule can be used to obtain the numerical fitting relationship between the number of multi-head attention mechanisms of the deformer and the rank of the matrix; the interpolation submodule can be used to perform interpolation processing on the numerical fitting relationship to determine the number of multi-head attention mechanisms corresponding to the rank of the prediction vector matrix as the number of multi-head attention mechanisms of the deformer.
[0215] Since the functions of the device 1600 have been described in detail in their corresponding method embodiments, they will not be repeated here.
[0216] The modules and / or sub-modules and / or units described in the embodiments of this application can be implemented in software or hardware. The described modules and / or sub-modules and / or units can also be located in a processor. The names of these modules and / or sub-modules and / or units do not, in some cases, constitute a limitation on the module and / or sub-module and / or unit itself.
[0217] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a portion of a module or program segment containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer program instructions.
[0218] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this application, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0219] Figure 17 A schematic diagram of an electronic device suitable for implementing embodiments of this application is shown. It should be noted that... Figure 17 The electronic device 1700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0220] like Figure 17 As shown, the electronic device 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1702 or a program loaded from a storage section 1708 into a random access memory (RAM) 1703. The RAM 1703 also stores various programs and data required for the operation of the electronic device 1700. The CPU 1701, ROM 1702, and RAM 1703 are interconnected via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.
[0221] The following components are connected to I / O interface 1705: an input section 1706 including a keyboard, mouse, etc.; an output section 1707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN card, modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. Drive 1710 is also connected to I / O interface 1705 as needed. Removable media 1711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1710 as needed so that computer programs read from it can be installed into storage section 1708 as needed.
[0222] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing computer program instructions for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1709, and / or installed from removable medium 1711. When the computer program is executed by central processing unit (CPU) 1701, it performs the functions defined above in the system of this application.
[0223] It should be noted that the computer-readable storage medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable computer program instructions. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Computer program instructions contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0224] In another aspect, this application also provides a computer-readable storage medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable storage medium carries one or more programs that, when executed by the device, enable the device to perform the following functions: acquiring IoT traffic association prediction data in an IoT system; performing word embedding processing on the IoT traffic association prediction data to obtain a prediction vector matrix of the IoT association data; determining the rank of the prediction vector matrix; determining the number of multi-head attention mechanisms in the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms in the deformer is proportional to the rank of the prediction vector matrix; setting the parameters of the deformer based on the number of multi-head attention mechanisms; and training the deformer using the IoT traffic association prediction data so as to determine whether network traffic is abnormal using the trained deformer.
[0225] According to one aspect of this application, a computer program product or computer program is provided, comprising computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and a processor executes the computer program instructions to implement the methods provided in various optional implementations of the above embodiments.
[0226] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive) and includes several computer program instructions to cause an electronic device (such as a server or terminal device) to execute the method according to the embodiments of this application.
[0227] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application filed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not claimed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0228] It should be understood that this application is not limited to the detailed structure, drawing style or implementation method shown herein; on the contrary, this application is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.
Claims
1. A method for determining network traffic anomalies, characterized in that, include: Obtain IoT traffic correlation prediction data from IoT systems; The IoT traffic association prediction data is subjected to word embedding processing to obtain the prediction vector matrix of IoT association data; Determine the rank of the prediction vector matrix; The number of multi-head attention mechanisms of the deformer is determined based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix; The parameters of the deformer are set according to the number of multi-head attention mechanisms; The deformer is trained using the IoT traffic correlation prediction data so that the trained deformer can determine whether network traffic is abnormal.
2. The method according to claim 1, characterized in that, The number of multi-head attention mechanisms in the deformer is determined based on the rank of the prediction vector matrix, including: Obtain the first training data of IoT traffic and the network anomaly label corresponding to the first training data of IoT traffic; The first training data of the Internet of Things traffic is subjected to word embedding processing to obtain the first training vector matrix; The deformer is trained using the first training vector matrix and the network anomaly labels to obtain the first number of multi-head attention mechanisms for the deformer determined by the prediction vector matrix. The first number of predicted multi-head attention mechanisms is obtained by predicting the rank of the first training vector matrix using the target network model. The target network model is trained using the first number of trained multi-head attention mechanisms and the first number of predicted multi-head attention mechanisms. The number of multi-head attention mechanisms in the deformer is determined by processing the rank of the prediction vector matrix using the trained target network model.
3. The method according to claim 2, characterized in that, The IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data from the IoT system; wherein, the prediction vector matrix of the IoT association data includes: sensor prediction vector matrix, operational prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and device characteristic prediction vector matrix; The process of performing word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of the IoT association data includes: The sensor data, the operation data, the maintenance record data, the environmental condition data, and the equipment feature data are each subjected to word embedding processing to obtain the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix. Determining the rank of the prediction vector matrix includes: The ranks of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix are determined respectively. The determination of the number of multi-head attention mechanisms in the deformer based on the rank of the prediction vector matrix includes: The number of multi-head attention mechanisms in the deformer is determined based on the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix.
4. The method according to claim 1, characterized in that, The number of multi-head attention mechanisms in the deformer is determined based on the rank of the prediction vector matrix, including: Perform a Fourier transform on the prediction vector matrix to obtain the Fourier spectrum matrix; The mean spectral value is determined based on the Fourier spectrum matrix. The number of multi-head attention mechanisms of the deformer is determined based on the mean of the spectrum and the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to both the rank of the prediction vector matrix and the mean of the spectrum.
5. The method according to claim 4, characterized in that, The number of multi-head attention mechanisms in the deformer is determined based on the spectral mean and the rank of the prediction vector matrix, including: Obtain the second training data of IoT traffic and the network anomaly labels corresponding to the second training data of IoT association; The second training data associated with the Internet of Things is subjected to word embedding processing to obtain a second training vector matrix; The deformer is trained using the second training vector matrix and the network anomaly labels to obtain the second number of multi-head attention mechanisms for the deformer determined by the prediction vector matrix. The second predicted number of multi-head attention mechanisms is obtained by predicting the rank of the second training vector matrix and the mean of the spectrum using the target network model. The target network model is trained using the second number of trained multi-head attention mechanisms and the second number of predicted multi-head attention mechanisms. The number of multi-head attention mechanisms in the deformer is determined by processing the rank of the prediction vector matrix and the mean of the spectrum using the trained target network model.
6. The method according to claim 4, characterized in that, The IoT traffic association prediction data includes sensor data, operational data, maintenance record data, environmental condition data, and device characteristic data from the IoT system; wherein, the prediction vector matrix of the IoT association data includes: sensor prediction vector matrix, operational prediction vector matrix, maintenance record prediction vector matrix, environmental condition prediction vector matrix, and device characteristic prediction vector matrix; wherein, the Fourier spectrum matrix includes: sensor word spectrum matrix, operational spectrum matrix, maintenance record spectrum matrix, environmental condition spectrum matrix, and device characteristic spectrum matrix; The process of performing word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of the IoT association data includes: The sensor data, the operation data, the maintenance record data, the environmental condition data, and the equipment feature data are each subjected to word embedding processing to obtain the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix. Determining the rank of the prediction vector matrix includes: Determine the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix; The Fourier transform of the prediction vector matrix to obtain the Fourier spectrum matrix includes: Fourier transform is performed on the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the equipment feature prediction vector matrix to obtain the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment feature spectrum matrix, respectively. Determining the spectral mean based on the Fourier spectrum matrix includes: The mean spectral values of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the equipment feature spectrum matrix are determined respectively. The determination of the number of multi-head attention mechanisms in the deformer based on the mean of the spectrum and the rank of the prediction vector matrix includes: The number of multi-head attention mechanisms in the deformer is determined based on the spectral mean of the sensor word spectrum matrix, the operation spectrum matrix, the maintenance record spectrum matrix, the environmental condition spectrum matrix, and the device feature spectrum matrix, as well as the rank of the sensor prediction vector matrix, the operation prediction vector matrix, the maintenance record prediction vector matrix, the environmental condition prediction vector matrix, and the device feature prediction vector matrix.
7. The method according to claim 1, characterized in that, The number of multi-head attention mechanisms in the deformer is determined based on the rank of the prediction vector matrix, including: Obtain the numerical fitting relationship between the number of multi-head attention mechanisms in the deformer and the rank of the matrix; The numerical fitting relationship is interpolated to determine the number of multi-head attention mechanisms corresponding to the rank of the prediction vector matrix, which is then used as the number of multi-head attention mechanisms in the deformer.
8. A device for determining network traffic anomalies, characterized in that, include: The correlation prediction data acquisition module is used to acquire correlation prediction data of IoT traffic in the IoT system; The word embedding module is used to perform word embedding processing on the IoT traffic association prediction data to obtain the prediction vector matrix of IoT association data. A rank determination module is used to determine the rank of the prediction vector matrix; The attention mechanism number determination module is used to determine the number of multi-head attention mechanisms of the deformer based on the rank of the prediction vector matrix, wherein the number of multi-head attention mechanisms of the deformer is proportional to the rank of the prediction vector matrix; The parameter setting module is used to set the parameters of the deformer according to the number of multi-head attention mechanisms; The training module is used to train the deformer using the IoT traffic association prediction data, so that the trained deformer can be used to determine whether network traffic is abnormal.
9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer program instructions; the processor calls the computer program instructions stored in the memory to implement the network traffic anomaly determination method as described in any one of claims 1-7.
10. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the network traffic anomaly determination method as described in any one of claims 1-7.
Citation Information
Patent Citations
Low-rank CSI feedback method for deep iterative neural network, storage medium and equipment
CN112468203A
Network traffic anomaly detection method and system based on self-attention mechanism
CN115630298A