Power communication service identification method and system based on adaptive multi-channel attention

By constructing a teacher and student network through an adaptive multi-channel attention mechanism and knowledge distillation technology, the problem of high-dimensional data processing in power communication networks is solved, achieving efficient and rapid service identification, improving identification accuracy and speed, and making it suitable for intelligent identification in power communication networks.

CN119862549BActive Publication Date: 2025-12-30STATE GRID FUJIAN POWER ELECTRIC CO ECONOMIC RESEARCH INSTITUTE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411770665.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-12-30
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing methods for identifying services in power communication networks struggle to handle the characteristics of high-dimensional, nonlinear data and cannot adaptively adjust the processing of data characteristics, resulting in low identification accuracy, high computational resource consumption, and failure to meet the efficiency and robustness requirements of power communication networks. Furthermore, traditional models are slow to train and infer, which cannot meet the requirements for fast and accurate service identification.

Method used

An adaptive multi-channel attention-based power communication service identification method is adopted. By constructing teacher and student networks, and utilizing adaptive knowledge distillation technology, the model is compressed and the knowledge transfer process is dynamically adjusted. Combined with adaptive multi-channel mechanism, attention mechanism and interactive attention mechanism, high-precision and fast service identification is achieved.

Benefits of technology

It achieves high recognition accuracy while reducing computing resource consumption, improves the speed and accuracy of power communication service recognition, and is suitable for intelligent recognition application scenarios with complex data characteristics and dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862549B_ABST
    Figure CN119862549B_ABST
Patent Text Reader

Abstract

The application discloses a power communication service identification method and system based on adaptive multi-channel attention. The constructed power communication service identification network based on adaptive multi-channel attention is used as a teacher network. Through multiple parallel feature channels, diversified feature information in input data can be effectively captured. Based on power communication service sample data, the teacher network is used for adaptive knowledge distillation of an initial student network to obtain a final student network. Through knowledge distillation, the model can be compressed, and the speed of service identification is improved. The adaptive knowledge distillation can dynamically adjust the knowledge transmission process with the number of iterations, so that the lightweight student network can significantly reduce the consumption of computing resources while still maintaining high recognition accuracy. The power communication network service identification with high precision, adaptive adjustment and fast reasoning is realized, and the speed and accuracy of power communication service identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power communication networks, and in particular to a method and system for identifying power communication services based on adaptive multi-channel attention. Background Technology

[0002] With the continuous expansion and increasing complexity of power systems, power communication networks play a crucial role in ensuring the safe, stable, and efficient operation of the power grid. However, existing service identification and classification technologies still face numerous challenges. On the one hand, the service flows in power communication networks are diverse, with varied and complex data dimensions. Traditional port-based and DPI (Deep Packet Inspection) methods struggle to effectively handle the characteristics of high-dimensional, nonlinear data, resulting in low identification accuracy. On the other hand, in real-world scenarios, traffic in power communication networks exhibits dynamic changes. Traditional methods cannot adaptively adjust their processing of data characteristics, leading to uneven resource allocation and increasingly severe network congestion. Furthermore, with the introduction of new technologies such as big data and the Internet of Things, the service traffic and data volume that power communication networks need to handle have increased dramatically. Existing methods show significant shortcomings in terms of computational resource consumption and processing efficiency, making it difficult to meet the efficiency and robustness requirements of modern power systems for service identification. Traditional methods exhibit significant limitations when dealing with complex data characteristics, dynamic network environments, and high computational resource demands, especially when facing diverse service flows and constantly changing network states in power communication networks. Existing technologies cannot provide sufficiently refined and efficient support.

[0003] Therefore, developing innovative power communication service identification methods that are adaptive, resource-efficient, and possess high recognition accuracy is crucial for improving the overall performance of power communication networks. Power communication service identification requires extremely high real-time performance and accuracy. Existing models, due to their complex network structures, suffer from slow training and inference speeds, failing to meet the requirements for fast and accurate service identification. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for identifying power communication services based on adaptive multi-channel attention, which can improve the speed and accuracy of power communication service identification.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A method for identifying power communication services based on adaptive multi-channel attention, comprising the following steps:

[0007] A power communication service identification network based on adaptive multi-channel attention is constructed, and the power communication service identification network based on adaptive multi-channel attention is used as the teacher network;

[0008] The initial student network is determined based on the teacher network, and sample data of power communication services are obtained.

[0009] Based on the power communication service sample data, the teacher network is used to perform adaptive knowledge distillation on the initial student network to obtain the final student network.

[0010] The power communication service flow to be identified is obtained, and the power communication service flow is input into the final student network for power communication service identification to obtain the identification result.

[0011] To solve the above-mentioned technical problems, another technical solution adopted by the present invention is as follows:

[0012] A power communication service identification system based on adaptive multi-channel attention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps:

[0013] A power communication service identification network based on adaptive multi-channel attention is constructed, and the power communication service identification network based on adaptive multi-channel attention is used as the teacher network;

[0014] The initial student network is determined based on the teacher network, and sample data of power communication services are obtained.

[0015] Based on the power communication service sample data, the teacher network is used to perform adaptive knowledge distillation on the initial student network to obtain the final student network.

[0016] The power communication service flow to be identified is obtained, and the power communication service flow is input into the final student network for power communication service identification to obtain the identification result.

[0017] The beneficial effects of this invention are as follows: The constructed power communication service identification network based on adaptive multi-channel attention is used as the teacher network. Multiple parallel feature channels effectively capture diverse feature information in the input data. An initial student network is then determined based on the teacher network, and power communication service sample data is acquired. Based on the power communication service sample data, the teacher network performs adaptive knowledge distillation on the initial student network to obtain the final student network. Knowledge distillation compresses the model, improving the speed of service identification. Furthermore, the adaptive knowledge distillation dynamically adjusts the knowledge transfer process with the number of iterations, enabling the lightweight student network to maintain high identification accuracy while significantly reducing computational resource consumption. The acquired power communication service stream is input into the final student network for power communication service identification, yielding the identification result. This achieves high-precision, adaptively adjustable, and fast-reasoning power communication network service identification, improving both the speed and accuracy of power communication service identification. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the steps of a power communication service identification method based on adaptive multi-channel attention, according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the structure of a power communication service identification system based on adaptive multi-channel attention according to an embodiment of the present invention;

[0020] Figure 3 This is a schematic diagram of the power communication service identification network based on adaptive multi-channel attention in the power communication service identification method based on adaptive multi-channel attention according to an embodiment of the present invention.

[0021] Figure 4 This is a schematic diagram of adaptive knowledge distillation in the power communication service identification method based on adaptive multi-channel attention according to an embodiment of the present invention. Detailed Implementation

[0022] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0023] Please refer to Figure 1 A method for identifying power communication services based on adaptive multi-channel attention, comprising the following steps:

[0024] A power communication service identification network based on adaptive multi-channel attention is constructed, and the power communication service identification network based on adaptive multi-channel attention is used as the teacher network;

[0025] The initial student network is determined based on the teacher network, and sample data of power communication services are obtained.

[0026] Based on the power communication service sample data, the teacher network is used to perform adaptive knowledge distillation on the initial student network to obtain the final student network.

[0027] The power communication service flow to be identified is obtained, and the power communication service flow is input into the final student network for power communication service identification to obtain the identification result.

[0028] As can be seen from the above description, the beneficial effects of the present invention are as follows: The constructed power communication service identification network based on adaptive multi-channel attention is used as the teacher network. Multiple parallel feature channels can effectively capture diverse feature information in the input data. An initial student network is then determined based on the teacher network, and power communication service sample data is obtained. Based on the power communication service sample data, the teacher network performs adaptive knowledge distillation on the initial student network to obtain the final student network. Knowledge distillation can compress the model and improve the speed of service identification. Furthermore, the adaptive knowledge distillation can dynamically adjust the knowledge transfer process with the number of iterations, enabling the lightweight student network to maintain high identification accuracy while significantly reducing computational resource consumption. The obtained power communication service stream is input into the final student network for power communication service identification, yielding the identification result. This achieves high-precision, adaptively adjustable, and fast-reasoning power communication network service identification, improving the speed and accuracy of power communication service identification.

[0029] Furthermore, the construction of the power communication service identification network based on adaptive multi-channel attention includes:

[0030] Construct a single-channel power communication service identification network based on attention mechanism and interactive attention mechanism;

[0031] Based on the single-channel power communication service identification network, an adaptive multi-channel attention-based power communication service identification network is constructed.

[0032] As described above, a single-channel power communication service identification network based on attention and interactive attention mechanisms is first constructed, and then an adaptive multi-channel attention-based power communication service identification network is constructed based on the single-channel power communication service identification network. This allows the power communication service identification network to integrate adaptive multi-channel mechanisms, attention mechanisms, and interactive attention mechanisms, which can effectively capture more comprehensive and diverse feature information of the input data and improve identification accuracy.

[0033] Furthermore, the construction of a single-channel power communication service identification network based on attention mechanisms and interactive attention mechanisms includes:

[0034] Generate an embedding layer, an attention layer, a first residual connection and normalization layer, an interactive attention layer, a second residual connection and normalization layer, a first linear layer, and a second linear layer that are connected in sequence.

[0035] A single-channel power communication service identification network based on attention mechanism and interactive attention mechanism is established according to the embedding layer, the attention layer, the first residual connection and normalization layer, the interactive attention layer, the second residual connection and normalization layer, the first linear layer and the second linear layer.

[0036] As described above, the structure of the single-channel power communication service identification network based on attention mechanism and interactive attention mechanism includes an embedding layer, an attention layer, a first residual connection and normalization layer, an interactive attention layer, a second residual connection and normalization layer, a first linear layer, and a second linear layer connected in sequence. This allows the model to dynamically focus on the most important parts when processing input data, thereby improving model performance.

[0037] Further, the generation of the sequentially connected embedding layer, attention layer, first residual connection and normalization layer, interactive attention layer, second residual connection and normalization layer, first linear layer, and second linear layer includes:

[0038] Generate an embedding layer and set the dimension of the embedding layer to 512;

[0039] Generate an attention layer, a first residual connection and normalization layer, an interactive attention layer, and a second residual connection and normalization layer;

[0040] Generate a first linear layer and a second linear layer, and set the hidden neurons of the first linear layer to 2048 and the hidden neurons of the second linear layer to 512;

[0041] The embedding layer, the attention layer, the first residual connection and normalization layer, the interactive attention layer, the second residual connection and normalization layer, the first linear layer, and the second linear layer are connected in sequence.

[0042] As described above, the higher the dimension of the embedding layer, the larger the feature space that the network can represent, and the better it can capture the complex features of the input data. Setting the dimension of the embedding layer to 512 allows the embedding layer to provide more information, helping the network to better understand the input data. Setting the hidden neurons of the first linear layer to 2048 and the hidden neurons of the second linear layer to 512 ensures the accuracy of the network's recognition.

[0043] Furthermore, the construction of an adaptive multi-channel attention-based power communication service identification network based on the single-channel power communication service identification network includes:

[0044] Generate an adaptive weight layer, a feature fusion layer, a third linear layer, and a softmax layer that are connected in sequence;

[0045] Multiple single-channel power communication service identification networks are arranged in parallel and their outputs are fused together and then connected to the adaptive weight layer, the feature fusion layer, the third linear layer and the Softmax layer to obtain a power communication service identification network based on adaptive multi-channel attention.

[0046] As described above, by arranging multiple single-channel power communication service identification networks in parallel and fusing their outputs, and then connecting them with an adaptive weight layer, a feature fusion layer, a third linear layer, and a Softmax layer, a power communication service identification network based on adaptive multi-channel attention is obtained. Different attention weights are assigned to the input data, allowing the model to focus more on key information, significantly improving efficiency and accuracy in processing complex data, and enhancing feature representation capabilities.

[0047] Furthermore, the fusion of the outputs of multiple single-channel power communication service identification networks includes:

[0048]

[0049] In the formula, F fused This represents the fused feature matrix, where M represents the number of channels, and α... m R represents the adaptive weight corresponding to the m-th channel. m This represents the feature representation of the m-th channel.

[0050] As described above, the purpose of fusing the outputs of multiple single-channel power communication service identification networks is to weight and sum the features of each channel output according to adaptive weights to generate a unified feature representation, which facilitates subsequent data processing.

[0051] Furthermore, the step of using the teacher network to perform adaptive knowledge distillation on the initial student network based on the power communication service sample data to obtain the final student network includes:

[0052] The power communication service sample data is input into the teacher network. The cross-entropy loss function is used to calculate the loss between the predicted label and the true label of the teacher network. The loss is then backpropagated to update the parameters of the teacher network. The training is repeated multiple times until the teacher network converges and the value of the cross-entropy loss function is reduced to the minimum, thus obtaining the trained teacher network.

[0053] In the first stage, the power communication service sample data is input into the initial student network, and the cross-entropy loss function is used to calculate the loss between the predicted label and the true label of the initial student network.

[0054] In the second stage, the knowledge of the trained teacher network is passed to the initial student network using an adaptive knowledge distillation algorithm. The distillation loss is calculated based on the output of the trained teacher network and the output of the initial student network. The total loss of the student network is calculated based on the distillation loss, the loss of the initial student network, and the dynamic weight parameters. The training is repeated multiple times until the initial student network converges and the total loss is minimized, thus obtaining the final student network.

[0055] As described above, by implementing adaptive knowledge distillation, the knowledge of the teacher network is gradually transferred in two stages, and the transfer process is dynamically adjusted with the number of iterations, thereby optimizing the accuracy and efficiency of the student network. This results in a lightweight student network that maintains high recognition accuracy while significantly reducing computational resource consumption.

[0056] Furthermore, the step of calculating the loss between the predicted labels and the true labels of the teacher network using the cross-entropy loss function includes:

[0057]

[0058] In the formula, L represents the loss between the predicted label and the true label of the teacher network, N represents the number of samples, C represents the number of sample categories, and y i,j This represents the true label of the i-th power communication service sample in the j-th class. This represents the probability that the teacher network predicts the i-th power communication service sample to be of the j-th class.

[0059] As described above, the cross-entropy loss function can directly measure the difference between the teacher network's prediction and the true label. By minimizing the cross-entropy loss, the prediction accuracy of the teacher network can be effectively improved. Furthermore, the cross-entropy loss function conforms to the maximum likelihood estimation principle, which can effectively solve the gradient vanishing problem and make the teacher network more stable during training.

[0060] Furthermore, the step of calculating the loss between the predicted labels and the true labels of the initial student network using the cross-entropy loss function includes:

[0061]

[0062] In the formula, L cls This represents the loss between the predicted labels and the actual labels of the initial student network. This represents the true label of the i-th power communication service sample in the j-th class. Let N represent the probability that the student network predicts the i-th power communication service sample as belonging to the j-th class, where N represents the number of samples and C represents the number of sample classes.

[0063] The calculation of distillation loss based on the output of the trained teacher network and the output of the initial student network includes:

[0064]

[0065] In the formula, L KD This indicates the distillation loss, and T represents the distillation temperature. This represents the probability that the teacher network predicts the i-th power communication service sample to be of the j-th class;

[0066] The calculation of the total loss of the student network based on the distillation loss, the initial student network loss, and the dynamic weight parameters includes:

[0067] L total =λ·L cls +(1-λ)·L KD ;

[0068] In the formula, L total Let λ represent the total loss of the student network, and λ represent the dynamic weight parameters.

[0069] As described above, calculating the total loss of the student network based on distillation loss, the initial student network loss, and dynamic weight parameters can overcome the limitations of relying solely on distillation loss for model training. This allows the student network to better learn the knowledge of the teacher network, ensuring that the final student network can achieve fast and accurate identification of power communication services.

[0070] Please refer to Figure 2 Another embodiment of the present invention provides a power communication service identification system based on adaptive multi-channel attention, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements each step of the power communication service identification method based on adaptive multi-channel attention described above.

[0071] The power communication service identification method and system based on adaptive multi-channel attention described above are applicable to power communication service identification scenarios. The following detailed embodiments illustrate these methods:

[0072] Please refer to Figure 1 , Figure 3 and Figure 4 Embodiment 1 of the present invention is as follows:

[0073] A method for identifying power communication services based on adaptive multi-channel attention, comprising the following steps:

[0074] S1. Construct a power communication service identification network based on adaptive multi-channel attention, and use the power communication service identification network based on adaptive multi-channel attention as the teacher network, specifically including S11 to S13:

[0075] S11. Construct a single-channel power communication service identification network based on attention mechanism and interactive attention mechanism, specifically including S111~S112:

[0076] S111. Generate an embedding layer, an attention layer, a first residual connection and normalization layer, an interactive attention layer, a second residual connection and normalization layer, a first linear layer, and a second linear layer that are connected in sequence.

[0077] Specifically, an embedding layer is generated, and the dimension of the embedding layer is set to 512; an attention layer, a first residual connection and normalization layer, an interactive attention layer, and a second residual connection and normalization layer are generated; a first linear layer and a second linear layer are generated, and the hidden neurons of the first linear layer are set to 2048, and the hidden neurons of the second linear layer are set to 512; the embedding layer, the attention layer, the first residual connection and normalization layer, the interactive attention layer, the second residual connection and normalization layer, the first linear layer, and the second linear layer are connected sequentially.

[0078] The embedding layer uses one-dimensional convolution to encode the features of the input samples. The size of the convolution kernel is k, the stride is s, and the number of channels is C. The corresponding convolutional layer is represented as follows:

[0079] X′=Conv1D(X,W,b);

[0080] In the formula, X represents the input sample. This represents the output features after convolution. Indicates the weights of the convolution kernel. This indicates the bias. After the convolution operation, the ReLU activation function is applied, and max pooling is added, specifically:

[0081] E=MaxPooling1D(RELU(X′));

[0082] In the formula, d represents the feature dimension after pooling, and d′ represents the feature length after pooling.

[0083] The attention layer obtains query components, key components, and value components based on the input embedding features E through linear mapping, specifically:

[0084] Q = EW q +b q ;

[0085] K = EW k +b k ;

[0086] V = EW v +b v ;

[0087] In the formula, Indicates the query component. Indicates the bond component. Represents the value component, q m k represents the column vector of the query component matrix. m v represents the column vector of the key component matrix. m Represents the column vectors of the value component matrix. This represents the first learnable weight vector. This represents the learnable second weight vector. Let b represent the learnable third weight vector. q Indicates the first bias term, b k Indicates the second bias term, b v This represents the third bias term. Scaled dot product attention is applied to Q, K, and V to compute this local information, specifically as follows:

[0088]

[0089] In the formula, O represents the output of the attention layer.

[0090] The first residual connection and normalization layer performs a residual connection between the output O of the attention layer and the original input E, and then performs layer normalization, specifically as follows:

[0091] O′=LayerNorm(O+E);

[0092] In the formula, O′ represents the output of the first residual connection and the normalization layer.

[0093] The interactive attention layer captures high-order feature interactions by passing the output O′ of the first residual connection and normalization layer through the AIN layer (AttentionInteraction Network), specifically:

[0094]

[0095] In the formula, This represents the interaction features. The attention weights α for the interaction features are also included. ij Represented as:

[0096] α ij =softmax(f(z) ij )).

[0097] The output H of the AIN layer is represented as:

[0098]

[0099] The second residual connection and normalization layer is similar to the first residual connection and normalization layer, performing a residual connection between the output of the AIN layer and the original input, and then normalizing it through the layer.

[0100] The first linear layer and the second linear layer are processed using the ReLU activation function, specifically:

[0101] L = RELU(W lin Output norm +b lin );

[0102] In the formula, L represents the final output of the network, and W... lin Represents the weight matrix, Output norm This represents the result of the output H of the AIN layer after passing through the second residual connection and the normalization layer, b lin This represents the bias vector.

[0103] S112. A single-channel power communication service identification network based on attention mechanism and interactive attention mechanism is established according to the embedding layer, the attention layer, the first residual connection and normalization layer, the interactive attention layer, the second residual connection and normalization layer, the first linear layer, and the second linear layer. Figure 3 As shown.

[0104] S12. Construct an adaptive multi-channel attention-based power communication service identification network based on the single-channel power communication service identification network, such as... Figure 3 As shown, specifically including S121 to S122:

[0105] S121. Generate an adaptive weight layer, feature fusion layer, third linear layer and softmax layer connected in sequence.

[0106] The number of hidden neurons in the third linear layer is set to 2048.

[0107] S122. The multiple single-channel power communication service identification networks are arranged in parallel and the outputs of the multiple single-channel power communication service identification networks are fused and then connected to the adaptive weight layer, the feature fusion layer, the third linear layer and the Softmax layer to obtain a power communication service identification network based on adaptive multi-channel attention.

[0108] Assuming the output of each channel is Where N represents the number of samples, and d′ represents the output dimension of the fully connected layer. For the output R of each channel, the global features of all channels are concatenated into a feature matrix G, specifically:

[0109] G = [R1; R2; ...; R m ];

[0110] In the formula, The weight h for each channel is generated using an MLP (Multilayer Perceptron) network with two hidden layers, specifically as follows:

[0111] h = MLP(G);

[0112] In the formula, Finally, the adaptive weights for the M channels are obtained through softmax, specifically:

[0113] α = softmax(h);

[0114] In the formula, α={α1,α2,...,α m} represents the adaptive weight of each channel.

[0115] The step of arranging multiple single-channel power communication service identification networks in parallel includes:

[0116] The three single-channel power communication service identification networks are arranged in parallel.

[0117] The features of all channels are combined through a weighted fusion operation and input into a fully connected layer and a Softmax layer to generate a classification result. The goal of weighted fusion is to sum the features output by each channel according to adaptive weights to generate a unified feature representation. The fusion of the outputs of multiple single-channel power communication service identification networks includes:

[0118]

[0119] In the formula, This represents the fused feature matrix, where M represents the number of channels, and α... m R represents the adaptive weight corresponding to the m-th channel. m Let F represent the feature representation of the m-th channel. The weighted fused feature F... fused Further feature extraction is performed after the fully connected layer, specifically:

[0120] F = RELU(F fused ·W f +b f );

[0121] In the formula, This represents the weight matrix of the fully connected layer. d represents the bias term. out The dimension of the output feature is represented by F, and the extracted feature is represented by F.

[0122] S13. The power communication service identification network based on adaptive multi-channel attention is used as the teacher network.

[0123] S2. Determine the initial student network based on the teacher network and obtain sample data of power communication services.

[0124] The initial student network is structured as follows: an embedding layer, an attention layer, a first residual connection and normalization layer, an interactive attention layer, a second residual connection and normalization layer, a first linear layer, and a second linear layer, all connected sequentially. The embedding layer has a dimension of 256, the attention layer has a dimension of 256, the interactive attention layer has a dimension of 256, the first linear layer has 1024 hidden neurons, and the second linear layer has 256 hidden neurons. Therefore, the initial student network is a simplification of the teacher network.

[0125] S3. Based on the power communication service sample data, use the teacher network to perform adaptive knowledge distillation on the initial student network to obtain the final student network, such as... Figure 4 As shown, specifically including S31 to S33:

[0126] S31. Input the power communication service sample data into the teacher network, calculate the loss between the predicted label and the real label of the teacher network using the cross-entropy loss function, and backpropagate to update the parameters of the teacher network. Train multiple times until the teacher network converges and the value of the cross-entropy loss function is reduced to the minimum, thus obtaining the trained teacher network.

[0127] The step of calculating the loss between the predicted labels and the true labels of the teacher network using the cross-entropy loss function includes:

[0128]

[0129] In the formula, L represents the loss between the predicted label and the true label of the teacher network, N represents the number of samples, C represents the number of sample categories, and y i,j This represents the true label of the i-th power communication service sample in the j-th class. This represents the probability that the teacher network predicts the i-th power communication service sample to be of the j-th class.

[0130] S32. In the first stage, the power communication service sample data is input into the initial student network, and the loss between the predicted label and the real label of the initial student network is calculated using the cross-entropy loss function.

[0131] The step of calculating the loss between the predicted labels and the true labels of the initial student network using the cross-entropy loss function includes:

[0132]

[0133] In the formula, L cls This represents the loss between the predicted labels and the actual labels of the initial student network. This represents the true label of the i-th power communication service sample in the j-th class. This represents the probability that the student network predicts the i-th power communication service sample as belonging to the j-th class.

[0134] S33. In the second stage, the knowledge of the trained teacher network is passed to the initial student network using an adaptive knowledge distillation algorithm. The distillation loss is calculated based on the output of the trained teacher network and the output of the initial student network. The total loss of the student network is calculated based on the distillation loss, the loss of the initial student network, and the dynamic weight parameters. The training is repeated multiple times until the initial student network converges and the total loss is minimized, thus obtaining the final student network.

[0135] The step of calculating the distillation loss based on the output of the trained teacher network and the output of the initial student network includes:

[0136]

[0137] In the formula, L KD This indicates the distillation loss, and T represents the distillation temperature. This represents the probability that the teacher network predicts the i-th power communication service sample to be of the j-th class.

[0138] Knowledge distillation is a technique for compressing deep learning models. It extracts knowledge from a "teacher" model and then transfers this knowledge to a smaller, lighter student model. In traditional knowledge distillation methods, the teacher model's prediction output is used as soft labels, and the student model learns the teacher model's knowledge by minimizing the loss between these soft labels. Power communication service identification requires extremely high real-time performance and accuracy. Existing models, due to their complex network structures, suffer from slow training and inference speeds, failing to meet the requirements for rapid service identification. Knowledge distillation can compress the model, improving the speed and accuracy of service identification. However, with increasing task complexity and the expansion of deep learning network structures, simple knowledge transfer methods cannot fully utilize the rich information in the teacher model. To address this issue and fully utilize the rich information in the teacher model, this invention proposes an adaptive knowledge distillation framework to compress the power communication service identification model into a lightweight model, thereby improving the speed and accuracy of service identification.

[0139] The adaptive knowledge distillation algorithm adjusts the focus on two loss functions using a weight parameter λ. Specifically, it operates on the loss of the model's output layer, controlling and switching stages through the weights of the loss layer. λ represents the current level of focus on the corresponding task, and the weights of the loss terms have a significant impact on the learning process; tasks with larger loss values ​​are updated faster than those with smaller loss values. In the first stage, λ is set to 1, and only the learning task on the original data is performed, as the weights of tasks in other stages are 0, stopping the backpropagation of their losses. The student network parameter update in the first stage is represented as follows:

[0140]

[0141] In the formula, Δw1 represents the student network parameters after the first stage update, and w1 represents the student network parameters before the first stage update.

[0142] In the second phase, λ gradually decreases until it reaches 0, depending on the number of training rounds. The goal is to learn from both the raw and predicted data simultaneously, gradually transferring the knowledge from the teacher network to the student network. λ is represented as:

[0143]

[0144] In the formula, B represents the current training round, B1 represents the number of training rounds in the first phase, and B2 represents the number of training rounds in the second phase. max This indicates the maximum number of training rounds set. Once the previous task is trained well, the focus on subsequent tasks will increase rapidly.

[0145] The calculation of the total loss of the student network based on the distillation loss, the initial student network loss, and the dynamic weight parameters includes:

[0146] L total =λ·L cls +(1-λ)·L KD ;

[0147] In the formula, L total Let λ represent the total loss of the student network, and λ represent the dynamic weight parameters.

[0148] The second phase of student network parameter updates is represented as follows:

[0149]

[0150] In the formula, Δw2 represents the student network parameters after the second stage update, and w2 represents the student network parameters before the second stage update.

[0151] S4. Obtain the power communication service flow to be identified, and input the power communication service flow into the final student network for power communication service identification to obtain the identification result.

[0152] The identification result is the type of power communication service. In one optional implementation, the identification result includes dispatch telephone, power metering, dispatch automation, relay protection, protection management system, safety and stability management system, video conferencing, video surveillance, administrative telephone, monitoring system, financial and marketing information, or administrative management, as shown in Table 1.

[0153] Table 1 Types of Power Communication Services

[0154]

[0155] This invention addresses the diverse service flows and constantly changing network states in power communication networks by integrating an adaptive multi-channel mechanism, an attention mechanism, and an interactive attention mechanism. It effectively captures diverse feature information from input data through multiple parallel feature channels and introduces an adaptive weighting mechanism to dynamically adjust the contributions of different channels. Furthermore, it employs an adaptive knowledge distillation method to progressively transfer knowledge from the teacher network in two stages, dynamically adjusting the transfer process with each iteration. This allows the lightweight student network to maintain high recognition accuracy while significantly reducing computational resource consumption. Thus, it achieves high recognition accuracy, adaptive adjustment, and fast inference for power communication network service recognition, making it particularly suitable for intelligent recognition applications with complex data features, dynamic network environments, and high computational resource requirements.

[0156] Please refer to Figure 2 Embodiment two of the present invention is as follows:

[0157] A power communication service identification system based on adaptive multi-channel attention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps of the aforementioned power communication service identification method based on adaptive multi-channel attention.

[0158] In summary, this invention provides a method and system for identifying power communication services based on adaptive multi-channel attention. The constructed power communication service identification network based on adaptive multi-channel attention serves as the teacher network. Multiple parallel feature channels effectively capture diverse feature information from the input data. An initial student network is then determined based on the teacher network, and power communication service sample data is acquired. Based on the power communication service sample data, the teacher network performs adaptive knowledge distillation on the initial student network to obtain the final student network. Knowledge distillation compresses the model, improving the speed of service identification. Furthermore, the adaptive knowledge distillation dynamically adjusts the knowledge transfer process with the number of iterations, enabling the lightweight student network to maintain high identification accuracy while significantly reducing computational resource consumption. The acquired power communication service stream is then input into the final student network for power communication service identification, resulting in the identified service. As a result, high-precision, adaptively adjustable, and fast-reasoning power communication network service identification was achieved, improving the speed and accuracy of power communication service identification. Furthermore, a single-channel power communication service identification network based on attention and interactive attention mechanisms was first constructed, followed by the construction of an adaptive multi-channel attention-based power communication service identification network. This integrated adaptive multi-channel, attention, and interactive attention mechanisms, effectively capturing more comprehensive and diverse feature information from the input data, thus improving identification accuracy. Moreover, by implementing adaptive knowledge distillation, the knowledge of the teacher network was gradually transferred in two stages, with the transfer process dynamically adjusted with each iteration, optimizing the accuracy and efficiency of the student network. This resulted in a lightweight student network that maintained high identification accuracy while significantly reducing computational resource consumption.

[0159] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for power communication service identification based on adaptive multi-channel attention, characterized in that, The method comprises the steps of: constructing an adaptive multi-channel attention-based power communication service identification network and taking the adaptive multi-channel attention-based power communication service identification network as a teacher network; determining an initial student network based on the teacher network and obtaining power communication service sample data; performing adaptive knowledge distillation on the initial student network using the teacher network based on the power communication service sample data to obtain a final student network; obtaining a power communication service stream to be identified and inputting the power communication service stream into the final student network for power communication service identification to obtain an identification result; the constructing an adaptive multi-channel attention-based power communication service identification network comprises: constructing a single-channel power communication service identification network based on an attention mechanism and an interactive attention mechanism; constructing an adaptive multi-channel attention-based power communication service identification network based on the single-channel power communication service identification network; the constructing an adaptive multi-channel attention-based power communication service identification network based on the single-channel power communication service identification network comprises: generating an adaptive weight layer, a feature fusion layer, a third linear layer and a Softmax layer connected in sequence; parallelly arranging a plurality of the single-channel power communication service identification networks, fusing outputs of the plurality of single-channel power communication service identification networks, and connecting the fused outputs with the adaptive weight layer, the feature fusion layer, the third linear layer and the Softmax layer to obtain an adaptive multi-channel attention-based power communication service identification network; the performing adaptive knowledge distillation on the initial student network using the teacher network based on the power communication service sample data to obtain a final student network comprises: inputting the power communication service sample data into the teacher network, calculating a loss between predicted labels and real labels of the teacher network by using a cross-entropy loss function, performing back propagation, updating parameters of the teacher network, training multiple times, until the teacher network converges, a value of the cross-entropy loss function is minimized, and a trained teacher network is obtained; in a first stage, inputting the power communication service sample data into the initial student network, and calculating a loss between predicted labels and real labels of the initial student network by using a cross-entropy loss function; in a second stage, transferring knowledge of the trained teacher network to the initial student network by using an adaptive knowledge distillation algorithm, calculating a distillation loss according to an output of the trained teacher network and an output of the initial student network, calculating a total loss of the student network according to the distillation loss, a loss of the initial student network and a dynamic weight parameter, training multiple times, until the initial student network converges, the total loss is minimized, and a final student network is obtained.

2. The power communication service recognition method based on adaptive multi-channel attention according to claim 1, characterized in that, the constructing a single-channel power communication service identification network based on an attention mechanism and an interactive attention mechanism comprises: generating an embedding layer, an attention layer, a first residual connection and a normalization layer, an interactive attention layer, a second residual connection and a normalization layer, a first linear layer and a second linear layer connected in sequence; According to the embedding layer, the attention layer, the first residual connection and normalization layer, the interaction attention layer, the second residual connection and normalization layer, the first linear layer, and the second linear layer, a single-channel power communication service identification network based on attention mechanism and interaction attention mechanism is established. 3.The power communication service identification method based on adaptive multi-channel attention of claim 2, wherein, The generating sequentially connected embedding layer, attention layer, first residual connection and normalization layer, interaction attention layer, second residual connection and normalization layer, first linear layer, and second linear layer comprises: An embedding layer is generated, and the dimension of the embedding layer is set to 512; An attention layer, a first residual connection and normalization layer, an interaction attention layer, and a second residual connection and normalization layer are generated; A first linear layer and a second linear layer are generated, the hidden neurons of the first linear layer are set to 2048, and the hidden neurons of the second linear layer are set to 512; The embedding layer, the attention layer, the first residual connection and normalization layer, the interaction attention layer, the second residual connection and normalization layer, the first linear layer, and the second linear layer are sequentially connected.

4. The power communication service recognition method based on adaptive multi-channel attention according to claim 1, characterized in that, The fusing of the outputs of multiple single-channel power communication service identification networks comprises: ; In the formula, denotes the feature matrix after fusion, M denotes the number of channels, denotes the adaptive weight corresponding to the mth channel, denotes the feature representation of the mth channel.

5. The power communication service recognition method based on adaptive multi-channel attention according to claim 1, characterized in that, The calculating of the loss between the predicted label and the real label of the initial student network by using the cross-entropy loss function comprises: ; In the formula, L represents the loss between the predicted label and the real label of the teacher network, N represents the number of samples, C represents the number of sample categories, y i,j represents the real label of the i th power communication service sample in the j th category, pi,j represents the probability that the teacher network predicts the i th power communication service sample to be in the j th category.

6. The power communication service recognition method based on adaptive multi-channel attention according to claim 1, characterized in that, The calculating of the loss between the predicted label and the real label of the initial student network by using the cross-entropy loss function comprises: ; In the formula, denotes the loss between the predicted label and the real label of the initial student network, denotes the real label of the i-th power communication service sample in the j-th category, denotes the probability that the student network predicts the i-th power communication service sample to be in the j-th category, N denotes the number of samples, and C denotes the number of sample categories; The calculating of the distillation loss according to the output of the trained teacher network and the output of the initial student network comprises: ; In the formula, represents the distillation loss, T represents the distillation temperature, represents the probability that the teacher network predicts the i-th power communication service sample to be the j-th class. The calculating of the total loss of the student network according to the distillation loss, the loss of the initial student network, and a dynamic weight parameter comprises: ; wherein denotes the total loss of the student network, denotes the dynamic weight parameter.

7. A power communication service recognition system based on adaptive multi-channel attention, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements each step of the power communication service identification method based on adaptive multi-channel attention in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Target identification method based on multi-fusion deep neural network

    CN114565856A

  • Lightweight distributed optical fiber sensing event identification method and system

    CN116091897A