Proton whistle sound wave cross frequency identification model and method
Through the proton whistle cross-frequency identification model, the backbone network and multi-scale feature fusion network are used, combined with the lightweight attention mechanism and pruning module, the problem of inefficiency of the existing methods is solved, and efficient proton whistle cross-frequency identification is achieved to adapt to TB-level satellite observation data.
Patent Information
- Application Number
- CN202510673134.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
The existing method of extracting the cross frequency of proton whistle waves based on manual analysis of time-frequency graphs is too inefficient and difficult to adapt to the challenge of TB-level satellite observation data. Research on automated processing of proton whistle waves has not been carried out.
The proton whistle cross-frequency identification model is adopted, including backbone network, multi-scale feature fusion network and prediction classifier, combined with the CBS convolution module, DepthSepConv depth separable convolution module, SE lightweight attention mechanism module and SPPF multi-scale feature fusion module, the model complexity is reduced through the LAMP pruning module, and the multi-scale feature fusion and lightweight attention mechanism are used to improve the recognition efficiency.
High-precision identification of the cross frequency of proton whistle waves is achieved, which significantly improves the recognition efficiency and can process TB-level satellite observation data.
Smart Images

Figure CN120596887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of proton whistle wave analysis, and in particular to a proton whistle wave cross-frequency recognition model and method. Background Art
[0002] Proton whistlers (PWs) are a common electromagnetic wave phenomenon in low-Earth orbit (LEO). They arise from variations in the propagation characteristics of electromagnetic waves within the ionospheric plasma. When the wave frequency approaches the cyclotron frequency of protons and other ions, right-handed and left-handed ion cyclotron waves coexist. A linear polarization crossover frequency occurs between adjacent cyclotron frequencies, causing a polarization reversal, where the right-handed whistler wave transforms into an ion cyclotron wave (Gurnett et al., 1965; Shawhan, 1966). This phenomenon manifests itself in a spectrum diagram as a downward bend in the electron whistler wave at high frequencies, while the proton whistler wave bends upward at low frequencies, with the two intersecting at the crossover frequency. This crossover frequency not only explains the gradual approach of the whistler wave frequency to the proton cyclotron frequency but also serves as a key physical quantity for estimating the relative hydrogen ion concentration in the plasma (Muzzio, 1968; Hughes, 1997), providing important insights into the composition of the ionosphere and wave energy transmission.
[0003] However, existing methods for extracting proton whistler crossover frequencies based on manual analysis of time-frequency plots are unable to cope with the challenges of terabyte-scale satellite observation data. Although significant progress has been made in intelligent identification of electron whistlers (EWs), research on automated processing of proton whistlers has yet to be conducted.
[0004] Therefore, it is urgent to invent a proton whistle wave cross-frequency identification method to solve the problem that the existing method of extracting proton whistle wave cross-frequency based on manual analysis of time-frequency diagrams is too inefficient and difficult to adapt to the challenges of massive satellite observation data. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a proton whistle wave cross-frequency recognition model and method to solve the problem of low efficiency of existing proton whistle wave cross-frequency recognition.
[0006] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.
[0007] In order to achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0008] According to a first aspect of an embodiment of the present invention, a proton whistle wave cross-frequency recognition model is provided, the model comprising:
[0009] The backbone network is used to extract the proton whistle wave characteristics and the cross-frequency characteristics of the proton whistle wave from the electric field time-frequency data collected by the Zhangheng satellite;
[0010] The backbone network includes a CBS convolution module, a DepthSepConv depth-separable convolution module, a SE lightweight attention mechanism module, and an SPPF multi-scale feature fusion module;
[0011] The multi-scale feature fusion network is set between the backbone network and the prediction classifier, and is used to fuse the different levels of feature maps output by the backbone network to output multi-scale and multi-level feature maps;
[0012] The prediction classifier is used to extract the proton whistle wave detection results and the proton whistle wave cross-frequency detection results from the multi-scale and multi-level feature map;
[0013] The LAMP pruning module is used to score the weights of each channel in the model's convolutional layer and determine the channels to be removed based on the weight scores of each channel to reduce the complexity of the model.
[0014] Furthermore, the CBS convolution module consists of a Conv convolution layer, a batch normalization layer and a ReLU activation function;
[0015] The calculation process of the CBS convolution module is: G0 = ReLU (BN (s k ·F w,h +b k )), for w,h∈[1,320],k∈[1,16], where S k is the kth 3×3 convolution kernel of the Conv convolution layer, F w,h The input time-frequency Figure X [m,n] eigenvector at position (w,h), b k is the bias term of the kth output channel, BN() is batch normalization, ReLU() is the activation function, and G0 is the output result of the CBS convolution module.
[0016] Furthermore, the calculation process of the DepthSepConv depth separable convolution module is: G i =φ i,p (φ i,d (G i-1 )), where G i is the output result of the i-th DepthSepConv depth-separable convolution module, φ i,p With φ i,d They are the point-by-point convolution and depth-wise convolution of the i-th DepthSepConv depth-wise separable convolution module respectively;
[0017] φ d It is used to increase the dimension of the features of a stage, and improve the expression ability of time-frequency features by increasing the number of channels or dimensions of the features. p It is used to reduce the dimensionality of features, aggregate important features together, and filter out minor features.
[0018] Furthermore, the SE lightweight attention mechanism module is set at the end of the DepthSepConv depth-separable convolution module;
[0019] The SE lightweight attention mechanism module is used to weight channels in the network, thereby enhancing important features and suppressing irrelevant information;
[0020] The calculation process of the SE lightweight attention mechanism module is: G i =F scale (F ex (F sq (F tr (G i-1 )),V)⊙U), where G i is the output result of the SE lightweight attention mechanism module, G i-1 is the feature input of the DepthSepConv depth-separable convolution module, F tr To extract the image feature function, F sq is the channel compression function, F ex (, V) is the channel excitation function, V is the weight matrix, F scale () is the channel scaling function, ⊙ is the element-by-element multiplication symbol, and U is F tr Feature input after feature extraction, U=F tr (G i-1 );
[0021] Among them, i and j are the position indexes of the feature map in the height H and width W directions, U C (i, j) represents the eigenvalue of the Cth channel at the spatial position (i, j), F sq The function is used to compress the information of each channel into a single value through global average pooling;
[0022] F ex (z,V)=σ(V2·δ(V1·z))=s, where δ is the ReLU activation function, σ is the Sigmoid activation function, V1 is the channel compression layer weight, V2 is the channel recovery layer weight, and F ex The function is used to obtain the weight vector of each channel through two layers of 1×1 convolution, ReLU activation and Sigmoid activation;
[0023] G i =F scale (U,s)=s⊙U,F scale Function is used to apply channel weights to F tr Output feature U to achieve channel recalibration.
[0024] Furthermore, the backbone network includes a CBS convolution module, nine DepthSepConv depth-separable convolution modules and an SPPF multi-scale feature fusion module connected in sequence;
[0025] Among them, the SE lightweight attention mechanism module is provided at the end of the eighth DepthSepConv depth-separable convolution module, and the SE lightweight attention mechanism module is provided at the end of the ninth DepthSepConv depth-separable convolution module.
[0026] Furthermore, the prediction classifier is composed of comprehensive convolution modules of different scales, wherein the comprehensive convolution module is composed of a Bbox sub-convolution module, a Cls sub-convolution module and a Pose sub-convolution module;
[0027] The Bbox sub-convolution module is used to predict the proton whistle wave bounding box, the Cls sub-convolution module is used to predict the feature category, and the Pose sub-convolution module is used to predict the proton whistle wave cross-frequency coordinates and visibility confidence.
[0028] Furthermore, the loss function of the proton whistle wave cross-frequency recognition model is:
[0029]
[0030] Loss term L Cls is the binary cross entropy loss, which is used to measure the difference between the predicted probability distribution of the proton whistle wave cross frequency recognition model and the true label distribution, λ Cls is the loss term L Cls The corresponding weight coefficient;
[0031] Loss term L CloU The degree of overlap between the predicted bounding box and the true bounding box for detecting proton whistler waves, λ CloU is the loss term L CloU The corresponding weight coefficient;
[0032] Loss term L DFL It is used to optimize the position prediction accuracy of the coordinate proton whistle wave and enhance the learning ability of the proton whistle wave cross-frequency recognition model for samples with unclear proton whistle wave characteristics. DFL is the loss term L DFL The corresponding weight coefficient;
[0033] Loss term L Pose is the cross-frequency position loss, which is used to measure the difference between the predicted position and the true position of the proton whistle wave cross-frequency, λ Pose is the loss term L Pose The corresponding weight coefficient;
[0034] Loss Item is the cross-frequency visibility loss, which is used to optimize the model's learning ability for cross-frequency difficult samples. The calculation formula is: Among them, y is the true label indicating whether the cross-frequency point exists, is the probability of the existence of the cross-frequency point predicted by the model, and the TopK selection function indicates the selection of the top K cross-frequency points with the largest loss value.
[0035] Furthermore, the scoring formula of the LAMP pruning module is:
[0036] Where W[u] represents the weight value at the u-th position in the weight tensor W, ∑ v≥u (W[v]) 2 Represents the sum of all squared weights connected after the current weight position u.
[0037] According to a second aspect of an embodiment of the present invention, a method for identifying cross-frequency of a proton whistle wave is provided, the method comprising:
[0038] Obtain electric field load data collected by the Zhangheng satellite;
[0039] performing data preprocessing on the electric field load data to obtain preprocessed time-frequency data;
[0040] The time-frequency data is input into a proton whistle wave cross-frequency recognition model as described in any one of the above items to obtain a proton whistle wave recognition result and a proton whistle wave cross-frequency recognition result corresponding to the time-frequency data.
[0041] Furthermore, the electric field load data is preprocessed to obtain preprocessed time-frequency data, including:
[0042] intercepting the electric field load data using a preset sliding time window to obtain waveform data segments;
[0043] Overlapping short-time Fourier transform processing is performed on the waveform data segment to obtain a time-frequency diagram corresponding to the waveform data segment.
[0044] The present invention discloses a proton whistle wave cross-frequency recognition model and method. The model includes: a backbone network for extracting proton whistle wave features and proton whistle wave cross-frequency features from electric field time-frequency data collected by the Zhang Heng satellite, the backbone network including a CBS convolution module, a depthwise separable convolution module, an SE lightweight attention mechanism module, and an SPPF multi-scale feature fusion module; a multi-scale feature fusion network for fusing feature maps of different levels output by the backbone network; a prediction classifier for extracting proton whistle wave cross-frequency detection results from the multi-scale and multi-level feature maps; and a LAMP pruning module for scoring the weights of each channel in the model convolution layer, determining the channels to be removed based on the weight scores of each channel, and reducing the complexity of the model. The embodiments of the present invention achieve high-precision recognition of proton whistle wave cross-frequency and significantly improve the recognition efficiency of proton whistle wave cross-frequency by reducing the amount of model calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.
[0046] Figure 1 A schematic diagram of the structure of a backbone network of a proton whistle wave cross-frequency recognition model provided by an embodiment of the present invention;
[0047] Figure 2 A comparative schematic diagram of the backbone network structure provided by an embodiment of the present invention;
[0048] Figure 3 A schematic diagram of the calculation flow of the SE lightweight attention mechanism module of the backbone network L8 layer provided by an embodiment of the present invention;
[0049] Figure 4 A schematic diagram of the network structure of a prediction classifier provided by an embodiment of the present invention;
[0050] Figure 5 A schematic diagram illustrating the principle of identifying cross-frequency characteristics of proton whistle waves provided in an embodiment of the present invention;
[0051] Figure 6 Schematic diagram of the principle of the LRM loss function provided by an embodiment of the present invention;
[0052] Figure 7 Schematic diagram of the principle of the LAMP pruning module provided in an embodiment of the present invention;
[0053] Figure 8A schematic diagram of proton whistle waveform data provided by an embodiment of the present invention;
[0054] Figure 9 A schematic diagram of the time-frequency data of a proton whistle wave provided in an embodiment of the present invention;
[0055] Figure 10 A schematic structural diagram of a proton whistle wave cross-frequency recognition model is provided for an embodiment of the present invention;
[0056] Figure 11 A PR curve diagram of proton whistle wave crossover frequency detection provided by an embodiment of the present invention;
[0057] Figure 12 One of the effect diagrams of the Yolov8 model provided in an embodiment of the present invention on the recognition of the crossover frequency of proton whistle waves;
[0058] Figure 13 One of the recognition effect diagrams of the proton whistle wave cross-frequency recognition model provided in an embodiment of the present invention;
[0059] Figure 14 Figure 2 shows the effect of the Yolov8 model provided in an embodiment of the present invention on the recognition of the crossover frequency of proton whistle waves;
[0060] Figure 15 The second diagram of the recognition effect of the proton whistle wave cross-frequency recognition model provided by an embodiment of the present invention;
[0061] Figure 16 This is a diagram showing the output features of the Conv convolution module in the L0 layer of the backbone network provided by an embodiment of the present invention;
[0062] Figure 17 The backbone network L1 to L1 provided by the embodiment of the present invention 10 Output feature rendering of the layer-wise lightweight convolution module;
[0063] Figure 18 Output feature rendering of the spatial pyramid pooling module provided by an embodiment of the present invention;
[0064] Figure 19 A comparison chart of the number of model channels before and after LAMP pruning provided by an embodiment of the present invention;
[0065] Figure 20 A schematic diagram of a module of a proton whistle wave cross-frequency recognition model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0066] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0067] In addition, the terms "comprises" and "having" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed but may include other steps or elements not expressly listed or inherent to such process, method, product or apparatus.
[0068] refer to Figure 10 and Figure 20 An embodiment of the present invention provides a proton whistle wave cross-frequency recognition model - PWNet. The embodiment of the present invention is constructed based on the principle of multi-scale feature pyramid, which specifically includes: a backbone network (Backbone), a multi-scale feature fusion network (Neck) and a prediction classifier (Head).
[0069] Specifically, the backbone network (Backbone) is used to extract proton whistle wave characteristics and proton whistle wave cross-frequency characteristics from the electric field time-frequency data collected by the Zhang Heng satellite.
[0070] The multi-scale feature fusion network (Neck) is set between the backbone network and the prediction classifier to fuse the feature maps of different levels at each stage output by the backbone network and output multi-scale and multi-level feature maps, thereby significantly enhancing the feature expression capability.
[0071] The above-mentioned prediction classifier (Head) adopts a decoupled design scheme to achieve multi-task learning, and extracts proton whistle wave detection results and proton whistle wave cross-frequency detection results from multi-scale and multi-level feature maps.
[0072] The model also includes a LAMP pruning module, which is used to score the weights of each channel in the model's convolutional layer, determine the channels to be removed based on the weight scores of each channel, reduce the model complexity, and build an efficient proton whistle wave cross-frequency recognition module.
[0073] The structure of the above backbone network is as follows Figure 1 As shown in (a).
[0074] The above backbone network includes a 3×3CBS convolution module, a DepthSepConv depth-separable convolution module, a SE lightweight attention mechanism module, and an SPPF multi-scale feature fusion module.
[0075] Preferably, the backbone network includes L0 to L 10 There are 11 layers of modules in total. The L0 layer has a 3×3 CBS convolution module, and the L1 to L9 layers each have a DepthSepConv depth-separable convolution module. 10 The layer is equipped with a SPPF multi-scale feature fusion module. Figure 2 A comparison diagram of the backbone network structure is shown. In addition, a SE lightweight attention mechanism module (Squeeze-and-Excitation) is also set at the end of the DepthSepConv depth-separable convolution module in the L8 and L9 layers.
[0076] like Figure 2 As shown in the figure, since the embodiment of the present invention replaces ordinary convolution with DepthSepConv convolution, the overall computational complexity and parameter count are greatly reduced, but the expressive power of a single DepthSepConv convolution module is also correspondingly weakened. To compensate for this change and ensure that the network can extract sufficient features, the design will appropriately add a layer of DepthSepConv. This helps ensure that the network can capture more global information in a lightweight design, and the impact on the overall inference speed is very limited.
[0077] The above 3×3 CBS convolution module consists of a Conv convolution layer, a batch normalization (BN) layer, and a ReLU activation function. The calculation process of the above CBS convolution module is: G0 = ReLU (BN (s k ·F w,h +b k )), for w,h∈[1,320],k∈[1,16], where S k is the kth 3×3 convolution kernel of the Conv convolution layer, F w,h The input time-frequency Figure X [m,n] eigenvector at position (w,h), b k is the bias term of the kth output channel, BN() is batch normalization, ReLU() is the activation function, and the output result of the L0 layer CBS convolution module
[0078] The above calculation process uses a linear combination method to complete the expansion and re-expression of features, thereby obtaining richer time-frequency features.
[0079] refer to Figure 1(b) The above DepthSepConv depth separable convolution module consists of depth convolution and point-by-point convolution structure, and its calculation process is: G i =φ i,p (φ i,d (G i-1 )), where G i is the output result of the i-th DepthSepConv depth-separable convolution module, φ i,p With φ i,d are the point-by-point convolution and depth-wise convolution of the i-th DepthSepConv depth-wise separable convolution module respectively; φ d It is used to increase the dimension of the features of a stage, and improve the expression ability of time-frequency features by increasing the number of channels or dimensions of the features. p It is used to reduce the dimensionality of features, aggregate important features together, and filter out minor features.
[0080] refer to Figure 1 (c) The SE lightweight attention mechanism module is mainly used to weight channels in the network, thereby enhancing important features and suppressing irrelevant information. This embodiment of the present invention places the SE lightweight attention mechanism module at the end of the DepthSepConv depthwise separable convolution module in the L8 and L9 layers, respectively.
[0081] by Figure 1 Taking the SE lightweight attention mechanism module of the L8 layer in (a) as an example, its calculation process is as follows Figure 3 shown.
[0082] The calculation process of the above SE lightweight attention mechanism module is: G8 = F scale (F ex (F sq (F tr (G7)),V)⊙U).
[0083] Among them, G8 is the output result of the SE lightweight attention mechanism module, is the feature input of L8 layer, F tr To extract the image feature function, F sq is the channel compression function, F ex (, V) is the channel excitation function, V is the weight matrix, F scale () is the channel scaling function, ⊙ is the element-by-element multiplication symbol, and U is F tr Feature input after feature extraction.
[0084] Combine Figure 3 As shown in (a), F tr The feature extraction formula is: U = F tr(G7), the feature map dimension (512, 20, 20) of the current feature input G7 after DepthSepConv convolution output corresponds to (C, W, H) in the figure.
[0085] Combine Figure 3 As shown in (b), F sq The information compression formula is:
[0086]
[0087] Among them, i and j are the position indexes of the feature map in the height H and width W directions, U C (i, j) represents the eigenvalue of the Cth channel at the spatial position (i, j), F sq The function is used to compress the information of each channel into a single value through global average pooling, i.e. (C, W, H) → (C, 1, 1).
[0088] Combine Figure 3 As shown in (c), F ex The formula for generating the weight of (, V) is:
[0089] F ex (z,V)=σ(V2·δ(V1·z))=s.
[0090] Among them, δ is the ReLU activation function, σ is the Sigmoid activation function, V1 is the channel compression layer weight, V2 is the channel recovery layer weight, F ex The function is used to obtain the weight vector of each channel through two layers of 1×1 convolution, ReLU activation and Sigmoid activation.
[0091] Combine Figure 3 As shown in (d), F scale () The channel calibration formula is: G8 = F scale (U,s)=s⊙U,F scale Function is used to apply channel weights to F tr Output feature U to achieve channel recalibration. The larger the weight, the more significant the feature enhancement and the stronger the dominant effect on the output; the smaller the weight, the more obvious the feature suppression and the lower the contribution to the output; the final output feature map dimension remains unchanged, still (512, 20, 20), that is, the output
[0092] The above SPPF multi-scale feature fusion module is used to extract the context features of feature G9 through three series-connected maximum pooling layers, complete the multi-scale feature fusion, and output the feature
[0093] The effectiveness of the backbone network of the proton whistle wave cross-frequency recognition model provided by the embodiment of the present invention is verified. Figure 16 The output feature effect diagram of the Conv convolution module of the backbone network L0 layer is shown. Figure 17 shows the backbone network L1 to L 10 Output feature rendering of the layer-wise lightweight convolution module. Figure 18 The output feature effect diagram of the spatial pyramid pooling module is shown.
[0094] Through analysis Figure 16 、 17 , 18, we can see that the image features output by the Conv convolution module can clearly present a proton whistle wave event ( Figure 16 In the lightweight convolution module, the features output by each submodule are as follows Figure 17 As shown in Figure 2, submodules 1, 2, and 3 mainly retain the global structural features of proton whistler wave events at different levels; the remaining submodules emphasize local features in turn. The output features after entering the spatial pyramid pooling module are as follows: Figure 18 As shown, the details of the proton whistle wave event gradually become abstract, which helps process input images of different sizes and avoids the precision loss and computational redundancy caused by image scaling. This shows that the superposition of multiple layers of lightweight convolution can successfully separate low-frequency and high-frequency information in the image, thereby obtaining the outline of the proton whistle wave event. At the same time, the use of large 5×5 convolution kernels in submodules 7, 8, and 9, and the addition of the SE attention mechanism in submodules 8 and 9 can effectively provide a more fine-grained receptive field of the proton whistle wave event. With the help of the fusion of modules such as lightweight convolution, the effectiveness of the backbone network of the PWNet model provided by the embodiment of the present invention in efficiently extracting features is verified.
[0095] The multi-scale feature fusion network (Neck) connects the Backbone and Head modules. The multi-scale feature fusion network (Neck) inputs information of three feature dimensions from the backbone network (Backbone), which are The multi-scale feature fusion network (Neck) fuses feature maps from different layers through top-down and bottom-up information transfer.
[0096] The multi-scale feature fusion network (Neck) is different from the SPPF multi-scale feature fusion module in that the multi-scale feature fusion network (Neck) not only processes different scales, but also strengthens the fusion of low-level features and high-level features, providing rich contextual information to improve the accuracy of target detection, especially in the case of complex background or occlusion. The multi-scale feature fusion network (Neck) outputs feature information of three scales (80×80, 40×40, and 20×20) to the prediction classifier (Head) layer, which are output as follows:
[0097] The prediction classifier (Head) is the output module of the embodiment of the present invention, responsible for extracting the final target detection result from the feature map output by the multi-scale feature fusion network (Neck).
[0098] The network structure of the above prediction classifier (Head) is as follows Figure 4 As shown in the figure, it consists of three convolution modules of different scales. Each convolution module is divided into three sub-convolution modules in a decoupled form, namely the Bbox sub-convolution module, the Cls sub-convolution module and the Pose sub-convolution module, which are used to predict three different task categories respectively.
[0099] Among them, Bbox convolution is used to predict the proton whistle wave bounding box, Cls convolution is used to predict the feature category raw score, and Pose convolution is used to predict the proton whistle wave cross-frequency coordinates and visibility confidence.
[0100] Figure 5 The schematic diagram shows the recognition principle of the three seed convolution modules for the proton whistle wave and its cross-frequency characteristics with an output scale of 40×40. The feature input of the current scale is The task categories are decoupled into Bbox, Cls, and Pose, sharing the same feature information. To facilitate subsequent unified calculations and simplify the model structure, the Bbox and Pose sub-convolution modules of the three different scale modules are compressed to 64 channels after the CBS*2 convolution module, and the Cls sub-convolution module is compressed to 256 channels. After Conv2d convolution, the feature dimensions of the three sub-module structures are output respectively.
[0101] Specifically, each Bbox coordinate (x, y, w, h) in the Bbox subconvolution module corresponds to reg_max (default is 16) probability distribution intervals, then the output feature dimension is (4reg_max, 40, 40), where the number of channels is 4reg_max: 64. Therefore, the structural feature output of the Bbox subconvolution module is expressed as
[0102] Specifically, the feature dimension of the Cls subconvolution module output is (nc, 40, 40), where the number of channels is nc:1, that is, Cls:PW. Therefore, the structural feature output of the Cls subconvolution module is expressed as
[0103] Specifically, the feature dimension of the Pose sub-convolution module output is (3nk, 40, 40), where the number of channels is 3nk, nk is the number of crossover frequencies, and 3 refers to the x, y coordinates of the crossover frequency and the visibility confidence. Therefore, the structural feature output of the Pose sub-convolution module is expressed as
[0104] In summary, the feature output of the proton whistle wave cross-frequency recognition model provided by the embodiment of the present invention comes from the prediction classifier (Head), which contains feature information of three task categories, and these three task categories have different output scales. When the output scale is 40×40, the feature output is
[0105] The output logic of the other two scales is the same as this scale.
[0106] The loss function of the proton whistle wave cross-frequency recognition model provided by the embodiment of the present invention is a complex multi-component function used to measure the difference between the model prediction result and the true label.
[0107] The loss function of the embodiment of the present invention mainly consists of three parts, namely classification (Cls) loss, regression (Bbox) loss and cross-frequency (Pose) loss. The regression loss includes CloU loss and DFL loss, and the key point (cross-frequency) loss includes cross-frequency position loss and cross-frequency visibility loss. There are a total of 5 loss functions.
[0108] During the model training process, the overall loss function of the lightweight proton whistle wave cross-frequency recognition model is the weighted sum of the above five types of losses, namely:
[0109]
[0110] Among them, the classification loss term L Cls The binary cross-entropy (BCE) loss is used to measure the difference between the predicted probability distribution of the proton whistle wave cross-frequency recognition model PWNet and the true PW label distribution. The smaller the difference, the lower the loss value. Cls is the loss term L Cls The corresponding weight coefficient, L Cls The formula is as follows:
[0111] L Cls =-[y c ·log(σ(p c ))+(1-y c )·log(1-σ(p c ))]
[0112] Where c is the PW category; σ is the sigmoid activation function; y c is the onehot encoding of the true label (0 or 1); p c is the raw (unnormalized) prediction for class c.
[0113] Loss term L CloU (Complete IoU Loss) is used to detect the degree of overlap between the predicted bounding box of the proton whistle wave and the true bounding box, while considering the center distance and aspect ratio of the proton whistle wave, λ CloU is the loss term L CloU The corresponding weight coefficient, L CloU The formula is as follows:
[0114]
[0115] Among them, A and B refer to the areas of the PW predicted bounding box and the true bounding box respectively; b and b gt are the center coordinates of the PW predicted box and the PW true box respectively; ρ is the Euclidean distance between the two points; c is the diagonal length of the minimum bounding rectangle containing its predicted box and the true box; α is a balance coefficient; v is a term related to the aspect ratio.
[0116] Loss term L DFL For the bounding box task of proton whistle wave detection, since the time-frequency graph data exists in a complex background, the proton whistle wave may have blurred and uncertain boundaries. Based on this situation, the DFL loss is used. The DFL loss optimizes the position prediction accuracy of the coordinate proton whistle wave through discrete probability distribution, thereby enhancing the learning ability of the proton whistle wave cross-frequency recognition model for samples with unclear proton whistle wave features. DFL is the loss term L DFL The corresponding weight coefficient, L DFL The formula is as follows:
[0117]
[0118] Among them, y is the position label value of the cross frequency; y i and y i+1 They are the adjacent discrete positions of y, corresponding to and It is y i and y i+1 The confidence of the position, this loss term makes the model more focused on learning the characteristics of proton whistle waves by adjusting the importance of samples.
[0119] Loss term L Pose The crossover frequency position loss is the mean square error (MSE), which is used to measure the difference between the predicted position and the true position of the proton whistle wave crossover frequency. Pose is the loss term L Pose The corresponding weight coefficient, L Pose The formula is as follows:
[0120]
[0121] Where N is the total number of key points, and k represents the kth crossover frequency. (x pred ,y pred ) is the predicted coordinate of the crossover frequency; (x true ,y true ) is the real coordinate of the crossover frequency.
[0122] Loss Item is the cross-frequency visibility loss, which is used to optimize the model's learning ability for cross-frequency difficult samples. The calculation formula is: Among them, y is the true label indicating whether the cross-frequency point exists, The probability of the crossover frequency point predicted by the model exists. The TopK selection function indicates the selection of the first K crossover frequency points with the largest loss value. For L Kobj , The TopK cross-frequency points with the largest loss will be selected for back propagation.
[0123] The traditional visibility loss uses binary cross entropy (BCE) loss to predict the existence of each cross frequency. In order to more effectively grasp the positioning accuracy of the cross frequency, the embodiment of the present invention modifies the cross frequency visibility loss and introduces LRM Loss to mine difficult samples to improve the performance of the lightweight proton whistle wave cross frequency recognition model in cross frequency detection. The traditional loss function will perform a weighted average or summation on the loss of each sample containing a cross frequency, while the LRM of the embodiment of the present invention will ignore easy samples (samples with smaller losses) and focus on difficult samples (samples with larger losses). Figure 6 The figure shows the principle diagram of the LRM loss function. First, the cross-frequency binary cross entropy (BCE) loss is calculated for each sample. Then, LRM is used to perform TopK selection, that is, to select the cross-frequency samples with the largest loss for optimization. LRM ignores the samples with the smallest loss (i.e., simple samples that are easy to classify), calculates the loss of the selected samples, and returns the weighted loss, thereby optimizing the model's learning ability for cross-frequency difficult samples.
[0124] The LAMP pruning module is based on layer adaptive amplitude pruning, which is used to remove filters or channels that contribute less to the model output, thereby further reducing model complexity, improving inference efficiency, and reducing memory usage.
[0125] The LAMP pruning module scores the weights of each channel in the model's convolutional layer, determines the channels to be removed based on the weight scores of each channel, reduces the model complexity, and constructs an efficient proton whistle wave cross-frequency recognition module.
[0126] The scoring formula for the above LAMP pruning module is: Where W[u] represents the weight value of the u-th position in the weight tensor W (i.e., the weight size of the position); (W[u]) 2 is the square of the weight, indicating the strength of the weight; ∑ v≥u (W[v]) 2 Represents the sum of squares of all weights connected after the current weight position u, which is used to measure the strength of other weights in the same layer.
[0127] The decision formula of the above LAMP pruning module is:
[0128]
[0129] If one weight W[u] is larger than another weight W[v], then score(u; W) will be larger than score(v; W), which means W[u] is more important and will be retained. During pruning, weights with lower LAMP scores are pruned, while weights with higher LAMP scores are retained.
[0130] Figure 7 A schematic diagram of the LAMP pruning module is shown, where Ci1, Ci2, and so on represent different channels in the i-th convolutional layer. In the original network (left), all channels in the convolutional layer have a corresponding scaling factor. The weight of each channel, or the size of the scaling factor, determines which channels are pruned. The pruning process shown in the figure removes channels with small scaling factors (such as 0.01, 0.08, and 0.02), leaving only those channels that have a significant impact on the network.
[0131] Figure 19 The figure shows the comparison of the number of model channels before and after LAMP pruning. Figure 19 The figure shows the change in the number of channels in each layer of the lightweight model (base) and the LAMP pruned model (prune). The yellow bar chart (base) represents the number of channels in each layer of the original model, and the red bar chart (prune) represents the number of channels in each layer of the pruned model. The horizontal axis is the channel number of the model. Pruning usually removes some redundant or unimportant channels. Therefore, in many layers, the number of channels after pruning is reduced compared to the original model. The figure shows the number of channels in each layer. The number of channels in most layers has been significantly reduced after pruning. Overall, the number of channels has decreased after pruning, indicating that the complexity and number of parameters of the model have been reduced through pruning, resulting in higher computational efficiency and lower storage requirements.
[0132] Corresponding to the above-disclosed proton whistle wave cross-frequency recognition model, an embodiment of the present invention further discloses a proton whistle wave cross-frequency recognition method. The following describes in detail a proton whistle wave cross-frequency recognition method disclosed in an embodiment of the present invention in conjunction with the above-described proton whistle wave cross-frequency recognition model.
[0133] The following describes the specific steps of a proton whistle wave cross-frequency identification method provided by an embodiment of the present invention.
[0134] Obtain electric field load data collected by the Zhangheng satellite;
[0135] The electric field load data is preprocessed to obtain the preprocessed time-frequency data, including: firstly intercepting the electric field load data using a 4s sliding time window to obtain the waveform data b of the component d ∈R n (4 seconds of ELF band data, n = 20480), Figure 8 A schematic diagram of the waveform data of the proton whistle wave event from 03:20:21 to 03:20:24 on February 10, 2020 is shown.
[0136] Then, the waveform data segments are processed by overlapping short-time Fourier transform (STFT) to convert the waveform data segments into corresponding time-frequency graphs. The calculation formula is:
[0137]
[0138] Among them, X[m,n] represents the energy of the data at the frequency point n corresponding to time m, b d [k] represents the kth sampling point in the input signal, L is the length of the input signal, and g[.] represents the window function.
[0139] Figure 9 A schematic diagram of the time-frequency data of the proton whistle wave event from 03:20:21 to 03:20:24 on February 10, 2020 is shown.
[0140] The pre-processed time-frequency data is input into a proton whistle wave cross-frequency recognition model as described above to obtain a proton whistle wave recognition result and a proton whistle wave cross-frequency recognition result corresponding to the time-frequency data.
[0141] The proton whistle wave cross-frequency recognition results were evaluated using the following evaluation indicators:
[0142] Precision is defined as follows: TP (TruePositives, TP) represents the number of accurately identified crossover frequencies; FP (FalsePositives, FP) represents the number of incorrectly identified crossover frequencies.
[0143] Recall is defined as follows: Wherein, FN (False Negatives, FN) represents the number of missed crossover frequencies.
[0144] mAP50 (Mean Average Precision at IoU=0.5) is the average precision (AP) when the IoU threshold is 0.5. Its calculation process includes:
[0145] Step 1: Calculate the key point similarity (OKS). OKS measures the similarity between the predicted crossover frequency and the actual crossover frequency (the default threshold is 0.5). The calculation formula is as follows:
[0146]
[0147] Among them, d i is the Euclidean distance of the i-th crossover frequency; δ(v i ) is an indicator function used to consider whether the cross-frequency true label is visible; (x pred ,y pred ) is the coordinate of the predicted key point; (x gt ,y gt ) are the coordinates of the true keypoint; σ is the area scale of the target. The OKS of each crossover frequency is compared with its threshold to determine the authenticity of the recognition result. If OKS ≥ the threshold of 0.5, it is TP, otherwise it is FP. A true keypoint that is not detected is FN.
[0148] Step 2: Draw the precision-recall (PR) curve. The precision-recall (PR) curve shows the relationship between precision (Precision) and recall (Recall). By calculating the area under the PR curve, the average precision (AP) of the cross-frequency can be obtained. Figure 11 The PR curve of proton whistle wave crossover frequency detection is shown. Figure 11 It can be seen that the proton whistle wave cross-frequency recognition model provided by the embodiment of the present invention still maintains a high precision under a high recall rate, indicating that the model has good balanced performance in the cross-frequency detection task.
[0149] Step 3: Calculate mAP50 (Mean Average Precision at IoU threshold 0.5). For tasks with multiple cross-frequency values, first calculate the AP value for each cross-frequency value, then take the average of these AP values to get mAP50. The calculation formula is:
[0150]
[0151] Among them, r represents the recall rate (Recall); p (r) is the precision under a given recall rate r; N is the total number of cross-frequency; AP i is the AP value of the i-th cross frequency, calculated when the IoU threshold is 0.5.
[0152] mAP50-95 (Mean Average Precision at IoU thresholds from 0.5 to 0.95) represents the average of the average precision (AP) calculated at multiple IoU thresholds (from 0.5 to 0.95, with a step size of 0.05); it can better evaluate the overall performance of the PWNet model under different conditions.
[0153] Parameters refer to the total number of trainable parameters in the model, including convolution kernels, fully connected layer weights, biases, etc. The fewer parameters, the lighter the model and the lower the computational overhead.
[0154] GFLOPs is a metric that measures the computational complexity of a model. It represents the number of floating-point operations performed by the model during inference. The calculation of GFLOPs is related to the model's structure (number of layers, convolution kernel size, number of channels), the size of the input data, and the batch size. Lower GFLOPs indicate a lighter model and faster processing speed.
[0155] Size refers to the memory or disk space occupied by the model. A smaller model requires less storage space when deployed, making it suitable for resource-limited environments such as embedded devices and mobile devices.
[0156] Based on the above evaluation indicators, the proton whistle wave cross frequency recognition model PWNet provided by the embodiment of the present invention is used to quantitatively evaluate the parameter quantity, computational complexity, model size, and accuracy of identifying the proton whistle wave cross frequency, along with the benchmark model Yolov8, the model Yolov-PP after lightweight convolution of the benchmark model, and the model Yolov-PLAMP after pruning Yolov-PP. The quantitative evaluation results are shown in Table 1:
[0157] Table 1 Quantitative evaluation results of PWNet model
[0158]
[0159] Analysis of Table 1 shows that the PWNet model of the present invention significantly reduces parameters, GFLOPs, and model size compared to Yolov8n, Yolov-PP, and Yolov-PLAMP. This achieves lightweight model parameters and accelerates model inference speed, accelerating inference by 3.03 times and reducing computational redundancy by 68.2%. Furthermore, the PWNet model of the present invention does not significantly reduce model accuracy, with a loss of only 0.006 in cross-frequency detection mAP50. Other evaluation metrics also show minimal degradation.
[0160] In summary, the proton whistle wave cross-frequency recognition model and method provided by the embodiments of the present invention not only lightweights the model's convolutional module, but also introduces local regression modulation loss (LRM Loss) to enhance the sub-pixel positioning accuracy of the cross-frequency and further lightweights the model structure using LAMP pruning technology. The proton whistle wave cross-frequency recognition model PWNet provided by the embodiments of the present invention achieves negligible performance loss while having only 31.7% of the parameters of the baseline model and 34.2% of the original storage size, accelerating the inference speed of automated proton whistle wave cross-frequency recognition.
[0161] Figure 12 One of the diagrams showing the Yolov8 model's recognition effect on the crossover frequency of proton whistle waves. Figure 13 One of the recognition effect diagrams of the proton whistle wave cross-frequency recognition PWNet model provided by an embodiment of the present invention is shown.
[0162] Figure 14 The second figure shows the recognition effect of the Yolov8 model on the crossover frequency of proton whistle waves. Figure 15 The second recognition effect diagram of the proton whistle wave cross-frequency recognition PWNet model provided by an embodiment of the present invention is shown.
[0163] Through comparative analysis Figures 12 to 15 It can be seen that the PWNet model of the embodiment of the present invention has no loss of cross-frequency recognition accuracy compared to the Yolov8 model, and the confidence of the five proton whistle wave detection frames only decreases by an average of 1.4%. Therefore, the proton whistle wave cross-frequency recognition model provided by the embodiment of the present invention basically achieves lossless accuracy on the basis of reducing the model calculation amount.
[0164] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A proton whistle wave cross-frequency recognition model, characterized in that: The model includes: The backbone network is used to extract the proton whistle wave characteristics and the cross-frequency characteristics of the proton whistle wave from the electric field time-frequency data collected by the Zhangheng satellite; The backbone network includes a CBS convolution module, a DepthSepConv depth-separable convolution module, a SE lightweight attention mechanism module, and an SPPF multi-scale feature fusion module; The multi-scale feature fusion network is set between the backbone network and the prediction classifier, and is used to fuse the different levels of feature maps output by the backbone network to output multi-scale and multi-level feature maps; The prediction classifier is used to extract the proton whistle wave detection results and the proton whistle wave cross-frequency detection results from the multi-scale and multi-level feature map; The LAMP pruning module is used to score the weights of each channel in the model's convolutional layer and determine the channels to be removed based on the weight scores of each channel to reduce the complexity of the model.
2. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The CBS convolution module consists of a Conv convolution layer, a batch normalization layer and a ReLU activation function; The calculation process of the CBS convolution module is: G0 = ReLU (BN (s k ·F w,h +b k )), for w,h∈[1,320],k∈[1,16], where S k is the kth 3×3 convolution kernel of the Conv convolution layer, F w,h is the eigenvector of the input time-frequency graph X[m,n] at position (w,h), b k is the bias term of the kth output channel, BN() is batch normalization, ReLU() is the activation function, and G0 is the output result of the CBS convolution module.
3. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The calculation process of the DepthSepConv depth separable convolution module is: G i =φ i,p (φ i,d (G i-1 ), where Gi is the output of the i-th DepthSepConv depth-separable convolution module, φ i,p With φ i,d They are the point-by-point convolution and depth-wise convolution of the i-th DepthSepConv depth-wise separable convolution module respectively; φ d It is used to increase the dimension of the features of a stage, and improve the expression ability of time-frequency features by increasing the number of channels or dimensions of the features. p It is used to reduce the dimensionality of features, aggregate important features together, and filter out minor features.
4. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The SE lightweight attention mechanism module is set at the end of the DepthSepConv depth separable convolution module; The SE lightweight attention mechanism module is used to weight channels in the network, thereby enhancing important features and suppressing irrelevant information; The calculation process of the SE lightweight attention mechanism module is: G i =F scale (F ex (F sq (F tr (G i-1 )),V)⊙U), where G i is the output result of the SE lightweight attention mechanism module, G i-1 is the feature input of the DepthSepConv depth-separable convolution module, F tr To extract the image feature function, F sq is the channel compression function, F ex (, V) is the channel excitation function, V is the weight matrix, F scale () is the channel scaling function, ⊙ is the element-by-element multiplication symbol, and U is F tr Feature input after feature extraction, U=F tr (G i-1 ); Among them, i and j are the position indexes of the feature map in the height H and width W directions, U C (i, j) represents the eigenvalue of the Cth channel at the spatial position (i, j), F sq The function is used to compress the information of each channel into a single value through global average pooling; F ex (z,V)=σ(V2·δ(V1·z))=s, where δ is the ReLU activation function, σ is the Sigmoid activation function, V1 is the channel compression layer weight, V2 is the channel recovery layer weight, and F ex The function is used to obtain the weight vector of each channel through two layers of 1×1 convolution, ReLU activation and Sigmoid activation; G i =F scale (U,s)=s⊙U,F scale Function is used to apply channel weights to F tr Output feature U to achieve channel recalibration.
5. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The backbone network includes a CBS convolution module, nine DepthSepConv depth-separable convolution modules and an SPPF multi-scale feature fusion module connected in sequence; Among them, the SE lightweight attention mechanism module is provided at the end of the eighth DepthSepConv depth-separable convolution module, and the SE lightweight attention mechanism module is provided at the end of the ninth DepthSepConv depth-separable convolution module.
6. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The prediction classifier is composed of comprehensive convolution modules of different scales, wherein the comprehensive convolution module is composed of a Bbox sub-convolution module, a Cls sub-convolution module and a Pose sub-convolution module; The Bbox sub-convolution module is used to predict the proton whistle wave bounding box, the Cls sub-convolution module is used to predict the feature category, and the Pose sub-convolution module is used to predict the proton whistle wave cross-frequency coordinates and visibility confidence.
7. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The loss function of the proton whistle wave cross-frequency recognition model is: Loss term L Cls is the binary cross entropy loss, which is used to measure the difference between the predicted probability distribution of the proton whistle wave cross frequency recognition model and the true label distribution, λ Cls is the loss term L Cls The corresponding weight coefficient; Loss term L CloU The degree of overlap between the predicted bounding box and the true bounding box for detecting proton whistler waves, λ CloU is the loss term L CloU The corresponding weight coefficient; Loss term L DFL It is used to optimize the position prediction accuracy of the coordinate proton whistle wave and enhance the learning ability of the proton whistle wave cross-frequency recognition model for samples with unclear proton whistle wave characteristics. DFL is the loss term L DFL The corresponding weight coefficient; Loss term L Pose is the cross-frequency position loss, which is used to measure the difference between the predicted position and the true position of the proton whistle wave cross-frequency, λ Pose is the loss term L Pose The corresponding weight coefficient; Loss Item is the cross-frequency visibility loss, which is used to optimize the model's learning ability for cross-frequency difficult samples. The calculation formula is: Among them, y is the true label indicating whether the cross-frequency point exists, is the probability of the existence of the cross-frequency point predicted by the model, and the TopK selection function indicates the selection of the top K cross-frequency points with the largest loss value.
8. A proton whistle wave cross-frequency recognition model according to claim 1, characterized in that: The scoring formula of the LAMP pruning module is: Where W[u] represents the weight value at the u-th position in the weight tensor W, ∑ v≥u (W[v]) 2 Represents the sum of all squared weights connected after the current weight position u.
9. A proton whistle wave crossover frequency identification method for proton whistle wave crossover frequency identification, characterized in that: The method comprises: Obtain electric field load data collected by the Zhangheng satellite; performing data preprocessing on the electric field load data to obtain preprocessed time-frequency data; The time-frequency data is input into a proton whistle wave cross-frequency recognition model according to any one of claims 1 to 8 to obtain a proton whistle wave recognition result and a proton whistle wave cross-frequency recognition result corresponding to the time-frequency data.
10. A proton whistle wave cross-frequency identification method according to claim 9, characterized in that: Performing data preprocessing on the electric field load data to obtain preprocessed time-frequency data includes: intercepting the electric field load data using a preset sliding time window to obtain waveform data segments; Overlapping short-time Fourier transform processing is performed on the waveform data segment to obtain a time-frequency diagram corresponding to the waveform data segment.