Online learning method and device for brain-like electromagnetic model

CN122548355APending Publication Date: 2026-08-11TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-06
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明提供一种类脑电磁模型的在线学习方法和装置,用以解决现有技术中用于进行电磁信号识别的大模型,难以在计算与存储资源受限的边缘设备上实现高效部署与在线学习的缺陷,实现电磁信号识别的大模型进行在线学习时,能够降低计算量、耗时以及对硬件资源的要求,在计算与存储资源受限的边缘设备上实现高效部署与在线学习

Benefits of technology

[0020] The present invention provides an online learning method and apparatus for a neuromorphic electromagnetic model. This method involves acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category. The negative electromagnetic signal samples are samples with different signal categories than the positive electromagnetic signal samples, and/or samples generated after perturbing the positive electromagnetic signal samples. The positive electromagnetic signal samples are input into a slow system network to obtain a first feature vector output by each feature extraction layer. The negative electromagnetic signal samples are also input into the slow system network to obtain a second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain a first modulation feature output by the feature modulation layer. The second feature vector output by the same feature extraction layer is also input into the feature modulation layer to obtain a second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first and second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning of a fast system network, which is used to identify the signal category of the electromagnetic signal. Because it employs a local learning mechanism without backpropagation, it avoids the cross-layer gradient calculations and storage of a large number of intermediate activation values ​​required by traditional backpropagation algorithms. This reduces computational load, time consumption, and hardware resource requirements when learning large models for electromagnetic signal recognition online, enabling efficient deployment and online learning on edge devices with limited computing and storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548355A_ABST
    Figure CN122548355A_ABST
Patent Text Reader

Abstract

This invention provides an online learning method and apparatus for a neuromorphic electromagnetic model, comprising: acquiring negative and positive electromagnetic signal samples; inputting the positive electromagnetic signal samples into a slow system network to obtain a first feature vector output by each feature extraction layer, and inputting the negative electromagnetic signal samples into the slow system network to obtain a second feature vector output by each feature extraction layer; for each feature modulation layer, inputting the first feature vector output by the corresponding feature extraction layer into the feature modulation layer to obtain a first modulation feature, and inputting the second feature vector output by the same feature extraction layer into the feature modulation layer to obtain a second modulation feature; determining a first goodness and a second goodness based on the first and second modulation features, determining a local loss based on the first and second goodness; and updating the network parameters of the feature modulation layer based on the local loss. This invention can reduce the computational load, time consumption, and hardware resource requirements when performing online learning of a neuromorphic electromagnetic model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an online learning method and apparatus for a brain-like electromagnetic model. Background Technology

[0002] With the rapid development of deep learning technology, large-scale neural network models (large models) have demonstrated powerful feature extraction and recognition capabilities in the field of electromagnetic signal processing. Existing large models typically employ backpropagation (BP) for training and fine-tuning, calculating the gradient of the loss function with respect to the parameters of each layer and updating the network weights using stochastic gradient descent (SGD) or its variants. In electromagnetic signal processing, one-dimensional radio frequency signals are usually converted into two-dimensional time-frequency images, and then pre-trained deep neural networks are used for transfer learning or fine-tuning to adapt to specific electromagnetic signal recognition tasks.

[0003] However, existing fine-tuning methods based on backpropagation are computationally intensive and time-consuming when dealing with large models with a huge number of parameters. They also require storing intermediate activation values, which places extremely high demands on hardware resources, making it difficult to achieve efficient deployment and online learning on edge devices with limited computing and storage resources. Summary of the Invention

[0004] This invention provides an online learning method and apparatus for neuromorphic electromagnetic models, which addresses the shortcomings of existing technologies for large models used in electromagnetic signal recognition, which are difficult to deploy and learn efficiently on edge devices with limited computing and storage resources. When large models for electromagnetic signal recognition are used for online learning, the computational load, time consumption, and hardware resource requirements can be reduced, enabling efficient deployment and online learning on edge devices with limited computing and storage resources.

[0005] This invention provides an online learning method for a neuromorphic electromagnetic model, wherein the neuromorphic electromagnetic model includes a fast system network and a slow system network, the fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer, wherein each feature modulation layer and each feature extraction layer corresponds one-to-one; the method includes: Acquire multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; The multiple positive electromagnetic signal samples are input into the slow system network to obtain the first feature vector output by each feature extraction layer, and the multiple negative electromagnetic signal samples are input into the slow system network to obtain the second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and the second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. A first quality is determined based on the first modulation feature, a second quality is determined based on the second modulation feature, and the local loss of the feature modulation layer is determined based on the first quality and the second quality. The network parameters of the feature modulation layer are updated based on the local loss to enable online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.

[0006] According to an online learning method for a neuromorphic electromagnetic model provided by the present invention, determining the local loss of the feature modulation layer based on the first goodness and the second goodness includes: Determine the first according to formula (1) or formula (2). Local loss of layer feature modulation layer : (1) (2) in, This indicates the first level of excellence. This indicates the second degree of excellence. Indicates the temperature coefficient. This indicates the target quality of the positive sample of the electromagnetic signal. This indicates the target quality of the negative sample of the electromagnetic signal.

[0007] According to the online learning method for a neuromorphic electromagnetic model provided by the present invention, the step of inputting the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer includes: The modulation features output from the previous feature modulation layer are input into the feature modulation layer, and the modulation vector is determined by the modulation generator in the feature modulation layer. Based on the modulation vector, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is modulated to obtain the first modulation feature.

[0008] According to the online learning method for a neuromorphic electromagnetic model provided by the present invention, the step of acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category includes: The plurality of electromagnetic signal negative samples are obtained by sampling from the negative sample pool; wherein, the negative sample pool includes electromagnetic signal samples of different signal categories, and / or electromagnetic signal samples generated by perturbation; Multiple electromagnetic signal positive samples are obtained by sampling from the positive sample pool corresponding to the target signal category.

[0009] According to the online learning method for a neuromorphic electromagnetic model provided by the present invention, the step of sampling multiple electromagnetic signal positive samples from the positive sample pool corresponding to the target signal category includes: If the number of electromagnetic signal samples in the positive sample pool corresponding to the target signal category is greater than or equal to a first preset threshold, multiple electromagnetic signal positive samples are sampled from the positive sample pool corresponding to the target signal category.

[0010] According to the present invention, an online learning method for a neuromorphic electromagnetic model is provided, the method further includes: The acquired electromagnetic signal samples are input into the slow system network to obtain the target features output by the slow system network; Determine the similarity between the target feature and each sample feature in the feature cache, wherein the sample features are features of stored electromagnetic signal samples; The similarities are sorted in descending order, and based on the first preset number of similarities, pseudo-labels and pseudo-label confidence scores of the electromagnetic signal samples are generated by weighted voting. If the confidence level of the pseudo-label is greater than or equal to the first preset confidence level, the electromagnetic signal sample and the target feature are added to the positive sample pool corresponding to the signal category represented by the pseudo-label.

[0011] According to the present invention, an online learning method for a neuromorphic electromagnetic model is provided, the method further includes: If the confidence level of the pseudo-label is less than the first preset confidence level, the electromagnetic signal sample and the target feature are added to the new category candidate sample pool. If the number of samples in the new category candidate sample pool is greater than or equal to a second preset threshold, the target features in the new category candidate sample pool are clustered. If at least one cluster is successfully clustered, a new positive sample pool for the new signal category corresponding to each cluster is created, and the candidate electromagnetic signal samples belonging to the cluster and the corresponding target features are stored in the new positive sample pool.

[0012] According to the present invention, an online learning method for a neuromorphic electromagnetic model is provided, the method further includes: If the number of electromagnetic signal samples in the target positive sample pool reaches a third preset threshold, the target positive sample pool is updated using a first-in-first-out (FIFO) method, or electromagnetic signal samples in the target positive sample pool with pseudo-label confidence levels lower than a second preset confidence level are deleted.

[0013] According to the present invention, an online learning method for a neuromorphic electromagnetic model is provided, the method further includes: The negative sample pool is constructed based on at least one of the following methods: Electromagnetic signal samples are randomly drawn from the positive sample pool corresponding to multiple signal categories, and the drawn electromagnetic signal samples of different signal categories are stored in the negative sample pool. The positive electromagnetic signal samples in the positive sample pool are perturbed to generate negative electromagnetic signal samples, and the generated negative electromagnetic signal samples are stored in the negative sample pool. Electromagnetic signal negative samples are obtained by randomly sampling from the environmental background signal and stored in the negative sample pool. Electromagnetic signal samples with a false label confidence level lower than the first preset confidence level are stored in the negative sample pool.

[0014] According to the present invention, an online learning method for a neuromorphic electromagnetic model, wherein updating the network parameters of the feature modulation layer based on the local loss to perform online learning of the fast system network includes: The modulation generator in the currently trained feature modulation layer is used as a policy network. The modulation vector is output according to the current state of the feature modulation layer. The current state includes at least one of the following: the first feature vector and the second feature vector output by the feature extraction layer corresponding to the feature modulation layer, the modulation feature output by the previous feature modulation layer, and the pseudo-label confidence of the current electromagnetic signal sample. Based on the multiple negative electromagnetic signal samples and the multiple positive electromagnetic signal samples, a reward signal for the reinforcement learning process is determined. This reward signal includes a main reward, an auxiliary reward, and a sparse reward. The main reward is determined based on the proportion of correctly classified samples and / or misclassified samples among the multiple positive electromagnetic signal samples, or based on the recognition confidence of the multiple positive electromagnetic signal samples after modulation by the modulation vector. The auxiliary reward is determined based on a first excellence score and a second excellence score. The sparse reward is a reward configured when a new category is successfully identified or a preset task is completed. The reinforcement learning process is used to characterize the parameter update process of the currently trained feature modulation layer. Based on the reward signal, a policy loss is constructed, and based on the local loss and the policy loss, the parameters of the policy network are updated to obtain the updated feature modulation layer network parameters, so as to perform online learning on the fast system network.

[0015] According to the present invention, an online learning method for a neuromorphic electromagnetic model is provided, the method further includes: At preset intervals, the slow system network is subjected to self-supervised learning based on multiple unlabeled sample features in the feature cache. During the self-supervised learning process of the slow system network, the fast system network keeps its current parameters unchanged. The feature cache includes sample features of multiple electromagnetic signal samples.

[0016] This invention also provides an online learning device for a neuromorphic electromagnetic model, the neuromorphic electromagnetic model comprising a fast system network and a slow system network, the fast system network comprising at least one feature modulation layer, and the slow system network comprising at least one feature extraction layer, wherein each feature modulation layer and each feature extraction layer corresponds one-to-one; the device comprises: The acquisition module is used to acquire multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; The input module is used to input the multiple positive electromagnetic signal samples into the slow system network to obtain the first feature vector output by each feature extraction layer, and to input the multiple negative electromagnetic signal samples into the slow system network to obtain the second feature vector output by each feature extraction layer; The input module is further configured to input the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and input the second feature vector output by the same feature extraction layer into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. The determination module is configured to determine a first quality level based on the first modulation feature, determine a second quality level based on the second modulation feature, and determine the local loss of the feature modulation layer based on the first quality level and the second quality level. An update module is used to update the network parameters of the feature modulation layer based on the local loss, so as to perform online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.

[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement an online learning method for a neuromorphic electromagnetic model as described above.

[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an online learning method for a neuromorphic electromagnetic model as described above.

[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements an online learning method for a neuromorphic electromagnetic model as described above.

[0020] The present invention provides an online learning method and apparatus for a neuromorphic electromagnetic model. This method involves acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category. The negative electromagnetic signal samples are samples with different signal categories than the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples. The positive electromagnetic signal samples are input into a slow system network to obtain a first feature vector output by each feature extraction layer. The negative electromagnetic signal samples are also input into the slow system network to obtain a second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain a first modulation feature output by the feature modulation layer. The second feature vector output by the same feature extraction layer is also input into the feature modulation layer to obtain a second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first and second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning of a fast system network, which is used to identify the signal category of the electromagnetic signal. Because it employs a local learning mechanism without backpropagation, it avoids the cross-layer gradient calculations and storage of a large number of intermediate activation values ​​required by traditional backpropagation algorithms. This reduces computational load, time consumption, and hardware resource requirements when learning large models for electromagnetic signal recognition online, enabling efficient deployment and online learning on edge devices with limited computing and storage resources. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the structure of a brain-like electromagnetic model provided in an embodiment of the present invention.

[0023] Figure 2 This is one of the flowcharts illustrating the online learning method for a neuromorphic electromagnetic model provided in an embodiment of the present invention.

[0024] Figure 3 The second flowchart illustrates the online learning method for a neuromorphic electromagnetic model provided in this embodiment of the invention.

[0025] Figure 4 The third flowchart illustrates the online learning method for a brain-like electromagnetic model provided in this embodiment of the invention.

[0026] Figure 5 This is a schematic diagram of the structure of an online learning device for a neuromorphic electromagnetic model provided in an embodiment of the present invention.

[0027] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0029] Currently, in the field of electromagnetic signal processing, one-dimensional radio frequency signals are typically converted into two-dimensional time-frequency images, and pre-trained large models based on the backpropagation (BP) algorithm (such as ResNet and VisionTransformer) are used for transfer learning or fine-tuning to complete specific electromagnetic signal recognition tasks. However, fine-tuning methods based on backpropagation have significant limitations when dealing with large models with a huge number of parameters. Traditional backpropagation algorithms require calculating the gradient of the loss function with respect to the parameters of each layer after forward propagation and then backpropagating the error signal layer by layer. For large models with a huge number of parameters, this process is not only computationally intensive and time-consuming, but also requires storing intermediate activation values, placing extremely high demands on hardware resources (especially GPU memory), making it difficult to achieve efficient deployment and online learning on edge devices with limited computing and storage resources.

[0030] In view of the above-mentioned problems, this invention proposes an online learning method and apparatus for a neuromorphic electromagnetic model. In this method, a forward-forward (FF) algorithm can be used to update the network parameters of each layer of the fast system network without relying on the error gradients from other network layers in the neuromorphic electromagnetic model. This enables layer-by-layer independent optimization of the fast system network, thereby reducing the computational load, time consumption, and hardware resource requirements when learning large models for electromagnetic signal recognition online. It also enables efficient deployment and online learning on edge devices with limited computing and storage resources.

[0031] The following is combined Figures 1 to 4 The online learning method for the neuromorphic electromagnetic model provided in this invention is described. This invention can be applied to scenarios requiring real-time performance, adaptability, and computational efficiency in open dynamic electromagnetic signal processing, such as UAV / unmanned vehicle-mounted signal detection platforms, adaptive interference and anti-interference, and spectrum monitoring and illegal signal detection.

[0032] The subject executing this method can be an electronic device such as a terminal device, computer, server, server cluster, or specially designed online learning device for neuromorphic electromagnetic models. It can also be an online learning device for neuromorphic electromagnetic models installed in such electronic devices. The online learning device for neuromorphic electromagnetic models can be implemented through software, hardware, or a combination of both.

[0033] Figure 1 This is a schematic diagram of the structure of the neuromorphic electromagnetic model provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this brain-like electromagnetic model has a dual-system architecture, including a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, such as L feature modulation layers. The slow system network includes at least one feature extraction layer, such as L feature extraction layers. Each feature modulation layer and each feature extraction layer corresponds one-to-one.

[0034] The Slow System network employs a deep neural network as its backbone, such as ResNet50, ResNet101, or Vision Transformer (ViT). The Slow System network learns general, task-independent feature representations through self-supervised learning on a large number of unlabeled electromagnetic signals. The loss function of the self-supervised learning objective is defined as shown in equation (3): (3) in, This represents the cross-correlation matrix of two enhanced view features. These are the tradeoff coefficients. All learnable parameters of a slow system network are denoted as... These learnable parameters are updated only in the background through self-supervised learning. It does not participate in online fine-tuning to ensure long-term knowledge stability. The output of the slow system network includes feature vectors h1, h2...h1 of each layer. L And the global feature vector z, where L represents the number of network layers, and the global feature vector z is usually the feature after global average pooling.

[0035] Fast System Network (FSN) is composed of The network consists of stacked feature modulation layers, each corresponding to a feature extraction layer in the slow system network. The first layer is... The structure of the layer feature modulation layer is as follows: Modulation generator A lightweight subnetwork whose input is the modulated output of the previous layer. (The first layer input is the original electromagnetic signal) The output is the modulation vector. . The specific implementation can be achieved by using two convolutional layers (1×1 kernel size) followed by a Sigmoid activation function, or by using two fully connected layers. The number of parameters is much smaller than the number of parameters in the corresponding feature extraction layer in the slow system network.

[0036] Modulation operation: ,in This represents element-wise multiplication. This represents the fixed features extracted by the l-th feature extraction layer in a slow system network. This represents the modulation vector output by the l-th feature modulation layer in the fast system network. The modulated feature vector... It serves as both the output of the current feature modulation layer and is passed to the next feature modulation layer.

[0037] The final modulation feature output by the feature modulation layer The input classification head identifies the type of electromagnetic signal. This classification head can employ a prototype classifier or a linear classification layer. The prototype classifier maintains a prototype vector for each category. Calculation during prediction The distance to each prototype vector is used to determine the category with the closest distance as the final output.

[0038] Among them, the total number of parameters of the fast system network is about 10% to 20% of that of the slow system network, and it can quickly adapt to new tasks.

[0039] The network parameters of the fast system network are denoted as The network parameters of the feature modulation layer in the fast-updating system network mentioned below are the updated network parameters. .

[0040] This invention constructs a dual-system architecture of a brain-like electromagnetic model, which includes a slow system network (for self-supervised learning of general features) and a fast system network (for rapid adaptation through feature modulation). This architecture simulates the complementary learning mechanism of the hippocampus and neocortex in the human brain, achieving a balance between long-term knowledge accumulation and short-term rapid adaptation.

[0041] Figure 2 This is one of the flowcharts illustrating the online learning method for a neuromorphic electromagnetic model provided in an embodiment of the present invention, such as... Figure 2 As shown, the method includes: Step 201: Obtain multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category. The multiple negative electromagnetic signal samples are samples with different signal categories than the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples.

[0042] In this step, the positive electromagnetic signal samples are a batch of electromagnetic signal samples with completely complete signal categories, while the negative electromagnetic signal samples are a batch of electromagnetic signal samples with different signal categories, and / or, generated by applying noise, masking, or other artificial perturbations to existing positive electromagnetic signal samples, resulting in altered characteristics.

[0043] In communication signal modulation identification, signal categories can include different modulation schemes such as Binary Phase Shift Keying (BPSK), Quadrature Phase Shift Keying (QPSK), and 16 Quadrature Amplitude Modulation (16QAM). In individual radiation source identification, signal categories can correspond to specific devices such as different models of radar, communication radios, or drone remote controllers.

[0044] Step 202: Input multiple positive electromagnetic signal samples into the slow system network to obtain the first feature vector output by each feature extraction layer, and input multiple negative electromagnetic signal samples into the slow system network to obtain the second feature vector output by each feature extraction layer.

[0045] In this step, such as Figure 1 As shown, each positive sample of electromagnetic signal is input into the slow system network. Each feature extraction layer of the slow system network sequentially performs nonlinear feature extraction and mapping on each positive sample of electromagnetic signal to obtain the first feature vector output by each feature extraction layer, such as a fixed feature h.

[0046] Similarly, multiple negative electromagnetic signal samples are input into the slow system network. Each feature extraction layer of the slow system network sequentially performs nonlinear feature extraction and mapping on each negative electromagnetic signal sample to obtain the second feature vector output by each feature extraction layer.

[0047] Step 203: For each feature modulation layer, input the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and input the second feature vector output by the same feature extraction layer into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer.

[0048] In this step, since each feature modulation layer in the fast system network corresponds one-to-one with each feature extraction layer in the slow system network, for each currently trained feature modulation layer in the fast system network, the first feature vector output by the corresponding feature extraction layer can be input into the feature modulation layer. A modulation vector is generated by the built-in modulation generator of the feature modulation layer, and this modulation vector is used to modulate the input first feature vector element-wise (e.g., multiply), thereby obtaining the first modulation feature output by the feature modulation layer. Here, the first modulation feature is a feature representation that enhances the signal category discriminative information of the current electromagnetic signal positive sample after task-adaptive modulation.

[0049] Similarly, the second feature vector output by the feature extraction layer corresponding to the feature modulation layer can be input into the feature modulation layer. A corresponding modulation vector is generated using the same modulation generator, and this modulation vector is used to perform the same element-wise modulation on the input second feature vector, thereby obtaining the second modulation feature output by the feature modulation layer. The second modulation feature is a feature representation of negative electromagnetic signal samples that have undergone adaptive modulation under the same task but originate from different signal categories or disturbances, and is used to contrast with the first modulation feature at the feature level.

[0050] Step 204: Determine the first quality based on the first modulation feature, determine the second quality based on the second modulation feature, and determine the local loss of the feature modulation layer based on the first quality and the second quality.

[0051] In this step, the first goodness can be defined as the sum of squares of the first modulation features, and the second goodness can be defined as the sum of squares of the second modulation features.

[0052] Among them, for the first The quality of the layer feature modulation layer can be determined based on the following formula (4) or formula (5): (4) (5) in, This represents the modulation feature output by the l-th modulation feature layer. This represents the j-th element in the modulation feature output of the l-th modulation feature layer.

[0053] The first quality can be calculated based on the first modulation feature and formula (4), or based on the first modulation feature and formula (5). The first quality can be used to characterize the activation intensity of the first modulation feature generated by the feature modulation layer or the degree of matching with the current learning task. The higher the first quality, the more effective the feature modulation layer is in modulating the positive samples of electromagnetic signals, and the better the output feature is.

[0054] The second goodness can be calculated based on the second modulation feature and formula (4), or based on the second modulation feature and formula (5). The second goodness can be used to characterize the activation intensity of the second modulation feature generated by the same feature modulation layer or the degree of deviation from the current learning task.

[0055] The local loss of the currently trained feature modulation layer can be determined based on the first and second goodness scores. This local loss can be a contrastive loss function (e.g., using the Softmax cross-entropy or mean squared error form). Its optimization objective is to widen the gap between the first and second goodness scores, that is, to encourage the currently trained feature modulation layer to make the goodness score of positive electromagnetic signal samples significantly higher than that of negative electromagnetic signal samples.

[0056] Step 205: Update the network parameters of the feature modulation layer based on local loss to perform online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.

[0057] In this step, after determining the local loss corresponding to the currently trained feature modulation layer, the network parameters of the modulation generator in this feature modulation layer are updated using only the local loss of that feature modulation layer. For example, local gradient descent or gradient-free optimization can be used for updating. ,in, Indicates the learning rate. This represents the gradient of the l-th feature modulation layer. Let represent the local loss of the l-th feature modulation layer. In the case of online learning of the fast system network, the network parameters of the slow system network are completely fixed. The fast system network adapts to the new task only through local modulation parameters, naturally decoupling the knowledge of the old and new tasks, fundamentally avoiding the problem of catastrophic forgetting.

[0058] The value of the aforementioned local loss depends only on the output of the currently trained feature modulation layer. Therefore, the update of the network parameters of the feature modulation layer based on this local loss is local and does not depend on the backpropagation of error gradients from other layers of the fast system network, thus meeting the real-time requirements of electromagnetic signal processing.

[0059] By updating the network parameters for each feature modulation layer in the manner described above, and by allowing multiple feature modulation layers to update their parameters in parallel without interference, the efficiency of online learning for fast system networks can be improved. Since slow system networks do not participate in online learning, performing the aforementioned local parameter updates without backpropagation on the lightweight fast system network enables the entire neuromorphic electromagnetic model to quickly adapt to new tasks or data, thus achieving online learning of the neuromorphic electromagnetic model.

[0060] Fast system networks can be used to identify the signal type of electromagnetic signals.

[0061] The online learning method for a brain-like electromagnetic model provided in this invention involves acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category. The negative electromagnetic signal samples are samples with different signal categories than the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples. The positive electromagnetic signal samples are input into a slow system network to obtain a first feature vector output by each feature extraction layer, and the negative electromagnetic signal samples are input into the slow system network to obtain a second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain a first modulation feature output by the feature modulation layer, and the second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain a second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first and second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning of a fast system network, which is used to identify the signal category of the electromagnetic signal. Because it employs a local learning mechanism without backpropagation, it avoids the cross-layer gradient calculations and storage of a large number of intermediate activation values ​​required by traditional backpropagation algorithms. This reduces computational load, time consumption, and hardware resource requirements when learning large models for electromagnetic signal recognition online, enabling efficient deployment and online learning on edge devices with limited computing and storage resources.

[0062] For example, based on the above embodiments, when determining the local loss of the feature modulation layer based on the first and second excellence, the first loss can be determined according to formula (1) or formula (2). Local loss of layer feature modulation layer : (1) (2) in, Indicates first-class excellence. Indicates the second best quality. This represents the temperature coefficient, which is typically between 0.1 and 1.0. This indicates the target quality of the positive sample of the electromagnetic signal; for example, it can be set to a value of 1. This indicates the target quality of the negative sample of the electromagnetic signal; for example, it can be set to 0.

[0063] In this embodiment, the contrast loss is determined as a local loss by formula (1) or formula (2), which can encourage the first goodness of positive electromagnetic signal samples to be higher and the second goodness of negative electromagnetic signal samples to be lower. This can effectively drive each feature modulation layer to learn a more discriminative feature representation, enabling the fast system network to distinguish signal categories more clearly and robustly when facing new data, thereby improving the online learning efficiency and recognition robustness of the fast system network in open electromagnetic environments.

[0064] For example, based on the above embodiments, when the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, the modulation feature output by the previous feature modulation layer can be input into the feature modulation layer, the modulation vector is determined by the modulation generator in the feature modulation layer, and the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is modulated based on the modulation vector to obtain the first modulation feature.

[0065] Specifically, the multiple feature modulation layers in the fast system network are connected in series. When training the l-th feature modulation layer, the modulation features output by the (l-1)-th feature modulation layer can be input into the l-th feature modulation layer. The modulation generator (a lightweight network, such as a two-layer fully connected network) in this layer processes and transforms the modulation vector, which is a weight vector with the same dimension as the first feature vector and whose element values ​​are usually in the range [0,1].

[0066] Furthermore, the first modulated feature can be obtained by modulating the first feature vector output by the l-th feature extraction layer by performing an element-wise multiplication (Hadamard product) operation between the modulation vector and the first feature vector output by the l-th feature extraction layer.

[0067] The method for determining the second modulation feature is similar to that for determining the first modulation feature, and will not be repeated here.

[0068] In this embodiment, the modulation vector is determined based on the modulation features output by the previous feature modulation layer, and the first modulation feature output by the feature modulation layer is determined based on the modulation vector. This enables the cascading and adaptive refinement of features, allowing the modulation of the current feature modulation layer to depend on the contextual information contained in the modulated features of the previous layer. This allows the fast system network to adaptively adjust the fixed features of the slow system network from coarse to fine in a layer-by-layer, progressive manner, effectively enhancing the brain-like electromagnetic model's ability to deeply analyze and rapidly adapt to complex electromagnetic signal features.

[0069] For example, based on the above embodiments, when acquiring multiple electromagnetic signal negative samples and multiple electromagnetic signal positive samples of the same signal category, multiple electromagnetic signal negative samples can be sampled from the negative sample pool; wherein, the negative sample pool includes electromagnetic signal samples of different signal categories, and / or electromagnetic signal samples generated by perturbation; multiple electromagnetic signal positive samples are sampled from the positive sample pool corresponding to the target signal category.

[0070] Specifically, the negative sample pool includes at least one of the following: samples randomly drawn from the positive sample pool corresponding to other signal categories different from the current target signal category; samples generated by perturbing known positive samples (e.g., adding noise, random masking); samples randomly sampled from environmental background signals; and low-quality samples with confidence levels below a threshold. Each positive sample pool includes electromagnetic signal samples of the same signal category; therefore, a corresponding positive sample pool can be created for each signal category.

[0071] By sampling multiple negative electromagnetic signal samples from the negative sample pool and multiple positive electromagnetic signal samples from the positive sample pool, high-quality and diverse training data pairs can be provided for subsequent local updates without backpropagation based on contrastive learning. This sampling mechanism can ensure the richness of electromagnetic signal negative samples, which is conducive to fast system networks learning more discriminative and robust feature representations, thereby improving the efficiency and effectiveness of online learning.

[0072] For example, based on the above embodiments, when sampling multiple electromagnetic signal positive samples from the positive sample pool corresponding to the target signal category, multiple electromagnetic signal positive samples can be sampled from the positive sample pool corresponding to the target signal category when the number of electromagnetic signal samples in the positive sample pool corresponding to the target signal category is greater than or equal to a first preset threshold.

[0073] Specifically, when adding electromagnetic signal samples to the positive sample pools corresponding to each signal category, if the number of electromagnetic signal samples in the positive sample pool corresponding to a certain target signal category is greater than or equal to a first preset threshold, such as greater than or equal to 32, it indicates that a sufficient number of high-quality learning samples have been accumulated for that target signal category, which can constitute a statistically significant training batch. At this time, multiple electromagnetic signal positive samples can be sampled from the positive sample pool corresponding to that target signal category, thereby triggering the online learning process of the fast system network in the neuromorphic electromagnetic model.

[0074] In this embodiment, once the number of electromagnetic signal samples in the positive sample pool reaches a first preset threshold, the fast system network is triggered to learn online. This enables efficient, on-demand learning driven by events, avoiding frequent updates for each new sample. This mechanism balances learning real-time performance with computational efficiency, avoids redundant computation, and ensures that the fast system network only adjusts its parameters when it has sufficient electromagnetic signal samples. This significantly reduces the average computational load and energy consumption on resource-constrained edge devices.

[0075] The following section will provide a detailed explanation of the construction process for the positive and negative sample pools.

[0076] For example, for the positive sample pool corresponding to each signal category, the acquired electromagnetic signal samples can be input into the slow system network to obtain the target features output by the slow system network. The similarity between the target features and the sample features in the feature cache is determined. The sample features are the features of the stored electromagnetic signal samples. The similarities are sorted in descending order. Based on the first preset number of similarities, pseudo-labels and pseudo-label confidence scores of the electromagnetic signal samples are generated by weighted voting. If the pseudo-label confidence score is greater than or equal to the first preset confidence score, the electromagnetic signal samples and target features are added to the positive sample pool corresponding to the signal category represented by the pseudo-label.

[0077] Specifically, for newly acquired, unlabeled electromagnetic signal samples, a self-labeling method based on feature self-similarity can be used to label the signal category of the unlabeled electromagnetic signal samples.

[0078] Among these features, a feature cache can be pre-built. , Let represent the sample feature of the i-th electromagnetic signal sample, which is the feature output by the slow system network. This represents a known label for the i-th electromagnetic signal sample, such as a real or fake label. This represents the timestamp when the i-th electromagnetic signal sample is stored in the feature buffer. This represents the confidence level when the label of the i-th electromagnetic signal sample is a pseudo-label. The capacity of this feature cache can be set to... For example, 10000, and update using a First In First Out (FIFO) strategy or an importance-based sampling strategy.

[0079] Newly collected unlabeled electromagnetic signal samples Inputting the target features into a slow system network yields the target features output by the slow system network. Then, the target feature is calculated using formula (6). Cosine similarity between all sample features in the feature cache: (6) in, Representing target features Features of the i-th sample Cosine similarity between them This indicates the total number of sample features in the feature cache. The feature cache stores all stored electromagnetic signal samples and their respective features.

[0080] After determining the similarity between the target feature and the features of each sample, all similarities are sorted in descending order, and the first preset number of similarities are selected. The preset number can be, for example, K, which is a hyperparameter and can be set to 5-10 based on experience.

[0081] Furthermore, the pseudo-labels and pseudo-label confidence levels of the electromagnetic signal samples are generated by weighted voting according to the following formula (7): (7) Calculate the highest weighted vote score The highest weighted voting score is also the false label confidence level. If the highest weighted voting score is greater than or equal to the first pre-set confidence level... Then pseudo-tags are accepted. and target features and pseudo-labels Stored in the positive sample pool corresponding to the signal category represented by the pseudo-label.

[0082] in, This represents the pseudo-label, i.e., the predicted signal category finally determined through a weighted voting mechanism; c represents the candidate signal category; and the first pre-set confidence level is... It is usually set to 0.6.

[0083] In this embodiment, by determining the similarity between the target feature of the electromagnetic signal sample and each sample feature in the feature cache, for a preset number of sample features with the highest similarity, pseudo-labels and pseudo-label confidence scores of the electromagnetic signal sample can be generated through weighted voting. When the pseudo-label confidence score is greater than or equal to a first preset confidence score, the electromagnetic signal sample and the target feature are added to the positive sample pool corresponding to the signal category represented by the pseudo-label. This allows for the automatic generation of reliable pseudo-labels for a large number of unlabeled signals in an open electromagnetic environment, thereby dynamically expanding the positive sample pool of the corresponding signal category. This approach not only fundamentally reduces the reliance on scarce and expensive manually labeled data, but also provides a high-quality training data source that can directly drive learning for subsequent backpropagation-free local fine-tuning (Forward-Forward algorithm) based on the comparison of positive and negative electromagnetic signal samples.

[0084] For example, if the confidence level of the pseudo-label is less than the first preset confidence level, the electromagnetic signal sample and the target feature are added to the new category candidate sample pool. If the number of samples in the new category candidate sample pool is greater than or equal to the second preset threshold, the target features in the new category candidate sample pool are clustered. If there is at least one cluster that is successfully clustered, a new positive sample pool for the new signal category corresponding to each cluster is created, and the candidate electromagnetic signal samples belonging to the cluster and the corresponding target features are stored in the new positive sample pool.

[0085] Specifically, if the pseudo-label confidence score of a newly acquired unlabeled electromagnetic signal sample is less than the first preset confidence score, it indicates that the target features of the electromagnetic signal sample are not sufficiently similar to the sample features of all known signal categories in the feature cache, and cannot be reliably classified using existing knowledge. Therefore, it may belong to a new signal category that the system has not yet learned. In this case, it is necessary to retrieve the unlabeled electromagnetic signal sample and its corresponding target features. Add it to the new category candidate sample pool, and set the label of the electromagnetic signal sample to None and the pseudo-label confidence to 0.

[0086] Through continuous accumulation, when the number of samples in the candidate pool for the new category is greater than or equal to a second preset threshold, such as greater than or equal to 10, it indicates that a sufficient number of electromagnetic signal samples have been accumulated for this potential new signal category. Cluster analysis can then be used to confirm the existence of a new signal category with statistical reliability. Therefore, the following operation can be performed: cluster the target features in the candidate pool for the new category, such as using density-based spatial clustering of applications with noise (DBSCAN). Set the value to 0.5 and the minimum sample size to 3.

[0087] If at least one cluster is successfully formed, for each cluster, the cluster center is taken as the prototype feature of the new signal category, and a new positive sample pool is created for that new signal category. Candidate electromagnetic signal samples belonging to this cluster and their corresponding target features are moved into the new positive sample pool. In, and assign new signal category labels. .

[0088] If the samples in the new category candidate sample pool are too scattered, resulting in no successfully clustered clusters, no processing can be done temporarily. Continue to wait for more samples to be stored in the new category candidate sample pool until the number of samples in the new category candidate sample pool reaches the fourth preset threshold, and then re-clustering can be performed.

[0089] In this embodiment, by clustering the target features in the candidate sample pool of new categories and creating a new positive sample pool for each successfully clustered cluster, the candidate electromagnetic signal samples belonging to the cluster and the corresponding target features are stored in the new positive sample pool. This enables automatic labeling of newly emerging or unknown electromagnetic signal categories. In this way, new electromagnetic signals that are significantly different from existing categories can be identified autonomously without relying on prior knowledge or manual relabeling, and an independent positive sample pool can be established for them. This greatly expands the adaptability of the brain-like electromagnetic model in dynamic electromagnetic environments and provides a data foundation for subsequent online learning of these new signal categories.

[0090] For example, based on the above embodiments, when the number of electromagnetic signal samples in the target positive sample pool reaches a third preset threshold, the target positive sample pool is updated using a FIFO method, or electromagnetic signal samples in the target positive sample pool with pseudo-label confidence levels lower than a second preset confidence level are deleted.

[0091] Specifically, an upper limit for the pool capacity can be set for the positive sample pool corresponding to each signal category. If the number of electromagnetic signal samples in a target positive sample pool exceeds a third preset threshold, it indicates that the target positive sample pool has reached or is about to reach its storage capacity limit and cannot continue to store new electromagnetic signal samples indefinitely. In this case, a preset number of electromagnetic signal samples and their corresponding features can be eliminated using a FIFO (First-In, First-Out) method, or electromagnetic signal samples in the target positive sample pool with pseudo-label confidence levels lower than a second preset confidence level can be deleted. When a new electromagnetic signal sample has a true label or a pseudo-label confidence level greater than the first preset confidence level, the new electromagnetic signal sample and its corresponding features can be stored in the target positive sample pool. The third preset threshold can be the pool capacity limit. It can also be less than the maximum pool capacity. The value.

[0092] In this embodiment, dynamic management and optimization of the positive sample pool can be achieved by updating the electromagnetic signal samples in the target positive sample pool. The FIFO strategy can ensure the timeliness of the samples in the positive sample pool, prioritizing the retention of the latest electromagnetic signal samples to reflect the latest changes in the environment; while the confidence-based elimination mechanism can actively filter out low-quality pseudo-label samples that may introduce noise, thereby improving the overall reliability of the electromagnetic signal samples in the positive sample pool.

[0093] Similarly, when the number of electromagnetic signal samples in the negative sample pool has reached or is about to reach its storage capacity limit, the negative sample pool can also be updated using the FIFO method.

[0094] For example, a negative sample pool can be constructed based on at least one of the following methods: Electromagnetic signal samples are randomly drawn from the positive sample pools corresponding to multiple signal categories, and the drawn electromagnetic signal samples of different signal categories are stored in the negative sample pool. Perturbations are applied to the positive electromagnetic signal samples in the positive sample pool, such as adding Gaussian noise or random masks, to generate negative electromagnetic signal samples, and the generated negative electromagnetic signal samples are stored in the negative sample pool. Electromagnetic signal negative samples are obtained by randomly sampling from the environmental background signal and then stored in the negative sample pool. Electromagnetic signal samples with a false label confidence level lower than the first preset confidence level are stored in the negative sample pool.

[0095] By dynamically constructing a global negative sample pool from multiple different sources, the diversity of electromagnetic signal negative samples can be ensured, enabling it to cover samples of other known categories different from the positive electromagnetic signal samples, perturbed samples, environmental background noise, and low-confidence samples. This diverse source of electromagnetic signal negative samples provides rich and comprehensive negative samples for contrastive learning-based backpropagation-free local fine-tuning (such as the Forward-Forward algorithm), thereby driving fast system networks to learn more discriminative and generalizable feature representations. This effectively improves the robustness and adaptability of neuromorphic electromagnetic models in distinguishing different signals and resisting interference in complex and open electromagnetic environments.

[0096] Furthermore, based on the above embodiments, in order to further improve the adaptive capability of fast system networks in open environments, when updating the network parameters of the feature modulation layer based on local loss to perform online learning of fast system networks, reinforcement learning (RL) can be combined with local fine-tuning without backpropagation.

[0097] For example, the modulation generator in the currently trained feature modulation layer is used as the policy network. The modulation vector is output based on the current state of the feature modulation layer. The current state includes at least one of the following: a first feature vector, a second feature vector output by the feature extraction layer corresponding to the feature modulation layer, the modulation feature output by the previous feature modulation layer, and the pseudo-label confidence of the current electromagnetic signal sample. Based on multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples, a reward signal for the reinforcement learning process is determined. The reward signal includes a main reward, an auxiliary reward, and a sparse reward. The main reward is determined based on the proportion of correctly classified samples and / or the proportion of incorrectly classified samples among the multiple positive electromagnetic signal samples, or based on the recognition confidence of the multiple positive electromagnetic signal samples after modulation of the modulation vector. The auxiliary reward is determined based on a first goodness and a second goodness. The sparse reward is a reward configured when a new category is successfully identified or a preset task is completed. The reinforcement learning process is used to characterize the parameter update process of the currently trained feature modulation layer. A policy loss is constructed based on the reward signal, and the parameters of the policy network are updated based on the local loss and the policy loss to obtain the updated feature modulation layer network parameters for online learning of the fast system network.

[0098] Specifically, the update process of the network parameters of each feature modulation layer is regarded as an independent reinforcement learning problem, with the first layer being the second layer. As for the current training feature modulation layer, the first layer will be the feature modulation layer. Modulation generator in the layer feature modulation layer As a policy network, according to the first The current state of the layer feature modulation layer outputs the modulation vector. The current state includes the [number]th [unit] in the slow system network. The first and second feature vectors output by the feature extraction layer, and the third feature vector... The modulated features output by the layer feature modulation layer, the pseudo-label confidence of the current electromagnetic signal sample, and at least one of the following: current electromagnetic signal sample includes positive and negative electromagnetic signal samples. The task context information can be, for example, a global task embedding, such as a task identifier. Modulation vector. That is, action It can be a continuous value, such as sampled through a Gaussian distribution, or a discrete value, such as sampled through a classification distribution.

[0099] Furthermore, the reward signal for the reinforcement learning process can be determined based on multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples. The reward signal fed back by this environment It is determined by performance in the current task. Specifically, the reward signal. The reward includes a primary reward, an auxiliary reward, and a sparse reward. The primary reward is determined based on the correctness and / or error rate of classification of multiple positive electromagnetic signal samples, such as +1 for correct and -1 for incorrect. It can also be determined based on the recognition confidence of multiple positive electromagnetic signal samples after modulation of the modulation vector, such as the softmax probability. The auxiliary reward can be determined based on the first and second excellence scores. For example, the first excellence score of the positive electromagnetic signal samples can be used as an immediate positive reward to encourage fast system networks to generate good feature activations. The sparse reward is a large reward configured for successfully identifying a new signal category or completing a preset task.

[0100] When performing policy optimization, policy gradient methods (such as REINFORCE) or their variants (such as PPO, A2C) are used to optimize the policy network parameters. Taking REINFORCE as an example, the policy loss can be expressed as shown in the following formula (8): (8) in, Indicates strategy loss. This represents the instantaneous reward signals at each time step. The cumulative reward, obtained by summing the discount factors, reflects the total expected return from the present moment to the future after taking a sequence of actions in the current state. This represents the action output by the policy network at time t, i.e., the modulation vector. , This represents the current state at time t. Represents the policy network.

[0101] By optimizing the policy network to maximize the cumulative reward, parameter updates can be achieved without backpropagation, since the policy gradient only requires forward sampling and reward calculation.

[0102] After determining the policy loss, the total loss is determined based on the local loss and the policy loss. The parameters of the policy network are then updated based on this total loss, which means updating the network parameters of the currently trained feature modulation layer, thereby enabling online learning of the fast system network.

[0103] In practical implementation, during FF update, the contrast loss between positive and negative electromagnetic signal samples can be considered as an immediate reward. Meanwhile, RL allows the system to optimize using long-term cumulative rewards, making it more suitable for sequential decision-making tasks (such as multi-step inference and task switching). In actual implementation, pure FF update, pure RL update, or a combination of both can be selected based on the task characteristics.

[0104] In this embodiment, a reinforcement learning-enhanced backpropagation-free fine-tuning mechanism can be adopted, which can endow the fast system network with long-term decision-making and adaptive optimization capabilities. This enables it to more effectively handle task switching, sequential decision-making, and learning scenarios affected by delayed feedback in a dynamic and open electromagnetic environment, thereby further improving the learning efficiency, robustness, and environmental adaptability of the fast system network.

[0105] For example, based on the above embodiments, the update of the slow system network can be performed by performing self-supervised learning on the slow system network every preset time interval, based on the unlabeled sample features of multiple electromagnetic signal samples in the feature cache. During the self-supervised learning process of the slow system network, the fast system network keeps its current parameters unchanged, and the feature cache includes the sample features of multiple electromagnetic signal samples.

[0106] Specifically, when the system is idle, or no new signals arrive, or when a preset time has elapsed, multiple unlabeled sample features from the feature cache, or all cached features, can be used to perform self-supervised learning on the slow system network, such as Barlow Twins, to update the system parameters of the slow system network. The update frequency can be set to every [number of] [minutes / days]. Once a time step is taken, such as once per hour, this method will not affect the online learning of the fast system network.

[0107] It should be noted that during the self-supervised learning process of the slow system network, the fast system network keeps its current parameters unchanged.

[0108] In this embodiment, the slow system network can undergo self-supervised learning at preset intervals based on multiple unlabeled sample features in the feature cache. This enables the fast and slow system networks to learn stably and collaboratively at different times. The fast system network performs rapid task adaptation at the millisecond / second level, while the slow system network undergoes slow evolution of general features at the hour / day level. This asynchronous, low-frequency background update mechanism ensures that the slow system network can continuously optimize its general feature representation using long-term accumulated and diverse historical data, making it more robust and discriminative. Furthermore, since the update process of the slow system network is completely decoupled from the online learning and real-time inference of the fast system network, the brain-like electromagnetic model achieves a combined fast and slow, continuously evolving learning capability without affecting the system's real-time response. This enhances the adaptability and robustness of the brain-like electromagnetic model for long-term deployment in open environments.

[0109] Figure 3 The second flowchart of the online learning method for the neuromorphic electromagnetic model provided in this embodiment of the invention is as follows: Figure 3As shown, this method constructs a control flow of "real-time inference - dynamic evaluation - conditional triggering - asynchronous update" to achieve efficient adaptive learning of the brain-like electromagnetic model in resource-constrained environments. First, the system continuously receives electromagnetic signals from the open environment and inputs them into a slow system network for feature extraction. Then, the extracted features generate pseudo-labels through a self-labeling module, and the fast system network predicts and completes the real-time signal recognition task. The recognition results are not only used for the current decision output but also enter a new category processing stage to identify potential new signal types, and the relevant sample data is accumulated and stored through a sample database entry stage.

[0110] To ensure real-time performance while also considering model evolution, a condition-triggered fine-tuning mechanism can be used: fine-tuning is triggered by decision nodes to evaluate the current system state, sample accumulation, or task requirements. For example, if the number of accumulated samples is less than a first preset threshold, the system directly returns to the signal input stage to continue real-time inference, ensuring low latency; while if the number of accumulated samples is greater than or equal to the first preset threshold, the system will activate the "local fine-tuning (without backpropagation)" module to update the parameters of the fast system network using newly added samples.

[0111] It is worth noting that this fine-tuning process only performs independent local optimization layer by layer for the fast system network, without relying on global error backpropagation, thus avoiding huge computational overhead while absorbing new knowledge. In addition, to ensure that the basic representation ability of the brain-like electromagnetic model continuously improves over time, the system also periodically performs self-supervised learning on the slow system network based on massive amounts of unlabeled historical data in the feature cache through a background slow system update path (such as using mechanisms like BarlowTwins). Moreover, this background update process is decoupled from the front-end online learning, thereby achieving efficient co-evolution of the fast and slow system networks in time and space.

[0112] In this embodiment of the invention, the Forward Propagation (FF) algorithm is used as an example to achieve backpropagation-free local fine-tuning of fast system networks. The FF algorithm optimizes the network layer by layer through two forward propagations (positive sample propagation and negative sample propagation of electromagnetic signals). Each hidden layer independently optimizes the local objective function, completely avoiding cross-layer gradient propagation. Based on this, the invention can also employ other backpropagation-free algorithms, such as Target Propagation (TP), Equilibrium Propagation (EP), NoProp, and Reinforcement Learning (RL), to achieve independent layer-by-layer optimization of fast systems. Each algorithm is compatible with the dual-system architecture and self-labeling module, enabling backpropagation-free continuous learning of fast systems.

[0113] In the Target Propagation (TP) algorithm, a target representation is assigned to each hidden layer. The top-level target is propagated back to each layer via a reverse mapping, ensuring that the outputs of each layer approximate the target. Specifically, a reverse network is trained. Backpropagate the top-level error to the first... Layer, generating the desired output of that layer. Then update To minimize the actual output With expected output Differences (such as mean squared error). In fast system networks, the feature modulation layer can be regarded as a hidden layer in target propagation, and the top-level target is generated by pseudo-labels provided by the self-labeling module.

[0114] In the Equilibrium Propagation (EP) algorithm, the network can be viewed as an energy model that reaches equilibrium through two forward propagations. Weight updates are then calculated using perturbations near the equilibrium point. Specifically, the network first reaches equilibrium in a free state (first stage). Then, the input is fixed, and the output is perturbed towards the target direction, allowing the network to reach equilibrium again (second stage). The difference between the two equilibrium states is used to update the weights. In fast system networks, the feature modulation layer parameters can be considered as part of an energy function. Weight updates are calculated by comparing the equilibrium states of positive and negative electromagnetic signal samples, without the need for explicit gradients.

[0115] The NoProp algorithm borrows the idea of ​​diffusion models, treating each layer as a denoising autoencoder. During training, noise is added to the input, and each layer independently learns the denoising target, with the loss function being the denoising error (such as mean squared error). NoProp does not require forward or backward propagation, relying solely on intra-layer computation. In fast system networks, the output of the feature modulation layer can be considered as the denoised result, and the noisy version can be used as input for training.

[0116] Figure 4 The third flowchart of the online learning method for the neuromorphic electromagnetic model provided in this embodiment of the invention is as follows: Figure 4 As shown, this method enables efficient adaptive learning of the system in resource-constrained environments. First, at each time step t, the system continuously receives electromagnetic signals from the open environment. The electromagnetic signal It may or may not be tagged. It transmits electromagnetic signals. Inputting the data into a slow system network for feature extraction yields the target features. and features of each layer Subsequently, the system proceeds to a decision node with a real label: if the electromagnetic signal... Carry real labels If no labels are found, proceed directly to the next step using the actual labels; otherwise, proceed to the self-labeling module and output pseudo labels using the self-labeling methods described in the aforementioned embodiments. And the confidence level of the pseudo-label. In the self-labeling module, if the confidence level of the pseudo-label is lower than the first preset confidence level... Then the electromagnetic signal Features Store in feature cache and wait for the next electromagnetic signal; if the confidence level of the fake tag is higher than the first preset confidence level... Then adopt the pseudo-tag. .

[0117] Next, the system performs the fast system network forward and prediction steps, using the current fast system network parameters to analyze the electromagnetic signal. Modulation and classification are performed, and prediction results are output. That is, electromagnetic signals The signal category. After prediction, the system performs a sample storage operation, adding samples with real labels. or pseudo-label The samples are added to the positive sample pool of the corresponding category and stored in the feature cache (if they already exist, the confidence level is updated).

[0118] Furthermore, the number of samples in the positive sample pool corresponding to each signal category can be determined. If the number of samples in the positive sample pool corresponding to a certain signal category is greater than or equal to the first preset threshold B, the local fine-tuning (without backpropagation) process is activated; otherwise, the system directly returns to the signal input stage to continue real-time inference, ensuring low latency.

[0119] The local fine-tuning process employs the Forward-Forward (FF) algorithm. It samples B positive electromagnetic signal samples from a positive sample pool whose sample count is greater than or equal to a first preset threshold B, and randomly samples B negative electromagnetic signal samples from a negative sample pool. FF updates are performed in parallel for each feature modulation layer to optimize the network parameters of each feature modulation layer. Alternatively, the positive sample pool can be cleared or some samples can be retained, for example, retaining the 10 most recent positive electromagnetic signal samples. After fine-tuning is complete, the next signal is processed.

[0120] In addition, the system is designed with an asynchronous background slow system update mechanism: when the system is idle (such as when no new signal arrives) or every preset time interval, the unlabeled data in the feature cache is used to perform self-supervised learning (such as Barlow Twins) on the slow system network to update the network parameters of the slow system network. The update frequency can be set to once every T_slow time steps (such as once per hour). When the slow system network is updated, the fast system network keeps its current parameters unchanged, thereby realizing the efficient co-evolution of the fast and slow systems in time and space. Simultaneously, if a new signal category is detected, the system will perform a "classifier expansion" operation: if the fast system network uses a prototype classifier, the new prototype is added to the prototype list; if a linear classifier head is used, the classification weight matrix is ​​expanded, the weight of the new signal category is initialized to a specific value (which needs to be linearly transformed), and the bias is initialized to 0; then the negative sample pool is updated, and the new signal category is added to the candidate set of the negative sample pool for subsequent negative sample sampling; finally, the feature cache is updated, the labels of this batch of samples in the cache are updated to the new signal category, and the confidence level is set to 1.0, thereby ensuring that the neuromorphic electromagnetic model can dynamically adapt to new signal types in the open environment and continuously improve the generalization ability and robustness of the neuromorphic electromagnetic model.

[0121] The online learning device for the neuromorphic electromagnetic model provided by the present invention will be described below. The online learning device for the neuromorphic electromagnetic model described below and the online learning method for the neuromorphic electromagnetic model described above can be referred to in correspondence.

[0122] Figure 5 This is a schematic diagram of the structure of an online learning device for a neuromorphic electromagnetic model provided in an embodiment of the present invention. The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer corresponds one-to-one. Figure 5 As shown, the online learning device 500 for this brain-like electromagnetic model includes: The acquisition module 11 is used to acquire multiple electromagnetic signal negative samples and multiple electromagnetic signal positive samples of the same signal category, wherein the multiple electromagnetic signal negative samples are samples whose signal categories are different from those of the electromagnetic signal positive samples, and / or samples generated after perturbing the electromagnetic signal positive samples; Input module 12 is used to input the plurality of positive electromagnetic signal samples into the slow system network to obtain the first feature vector output by each feature extraction layer, and to input the plurality of negative electromagnetic signal samples into the slow system network to obtain the second feature vector output by each feature extraction layer; The input module 12 is further configured to input the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and input the second feature vector output by the same feature extraction layer into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. The determining module 13 is used to determine a first quality based on the first modulation feature, determine a second quality based on the second modulation feature, and determine the local loss of the feature modulation layer based on the first quality and the second quality. The update module 14 is used to update the network parameters of the feature modulation layer based on the local loss, so as to perform online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.

[0123] In one example embodiment, the determining module 13 is specifically used for: Determine the first according to formula (1) or formula (2). Local loss of layer feature modulation layer : (1) (2) in, This indicates the first level of excellence. This indicates the second degree of excellence. Indicates the temperature coefficient. This indicates the target quality of the positive sample of the electromagnetic signal. This indicates the target quality of the negative sample of the electromagnetic signal.

[0124] In one example embodiment, the input module 12 is specifically used for: The modulation features output from the previous feature modulation layer are input into the feature modulation layer, and the modulation vector is determined by the modulation generator in the feature modulation layer. Based on the modulation vector, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is modulated to obtain the first modulation feature.

[0125] In one example embodiment, the acquisition module 11 is specifically used for: The plurality of electromagnetic signal negative samples are obtained by sampling from the negative sample pool; wherein, the negative sample pool includes electromagnetic signal samples of different signal categories, and / or electromagnetic signal samples generated by perturbation; Multiple electromagnetic signal positive samples are obtained by sampling from the positive sample pool corresponding to the target signal category.

[0126] In one example embodiment, the acquisition module 11 is specifically used for: If the number of electromagnetic signal samples in the positive sample pool corresponding to the target signal category is greater than or equal to a first preset threshold, multiple electromagnetic signal positive samples are sampled from the positive sample pool corresponding to the target signal category.

[0127] In one example embodiment, the apparatus further includes a generation module and an adding module, wherein: The input module 12 is also used to input the acquired electromagnetic signal samples into the slow system network to obtain the target features output by the slow system network; The determining module 13 is also used to determine the similarity between the target feature and each sample feature in the feature cache, wherein the sample feature is the feature of the stored electromagnetic signal sample; The generation module is used to sort the similarities in descending order, and generate pseudo-labels and pseudo-label confidence scores for the electromagnetic signal samples by weighted voting based on the first preset number of similarities. An adding module is used to add the electromagnetic signal sample and the target feature to the positive sample pool corresponding to the signal category represented by the pseudo-label when the confidence level of the pseudo-label is greater than or equal to a first preset confidence level.

[0128] In one example embodiment, the device further includes a clustering module and a storage module, wherein: The addition module is also used to add the electromagnetic signal sample and the target feature to the new category candidate sample pool when the confidence of the pseudo-label is less than the first preset confidence. The clustering module is used to cluster the target features in the new category candidate sample pool when the number of samples in the new category candidate sample pool is greater than or equal to a second preset threshold. The storage module is configured to, in the case that at least one cluster has been successfully clustered, create a new positive sample pool for each cluster corresponding to a new signal category, and store the candidate electromagnetic signal samples belonging to the cluster and the corresponding target features into the new positive sample pool.

[0129] In one example embodiment, the device further includes an update module, wherein: The update module is used to update the target positive sample pool in a first-in-first-out (FIFO) manner when the number of electromagnetic signal samples in the target positive sample pool reaches a third preset threshold, or to delete electromagnetic signal samples in the target positive sample pool whose false label confidence is lower than a second preset confidence.

[0130] In one example embodiment, the negative sample pool is constructed based on at least one of the following methods: Electromagnetic signal samples are randomly drawn from the positive sample pool corresponding to multiple signal categories, and the drawn electromagnetic signal samples of different signal categories are stored in the negative sample pool. The positive electromagnetic signal samples in the positive sample pool are perturbed to generate negative electromagnetic signal samples, and the generated negative electromagnetic signal samples are stored in the negative sample pool. Electromagnetic signal negative samples are obtained by randomly sampling from the environmental background signal and stored in the negative sample pool. Electromagnetic signal samples with a false label confidence level lower than the first preset confidence level are stored in the negative sample pool.

[0131] In one example embodiment, the update module 14 is specifically used for: The modulation generator in the currently trained feature modulation layer is used as a policy network. The modulation vector is output according to the current state of the feature modulation layer. The current state includes at least one of the following: the first feature vector and the second feature vector output by the feature extraction layer corresponding to the feature modulation layer, the modulation feature output by the previous feature modulation layer, and the pseudo-label confidence of the current electromagnetic signal sample. Based on the multiple negative electromagnetic signal samples and the multiple positive electromagnetic signal samples, a reward signal for the reinforcement learning process is determined. This reward signal includes a main reward, an auxiliary reward, and a sparse reward. The main reward is determined based on the proportion of correctly classified samples and / or misclassified samples among the multiple positive electromagnetic signal samples, or based on the recognition confidence of the multiple positive electromagnetic signal samples after modulation by the modulation vector. The auxiliary reward is determined based on a first excellence score and a second excellence score. The sparse reward is a reward configured when a new category is successfully identified or a preset task is completed. The reinforcement learning process is used to characterize the parameter update process of the currently trained feature modulation layer. Based on the reward signal, a policy loss is constructed, and based on the local loss and the policy loss, the parameters of the policy network are updated to obtain the updated feature modulation layer network parameters, so as to perform online learning on the fast system network.

[0132] In one example embodiment, the update module 14 is further configured to perform self-supervised learning on the slow system network at preset intervals based on multiple unlabeled sample features in the feature cache, wherein, during the self-supervised learning process of the slow system network, the fast system network keeps its current parameters unchanged, and the feature cache includes sample features of multiple electromagnetic signal samples.

[0133] The apparatus of this embodiment can be used in any embodiment of the online learning method for brain-like electromagnetic models. Its specific implementation process and technical effects are similar to those in the online learning method for brain-like electromagnetic models. For details, please refer to the detailed description in the online learning method for brain-like electromagnetic models, which will not be repeated here.

[0134] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an online learning method for a neuromorphic electromagnetic model. The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer corresponds one-to-one. The method includes: acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; inputting the multiple positive electromagnetic signal samples into the slow system network to obtain the first feature vector output by each feature extraction layer, and inputting the multiple negative electromagnetic signal samples into... The network is fed into the slow system network to obtain the second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer. The second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first goodness and the second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning on the fast system network, which is used to identify the signal category of electromagnetic signals.

[0135] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0136] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the online learning method of the neuromorphic electromagnetic model provided by the above methods. The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer corresponds one-to-one. The method includes: acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from those of the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; inputting the multiple positive electromagnetic signal samples into the slow system network to obtain... The network obtains a first feature vector output by each feature extraction layer and inputs the multiple electromagnetic signal negative samples into the slow system network to obtain a second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain a first modulation feature output by the feature modulation layer, and the second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain a second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first goodness and the second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning of the fast system network, which is used to identify the signal category of the electromagnetic signal.

[0137] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an online learning method for a neuromorphic electromagnetic model provided by the methods described above. The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer corresponds one-to-one. The method includes: acquiring multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories differ from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; inputting the multiple positive electromagnetic signal samples into the slow system network to obtain a first [sample] output by each of the feature extraction layers. The feature vectors are processed, and the multiple negative samples of electromagnetic signals are input into the slow system network to obtain the second feature vectors output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and the second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. A first goodness is determined based on the first modulation feature, a second goodness is determined based on the second modulation feature, and a local loss of the feature modulation layer is determined based on the first goodness and the second goodness. The network parameters of the feature modulation layer are updated based on the local loss to perform online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.

[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0139] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An online learning method of a brain-like electromagnetic model, characterized in that, The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer corresponds one-to-one. The method includes: Acquire multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; The multiple positive electromagnetic signal samples are input into the slow system network to obtain the first feature vector output by each feature extraction layer, and the multiple negative electromagnetic signal samples are input into the slow system network to obtain the second feature vector output by each feature extraction layer. For each feature modulation layer, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is input into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and the second feature vector output by the same feature extraction layer is input into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. A first quality is determined based on the first modulation feature, a second quality is determined based on the second modulation feature, and the local loss of the feature modulation layer is determined based on the first quality and the second quality. The network parameters of the feature modulation layer are updated based on the local loss to enable online learning of the fast system network, which is used to identify the signal category of electromagnetic signals. 2.The online learning method of brain-like electromagnetic model according to claim 1, characterized in that, The step of determining the local loss of the feature modulation layer based on the first goodness and the second goodness includes: The local loss of the layer feature modulation layer is determined according to formula (1) or formula (2) layer feature modulation layer : (1) (2) in, This indicates the first level of excellence. This indicates the second degree of excellence. Indicates the temperature coefficient. This indicates the target quality of the positive sample of the electromagnetic signal. This indicates the target quality of the negative sample of the electromagnetic signal.

3. The online learning method of brain-like electromagnetic model according to claim 1, wherein, The step of inputting the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer includes: The modulation features output from the previous feature modulation layer are input into the feature modulation layer, and the modulation vector is determined by the modulation generator in the feature modulation layer. Based on the modulation vector, the first feature vector output by the feature extraction layer corresponding to the feature modulation layer is modulated to obtain the first modulation feature.

4. The online learning method of brain-like electromagnetic models according to any one of claims 1-3, characterized in that, The acquisition of multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category includes: The plurality of electromagnetic signal negative samples are obtained by sampling from the negative sample pool; wherein, the negative sample pool includes electromagnetic signal samples of different signal categories, and / or electromagnetic signal samples generated by perturbation; If the number of electromagnetic signal samples in the positive sample pool corresponding to the target signal category is greater than or equal to a first preset threshold, multiple electromagnetic signal positive samples are sampled from the positive sample pool corresponding to the target signal category.

5. The online learning method of brain-like electromagnetic model according to claim 4, wherein, The method further includes: The acquired electromagnetic signal samples are input into the slow system network to obtain the target features output by the slow system network; Determine the similarity between the target feature and each sample feature in the feature cache, wherein the sample features are features of stored electromagnetic signal samples; The similarities are sorted in descending order, and based on the first preset number of similarities, pseudo-labels and pseudo-label confidence scores of the electromagnetic signal samples are generated by weighted voting. If the confidence level of the pseudo-label is greater than or equal to the first preset confidence level, the electromagnetic signal sample and the target feature are added to the positive sample pool corresponding to the signal category represented by the pseudo-label.

6. The online learning method of brain-like electromagnetic model according to claim 5, wherein, The method further includes: If the confidence level of the pseudo-label is less than the first preset confidence level, the electromagnetic signal sample and the target feature are added to the new category candidate sample pool. If the number of samples in the new category candidate sample pool is greater than or equal to a second preset threshold, the target features in the new category candidate sample pool are clustered. If at least one cluster is successfully clustered, a new positive sample pool for the new signal category corresponding to each cluster is created, and the candidate electromagnetic signal samples belonging to the cluster and the corresponding target features are stored in the new positive sample pool.

7. The online learning method of brain-like electromagnetic model according to claim 4, wherein, The method further includes: The negative sample pool is constructed based on at least one of the following methods: Electromagnetic signal samples are randomly drawn from the positive sample pool corresponding to multiple signal categories, and the drawn electromagnetic signal samples of different signal categories are stored in the negative sample pool. The positive electromagnetic signal samples in the positive sample pool are perturbed to generate negative electromagnetic signal samples, and the generated negative electromagnetic signal samples are stored in the negative sample pool. Electromagnetic signal negative samples are obtained by randomly sampling from the environmental background signal and stored in the negative sample pool. Electromagnetic signal samples with a false label confidence level lower than the first preset confidence level are stored in the negative sample pool.

8. The online learning method of brain-like electromagnetic models according to any one of claims 1-3, characterized in that, The step of updating the network parameters of the feature modulation layer based on the local loss to perform online learning of the fast system network includes: The modulation generator in the currently trained feature modulation layer is used as a policy network. The modulation vector is output according to the current state of the feature modulation layer. The current state includes at least one of the following: the first feature vector and the second feature vector output by the feature extraction layer corresponding to the feature modulation layer, the modulation feature output by the previous feature modulation layer, and the pseudo-label confidence of the current electromagnetic signal sample. Based on the multiple negative electromagnetic signal samples and the multiple positive electromagnetic signal samples, a reward signal for the reinforcement learning process is determined. The reward signal includes a main reward, an auxiliary reward, and a sparse reward. The main reward is determined based on the proportion of correctly classified samples and / or misclassified samples among the multiple positive electromagnetic signal samples, or based on the recognition confidence of the multiple positive electromagnetic signal samples after modulation by the modulation vector. The auxiliary reward is determined based on a first goodness and a second goodness. The sparse reward is a reward configured when a new category is successfully identified or a preset task is completed. The reinforcement learning process is used to characterize the parameter update process of the currently trained feature modulation layer. Based on the reward signal, a policy loss is constructed, and based on the local loss and the policy loss, the parameters of the policy network are updated to obtain the updated feature modulation layer network parameters, so as to perform online learning on the fast system network.

9. The online learning method of brain-like electromagnetic models according to any one of claims 1-3, characterized in that, The method further includes: At preset intervals, the slow system network is subjected to self-supervised learning based on multiple unlabeled sample features in the feature cache. During the self-supervised learning process of the slow system network, the fast system network keeps its current parameters unchanged. The feature cache includes sample features of multiple electromagnetic signal samples.

10. An on-line learning device of a brain-like electromagnetic model, characterized by comprising: The neuromorphic electromagnetic model includes a fast system network and a slow system network. The fast system network includes at least one feature modulation layer, and the slow system network includes at least one feature extraction layer. Each feature modulation layer and each feature extraction layer correspond one-to-one. The device includes: The acquisition module is used to acquire multiple negative electromagnetic signal samples and multiple positive electromagnetic signal samples of the same signal category, wherein the multiple negative electromagnetic signal samples are samples whose signal categories are different from the positive electromagnetic signal samples, and / or samples generated after perturbing the positive electromagnetic signal samples; The input module is used to input the multiple positive electromagnetic signal samples into the slow system network to obtain the first feature vector output by each feature extraction layer, and to input the multiple negative electromagnetic signal samples into the slow system network to obtain the second feature vector output by each feature extraction layer; The input module is further configured to input the first feature vector output by the feature extraction layer corresponding to the feature modulation layer into the feature modulation layer to obtain the first modulation feature output by the feature modulation layer, and input the second feature vector output by the same feature extraction layer into the feature modulation layer to obtain the second modulation feature output by the feature modulation layer. The determination module is configured to determine a first quality level based on the first modulation feature, determine a second quality level based on the second modulation feature, and determine the local loss of the feature modulation layer based on the first quality level and the second quality level. An update module is used to update the network parameters of the feature modulation layer based on the local loss, so as to perform online learning of the fast system network, which is used to identify the signal category of electromagnetic signals.