Training methods for wireless sensing pre-trained models
By using local consistency rearrangement data augmentation and self-supervised comparative learning of Siamese networks, the problem of utilizing large-scale unlabeled wireless signal data is solved, improving the pre-training effect of wireless sensing models and the performance of downstream tasks.
Patent Information
- Application Number
- CN202411607710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing technologies struggle to effectively utilize large-scale unlabeled data from wireless signals for self-supervised comparative learning, resulting in poor performance of downstream tasks. Furthermore, existing methods suffer from synchronization and calibration issues in wireless signal labeling, increasing deployment complexity.
We employ a local consistency rearrangement data augmentation method, which constructs positive and negative sample pairs through phase alignment and local feature rearrangement. We then utilize Siamese networks for self-supervised comparative learning to avoid shortcut information and improve the model pre-training effect.
It significantly improves the performance of downstream tasks, reduces the need for large-scale labeled data, and improves the pre-training efficiency of wireless sensing models.
Smart Images

Figure CN119808876B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of signal processing and artificial intelligence, and more specifically, to a training method for a wireless sensing pre-training model based on asymmetric local consistency rearrangement. Background Art
[0002] Currently, the biggest obstacle to developing learning-based wireless sensing models is obtaining large-scale, high-quality manually labeled datasets. However, the lack of intuitiveness and interpretability of wireless signals makes them difficult to annotate manually like visual image data. Furthermore, the high correlation between wireless signals and the data acquisition environment necessitates large-scale data collection under various conditions, further increasing the difficulty of dataset acquisition.
[0003] To address these issues, a common approach is to use sensors of other modalities (such as visual cameras) for simultaneous data acquisition. By labeling the data from these other modalities, the wireless signal can be indirectly labeled. However, the additional overhead caused by synchronization and calibration issues between different modalities still limits the practical deployment and application of wireless systems.
[0004] While acquiring large-scale, high-quality manually labeled wireless signal data is challenging, large-scale unlabeled wireless data is relatively readily available. In recent years, self-supervised contrastive learning has facilitated the development of downstream tasks by learning general semantic representations from vast amounts of unlabeled data and fine-tuning pre-trained models using a small amount of labeled data. This approach has shown great potential in improving downstream task performance and reducing reliance on large-scale labeled data. However, existing self-supervised contrastive learning techniques are primarily designed for conventional data such as images and language. When directly applied to wireless signal data, they are prone to learning meaningless shortcuts, offering no benefit to downstream task development.
[0005] Therefore, based on the characteristics of wireless signals, designing a general and efficient self-supervised contrastive learning pre-training method for wireless signals to obtain an efficient wireless signal pre-training model is of great significance for large-scale wireless sensing applications. Summary of the Invention
[0006] In view of this, this disclosure provides a training method for a wireless sensing pre-trained model.
[0007] One aspect of this disclosure provides a training method for a wireless sensing pre-training model, comprising: performing phase alignment on the original signal information to obtain a sample original feature map of the original signal information; inputting the sample original feature map into a first neural network of the wireless sensing pre-training model to obtain sample original features; performing a first downsampling on the sample original feature map to obtain multiple initial local region features of the sample original feature map; performing feature rearrangement on the multiple initial local region features based on a first rearrangement rule to obtain multiple enhanced local region features; and inputting a sample enhanced feature map generated from the multiple enhanced local region features into a second neural network of the wireless sensing pre-training model to obtain sample enhanced features; constructing a similarity loss based on the sample original features and sample enhanced features, and training the wireless sensing pre-training model to obtain a trained wireless sensing pre-training model.
[0008] Another aspect of this disclosure provides a training method for a wireless sensing model, wherein the wireless sensing model includes functional modules and a first neural network module in the wireless sensing pre-trained model trained based on the training method of the wireless sensing pre-trained model of this disclosure, the functional modules being set after the first neural network module, and the training method for the wireless sensing model including: fine-tuning the first neural network module and the functional modules using labeled sample wireless signals.
[0009] Another aspect of this disclosure provides a training apparatus for a wireless sensing pre-training model, comprising: a phase alignment module for performing phase alignment on original signal information to obtain a sample original feature map of the original signal information; a first neural network module for inputting the sample original feature map into the first neural network of the wireless sensing pre-training model to obtain sample original features; a first downsampling module for performing a first downsampling on the sample original feature map to obtain multiple initial local region features of the sample original feature map; a first rearrangement module for performing feature rearrangement on the multiple initial local region features based on a first rearrangement rule to obtain multiple enhanced local region features; a second neural network module for inputting a sample enhanced feature map generated from the multiple enhanced local region features into the second neural network of the wireless sensing pre-training model to obtain sample enhanced features; and a training module for constructing a similarity loss based on the sample original features and sample enhanced features, training the wireless sensing pre-training model to obtain a trained wireless sensing pre-training model.
[0010] According to the above embodiments of this disclosure, a local consistency rearrangement data augmentation method is proposed. By introducing only a small amount of interference in a limited area, it can prevent the destruction of global semantics in the signal and maximize the shared information between the original signal and its augmented version. It is suitable for self-supervised contrastive learning methods for wireless signals and can be used for model pre-training on large-scale unlabeled wireless datasets, significantly improving the performance of downstream tasks and reducing the need for large-scale labeled data. Attached Figure Description
[0011] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 The illustration schematically shows an exemplary system architecture in which at least one of the methods of training a wireless sensing pre-trained model and training a wireless sensing model can be applied according to embodiments of the present disclosure.
[0013] Figure 2 A flowchart illustrating a training method for a wireless sensing pre-trained model according to an embodiment of the present disclosure is shown schematically.
[0014] Figure 3A A schematic diagram illustrating signal local consistency rearrangement data enhancement according to an embodiment of the present disclosure is shown.
[0015] Figure 3B The illustration schematically shows the effect of signal local consistency rearrangement data enhancement according to an embodiment of the present disclosure;
[0016] Figure 4 This illustration schematically shows the effect of data enhancement through local consistency rearrangement of the original signal according to an embodiment of the present disclosure;
[0017] Figure 5 The diagram illustrates the overall training process of a wireless sensing pre-trained model according to an embodiment of the present disclosure.
[0018] Figure 6 A flowchart illustrating a method for training a wireless sensing model according to an embodiment of the present disclosure is shown schematically.
[0019] Figure 7 A schematic diagram illustrating the development of downstream tasks based on a wireless sensing model according to an embodiment of the present disclosure is shown.
[0020] Figure 8 A block diagram of a training apparatus for a wireless sensing pre-trained model according to an embodiment of the present disclosure is shown schematically.
[0021] Figure 9 A block diagram schematically illustrates a training apparatus for a wireless sensing model according to an embodiment of the present disclosure; and
[0022] Figure 10 The diagram illustrates an electronic device suitable for implementing at least one of the following methods: a training method for a wireless sensing pre-trained model and a training method for a wireless sensing model, according to embodiments of the present disclosure. Detailed Implementation
[0023] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0026] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0027] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0028] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.
[0029] This disclosure proposes a training method for a wireless sensing pre-training model based on asymmetric local consistency rearrangement. Its core content includes the following parts: (1) constructing positive and negative sample pairs by using local consistency rearrangement data augmentation; (2) using asymmetric data augmentation to avoid shortcut information; (3) bringing positive sample pairs closer and moving negative sample pairs further apart during the comparative learning process.
[0030] Figure 1 An exemplary system architecture 100 is illustrated, according to embodiments of the present disclosure, in which at least one of a training method for a wireless sensing pre-trained model and a training method for a wireless sensing model can be applied. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] It should be noted that the training method for the wireless sensing pre-training model provided in this embodiment can generally be executed by server 105. Correspondingly, the training device for the wireless sensing pre-training model provided in this embodiment can generally be located in server 105. The training method for the wireless sensing pre-training model provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the training device for the wireless sensing pre-training model provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the training device method for the wireless sensing pre-training model provided in this embodiment can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the training device for the wireless sensing pre-training model provided in this embodiment can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0036] For example, the original signal information may be originally stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (e.g., the first terminal device 101, but not limited thereto), or it may be stored on an external storage device and imported into the first terminal device 101. Then, the first terminal device 101 may locally execute the training method of the wireless sensing pre-training model provided in the embodiments of this disclosure, or send the original signal information to other terminal devices, servers, or server clusters, and have the other terminal devices, servers, or server clusters that receive the original signal information execute the training method of the wireless sensing pre-training model provided in the embodiments of this disclosure.
[0037] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0038] Figure 2 A flowchart illustrating a training method for a wireless sensing pre-trained model according to an embodiment of the present disclosure is shown.
[0039] like Figure 2As shown, the method includes operations S201 to S206.
[0040] In operation S201, phase alignment is performed on the original signal information to obtain the original feature map of the original signal information sample.
[0041] According to embodiments of this disclosure, the original signal information can characterize information about the wireless signal received by the receiving antenna. Phase alignment can include, but is not limited to, alignment of frequency point phases on a single signal and alignment of adjacent antenna signals among multiple signals. The original feature map of the sample can characterize a two-dimensional feature map of the wireless signal after phase alignment recorded in a coordinate system with frequency as the horizontal axis and antenna as the vertical axis.
[0042] In operation S202, the original feature map of the sample is input into the first neural network of the wireless sensing pre-trained model to obtain the original features of the sample.
[0043] According to embodiments of this disclosure, the wireless sensing pre-training model can employ a Siamese network model or other neural network models with two branches, without limitation herein. The first neural network can extract features from the original feature map of the sample to obtain the original features of the sample.
[0044] In operation S203, the original feature map of the sample is downsampled for the first time to obtain multiple initial local region features of the original feature map of the sample.
[0045] According to embodiments of this disclosure, the first downsampling can be achieved by downsampling the matrix or by convolution, and is not limited to these methods.
[0046] For example, for a single frame of radar signal Where W and H represent the dimensions of the matrix, and S×S represents a small local region, the radar signal can be first downsampled as follows: ,in S is a global description of X. 2 It contains a local detail of X.
[0047] In operation S204, based on the first rearrangement rule, the features of multiple initial local regions are rearranged to obtain multiple enhanced local region features.
[0048] According to embodiments of this disclosure, the first rearrangement rule can characterize a pre-defined rearrangement rule, such as rearranging a feature at a certain position to another preset position, and is not limited to this. The aforementioned set rule is followed when rearranging features for each initial local region feature. Through feature rearrangement, data augmentation of the initial local region features can be achieved, resulting in enhanced local region features.
[0049] For example, perform the same rearrangement on each initial local region feature. The rearranged enhanced local region features can provide This can be obtained by executing formula (1):
[0050] Formula (1)
[0051] In formula (1), This represents the index of the local region within the global region. It can characterize the features of the initial local region. It can represent enhanced local region features.
[0052] In operation S205, the sample enhancement feature map generated by multiple enhanced local region features is input into the second neural network of the wireless sensing pre-trained model to obtain sample enhancement features.
[0053] According to embodiments of this disclosure, when the wireless sensing pre-training model employs a Siamese network model, the first neural network and the second neural network can represent two branches with shared model parameters. When the wireless sensing pre-training model employs a neural network model with two branches, the initial model parameters of the first neural network and the second neural network can be set to be the same.
[0054] According to embodiments of this disclosure, multiple enhanced local region features are recombined to obtain a sample enhancement feature map. A second neural network can extract features from the sample enhancement feature map to obtain sample enhancement features.
[0055] Figure 3A A schematic diagram illustrating signal local consistency rearrangement data enhancement according to an embodiment of the present disclosure is shown.
[0056] like Figure 3A As shown, for radar signals After the original feature map 310 is downsampled, it is locally expanded to obtain a set 320 of multiple initial local region features, including initial local region features 321. Based on the first rearrangement rule, the initial local region features 321 are locally rearranged to obtain enhanced local region features 331. Each initial local region feature in the set 320 of multiple initial local region features is subjected to the aforementioned local rearrangement operation according to the first rearrangement rule to obtain a set 330 of multiple enhanced local region features, including the aforementioned enhanced local region features 331. The sets 330 of multiple enhanced local region features are combined to obtain the enhanced feature map 340 of the original feature map 310 after local consistency rearrangement data enhancement.
[0057] Figure 3BThe illustration shows the effect of signal local consistency rearrangement data enhancement according to an embodiment of the present disclosure.
[0058] like Figure 3A and Figure 3B As shown, the aforementioned local consistency rearrangement method divides the signal into multiple non-overlapping local regions of the same size and applies the same rearrangement method to all local regions. This allows the locally consistent rearrangement data enhancement to satisfy singular value consistency, thereby maintaining the semantic consistency of the wireless signal before and after enhancement.
[0059] In operation S206, a similarity loss is constructed based on the original features and enhanced features of the samples, and the wireless sensing pre-trained model is trained to obtain the trained wireless sensing pre-trained model.
[0060] Through the above embodiments of this disclosure, a local consistency rearrangement data augmentation method is proposed. By introducing only a small amount of interference in a limited area, it can prevent the destruction of global semantics in the signal and maximize the shared information between the original signal and its augmented version. It is suitable for self-supervised contrastive learning methods for wireless signals and can be used for model pre-training on large-scale unlabeled wireless datasets, significantly improving the performance of downstream tasks and reducing the need for large-scale labeled data.
[0061] The following describes specific embodiments. Figure 2 The method shown will be further explained.
[0062] According to embodiments of this disclosure, in addition to performing local consistency rearrangement data enhancement in the spectral space of the signal, local consistency rearrangement data enhancement can also be performed in the original space of the signal. That is, before operation S201, local consistency rearrangement data enhancement of the original signal is first performed. This method may include: performing a second downsampling on the original signal information to obtain multiple initial local region information of the original signal information; rearranging the multiple initial local region information according to a second rearrangement rule to obtain multiple enhanced local region information; and generating data-enhanced original signal information based on the multiple enhanced local region information.
[0063] It should be noted that the second downsampling, the second rearrangement rule, and the information rearrangement in this method can have the same or similar implementation methods as the aforementioned first downsampling, the first rearrangement rule, and the feature rearrangement, which will not be elaborated here.
[0064] Figure 4 The illustration shows the effect of local consistency rearrangement data enhancement of the original signal according to an embodiment of the present disclosure.
[0065] like Figure 4As shown, performing a phase alignment operation on the original signal information 410 yields a first sample feature map 411. Performing local consistency rearrangement data enhancement on the original signal information 410 in its original space, followed by a phase alignment operation, yields a second sample feature map 412.
[0066] Through the above embodiments of this disclosure, local consistency rearrangement is performed in the original space of the signal, which is equivalent to distributing the energy of the target to multiple locations in the original space. This can enhance the spatial diversity of wireless signals and improve the feature learning performance of the pre-trained model.
[0067] According to embodiments of this disclosure, the above operation S201 may include: acquiring the frequency point phase difference representing the same received signal in the original signal information and the antenna phase difference between different received signals represented in the original signal information. Based on the frequency point phase difference and the antenna phase difference, the original signal information is phase-aligned to obtain a sample original feature map of the original signal information.
[0068] For example, consider a wireless signal originating from a transmitting antenna, reflected by a target located at (x, y), and finally received by a receiving antenna. The phase difference between adjacent frequency points in the received signal can be calculated using formula (2):
[0069] Formula (2)
[0070] In formula (2), The frequency difference between adjacent frequency points is represented by c, which represents the speed of light. This indicates the phase difference at a frequency point.
[0071] The phase difference between adjacent antennas can be calculated using formula (3):
[0072] Formula (3)
[0073] In formula (3), d represents the spacing between adjacent receiving antennas. Indicates the wavelength of adjacent receiving antennas. This indicates the antenna phase difference.
[0074] By compensating for phase shift and superimposing wireless signals from different frequencies and antennas, the signal from the target (x, y) will be enhanced, while signals from other locations will be suppressed. Therefore, the signal from the target (x, y) will be separated, and the extracted signal can be expressed as formula (4):
[0075] Formula (4)
[0076] In formula (4), This represents the wireless signal received by the device. 'a' and 'e' represent the receiving antenna index and signal frequency index, respectively, while 'A' and 'E' represent the total number of receiving antennas and the total number of signal frequencies, respectively. Represents the original feature map of the sample. It is an imaginary number.
[0077] Through the above embodiments of this disclosure, signal information can be preprocessed into original feature maps of samples, which is beneficial for subsequent local consistency rearrangement data enhancement.
[0078] According to embodiments of this disclosure, during the execution of the above operation S204, the possible number of rearrangements obtained by random feature rearrangement includes S. 2 Since not all rearrangements have the same importance, and considering that the proximity of signals is important, signal rearrangements can be restricted to column rearrangements and row rearrangements.
[0079] Based on this, the above operation S204 may include: for each initial local region feature, performing matrix multiplication with at least one of the row permutation matrix and column permutation matrix and the initial local region feature to obtain the enhanced local region feature.
[0080] For example, ,in Represents a row permutation matrix. This represents the row and column permutation matrix, and the rearranged signal can be expressed as formula (5):
[0081] Formula (5)
[0082] In formula (5), The initial global matrix represents the features of multiple initial local regions. This represents the global representation of the row permutation matrix corresponding to the initial global matrix. This represents the overall representation of the column permutation matrix corresponding to the overall matrix. This represents an overall enhancement matrix that enhances the features of multiple local regions.
[0083] Furthermore, the same rearrangement method is applied to all non-overlapping local regions, which is equivalent to rearranging the convolutional kernels of the first convolutional layer of the model. Therefore, the above formula (5) can also be replaced by formula (6):
[0084] Formula (6)
[0085] In formula (6), W is the convolution kernel parameter, and * refers to the convolution operation. yes The result of rearranging a local region is the aforementioned enhancement of local region features.
[0086] Through the above embodiments of this disclosure, the same rearrangement method is used in all local regions, making the local consistency rearrangement operation on the signal equivalent to rearranging the convolution kernel of the first layer of the model, thus providing an efficient local consistency rearrangement execution method.
[0087] Furthermore, during the aforementioned process of performing local consistency rearrangement data enhancement of the original signal, the rearrangement operation shown in formula (5) or formula (6) can be applied to the original wireless signal R. The rearranged original signal can be expressed as, for example, formula (7):
[0088] Formula (7)
[0089] Then, the phase alignment operation of S201 is performed. The essence of the original signal local consistency rearrangement enhancement is to spread the target energy to other locations, which can further enhance the spatial richness of the wireless signal.
[0090] According to embodiments of this disclosure, based on the above method, the original features of the signal and the enhanced features after local consistency rearrangement data enhancement can be obtained. Furthermore, positive and negative sample pairs can be constructed based on these features to support the training of the above operation S206.
[0091] In realizing the concept disclosed herein, the inventors discovered that the core of self-supervised contrastive learning lies in constructing positive and negative sample pairs to shorten the distance between positive sample pairs output by the network and widen the distance between negative sample pairs, thereby ensuring a certain consistency in the feature space and learning general semantic features. Existing positive and negative sample pairs are typically constructed based on data augmentation methods, which are mostly designed for image and vision data. When directly applied to wireless signal data, they can only learn some meaningless shortcut information. Therefore, how to design efficient data augmentation for wireless signals while avoiding shortcuts is the core challenge of applying self-supervised contrastive learning to wireless signals.
[0092] According to embodiments of this disclosure, the number of original signal information is multiple. Corresponding to the process of constructing positive and negative sample pairs, the above operation S206 may include: determining positive sample pairs based on the original sample features of the target original signal information and the sample enhancement features of the target original signal information, wherein the target original signal information represents any one of the multiple original signal information. Determining negative sample pairs based on the original sample features of the target original signal information and the sample enhancement features of other original signal information, wherein the other original signal information represents any one of the multiple original signal information other than the target original signal information. Constructing a similarity loss based on the first similarity metric of the positive sample pairs and the second similarity metric of the negative sample pairs, and training the wireless sensing pre-training model.
[0093] For example, an unlabeled dataset containing N raw signal information can be represented as: First, by using the aforementioned data augmentation operation of local consistency rearrangement of the original signal, a dataset of the original signal information can be obtained for data augmentation. Next, performing the phase alignment operation in operation S201 yields a dataset of the original feature maps of the data-augmented samples. Then, by combining the local consistency rearrangement data augmentation operations S203~S204, a dataset with augmented feature maps can be obtained. ,in This indicates that the data underwent two local consistency rearrangement operations: one for data augmentation based on the original signal's local consistency and the other for data augmentation based on the original feature maps of the samples. Furthermore, for unlabeled datasets... Performing only the phase alignment operation in operation S201 can yield a dataset of original feature maps of the samples without data augmentation. This embodiment can... These are called positive sample pairs. These are called negative sample pairs.
[0094] It is worth noting that conventional methods for constructing positive and negative sample pairs use symmetric data augmentation, where the sample pairs consist of augmented signals. This can lead to the model learning shortcut features. To address this issue, this disclosure employs an asymmetric data augmentation method, using the wireless signal and its augmented version to construct positive and negative sample pairs instead of using only augmented signals.
[0095] According to embodiments of this disclosure, the goal of pre-training a self-supervised contrastive learning model is to narrow the distance between features of positive sample pairs and exclude the distance between features of negative sample pairs, so that the feature space satisfies a certain consistency. To this end, the wireless sensing pre-training model can employ two branches f that share model parameters or have consistent initial parameter information. q (.) and f k (.)constitute.
[0096] For example, the first neural network is f q (.), the second neural network is f k (.). Corresponding to the similarity loss mentioned above, the objective of model training can be to minimize the InfoNCE loss function as shown in formula (8):
[0097] Formula (8)
[0098] In formula (8), , , , These represent the positive and negative sample pairs of features extracted by the model, respectively. For feature similarity measurement function, Here, K represents the temperature coefficient, and K represents the number of negative sample pairs.
[0099] Through the above embodiments of this disclosure, the asymmetric data augmentation method can significantly reduce shortcut features in the pre-trained model by maximizing the shared information between the original wireless signal and its augmented version, rather than maximizing the shared information between the two augmented versions, which is beneficial for training an efficient wireless sensing pre-trained model.
[0100] According to embodiments of this disclosure, corresponding to the process of adjusting model parameters, the above operation S206 may further include: during the t-th training round, adjusting the first model parameters in the first neural network based on the t-th similarity loss value of the similarity loss to obtain the first adjusted model parameters. Based on the momentum coefficient, the first adjusted parameters, and the second model parameters in the second neural network, performing momentum summation calculation to obtain the second adjusted model parameters in the second neural network, wherein the initial parameter information of the first model parameters is the same as the initial parameter information of the second model parameters.
[0101] According to embodiments of this disclosure, a Siamese network is typically a weight-sharing model that processes multiple inputs, including a gradient-computable branch for updating network parameters and a gradient-untrainable branch. The data domain of the entire pre-trained model depends on the input data of the gradient-trainable branch. Therefore, the raw signal information can be input into the gradient-trainable branch, i.e., the first neural network f. q (.), which inputs the enhanced version into the gradient-untrainable branch, i.e., the second neural network f. k (.), this operation avoids the model fitting to potential shortcut features introduced by data augmentation.
[0102] For f q The model parameters in (.) can be updated by backpropagation algorithm based on the InfoNCE loss function shown in formula (8).
[0103] For f k The model parameters in (.) can be updated by summing the momentum in formula (9):
[0104] Formula (9)
[0105] In formula (9), They represent f respectively q (.) and f k The model parameters of (.). =0.999 represents the momentum coefficient, and its value can be adjusted according to actual business needs, and is not limited to this.
[0106] Figure 5 The diagram illustrates the overall training process of a wireless sensing pre-trained model according to an embodiment of the present disclosure.
[0107] like Figure 5 As shown, the original feature map of the sample obtained after performing phase alignment on the original signal information 510 can be input into model f. q (.) 520, Extracting original features of the sample 521. The sample enhancement feature maps obtained by sequentially performing operations S501~S503 on the original signal information 510 can be input into model f. k (.) 530, Extract sample enhancement features (or 531.
[0108] In operation S501, local consistency rearrangement data enhancement is performed in the original space of the signal.
[0109] In operation S502, phase alignment is performed.
[0110] In operation S503, local consistency rearrangement data enhancement is performed in the spectral space of the signal.
[0111] According to embodiments of this disclosure, the number of negative sample pairs K is typically a large value, for example, K=16385. This results in the need for large batches of negative sample pair data, leading to high memory requirements and hardware costs. Therefore, as... Figure 5 As shown, a queue 540 can be used to store f during the training process. k The output of (.) only needs to be sampled from queue 540 each time. A sample, then with Negative sample pairs are formed, and the model is trained using InfoNCE loss.
[0112] It should be noted that, corresponding to the process of performing local consistency rearrangement data enhancement on the original signal information 510 to obtain the sample enhanced feature map, only operations S501~S502 or only operations S502~S503 can be performed, and no limitation is made here.
[0113] According to an embodiment of this disclosure, corresponding to the model determination process, the above operation S206 may further include: determining a trained wireless sensing pre-trained model based on the trained first neural network.
[0114] For example, after pre-training on large-scale unlabeled data, the final f q (.) represents the wireless sensing model officially used in actual business operations.
[0115] Through the above embodiments of this disclosure, a training method for a wireless sensing pre-trained model based on asymmetric local consistency rearrangement is realized. This method maximizes the shared information between the original radio frequency signal and its enhanced version by utilizing local consistency rearrangement data augmentation, introducing only minor interference in a limited area, thus preventing it from destroying the global semantics of the wireless signal. By utilizing asymmetric augmentation to eliminate shortcuts, an efficient wireless signal pre-trained model is obtained.
[0116] Figure 6 A flowchart illustrating a method for training a wireless sensing model according to an embodiment of the present disclosure is shown.
[0117] According to embodiments of this disclosure, the wireless sensing model includes functional modules and a first neural network module in a wireless sensing pre-trained model trained based on the above method. The functional modules are positioned after the first neural network module.
[0118] like Figure 6 As shown, the method includes operation S601.
[0119] When operating the S601, the first neural network module and functional modules are fine-tuned and trained using tagged sample wireless signals.
[0120] According to embodiments of this disclosure, for different downstream tasks, only a small amount of labeled data needs to be collected for f. q The development of downstream tasks can be completed with minor adjustments.
[0121] Figure 7 The illustration shows a schematic diagram of developing downstream tasks based on a wireless sensing model according to an embodiment of the present disclosure.
[0122] like Figure 7 As shown, the input can be a wireless signal 701. The pre-trained model 710 can use the first neural network module f in the wireless sensing pre-trained model trained based on the above method. q (.). Corresponding to different downstream tasks, functional module 720 can adopt any one of the following: classifier module 721, regression module 722, and decoder module 723.
[0123] For example, the first embodiment of this disclosure verifies the performance improvement and reduced requirement for labeled samples in the wireless sensing pre-trained model obtained by this disclosure in the action recognition task. Action recognition is a classification task, which only requires adding a classifier to the end of the pre-trained model and fine-tuning it with a small amount of labeled data. The performance metric is classification accuracy. The dataset contains four types of actions: single-person walking, standing, squatting, and multiple-person walking. The overall dataset is divided into three parts: pre-training dataset, training set, and test set. First, the model is pre-trained on the unlabeled pre-training dataset, then fine-tuned on the labeled training set, and finally the model performance is evaluated on the test set. The pre-training dataset contains 226,008 samples, the training set contains 11,760 samples, and the test set contains 4,557 samples.
[0124] Experimental results show that, without using the pre-trained model disclosed herein, training directly on the training set resulted in an action recognition accuracy of 85.211% on the test set. However, using the method disclosed herein—pre-training the model on the pre-training dataset, fine-tuning it on the training set, and finally testing it on the test set—the action recognition accuracy improved to 92.298%, a significant improvement of 7.087%. Further reducing the training set size to 50% (5880 samples) for fine-tuning the pre-trained model resulted in an action recognition accuracy of 91.703%; reducing the training set size to 10% (1176 samples) resulted in an accuracy of 89.217%, still 4.006% higher than the model trained from scratch. Therefore, the pre-trained model disclosed herein significantly improves the accuracy of action recognition while reducing the need for labeled data.
[0125] For example, the second embodiment of this disclosure verifies the performance improvement and reduced requirement for labeled samples in the 3D pose estimation task of the wireless sensing pre-trained model obtained in this disclosure. 3D pose estimation is a structure prediction task; it only requires adding a local proposal module to the end of the pre-trained model to extract the features of each individual, and then inputting the individual features into the 3D pose estimation network to complete the pose estimation. The performance metric is the mean joint error (in mm). The dataset contains 226,008 samples in the pre-training dataset, 47,304 samples in the training set, and 22,192 samples in the test set. The model is first pre-trained on the unlabeled pre-training dataset, then fine-tuned on the labeled training set, and finally evaluated on the test set.
[0126] Experimental results show that without using the pre-trained model of this disclosure, the pose estimation error trained directly on the training set is 134.90 mm. However, using the method of this disclosure, which first pre-trains on the pre-training dataset and then fine-tunes it on the training set, the pose estimation error on the final test set is 121.13 mm, a significant improvement of 13.77 mm in accuracy. If the training set size is reduced to 50% (23,652 samples), the fine-tuned pose estimation error is 126.76 mm; if the training set size is reduced to 10% (4,730 samples), the error is 147.28 mm. Even with a 50% reduction in the training set size, the model performance is still 8.14 mm higher than the model trained from scratch. Therefore, the pre-trained model of this disclosure significantly improves the accuracy of 3D pose estimation and reduces the need for labeled data.
[0127] For example, the third embodiment of this disclosure verifies the performance improvement and reduced requirement for labeled samples in the human contour generation task obtained by this disclosure using the wireless sensing pre-trained model. Human contour generation is a dense prediction task, requiring only the addition of a decoder module to the end of the pre-trained model and fine-tuning with a small amount of labeled data. The performance metric is the Intersection over Union (IoU). The dataset contains 226,008 samples in the pre-training dataset, 17,520 samples in the training set, and 4,675 samples in the test set. The model is first pre-trained on the unlabeled pre-training dataset, then fine-tuned on the labeled training set, and finally evaluated on the test set.
[0128] Experimental results show that, without using the pre-trained model of this disclosure, the IoU of human contour generation trained directly on the training set is 0.699. However, using the method of this disclosure, which first pre-trains on the pre-training dataset and then fine-tunes it on the training set, the IoU on the final test set increases to 0.727, a significant improvement of 0.028. If the training set size is reduced to 50% (8760 samples), the fine-tuned IoU is 0.712; if the training set size is reduced to 10% (1752 samples), the IoU is 0.686, only slightly lower than the model trained directly from scratch by 0.013. Therefore, the pre-trained model of this disclosure significantly improves the accuracy of human contour generation while reducing the need for labeled data.
[0129] According to the embodiments described above, the effectiveness of the proposed method can be verified through three radio frequency sensing tasks. Experimental results show that the present disclosure significantly improves the performance of various downstream tasks. Even using only 10% of the training dataset, the performance of the wireless sensing pre-trained model based on the present disclosure is still comparable to that of a model trained from scratch using 100% of the training dataset, demonstrating that the present disclosure can significantly reduce the need for large-scale labeled datasets for downstream task models.
[0130] Figure 8 A block diagram of a training apparatus for a wireless sensing pre-trained model according to an embodiment of the present disclosure is shown schematically.
[0131] like Figure 8 As shown, the training device 800 for the wireless sensing pre-trained model includes a phase alignment module 810, a first neural network module 820, a first downsampling module 830, a first rearrangement module 840, a second neural network module 850, and a training module 860.
[0132] The phase alignment module 810 is used to perform phase alignment on the original signal information to obtain the original feature map of the original signal information.
[0133] The first neural network module 820 is used to input the original feature map of the sample into the first neural network of the wireless sensing pre-trained model to obtain the original features of the sample.
[0134] The first downsampling module 830 is used to perform a first downsampling on the original feature map of the sample to obtain multiple initial local region features of the original feature map of the sample.
[0135] The first rearrangement module 840 is used to rearrange the features of multiple initial local regions based on the first rearrangement rule to obtain multiple enhanced local region features.
[0136] The second neural network module 850 is used to input the sample enhancement feature map generated by multiple enhanced local region features into the second neural network of the wireless sensing pre-trained model to obtain sample enhancement features.
[0137] Training module 860 is used to construct a similarity loss based on the original features and enhanced features of the samples, and to train the wireless sensing pre-trained model to obtain the trained wireless sensing pre-trained model.
[0138] According to embodiments of this disclosure, the training apparatus for the wireless sensing pre-trained model further includes a second downsampling module, a second rearrangement module, and an original signal data enhancement module.
[0139] The second downsampling module is used to perform a second downsampling on the original signal information to obtain multiple initial local region information of the original signal information.
[0140] The second rearrangement module is used to rearrange the information of multiple initial local regions based on the second rearrangement rules to obtain multiple enhanced local region information.
[0141] The original signal data enhancement module is used to generate data-enhanced original signal information based on multiple enhanced local area information.
[0142] According to embodiments of this disclosure, the phase alignment module includes a phase difference acquisition unit and a phase alignment unit.
[0143] The phase difference acquisition unit is used to acquire the frequency phase difference of the signal frequency points representing the same received signal in the original signal information, as well as the antenna phase difference between different received signals represented in the original signal information.
[0144] The phase alignment unit is used to perform phase alignment on the original signal information based on the frequency phase difference and the antenna phase difference, so as to obtain the original feature map of the original signal information.
[0145] According to embodiments of this disclosure, the first rearrangement module includes a locally consistent rearrangement unit.
[0146] The local consistency rearrangement unit is used to perform matrix multiplication with at least one of the row permutation matrix and column permutation matrix for each initial local region feature to obtain the enhanced local region feature.
[0147] According to embodiments of this disclosure, the number of original signal information is multiple. The training module includes a positive sample pair determination unit, a negative sample pair determination unit, and a similarity loss construction unit.
[0148] The positive sample pair determination unit is used to determine positive sample pairs based on the original sample features of the target original signal information and the sample enhancement features of the target original signal information, wherein the target original signal information represents any one of multiple original signal information.
[0149] The negative sample pair determination unit is used to determine negative sample pairs based on the original features of the target original signal information and the sample enhancement features of other original signal information. The other original signal information represents any one of the multiple original signal information other than the target original signal information.
[0150] The similarity loss construction unit is used to construct a similarity loss based on the first similarity metric of positive sample pairs and the second similarity metric of negative sample pairs, and to train the wireless sensing pre-trained model.
[0151] According to embodiments of this disclosure, the training module includes a first model parameter adjustment unit and a second model parameter calculation unit.
[0152] The first model parameter adjustment unit is used to adjust the first model parameters in the first neural network according to the t-th similarity loss value of the similarity loss during the t-th training round, so as to obtain the first adjusted model parameters.
[0153] The second model parameter calculation unit is used to calculate the momentum summation based on the momentum coefficient, the first adjusted parameters, and the second model parameters in the second neural network, so as to obtain the second adjusted model parameters in the second neural network. The initial parameter information of the first model parameters is the same as that of the second model parameters.
[0154] According to embodiments of this disclosure, the training module includes a pre-trained model determination unit.
[0155] The pre-trained model determination unit is used to determine the trained wireless sensing pre-trained model based on the trained first neural network.
[0156] Figure 9 A block diagram schematically illustrates a training apparatus for a wireless sensing model according to an embodiment of the present disclosure. The wireless sensing model 900 includes functional modules and a first neural network module in a wireless sensing pre-trained model trained using a training method based on a wireless sensing pre-trained model. The functional modules are disposed after the first neural network module.
[0157] like Figure 9 As shown, the training device 900 for the wireless sensing model includes a fine-tuning module 910.
[0158] The fine-tuning module 910 is used to fine-tune the training of the first neural network module and the functional module using labeled sample wireless signals.
[0159] According to embodiments of this disclosure, the functional modules include any one of the following: a classifier module, a regression module, and a decoder module.
[0160] Any one or more of the modules or units according to embodiments of this disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules or units according to embodiments of this disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules or units according to embodiments of this disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented by hardware or firmware in any other reasonable manner by integrating or packaging the circuitry, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, one or more of the modules or units according to embodiments of this disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0161] For example, any and multiple of the phase alignment module 810, the first neural network module 820, the first downsampling module 830, the first rearrangement module 840, the second neural network module 850, and the training module 860, or the fine-tuning module 910, can be combined into one module / unit, or any one of these modules / units can be split into multiple modules / units. Alternatively, at least some of the functionality of one or more of these modules / units can be combined with at least some of the functionality of other modules / units and implemented in one module / unit. According to embodiments of this disclosure, at least one of the phase alignment module 810, the first neural network module 820, the first downsampling module 830, the first rearrangement module 840, the second neural network module 850, and the training module 860, or the fine-tuning module 910, can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable method of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three methods. Alternatively, at least one of the phase alignment module 810, the first neural network module 820, the first downsampling module 830, the first rearrangement module 840, the second neural network module 850, and the training module 860, or the fine-tuning module 910, can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0162] It should be noted that the training device part of the wireless sensing pre-training model in the embodiments of this disclosure corresponds to the training method part of the wireless sensing pre-training model in the embodiments of this disclosure. For a detailed description of the training device part of the wireless sensing pre-training model, please refer to the training method part of the wireless sensing pre-training model, which will not be repeated here.
[0163] The training device part of the wireless sensing model in the embodiments of this disclosure corresponds to the training method part of the wireless sensing model in the embodiments of this disclosure. For a detailed description of the training device part of the wireless sensing model, please refer to the training method part of the wireless sensing model, which will not be repeated here.
[0164] Figure 10 The diagram illustrates an electronic device suitable for implementing at least one of the following methods: a training method for a wireless sensing pre-trained model and a training method for a wireless sensing model, according to embodiments of the present disclosure. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0165] like Figure 10 As shown, an electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage portion 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0166] RAM 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Processor 1001 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1002 and / or RAM 1003. It should be noted that the programs may also be stored in one or more memories other than ROM 1002 and RAM 1003. Processor 1001 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0167] According to embodiments of this disclosure, the electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to a bus 1004. The system 1000 may also include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.
[0168] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by processor 1001, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0169] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement at least one of the training methods for a wireless sensing pre-trained model and a training method for a wireless sensing model according to embodiments of this disclosure.
[0170] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0171] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 1002 and / or RAM 1003 described above and / or one or more memories other than ROM 1002 and RAM 1003.
[0172] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement at least one of the training methods for the wireless sensing pre-trained model and the training method for the wireless sensing model provided in the embodiments of this disclosure.
[0173] When the computer program is executed by the processor 1001, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0174] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1009, and / or installed from a removable medium 1011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0175] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not expressly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0177] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A behavior recognition method based on a wireless sensing pre-trained model, comprising: Phase alignment is performed on the original signal information to obtain the original feature map of the original signal information; The original feature map of the sample is input into the first neural network of the wireless sensing pre-trained model to obtain the original features of the sample. The original feature map of the sample is downsampled to obtain multiple initial local region features of the original feature map of the sample; Based on the first rearrangement rule, the features of the multiple initial local regions are rearranged to obtain multiple enhanced local region features; and The sample augmentation feature map generated from multiple enhanced local region features is input into the second neural network of the wireless sensing pre-trained model to obtain sample augmentation features; A similarity loss is constructed based on the original features and the enhanced features of the samples. The wireless sensing pre-training model is then trained to obtain a trained wireless sensing pre-training model. A classifier is added to the end of the pre-training model and fine-tuned using a dataset to obtain a behavior recognition model. The dataset contains four types of behaviors: single person walking, standing, squatting, and multiple people walking. The step of performing phase alignment on the original signal information to obtain a sample original feature map of the original signal information includes: obtaining the frequency point phase difference of the signal frequency points representing the same received signal in the original signal information and the antenna phase difference between different received signals represented in the original signal information; performing phase alignment on the original signal information based on the frequency point phase difference and the antenna phase difference to obtain a sample original feature map of the original signal information; The step of rearranging the features of the multiple initial local region features to obtain multiple enhanced local region features includes: for each initial local region feature, performing matrix multiplication with at least one of the row permutation matrix and column permutation matrix on the initial local region feature to obtain the enhanced local region feature.
2. The method according to claim 1, further comprising: Before performing phase alignment on the original signal information. The original signal information is downsampled a second time to obtain multiple initial local region information of the original signal information; Based on the second rearrangement rule, the information of multiple initial local regions is rearranged to obtain multiple enhanced local region information; as well as Based on the multiple enhanced local region information, the original signal information for data enhancement is generated.
3. The method according to claim 1, wherein, The number of original signal information is multiple; the step of constructing a similarity loss based on the original features of the samples and the enhanced features of the samples, and training the wireless sensing pre-training model includes: Positive sample pairs are determined based on the original features of the target original signal information and the sample enhancement features of the target original signal information, wherein the target original signal information represents any one of the multiple original signal information; Negative sample pairs are determined based on the original features of the target original signal information and the sample enhancement features of other original signal information, wherein the other original signal information represents any one of the plurality of original signal information other than the target original signal information; and The similarity loss is constructed based on the first similarity metric of the positive sample pairs and the second similarity metric of the negative sample pairs, and the wireless sensing pre-training model is trained.
4. The method according to claim 1 or 3, wherein, The step of constructing a similarity loss based on the original features and enhanced features of the samples, and training the wireless sensing pre-training model, includes: During the t-th training round, the first model parameters in the first neural network are adjusted based on the t-th similarity loss value of the similarity loss, resulting in the first adjusted model parameters; and Based on the momentum coefficient, the first adjusted parameter, and the second model parameter in the second neural network, momentum summation is performed to obtain the second adjusted model parameter in the second neural network. The initial parameter information of the first model parameter is the same as the initial parameter information of the second model parameter.
5. The method according to claim 1, wherein, The step of training the wireless sensing pre-trained model to obtain the trained wireless sensing pre-trained model includes: Based on the trained first neural network, a trained wireless sensing pre-trained model is determined.
6. A training method for a wireless sensing model, wherein, The wireless sensing model includes functional modules and a first neural network module in a wireless sensing pre-trained model trained according to the method of any one of claims 1-5, wherein the functional modules are disposed after the first neural network module, and the method includes: The first neural network module and the functional module are fine-tuned and trained using tagged sample wireless signals.
7. The method according to claim 6, wherein, The functional modules include any one of the following: classifier module, regression module, and decoder module.
8. A behavior recognition device based on a wireless sensing pre-trained model, comprising: A phase alignment module is used to perform phase alignment on the original signal information to obtain a sample original feature map of the original signal information. The first neural network module is used to input the original feature map of the sample into the first neural network of the wireless sensing pre-trained model to obtain the original features of the sample. The first downsampling module is used to perform a first downsampling on the original feature map of the sample to obtain multiple initial local region features of the original feature map of the sample; The first rearrangement module is used to rearrange the features of multiple initial local regions based on a first rearrangement rule, thereby obtaining multiple enhanced local region features; and The second neural network module is used to input the sample enhancement feature map generated by multiple enhanced local region features into the second neural network of the wireless sensing pre-trained model to obtain sample enhancement features; The training module is used to construct a similarity loss based on the original features of the sample and the enhanced features of the sample, train the wireless sensing pre-training model to obtain the trained wireless sensing pre-training model, add a classifier to the end of the pre-training model, and fine-tune it using a dataset to obtain a behavior recognition model. The dataset contains four types of behaviors: single person walking, standing, squatting, and multiple people walking. The phase alignment module includes a phase difference acquisition unit and a phase alignment unit; The phase difference acquisition unit is used to acquire the frequency point phase difference of the signal frequency points representing the same received signal in the original signal information and the antenna phase difference between different received signals represented in the original signal information. The phase alignment unit is used to perform phase alignment on the original signal information based on the frequency point phase difference and the antenna phase difference to obtain a sample original feature map of the original signal information; The first rearrangement module includes a local consistency rearrangement unit; the local consistency rearrangement unit is used to perform matrix multiplication calculation with at least one of the row permutation matrix and column permutation matrix and the initial local region feature for each initial local region feature to obtain the enhanced local region feature.
Citation Information
Patent Citations
Remote sensing intelligent extraction method for disturbance range of large-scale artificial water and soil loss
CN116580320A
Neural function prediction model training method and device and neural function prediction method and device
CN117520851A