Wireless Authentication Method and System Based on Sparse Attention Mask
By converting wireless radio frequency data into 4D signal tensor samples and using sparse encoder-decoder module for self-recovery pre-training, the problem of lack of generalization and robustness in wireless identity authentication technology is solved, and higher identity authentication accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510329146.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing wireless identity authentication technology is difficult to physically model in the actual environment, and the labeling of wireless signal data is difficult, resulting in a lack of generalization and robustness of learning-based methods.
Using a sparse attention mask-based method, the original RF data sample is characterized as a 4D signal tensor sample, and self-recovery iterative pre-training is performed through the sparse encoder-decoder module to generate a pre-trained encoder for identity authentication model.
It improves the accuracy and generalization of wireless identity authentication, enhances the model's processing performance of RF signals, saves storage space and speeds up computing speed.
Smart Images

Figure CN119854789B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of signal processing, artificial intelligence, and network security, and particularly relates to a wireless identity authentication method and system, an electronic device, and a storage medium based on a sparse attention mask. Background Art
[0002] At present, identity authentication has become an important research direction in the field of network security. Different from traditional authentication methods such as fingerprints or vision, wireless signals have the characteristics of non-contact operation, natural penetrability, and high privacy in the network space. Based on these advantages, wireless sensing technology shows great application potential in the direction of secure identity authentication, and is used to explore potential ways to complete secure identity authentication tasks in a more concealed manner in actual scenarios.
[0003] Wireless identity authentication can utilize signals to sense the unique behavioral information or physiological information of different people, extract individual unique features therefrom, and ultimately perform secure identity authentication. However, existing technical solutions have many technical problems. For example, it is difficult to perform physical modeling for wireless identity authentication in an actual environment, and deep learning technology usually needs to be combined to meet more complex sensing requirements such as obtaining identity features. In addition, wireless signal data is different from vision or text, which is difficult to intuitively understand and label, resulting in the learning-based method being easily limited by the scale of labeled data. Therefore, under limited data samples, existing learning-based wireless identity authentication technical solutions lack generalization and robustness in actual applications. Summary of the Invention
[0004] In view of the above technical problems, the present invention provides a wireless identity authentication method and system, an electronic device, and a storage medium based on a sparse attention mask, which are used to solve at least one of the above technical problems.
[0005] According to a first aspect of the present invention, there is provided a wireless identity authentication method based on a sparse attention mask, including:
[0006] Using signal processing technology to represent the original radio frequency data sample as a 4D signal tensor sample;
[0007] Randomly generating a preliminary set of dense regions on the 4D signal tensor sample, sorting and screening the preliminary set of dense regions to obtain a set of target information-dense regions, and using the set of target information-dense regions to generate a sparse attention mask for the 4D signal tensor sample to obtain a masked 4D signal tensor sample;
[0008] Using the masked 4D signal tensor sample and a preset reconstruction loss function to perform self-recovery iterative pre-training on the sparse encoder-decoder module to obtain a pre-trained encoder, and using the pre-trained encoder as the encoder of the identity authentication model;
[0009] Iteratively train an identity authentication model using a wireless training dataset with human body contour labels and pose labels and an identity authentication loss function to obtain a trained identity authentication model;
[0010] Process the radio frequency signals in the target area using the trained identity authentication model to obtain the identity authentication result of the person in the target area.
[0011] According to an embodiment of the present invention, the above-mentioned characterizing the original radio frequency data sample as a 4D signal tensor sample using signal processing technology includes:
[0012] Perform signal processing on the original radio frequency data sample to obtain the angle of arrival-time of flight and Doppler frequency shift of the original radio frequency data sample, where the original radio frequency data sample is the radio frequency data of a dynamic target in the experimental area;
[0013] Use the amplitude image of the original radio frequency data sample as the reconstruction target, perform vector splicing based on the time dimension on the angle of arrival-time of flight and Doppler frequency shift to obtain a 4D signal tensor sample; or
[0014] Use the amplitude image of the original radio frequency data sample as the reconstruction target, process the angle of arrival-time of flight to obtain a 4D signal tensor sample.
[0015] According to an embodiment of the present invention, the above-mentioned sorting and screening the initial dense area set to obtain the target information dense area set includes:
[0016] Based on the correspondence between the signal high-energy area and the information density, sort the areas in the initial dense area set based on the total signal energy to obtain the sorted initial dense area set;
[0017] Based on a greedy strategy, screen the sorted initial dense area set, and select multiple non-overlapping areas with higher rankings in the sorted initial dense area set to construct the target information dense area set.
[0018] According to an embodiment of the present invention, the above-mentioned generating a sparse attention mask for the 4D signal tensor sample using the target information dense area set to obtain the masked 4D signal tensor sample includes:
[0019] Apply a preset mask ratio to each area in the target information dense area set, and mask the areas in the 4D signal tensor sample that do not belong to the target information dense area set to obtain the masked 4D signal tensor sample.
[0020] According to an embodiment of the present invention, the above-mentioned iterative pre-training of the sparse encoder-decoder module with self-recovery using the masked 4D signal tensor samples and the preset reconstruction loss function to obtain the pre-trained encoder includes:
[0021] Construct a sparse encoder-decoder module for human feature processing, and improve the reconstruction loss function to obtain a preset reconstruction loss function;
[0022] Use the encoder of the sparse encoder-decoder module to extract semantic features from the masked 4D signal tensor samples to obtain semantic feature samples;
[0023] Use the decoder of the sparse encoder-decoder module to map and reconstruct the semantic feature samples to obtain the 4D signal tensor samples after mask self-recovery;
[0024] Use the preset reconstruction loss function to process the 4D signal tensor samples and the 4D signal tensor samples after mask self-recovery to obtain a reconstruction loss value;
[0025] Use the reconstruction loss value to update the parameters of the sparse encoder-decoder module to obtain the sparse encoder-decoder module with updated parameters;
[0026] Iteratively perform semantic feature extraction operations, mapping reconstruction operations, reconstruction loss value calculation operations, and model parameter update operations until the first preset training condition is met to obtain the pre-trained encoder.
[0027] According to an embodiment of the present invention, the above-mentioned preset reconstruction loss function is shown in the following formula:
[0028] ,
[0029] Where, represents the set of target information dense regions, represents the decoder in the sparse encoder-decoder module, represents the encoder in the sparse encoder-decoder module, represents the mask corresponding to the set of target information dense regions, represents the L2 norm.
[0030] According to an embodiment of the present invention, the above-mentioned iterative training of the identity authentication model using the wireless training dataset with human contour labels and pose labels and the identity authentication loss function to obtain the trained identity authentication model includes:
[0031] Based on the preset weights, weight the identity classification loss function, the human unique pose information estimation error loss function, and the human unique contour information generation loss function to obtain a preset identity authentication loss function;
[0032] Process the 4D signal tensor sample using an identity authentication model based on a pre-trained encoder to obtain the identity authentication result of the dynamic experimental target, where the identity authentication result of the dynamic experimental target includes human contour information and pose information;
[0033] Process the identity authentication result of the dynamic experimental target and the human contour label and pose label of the dynamic experimental target in the wireless data sample using a preset identity authentication loss function to obtain an identity authentication loss value;
[0034] Update the parameters of the identity authentication model using the identity authentication loss value to obtain an identity authentication model with updated parameters;
[0035] Iteratively perform model data processing operations, loss value calculation operations, and model parameter update operations until the second preset training condition is met to obtain a trained identity authentication model.
[0036] The second aspect of the present invention provides a wireless identity authentication system based on a sparse attention mask, including:
[0037] A signal characterization module for characterizing the original radio frequency data sample as a 4D signal tensor sample using signal processing techniques;
[0038] A sparse attention mask module for randomly generating an initial set of dense regions on the 4D signal tensor sample, sorting and screening the initial set of dense regions to obtain a set of target information dense regions, and using the set of target information dense regions to generate a sparse attention mask for the 4D signal tensor sample to obtain a masked 4D signal tensor sample;
[0039] A self-recovery pre-training module for iteratively pre-training the sparse encoder-decoder module in a self-recovery manner using the masked 4D signal tensor sample and a preset reconstruction loss function to obtain a pre-trained encoder, and using the pre-trained encoder as the encoder of the identity authentication model;
[0040] An identity authentication model training module for iteratively training the identity authentication model using a wireless training data set with human contour labels and pose labels and an identity authentication loss function to obtain a trained identity authentication model;
[0041] A wireless identity authentication module for processing the wireless radio frequency signal of the target area using the trained identity authentication model to obtain the identity authentication result of the person in the target area.
[0042] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0043] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0044] The wireless identity authentication method based on a sparse attention mask provided by the present invention is a novel sparse perception mask auto-encoding framework for radio frequency sensing tasks; a pre-training framework that uses self-supervised learning based on (Masked AutoEncoder) MAE to obtain general semantic representations from large-scale radio frequency signal datasets, especially improving the ability to obtain unique human behavior characteristics. In addition, the present invention proposes a sparse perception masking strategy that can effectively bridge the information gap in radio frequency signals. By focusing on small information-dense areas and discarding less-information areas, not only the radio frequency sensing performance is improved, but also the computational and memory resource consumption is optimized. At the same time, the present invention redesigns the above sparse encoder-decoder module for high-dimensional RF data cubes, improving the processing performance of the identity authentication model for RF signals, saving storage space and accelerating the calculation speed; and considering the inherent sparsity of RF signals, a sparse mask scheme is provided, thus significantly improving the ability of the model pre-training to extract individual characteristics of personnel, and then showing higher accuracy and stronger generalization ability in the identity authentication model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0046] Figure 1 is an application scenario diagram of the wireless identity authentication method based on a sparse attention mask according to an embodiment of the present invention;
[0047] Figure 2 is a flowchart of the wireless identity authentication method based on a sparse attention mask according to an embodiment of the present invention;
[0048] Figure 3 is a training framework diagram of a sparse encoder-decoding module and an identity authentication model according to an embodiment of the present invention;
[0049] Figure 4 is a schematic flowchart of generating a sparse attention mask according to an embodiment of the present invention;
[0050] Figure 5Schematic diagram of the pre-training framework of the sparse encoder-decoder module according to an embodiment of the present invention;
[0051] Figure 6 Block diagram of the encoder shared by the pre-training framework and the identity authentication model according to an embodiment of the present invention;
[0052] Figure 7 Effect diagram of the comparative experiment of the wireless identity authentication method according to an embodiment of the present invention;
[0053] Figure 8 Block diagram of the structure of the wireless identity authentication system based on the sparse attention mask according to an embodiment of the present invention;
[0054] Figure 9 Block diagram of the electronic device suitable for implementing the wireless identity authentication method based on the sparse attention mask according to an embodiment of the present invention. Detailed implementation manners
[0055] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0056] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0057] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0058] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0059] To address the existing challenges in wireless identity authentication tasks, it is necessary to utilize self-supervised pre-training methods to enhance the model's ability to perceive individual identity features from large-scale unlabeled data, and ultimately achieve the goal of significantly improving the accuracy, generalization, and robustness of wireless identity authentication, better promoting large-scale wireless sensing applications represented by wireless identity authentication, and demonstrating its good development prospects in the fields of network security and wireless sensing. Therefore, the present invention designs an efficient self-supervised pre-training framework based on a masked autoencoder specific to wireless signals, and uses this framework to perform efficient self-supervised learning on large-scale unlabeled wireless signal data, thereby obtaining an encoder-decoder that has been pre-trained to learn human unique features, and further applying it to the identity authentication model. The model directly generates the unique action postures and body contours of personnel from the input radar signals, and performs identity authentication and recognition based on the subtle differences in personal body shapes and movement habits. With the help of the prior knowledge obtained through pre-training, the model can achieve a high identity authentication accuracy with only a small amount of labeled data for fine-tuning, as well as stronger robustness across different environments and different personnel. Obviously, this contactless method can potentially complete the identity authentication task, and the efficient self-supervised pre-training framework based on the masked autoencoder proposed by the present invention can obtain deep human unique features at low cost, improve the identity authentication accuracy and generalization, and demonstrate greater application value in practical scenarios.
[0060] Figure 1 FIG. is an application scenario diagram of a wireless identity authentication method based on a sparse attention mask according to an embodiment of the present invention.
[0061] As Figure 1 shown, the application scenario 100 according to this embodiment may include signal processing, artificial intelligence, network security, etc. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0062] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0063] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0064] Server 105 may be a server that provides various services, such as a background management server (merely for example) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0065] It should be noted that the wireless identity authentication method based on a sparse attention mask provided in the embodiments of the present invention can generally be executed by server 105. Correspondingly, the wireless identity authentication system based on a sparse attention mask provided in the embodiments of the present invention can generally be set in server 105. The wireless identity authentication method based on a sparse attention mask provided in the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the wireless identity authentication system based on a sparse attention mask provided in the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0066] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0067] It should be specifically noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0068] Meanwhile, in the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. And the processing of collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the relevant data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0069] In addition, in the scenario of making automated decisions using personal information, the methods, devices, and systems provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making decisions. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in work in a specific field, have specialized experience, knowledge, and skills, and have reached a certain professional level.
[0070] Based on the Figure 1 scenario described below, through Figures 2 to 7 a detailed description of the wireless identity authentication method based on a sparse attention mask for the disclosed embodiments will be given.
[0071] Figure 2 is a flowchart of the wireless identity authentication method based on a sparse attention mask according to an embodiment of the present invention.
[0072] As Figure 2 shown, the wireless identity authentication method based on a sparse attention mask of this embodiment includes operations S210 to S250.
[0073] In operation S210, the original radio frequency data samples are characterized as 4D signal tensor samples using signal processing techniques.
[0074] The above-mentioned original radio frequency data samples can be wireless WiFi signals or radar signals.
[0075] By performing signal processing on the above-mentioned original radio frequency data samples, the characteristic information of dynamic experimental targets in a specific experimental area can be obtained, such as the azimuth of the dynamic experimental targets.
[0076] In operation S220, a dense area initial selection set is randomly generated on the 4D signal tensor samples, sorted and screened to obtain a target information dense area set, and a sparse attention mask is generated for the 4D signal tensor samples using the target information dense area set to obtain a masked 4D signal tensor sample.
[0077] In the above operation S220, masking through a sparse mask enables relevant technical personnel to focus more on the information dense areas of dynamic experimental targets, and this method can greatly reduce the consumption of computing and memory resources.
[0078] In operation S230, the sparse encoder-decoder module is iteratively pre-trained self-recoverably using the masked 4D signal tensor samples and a preset reconstruction loss function to obtain a pre-trained encoder, and the pre-trained encoder is used as the encoder of the identity authentication model.
[0079] The sparse encoder-decoder module can self-recover the mask in the masked 4D signal tensor samples to obtain the 4D signal tensor samples after mask self-recovery.
[0080] The above-mentioned sparse encoder-decoder is redesigned according to the characteristics of RF data to reduce the consumption of computing and memory resources and improve the generalization of the pre-trained encoder.
[0081] In operation S240, the identity authentication model is iteratively trained using a wireless training dataset with human contour labels and pose labels and an identity authentication loss function to obtain a trained identity authentication model.
[0082] The encoder of the above identity authentication model utilizes the encoder trained in operation S230.
[0083] In operation S250, the wireless radio frequency signals in the target area are processed using the trained identity authentication model to obtain the person identity authentication result in the target area.
[0084] In an embodiment of the present invention, before obtaining the user's information, the consent or authorization of the user can be obtained. For example, before operation S250, a request to obtain the user's information can be sent to the user. When the user consents or authorizes to obtain the user's information, the operation S250 is executed.
[0085] The wireless authentication method based on sparse attention masks provided by the present invention is a novel sparse-aware mask auto-encoding framework for radio frequency sensing tasks; it is a pre-training framework that uses self-supervised learning based on Masked AutoEncoder (MAE) to obtain general semantic representations from large-scale radio frequency signal datasets, especially improving the ability to obtain unique human behavior characteristics. In addition, the present invention proposes a sparse-aware masking strategy that can effectively bridge the information gap in radio frequency signals. By focusing on small information-dense regions and discarding less-informative regions, it not only improves radio frequency sensing performance but also optimizes computational and memory resource consumption. At the same time, the present invention redesigned the above-mentioned sparse encoder-decoder module for high-dimensional RF data cubes, improving the processing performance of the authentication model for RF signals, saving storage space and accelerating calculation speed; and because the inherent sparsity of RF signals is considered, a sparse mask scheme is provided, thus significantly improving the ability of the model to pre-train and extract individual characteristics of personnel, and then showing practical authentication advantages of higher accuracy and stronger generalization ability in the authentication model.
[0086] The following is a further detailed description of the architecture of the wireless authentication method based on sparse attention masks provided by the present invention through specific embodiments in combination with the attached Figure 3 drawings.
[0087] Figure 3 It is a training framework diagram of a sparse encoder-decoder module and an authentication model according to an embodiment of the present invention.
[0088] As Figure 3 shown, the present invention first designs a general mask auto-encoder pre-training framework specific to wireless signals. Given an RF (Radio Freqency) signal frame, signal processing algorithms are used to create signal representations consistent with the requirements of downstream RF sensing tasks. The signal representations are processed through sparse attention masks to obtain their masked versions. The masked signal representations are input into the encoder backbone to obtain high-level semantic representations. At the same time, the decoder attempts to predict the masked part. Therefore, the reconstruction objective is to capture high-quality unique individual information representations contained in RF signals by minimizing the loss. After an efficient pre-training process, the encoder that has obtained deep individual unique behavior characteristics is directly applied to the training of the authentication model to greatly improve the accuracy of authentication. Specifically, the present invention is mainly implemented in the following five technical modules, namely: (1) signal representation; (2) sparse attention mask; (3) encoder-decoder construction; (4) reconstruction loss; (5) authentication model training module.
[0089] According to an embodiment of the present invention, the above-mentioned characterization of the original RF data samples as 4D signal tensor samples by using signal processing technology includes: performing signal processing on the original RF data samples to obtain the angle of arrival-time of flight and Doppler frequency shift of the original RF data samples, where the original RF data samples are wireless RF data of dynamic targets in the experimental area; using the amplitude image of the original RF data samples as the reconstruction target, performing vector splicing based on the time dimension on the angle of arrival-time of flight and Doppler frequency shift to obtain 4D signal tensor samples; or using the amplitude image of the original RF data samples as the reconstruction target, processing the angle of arrival-time of flight to obtain 4D signal tensor samples.
[0090] In the process of obtaining the 4D signal tensor samples, they can be obtained by splicing the angle of arrival-time of flight and Doppler frequency shift; or they can be obtained only by using the angle of arrival-time of flight.
[0091] The following further elaborates on the above process of obtaining the 4D signal tensor samples through specific embodiments.
[0092] Given the original RF data, such as radar signals, the present invention can use various mature signal processing techniques to obtain their corresponding signal representations, such as the angle of arrival-time of flight (Angle of Arrival (AoA)-Time of Flight (ToF)) and Doppler frequency shift (Doppler Frequency Shift (DFS)).
[0093] Among them, the angle of arrival-time of flight determines the angle at which the signal arrives through a multi-antenna array for direction estimation, and estimates the distance through the time difference between the signal traveling from the base station to the moving object and then back. It can be expressed by formula (1):
[0094] (1),
[0095] Where is the AoA, is the ToF, is the spatial interval between two adjacent antennas, is the signal frequency, is the speed of light, is the difference between adjacent frequencies, is the received signal, and are the indices of the receiver antenna and signal frequency. The output of AoA-ToF is a two-dimensional matrix. Note that due to differences in resolution, the matrices obtained from radar and WiFi usually have significant differences.
[0096] Among them, the Doppler frequency shift (DFS) represents the change in the signal path length reflected by the target due to the movement of the target, thereby generating a certain frequency shift in the observed signal. By performing time-frequency analysis on the CSI power, the present invention can extract the DFS caused by the target movement from the dynamic change of the power. Operations such as conjugate division, principal component analysis (PCA) algorithm, and short-time Fourier transform (STFT) are applied to the CSI data to extract Doppler information. The output of the DFS is a frequency-time matrix with a dimension of , where represents the frequency domain, and represents the time domain.
[0097] However, compared with visual images and language, the signal representations of most RF signals are globally information-sparse and locally information-dense, that is, the information carried by the RF signals is concentrated in a few regions, which makes it difficult to perform RF pre-training based on MAE. In addition, different signal representations have different characteristics (such as dimensions, features, etc.). For example, AoA-ToF is a 3D matrix (Time-AoA-ToF), while DFS is a 2D matrix (frequency-time). For simplicity, the signal representation is abstracted into a general 4D tensor form. In particular, the present invention does not process the spectral image as a conventional real-valued RGB image, because this will cause the loss of finer phase information in the signal. For example, the channels of AOA-TOF are 2, including the real part and the imaginary part, instead of taking the modulus and using the amplitude image as the input. However, during the generation process, the amplitude image is used as the reconstruction target because it is found that this technique accelerates the convergence process. The input tensor is divided into a set of small blocks, and each small block contains a representation tensor. The present invention aims to use the partially information-dense small blocks to predict other masked small blocks, so as to obtain a high-level understanding of the RF signal.
[0098] According to an embodiment of the present invention, the above sorting and screening of the initial set of dense regions to obtain the set of target information-dense regions includes: sorting the regions in the initial set of dense regions based on the total signal energy according to the correspondence between the signal high-energy region and the information density, to obtain the sorted initial set of dense regions; screening the sorted initial set of dense regions based on the greedy strategy, and selecting multiple non-overlapping regions with higher rankings in the sorted initial set of dense regions to construct the set of target information-dense regions.
[0099] According to an embodiment of the present invention, the above generation of the sparse attention mask for the 4D signal tensor sample using the set of target information-dense regions to obtain the masked 4D signal tensor sample includes: applying a preset mask ratio to each region in the set of target information-dense regions, and masking the regions in the 4D signal tensor sample that do not belong to the set of target information-dense regions, to obtain the masked 4D signal tensor sample.
[0100] The following is a further detailed description of the process for obtaining the above-mentioned masked 4D signal tensor samples after masking in the specific embodiments in conjunction with the accompanying Figure 4 drawings for the present invention.
[0101] Figure 4 FIG. is a schematic flow chart of generating a sparse attention mask according to an embodiment of the present invention.
[0102] To solve the problem of sparse global information and dense local information in radio frequency signals, the present invention proposes a customized masking strategy FOCUS. Specifically, the present invention dynamically tracks information-dense regions, only retains some patches in these information-dense regions, and discards all other patches. In this way, the pre-training process can observe some information-dense patches and try to predict other patches to obtain a deep understanding of the RF signal representation. In addition, many information-sparse patches do not participate in model training, which greatly reduces the consumption of computing and memory resources.
[0103] The FOCUS masking process mainly includes the following two steps, as Figure 4 shown:
[0104] Selection of information-dense regions: The present invention first randomly generates an initial set of dense regions, denoted as , where the initial set can be generated over the entire region and contains patches, where and . Then, since high-energy regions are usually information-dense, the region proposal set is sorted according to the sum of signal energies. Following the greedy strategy, the top non-overlapping regions are selected as the target information-dense region set ;
[0105] Mask generation: The present invention only applies a masking ratio to each information-dense region and masks other blocks, including patches in information-sparse regions. Generally speaking, this is equivalent to applying a larger masking ratio to the entire region to obtain a masked version of the signal representation.
[0106] According to an embodiment of the present invention, the above-mentioned iterative pre-training of the sparse encoder-decoder module with self-recovery using the masked 4D signal tensor samples and the preset reconstruction loss function to obtain a pre-trained encoder includes: constructing a sparse encoder-decoder module for human feature processing, and improving the reconstruction loss function to obtain a preset reconstruction loss function; using the encoder of the sparse encoder-decoder module to extract semantic features from the masked 4D signal tensor samples to obtain semantic feature samples; using the decoder of the sparse encoder-decoder module to map and reconstruct the semantic feature samples to obtain the 4D signal tensor samples after mask self-recovery; using the preset reconstruction loss function to process the 4D signal tensor samples and the 4D signal tensor samples after mask self-recovery to obtain a reconstruction loss value; using the reconstruction loss value to update the parameters of the sparse encoder-decoder module to obtain the sparse encoder-decoder module with updated parameters; iteratively performing semantic feature extraction operations, mapping reconstruction operations, reconstruction loss value calculation operations, and model parameter update operations until the first preset training condition is satisfied to obtain a pre-trained encoder.
[0107] The following is a further detailed description of the training process of the above-mentioned sparse encoder-decoder module provided by the present invention through specific embodiments in combination with the attached Figure 5 and 6 drawings.
[0108] Figure 5 FIG. is a schematic diagram of the pre-training framework of the sparse encoder-decoder module according to an embodiment of the present invention.
[0109] Figure 6 FIG. is a diagram sharing the encoder block between the pre-training framework and the identity authentication model according to an embodiment of the present invention.
[0110] As Figure 5 shown, the goal of the encoder is to convert the masked signal representation into a high-semantic representation, which can be any classical convolutional network, such as ResNet or ConvNeXt. However, compared with ordinary convolutional networks, the present invention uses a sparse version of the convolutional layer. Sparse convolution not only reduces the consumption of computing resources in the masked patches, but more importantly, avoids information leakage from the observed patches to their adjacent masked patches and representation collapse caused by ordinary convolution calculations in many zero-valued masked regions.
[0111] The role of the decoder is to map the high-semantic representation to the reconstructed masked patches. Therefore, the present invention adopts a lightweight design of a common convolutional decoder, allowing communication between information-dense patches and masked patches, while forcing the encoder to have a high-level understanding, rather than simply transmitting the input to the decoder, which would lead to representation collapse. The lightweight decoder also improves the efficiency of pre-training. In addition, it should be noted that the decoder is only used in the pre-training stage. Once the pre-training is completed, the decoder will be discarded and replaced with a different sub-task network adapted for the purpose of downstream tasks.
[0112] In terms of the reconstruction loss function: The classical masked pre-training framework focuses on recovering all masked regions. Although the framework of the present invention also recovers all masked patches, the difference is that PRISM mainly focuses on recovering the masked patches in information-dense regions.
[0113] According to an embodiment of the present invention, the above preset reconstruction loss function is shown in formula (2):
[0114] (2),
[0115] where, represents the set of target information-dense regions, represents the decoder in the sparse encoder-decoder module, represents the encoder in the sparse encoder-decoder module, represents the mask corresponding to the set of target information-dense regions, represents the L2 norm.
[0116] At the same time, the present invention redesigned the original encoder architecture to adapt to high-dimensional RF data cubes, as Figure 6 shown, which involves optimizing the sampling rate, adjusting the convolutional kernel size, and reforming the spatio-temporal separable convolution, all of which are expected to synergistically improve the model's processing performance for RF signals, save storage space, and speed up the calculation. In particular, a new masking strategy needs to be designed considering the inherent sparsity of RF signals, with special attention paid to meaningful sparse regions, thereby significantly improving the model's ability to extract individual characteristics during pre-training, and further demonstrating practical authentication advantages of higher accuracy and stronger generalization in the identity authentication model.
[0117] According to an embodiment of the present invention, the iterative training of the identity authentication model by using the wireless training data set with human contour tags and pose tags and the identity authentication loss function includes: weighting the identity classification loss function, the human unique pose information estimation error loss function, and the human unique contour information generation loss function based on a preset weight to obtain a preset identity authentication loss function; using the identity authentication model based on the pre-trained encoder to process the 4D signal tensor sample to obtain the identity authentication result of the dynamic experimental target, where the identity authentication result of the dynamic experimental target includes human contour information and pose information; using the preset identity authentication loss function to process the identity authentication result of the dynamic experimental target and the human contour tags and pose tags of the dynamic experimental target in the wireless data sample to obtain an identity authentication loss value; using the identity authentication loss value to update the parameters of the identity authentication model to obtain the identity authentication model with updated parameters; iteratively performing the model data processing operation, the loss value calculation operation, and the model parameter update operation until the second preset training condition is satisfied to obtain the trained identity authentication model.
[0118] The training process of the above identity authentication model will be further described in detail through specific embodiments below.
[0119] The encoder of the identity authentication model is shared with the pre-trained model to directly obtain the prior of human unique information. The present invention mainly performs identity authentication and recognition based on the contour and pose information of the human body. Therefore, the identity authentication network requires the human contour and pose features generated by the model to assist in identity recognition. Therefore, the present invention simultaneously retains the above human contour segmentation and 3D pose estimation intermediate task networks and the corresponding output results, combines the identity classification network constructed by the linear layer, and adds the loss functions of the first two as auxiliary losses to the identity classification loss to ensure the effective feature extraction in the model learning process and improve the classification loss accuracy. The loss function is shown in formula (3):
[0120] (3),
[0121] where, is the identity classification loss, is the human unique pose information estimation error loss, is the human unique contour information generation loss, are the preset weights of the three losses respectively, is the total loss of the final identity authentication model training.
[0122] To better illustrate the advantages and effectiveness of the above method provided by the present invention, the following will further verify the above method provided by the present invention through specific experiments and in combination with the attached Figure 7 Do a further verification of the above method provided by the present invention.
[0123] Figure 7 It is the comparative experimental effect diagram of the wireless identity authentication method according to the embodiment of the present invention.
[0124] In the specific experimental settings, the present invention uses the open-source dataset HIBER based on radar, including various environments, users, occlusions, and actions. It is collected by capturing RF signals from the horizontal and vertical planes using two vertical FMCW radars. This dataset provides RF heatmaps, RGB images, 2D / 3D skeletons, bounding boxes, and contour ground truths.
[0125] In addition, the present invention uses the identity authentication model enhanced by the masked autoencoder pre-training framework proposed by the present invention to better extract relevant unique identity information such as human body contours or postures. Further, the present invention needs to verify the feasibility of performing identity authentication based on Radar sensing information such as human body contours and postures, and the PRISM framework can enhance the accuracy of identity authentication. For this purpose, the present invention also selects views 2, 3, 4, 5, 7, and 9 in the dataset as pre-training samples, totaling 218,650 samples.
[0126] As Figure 7 shown, the present invention first tests the information extraction effect of the pre-training framework on the model to extract unique human behavior characteristics. For contour segmentation, the present invention uses view6 for training (17,520 samples) and testing (4,664 samples). For 3D pose estimation, the present invention uses views Figure 1 , 6 , 8, and 10 for training (32,584 samples) and testing (8,584 samples). The experiments show that the IoU (Intersection of Union) of the models trained from scratch for extracting unique contour features are 0.7444, 0.7485, and 0.7473 respectively, which are 0.0425, 0.0384, and 0.0413 higher than the classical supervised baseline. After passing through the pre-training framework proposed by the present invention, the IoUs of each backbone are 0.7642, 0.7635, and 0.7630 respectively. Compared with their respective supervised training, the IoUs are significantly improved by 0.0198, 0.015, and 0.0157. At the same time, when comparing the effect of extracting unique pose features after pre-training, the Euclidean distance errors of the joint point estimations of each backbone network are improved by 23.79 mm, 25.44 mm, and 19.32 mm respectively in 3D pose estimation. In addition, the overall best result is 72.67 mm.
[0127] As can be seen from the above, after the model is pre-trained by the masked autoencoder proposed in the present invention, the ability to extract human unique behavioral features (contours or postures) is significantly improved, which can further improve the recognition accuracy of identity authentication. For the training set and test set of downstream tasks, in order to ensure fairness, the present invention first maintains the same partitioning method as the pose estimation task, and at the same time uses views 1, 6, 8, and 10 as the test environments for k-fold cross-validation respectively. After model training and testing, the final identity recognition accuracy results show that when using the same backbone network, training the network from scratch can only achieve an identity verification accuracy of about 60% in 4 environments, and the overall average accuracy is 61.2%. However, after extracting more accurate human contour and pose information through the pre-training framework, the identity verification accuracy can reach 75% in 4 environments, and the overall average accuracy is 75.7%, significantly improving the accuracy by 14.5%. Therefore, through experiments, it can be shown that the pre-trained enhanced model can extract more accurate individual identity feature information such as human contours and postures to obtain a higher identity verification accuracy.
[0128] In addition, the present invention further tests the identity verification accuracy results under three different external conditions (normal lighting without occlusion environment, occlusion environment, dark environment), as Figure 7 shown, training the network from scratch can only achieve an identity verification accuracy of 60.4%, 51.2%, and 57.6% respectively under 3 conditions. However, after extracting more accurate human contour and pose information through the pre-trained network, the identity verification accuracy can reach 77.6%, 66.1%, and 75.3% under 3 conditions. Finally, the accuracy is significantly improved by 17.2%, 14.9%, and 17.7% respectively, fully highlighting the significant improvement of the model after pre-training by the PRISM framework, and also demonstrating the potential advantages and development prospects of wireless signals for identity verification in special scenarios.
[0129] Based on the above wireless identity authentication method based on sparse attention masks, the present invention also provides a wireless identity authentication system based on sparse attention masks. The following will be combined with Figure 8 to describe the system in detail.
[0130] Figure 8 is a structural block diagram of a wireless identity authentication system based on sparse attention masks according to an embodiment of the present invention.
[0131] As Figure 8 shown, the above wireless identity authentication system 800 based on sparse attention masks includes a signal characterization module 810, a sparse attention mask module 820, a self-recovery pre-training module 830, an identity authentication model training module 840, and a wireless identity authentication module 850.
[0132] A signal characterization module 810, which is used to characterize the original RF data samples as 4D signal tensor samples by using signal processing techniques; in one embodiment, the signal characterization module 810 can be used to perform the operation S210 described above, which will not be elaborated here.
[0133] A sparse attention mask module 820, which is used to randomly generate an initial set of dense regions on the 4D signal tensor samples, sort and filter the initial set of dense regions to obtain a set of target information-dense regions, and use the set of target information-dense regions to generate a sparse attention mask for the 4D signal tensor samples, so as to obtain a masked 4D signal tensor sample; in one embodiment, the sparse attention mask module 820 can be used to perform the operation S220 described above, which will not be elaborated here.
[0134] A self-recovery pre-training module 830, which is used to perform self-recovery iterative pre-training on the sparse encoder-decoder module by using the masked 4D signal tensor samples and a preset reconstruction loss function, so as to obtain a pre-trained encoder, and use the pre-trained encoder as the encoder of the identity authentication model; in one embodiment, the self-recovery pre-training module 830 can be used to perform the operation S230 described above, which will not be elaborated here.
[0135] An identity authentication model training module 840, which is used to iteratively train the identity authentication model by using a wireless training data set with human contour labels and pose labels and an identity authentication loss function, so as to obtain a trained identity authentication model; in one embodiment, the identity authentication model training module 840 can be used to perform the operation S240 described above, which will not be elaborated here.
[0136] A wireless identity authentication module 850, which is used to process the wireless RF signals in the target area by using the trained identity authentication model to obtain the identity authentication result of the person in the target area; in one embodiment, the wireless identity authentication module 850 can be used to perform the operation S250 described above, which will not be elaborated here.
[0137] According to an embodiment of the present invention, any multiple modules among the signal characterization module 810, the sparse attention mask module 820, the self-recovery pre-training module 830, the identity authentication model training module 840, and the wireless identity authentication module 850 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the signal characterization module 810, the sparse attention mask module 820, the self-recovery pre-training module 830, the identity authentication model training module 840, and the wireless identity authentication module 850 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the signal characterization module 810, the sparse attention mask module 820, the self-recovery pre-training module 830, the identity authentication model training module 840, and the wireless identity authentication module 850 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0138] Figure 9 It is a block diagram of an electronic device suitable for implementing a wireless identity authentication method based on a sparse attention mask according to an embodiment of the present invention.
[0139] As Figure 9 shown, the electronic device 900 according to an embodiment of the present invention includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 901 can also include on-board memory for caching purposes. The processor 901 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0140] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0141] According to an embodiment of the present invention, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the input / output (I / O) interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage portion 908 as needed.
[0142] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0143] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the ROM 902 and / or RAM 903 and / or ROM 902 and RAM 903 described above.
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0145] Those skilled in the art can understand that the features described in various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0146] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and not for limiting the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A wireless identity authentication method based on sparse attention mask, characterized in that: The method comprises: Using signal processing technology to characterize the original RF data samples into 4D signal tensor samples; Randomly generating a preliminary set of dense regions on the 4D signal tensor samples, sorting and screening the preliminary set of dense regions to obtain a target information dense region set, and using the target information dense region set to generate a sparse attention mask for the 4D signal tensor samples to obtain a masked 4D signal tensor sample; Using the masked 4D signal tensor samples and a preset reconstruction loss function to self-recoveringly iterate pre-train a sparse encoder-decoder module to obtain a pre-trained encoder, and using the pre-trained encoder as an encoder of the identity authentication model; Iteratively training the identity authentication model using a wireless training data set with human body contour labels and posture labels and an identity authentication loss function to obtain a trained identity authentication model; The trained identity authentication model is used to process the wireless radio frequency signal of the target area to obtain the identity authentication result of the person in the target area.
2. The method according to claim 1, characterized in that The signal processing techniques used to characterize the raw RF data samples into 4D signal tensor samples include: Performing signal processing on the original radio frequency data samples to obtain the arrival angle-flight time and Doppler frequency shift of the original radio frequency data samples, wherein the original radio frequency data samples are the radio frequency data of the dynamic targets in the experimental area; Taking the amplitude image of the original RF data sample as a reconstruction target, performing vector concatenation based on the time dimension on the arrival angle-flight time and the Doppler frequency shift to obtain the 4D signal tensor sample; or The amplitude image of the original radio frequency data sample is used as a reconstruction target, and the arrival angle-flight time is processed to obtain the 4D signal tensor sample.
3. The method according to claim 1, characterized in that The preliminary set of dense areas is sorted and screened to obtain a target information dense area set including: Based on the correspondence between the signal high energy area and the information density, the areas in the preliminary selection set of dense areas are sorted based on the sum of signal energies to obtain a sorted preliminary selection set of dense areas; The sorted preliminary set of dense regions is screened based on a greedy strategy, and a plurality of non-overlapping regions with top rankings in the preliminary set of dense regions are selected to construct the target information dense region set.
4. The method according to claim 3, characterized in that Generating a sparse attention mask for the 4D signal tensor sample using the target information dense region set to obtain the masked 4D signal tensor sample includes: A preset mask ratio is applied to each region in the target information dense region set, and regions in the 4D signal tensor samples that do not belong to the target information dense region set are masked to obtain the masked 4D signal tensor samples.
5. The method according to claim 1, characterized in that The sparse encoder-decoder module is self-recoveringly pre-trained iteratively using the masked 4D signal tensor samples and the preset reconstruction loss function, and the pre-trained encoder includes: Constructing a sparse encoder-decoder module for character feature processing, and improving the reconstruction loss function to obtain the preset reconstruction loss function; Using the encoder of the sparse encoder-decoder module to extract semantic features from the masked 4D signal tensor samples to obtain semantic feature samples; Using the decoder of the sparse encoder-decoder module to map and reconstruct the semantic feature samples to obtain 4D signal tensor samples after mask self-recovery; Processing the 4D signal tensor samples and the masked self-recovered 4D signal tensor samples using the preset reconstruction loss function to obtain a reconstruction loss value; Using the reconstruction loss value, the parameters of the sparse encoder-decoder module are updated to obtain a sparse encoder-decoder module after parameter update; The semantic feature extraction operation, the mapping reconstruction operation, the reconstruction loss value calculation operation, and the model parameter update operation are iterated until the first preset training condition is met to obtain the pre-trained encoder.
6. The method according to claim 5, characterized in that The preset reconstruction loss function is shown in the following formula: , in, represents the target information dense area set, represents a decoder in the sparse encoder-decoder module, represents the encoder in the sparse encoder-decoder module, represents the mask corresponding to the target information dense area set, represents the L2 norm.
7. The method according to claim 1, characterized in that The identity authentication model is iteratively trained using a wireless training data set with human body contour labels and posture labels and an identity authentication loss function, and the trained identity authentication model includes: Based on the preset weights, the identity classification loss function, the human body unique posture information estimation error loss function and the human body unique contour information generation loss function are weighted to obtain a preset identity authentication loss function; Processing the 4D signal tensor samples using an identity authentication model based on the pre-trained encoder to obtain an identity authentication result of a dynamic experimental target, wherein the identity authentication result of the dynamic experimental target includes human body contour information and posture information; The preset identity authentication loss function is used to process the identity authentication result of the dynamic experimental target and the human body contour label and posture label of the dynamic experimental target in the wireless data sample to obtain an identity authentication loss value; Using the identity authentication loss value to update the parameters of the identity authentication model to obtain an identity authentication model with updated parameters; The model data processing operation, the loss value calculation operation and the model parameter update operation are iterated until the second preset training condition is met to obtain a trained identity authentication model.
8. A wireless identity authentication system based on sparse attention mask, characterized in that: The system comprises: A signal characterization module, used for characterizing the original RF data samples into 4D signal tensor samples using signal processing technology; A sparse attention mask module is used to randomly generate a preliminary set of dense regions on the 4D signal tensor samples, sort and screen the preliminary set of dense regions to obtain a target information dense region set, and use the target information dense region set to generate a sparse attention mask for the 4D signal tensor samples to obtain the masked 4D signal tensor samples; A self-recovery pre-training module, used to perform self-recovery iterative pre-training on a sparse encoder-decoder module using the masked 4D signal tensor samples and a preset reconstruction loss function to obtain a pre-trained encoder, and use the pre-trained encoder as an encoder of an identity authentication model; An identity authentication model training module, used to iteratively train the identity authentication model using a wireless training data set with human body contour labels and posture labels and an identity authentication loss function to obtain a trained identity authentication model; The wireless identity authentication module is used to process the wireless radio frequency signal of the target area by using the trained identity authentication model to obtain the identity authentication result of the person in the target area.
9. An electronic device, comprising: one or more processors; A memory for storing one or more instructions, Wherein, when the one or more instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Natural language understanding method and device fusing dialogue context information
CN116542256A
Semi-supervised radio frequency fingerprint identification method and system based on mask contrast training
CN118196845A