Human body behavior recognition method and device, equipment and medium
The CSI data is converted into two-dimensional images through Gram angular difference field and fractional Fourier transform. Combined with the repeated cross attention module and the fully connected layer, the problem of insufficient accuracy in human behavior recognition is solved, and efficient behavior recognition effect is achieved.
Patent Information
- Application Number
- CN202510740202.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Existing lightweight models are insufficient in human behavior recognition, and traditional methods require strong computing resources and long response time.
The Gram angular difference field is used to convert CSI data into two-dimensional images, and features are extracted by fractional Fourier transform and repeated cross attention modules. Features are fused through multi-layer perceptron and fully connected layers to achieve accurate recognition.
Accurate recognition of human behavior under light models is achieved, improving recognition accuracy and reducing computing resource requirements.
Smart Images

Figure CN120260139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of behavior recognition, and in particular, to a human behavior recognition method, device, equipment and medium. Background Art
[0002] With the booming rise of artificial intelligence technology and the wide popularization of the Internet of Things, the importance of human behavior recognition in intelligent environments is increasing day by day. In the field of smart home, by recognizing the behaviors of residents, the automatic control of household appliances can be realized, improving the comfort and convenience of living; in health monitoring, the motion state and physiological activities of the human body can be captured in real time, providing data support for disease prevention and treatment; in security scenarios, abnormal behaviors can be detected in time and alarms can be issued to ensure public safety; in the field of sports tracking, the movements of athletes can be accurately analyzed to help improve training and competitive levels.
[0003] Traditional human behavior recognition methods mainly rely on traditional sensors for data collection and analysis. Although behavior recognition can be achieved to a certain extent, there are many limitations. Human activity recognition based on WiFi Channel State Information (CSI) can ensure personal privacy and security, eliminating the requirement of equipping any device for the target, and has become a research hotspot. With the maturity of deep learning technology, existing research has used deep learning methods to achieve human activity recognition based on CSI. For example, deep learning networks such as Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) have been used in CSI-based activity recognition. Although the recognition accuracy has been improved to a certain extent, the performance of these methods depends to a large extent on complex feature processing and model architecture. Strong computing resources are required and the response time is long. While using lightweight models will reduce the accuracy of human behavior recognition. Summary of the Invention
[0004] Embodiments of the present invention provide a human behavior recognition method, device, equipment and medium to solve the technical problem that the accuracy of human behavior recognition using CSI with a lightweight model is reduced in the prior art.
[0005] In a first aspect, embodiments of the present invention provide a human behavior recognition method, including: Converting the collected human behavior into a two-dimensional image by using the Gram angular difference field, so that the two-dimensional image retains the time information of the CSI data; Dividing the two-dimensional image to form a plurality of token sequences. When encoding the token sequences, perform a fractional Fourier transform of the transformed angle on the normalized processing result of the output of each encoding layer, and add the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; Normalize the output of the encoding layer and input the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result; Apply a one-dimensional discrete fractional Fourier transform along the hidden dimension to the processing result for the first transformation to obtain a first transformation result, and apply a one-dimensional discrete fractional Fourier transform to the first transformation result along the sequence dimension to obtain a second transformation result, and obtain the real part of the second transformation result as the first output feature; Input the two-dimensional image into the repeated cross-attention module to obtain a second output feature with dense and rich context information; Fuse the first output feature and the second output feature, and input the fusion result into a fully connected layer to output the human behavior recognition result.
[0006] In a second aspect, an embodiment of the present invention further provides a human behavior recognition device, including: A two-dimensional image conversion module for converting the collected human behavior into a two-dimensional image by using the Gram angular difference field, so that the two-dimensional image retains the time information of the CSI data; A partitioning module for partitioning the two-dimensional image to form a plurality of token sequences. When encoding the token sequences, perform a fractional Fourier transform of the transformation angle on the normalization processing result of the output of each encoding layer, and add the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; A normalization module for normalizing the output of the encoding layer and inputting the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result; A transformation module for applying a one-dimensional discrete fractional Fourier transform along the hidden dimension to the processing result for the first transformation to obtain a first transformation result, and applying a one-dimensional discrete fractional Fourier transform to the first transformation result along the sequence dimension to obtain a second transformation result, and obtaining the real part of the second transformation result as the first output feature; An input module for inputting the two-dimensional image into the repeated cross-attention module to obtain a second output feature with dense and rich context information; An output module for fusing the first output feature and the second output feature, and inputting the fusion result into a fully connected layer to output the human behavior recognition result.
[0007] In a third aspect, an embodiment of the present invention further provides a device, including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the human behavior recognition method provided in the above embodiment.
[0008] In a fourth aspect, an embodiment of the present invention further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the human behavior recognition method provided in the above embodiment.
[0009] The human behavior recognition method, device, equipment and medium provided by the embodiment of the present invention convert the collected human behavior into a two-dimensional image by using the Gram angle difference field, so that the two-dimensional image retains the time information of the CSI data; the two-dimensional image is divided to form a plurality of token sequences, and when encoding the token sequence, the output normalization processing result of each coding layer is subjected to the fractional Fourier transform of the transformation angle, and the transformation result is added to the output result of the previous layer as the input of the next layer or the final output of the coding layer; the output of the coding layer is normalized, and the normalization processing result is input into a multi-layer perceptron for processing to obtain the processing result; the processing result is subjected to a first transformation by applying a one-dimensional discrete fractional Fourier transform along the hidden dimension to obtain a first transformation result, and the first transformation result is subjected to a one-dimensional discrete fractional Fourier transform along the sequence dimension to obtain a second transformation result, and the real part of the second transformation result is obtained as the first output feature; the two-dimensional image is input into a repeated cross attention module to obtain a second output feature with dense and rich context information; the first output feature and the second output feature are fused, and the fusion result is input into a fully connected layer to output the human behavior recognition result. The collected data is converted into a two-dimensional feature map using GADF technology, making the data features more intuitive and easier to analyze. Feature extraction is performed on the two-dimensional feature map from multiple angles to mine rich behavioral feature information. The extracted features are integrated using the fusion module to make the feature set more comprehensive and effective, and to achieve accurate recognition of human behavior using a lightweight model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 1 is a flow chart of a human behavior recognition method provided by Embodiment 1 of the present invention; Figure 2 is a schematic diagram of a test scenario in the human behavior recognition method provided in the first embodiment; Figure 3 is a flow chart of a human behavior recognition method provided by Embodiment 2 of the present invention; Figure 4 is a schematic diagram of the structure of a human behavior recognition device provided by Embodiment 3 of the present invention; Figure 5 It is a structural diagram of the device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0011] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0012] Embodiment 1 Figure 1 : is a flow chart of a human behavior recognition method provided in Embodiment 1 of the present invention. This embodiment is applicable to the case where human behavior is recognized using CSI. The method can be executed by a human behavior recognition device and specifically includes the following steps: Step 110: using the Gram angle difference field to convert the collected human behavior into a two-dimensional image, so that the two-dimensional image retains the time information of the CSI data.
[0013] Exemplarily, the method of using the Gram angle difference field to convert the collected human behavior into a two-dimensional image includes: obtaining a CSI data matrix generated when a person performs an action, wherein the elements in the matrix include an in-phase component, an orthogonal component, a phase value, and an amplitude value corresponding to each signal; normalizing all amplitude values, and converting the normalized signal into polar coordinates, and calculating the angle of the polar coordinates; and using GADF to transform the signal converted into polar coordinates into an image.
[0014] Specifically, a human behavior monitoring system based on Wi-Fi signals can be built as the basic platform. The tester performs actions at the specified location and collects data to ensure that the data source and collection method are standardized and reliable. The collected data is converted into a two-dimensional feature map through GADF, and the feature map is input into the VIT-FrFT module and the RCCA module respectively to extract the corresponding features. The extracted features are then input into the fusion module for feature fusion. The fused features are input into the output module for training data, so that the system can learn the mapping relationship between behavioral features and categories.
[0015] Figure 2 Schematic diagram of the test scene in the human behavior recognition method provided in the first embodiment. Figure 2 , an indoor environment can be selected for CSI data collection. Multiple sampling points with fixed positions and even arrangement are set in the environment to ensure that the sampling points can fully and reasonably cover the collection area. The tester performs actions such as waving, clapping, squatting, front kicking, side kicking, sitting down, and standing up at each point. A mobile device equipped with a network card that supports CSI data collection is placed in front of the sampling range, and a router is placed behind the sampling range. The CSI data matrix generated when the tester performs the action is collected through the data collection tool that is matched with the network card that supports CSI data collection. , is an information matrix containing data packets, that is , where , , represents the imaginary part, and respectively represent the in-phase component and the quadrature component, and respectively represent the phase value and the amplitude value.
[0016] The collected CSI data is transformed into a two-dimensional image by using the Gramian Angular Difference Field (GADF) technique. The extracted amplitude sequence is normalized so that all input values are within range, obtaining . The normalized signal is converted to polar coordinates , where is the signal length, represents the sample corresponding angle value. Then, the GADF is used to transform the signal into an image.
[0017] The transformed image can be represented as a matrix: , is a matrix composed of sine functions of the difference between any two angles with a size of Among them, reflects the sine value of the difference between different angle values, and the transformed two-dimensional image is represented by the matrix formed by these elements. , , and are the normalized signal values, , respectively represent the sample , corresponding angle values, The value ranges of
[0018] Step 120: Divide the two-dimensional image to form multiple token sequences. When encoding the token sequences, the normalized processing results of the outputs of each encoding layer are subjected to a fractional Fourier transform of the transformation angle, and the transformation results are added to the output results of the previous layer as the input of the next layer or the final output of the encoding layer.
[0019] Optionally, the dividing of the two-dimensional image to form a plurality of token sequences may include: dividing the two-dimensional feature map into a plurality of blocks; mapping each block through a linear layer to obtain a plurality of token sequences of a fixed size.
[0020] Exemplarily, the two-dimensional image obtained by the GADF transform may be used as the input feature map of the VIT-FrFT module , and its dimension is represented as , is the number of channels of the image, is the image size.
[0021] In this embodiment, in order to improve the accuracy of human behavior recognition, it is necessary to extract invisible features as much as possible. Therefore, in this embodiment, an improved Transformer encoder structure may be adopted. Transformer usually processes sequence data. To make the image data meet the input requirements of Transformer, the two-dimensional feature map is divided into token sequences. After the Patch Embedding operation on the feature map, a fixed-size token sequence is obtained, and the dimension of each token is , and the length of the entire sequence is , and finally reshaped into , that is, the input image is divided into multiple tokens .
[0022] Correspondingly, the method may further include the following steps: performing a D-dimensional linear mapping on each token, mapping it to the D dimension, and expressing the mapped result using a linear layer. The expression result includes: a mapping matrix and a position encoding to enhance the perception ability of the position information of the elements in the image.
[0023] Exemplarily, it can be represented using a linear layer as , where is the learnable sequence of the embedding, is the mapping matrix, is the position encoding to enhance the model's perception ability of the position information of the elements in the image, and then the embedded sequence is sent to the first half of the Transformer encoder, and the calculation method is . Specifically, for each layer ( , is the number of encoders), first perform layer normalization (Layer Normalization, LN) on the output of the previous layer, and then perform a transformation with an angle of Fractional Fourier Transform (FrFT) Add the result of the transformation to the output of the previous layer to obtain the output of the current layer Through the above series of operations, the encoding process of the input is gradually completed, and a feature representation containing context information and position information is extracted. Among them, The specific calculation method is as follows: Given a sequence The Discrete Fractional Fourier transform (DFrFT) can be defined as: When D is an integer, then When D is an integer, then ; Step 130: Normalize the output of the encoding layer and input the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result
[0024] Normalize the features extracted from the first half of the Transformer encoder and input the result into a Multi-Layer Perceptron (MLP) block
[0025] Correspondingly, the step of inputting the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result includes: The multi-layer perceptron includes: two linear layers and a non-linear activation layer Gaussian error linear unit. Perform a two-dimensional discrete fractional Fourier transform operation on the features obtained after being processed by the multi-layer perceptron, and perform the transformation separately from the sequence dimension and the hidden dimension to obtain a processing result
[0026] Normalize the extracted features further and input the result into a Multi-Layer Perceptron (MLP) block, which can be expressed as The MLP block consists of two linear layers and a non-linear activation layer Gaussian error linear unit
[0027] Step 140: Apply a one-dimensional discrete fractional Fourier transform to the processing result along the hidden dimension for the first transformation to obtain a first transformation result, and perform a one-dimensional discrete fractional Fourier transform on the first transformation result along the sequence dimension to obtain a second transformation result. Take the real part of the second transformation result as the first output feature
[0028] Perform a two-dimensional discrete fractional Fourier transform (2D-DFrFT) operation on the features processed by the MLP block. Specifically, After passing through the LN and MLP blocks, we get , and then the 2D-DFrFT is performed separately from two dimensions (sequence dimension and hidden dimension) to capture the dynamic change information of human behavior in the temporal and spatial order, explore the potential correlations between different feature channels, so as to extract more comprehensive and higher-level features, and provide multi-domain context information for improving the performance of the human behavior recognition model. Specifically, the system first performs a one-dimensional discrete fractional Fourier transform (1D-DFrFT) on the input features, that is, the features processed by the LN and MLP blocks, , along the hidden dimension, that is , apply 1D-DFrFT to the result of the previous step, along the sequence dimension for transformation. Finally, by taking the real part, the output feature Z is obtained. That is, 2D-DFrFT can be expressed as . Since the FrFT has an additive property, that is , each encoder block transforms the input features at in the spatial frequency plane. Therefore, the addition of two FrFTs with different orders results in a larger transformation angle. In each block, the output of the FrFT is multiplied in the MLP sublayer in the Fourier domain. The multiplication in the Fourier domain is a fractional convolution in the spatial domain, that is . The 2D-DFrFT block represents a structure that performs fractional convolutions in the spatial domain with different transformation angles. The connections between the FrFT blocks are used to provide multi-domain context information. Using the above method, the first output feature can be obtained.
[0029] Step 150: Input the two-dimensional image into the repeated cross-attention module to obtain the second output feature with dense and rich context information.
[0030] Exemplarily, the feature map can be subjected to a cross-attention (CCA) operation. Specifically, is applied to two convolutional layers to generate two feature maps and , where , is the number of channels, and the dimension ratio is smaller. For each position in the feature map , extract the feature vector at this position. Extract all the feature vectors in the same row or the same column as the position from the feature map to form a set , calculate Q u and K u to obtain the dot product of each eigenvector in, getting the correlation score , where , is the correlation degree between Q u and K u . is the correlation matrix of all positions, applying the softmax operation on to obtain the attention map , that is . Then, apply another convolutional layer to generate the feature map . For each position , extract all eigenvectors in that are in the same row or the same column as the position to form the set . , is the set of eigenvectors in the same row or the same column as the position . Using the attention map to perform weighted summation on , and adding the result to the original feature X u to obtain the enhanced feature map , where is the eigenvector at the position u in the output feature map. Context information is added to the local feature to enhance the representation of relevant information. Therefore has a wide context receptive field and selectively aggregates context information according to the spatial attention map.
[0031] Furthermore, to solve the connection problem between the Figure 1 pixels in the CCA module and the pixels around them that are not on the cross path, the Recurrent Cross Attention RCCA module is introduced. The RCCA module is equipped with two loops and can collect full-image context information from all pixels to generate new features with dense and rich context information.
[0032] Step 160, fuse the first output feature and the second output feature, and input the fusion result into the fully connected layer to output the human behavior recognition result.
[0033] Exemplarily, through existing fusion methods, the first output feature and the second output feature can be fused using a fusion layer, and the fused features can be input into the fully connected layer to implement subsequent classification and regression tasks. Output the human behavior recognition result.
[0034] This embodiment converts the collected human behavior into a two-dimensional image by using the Gram angle difference field, so that the two-dimensional image retains the time information of the CSI data; divides the two-dimensional image to form multiple token sequences, and when encoding the token sequence, performs a fractional Fourier transform of the transformation angle on the output normalization processing result of each encoding layer, and adds the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; normalizes the output of the encoding layer, and inputs the normalization processing result into a multi-layer perceptron for processing to obtain a processing result; applies a one-dimensional discrete fractional Fourier transform to the processing result along the hidden dimension for a first transformation to obtain a first transformation result, performs a one-dimensional discrete fractional Fourier transform on the first transformation result along the sequence dimension to obtain a second transformation result, and obtains the real part of the second transformation result as the first output feature; inputs the two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information; fuses the first output feature and the second output feature, and inputs the fusion result into a fully connected layer to output a human behavior recognition result. The collected data is converted into a two-dimensional feature map using GADF technology, making the data features more intuitive and easier to analyze. Feature extraction is performed on the two-dimensional feature map from multiple angles to mine rich behavioral feature information. The extracted features are integrated using the fusion module to make the feature set more comprehensive and effective, and to achieve accurate recognition of human behavior using a lightweight model.
[0035] Embodiment 2 Figure 3 It is a flow chart of the human behavior recognition method provided in the second embodiment of the present invention. This embodiment is optimized based on the above embodiment, and the first output feature and the second output feature are merged. The specific optimization is: the first output feature and the second output feature are connected into a connection matrix, and the connection matrix is linearly transformed by using a fully connected layer, and the attention weight is calculated.
[0036] See also Figure 3 , the human behavior recognition method comprises: Step 210: Convert the collected human behavior into a two-dimensional image using the Gram angle difference field, so that the two-dimensional image retains the time information of the CSI data.
[0037] Step 220, divide the two-dimensional image to form multiple token sequences. When encoding the token sequence, perform a fractional Fourier transform of the transformation angle on the output normalization processing result of each coding layer, and add the transformation result to the output result of the previous layer as the input of the next layer or the final output of the coding layer.
[0038] Step 230: Normalize the output of the encoding layer and input the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result.
[0039] Step 240: Apply a one-dimensional discrete fractional Fourier transform along the hidden dimension to the processing result for the first transformation to obtain a first transformation result. Then, apply a one-dimensional discrete fractional Fourier transform to the first transformation result along the sequence dimension to obtain a second transformation result, and take the real part of the second transformation result as the first output feature.
[0040] Step 250: Input a two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information.
[0041] Step 260: Concatenate the first output feature and the second output feature into a concatenated matrix, perform a linear transformation on the concatenated matrix using a fully connected layer, calculate the attention weights, and input the fusion result into a fully connected layer to output the human behavior recognition result.
[0042] Concatenate the features F1 and F2 extracted by the two branches into a matrix , where , is the dimension of the feature. Perform a linear transformation on through a fully connected layer to obtain a matrix , where W F is a learnable weight matrix, , and then calculate the attention weights , where is the transpose, W ω is a learnable vector generated by the model. During the training process, the model adjusts the value of W ω through backpropagation so that it can better focus on the most important parts of the input features for the task. Apply the weight to , to obtain the fused feature . In this way, the network can adaptively adjust the contribution degrees of the features extracted by RCCA and VIT-FrFT according to the self-attention mechanism, thus achieving more effective feature fusion.
[0043] The fused feature is input into a fully connected layer to obtain a score vector x for representing each category i , , represents the set of all behavior categories. The output x of the fully connected layer iIt will further calculate the prediction probability of each category through the softmax function , that is , where e is the natural constant; the model is trained by minimizing the loss function , where y i represents the true label of the sample; enabling the trained model to make accurate predictions when facing new data.
[0044] In this embodiment, by fusing the first output feature and the second output feature, it is specifically optimized as follows: connecting the first output feature and the second output feature into a connection matrix, performing a linear transformation on the connection matrix using a fully connected layer, and calculating the attention weights. By using the above method, a flexible weight adjustment mechanism can be utilized, and the two output features can be adaptively adjusted according to the self-attention mechanism. It can effectively fuse the features and further enhance the accuracy of human behavior recognition.
[0045] In a preferred embodiment of this embodiment, the method may further include the following steps: training using the minimum loss function to update the transformation angle and the attention weights. Exemplarily, the model is trained by minimizing the loss function so that the trained model can make accurate predictions when facing new data. During the training process, the transformation angle α and W ω value can be adjusted through backpropagation, enabling the entire model to accurately recognize human behaviors.
[0046] Embodiment III Figure 4 is a schematic structural diagram of the human behavior recognition device provided in Embodiment III of the present invention. Refer to Figure 4 , the human behavior recognition device includes:[[]] A two-dimensional image conversion module 310, configured to convert the collected human behavior into a two-dimensional image using the Gram angle difference field, so that the two-dimensional image retains the time information of the CSI data; A partitioning module 320, configured to partition the two-dimensional image to form a plurality of token sequences. When encoding the token sequences, perform a fractional Fourier transform of the transformation angle on the normalized processing result of the output of each encoding layer, and add the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; A normalization module 330, configured to normalize the output of the encoding layer and input the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result; A transformation module 340 is used to apply a one-dimensional discrete fractional Fourier transform to the processing result along the hidden dimension to perform a first transformation to obtain a first transformation result, and to perform a one-dimensional discrete fractional Fourier transform on the first transformation result along the sequence dimension to obtain a second transformation result, and to obtain a real part of the second transformation result as a first output feature; An input module 350, configured to input the two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information; The output module 360 is used to fuse the first output feature and the second output feature, input the fusion result into the fully connected layer, and output the human behavior recognition result.
[0047] The human behavior recognition device provided by this embodiment converts the collected human behavior into a two-dimensional image by using the Gram angle difference field, so that the two-dimensional image retains the time information of the CSI data; divides the two-dimensional image to form multiple token sequences, and when encoding the token sequence, performs a fractional Fourier transform of the transformation angle on the output normalization processing result of each encoding layer, and adds the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; normalizes the output of the encoding layer, and inputs the normalization processing result into a multi-layer perceptron for processing to obtain a processing result; applies a one-dimensional discrete fractional Fourier transform to the processing result along the hidden dimension for a first transformation to obtain a first transformation result, performs a one-dimensional discrete fractional Fourier transform on the first transformation result along the sequence dimension to obtain a second transformation result, and obtains the real part of the second transformation result as the first output feature; inputs the two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information; fuses the first output feature and the second output feature, and inputs the fusion result into a fully connected layer to output a human behavior recognition result. The collected data is converted into a two-dimensional feature map using GADF technology, making the data features more intuitive and easier to analyze. Feature extraction is performed on the two-dimensional feature map from multiple angles to mine rich behavioral feature information. The extracted features are integrated using the fusion module to make the feature set more comprehensive and effective, and to achieve accurate recognition of human behavior using a lightweight model.
[0048] Based on the above embodiments, the two-dimensional image conversion module includes: An acquisition unit is used to acquire a CSI data matrix generated when a person performs an action. The elements in the matrix include an in-phase component, an orthogonal component, a phase value, and an amplitude value corresponding to each signal. A conversion unit, used for normalizing all amplitude values, converting the normalized signals into polar coordinates, and calculating the angles of the polar coordinates; A transformation module for transforming a signal converted to polar coordinates into an image using GADF.
[0049] Based on the above embodiments, the partitioning module includes: A segmentation unit for segmenting a two-dimensional feature map into multiple blocks; A mapping unit for mapping each block through a linear layer to obtain multiple token sequences of a fixed size.
[0050] Based on the above embodiments, the apparatus further includes: An expression module for performing a D-dimensional linear mapping on each token, mapping it to the D dimension, and expressing the mapped result using a linear layer. The expression result includes a mapping matrix and a positional encoding to enhance the ability to perceive the positional information of elements in the image.
[0051] Based on the above embodiments, the normalization module includes: A separate transformation unit for performing a two-dimensional discrete fractional Fourier transform operation on the features obtained after being processed by a multi-layer perceptron, performing transformations separately from the sequence dimension and the hidden dimension to obtain a processing result. The multi-layer perceptron includes two linear layers and a non-linear activation layer, Gaussian error linear unit.
[0052] Based on the above embodiments, the fusion module includes: An attention weight calculation unit for concatenating the first output feature and the second output feature into a concatenation matrix, performing a linear transformation on the concatenation matrix using a fully connected layer, and calculating the attention weight.
[0053] Based on the above embodiments, the apparatus further includes: An update module for training using a minimum loss function to update the transformation angle and the attention weight.
[0054] The human behavior recognition device provided by the embodiments of the present invention can execute the human behavior recognition method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0055] Embodiment IV Figure 5 It is a schematic structural diagram of a device provided by Embodiment IV of the present invention. Figure 5 It shows a block diagram of an exemplary device 12 suitable for implementing the embodiments of the present invention. Figure 5 The shown device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0056] As Figure 5As shown, device 12 is embodied in the form of a general-purpose computing device. The components of device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that couples different system components including the system memory 28 and the processing unit 16.
[0057] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0058] Device 12 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by device 12, including volatile and nonvolatile media, removable and non-removable media.
[0059] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, a storage system 34 can be used for reading from and writing to a non-removable, nonvolatile magnetic medium ( Figure 5 not shown, typically referred to as a "hard disk drive"). Although Figure 5 not shown in the figures, a disk drive for reading from and writing to a removable nonvolatile disk (such as a "floppy disk"), and an optical disk drive for reading from and writing to a removable nonvolatile optical disk (such as a CD-ROM, DVD-ROM, or other optical medium) can be provided. In these instances, each drive can be connected to bus 18 by one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the various embodiments of the present invention.
[0060] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. The program modules 42 generally carry out the functions and / or methods of the embodiments described herein.
[0061] Device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the device 12, and / or communicate with any device that enables the device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of device 12 through a bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0062] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the human behavior recognition method provided in the embodiments of the present invention.
[0063] Embodiment 5 Embodiment 5 of the present invention also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute any one of the human behavior recognition methods provided in the above embodiments when executed by a computer processor.
[0064] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0065] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0066] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0067] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or device. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet).
[0068] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A human behavior recognition method, characterized in that, Comprising: Converting the collected human behaviors into two-dimensional images by using the Gram angular difference field, so that the two-dimensional images retain the time information of the CSI data; Dividing the two-dimensional image to form a plurality of token sequences. When encoding the token sequences, performing a fractional Fourier transform of the transformed angle on the normalized processing result of the output of each encoding layer, and adding the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; Normalizing the output of the encoding layer and inputting the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result; Performing a first transform on the processing result by applying a one-dimensional discrete fractional Fourier transform along the hidden dimension to obtain a first transform result, and performing a one-dimensional discrete fractional Fourier transform on the first transform result along the sequence dimension to obtain a second transform result, and obtaining the real part of the second transform result as a first output feature; Inputting the two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information; Fusing the first output feature and the second output feature, and inputting the fusion result into a fully connected layer to output the human behavior recognition result.
2. The method according to claim 1, wherein The converting the collected human behaviors into two-dimensional images by using the Gram angular difference field comprises: Obtaining a CSI data matrix generated when a person performs an action, and the elements in the matrix include the in-phase component, quadrature component, phase value, and amplitude value corresponding to each signal; Performing a normalization process on all amplitude values, converting the normalized signals into polar coordinates, and calculating the angles of the polar coordinates; Using the GADF to transform the signals converted into polar coordinates into images.
3. The method according to claim 1, characterized in that, The dividing the two-dimensional image to form a plurality of token sequences comprises: Dividing the two-dimensional feature map into a plurality of blocks; Mapping each block through a linear layer to obtain a plurality of token sequences of a fixed size.
4. The method according to claim 3, characterized in that, The method further comprises: Performing a D-dimensional linear mapping on each token, mapping it to the D dimension, and expressing the mapped result using a linear layer. The expression result includes: a mapping matrix and a position encoding to enhance the perception ability of the position information of the elements in the image.
5. The method according to claim 1, characterized in that, The inputting the result of the normalization process into a multi-layer perceptron for processing to obtain a processing result comprises: The multi-layer perceptron comprises: two linear layers and a non-linear activation layer Gaussian error linear unit. Performing a two-dimensional discrete fractional Fourier transform operation on the features obtained after being processed by the multi-layer perceptron, and performing transforms respectively from the sequence dimension and the hidden dimension to obtain a processing result.
6. The method according to claim 1, characterized in that The fusing the first output feature and the second output feature comprises: Connecting the first output feature and the second output feature into a connection matrix, and performing a linear transformation on the connection matrix by using a fully connected layer and calculating the attention weights.
7. The method according to claim 6, characterized in that, The method further comprises: Training by using a minimum loss function to update the transformed angle and the attention weights.
8. A human behavior recognition device, characterized in that, Comprising: A two-dimensional image conversion module for converting the collected human behaviors into two-dimensional images by using the Gram angular difference field, so that the two-dimensional images retain the time information of the CSI data; A partitioning module, configured to partition the two-dimensional image to form multiple token sequences. When encoding the token sequences, perform a fractional Fourier transform of the transformed angle on the normalized processing result of the output of each encoding layer, and add the transformation result to the output result of the previous layer as the input of the next layer or the final output of the encoding layer; A normalization module, configured to normalize the output of the encoding layer and input the normalized processing result into a multi-layer perceptron for processing to obtain a processing result; A transformation module, configured to perform a first transformation on the processing result by applying a one-dimensional discrete fractional Fourier transform along the hidden dimension to obtain a first transformation result, and perform a one-dimensional discrete fractional Fourier transform on the first transformation result along the sequence dimension to obtain a second transformation result, and take the real part of the second transformation result as the first output feature; An input module, configured to input the two-dimensional image into a repeated cross-attention module to obtain a second output feature with dense and rich context information; An output module, configured to fuse the first output feature and the second output feature, and input the fusion result into a fully-connected layer to output a human behavior recognition result.
9. A computer device, characterized in that, The computer device includes: One or more processors; A storage device, configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the human behavior recognition method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to execute the human behavior recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Non-contact multi-target behavior recognition method
CN114358103A
Human body action recognition system based on attention mechanism feature fusion and irrelevant to position
CN114639169A
Novel human body behavior recognition method based on channel feature enhancement and deep learning
CN114694260A
Human body action recognition method, system and device based on WiFi channel state information imaging and readable storage medium
CN115830705A
Millimeter wave radar human body tumble behavior identification method and system based on multi-class three-dimensional features and Transform
CN115982620A