Motion recognition method for millimeter wave and image signals

By combining the feature fusion method of millimeter-wave radar and RGB image data, the problem of low accuracy in human motion recognition in complex environments is solved, and stable perception and accurate recognition in low-light or occluded scenes are achieved. It is suitable for security monitoring, medical care and intelligent interaction.

CN120599702APending Publication Date: 2025-09-05珠海城市职业技术学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510768587.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in human motion recognition in complex environments, especially in poor lighting or occlusion conditions, where it is difficult to stably perceive human motion.

Method used

Combining millimeter-wave radar and RGB image data, features are extracted through one-dimensional and two-dimensional feature encoders, and dual-path feature fusion and predictor are used to generate human action probability distribution to achieve action recognition.

Benefits of technology

It improves the accuracy and stability of human motion recognition in low-light or obscured scenes while ensuring privacy, making it suitable for security monitoring, medical care, and intelligent interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599702A_ABST
    Figure CN120599702A_ABST
Patent Text Reader

Abstract

The invention discloses an action recognition method for millimeter wave and image signals, and relates to the technical field of action recognition, and the method comprises the steps: obtaining corresponding millimeter wave point clouds and RGB images of a human body at the same time; extracting features of the millimeter wave point cloud through a one-dimensional feature encoder, and further processing the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; extracting features of the RGB image through a two-dimensional feature encoder, and further processing the extracted features through a second global feature pooling layer to obtain an image feature vector; splicing the point cloud feature vector and the image feature vector by using two-way feature fusion and a predictor to obtain an integrated feature vector; generating human body action probability distribution according to the integrated feature vector by using a double-path feature fusion and predictor; and recognizing human body actions according to the human body action probability distribution. According to the invention, the human body motion can be accurately recognized by using the characteristic that the millimeter waves can stably perceive the human body motion in a low-light or shielding scene and combining image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of motion recognition technology, and in particular to a motion recognition method for millimeter wave and image signals. Background Art

[0002] Human action recognition technology based on image and millimeter wave fusion aims to combine the rich semantic information of visual images with the spatial dynamic characteristics of millimeter wave radar to enhance the robust recognition of human behavior in complex environments. Image data can accurately capture posture and action details under good lighting conditions, but it is easily affected by factors such as occlusion and lighting changes. Summary of the Invention

[0003] In view of this, an embodiment of the present application provides a motion recognition method for millimeter waves and image signals to improve the accuracy of motion recognition.

[0004] An aspect of an embodiment of the present application provides a method for motion recognition based on millimeter wave and image signals, the method comprising the following steps: Obtain millimeter wave point cloud and RGB image corresponding to the human body at the same time; Extract features from the millimeter-wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; Extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; Using a dual-path feature fusion and predictor to concatenate the point cloud feature vector and the image feature vector to obtain an integrated feature vector; generating a human motion probability distribution based on the integrated feature vector using the dual-path feature fusion and predictor; Human body movements are identified according to the human body movement probability distribution.

[0005] In some embodiments, extracting features from the millimeter wave point cloud through a one-dimensional feature encoder includes the following steps: Passing the millimeter wave point cloud through the alternating column convolution layers and row convolution layers in the one-dimensional feature encoder in sequence to extract features; wherein each of the column convolution layers is used to extract features of each individual attribute of each signal point in the millimeter wave point cloud; and each of the row convolution layers is used to extract features of all attributes of each signal point in the millimeter wave point cloud; The step of obtaining a point cloud feature vector through the first global feature pooling layer includes the following steps: The extracted features are passed through the row average pooling layer and the column average pooling layer in the first global feature pooling layer, and the extracted features are respectively encoded in one dimension row features and one dimension column features, and then the encoded row feature vectors and column feature vectors are spliced ​​into a global feature vector as the point cloud feature vector.

[0006] In some embodiments, extracting features from the RGB image through a two-dimensional feature encoder comprises the following steps: The RGB image is sequentially subjected to repeated alternating downsampling operations and two-dimensional convolutional layers in the two-dimensional feature encoder to extract features; wherein each of the downsampling operations is used to downsample the image data or image features by two times in length and width, and each of the two-dimensional convolutional layers is used to perform sliding window feature encoding on the two-dimensional spatial feature map obtained by the downsampling operation and generate a feature vector with a set number of channels; The extracted features are passed through a second global feature pooling layer to obtain an image feature vector, comprising the following steps: The feature vector extracted by the two-dimensional convolution layer is passed through the row average pooling layer and the column average pooling layer in the second global feature pooling layer, and the feature vector extracted by the two-dimensional convolution layer is respectively subjected to one-dimensional row feature encoding and one-dimensional column feature encoding, and then the encoded row feature vector and column feature vector are spliced ​​into a global feature vector as the image feature vector.

[0007] In some embodiments, the step of combining the point cloud feature vector and the image feature vector using a dual-path feature fusion and predictor to obtain an integrated feature vector comprises the following steps: Using the first fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the point cloud feature vector to generate a first dimensionality reduction feature; using the first sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the first fully connected layer, reset some weights to zero, and filter the first dimensionality reduction feature to obtain a first filtered feature; Using the second fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the image feature vector to generate a second dimensionality reduction feature; using the second sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the second fully connected layer, reset some weights to zero, and filter the second dimensionality reduction feature to obtain a second filtered feature; The first screening features and the second screening features are concatenated into the integrated feature vector using the two-way feature fusion and predictor.

[0008] In some embodiments, the step of generating a human motion probability distribution based on the integrated feature vector using the dual-path feature fusion and predictor comprises the following steps: The integrated feature vector is feature encoded using the second fully connected layer in the dual-path feature fusion and predictor, thereby generating an action category probability vector as the human action probability distribution.

[0009] In some embodiments, the identifying of human body motions according to the human body motion probability distribution comprises the following steps: The occurrence probabilities of multiple human motion behaviors are predicted according to the human motion probability distribution, and then the human motion is identified according to each predicted occurrence probability.

[0010] In some embodiments, the method further comprises the following steps: The identified human body motion is applied to at least one of security monitoring, medical care or intelligent interaction.

[0011] Another aspect of the present application further provides a motion recognition device for millimeter wave and image signals, the device comprising: A data acquisition unit, used to acquire millimeter wave point clouds and RGB images corresponding to the human body at the same time; A first feature encoding unit is configured to extract features from the millimeter wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; A second feature encoding unit is used to extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; a feature fusion unit, configured to combine the point cloud feature vector and the image feature vector using a dual-path feature fusion and predictor to obtain an integrated feature vector; A probability distribution generating unit, configured to generate a probability distribution of human motion according to the integrated feature vector using the dual-path feature fusion and predictor; The action recognition unit is used to recognize human actions according to the human action probability distribution.

[0012] Another aspect of the embodiments of the present application further provides an electronic device, including a processor and a memory; The memory is used to store programs; The processor executes the program to implement any of the above methods.

[0013] Another aspect of the embodiments of the present application further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement any of the above methods.

[0014] This application has at least the following beneficial effects: This application can obtain the millimeter wave point cloud and RGB image corresponding to the human body at the same time; extract features from the millimeter wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; use two-way feature fusion and predictor to splice the point cloud feature vector and the image feature vector to obtain an integrated feature vector; use two-way feature fusion and predictor to generate a human motion probability distribution based on the integrated feature vector; and identify human motion based on the human motion probability distribution. This application utilizes the characteristic of millimeter waves that can stably perceive human motion in low-light or occluded scenes, combined with image recognition, to achieve more comprehensive and accurate human motion recognition while ensuring privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0016] Figure 1 A flowchart of a method for motion recognition based on millimeter waves and image signals provided in an embodiment of the present application; Figure 2 An example flow chart of a method for motion recognition based on millimeter waves and image signals provided in an embodiment of the present application; Figure 3 An example flow chart of another method for motion recognition based on millimeter waves and image signals provided in an embodiment of the present application; Figure 4 This is a structural block diagram of a motion recognition device for millimeter waves and image signals provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0018] Before describing the embodiments of the present application in detail, some of the related technologies involved in the embodiments of the present application are first described as follows: Human action recognition technology based on image and millimeter wave fusion aims to combine the rich semantic information of visual images with the spatial dynamic characteristics of millimeter wave radar to enhance robust recognition of human behavior in complex environments. Image data can accurately capture posture and movement details under good lighting conditions, but is easily affected by factors such as occlusion and lighting changes. Millimeter wave radar, on the other hand, has strong penetration and anti-interference capabilities, allowing for stable perception of human movement in low-light or occluded scenarios. Effectively integrating the data from the two sensors not only complements their respective perception deficiencies but also enables more comprehensive human behavior analysis while protecting privacy. This application proposes a feature integration scheme and action recognition method for millimeter wave and image signals. By deeply fusing multi-source heterogeneous features and combining them with a spatial modeling mechanism, this method improves the recognition accuracy of key behaviors such as falls, walking, sitting, and lying down. It has strong versatility and deployment adaptability, making it particularly suitable for scenarios such as security monitoring, medical care, and intelligent interaction.

[0019] Reference Figure 1 The embodiment of the present application provides a method for motion recognition based on millimeter wave and image signals, which specifically includes the following steps S100 to S150: S100: Acquire the millimeter wave point cloud and RGB image corresponding to the human body at the same time; S110: Extracting features from the millimeter-wave point cloud through a one-dimensional feature encoder, and then passing the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; S120: Extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; S130: Using a dual-path feature fusion and predictor to concatenate the point cloud feature vector and the image feature vector to obtain an integrated feature vector; S140: Generate a human motion probability distribution according to the integrated feature vector using the dual-path feature fusion and predictor; S150: Identify human body movements according to the human body movement probability distribution.

[0020] Optionally, extracting features from the millimeter wave point cloud through a one-dimensional feature encoder comprises the following steps: Passing the millimeter wave point cloud through the alternating column convolution layers and row convolution layers in the one-dimensional feature encoder in sequence to extract features; wherein each of the column convolution layers is used to extract features of each individual attribute of each signal point in the millimeter wave point cloud; and each of the row convolution layers is used to extract features of all attributes of each signal point in the millimeter wave point cloud; The step of obtaining a point cloud feature vector through the first global feature pooling layer includes the following steps: The extracted features are passed through the row average pooling layer and the column average pooling layer in the first global feature pooling layer, and the extracted features are respectively encoded in one dimension row features and one dimension column features, and then the encoded row feature vectors and column feature vectors are spliced ​​into a global feature vector as the point cloud feature vector.

[0021] Optionally, extracting features from the RGB image through a two-dimensional feature encoder comprises the following steps: The RGB image is sequentially subjected to repeated alternating downsampling operations and two-dimensional convolutional layers in the two-dimensional feature encoder to extract features; wherein each of the downsampling operations is used to downsample the image data or image features by two times in length and width, and each of the two-dimensional convolutional layers is used to perform sliding window feature encoding on the two-dimensional spatial feature map obtained by the downsampling operation and generate a feature vector with a set number of channels; The extracted features are passed through a second global feature pooling layer to obtain an image feature vector, comprising the following steps: The feature vector extracted by the two-dimensional convolution layer is passed through the row average pooling layer and the column average pooling layer in the second global feature pooling layer, and the feature vector extracted by the two-dimensional convolution layer is respectively subjected to one-dimensional row feature encoding and one-dimensional column feature encoding, and then the encoded row feature vector and column feature vector are spliced ​​into a global feature vector as the image feature vector.

[0022] Optionally, the step of using a dual-path feature fusion and predictor to splice the point cloud feature vector and the image feature vector to obtain an integrated feature vector comprises the following steps: Using the first fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the point cloud feature vector to generate a first dimensionality reduction feature; using the first sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the first fully connected layer, reset some weights to zero, and filter the first dimensionality reduction feature to obtain a first filtered feature; Using the second fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the image feature vector to generate a second dimensionality reduction feature; using the second sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the second fully connected layer, reset some weights to zero, and filter the second dimensionality reduction feature to obtain a second filtered feature; The first screening features and the second screening features are concatenated into the integrated feature vector using the two-way feature fusion and predictor.

[0023] Optionally, the generating of a human motion probability distribution according to the integrated feature vector by using the dual-path feature fusion and predictor comprises the following steps: The integrated feature vector is feature encoded using the second fully connected layer in the dual-path feature fusion and predictor, thereby generating an action category probability vector as the human action probability distribution.

[0024] Optionally, identifying a human motion according to the human motion probability distribution comprises the following steps: The occurrence probabilities of multiple human motion behaviors are predicted according to the human motion probability distribution, and then the human motion is identified according to each predicted occurrence probability.

[0025] Optionally, the method further comprises the following steps: The identified human body motion is applied to at least one of security monitoring, medical care or intelligent interaction.

[0026] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.

[0027] like Figure 2 As shown, this embodiment method simultaneously acquires millimeter-wave signals and RGB images corresponding to a human body. The millimeter-wave signal is a point cloud, where N represents the number of points and S represents the length of the point cloud's feature vector. In millimeter-wave radar, this embodiment acquires point cloud features including the spatial coordinates (x, y, z), radial velocity, and echo power or intensity of each signal point. The RGB image is a common three-channel image, with W and H representing the image width and height. Both types of data first enter their respective feature encoders for continuous feature encoding. Then, they pass through a global feature pooling layer to generate corresponding feature vectors. The point cloud feature vector has a length of (N + S) × C, and the image feature vector has a length of (W / n + H / n) × C, where C represents the number of feature channels and n represents the downsampling factor of the original image. Finally, these two feature vectors are converted into a set of class probability vectors related to human motion through a dual-path feature integration and predictor, which is used to predict the current target's motion behavior.

[0028] Next, combine Figure 3 The solution of this embodiment is described in detail.

[0029] Specifically, this embodiment includes the following solutions: 1) One-dimensional / two-dimensional feature encoder: This embodiment designs feature encoders for millimeter-wave point cloud and RGB image data, respectively. The one-dimensional feature encoder consists of alternating column and row convolution layers. The column convolution layer has a convolution kernel dimension of (3×1×8), which encodes the one-dimensional column features of the point cloud (i.e., the features of each individual attribute of each signal point). The row convolution layer has a convolution kernel dimension of (1×3×8), which encodes the one-dimensional row features of the point cloud (i.e., the features of all attributes of each signal point). Assuming the input point cloud data has a dimension of (64×5), the one-dimensional feature encoder generates a 64×5×8 feature vector. The two-dimensional feature encoder consists of alternating downsampling operations and two-dimensional convolution layers. The downsampling operation downsamples the image data or image features by a factor of two in each downsampling operation. The convolution layer then performs sliding window feature encoding on the two-dimensional spatial feature map, generating a feature vector with 8 channels. Assuming the input image data has dimensions of (480 × 320 × 3) and is encoded three times in alternation, the two-dimensional feature encoder will generate a feature vector of 60 × 40 × 8. The number of alternations here can be used as a configuration item to balance network performance and computational complexity.

[0030] 2) Global feature pooling layer: The global feature pooling layer primarily consists of a row average pooling layer and a column average pooling layer, which perform one-dimensional row and column feature encoding on the extracted features, respectively. Assuming the input feature vector dimensions are 64×5×8, the row average pooling layer averages each 1×5 feature vector to generate a 64×8 row feature vector. The column average pooling layer averages each 64×1 feature vector to generate a 5×8 column feature vector. Finally, the encoded row and column feature vectors are concatenated to form a 1×69×8 global feature vector.

[0031] 3) Dual-path feature fusion and predictor: In order to fuse the point cloud feature vector and the image feature vector, this embodiment designs a dual-path feature integration and prediction device. This module mainly consists of two layers of fully connected layers and sparse coding operations. The first layer of fully connected layers mainly performs weighted summation of the input feature vector across feature points to generate a more compact feature vector, which is a feature dimensionality reduction process. The sparse coding operation introduces L1 regularization constraints. During the network training process, all weights of the fully connected layer are constrained to be sparsity, forcing some weights to be reset to zero, thereby achieving the purpose of feature selection. Figure 3 As can be seen in Figure 2, both the point cloud feature vector and the image feature vector are ultimately converted to 1×128 feature vectors and concatenated into a 1×256 integrated feature vector. Finally, the second fully connected layer performs final feature encoding on the integrated feature vector to generate a 1×Nc action category probability vector, where Nc represents the number of categories. This vector can be used to simultaneously predict the probability of multiple human actions.

[0032] Reference Figure 4 , an embodiment of the present application provides a motion recognition device for millimeter wave and image signals, including: A data acquisition unit, used to acquire millimeter wave point clouds and RGB images corresponding to the human body at the same time; A first feature encoding unit is configured to extract features from the millimeter wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; A second feature encoding unit is used to extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; a feature fusion unit, configured to combine the point cloud feature vector and the image feature vector using a dual-path feature fusion and predictor to obtain an integrated feature vector; A probability distribution generating unit, configured to generate a probability distribution of human motion according to the integrated feature vector using the dual-path feature fusion and predictor; The action recognition unit is used to recognize human actions according to the human action probability distribution.

[0033] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0034] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.

[0035] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0036] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0037] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0038] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.

[0039] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0040] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0041] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.

[0042] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A method for motion recognition based on millimeter wave and image signals, characterized in that: The method comprises the following steps: Obtain millimeter wave point cloud and RGB image corresponding to the human body at the same time; Extract features from the millimeter-wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; Extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; Using a dual-path feature fusion and predictor to concatenate the point cloud feature vector and the image feature vector to obtain an integrated feature vector; generating a human motion probability distribution based on the integrated feature vector using the dual-path feature fusion and predictor; Human body movements are identified according to the human body movement probability distribution.

2. The method for motion recognition based on millimeter wave and image signals according to claim 1, characterized in that: The step of extracting features from the millimeter wave point cloud through a one-dimensional feature encoder comprises the following steps: Passing the millimeter wave point cloud through the alternating column convolution layers and row convolution layers in the one-dimensional feature encoder in sequence to extract features; wherein each of the column convolution layers is used to extract features of each individual attribute of each signal point in the millimeter wave point cloud; and each of the row convolution layers is used to extract features of all attributes of each signal point in the millimeter wave point cloud; The step of obtaining a point cloud feature vector through the first global feature pooling layer includes the following steps: The extracted features are passed through the row average pooling layer and the column average pooling layer in the first global feature pooling layer, and the extracted features are respectively encoded in one dimension row features and one dimension column features, and then the encoded row feature vectors and column feature vectors are spliced ​​into a global feature vector as the point cloud feature vector.

3. The method for motion recognition based on millimeter wave and image signals according to claim 1, characterized in that: The step of extracting features from the RGB image through a two-dimensional feature encoder comprises the following steps: The RGB image is sequentially subjected to repeated alternating downsampling operations and two-dimensional convolutional layers in the two-dimensional feature encoder to extract features; wherein each of the downsampling operations is used to downsample the image data or image features by two times in length and width, and each of the two-dimensional convolutional layers is used to perform sliding window feature encoding on the two-dimensional spatial feature map obtained by the downsampling operation and generate a feature vector with a set number of channels; The extracted features are passed through a second global feature pooling layer to obtain an image feature vector, comprising the following steps: The feature vector extracted by the two-dimensional convolution layer is passed through the row average pooling layer and the column average pooling layer in the second global feature pooling layer, and the feature vector extracted by the two-dimensional convolution layer is respectively subjected to one-dimensional row feature encoding and one-dimensional column feature encoding, and then the encoded row feature vector and column feature vector are spliced ​​into a global feature vector as the image feature vector.

4. The method for motion recognition based on millimeter wave and image signals according to claim 1, characterized in that: The step of using a dual-path feature fusion and predictor to splice the point cloud feature vector and the image feature vector to obtain an integrated feature vector includes the following steps: Using the first fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the point cloud feature vector to generate a first dimensionality reduction feature; using the first sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the first fully connected layer, reset some weights to zero, and filter the first dimensionality reduction feature to obtain a first filtered feature; Using the second fully connected layer in the dual-path feature fusion and predictor to perform a weighted summation across feature points on the image feature vector to generate a second dimensionality reduction feature; using the second sparse coding operation in the dual-path feature fusion and predictor to introduce an L1 regularization constraint, perform a sparsity constraint on all weights of the second fully connected layer, reset some weights to zero, and filter the second dimensionality reduction feature to obtain a second filtered feature; The first screening features and the second screening features are concatenated into the integrated feature vector using the two-way feature fusion and predictor.

5. The method for motion recognition based on millimeter wave and image signals according to claim 1, characterized in that: The method of generating a human motion probability distribution according to the integrated feature vector by using the dual-path feature fusion and predictor comprises the following steps: The integrated feature vector is feature encoded using the second fully connected layer in the dual-path feature fusion and predictor, thereby generating an action category probability vector as the human action probability distribution.

6. The method for motion recognition based on millimeter waves and image signals according to claim 1, characterized in that: The method of identifying human body movements according to the human body movement probability distribution comprises the following steps: The occurrence probabilities of multiple human motion behaviors are predicted according to the human motion probability distribution, and then the human motion is identified according to each predicted occurrence probability.

7. The method for motion recognition based on millimeter wave and image signals according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: The identified human body motion is applied to at least one of security monitoring, medical care or intelligent interaction.

8. A motion recognition device for millimeter waves and image signals, characterized in that: The device comprises: A data acquisition unit, used to acquire millimeter wave point clouds and RGB images corresponding to the human body at the same time; A first feature encoding unit is configured to extract features from the millimeter wave point cloud through a one-dimensional feature encoder, and then pass the extracted features through a first global feature pooling layer to obtain a point cloud feature vector; A second feature encoding unit is used to extract features from the RGB image through a two-dimensional feature encoder, and then pass the extracted features through a second global feature pooling layer to obtain an image feature vector; a feature fusion unit, configured to combine the point cloud feature vector and the image feature vector using a dual-path feature fusion and predictor to obtain an integrated feature vector; A probability distribution generating unit, configured to generate a probability distribution of human motion according to the integrated feature vector using the dual-path feature fusion and predictor; The action recognition unit is used to recognize human actions according to the human action probability distribution.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.