Multi-modal human recognition system and method based on photoelectric super surface and radar fusion
The multimodal human body recognition system, which integrates optoelectronic metasurfaces and radar, solves the problems of human body recognition accuracy and robustness in complex scenarios, and achieves low power consumption and high efficiency recognition results. It is suitable for scenarios such as smart healthcare, intelligent security and disaster relief.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2025-08-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing human body recognition technologies face problems such as low spatial resolution, high hardware power consumption, and insufficient sensitivity to micro-motion features in complex scenarios, especially with a sharp drop in recognition rate in occluded, low-light, and multi-target scenarios.
A multimodal human recognition system based on optoelectronic metasurface and radar fusion is adopted, including a modulation module, a radar detection module, a spectral sensing module and a multimodal feature fusion module. The system uses a programmable metasurface for optical preprocessing, and continuous wave radar and FMCW lidar to acquire micro-Doppler spectrum and three-dimensional point cloud. Semantic alignment and fusion are performed through cross-modal Transformer, and recognition is completed on a heterogeneous edge chip.
It improves recognition accuracy in occluded, low-light, and multi-target scenarios, reduces reliance on large-scale labeled data, achieves low-power and low-latency recognition, and significantly improves the robustness and accuracy of pose recognition and vital sign detection.
Smart Images

Figure CN121008263B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of optoelectronics and radar sensing technology, and in particular to a multimodal human body recognition system and method based on the fusion of optoelectronic metasurface and radar, which can be widely used in smart healthcare, intelligent security, human-computer interaction, disaster relief and other scenarios. Background Technology
[0002] Current human recognition technologies primarily rely on single-modal perception, which faces significant limitations in complex scenarios. Specific limitations are as follows:
[0003] The "CN109379454A Millimeter Wave Radar Fall Detection Method" discloses a single-mode scheme for millimeter wave / terahertz radar, which has low spatial resolution, difficulty in separating obstructing targets, and insufficient sensitivity to micro-motion characteristics (breathing / heart rate).
[0004] The paper “US20210267530A1 3D Pose Estimation Using LiDAR” discloses FMCW LiDAR point cloud recognition, but its point cloud is sparse in low reflectivity scenes and its hardware power consumption is high.
[0005] The paper “Science Advances 2021,7:eabd7690” discloses an optical metasurface edge detection method that only outputs a single-channel edge and lacks high-level semantic and motion information coupling.
[0006] The “CVPR 2022 MVPNet” project revealed a multimodal CNN fusion scheme that relies on a large amount of labeled data but does not utilize low-frequency information such as vital signs.
[0007] Therefore, there is an urgent need for a unified sensing framework that integrates spatial geometry, dynamic spectrum and physiological micro-vibration information, and can operate at low power consumption at the edge, in order to improve the accuracy and robustness of human body recognition in complex environments (occlusion, low light, multiple targets). Summary of the Invention
[0008] The purpose of this invention is to provide a multimodal human body recognition system and method based on the fusion of optoelectronic metasurface and radar, which solves the problem of sharp drop in recognition rate of single-modal systems in occlusion, low light and multi-person scenarios; reduces the dependence on large-scale labeled data; and breaks through the bottleneck of high power consumption and high latency at the edge.
[0009] To achieve the above objectives, the present invention provides a multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar, comprising:
[0010] The modulation module includes a programmable metasurface for optical preprocessing to enhance edges;
[0011] The radar detection module, including continuous wave radar and FMCW lidar, is used to acquire the target's micro-Doppler spectrum and three-dimensional point cloud.
[0012] The spectral sensing module, including a tunable laser source and a two-dimensional material detector, is used to capture the near-infrared reflectance spectrum of the target;
[0013] The multimodal feature fusion module performs semantic alignment and fusion of multi-source signals based on cross-modal Transformer;
[0014] The recognition module includes a multi-task learning network for pose recognition, vital sign estimation, and anomaly detection.
[0015] Heterogeneous edge chips are used to perform optical neural network inference and low-latency recognition.
[0016] Preferably, the design and fabrication of the metasurface cascaded optoelectronic device involves an optical metasurface comprising a phase-tunable cascaded structure. The metasurface cascaded optoelectronic device is designed using a reverse design method, starting with phase layering: the first layer undergoes radial correction, and its phase distribution is as follows...
[0017] φ1(r)=sin(6r)e -2r ;
[0018] Enhance center-edge contrast.
[0019] The phase distribution of the second layer (angular petals) is as follows:
[0020] φ2(θ)=0.5cos(10θ);
[0021] (Effective only in the region r < 0.8), producing multi-lobed diffraction.
[0022] Then comes the superposition and clipping. Its reflection phase is jointly defined by the following functions:
[0023]
[0024] θ = atan2(y,x).
[0025] Using the above parameter values, electron beam (E-beam) or photolithography + etching are used to perform micro-nano fabrication on the substrate to create two metasurface structures, each corresponding to a different depth / material, in order to achieve the desired phase distribution.
[0026] Preferably, the point cloud encoding module adopts a KP-Conv or Sparse Voxel Transformer structure to tokenize the 3D point cloud data for semantic fusion.
[0027] Preferably, the micro-Doppler vital signs channel uses a graph wavelet convolutional network to periodically model respiratory and heart rate signals.
[0028] Preferably, the multimodal fusion module further incorporates tensor decomposition (CP) and graph attention network (GAT) to assign interpretable weights to each modality feature.
[0029] Preferably, the heterogeneous chip has a photoelectric hybrid computing unit and uses a Mach-Zehnder interference structure to perform simulated neural computation, achieving sub-10 microsecond level edge reasoning capability.
[0030] Preferably, the system has the ability to maintain robustness of recognition under different environmental conditions, including but not limited to: scenes where the target is partially occluded; scenes with low light or drastic changes in lighting; complex scenes with multiple overlapping or dynamic interactions in a group; and improves the accuracy of target boundary determination and dynamic behavior separation by fusing optical edge and micro-Doppler spectrum.
[0031] Preferably, the fusion model has multi-task transfer learning capability and can share feature representations and parameter structures among the following tasks: pose recognition and skeletal keypoint detection;
[0032] Non-contact vital sign detection (respiratory rate, heart rate); fall detection and abnormal action recognition; the model adopts a cross-task shared encoder and combines task-specific loss branch optimization to improve transfer ability and generalization performance under small sample conditions.
[0033] This invention also provides a multimodal human body recognition method based on the fusion of optoelectronic metasurface and radar, comprising the following steps:
[0034] S1. Optical high-pass filtering is performed using a phase-programmable optical metasurface to extract target edge features;
[0035] S2. Use continuous wave radar to capture the micro-Doppler spectrum and use FMCW lidar to obtain three-dimensional point clouds;
[0036] S3. Scan the near-infrared spectrum of the target using a tunable laser source, and generate a spectral vector using a two-dimensional material detector;
[0037] S4. A cross-modal Transformer structure is used to perform semantic fusion of edge features, point clouds, micro-Doppler spectra, and spectral features.
[0038] S5. Performs pose recognition, vital sign estimation, and anomaly detection based on a multi-task learning network;
[0039] S6. Low-latency inference is achieved through heterogeneous edge computing chips.
[0040] Preferably, the edge detection module employs an optical edge extraction method based on a metasurface adaptive kernel function, using a Laplace-Gaussian (S-LoG) kernel with directional second derivatives, introducing a direction parameter θ and cross-term weight w, and a Gaussian term.
[0041]
[0042] Second derivative components
[0043]
[0044] Directed nuclei
[0045]
[0046] Then, the kernel function is zero-mean normalized.
[0047]
[0048] Construct two point spread functions, positive and negative PSFK. + K - ;
[0049] K + =max(K,0), K — =max(-K,0),
[0050] The following optical filter is used to output simulated edge light intensity values.
[0051] E 仿真 (x,y)=I*K + (x,y)-I*K - (x,y);
[0052] The features also include updates to the loss function and parameters.
[0053] The loss function is commonly represented by mean squared error (MSE).
[0054]
[0055] During gradient calculation, σ, θ, and w all participate in the convolution operation, thus allowing the gradient to propagate backward.
[0056]
[0057] Parameter updates are performed using stochastic gradient descent (SGD) or adaptive moment estimation (Adam) optimization algorithms.
[0058]
[0059] Simultaneously, joint optimization of metasurface phase design parameters can be performed.
[0060] Preferably, the kernel function is trained using a backpropagation algorithm based on stochastic gradient descent (SGD) or adaptive moment estimation (Adam) optimizer to dynamically adjust the metasurface phase function and kernel parameters, so that the final simulated edge is as close as possible to the target edge in terms of spatial frequency domain and structural details.
[0061] Preferably, the spectral feature extraction module includes: a tunable laser that emits a narrowband scanning spectrum covering the range of 400–2000 nm; a photodetector array based on two-dimensional materials (such as MoS2, Graphene, etc.); combining the reflectivity of different wavelengths of the target into a spectral vector; and inputting the spectral vector into a one-dimensional convolutional autoencoder (1D-CNN-AE) or a spectral Transformer for high-dimensional feature embedding and nonlinear compression.
[0062] Preferably, the embedded spectral features have a dimension of 128–512, and are aligned with other modalities (radar, point cloud, vital signs) through the Cross-Attention module to construct a unified multimodal semantic space.
[0063] Preferably, the vital signs detection module uses a graph wavelet CNN to perform frequency domain modeling on the low-frequency vibration components in the radar signal, in order to extract the target's periodic breathing, heartbeat and other vital micro-vibration features.
[0064] Preferably, the multimodal fusion module adopts the following structure: it performs feature alignment in the temporal and spatial dimensions between modalities based on the Transformer Encoder architecture; it uses Graph Attention Network (GAT) and Tensor Decomposition (Canonical Polyadic) to assign interpretable dynamic weights to the multimodal inputs; and the multi-task outputs correspond to downstream tasks such as pose angle prediction, anomaly detection and classification, and vital sign regression.
[0065] Preferably, the system is integrated into an edge heterogeneous chip and uses a Mach-Zehnder interference structure or a photoelectric hybrid ReRAM array to achieve low-power, low-latency feature extraction and inference. The overall system latency is less than 25ms and the power consumption is no more than 4W, making it suitable for indoor intelligent monitoring, disaster relief and wearable health monitoring scenarios.
[0066] Therefore, the multimodal human body recognition system and method based on the fusion of optoelectronic metasurface and radar, which adopts the above-mentioned structure, has the following beneficial effects:
[0067] (1) This invention reduces the pose angle error in occluded scenes by 38%, improves the F1-score in low-light environments by 42%, and improves the ID-F1 of multi-target tracking by 23%. It also enhances target boundary determination through the fusion of optical edge and micro-Doppler. The feature encoder is shared for tasks such as pose recognition and vital sign detection, and the average transfer performance is improved by 8–15% under small sample conditions.
[0068] (2) The present invention uses a photoelectric hybrid chip to achieve sub-millisecond inference with a power consumption of only 4W. Compared with the traditional solution, the unit power consumption recognition efficiency is improved by 3.2 times, which is suitable for wearable devices and real-time monitoring.
[0069] (3) The present invention dynamically allocates modal weights across modal Transformer-GAT, improves feature clustering separation, and achieves an anomaly detection F1-score of 0.92, which is better than the single-modal scheme.
[0070] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0071] Figure 1 This is a schematic diagram of the overall system architecture and a diagram of the underlying acceleration module and multimodal fusion structure.
[0072] Figure 2 The flowchart shows the human target edge detection algorithm based on metasurface cascade + adaptive kernel.
[0073] Figure 3 This is a diagram of the multimodal fusion network structure.
[0074] Figure 4 The diagram shows a block diagram of a multimodal human pose recognition algorithm based on the fusion of edge images and Doppler radar spectrum.
[0075] Figure 5 The image shows the effect of human posture recognition based on metasurface spectroscopy and radar spectroscopy (as an example);
[0076] Figure 6 A comparison chart of recognition performance between multimodal fusion and single-modal models;
[0077] Figure 7 A comparison chart of inference latency and power consumption for different hardware platforms;
[0078] Figure 8 Clustering visualization (left image shows features before fusion; right image shows features after fusion);
[0079] Figure 9 Heatmaps of different detection mechanisms in different scenarios (mainly accuracy and transfer score);
[0080] Figure 10This is a high-resolution detection result image of the algorithm proposed in this invention. Detailed Implementation
[0081] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0082] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0083] Example
[0084] like Figure 1-8 As shown, this invention provides a multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar, comprising:
[0085] The modulation module includes a programmable metasurface for optical preprocessing to enhance edges;
[0086] The radar detection module, including continuous wave radar and FMCW lidar, is used to acquire the target's micro-Doppler spectrum and three-dimensional point cloud.
[0087] The spectral sensing module, including a tunable laser source and a two-dimensional material detector, is used to capture the near-infrared reflectance spectrum of the target;
[0088] The multimodal feature fusion module performs semantic alignment and fusion of multi-source signals based on cross-modal Transformer;
[0089] The recognition module includes a multi-task learning network for pose recognition, vital sign estimation, and anomaly detection.
[0090] Heterogeneous edge chips are used to perform optical neural network inference and low-latency recognition.
[0091] Design and fabrication of metasurface cascaded optoelectronic devices: Optical metasurfaces include phase-tunable cascaded structures. A reverse design approach is used to design these cascaded optoelectronic devices. The first step is phase layering: the first layer undergoes radial correction, and its phase distribution is as follows...
[0092] φ1(r)=sin(6r)w -2r ;
[0093] Enhance center-edge contrast.
[0094] The phase distribution of the second layer (angular petals) is as follows:
[0095] φ2(θ)=0.5cos(10θ);
[0096] (Effective only in the region r < 0.8), producing multi-lobed diffraction.
[0097] Then comes the superposition and clipping. Its reflection phase is jointly defined by the following functions:
[0098]
[0099] θ = atan2(y,x).
[0100] Using the above parameter values, electron beam (E-beam) or photolithography + etching are used to perform micro-nano fabrication on the substrate to create two metasurface structures, each corresponding to a different depth / material, in order to achieve the desired phase distribution.
[0101] The point cloud encoding module uses a KP-Conv or Sparse VoxelTransformer structure to tokenize 3D point cloud data for semantic fusion.
[0102] The micro-Doppler vital signs channel uses a graph wavelet convolutional network to periodically model respiratory and heart rate signals.
[0103] The multimodal fusion module further introduces tensor decomposition (CP) and graph attention network (GAT) to assign interpretable weights to each modality feature.
[0104] The heterogeneous chip has a hybrid optoelectronic computing unit and uses a Mach-Zehnder interference structure to simulate neural computation, achieving sub-10 microsecond level edge inference capabilities.
[0105] The system is capable of maintaining robust recognition under different environmental conditions, including but not limited to: scenes where the target is partially occluded; scenes with low light or drastic changes in lighting; complex scenes with multiple overlapping or dynamic interactions in a group; and by fusing optical edges and micro-Doppler spectrum, it improves the accuracy of target boundary determination and dynamic behavior separation.
[0106] The fusion model has multi-task transfer learning capabilities and can share feature representations and parameter structures across the following tasks: pose recognition and skeletal keypoint detection; non-contact vital sign detection (respiratory rate, heart rate); fall detection and abnormal action recognition; the model adopts a cross-task shared encoder and combines task-specific loss branch optimization to improve transfer ability and generalization performance under small sample conditions.
[0107] This invention also provides a multimodal human body recognition method based on the fusion of optoelectronic metasurface and radar, comprising the following steps:
[0108] S1. Optical high-pass filtering is performed using a phase-programmable optical metasurface to extract target edge features;
[0109] S2. Use continuous wave radar to capture the micro-Doppler spectrum and use FMCW lidar to obtain three-dimensional point clouds;
[0110] S3. Scan the near-infrared spectrum of the target using a tunable laser source, and generate a spectral vector using a two-dimensional material detector;
[0111] S4. A cross-modal Transformer structure is used to perform semantic fusion of edge features, point clouds, micro-Doppler spectra, and spectral features.
[0112] S5. Performs pose recognition, vital sign estimation, and anomaly detection based on a multi-task learning network;
[0113] S6. Low-latency inference is achieved through heterogeneous edge computing chips.
[0114] The edge detection module employs an optical edge extraction method based on a metasurface adaptive kernel function. It utilizes a Laplace-Gaussian (S-LoG) kernel with directional second derivatives, introducing a direction parameter θ and cross-term weights w, along with Gaussian terms.
[0115]
[0116] Second derivative components
[0117]
[0118] Directed nuclei
[0119]
[0120] Then, the kernel function is zero-mean normalized.
[0121]
[0122] Construct two point spread functions, positive and negative PSFK. + K - ;
[0123] K + =max(K,0), K — =max(-K,0),
[0124] The following optical filter is used to output simulated edge light intensity values.
[0125] E 仿真 (x,y)=I*K + (x,y)-I*K - (x,y);
[0126] The features also include updates to the loss function and parameters.
[0127] The loss function is commonly represented by mean squared error (MSE).
[0128]
[0129] During gradient calculation, σ, θ, and w all participate in the convolution operation, thus allowing the gradient to propagate backward.
[0130]
[0131] Parameter updates are performed using stochastic gradient descent (SGD) or adaptive moment estimation (Adam) optimization algorithms.
[0132]
[0133] Simultaneously, joint optimization of metasurface phase design parameters can be performed.
[0134] The kernel function is trained using a backpropagation algorithm based on stochastic gradient descent (SGD) or adaptive moment estimation (Adam) optimizer, dynamically adjusting the metasurface phase function and kernel parameters so that the final simulated edge is as close as possible to the target edge in terms of spatial frequency domain and structural details.
[0135] The spectral feature extraction module includes: a tunable laser that emits a narrowband scanning spectrum covering the range of 400–2000 nm; a photodetector array based on two-dimensional materials (such as MoS2, Graphene, etc.); combining the reflectivity of different wavelengths of the target into a spectral vector; and inputting the spectral vector into a one-dimensional convolutional autoencoder (1D-CNN-AE) or a spectral Transformer for high-dimensional feature embedding and nonlinear compression.
[0136] The embedded spectral features have a dimension of 128–512 and are aligned with other modalities (radar, point cloud, vital signs) through the Cross-Attention module to construct a unified multimodal semantic space.
[0137] The vital signs detection module uses a graph wavelet CNN to model the low-frequency vibration components in the radar signal in the frequency domain, in order to extract the target's periodic breathing, heartbeat and other vital micro-vibration features.
[0138] The multimodal fusion module adopts the following structure: it performs feature alignment in the temporal and spatial dimensions between modalities based on the Transformer Encoder architecture; it uses Graph Attention Network (GAT) and Canonical Polyadic to assign interpretable dynamic weights to the multimodal inputs; and the multi-task outputs correspond to downstream tasks such as pose angle prediction, anomaly detection and classification, and vital sign regression.
[0139] The system is integrated into an edge heterogeneous chip and achieves low-power, low-latency feature extraction and inference through a Mach-Zehnder interference structure or a photoelectric hybrid ReRAM array. The overall system latency is less than 25ms and the power consumption is no more than 4W, making it suitable for indoor intelligent monitoring, disaster relief and wearable health monitoring scenarios.
[0140] Hardware layer: Two-layer phase-programmable metasurfaces are used as optical high-pass filters, which are combined with continuous wave radar (CWRadar) and VCSEL-FMCWLiDAR to form a four-stream parallel sensing link;
[0141] Algorithm layer: Construct adaptive kernel function optical edge extraction + micro-Doppler time-frequency analysis + spectral Transformer embedding + KP-Conv point cloud encoding, and complete semantic fusion through cross-modal Transformer-GAT;
[0142] The mathematical model of KP-Conv+Transformer-GAT includes the KP-Conv (Kernel Point Convolution) mathematical model and the Transformer-GAT cross-modal semantic fusion model formula.
[0143] The KP-Conv (Kernel Point Convolution) mathematical model is shown below:
[0144] KP-Conv is a method for performing local convolution on each center point p in a point cloud. Its core idea is to use kernel points as sampling centers and perform a weighted summation of neighboring points. For each center point p, let its neighboring points be {p...}j If the kernel point is {km} and the corresponding weight function is Wm(·), then the output feature is:
[0145]
[0146] Where N(p) is the neighborhood of point p; f(p) j ) represents the input features of the neighborhood points; k m For the m-th core point (fixed or learnable); W m This represents the weighting function (usually a Gaussian or linear function) for that kernel point. Key point: KP-Conv preserves Euclidean geometry, making it suitable for unstructured point clouds.
[0147] The formula for the Transformer-GAT cross-modal semantic fusion model is shown below:
[0148] This module integrates two attention mechanisms: Transformer handles long-range modeling, capturing global context; GAT (GraphAttentionNetwork) enhances the structural relationships between points, adapting to cross-modal graph representations. The attention mechanism in Transformer, for the input feature sequence X∈R... n×d Where n is the number of tokens and d is the feature dimension, defined as: Q = XW Q K = XW K V = XW V Attention output is
[0149]
[0150] GAT's graph attention mechanism:
[0151] Given a graph G = (V, E), where each node has a feature hi ∈ RF, the influence of adjacent node j ∈ N(i) on i is:
[0152] e ij =LeakyReLU(a T Wh i |Wh j ));
[0153]
[0154] h″ i =σ(∑ k∈N(i) α ij Wh j );
[0155] Where W is a learnable linear transformation matrix, a is a learnable attention vector, and σ is an activation function. Transformer-GAT is embedded in concatenation or parallel for cross-modal semantic fusion of point clouds and spectral images.
[0156] Computation Layer: Explanation of Inference Acceleration Mapping Mechanisms on MZI-NPU / ReRAM
[0157] (1)MZI-NPU (Mach-Zehnder Interferometer-based Neural Processing Unit)
[0158] MZI is an interferometric optical computing element that can be used for matrix-vector multiplication (MAC) operations.
[0159] Calculation mechanism: Encode the weights W in the neural network into the phase control matrix θ of the MZI. ij The input vector x is converted into a light intensity signal input; the multiplication y = Wx is realized through the MZI grid structure to complete one inference.
[0160] Optical MAC advantages: No electrical storage required, latency < nanoseconds, power consumption approximately 1% of CMOS (no heat loss), high parallelism, up to 10 times per second. 15 MACs (reported in Nature 2025). Implementing MZI-NPU inference on a hybrid optoelectronic-ReRAM chip accelerates convolution and attention operations with a total system latency of <25ms and power consumption of ≤4W. Key advantages include: in-memory computing to avoid memory bandwidth bottlenecks; support for parallel batch inference; and the ability to be embedded in edge devices while meeting ≤4W power consumption design requirements.
[0161] (2) ReRAM (Resistive RAM) acceleration mechanism
[0162] ReRAM is used to construct programmable non-volatile weight arrays and supports the following operations:
[0163] Weight mapping: The neural network weights W∈Rm×n are mapped to the conduction states Rij in the ReRAM cross array; the input signal is applied to the horizontal row of the ReRAM, and the output current is obtained by integrating the column currents.
[0164]
[0165] Where yj is the output current, V i x is the input voltage. i For the input signal, R ij W represents the resistance of the cross array. ij These are the weights of the neural network.
[0166] Taking a metasurface reflector fabricated as an example, with a double-layer TiO2 / SiO2 nanopillars, a unit period of 200nm, and a thickness difference of 30nm.
[0167] Its phase function is
[0168] φ1(r)=sin(6r)e -2r φ2(θ)=0.5cos(10θ);
[0169] After being stacked, they are formed in one step by photolithography and reactive ion etching.
[0170] Specifically, the continuous wave radar has a carrier frequency of 77 GHz, a transmit power of 10 dBm, a bandwidth of 1 GHz, and a frame rate of 50 fps; the VCSEL-FMCWLiDAR has a center wavelength of 905 nm, a frequency modulation rate of 200 MHz / μs, and a point cloud density of 30 kpts / frame.
[0171] Specifically, the tunable laser in the spectral detection channel is stepped by 1 nm @ 400–2000 nm; the MoS2-InGaAs detector array is 512 ch.
[0172] Specifically, the mixed-signal SoC + MZI-NPU array is 64×64, with an INT4 throughput of >2 TOPS / W.
[0173] In the adaptive kernel function optical edge extraction, the kernel is initialized with σ = 0.2, θ = 42°, and w = 0.4.
[0174] Normalization of positive / negative kernels: K + =max(K,0), K — =max(-K,0),
[0175] Iteration: Adam (β1 = 0.9, β2 = 0.999, lr = 10-3), converged in 20 epochs, PSNR increased by 4.6dB.
[0176] Specifically, for micro-Doppler and point cloud encoding, the radar signal is processed by STFT (window length 256, hop 64) → 128×128 amplitude spectrum; the point cloud is used to extract 256-dim local features using KP-Conv (r=0.1m, K=16); the spectrum is compressed to 128-dim using 1D-CNN-AE.
[0177] Specifically, the key code for cross-modal Transformer-GAT fusion is as follows:
[0178] Input:{E_edge,S_radar,P_point,L_spec}
[0179] Encoder:4×{MultiHeadAttn(d_model=512,n_head=8),FFN}
[0180] GAT: K = 4 modal nodes, adjacency matrix dynamically learned.
[0181] Output: 1024-dim Semantic Token → Multitasking Header
[0182] Task 1: Posture regression – MSELoss; Task 2: Fall / abnormality classification – FocalLoss (γ=2); Task 3: Respiration / heart rate regression – SmoothL1Loss.
[0183] Specifically, in order to adapt to different scenarios, the present invention adopts different algorithms and mechanisms, as shown in Table 1.
[0184] Table 1 Processing Mechanisms for Different Scenarios
[0185]
[0186] Figure 9 This presentation showcases the thermodynamic maps of different detection mechanisms (such as point clouds, spectra, and multimodal detection) in various scenarios, primarily focusing on the accuracy and transfer learning capabilities of the detection algorithms. The detection mechanisms are mainly considered in three scenarios: 1) single-modal recognition performance using 3D point cloud data; 2) single-modal recognition performance using spectral images; and 3) recognition performance after fusing point clouds and spectral images. The analysis focuses on the detection accuracy and transfer learning capabilities of three algorithms across five different scenarios. These five scenarios include: fall detection—the system's accuracy in detecting human falls; abnormal behavior—the system's ability to detect abnormal behavior (such as violent twisting or paralysis); occlusion scenarios—the robustness of target recognition when multiple people / objects are occluded; low-light environments—the ability to preserve structural information in low-light images; and multi-person scenarios—the ability to detect multiple people simultaneously appearing in the field of view. In all tasks, multimodal fusion outperforms single-modal input, with a maximum improvement of +0.25. Spectral images show performance degradation in low-light and occluded scenarios, indicating a strong dependence on structural information. Point clouds perform well in multi-person scenarios but are prone to misjudgment in occlusion and dynamic behavior. The KP-Conv+Transformer-GAT architecture proposed in this invention demonstrates strong robustness in multi-task transfer and generalization. Compared with existing technologies, this invention has significant advantages in five aspects, as shown in Table 2. Figure 10 The results demonstrated by this invention show that the detection effect is extremely high resolution and very clear texture information.
[0187] Table 2 Comparison of the present invention with the prior art
[0188]
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar, characterized in that, include: Modulation module: Includes a phase-programmable optical metasurface for optical high-pass filtering of the target to achieve edge enhancement; Radar detection module: includes continuous wave radar and frequency modulated continuous wave (FMCW) lidar, used to acquire the target's micro-Doppler spectrum and three-dimensional point cloud data, respectively; Spectral sensing module: Includes a tunable laser source and a two-dimensional material detector, used to capture the near-infrared reflectance spectrum of the target; Multimodal feature fusion module: Employs a cross-modal Transformer structure to perform semantic-level fusion of optical edges, point clouds, micro-Doppler spectra, and spectral features; Recognition module: Based on a multi-task learning network, it outputs pose recognition, vital sign estimation, and anomaly detection results; Heterogeneous edge computing chip: integrates optoelectronic hybrid computing units for low-latency feature extraction and inference.
2. The multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar according to claim 1, characterized in that: The optical metasurface has a two-layer cascaded structure, and its reflection phase distribution satisfies: Φ casc (𝑥,𝑦)= ; in, r= , .
3. The multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar according to claim 1, characterized in that: The point cloud data is encoded using kernel-point convolution (KP-Conv) or sparse voxel Transformer to generate tokenized features for semantic fusion.
4. The multimodal human body recognition system based on the fusion of optoelectronic metasurface and radar according to claim 1, characterized in that: The micro-Doppler spectrum is modeled periodically using a graph wavelet CNN to model respiratory and heart rate signals.
5. A multimodal human body recognition system based on photoelectric metasurface and radar fusion according to claim 1, characterized in that: The multimodal feature fusion module further introduces Graph Attention Network (GAT) and Tensor Decomposition (CP) to assign dynamic and interpretable weights to each modality feature.
6. A multimodal human body recognition system based on photoelectric metasurface and radar fusion according to claim 1, characterized in that: The heterogeneous edge computing chip uses a Mach-Zehnder interferometer (MZI) to achieve optical neural computing, or accelerates weight mapping through a ReRAM cross array.
7. A multimodal human body recognition method based on the fusion of optoelectronic metasurface and radar, applied to the system described in any one of claims 1-6, characterized in that: Includes the following steps: S1. Optical high-pass filtering is performed using a phase-programmable optical metasurface to extract target edge features; S2. Use continuous wave radar to capture the micro-Doppler spectrum and use FMCW lidar to obtain three-dimensional point clouds; S3. Scan the near-infrared spectrum of the target using a tunable laser source, and generate a spectral vector using a two-dimensional material detector; S4. A cross-modal Transformer structure is used to perform semantic fusion of edge features, point clouds, micro-Doppler spectra, and spectral features. S5. Performs pose recognition, vital sign estimation, and anomaly detection based on a multi-task learning network; S6. Low-latency inference is achieved through heterogeneous edge computing chips.
8. The multimodal human body recognition method based on photoelectric metasurface and radar fusion according to claim 7, characterized in that: In step S1, the edge extraction employs an adaptive kernel function optical edge extraction method, whose directional convolution kernel is defined as: w ; in, For Gaussian terms; , , It is the second derivative component.
9. A multimodal human body recognition method based on photoelectric metasurface and radar fusion according to claim 8, characterized in that: The parameters of the adaptive kernel function Optimized using the backpropagation algorithm, the loss function is: ; The optimizer employs stochastic gradient descent (SGD) or adaptive moment estimation (Adam).
10. A multimodal human body recognition method based on photoelectric metasurface and radar fusion according to claim 7, characterized in that: In step S4, the spectral vector is compressed into high-dimensional embedded features by a one-dimensional convolutional autoencoder 1D-CNN-AE or a spectral Transformer, with dimensions of 128–512.
Citation Information
Patent Citations
Electronic equipment
CN109379454A
Human body posture recognition method based on FMCW radar signal
CN113313040A
Millimeter wave radar behavior identification method based on multi-task cross-modal attention
CN120375473A