Multi-modal human body recognition system and method based on photoelectric metasurface and radar fusion
By integrating optoelectronic metasurfaces with radar, a multimodal human body recognition system has been developed, which solves the problems of human body recognition accuracy and robustness in complex scenarios. It achieves high-precision recognition with low latency and low power consumption, and is suitable for scenarios such as smart healthcare, intelligent security, and disaster relief.
Patent Information
- Application Number
- CN202511110905.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing human recognition technologies face challenges such as occlusion, low light, and multiple targets in complex scenarios. Single-modal perception systems struggle to separate occluded targets, consume a lot of power, rely on large amounts of labeled data, and lack sufficient recognition accuracy and robustness.
A multimodal human recognition system based on optoelectronic metasurface and radar fusion is adopted, including a modulation module, a radar detection module, a spectral sensing module and a multimodal feature fusion module. The system utilizes a programmable metasurface for optical preprocessing, continuous wave radar and FMCW lidar to acquire micro-Doppler spectrum and three-dimensional point cloud, and combines a tunable laser source and a two-dimensional material detector to capture near-infrared reflection spectrum. Semantic alignment and fusion are performed through a cross-modal Transformer, and low-latency recognition is achieved through heterogeneous edge chips.
Improved recognition rate in occluded, low-light, and multi-person scenarios; reduced reliance on large-scale labeled data; sub-millisecond inference and low power consumption; improved recognition accuracy and robustness; cross-modal Transformer-GAT dynamic allocation of modal weights; improved feature clustering separation; and anomaly detection F1-score of 0.92.
Smart Images

Figure CN121008263A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of optoelectronics and radar perception, in particular to a multi-modal human body recognition system and method based on the fusion of optoelectronic super surface and radar, which can be widely applied to intelligent medical treatment, intelligent security, human-computer interaction, disaster rescue and the like. BACKGROUND
[0002] The existing human body recognition technology mainly relies on single modal perception, and faces serious bottlenecks in complex scenes. The specific limitations are as follows:
[0003] The “CN109379454A millimeter wave radar fall detection method” discloses a millimeter wave / terahertz radar single modal scheme, which has low spatial resolution and is difficult to separate the occluded target, and is insufficiently sensitive to micro-motion characteristics (respiration / heart rate);
[0004] “US20210267530A1 3D Pose Estimation Using LiDAR” discloses FMCW LiDAR point cloud recognition, but the point cloud is sparse in a low reflectivity scene, and the hardware power consumption is high;
[0005] “Science Advances 2021,7:eabd7690” discloses optical super surface edge detection: only outputs a single channel edge, and lacks coupling of high-level semantics and motion information;
[0006] “CVPR 2022 MVPNet” discloses a multi-modal CNN fusion scheme: relies on a large amount of labeled data, and does not utilize low-frequency information such as vital signs.
[0007] Therefore, there is an urgent need for a unified perception framework that simultaneously fuses spatial geometry, dynamic spectrum and physiological micro-vibration information, and can run on an edge with low power consumption, to improve the accuracy and robustness of human body recognition in complex environments (occlusion, low light, multiple targets). SUMMARY
[0008] The present application aims to provide a multi-modal human body recognition system and method based on the fusion of optoelectronic super surface and radar, which solves the problem of sudden drop in recognition rate of single modal system in occlusion, weak light and multiple person scenes, reduces the dependence on large-scale labeled data, and breaks through the bottleneck of high power consumption and high delay on the edge.
[0009] To achieve the above-mentioned purpose, the present application provides a multi-modal human body recognition system based on the fusion of optoelectronic super surface and radar, comprising:
[0010] A modulation module comprising a programmable super surface for realizing edge-enhanced optical preprocessing;
[0011] A radar detection module, including a continuous wave radar and an FMCW laser radar, is used to obtain the micro-Doppler spectrum and three-dimensional point cloud of the target;
[0012] A spectrum sensing module, including a tunable laser source and a two-dimensional material detector, is used to capture the near-infrared reflection spectrum of the target;
[0013] A multi-modal feature fusion module is based on a cross-modal Transformer to perform semantic alignment and fusion on multi-source signals;
[0014] An identification module includes a multi-task learning network for gesture recognition, vital sign estimation, and anomaly detection;
[0015] A heterogeneous edge chip is used to complete optical neural network inference and low-latency identification.
[0016] Preferably, the metasurface cascaded optoelectronic device is designed and manufactured, and the optical metasurface includes a phase-adjustable cascaded structure. The metasurface cascaded optoelectronic device is designed by using a reverse design method. First, phase layering is performed: the first layer corrects the radial direction, and the phase distribution is
[0017] φ1(r)=sin(6r)e -2r ;
[0018] Enhance the center-edge contrast.
[0019] The phase distribution of the second layer (angular petal) is
[0020] φ2(θ)=0.5cos(10θ);
[0021] (valid only in the r<0.8 region), which generates multi-beam petal diffraction.
[0022] Then, superposition and clipping are performed. The reflection phase is jointly defined by the following functions:
[0023]
[0024] θ=atan2(y,x)。
[0025] By using the above parameter values, micro-nano processing is performed on the substrate by using an electron beam (E-beam) or lithography + etching to manufacture two-layer metasurface structures corresponding to different depths / materials to achieve the required phase distribution.
[0026] Preferably, the point cloud encoding module adopts a KP-Conv or Sparse Voxel Transformer structure to perform Tokenization processing on three-dimensional point cloud data for semantic fusion.
[0027] Preferably, the micro-Doppler vital sign channel adopts a graph wavelet convolution network to model the periodicity of the respiration and heart rate signals.
[0028] Preferably, the multi-modal fusion module further introduces tensor decomposition (CP) and a graph attention network (GAT) to assign interpretable weights to each modality feature.
[0029] Preferably, the heterogeneous chip has an optoelectronic hybrid computing unit that adopts a Mach-Zehnder interference structure to perform analog neural operations, achieving an edge inference capability of sub-10 microseconds.
[0030] Preferably, the system has the ability to maintain recognition robustness under different environmental conditions, including but not limited to: scenes where the target part is partially occluded; scenes with weak light or dramatic light changes; complex scenes with multiple people overlapping or group dynamic interactions; and by fusing optical edges and micro-Doppler spectra, the accuracy of target boundary determination and dynamic behavior separation is improved.
[0031] Preferably, the fusion model has multi-task transfer learning capability, which can share feature representation and parameter structure between the following tasks: gesture recognition and skeletal key point detection.
[0032] Non-contact vital sign detection (respiration rate, heart rate); fall detection and abnormal motion recognition; the model uses a cross-task shared encoder and combines task-specific loss branch optimization to improve transferability and generalization performance under small samples.
[0033] The present application also provides a multi-modal human recognition method based on optical-electric super surface and radar fusion, comprising the following steps:
[0034] S1, performing optical high-pass filtering through a phase programmable optical super surface to extract target edge features;
[0035] S2, capturing micro-Doppler spectra using a continuous wave radar, and obtaining three-dimensional point clouds using an FMCW laser radar;
[0036] S3, scanning the near-infrared spectrum of the target through a tunable laser source, and generating a spectrum vector from a two-dimensional material detector;
[0037] S4, performing semantic fusion of edge features, point clouds, micro-Doppler spectra, and spectral features using a cross-modal Transformer structure;
[0038] S5, performing gesture recognition, vital sign estimation, and anomaly detection based on a multi-task learning network;
[0039] S6, completing low-latency inference through a heterogeneous edge computing chip.
[0040] Preferably, the edge detection module adopts an optical edge extraction method based on a super-surface adaptive kernel function, a Laplace-Gaussian (S-LoG) kernel based on directional second-order derivative, and introduces a direction parameter θ and a cross-term weight w, a Gaussian term
[0041]
[0042] Second-order derivative component
[0043]
[0044] Directional kernel
[0045]
[0046] And make the kernel function zero mean and normalized:
[0047]
[0048] Construct two positive and negative point spread functions PSFK + , K - ;
[0049] K + = max(K, 0), K — = max(-K, 0),
[0050] The following optical filter is used to output the simulated edge light intensity value
[0051] E 仿真 (x, y) = I * K + (x, y) - I * K - (x, y);
[0052] The features also include loss function and parameter update
[0053] The loss function is commonly expressed as mean square error (MSE)
[0054]
[0055] When calculating the gradient, σ, θ and w all participate in the convolution operation, so the gradient can be back-propagated:
[0056]
[0057] The random gradient descent (Stochastic Gradient Descent, SGD) or adaptive moment estimation optimization algorithm (Adaptive Moment Estimation, Adam) is used for parameter update.
[0058]
[0059] Meanwhile, the super surface phase design parameters can be jointly optimized.
[0060] Preferably, the training of the kernel function adopts a back propagation algorithm based on a stochastic gradient descent (SGD) or an adaptive moment estimation (Adam) optimizer, and the super surface phase function and the kernel parameters are dynamically adjusted so that the final simulation edge is maximally close to the target edge in the spatial frequency domain and structural details.
[0061] Preferably, the spectral feature extraction module comprises: a tunable laser emitting a narrow-band scanning spectrum covering a range of 400-2000 nm; a photodetector array based on two-dimensional materials (such as MoS2, Graphene, etc.); a combination of different wavelengths of the target reflectivity into a spectral vector; and the spectral vector is input into a one-dimensional convolutional autoencoder (1D-CNN-AE) or a spectral Transformer for high-dimensional feature embedding and nonlinear compression.
[0062] Preferably, the embedded spectral feature dimension is 128-512, and is aligned with other modalities (radar, point cloud, vital signs) through a Cross-Attention module, for constructing a unified multi-modal semantic space.
[0063] Preferably, the vital sign detection module adopts a graph wavelet convolutional network (Graph Wavelet CNN) to model the low-frequency vibration components in the radar signal in the frequency domain, so as to extract the periodic breathing, heartbeat and other life micro-vibration features of the target.
[0064] Preferably, the multi-modal fusion module adopts the following structure: a Transformer Encoder architecture based feature alignment in time and space dimensions between modalities; a graph attention network (GAT) and a tensor decomposition (Canonical Polyadic) are used to assign interpretable dynamic weights to the multi-modal input; and the multi-task output end corresponds to downstream tasks such as pose angle prediction, anomaly detection classification, and vital sign regression.
[0065] Preferably, the system is integrated in an edge heterogeneous chip, and the feature extraction and inference are realized through a Mach-Zehnder interference structure or an optoelectronic hybrid ReRAM array with low power consumption and low delay, and the overall system delay is less than 25 ms and the power consumption is not higher than 4 W, which is suitable for indoor intelligent monitoring, disaster rescue and wearable health monitoring scenes.
[0066] Therefore, the multi-modal human body recognition system and method based on the photoelectric super surface and radar fusion have the following beneficial effects:
[0067] (1) The present application reduces the occlusion scene pose angle error by 38%, improves the F1-score in weak light environment by 42%, and improves the ID-F1 in multi-target tracking by 23%. The optical edge and micro-Doppler fusion enhances the target boundary determination. The pose recognition and vital sign detection share the feature encoder, and the migration performance is improved by an average of 8-15% under small sample.
[0068] (2) The present application uses a photoelectric hybrid chip to realize sub-millisecond reasoning, with power consumption of only 4W, which is 3.2 times higher than the recognition efficiency per unit power consumption of the traditional scheme, and is suitable for wearable devices and real-time monitoring.
[0069] (3) The present application dynamically allocates modal weights through cross-modal Transformer-GAT, improves the feature clustering separation degree, and the F1-score of anomaly detection reaches 0.92, which is better than the single-modal scheme.
[0070] The technical solutions of the present application will be further described in detail below through the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0071] Figure 1 It is a schematic diagram of the overall architecture of the system and the bottom acceleration module and multi-modal fusion structure diagram;
[0072] Figure 2 It is a flowchart of a human target edge detection algorithm based on a metasurface cascade + adaptive kernel;
[0073] Figure 3 It is a multi-modal fusion network structure diagram;
[0074] Figure 4 It is a multi-modal human pose recognition algorithm block diagram based on edge image and Doppler radar spectrum fusion;
[0075] Figure 5 It is a human pose recognition effect diagram based on metasurface spectrum and radar spectrum (as an example);
[0076] Figure 6 It is a comparison diagram of multi-modal fusion and single-modal model recognition performance;
[0077] Figure 7 It is a comparison diagram of inference delay and power consumption of different hardware platforms;
[0078] Figure 8 It is a clustering visualization (left image is the feature before fusion; right image is the feature after fusion);
[0079] Figure 9 It is a heat map of different detection mechanisms in different scenes (mainly accuracy and migration score);
[0080] Figure 10The algorithm provided in the present application has a high-resolution detection result graph. DETAILED DESCRIPTION
[0081] The technical solutions of the present application are further described below through the drawings and examples.
[0082] Unless otherwise defined, the technical terms or scientific terms used in the present application shall be understood as the usual meaning understood by a person with ordinary skills in the art to which the present application belongs. The terms "first", "second", and similar words used in the present application do not represent any order, number, or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar words mean that the elements or objects appearing before the words cover the elements or objects listed after the words and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right", and the like are only used to represent relative positional relationships, and when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0083] EMBODIMENT
[0084] As shown in Figures 1-8 The present application provides a multi-modal human recognition system based on the fusion of optoelectronic super surface and radar, which includes:
[0085] A modulation module including a programmable super surface for realizing edge-enhanced optical preprocessing;
[0086] A radar detection module including a continuous wave radar and an FMCW laser radar for obtaining the micro-Doppler spectrum and three-dimensional point cloud of the target;
[0087] A spectral sensing module including a tunable laser source and a two-dimensional material detector for capturing the near-infrared reflection spectrum of the target;
[0088] A multi-modal feature fusion module based on cross-modal Transformer for semantic alignment and fusion of multi-source signals;
[0089] A recognition module including a multi-task learning network for posture recognition, vital sign estimation, and anomaly detection;
[0090] A heterogeneous edge chip for completing optical neural network inference and low-latency recognition.
[0091] Super surface cascaded optoelectronic device design and fabrication, the optical super surface includes a phase-adjustable cascaded structure, the inverse design method is used to design the super surface cascaded optoelectronic device, first is phase stratification: the first layer corrects radially, and the phase distribution is
[0092] φ1(r) = sin(6r)w -2r ;
[0093] Enhanced center-to-edge contrast.
[0094] The 2nd layer (angular petals) phase distribution is
[0095] φ2(θ) = 0.5cos(10θ);
[0096] (only valid in the region r<0.8), resulting in multi-beam petal-like diffraction.
[0097] Then comes the superposition and cropping. Its reflection phase is jointly defined by the following functions:
[0098]
[0099] θ = atan2(y, x).
[0100] With the above parameter values, micro-nano processing is performed on the substrate using E-beam or lithography + etching to fabricate two-layer metasurface structures corresponding to different depths / materials to achieve the required phase distribution.
[0101] The point cloud encoding module adopts the KP-Conv or Sparse Voxel Transformer structure to Tokenize the three-dimensional point cloud data for semantic fusion.
[0102] The micro-Doppler vital sign channel uses a graph wavelet convolution network to model the periodicity of the respiratory and heart rate signals.
[0103] The multi-modal fusion module further introduces tensor decomposition (CP) and graph attention network (GAT) to assign interpretable weights to each modal feature.
[0104] The heterogeneous chip has an optoelectronic hybrid computing unit that uses a Mach-Zehnder interference structure for analog neural operations, achieving sub-10 microsecond edge inference capabilities.
[0105] The system has the ability to maintain recognition robustness under different environmental conditions, including but not limited to: scenes where the target part is partially occluded; scenes with weak light or severe light changes; complex scenes with multiple people overlapping or group dynamic interactions; and by fusing optical edges and micro-Doppler spectra, the accuracy of target boundary determination and dynamic behavior separation is improved.
[0106] The fusion model has multi-task transfer learning capability, and can share feature representation and parameter structure between the following tasks: gesture recognition and skeleton key point detection, non-contact vital sign detection (respiratory rate, heart rate), fall detection and abnormal action recognition, the model adopts a cross-task shared encoder and combines a task-specific loss branch optimization to improve the transfer ability and generalization performance under small samples.
[0107] The application also provides a multi-modal human body recognition method based on photoelectric super surface and radar fusion, comprising the following steps:
[0108] S1, optical high-pass filtering is performed through a phase programmable optical super surface to extract target edge features;
[0109] S2, a continuous wave radar is used to capture a micro-Doppler spectrum, and an FMCW laser radar is used to obtain a three-dimensional point cloud;
[0110] S3, a tunable laser source is used to scan a near-infrared spectrum of a target, and a two-dimensional material detector is used to generate a spectrum vector;
[0111] S4, a cross-modal Transformer structure is used to perform semantic fusion on edge features, point clouds, micro-Doppler spectra and spectral features;
[0112] S5, gesture recognition, vital sign estimation and anomaly detection are performed based on a multi-task learning network;
[0113] S6, low-delay inference is completed through a heterogeneous edge computing chip.
[0114] The edge detection module adopts an optical edge extraction method based on a super surface adaptive kernel function, a Laplace-Gaussian (S-LoG) kernel based on a directional second derivative, a directional parameter θ and a cross term weight w are introduced, and the Gaussian term
[0115]
[0116] Second derivative component
[0117]
[0118] Directional kernel
[0119]
[0120] And make the kernel function zero mean and normalized:
[0121]
[0122] Construct two positive and negative point spread functions PSFK + , K - ;
[0123] K + = max(K, 0), K — = max(-K, 0),
[0124] The following optical filter is used to output the simulated edge intensity value
[0125] E 仿真 (x, y) = I * K + (x, y) - I * K - (x, y);
[0126] The features also include loss function and parameter update
[0127] The loss function is commonly expressed as mean square error (MSE)
[0128]
[0129] When calculating the gradient, σ, θ and w all participate in the convolution operation, so the gradient can be backpropagated:
[0130]
[0131] The random gradient descent (Stochastic Gradient Descent, SGD) or adaptive moment estimation optimization algorithm (Adam) is used for parameter update.
[0132]
[0133] The metasurface phase design parameters can also be jointly optimized.
[0134] The training of the kernel function uses a backpropagation algorithm based on the stochastic gradient descent (SGD) or adaptive moment estimation (Adam) optimizer, which dynamically adjusts the metasurface phase function and kernel parameters, so that the final simulated edge is maximally close to the target edge in spatial frequency domain and structural details.
[0135] The spectral feature extraction module includes: a tunable laser that emits a narrow-band scanning spectrum covering the range of 400-2000 nm; a photodetector array based on two-dimensional materials (such as MoS2, Graphene, etc.); combining the reflectivity of different wavelengths of the target into a spectral vector; the spectral vector is input into a one-dimensional convolutional autoencoder (1D-CNN-AE) or a spectral Transformer for high-dimensional feature embedding and nonlinear compression.
[0136] The embedded spectral feature dimension is 128-512 dimensions, and is aligned with other modalities (radar, point cloud, vital signs) through the Cross-Attention module to construct a unified multi-modal semantic space.
[0137] The vital sign detection module adopts a graph wavelet convolutional network (Graph Wavelet CNN) to model the low-frequency vibration components in the radar signal in the frequency domain to extract the periodic breathing, heartbeat and other life micro-vibration features of the target.
[0138] The multi-modal fusion module adopts the following structure: based on the Transformer Encoder architecture to align the features in time and space dimensions between modalities; using the Graph Attention Network (GAT) and the Canonical Polyadic tensor decomposition to assign interpretable dynamic weights to the multi-modal input; the multi-task output end corresponds to the pose angle prediction, anomaly detection classification, vital sign regression and other downstream tasks.
[0139] The system is integrated in an edge heterogeneous chip, and through the Mach-Zehnder interference structure or the optoelectronic hybrid ReRAM array, low-power and low-delay feature extraction and inference are realized, the overall system delay is less than 25ms, and the power consumption is not higher than 4W, which is suitable for indoor intelligent monitoring, disaster rescue and wearable health monitoring scenes.
[0140] Hardware layer: two layers of phase programmable metasurfaces are used as optical high-pass filters, which are combined with continuous wave radar (CWRadar) and VCSEL-FMCW LiDAR to form a four-stream parallel perception link;
[0141] Algorithm layer: construct adaptive kernel function optical edge extraction + micro-Doppler time-frequency analysis + spectral Transformer embedding + KP-Conv point cloud encoding, and complete semantic fusion through cross-modal Transformer-GAT;
[0142] The mathematical model of KP-Conv+Transformer-GAT includes the KP-Conv (Kernel Point Convolution) mathematical model and the Transformer-GAT cross-modal semantic fusion model formula.
[0143] The KP-Conv (Kernel Point Convolution) mathematical model is as follows:
[0144] KP-Conv is a method of local convolution for each center point p in the point cloud, and its core idea is to use kernel points as sampling centers to weight and sum the neighborhood points. For each center point p, let its neighborhood points bej}, the core point is {km}, and the corresponding weight function is Wm(·), then the output feature is:
[0145]
[0146] Where N(p) is the neighborhood of point p; f(p j ) is the input feature of the neighborhood point; k m is the mth core point (fixed or learnable) ; W m represents the weight function (usually Gaussian function or linear) of the core point. Key points: KP-Conv preserves the Euclidean geometric structure and is suitable for unstructured point clouds.
[0147] The formula of the Transformer-GAT cross-modal semantic fusion model is as follows:
[0148] This module combines two attention mechanisms: Transformer is responsible for long-range modeling and capturing global context; GAT (Graph Attention Network) enhances the structural relationship between points and adapts to cross-modal graph representation. The attention mechanism in Transformer is for input feature sequence X∈R n×d , where n is the number of tokens and d is the feature dimension, defined as: Q=XW Q , K=XW K , V=XW V . The attention output is
[0149]
[0150] The graph attention mechanism of GAT:
[0151] Given a graph G=(V,E), each node feature is hi∈RF, then the influence of adjacent node j∈N(i) on i is:
[0152] e ij =LeakyReLU(a T Wh i |Wh j ) ;
[0153]
[0154] h″ i =σ(∑ k∈N(i) α ij Wh j ) ;
[0155] where W is a learnable linear transformation matrix, a is a learnable attention vector, and σ is an activation function. The Transformer-GAT is embedded in series or parallel to realize cross-modal semantic fusion between point cloud and spectral image.
[0156] Computing layer: Inference acceleration mapping mechanism on MZI-NPU / ReRAM
[0157] (1) MZI-NPU (Mach-Zehnder Interferometer-based Neural Processing Unit)
[0158] MZI is an interference optical computing element that can be used for matrix-vector multiplication (MAC) operations:
[0159] Computing mechanism: encode the weight W in the neural network as the phase control matrix θ of MZI ij ; the input vector x is converted into an optical intensity signal input; the multiplication y=Wx is realized through the MZI grid structure, completing one inference.
[0160] Optical MAC advantages: no need for electrical storage operation, delay < nanosecond level, energy consumption about 1% of CMOS (no heat loss), high parallelism, up to 10 15 MACs (Nature 2025 report). Realize MZI-NPU inference on the optoelectronic-ReRAM hybrid heterogeneous chip, accelerate convolution and attention operations, system total delay < 25 ms, power consumption ≤ 4 W. Main advantages: in-memory computing avoids memory bandwidth bottleneck; supports parallel batch inference; can be embedded in edge devices to meet ≤ 4 W power consumption design.
[0161] (2) ReRAM (Resistive RAM) acceleration mechanism
[0162] ReRAM is used to build a programmable non-volatile weight array, supporting the following operations:
[0163] Weight mapping: the neural network weight W∈Rm×n is mapped to the conductive state Rij in the ReRAM cross array; the input signal is applied to the horizontal row of ReRAM, and the output current is integrated by column current to get the result:
[0164]
[0165] where yj is the output current, V i is the input voltage, x i is the input signal, R ij is the conductive resistance in the cross array; W ij is the neural network weight.
[0166] Take the process of the super surface reflector as an example, the double-layer TiO2 / SiO2 nanocolumn, the unit cycle is 200 nm, and the thickness difference is 30 nm.
[0167] The phase function is
[0168] φ1(r)=sin(6r)e -2r , φ2(θ)=0.5cos(10θ);
[0169] After superposition, one-time forming is performed through photolithography + reactive ion etching.
[0170] Specifically, the carrier frequency of the continuous wave radar is 77 GHz, the transmission power is 10 dBm, the bandwidth is 1 GHz, and the frame rate is 50 fps; the VCSEL-FMCW LiDAR: center wavelength 905 nm, frequency modulation rate 200 MHz / μs, point cloud density 30 kpts / frame.
[0171] Specifically, the tunable laser of the spectral detection channel steps 1 nm @ 400-2000 nm; the MoS2-InGaAs detection array is 512 ch,
[0172] Specifically, the digital-analog hybrid SoC + MZI-NPU array is 64x64, and the INT4 throughput is >2TOPS / W.
[0173] In adaptive kernel function optical edge extraction, the kernel initialization is σ=0.2, θ=42°, and w=0.4.
[0174] Positive / negative kernel normalization: K + =max(K,0), K — =max(-K,0),
[0175] Iteration: Adam (β1=0.9, β2=0.999, lr=10-3), 20 epochs, convergence, PSNR↑4.6dB.
[0176] Specifically, the micro-Doppler and point cloud coding, the radar signal is subjected to STFT (window length 256, hop 64)→128x128 amplitude spectrum; the point cloud is subjected to KP-Conv (r=0.1m, K=16) to extract 256-dim local features; the spectrum is compressed to 128-dim by 1D-CNN-AE.
[0177] Specifically, the cross-modal Transformer-GAT fusion key code is as follows
[0178] Input: {E_edge, S_radar, P_point, L_spec}
[0179] Encoder: 4x{MultiHeadAttn(d_model=512, n_head=8), FFN}
[0180] GAT: K=4 modal nodes, adjacency matrix dynamic learning
[0181] Output: 1024-dim semantic Token→multi-task head
[0182] Task 1: pose regression - MSELoss; Task 2: fall / abnormal classification - FocalLoss(gamma=2); Task 3: breath / heart rate regression - SmoothL1Loss.
[0183] Specifically, in order to adapt to different scenes, different algorithms and mechanisms are adopted, as shown in Table 1.
[0184] Table 1 Different scene processing mechanism
[0185]
[0186] Figure 9 The thermodynamic diagram of different detection mechanisms (such as point cloud, spectrum, and multi-modal) in different scenes is shown, mainly the accuracy and transfer learning ability of the detection algorithm. The detection mechanism mainly considers three cases, one is the single modal recognition performance using 3D graph point cloud data, the second is the single modal recognition performance using spectral images, and the third is the recognition performance after fusing point cloud and spectral images. The detection accuracy and transfer learning ability of three algorithms for five different scenes are mainly analyzed. The five scenes include fall detection - the accuracy of the system in detecting human fall events, abnormal behavior - the detection ability of the system for abnormal behavior (such as violent twisting and paralysis), occlusion scene - the robustness of target recognition in multi-person / occlusion, low light environment - the ability to retain structural information in low light images, and multi-person scene - the detection ability of multiple people appearing in the field of view at the same time. In all tasks, the multi-modal fusion performance is better than that of single modal input, with the largest improvement of +0.25; the spectral image performance decreases in the "low light environment" and "occlusion" scenes, indicating that it is strongly dependent on structural information; the point cloud performs well in the multi-person scene, but is prone to misjudgment in occlusion and dynamic behavior; the KP-Conv+Transformer-GAT architecture proposed in the present application shows strong robustness in multi-task transfer generalization. Compared with the prior art, the present application has obvious advantages in five aspects, as shown in Table 2. Figure 10 The detection effect of the present application is shown, which is extremely high resolution and very clear texture information.
[0187] Table 2 Comparison of the present application with prior art
[0188]
[0189] Finally, it should be noted that the above examples are intended to illustrate the technical solutions of the present application but not to limit the same. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still make modifications or equivalent replacements to the technical solutions of the present application, and these modifications or equivalent replacements should not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.
Claims
1. A multi-modal human recognition system based on plasmonic metasurface and radar fusion, characterized in that, Comprising: Modulation module: containing phase programmable optical metasurface, for optical high-pass filtering of the target to achieve edge enhancement; Radar detection module: including continuous wave radar and frequency modulated continuous wave FMCW laser radar, respectively for acquiring micro-Doppler spectrum and three-dimensional point cloud data of the target; Spectral sensing module: containing tunable laser source and two-dimensional material detector, for capturing near-infrared reflectance spectrum of the target; Multi-modal feature fusion module: adopting cross-modal Transformer structure, for semantic-level fusion of optical edge, point cloud, micro-Doppler spectrum and spectral features; Identification module: based on multi-task learning network, outputting posture recognition, vital sign estimation and anomaly detection results; Heterogeneous edge computing chip: integrating optoelectronic hybrid computing unit, for low-delay feature extraction and inference.
2. The multi-modal human recognition system based on the fusion of plasmonic metasurfaces and radar according to claim 1, characterized in that: The optical metasurface is a double-layer cascaded structure, and its reflection phase distribution satisfies: wherein Θ = atan2(y, x).
3. The multi-modal human recognition system based on the fusion of plasmonic metasurfaces and radar according to claim 1, characterized in that: The point cloud data is encoded by kernel point convolution KP-Conv or sparse voxel Transformer to generate Tokenized features for semantic fusion.
4. The multi-modal human recognition system based on the fusion of plasmonic metasurfaces and radar according to claim 1, characterized in that: The micro-Doppler spectrum adopts a graph wavelet convolution network Graph Wavelet CNN to model the periodicity of respiratory and heart rate signals.
5. The multi-modal human recognition system based on the fusion of plasmonic metasurfaces and radar according to claim 1, characterized in that: The multi-modal feature fusion module further introduces a graph attention network GAT and tensor decomposition CP to assign dynamic interpretable weights to each modal feature.
6. The multi-modal human recognition system based on the fusion of metasurface and radar according to claim 1, characterized in that: The heterogeneous edge chip adopts a Mach-Zehnder interferometer MZI to realize optical neural computing, or completes weight mapping acceleration through a ReRAM cross array.
7. A multi-modal human body recognition method based on the fusion of plasmonic metasurfaces and radar, applied to the system of any one of claims 1-6, characterized in that: Comprising the following steps: S1, optical high-pass filtering through a phase programmable optical metasurface to extract edge features of the target; S2, capturing micro-Doppler spectrum using continuous wave radar and acquiring three-dimensional point cloud using FMCW laser radar; S3, scanning near-infrared spectrum of the target through a tunable laser source and generating a spectral vector from a two-dimensional material detector; S4, adopting a cross-modal Transformer structure for semantic fusion of edge features, point cloud, micro-Doppler spectrum and spectral features; S5, based on a multi-task learning network, performing posture recognition, vital sign estimation and anomaly detection; S6, completing low-delay inference through a heterogeneous edge computing chip.
8. The multi-modal human recognition method based on the fusion of plasmonic metasurfaces and radar according to claim 7, characterized in that: In step S1, the edge extraction adopts an adaptive kernel function optical edge extraction method, and the directional convolution kernel is defined as: wherein is a Gaussian term; is the second derivative component.
9. The multi-modal human recognition method based on the fusion of metasurface and radar according to claim 8, characterized in that: The parameters σ, θ, w of the adaptive kernel function are optimized through a back propagation algorithm, and the loss function is The optimizer adopts stochastic gradient descent SGD or adaptive moment estimation Adam.
10. The multi-modal human recognition method based on the fusion of plasmonic metasurfaces and radar according to claim 7, characterized in that: In step S4, the spectral vector is compressed to high-dimensional embedding features with a dimension of 128-512 dimensions through a one-dimensional convolutional autoencoder 1D-CNN-AE or a spectral Transformer.
Citation Information
Patent Citations
Electronic equipment
CN109379454A
Human body posture recognition method based on FMCW radar signal
CN113313040A
Millimeter wave radar behavior identification method based on multi-task cross-modal attention
CN120375473A
Cited By
Human body falling risk early warning method and system, equipment and storage medium
CN121354282A