Multi-object vibration fusion voice perception and reconstruction method based on millimeter wave radar

By employing a multi-object vibration fusion speech perception method based on millimeter-wave radar and utilizing frequency response functions and generative adversarial networks for speech reconstruction, the problem of information loss in speech reconstruction under complex environments is solved, achieving high-fidelity and robust speech recovery results.

CN121438801APending Publication Date: 2026-01-30SOUTHEAST UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511819478.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing millimeter-wave speech perception systems ignore the object-specific frequency response differences in complex environments, leading to information loss and spectral distortion during speech reconstruction, especially resulting in low speech recovery quality in multi-object scenarios.

Method used

This paper proposes a multi-object vibration fusion speech perception method based on millimeter-wave radar. It uses a trained feature extraction network to obtain the frequency response function of each object, and performs dynamic alignment and feature complementarity through cross-channel frequency attention matching and generative adversarial network to achieve spectrogram fusion. Finally, it uses a speech reconstruction model to generate a high-quality speech spectrum.

Benefits of technology

It significantly improves the integrity and clarity of speech reconstruction, can adapt to different wall thicknesses, materials and multi-object resonance scenarios, improves the robustness and noise resistance of speech recovery, and is suitable for speech perception tasks in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121438801A_ABST
    Figure CN121438801A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-object vibration fusion voice perception and reconstruction method based on a millimeter-wave radar. The method comprises the following steps: S1, carrying out range frequency conversion on millimeter-wave radar receiving signals to obtain a plurality of discrete object vibration signals; s2, obtaining a frequency response function of each object by using the trained feature extraction network based on object vibration signals; s3, fusing the object vibration signal of each object and the corresponding frequency response function along a frequency axis to obtain a first spectrogram of each object; s4, performing cross-channel frequency attention matching on the first spectrograms of the objects to realize dynamic alignment and feature complementation so as to obtain a fused spectrogram; and S5, based on the fused spectrogram, obtaining a speech spectrum by using the trained speech reconstruction model. Compared with the prior art, the method has the advantages that the inherent frequency response difference of an object is utilized, the physical compensation of acoustic propagation characteristics is realized, and the voice reconstruction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, and in particular to a method for multi-object vibration fusion speech perception and reconstruction based on millimeter-wave radar. Background Technology

[0002] With the rapid development of millimeter-wave sensing and signal processing technologies, non-contact voice information extraction has gradually become a reality. Millimeter-wave signals possess excellent spatial resolution and penetration capabilities in the high-frequency band, enabling them to capture the subtle vibrations generated on the surface of objects during voice propagation, thus reconstructing acoustic information without relying on microphones. In recent years, this technology has been widely applied in scenarios such as security monitoring, privacy protection research, and microphone-free voice sensing systems.

[0003] Existing millimeter-wave speech sensing systems, such as MILLIEAR and mmEcho, utilize the phase-change characteristics of frequency-modulated continuous wave (FMCW) radar to penetrate obstacles like walls and recover speech signals from target surfaces. These systems capture signals across a speech frequency range (approximately 100Hz–4kHz) by detecting sound-driven nanoscale vibrations, without requiring speaker constraints or vocabulary limitations, thus validating the feasibility of millimeter waves in through-wall acoustic eavesdropping. Furthermore, some studies, such as VibSpeech and piezoelectric-based side-channel attack methods, have also indirectly recovered speech information through acoustic vibration paths. These studies collectively demonstrate the immense potential of millimeter waves and related physical side-channels in the field of covert speech sensing.

[0004] For example, Chinese patent CN117935765A provides a speech synthesis method, system, and terminal based on millimeter-wave radar. It removes static signals from the received signals from the millimeter-wave radar, applies the Range-FFT algorithm to determine target units and calculate the phase signal corresponding to each target unit, then identifies the phase signal related to sound vibrations generated by the emitting speech object in each phase signal, performs signal enhancement and preprocessing, and generates the corresponding phase Mel spectrum. Based on a trained spectrum reconstruction model, it obtains the corresponding reconstructed phase Mel spectrum from each phase Mel spectrum, and then uses a vocoder to obtain the speech waveform corresponding to each emitting speech object. This implements a radar microphone using millimeter-wave radar as a sound sensor, fully utilizing the characteristics of millimeter-wave radar to recover higher quality, more understandable speech audio, achieving automated sound source location detection and multi-source separation.

[0005] However, existing technologies, including the methods described above, generally suffer from a critical flaw in their assumptions—treating walls, tabletops, glass, or objects in the environment as homogeneous passive reflectors. This simplification may be feasible in ideal scenarios, but it leads to significant performance degradation in real-world, complex environments. In fact, objects of different materials, shapes, and structures possess unique frequency response characteristics; that is, they respond differently to different frequency components in a speech signal. Some objects may produce strong resonances in the low-to-mid-frequency range, while others exhibit strong vibrational sensitivity in the high-frequency range. These differences cause each object to act as a selective filter in the speech propagation path, amplifying or attenuating the speech signal to varying degrees.

[0006] Therefore, ignoring these "object-specific frequency response differences" will lead to information loss and spectral distortion during speech reconstruction. This is especially true in scenarios with multiple objects, where the frequency domain information carried by each object is often complementary: one object may have a better response in the low-frequency range, while another object provides clearer features in the high-frequency range. If these differentiated frequency response features can be selectively fused, the completeness and clarity of speech reconstruction can be significantly improved. Summary of the Invention

[0007] The purpose of this invention is to provide a multi-object vibration fusion speech perception and reconstruction method based on millimeter-wave radar, which aims to solve the problems of weak speech signals, target vibration feature aliasing, and low speech recovery quality in existing millimeter-wave speech perception systems in wall penetration and complex environments.

[0008] The objective of this invention can be achieved through the following technical solutions: A method for multi-object vibration fusion speech perception and reconstruction based on millimeter-wave radar includes: Step S1: Perform range-frequency transformation on the millimeter-wave radar received signal to obtain multiple discrete object vibration signals; Step S2: Based on the vibration signal of the object, obtain the frequency response function of each object using the trained feature extraction network; Step S3: After fusing the vibration signals and corresponding frequency response functions of each object along the frequency axis, the first spectrum diagram of each object is obtained. Step S4: The first spectrograms of each object are dynamically aligned and feature complementarity is achieved through cross-channel frequency attention matching to obtain a fused spectrogram; Step S5: Based on the fused spectrogram, obtain the speech spectrum using the trained speech reconstruction model.

[0009] Step S1 includes: Step S1-1: Acquire the millimeter-wave radar received signal; Step S1-2: Perform range-frequency transformation on the millimeter-wave radar received signal to achieve spatial segmentation, dividing the monitoring area into multiple discrete range units; Step S1-3: By analyzing the phase change characteristics of each distance unit, the vibration signal of each object is obtained.

[0010] Step S2 includes: Step S2-1: Enhance and reduce noise in the object vibration signal; Step S2-2: Based on the object vibration signal after signal enhancement and noise reduction, the frequency response function of each object is obtained using the trained feature extraction network; In step S3, the object vibration signal after signal enhancement and noise reduction and the corresponding frequency response function are fused along the frequency axis to obtain the first spectrum of each object.

[0011] The feature extraction network estimates the frequency response function through an optimization process, with the optimization objective being: in: The vibration signal of the i-th object after signal enhancement and noise reduction. Let i be the frequency response function of the i-th object. The hypothetical sound wave signal received by the object. Let be the gradient of the frequency response function of the i-th object.

[0012] Step S4 includes: Step S4-1: Select one of the objects as the main object; Step S4-2: Calculate the correlation between each frequency band of the main object and each frequency band of other objects in the first spectrum diagram; Step S4-3: Superimpose the obtained correlation onto the first spectrum map of the main object to obtain the fused spectrum map: in: For fusion spectrum, This is the first frequency band in the fused spectrum diagram. For the j-th frequency band of the fused spectrogram, For the Mth frequency band in the fused spectrogram, Is D the j-th frequency band of the first spectrum of the main object, and is D the dimension of the first spectrum of the main object? For the first l The k-th frequency band of the first spectrogram of another object. For correlation, M is the number of frequency bands, and B is the number of other objects.

[0013] The speech reconstruction model is a generative adversarial network.

[0014] The method further includes: The millimeter-wave radar is controlled to send frequency sweep signals to the target area, and multiple discrete object vibration signals are obtained by performing distance-frequency transformation based on the received signals of the millimeter-wave radar, which serve as the preliminary positioning information of the object.

[0015] The fusion method in step S3 is splicing.

[0016] A multi-object vibration fusion speech perception and reconstruction device based on millimeter-wave radar includes a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the method described above.

[0017] A storage medium having a program stored thereon, which, when executed, implements the method described above.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. By introducing object frequency response modeling, the resonance characteristics of different materials and structures can be distinguished. Through the physical property-driven modeling method, the spectral distortion of speech during propagation can be effectively compensated.

[0019] 2. The multi-object vibration fusion mechanism can utilize cross-attention modules to achieve feature association and weighted fusion in the frequency dimension, which is significantly better than the traditional simple splicing method, improving the detail fidelity and noise resistance of speech reconstruction.

[0020] 3. Generative adversarial optimization framework enhances the authenticity and consistency of generated spectra through adversarial learning, significantly improving both subjective intelligibility and objective spectral quality of reconstructed speech.

[0021] 4. It exhibits strong robustness and adaptability, demonstrating stability under varying wall thicknesses, materials, and multi-object resonance scenarios. It can adapt to the multipath, attenuation, and noise interference issues of speech signals in complex environments.

[0022] 5. It has good scalability, and the system framework can be extended to other non-contact acoustic sensing tasks, such as structural health monitoring, privacy-preserving voice acquisition, and intelligent environmental sensing. Attached Figure Description

[0023] Figure 1 This is a system scene diagram of the present invention; Figure 2 System flowchart; Figure 3 Schematic diagram for frequency response modeling; Figure 4 This is a schematic diagram illustrating the effect of speech enhancement. Figure 5 A network structure diagram for the feature fusion of conditional generative adversarial networks and cross-attention mechanisms; Figure 6 A comparison chart of reconstruction metrics for single-object and multi-object fusion signals under different distance conditions; Figure 7 A comparison of the reconstructed spectra of single-object and multi-object fusion signals under different distance conditions; Figure 8 This is a schematic diagram of the main steps of the method of the present invention. Detailed Implementation

[0024] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0025] Existing millimeter-wave-based acoustic eavesdropping techniques primarily focus on optimizing perception algorithms while neglecting the physical characteristics of the perceived objects. Differences in frequency responses between different objects are often treated as noise interference, failing to effectively utilize their potential information gain. Furthermore, traditional methods typically rely solely on the vibration signals of a single object for speech reconstruction, lacking a joint modeling and fusion mechanism for the frequency features of multiple objects, resulting in severe distortion and poor semantic intelligibility in speech recovery. On the other hand, existing signal reconstruction methods struggle to capture effective speech spectral features under low signal-to-noise ratio conditions and lack adaptive recovery capabilities for complex frequency responses. Therefore, a millimeter-wave perception method capable of fusing the frequency response characteristics of multiple objects and incorporating deep generative models to achieve high-fidelity speech reconstruction is needed to significantly improve the clarity and intelligibility of speech recovery in scenarios involving walls and multiple targets.

[0026] This application proposes a multi-object vibration fusion speech perception and reconstruction method based on millimeter-wave radar, constructing a high-fidelity speech perception system for complex scenarios. The method collects weak vibration signals generated by various objects (such as desktops, glass, and metal) in the environment under speech excitation using millimeter-wave radar. First, time-frequency segmentation and phase change detection modules are used to locate effective speech-related vibration sources. Then, based on a linear time-invariant (LTI) system model, the frequency response characteristics of each object are estimated, and their resonance and attenuation patterns in different frequency bands are extracted, thereby obtaining the physical feature representation of the speech propagation path. In the feature fusion stage, a cross-attention mechanism is introduced to perform feature alignment and weighted fusion of the multi-object vibration signals in the frequency dimension, strengthening the complementary information expression between different objects. Finally, a conditional generative adversarial network is used to reconstruct the speech spectrum of the fused vibration features, generating a high-quality speech spectrogram and achieving end-to-end speech recovery. This method can significantly improve the integrity and intelligibility of speech reconstruction in complex environments such as low signal-to-noise ratio, wall penetration, and multi-object reflection, providing an efficient and reliable technical solution for millimeter-wave acoustic perception and passive speech reconstruction.

[0027] like Figure 2 and Figure 8 As shown, it includes: Step S1: Perform range-frequency transformation on the millimeter-wave radar received signal to obtain multiple discrete object vibration signals; like Figure 1 The diagram shown is a scene illustration of this application. The millimeter-wave radar is located next to the sound source, and there are multiple objects near the sound source.

[0028] Specifically, step S1 includes: Step S1-1: Acquire the millimeter-wave radar received signal; Step S1-2: Perform range-frequency transformation on the millimeter-wave radar received signal to achieve spatial segmentation, dividing the monitoring area into multiple discrete range units; Step S1-3: By analyzing the phase change characteristics of each distance unit, the vibration signal of each object is obtained. Specifically, by analyzing the phase change characteristics of each distance unit, low-frequency (less than 1kHz) phase disturbances are found only in the distance unit corresponding to the target object, which reflects the surface micro-vibration characteristics caused by speech.

[0029] In some embodiments, after locating the target distance unit, the method further includes: The system controls a millimeter-wave radar to send frequency-sweeping signals to the target area, and performs range-frequency transformation on the received signals to obtain multiple discrete object vibration signals, which serve as preliminary object location information. This assists in object location during subsequent attack phases.

[0030] Step S2: As Figure 3As shown, based on the vibration signals of objects, the frequency response functions of each object are obtained using a trained feature extraction network, including: Step S2-1: Enhance and reduce noise in the object vibration signal; like Figure 4 As shown, the system further performs geometric enhancement processing on the phase signal to improve vibration perceptibility under low signal-to-noise ratio conditions. Specifically, by geometrically rotating and normalizing the phase trajectory on the IQ plane, its phase distribution is adjusted to a fixed angular range, thereby amplifying the angular changes corresponding to speech vibrations and improving signal stability and intelligibility. Through these steps, the system can accurately extract object vibration signals caused by through-wall speech in complex environments, providing reliable input for subsequent frequency domain modeling and signal reconstruction.

[0031] Step S2-2: Based on the object vibration signal after signal enhancement and noise reduction, the frequency response function of each object is obtained using the trained feature extraction network; Subsequently, the system characterizes the frequency response characteristics of different environmental objects to speech waves, thereby achieving distortion compensation and quality improvement during speech reconstruction. Because everyday objects exhibit significant differences in resonance and attenuation characteristics across different frequency bands, their surface vibrations can cause spectral distortion in the original speech signal. Therefore, by estimating and correcting the frequency response functions of each object, the true spectral structure of the speech can be effectively recovered.

[0032] The system models the interaction between sound waves and the object surface as a linear time-invariant (LTI) system. For the i-th object, its vibration signal y i The relationship between x(t) and the input speech x(t) can be expressed as: in: To represent the system's impulse response function, after transforming this relationship to the frequency domain, we obtain: Among them, H i (f) represents the frequency response function of the object. The feature extraction network estimates the frequency response function through an optimization process, with the optimization objective being: in: The vibration signal of the i-th object after signal enhancement and noise reduction. Let i be the frequency response function of the i-th object. The hypothetical sound wave signal received by the object. Let be the gradient of the frequency response function of the i-th object.

[0033] Where λ=0.1 is used to control H i(f) Smoothness. The optimization uses gradient descent to solve. Considering that the frequency resolution of the Fast Fourier Transform (FFT) is usually higher than the requirement of the subsequent Short Time Fourier Transform (STFT), the system performs a certain optimization after estimation. Downsampling is performed to match the resolution of the STFT.

[0034] Step S3: The vibration signals of each object and the corresponding frequency response functions are fused along the frequency axis to obtain the first spectrum of each object. In this embodiment, the vibration signals of the objects after signal enhancement and noise reduction and the corresponding frequency response functions are fused along the frequency axis to obtain the first spectrum of each object. In this embodiment, the fusion method is splicing.

[0035] Step S4: As Figure 5 As shown, the first spectrograms of each object are dynamically aligned and feature-complementary through cross-channel frequency attention matching to obtain a fused spectrogram, including: Step S4-1: Select one of the objects as the main object; Step S4-2: Calculate the correlation between each frequency band of the main object and each frequency band of other objects in the first spectrum diagram; Step S4-3: Superimpose the obtained correlation onto the first spectrum map of the main object to obtain the fused spectrum map: in: For fusion spectrum, This is the first frequency band in the fused spectrum diagram. For the j-th frequency band of the fused spectrogram, For the Mth frequency band in the fused spectrogram, Is D the j-th frequency band of the first spectrum of the main object, and is D the dimension of the first spectrum of the main object? For the first l The k-th frequency band of the first spectrogram of another object. For correlation, M is the number of frequency bands, and B is the number of other objects.

[0036] Step S5: Based on the fused spectrogram, obtain the speech spectrum using the trained speech reconstruction model, which is a generative adversarial network.

[0037] Finally, the system designs a fusion structure based on conditional generative adversarial networks and cross-attention mechanisms. The system jointly encodes vibration signals from different object surfaces with their corresponding frequency response features, performing fine-grained fusion along the frequency axis to fully utilize the complementary frequency domain information between the two objects. This structure achieves dynamic alignment and feature complementarity between the two sets of input signals through cross-channel frequency attention matching. The generator is responsible for generating an enhanced speech spectrogram from the fused features, while the discriminator is used to determine the authenticity of the generated result, thereby driving the system to learn higher-fidelity speech reconstruction features through adversarial training. Through the fusion and transformation of this module, the system achieves effective information integration of multi-object vibration signals, significantly improving the clarity and robustness of speech reconstruction in complex scenarios such as through-walls.

[0038] When using the millimeter-wave-based multi-object through-wall speech reconstruction system of this invention, the model pre-training process must first be completed in a known environment. The system uses millimeter-wave radar to collect micro-vibration signals generated by different common objects (such as metal plates, glass windows, and plastic bottles) under speech stimulation, and combines this with measured reference speech data to establish a training dataset. During the training phase, the system sequentially executes vibration capture, frequency response estimation, and multi-object vibration fusion modules. Through an adversarial generative model, it jointly learns the frequency response characteristics and fusion strategies of different objects, thereby obtaining a generalizable frequency response estimation model and a generative reconstruction network. During training, the system automatically models the resonance characteristics and bandwidth response patterns of various objects, independent of specific wall or environmental layouts, and adaptable to various materials and distance conditions.

[0039] During the deployment phase, specifically in the "through-wall speech reconstruction" scenario, the system does not require model retraining. The millimeter-wave radar first performs spatial segmentation and low-frequency phase change detection on the target area to locate potential speech-induced vibration sources. Subsequently, the system models the detected object vibration signals in the frequency domain and recovers their frequency characteristics using the frequency response compensation model obtained during training. Finally, a cross-attention-based fusion and generation module jointly reconstructs the vibration signals from multiple objects, outputting a clear speech spectrum and generating audible speech. This system can stably achieve high-fidelity speech recovery under complex conditions such as non-contact and through-wall scenarios, exhibiting low invasiveness and high robustness. It is suitable for applications such as security monitoring, voice-assisted perception, and passive sound source identification in complex environments, demonstrating good practicality and potential for widespread adoption.

[0040] like Figure 6 and Figure 7As shown, this invention systematically evaluates the performance of millimeter-wave-based multi-object speech reconstruction under various realistic wall-penetrating and complex reflection environments. The experimental system is implemented based on the PyTorch framework. In the experimental phase, four common objects (tissue box, paper bag, water cup, and laptop) were selected as known training objects, and their surface vibration data were collected by playing speech signals to complete model training. In the testing phase, an earphone box and a potato chip bag were selected as unknown test objects to verify the generalization ability of this invention in cross-object scenarios.

[0041] The reconstruction performance of single-object and multi-object fused signals was compared under different distance conditions. The results show that at a distance of 200cm, the speech reconstructed by multi-object fusion improves the intelligibility index (STOI) by 0.048 compared to the single-object (Object A) reconstruction, and the speech quality index (PESQ) also improves by 0.01–0.03 under most distance conditions. This indicates that the method of the present invention can effectively fuse complementary frequency information between different objects, significantly improving the quality and stability of speech reconstruction.

[0042] Furthermore, the speech spectrograms of different objects were visualized and compared. The results show that the speech spectrogram generated by multi-object fusion outperforms single-object signals in both detail recovery and structural fidelity: in some frequency bands, the fusion result successfully reconstructs the high-frequency detail features missing in single objects; while in other regions, the fusion model inherits the accurate speech texture from specific objects, thus obtaining reconstruction results closer to real speech. Overall, this invention, by introducing a vibration fusion mechanism based on cross-attention, fully exploits the frequency response differences among multiple objects, achieving high-fidelity and robust speech reconstruction performance, and verifying the effectiveness and practicality of the proposed system.

[0043] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar, characterized in that, The method comprises: Step S1: distance-frequency transformation is performed on the millimeter wave radar receiving signal to obtain a plurality of discrete object vibration signals; Step S2: based on the object vibration signal, a trained feature extraction network is used to obtain the frequency response function of each object; Step S3: the object vibration signal and the corresponding frequency response function of each object are fused along the frequency axis to obtain a first frequency spectrum of each object; Step S4: the first frequency spectrum of each object is matched through cross-channel frequency attention to realize dynamic alignment and feature complementation, and a fused frequency spectrum is obtained; Step S5: based on the fused frequency spectrum, a trained speech reconstruction model is used to obtain a speech spectrum.

2. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The step S1 comprises: Step S1-1: obtaining the millimeter wave radar receiving signal; Step S1-2: distance-frequency transformation is performed on the millimeter wave radar receiving signal to realize spatial segmentation, and the monitoring area is divided into a plurality of discrete distance units; Step S1-3: by analyzing the phase change characteristics of each distance unit, the object vibration signal of each object is obtained.

3. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The step S2 comprises: Step S2-1: signal enhancement and noise reduction are performed on the object vibration signal; Step S2-2: based on the signal-enhanced and noise-reduced object vibration signal, a trained feature extraction network is used to obtain the frequency response function of each object; In the step S3, the signal-enhanced and noise-reduced object vibration signal and the corresponding frequency response function are fused along the frequency axis to obtain the first frequency spectrum of each object.

4. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 3, characterized in that, The feature extraction network estimates the frequency response function through an optimization process, and the optimization target is: wherein: is the object vibration signal of the i-th object after signal enhancement and noise reduction, is the frequency response function of the i-th object, is the assumed sound wave signal received by the object, is the gradient of the frequency response function of the i-th object.

5. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The step S4 comprises: Step S4-1: selecting one of the objects as the main object; Step S4-2: the correlation of each frequency band in the first frequency spectrum of the main object with each frequency band of the other objects is calculated respectively; Step S4-3: the obtained correlation is superimposed into the first frequency spectrum of the main object to obtain a fused frequency spectrum: wherein: is a fusion spectrum map, is a first frequency band of the fusion spectrum map, is a jth frequency band of the fusion spectrum map, is an Mth frequency band of the fusion spectrum map, is a jth frequency band of a first spectrum map of a primary object, D is a dimension of the first spectrum map of the primary object, is a kth frequency band of a first spectrum map of a jth other object, l is a correlation, M is a number of frequency bands, and B is a number of other objects.​ 6. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The speech reconstruction model is a generative adversarial network.

7. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The method further comprises: The millimeter wave radar sends a frequency sweeping signal to the target area, and based on the millimeter wave radar receiving signal, distance-frequency transformation is performed to obtain a plurality of discrete object vibration signals as preliminary positioning information of the object.

8. The multi-object vibration fusion speech perception and reconstruction method based on millimeter wave radar according to claim 1, characterized in that, The fusion method in step S3 is splicing. 9.A multi-object vibration fusion speech perception and reconstruction device based on millimeter wave radar, comprising a memory, a processor, and a program stored in the memory, wherein, The processor executes the program to realize the method of any one of claims 1-8.

10. A storage medium having stored thereon a program, characterized by The program is executed to realize the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Speech synthesis method, system and terminal based on millimeter wave radar

    CN117935765A

  • Millimeter wave interception method and system

    CN113192518A

  • Voice signal reconstruction method and device, electronic equipment and storage medium

    CN115376549A

  • Hybrid speech processing method, electronic equipment and computer readable medium

    CN120236599A

  • Vibrational devices as sound sensors

    US20180336274A1