Long-distance non-visual indoor sound source localization method and system based on laser Doppler

Through the Mach-Zehnder heterodyne interference structure and the improved PHAT weighted generalized cross-correlation delay estimation algorithm, the problems of single listening medium and inaccurate positioning in laser Doppler frequency shift interferometry technology are solved, achieving higher-precision sound source positioning and wider listening applications.

CN116449293BActive Publication Date: 2025-09-19CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210011753.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-06
Publication Date
2025-09-19
Estimated Expiration
2042-01-06

AI Technical Summary

Technical Problem

The existing laser Doppler frequency shift interferometry technology can only listen to a single type of forced vibration medium, cannot accurately locate the target sound source signal, and has a limited scope of application.

Method used

The Mach-Zehnder heterodyne interferometer structure and the improved PHAT weighted generalized cross-correlation delay estimation algorithm are adopted. The vibration signals are synchronously acquired through multiple laser listening devices, the time difference and distance difference are calculated, and the three-dimensional coordinate positioning of the sound source point is realized in combination with the TDOA algorithm.

Benefits of technology

It improves the listening sensitivity and accuracy, expands the types of listening media, simplifies operation, enhances concealment, and achieves higher-precision sound source positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116449293B_ABST
    Figure CN116449293B_ABST
Patent Text Reader

Abstract

The present invention discloses a long-distance, non-visual indoor sound source localization method and system based on laser Doppler. The method comprises the following steps: Step 1: M laser listening devices are respectively set at M outdoor listening monitoring points, and the laser listening devices at the M listening monitoring points are used to synchronize the time to acquire vibration signals from M vibrating media located indoors; Step 2: The vibration signals of the vibrating media are converted into digital voice signals; Step 3: Using an improved PHAT weighted generalized cross-correlation time delay estimation algorithm, the time difference between the reference vibrating medium receiving the sound source signal and the time difference between the reference vibrating medium receiving the sound source signal and each other vibrating medium receiving the sound source signal is calculated, and then converted into the distance difference between the reference vibrating medium and the sound source point and each other vibrating medium and the sound source point; Step 4: Based on the distance difference between each vibrating medium and the sound source, the coordinate position of the sound source point is calculated. The present invention has higher accuracy for long-distance, non-visual indoor sound source localization and has a wide range of applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of laser communication technology, and in particular to a long-distance non-visual indoor sound source positioning method and system based on laser Doppler. Background Art

[0002] Theoretical research on laser audio analysis technology has yielded promising results, and corresponding products have been developed in the field of laser voice detection, enabling long-distance, non-visual target voice detection. Currently, domestic research methods for this technology include the optical lever method (reflective spot movement method), semiconductor laser self-mixing interferometry, and laser Doppler shift interferometry. Compared to the other two methods, laser Doppler shift interferometry offers advantages such as higher detection accuracy, longer detection range, and stronger anti-interference capabilities, making it a more mainstream method. However, currently used laser Doppler shift interferometry technology still has drawbacks. First, the type of forced vibration medium used for detection is limited, primarily window glass. Second, it cannot accurately locate the target sound source signal. Summary of the Invention

[0003] The present invention provides a long-distance non-visual indoor sound source positioning method and system based on laser Doppler, which has higher positioning accuracy and wide application range.

[0004] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0005] A long-distance non-visual indoor sound source localization method based on laser Doppler, comprising:

[0006] Step 1: M laser listening devices are respectively set at M outdoor listening monitoring points, and the laser listening devices at the M listening monitoring points are used to synchronize time to obtain vibration signals of M vibrating media located indoors; M ≥ 4;

[0007] Step 2, converting the vibration signal of the vibration medium into a voice digital signal;

[0008] Step 3: Using the improved PHAT weighted generalized cross-correlation delay estimation algorithm, the time difference between the reference vibration medium receiving the sound source signal and the time difference between the other vibration media receiving the sound source signal is calculated, and then converted into the distance difference between the reference vibration medium and the sound source point and the other vibration media and the sound source point;

[0009] Step 4: Calculate the coordinate position of the sound source point based on the distance difference between each vibration medium and the sound source point.

[0010] Furthermore, the laser listening device includes a laser transmitter, a laser beam splitter, a laser beam expander, a laser reflector, a laser polarization beam splitter prism, a laser beam combiner, a laser focusing mirror, two laser detectors, a laser demodulation photoelectric converter, a digital signal acquisition card, an acousto-optic modulator, a λ / 2 wave plate, and a transmitter and receiver;

[0011] The process of the laser listening device acquiring the vibration signal of the vibrating medium is as follows: the laser transmitter outputs laser light, which is split into two by a laser beam splitter, one of which is modulated by an acousto-optic modulator (AOM) to form a reference light, and the other is passed through a laser polarization splitter prism and a transmitter-receiver to acquire the signal light formed by the reflection of the vibrating medium; the signal light and the reference light are combined and mixed by a laser beam combiner, and then received by a laser detector, and the optical signal is converted into an electrical signal by a laser demodulation photoelectric converter, and finally the vibration signal is obtained by sampling through a digital signal acquisition card.

[0012] Furthermore, the vibration signal acquired by the laser listening device is expressed as:

[0013]

[0014] Where L is the forced vibration displacement of the vibrating medium caused by the excitation of the sound wave from the sound source, λ is the laser wavelength, and I1 and I2 are the electrical signals converted from the signal light signal received by the two laser detectors in the laser listening device and the reference light signal through the laser demodulation photoelectric converter.

[0015] Furthermore, the specific calculation process of step 3 is:

[0016] Step 3.1: for each voice digital signal I i (n) Perform discrete Fourier transform to obtain I i (k):

[0017] I i (k)=DFT{I i (n)}

[0018] Among them, i=1, 2, 3, 4 correspond to the four vibration media A, B, C, and D respectively;

[0019] Step 3.2: Compare the discrete Fourier signal I1(k) corresponding to the reference vibration medium A with the discrete Fourier signals I1(k) corresponding to the other vibration media B, C, and D. j Multiply the conjugate complex numbers of (k) by two or more to get the complex product I 1j (k):

[0020]

[0021] Where j = 2, 3, 4;

[0022] Step 3.3, multiply the complex product I 1j (k) Divide its modulus value to obtain the intermediate parameter

[0023]

[0024] Step 3.4, retain the speech interest frequency band ROI∈[f roi_min ,f roi_max ], set other frequency band values ​​to zero:

[0025]

[0026] Among them, f roi_min ,f roi_max are the lower and upper limits of the frequency band ROI of the speech interest domain, hz(k) is the actual frequency corresponding to frequency point k, hz(k) = k*(fs / N), fs is the sampling rate of the speech digital signal, and N is the length of the speech digital signal, that is, the number of sampling points;

[0027] Step 3.5, Perform inverse discrete Fourier transform to obtain intermediate parameters

[0028]

[0029] Step 3.6, calculate the distance difference between the reference vibration medium A and the sound source point S and the vibration medium B, C, D and the sound source point S:

[0030]

[0031]

[0032]

[0033] Where R a,b 、R a,c 、R a,d are the distance differences between the reference vibration medium A and the sound source point S and the vibration medium B, C, and D and the sound source point; v is the speed of sound, fs is the sampling rate; ind max {·} indicates the position where the maximum value is taken.

[0034] Furthermore, step 4 uses the TODA positioning algorithm to calculate the coordinate position of the sound source point:

[0035]

[0036] Where, (s x ,s y ,s z ) is the three-dimensional space coordinate of the sound source point S, (ax ,a y ,a z )、(b x ,b y ,b z )、(c x ,c y ,c z )、(d x ,d y ,d z ) are the three-dimensional space coordinates of the vibration media A, B, C, and D, respectively, and the three-dimensional space coordinates of the vibration media A, B, C, and D are known; R a,b 、R a,c 、R a,d They are the distance differences from the reference vibration medium A to the sound source point and the distances from the vibration media B, C, and D to the sound source point.

[0037] Furthermore, the laser listening device uses a Mach-Zehnder heterodyne interference structure to acquire the vibration signal of the vibrating medium.

[0038] A long-distance non-visual indoor sound source positioning system based on laser Doppler, comprising M laser listening devices, M vibration media and a computer;

[0039] The M vibration media are located at M different monitoring points indoors, and the M laser monitoring devices are respectively set at M monitoring points outdoors, for synchronously acquiring vibration signals of the M vibration media indoors and converting them into voice digital signals accordingly; M ≥ 4;

[0040] The computer includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor implements the following steps:

[0041] (1) Using the improved PHAT weighted generalized cross-correlation delay estimation algorithm, the time difference between the reference vibration medium receiving the sound source signal and the time difference between the other vibration media receiving the sound source signal is calculated, and then converted into the distance difference between the reference vibration medium and the sound source point and the other vibration media and the sound source point;

[0042] (2) Calculate the coordinate position of the sound source point based on the distance difference between each vibration medium and the sound source.

[0043] Furthermore, the laser listening equipment includes a laser transmitter, a laser beam splitter, a laser beam expander, a laser reflector, a laser polarization beam splitter prism, a laser beam combiner, a laser focusing mirror, two laser detectors, a laser demodulation photoelectric converter, a digital signal acquisition card, an acousto-optic modulator, a λ / 2 wave plate, and a transmitter and receiver; the laser transmitter outputs a laser, which is split into two by the laser beam splitter, one of which is modulated by the acousto-optic modulator AOM to form a reference light, and the other is passed through the laser polarization beam splitter prism and the transmitter and receiver to obtain a signal light formed by reflection of the vibrating medium; the signal light and the reference light are combined and mixed by the laser beam combiner, and then received by the laser detector, and the optical signal is converted into an electrical signal by the laser demodulation photoelectric converter, and finally the vibration signal is obtained by sampling the digital signal acquisition card.

[0044] Furthermore, the vibration signal acquired by the laser listening device is expressed as:

[0045]

[0046] Wherein, L is the forced vibration displacement of the vibrating medium caused by the excitation of the sound wave of the sound source, λ is the laser wavelength, and I1 and I2 are the electrical signals converted by the laser demodulation photoelectric converter from the signal light signal received by the two laser detectors and the reference light signal.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1. Higher sensitivity. The present invention is based on the Doppler frequency shift vibration measurement principle of the Mach-Zehnder interferometer structure. Compared with other reflection vibration measurement technologies, it has higher sensitivity and can obtain higher precision target vibration accuracy.

[0049] 2. A wider range of detection media. The present invention can detect a wider range of vibration media, including but not limited to paper materials, wood materials, plastic products, aluminum alloy products, etc., greatly expanding the detection range of laser detection technology.

[0050] 3. Easier operation: The present invention uses a Mach-Zehnder interference structure to integrate a laser transmitting device and a laser receiving device into one, thus achieving a simpler voice monitoring technology.

[0051] 4. Greater concealment: The present invention effectively simplifies monitoring requirements, reduces monitoring difficulty, and has greater concealment.

[0052] 5. Higher sound source positioning accuracy: The present invention adopts the TDOA sound source coordinate positioning algorithm and proposes a new generalized cross-correlation time delay estimation algorithm to obtain higher-precision sound source coordinate positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the Mach-Zehnder heterodyne interference structure of the present invention.

[0054] Figure 2 It is a structural diagram of the laser listening system of the present invention.

[0055] Figure 3 This is a diagram of the application environment of the long-distance laser listening sound source positioning of the present invention;

[0056] Figure 4 This is a flow chart of the improved PHAT weighted generalized cross-correlation delay estimation algorithm proposed in the present invention;

[0057] Figure 5 This is a flow chart of indoor sound source positioning in remote laser listening according to the present invention.

[0058] Figure 6 It is the time domain waveform diagram of the voice monitoring of the present invention.

[0059] Figure 7 It is the time domain waveform diagram of the vibration medium noise of the present invention.

[0060] Figure 8 This is a time difference matching result diagram obtained by the generalized cross-correlation algorithm of the present invention.

[0061] Figure 9 It is a top view of the real coordinate points and estimated coordinate points of the sound source point positioning of the present invention.

[0062] Figure 10 This is a Euclidean distance error diagram between the real coordinate points and the estimated coordinate points of the sound source point positioning of the present invention. DETAILED DESCRIPTION

[0063] The following is a detailed description of an embodiment of the present invention. This embodiment is based on the technical solution of the present invention, provides a detailed implementation method and a specific operation process, and further explains the technical solution of the present invention.

[0064] The present invention provides a long-distance non-visual indoor sound source positioning method based on laser Doppler, such as Figure 5As shown, the method includes: step 1, setting M laser listening devices at M outdoor listening and monitoring points respectively, and using the laser listening devices at the M listening and monitoring points to synchronize time to obtain vibration signals of M vibration media located indoors, where M ≥ 4; step 2, converting the vibration signals of the vibration media into voice digital signals; step 3, using an improved PHAT weighted generalized cross-correlation time delay estimation algorithm to calculate the time difference between the reference vibration medium receiving the sound source signal and the time difference between the other vibration media receiving the sound source signal, and then converting the time difference into the distance difference between the reference vibration medium and the sound source point and the other vibration media and the sound source point; step 4, calculating the coordinate position of the sound source point according to the distance difference between each vibration medium and the sound source.

[0065] The laser listening device in this embodiment can be found in Figure 1 As shown, a non-contact laser vibration measurement structure is adopted, specifically a Mach-Zehnder heterodyne interference structure. The following briefly introduces the Doppler vibration measurement technology of the laser listening device of this embodiment. The laser Doppler vibration measurement technology is developed based on the Doppler frequency shift principle. The Doppler frequency shift formula is as follows:

[0066]

[0067] Where Δf is the Doppler frequency shift, c is the speed of light, f is the frequency of the wave source, and the object moves at a speed v at an angle θ relative to the sound source.

[0068] From the above formula, we can know that the relative vibration velocity v of the object under test can be calculated by obtaining the Doppler frequency shift, collecting the reflected light path, and setting it to θ = 0. Assume that the frequencies of the signal light and reference light are fs1 and fs2, and after photoelectric conversion, the electric field intensities E1 and E2 of the signal light and reference light are

[0069]

[0070]

[0071] Among them, E 01 and E 02 is the amplitude of the signal light and the reference light, and is the initial phase of the signal light and the reference light. After the two beams interfere with each other, the relationship between their electric field intensity and light intensity is:

[0072]

[0073] in, Is a constant. In heterodyne interference, Δf is the vibration frequency of the forced vibration target object, Δf=f s1 -f s2Thus, the correlation between the vibration frequency of the forced vibration medium and the electric field strength of the sensor is realized, and the vibration frequency fluctuation of the forced vibration target can be measured by the fluctuation of the electric field strength.

[0074] Specifically, the laser listening device in this embodiment can be found in Figure 2 As shown, the system includes a laser transmitter, a laser beam splitter, a laser beam expander, a laser reflector, a laser polarization beam splitter prism, a laser beam combiner, a laser focusing lens, a laser detector, a laser demodulation and photoelectric conversion system, a digital signal acquisition card, an acousto-optic modulator (AOM), a λ / 2 wave plate, a transmitting and receiving system, and a sealed box. Its operating principle is that a continuous laser emits laser light, which is split into two by a beam splitter. One optical path is modulated by an acousto-optic modulator (AOM) to form a reference beam, while the other optical path passes through a polarization beam splitter prism and a transmitting and receiving system to obtain the signal light reflected by the target vibrating medium. After the signal and reference beams are combined and mixed by the beam combiner, they are received by a detector. The laser demodulation and photoelectric conversion system converts the optical signal into an electrical signal, which is then sampled by a digital signal acquisition card to obtain a digital signal, which is ultimately transmitted to a computer for processing. The basic principle of this laser listening device to obtain voice digital signals is: using the laser difference frequency method to detect the slight vibration of the forced vibration medium caused by sound propagation, extracting the laser Doppler difference frequency changes caused by the forced vibration of the medium from the detected signal, demodulating and filtering the signal to restore the sound information.

[0075] Assuming the laser emission frequency F, the frequency shifter modulation frequency f, and the movement distance of the vibrating medium is L, the vibration equation after the laser is emitted and reflected from the vibrating medium is:

[0076]

[0077] Where A is the amplitude of the emission light path, is the initial phase, and λ is the laser wavelength.

[0078] Similarly, the vibration equation of the reference light path is as follows:

[0079]

[0080] After heterodyne interference, the interference signals on the detector target surface are as follows:

[0081]

[0082]

[0083] After the square rate of the detector is collected, the current signals output by detectors 1 and 2 are respectively as follows after DC isolation:

[0084]

[0085]

[0086] The modulated signal f of the frequency shifter is down-converted and then low-pass filtered to obtain the following two optical signals:

[0087]

[0088]

[0089] The motion displacement of the vibrating medium can be calculated by the inverse tangent transformation

[0090]

[0091] The above formula is the expression of the forced vibration displacement generated by the vibrating medium when it is excited by the sound wave of the sound source, L is the forced vibration displacement generated by the vibrating medium when it is excited by the sound wave of the sound source, λ is the laser wavelength, I1 and I2 are respectively: the signal light signal received by the two laser detectors and the reference light signal converted by the laser demodulation photoelectric converter to obtain the electrical signals.

[0092] Since the vibration signal obtained by the laser listening device is the forced vibration displacement generated by the vibration medium being excited by the sound wave of the sound source, the vibration signal can be directly equivalent to a voice digital signal.

[0093] See also Figure 3 , showing a major application environment of the present invention. The left side of the figure is where the sound source signal is located, and the right side is where the laser listening device is located. There are obstacles such as walls in the middle, which realizes a long-distance non-visual environment. It is necessary to use 3-4 laser listening and observation devices. When using 3 laser listening and observation devices, the TDOA two-dimensional positioning algorithm can be applied to realize the positioning of the sound source coordinate point in the two-dimensional plane. When using 4 laser listening and observation devices, the TDOA three-dimensional space positioning algorithm can be applied to realize the positioning of the sound source coordinate point in the three-dimensional space.

[0094] The following explains the solution of the present invention for locating the three-dimensional spatial coordinates of the indoor sound source point S using four forced vibration media.

[0095] The computer receives a voice digital signal from the laser listening device. i After (n), the sound source is located by calculating the distance difference between each vibration medium and the sound source point. Where i = 1, 2, 3, 4 corresponds to the four vibration media A, B, C, and D respectively.

[0096] Among them, please read Figure 4 As shown in Figure 2, the calculation process of the distance difference between each vibration medium and the sound source point is:

[0097] Step 3.1: for each voice digital signal I i(n) Perform discrete Fourier transform to obtain I i (k):

[0098] I i (k)=DFT{I i (n)}

[0099] Among them, i=1, 2, 3, 4 correspond to the four vibration media A, B, C, and D respectively;

[0100] Step 3.2: Compare the discrete Fourier signal I1(k) corresponding to the reference vibration medium A with the discrete Fourier signals I1(k) corresponding to the other vibration media B, C, and D. j Multiply the conjugate complex numbers of (k) by two or more to get the complex product I 1j (k):

[0101]

[0102] Where j = 2, 3, 4;

[0103] Step 3.3, multiply the complex product I 1j (k) Divide its modulus value to obtain the intermediate parameter

[0104]

[0105] Step 3.4, retain the speech interest frequency band ROI∈[f roi_min ,f roi_max ], set other frequency band values ​​to zero:

[0106]

[0107] Among them, f roi_min ,f roi_max are the lower and upper limits of the frequency band ROI of the speech interest domain, hz(k) is the actual frequency corresponding to frequency point k, hz(k) = k*(fs / N), fs is the sampling rate of the speech digital signal, and N is the length of the speech digital signal, that is, the number of sampling points;

[0108] Step 3.5, Perform inverse discrete Fourier transform to obtain intermediate parameters

[0109]

[0110] Step 3.6, calculate the distance difference between the reference vibration medium A and the sound source point S and the vibration medium B, C, D and the sound source point S:

[0111]

[0112]

[0113]

[0114] Where R a,b 、R a,c 、R a,d are the distance differences between the reference vibration medium A and the sound source point S and the vibration medium B, C, and D and the sound source point; v is the speed of sound, fs is the sampling rate; ind max {·} indicates the position where the maximum value is taken.

[0115] Step 4 specifically uses the TODA positioning algorithm to calculate the coordinate position of the sound source point:

[0116]

[0117] Where, (s x ,s y ,s z ) is the three-dimensional space coordinate of the sound source point S, (a x ,a y ,a z )、(b x ,b y ,b z )、(c x ,c y ,c z )、(d x ,d y ,d z ) are the three-dimensional spatial coordinates of vibration media A, B, C, and D, respectively, and the three-dimensional spatial coordinates of vibration media A, B, C, and D are known.

[0118] The laser detection technology used in this invention, based on the Doppler frequency shift principle and Mach-Zehnder interference structure, has higher accuracy and anti-interference ability than existing laser vibration measurement technology. Its basic principle has also been effectively verified and applied in many industrial and scientific research fields. In addition, the method and device proposed in this invention have been implemented in a laboratory environment. Figure 6 , is the voice interception time domain waveform obtained by the present invention. Figure 7 , is the time domain waveform of the vibration medium noise obtained by the present invention. Figure 8 , is the time difference matching result diagram obtained by the improved PHAT weighted generalized cross-correlation delay estimation algorithm proposed in this invention. Figure 9 , is the top view of the real coordinates and estimated coordinates of the sound source point obtained by the present invention. Figure 10 , is a Euclidean distance error diagram between the real coordinate points and the estimated coordinate points of the sound source point positioning result obtained by the present invention.

[0119] The above embodiments are preferred embodiments of the present application. Ordinary technicians in this field can also make various changes or improvements on this basis. Without departing from the overall concept of the present application, these changes or improvements should fall within the scope of protection required by the present application.

Claims

1. A long-distance non-visual indoor sound source localization method based on laser Doppler, characterized in that: include: Step 1, Laser listening devices are set up outdoors Listening monitoring points, use The laser listening equipment at each listening monitoring point synchronizes the time to obtain the information in the indoor A vibration signal of a vibrating medium; ; Step 2, converting the vibration signal of the vibration medium into a voice digital signal; Step 3: Using the improved PHAT weighted generalized cross-correlation delay estimation algorithm, calculate the time difference between the reference vibration medium receiving the sound source signal and each other vibration medium receiving the sound source signal, and then convert it into the distance difference between the reference vibration medium and the sound source point and each other vibration medium and the sound source point. The specific calculation process of step 3 is as follows: Step 3.1: For each voice digital signal , perform discrete Fourier transform to get : ; in, They correspond to the four vibration media A, B, C, and D respectively; Step 3.2: The discrete Fourier transform signal corresponding to the reference vibration medium A is Discrete Fourier signals corresponding to other vibration media B, C, and D Multiply the conjugate complex numbers of two by two to get the complex product : ; in, ; Step 3.3, multiply the complex product Divide by its modulus value to get the intermediate parameter : ; Step 3.4: Retain the frequency band of interest in the speech domain , set the other frequency band values ​​to zero: ; in, The frequency bands of speech interest domain are The lower and upper limits of Frequency point The corresponding actual frequency, , is the sampling rate of the voice digital signal, is the length of the voice digital signal, that is, the number of sampling points; Step 3.5, Perform inverse discrete Fourier transform to obtain intermediate parameters : ; Step 3.6, calculate the distance difference between the reference vibration medium A and the sound source point S and the vibration medium B, C, D and the sound source point S: ; ; ; Where, 、 、 are the distance differences between the reference vibration medium A and the sound source point S and the vibration media B, C, and D and the sound source point respectively; is the speed of sound, is the sampling rate; Indicates the location of the maximum value; Step 4, calculating the coordinate position of the sound source point based on the distance difference between each vibration medium and the sound source point; Specifically, the TODA positioning algorithm is used to calculate the coordinate position of the sound source point: ; Where, The sound source point The three-dimensional space coordinates of 、 、 、 are the three-dimensional spatial coordinates of the vibration media A, B, C, and D, respectively, and the three-dimensional spatial coordinates of the vibration media A, B, C, and D are known; 、 、 They are the distance differences from the reference vibration medium A to the sound source point and the distances from the vibration media B, C, and D to the sound source point.

2. The method according to claim 1, characterized in that The laser listening equipment includes a laser transmitter, a laser beam splitter, a laser beam expander, a laser reflector, a laser polarization splitter prism, a laser beam combiner, a laser focusing mirror, two laser detectors, a laser demodulation photoelectric converter, a digital signal acquisition card, an acousto-optic modulator, a λ / 2 wave plate, and a transmitter and receiver; The process of the laser listening device acquiring the vibration signal of the vibrating medium is as follows: the laser transmitter outputs laser light, which is split into two by a laser beam splitter, one of which is modulated by an acousto-optic modulator (AOM) to form a reference light, and the other is passed through a laser polarization splitter prism and a transmitter-receiver to acquire the signal light formed by the reflection of the vibrating medium; the signal light and the reference light are combined and mixed by a laser beam combiner, and then received by a laser detector, and the optical signal is converted into an electrical signal by a laser demodulation photoelectric converter, and finally the vibration signal is obtained by sampling through a digital signal acquisition card.

3. The method according to claim 2, characterized in that The vibration signal acquired by the laser listening device is expressed as: ; Where, It is the forced vibration displacement generated by the vibration medium being excited by the sound wave of the sound source. is the laser wavelength, and They are: the signal light signal received by the two laser detectors in the laser listening device and the reference light signal, which are converted into electrical signals by the laser demodulation photoelectric converter.

4. The method according to claim 1, wherein The laser listening device adopts a Mach-Zehnder heterodyne interference structure to acquire the vibration signal of the vibrating medium.

5. A long-distance non-visual indoor sound source localization system based on laser Doppler, characterized in that: It includes M laser listening devices, M vibration media and computers; The M vibration media are located at M different listening and monitoring points indoors, and the M laser listening devices are respectively set at M listening and monitoring points outdoors, and are respectively used to synchronously acquire vibration signals of the M vibration media indoors and convert them into voice digital signals accordingly; ; The computer includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor implements the following steps: (1) The improved PHAT weighted generalized cross-correlation delay estimation algorithm is used to calculate the time difference between the reference vibration medium receiving the sound source signal and the time difference between the reference vibration medium receiving the sound source signal and each other vibration medium receiving the sound source signal, and then convert it into the distance difference between the reference vibration medium and the sound source point and each other vibration medium and the sound source point; the specific calculation process is as follows: Step 3.1: For each voice digital signal , perform discrete Fourier transform to get : ; in, They correspond to the four vibration media A, B, C, and D respectively; Step 3.2: The discrete Fourier transform signal corresponding to the reference vibration medium A is Discrete Fourier signals corresponding to other vibration media B, C, and D Multiply the conjugate complex numbers of two by two to get the complex product : ; in, ; Step 3.3, multiply the complex product Divide by its modulus value to get the intermediate parameter : ; Step 3.4: Retain the frequency band of speech interest , set the other frequency band values ​​to zero: ; in, The frequency bands of speech interest domain are The lower and upper limits of Frequency point The corresponding actual frequency, , is the sampling rate of the digital voice signal, is the length of the voice digital signal, that is, the number of sampling points; Step 3.5, Perform inverse discrete Fourier transform to obtain intermediate parameters : ; Step 3.6, calculate the distance difference between the reference vibration medium A and the sound source point S and the vibration medium B, C, D and the sound source point S: ; ; ; Where, 、 、 are the distance differences between the reference vibration medium A and the sound source point S and the vibration media B, C, and D and the sound source point respectively; is the speed of sound, is the sampling rate; Indicates the location of the maximum value; (2) Calculate the coordinate position of the sound source point based on the distance difference between each vibration medium and the sound source; Specifically, the TODA positioning algorithm is used to calculate the coordinate position of the sound source point: ; Where, The sound source point The three-dimensional space coordinates of 、 、 、 are the three-dimensional spatial coordinates of the vibration media A, B, C, and D, respectively, and the three-dimensional spatial coordinates of the vibration media A, B, C, and D are known; 、 、 They are the distance differences from the reference vibration medium A to the sound source point and the distances from the vibration media B, C, and D to the sound source point.

6. The system according to claim 5, characterized in that The laser listening device includes a laser transmitter, a laser beam splitter, a laser beam expander, a laser reflector, a laser polarization beam splitter prism, a laser beam combiner, a laser focusing mirror, two laser detectors, a laser demodulation photoelectric converter, a digital signal acquisition card, an acousto-optic modulator, a λ / 2 wave plate, and a transmitter and receiver. The laser transmitter outputs a laser, which is split into two by the laser beam splitter. One optical path is modulated by the acousto-optic modulator AOM to form a reference light, and the other optical path is passed through the laser polarization beam splitter prism and the transmitter and receiver to obtain a signal light formed by reflection from a vibrating medium. After the signal light and the reference light are combined and mixed by the laser beam combiner, they are received by the laser detector, and the optical signal is converted into an electrical signal by the laser demodulation photoelectric converter, and finally the vibration signal is obtained by sampling the digital signal acquisition card.

7. The system according to claim 6, characterized in that The vibration signal acquired by the laser listening device is expressed as: ; Where, It is the forced vibration displacement generated by the vibration medium being excited by the sound wave of the sound source. is the laser wavelength, and They are: the signal light signal received by the two laser detectors and the reference light signal are converted into electrical signals by a laser demodulation photoelectric converter.

Citation Information

Patent Citations

  • Method and device for voice interception, and laser-bounce voice source locating method

    CN104036787A

  • High-precision laser echo frequency modulation system and method

    CN109375230A