A multi-dimensional noise reduction method based on combining noise cancellation with a classifier
By combining an outdoor sound-receiving unit, a visual positioning device, and an indoor vibration source, the cancellation sound waves are generated and adjusted in real time, solving the problem of inaccurate noise reduction in existing technologies and achieving precise noise reduction and safety sound wave exemption in both indoor and outdoor environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU HOLLEY COLLEGE
- Filing Date
- 2026-02-14
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies cannot accurately reduce noise in indoor and outdoor environments, especially since they cannot account for noise generated by the vibration of objects such as window glass and doors, and they cannot adjust in real time to offset the dynamic changes in the position of sound waves and human ears, resulting in inaccurate noise reduction effects.
Noise signals are collected by an outdoor sound unit, and the position of the human ear is obtained in real time by a visual positioning device and a classifier. The phase compensation amount is calculated, and a canceling sound wave is generated. An anti-phase vibration is generated by an indoor vibration source, and the pan-tilt attitude and the extension and retraction of the sound unit are adjusted in real time to achieve multi-dimensional noise reduction.
It ensures precise coupling between the canceled sound waves and noise at the human ear, effectively canceling indoor and outdoor noise while taking into account the exemption of safe sound, and achieves dynamic adjustment to improve the reliability and stability of noise reduction.
Smart Images

Figure CN122090813A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent noise reduction technology, specifically to a multi-dimensional noise reduction method based on the combination of noise cancellation and classifier. Background Technology
[0002] Reducing indoor noise has many important benefits, including improving physical and mental health, quality of life, and living environment. By blocking out external noise, it can help relieve stress and fatigue, improve the body's immunity and resistance, and avoid the anxiety and irritability that can be caused by exposure to noise. This creates a quiet indoor environment that promotes relaxation and enhances happiness and life satisfaction.
[0003] For example, Chinese patent application number 202211128319.3, published on December 27, 2022, discloses an indoor noise reduction system, method, apparatus, device, and readable storage medium. The indoor noise reduction system includes a microphone array, a camera device, a processor, and at least one speaker. The microphone array collects indoor ambient noise, the camera device collects images of the listener, the processor determines the listener's ear position based on the image collected by the camera device, and determines the noise at the listener's ear position based on the indoor ambient noise collected by the microphone array and the listener's ear position; it generates a sound wave with the same amplitude and frequency but opposite phase as the noise at the listener's ear position as a noise reduction frequency signal; and it sends the noise reduction frequency signal to the speaker for playback.
[0004] The aforementioned literature collects noise using a microphone array, captures images of the listener using a camera device, determines the position of the listener's ear using a processor, and generates a sound wave with the same amplitude and frequency as the noise but opposite phase as a noise reduction frequency signal. This signal is then played through a speaker to cancel the noise in reverse. However, it only achieves noise reduction through the noise reduction frequency signal and only considers indoor environmental noise. It does not consider the noise generated by the vibration of objects such as windows and doors in indoor and outdoor environments. It cannot accurately reduce environmental noise based on the audio parameter characteristics of the noise, rather than simply reducing it in the opposite direction of amplitude and frequency. In addition, after noise processing, it does not dynamically update the acquired noise and make real-time adjustments based on the dynamic changes in the position of the sound unit and the listener's ear. Therefore, it cannot ensure that the canceling sound wave and noise are accurately coupled at the listener's ear during the dynamic process. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-dimensional noise reduction method based on the combination of noise cancellation and classifier, which can ensure that the cancellation sound wave and noise are accurately coupled at the human ear, while realizing vibration transmission cancellation and safety sound exemption.
[0006] To achieve the above objectives, this invention provides a multi-dimensional noise reduction method based on a combination of noise cancellation and a classifier. This method is implemented through a noise reduction system, which includes an outdoor recording unit, an indoor vibration source, a pan-tilt unit positioned along the path between the outdoor and indoor areas, a visual positioning device rotating on the pan-tilt unit, and a retractable sound-emitting unit. The method includes the following steps: S1. Noise signals from different directions are collected by an outdoor receiver unit, and the noise signals are preprocessed to form preprocessed noise signals and noise parameters are obtained, including the dominant frequency and amplitude. S2. The position of the human ear is obtained in real time by a visual positioning device and classifier set up indoors, and the distance between the human ear and the visual positioning device is calculated. S3. Determine the phase compensation amount based on the distance between the human ear and the visual positioning device and the speed of sound. Determine the cancellation sound wave parameters based on the wavelength, amplitude and phase compensation amount in the noise parameters. S4. Obtain the vibration parameters of the indoor environment, and then make the vibration source set up indoors generate anti-phase vibration according to the vibration parameters; S5. Adjust the gimbal posture and the extension / retraction of the sound unit in real time based on the relationship between the position of the human ear and the position of the sound unit, and generate a cancellation sound wave for noise reduction based on the cancellation sound wave parameters. S6. Obtain the noise reduction amount by the difference between the noise sound pressure level before noise reduction and the residual noise sound pressure level after noise reduction. Then, based on whether the noise reduction amount is within the preset threshold range, readjust the noise parameters in step S1, the vibration parameters of the vibration source in step S4, and the positions of the human ear and visual positioning device in step S5.
[0007] The above method preprocesses the noise using an outdoor noise receiving unit to obtain noise parameters, including the dominant frequency and amplitude. The dominant frequency is the frequency corresponding to the maximum value in the noise frequency domain, thus enabling more accurate determination of the noise type and noise cancellation. Furthermore, it determines the phase difference based on the distance between the human ear and the visual positioning device, ensuring better coupling between the human ear and the outdoor noise sound waves at the intersection point. Simultaneously, it detects vibration parameters in the indoor environment and generates opposite vibrations based on these parameters to achieve vibration cancellation, preventing indoor noise from being affected by outdoor noise vibrations and further ensuring the reliability of noise reduction. After noise reduction, the difference between the noise before and after reduction is compared to determine if the noise reduction requirements are met, leading to further adjustments to the positions of the human ear and visual positioning device, indoor environmental vibration parameters, and noise parameters, thus ensuring active and closed-loop noise reduction processing.
[0008] Furthermore, step S5 also includes: extracting the features of the preprocessed noise signal and converting them into a quantifiable first feature vector, and then calculating the cosine similarity between the first feature vector and the second feature vector in the preset sound sample library. If the similarity exceeds a preset value, then no noise reduction processing is performed.
[0009] The above settings, by matching noise samples, can effectively avoid blocking sounds that users need or that are related to safety, such as rain sounds and cries for help, thus balancing noise reduction effectiveness and safety.
[0010] Furthermore, step S1 includes: steps S1.1 to S1.3. S1.1. Noise signals from different directions are collected through the radio unit, and the time domain signal of the noise signal is denoted as si(t), i=1,2,...,n,n≥3,t is time. Then, the discrete time domain data sequence si[k] is output through the radio unit, where k=1,2,...,N,N≥3,N is the number of sampling points; S1.2, Denoising using wavelet thresholding, threshold ,but (1), w j,k These are the wavelet decomposition coefficients. Denoising coefficients These are preset parameters; Then, after filtering with a bandpass filter, a clean noise signal s is obtained. clean (t); S1.3, Processing the noise signal s clean (t) is subjected to a Discrete Fourier Transform, and then transformed to the frequency domain. (2), m=0,1,...,N 1. S[m] represents the frequency domain coefficients, j represents the imaginary unit, and k represents a value; Noise parameters are extracted, including the dominant frequency f0, amplitude A0, and wavelength λ0, where the dominant frequency f0 is the frequency corresponding to the maximum value of |S[m]|. (3), m0 is the frequency index corresponding to the maximum amplitude, f s f is the sampling frequency. s =48kHz; (4), (5), v is the speed of sound in air at standard atmospheric pressure.
[0011] The above settings, by preprocessing the noise signal, can effectively eliminate interference signals, ensuring the accuracy of subsequent noise parameter calculations. After the noise parameters are calculated and obtained, they can provide a basis for subsequent cancellation of sound wave generation and gimbal adjustment.
[0012] Furthermore, in step S5, "extracting the preprocessed noise signal features and converting them into a quantifiable first feature vector, then calculating the cosine similarity between the first feature vector and the second feature vector in the preset sound sample library; if the similarity exceeds a preset value, then no noise reduction processing is performed," includes steps S5.1 to S5.2. S5.1, For the preprocessed noise signal s clean (t) Extract the Mel frequency cepstral coefficients and generate the first eigenvector X=[x1,x2,...,x 13 ]; S2.2, Combine the first feature vector X with the second feature vector Yi = [Y] in the preset sound sample library. 1,k ,Y 2,k ,...,Y 13,k Calculate the cosine similarity S, where i is the sample index, i=1,2,... (6), Then, it is determined whether the noise signal is exempted. If S≥T, then no cancellation sound wave is generated; otherwise, a cancellation sound wave is generated for noise reduction. T is a preset cosine similarity threshold.
[0013] The above settings convert noise signals into quantifiable feature vectors, which are convenient for subsequent noise sample matching. By matching noise using cosine similarity, it is possible to effectively avoid blocking rain sounds, cries for help, and other sounds that users need or that are related to safety, thus balancing noise reduction effect with safety and personalized needs.
[0014] Furthermore, step S4 includes: steps S4.1-S4.2, S4.1. Vibration acceleration signals a(t) of the indoor environment are collected by vibration sensors installed along the length and width directions of the indoor environment, where a(k) is the value at point k in a(t). Vibration parameters are extracted by discrete Fourier transform, including the dominant vibration frequency f. v With amplitude A v ,but (7); (8); m v This is the frequency index corresponding to the maximum value of |A[m]|. (9); S4.2 Calculate the driving voltage U(t) of the vibration source based on the vibration parameters, so that the vibration source produces vibrations with the same frequency, matching amplitude, and opposite direction as the indoor environment. (10) K is the vibration source voltage-vibration conversion coefficient.
[0015] The above settings, by acquiring key parameters of indoor environmental vibrations such as those on doors and windows, can provide a basis for the vibration source to generate anti-phase vibrations, thereby driving the vibration source to generate anti-phase vibrations, offsetting the vibrations caused by external noise in the indoor environment, and initially reducing noise entering the room.
[0016] Furthermore, step S2 includes: steps S2.1 to S2.3. S2.1. Indoor images are acquired using a binocular camera. Then, the human ear region in the images is identified through Haar-like features and an Adaboost classifier. The pixel coordinates (u) of the human ear center in the left camera image are extracted. L ,v L ), pixel coordinates of the right camera (u R ,v R ); S2.2 Calculate the parallax d based on the pixel coordinates of the left and right cameras, and calculate the straight-line distance Z between the human ear and the camera based on the intrinsic parameters of the binocular camera. (11), p is the pixel size, and f1 is the focal length; (12); S2.3. Combining the intrinsic parameter matrix K of the stereo camera, the pixel coordinates are converted into three-dimensional coordinates (X, Y, Z) in the camera coordinate system, where... (13) u0 and v0 are the pixel coordinates of the camera's principal point. (14) (15).
[0017] The above settings first use Haar-like features, which are rectangular features based on image grayscale differences. Adaboost (adaptive boosting algorithm) is a weak classifier ensemble algorithm. In the "suspected human ear region" filtered by Haar-like features, interference is eliminated and "whether it is a human ear" is accurately determined. This allows for more accurate identification of the human ear region and precise positioning of the human ear in the image. Then, the two-dimensional image coordinates are converted into distance information in three-dimensional space to determine the relative distance between the human ear and the visual positioning device. In this way, obtaining the position of the human ear in three-dimensional space can provide target coordinates for subsequent gimbal attitude adjustment.
[0018] Furthermore, step S3 includes: steps S3.1 to S3.2. S3.1 Generate an anti-phase cancelling sound wave based on the noise parameters, ensuring that the cancelling sound wave and the noise have the same vibration frequency, matching amplitude, and opposite phase. Let the time-domain signal of the cancelling sound wave be s. cancel (t), (16) Indicates opposite phase, Δ This is the phase compensation amount. (17) L is the distance from the sound-producing unit to the human ear, and v is the speed of sound in air under standard atmospheric pressure; S3.2. Based on the noise wavelength λ0, adjust the extension / retraction amount ΔL of the sound unit so that the distance L from the sound unit to the ear is an integer multiple of λ0 / 2. (18) in (19) L current L represents the current distance from the sound-producing unit to the human ear. target Let k be the target distance from the sound-producing unit to the human ear, and k be a preset positive integer, taking the closest value to L. current The value of .
[0019] The above settings, by generating precise canceling sound waves, ensure that the canceling sound waves and noise meet the "destructive interference" condition at the human ear. At the same time, by determining the target distance as a preset positive integer and half of the wavelength, the phase difference caused by the sound wave propagation delay is compensated to ensure that the phase difference π between the canceling sound waves and the noise is π when the sound waves reach the human ear. The scaling adjustment optimizes the sound wave propagation path and further improves the stability of the cancellation effect.
[0020] Furthermore, step S5 includes: steps S5.1 to S5.2. S5.1 Calculate the horizontal and vertical rotation angles of the pan-tilt unit based on the three-dimensional coordinates of the human ear (X,Y,Z) and the initial coordinates of the sound-generating unit (X0,Y0,Z0). , (20) (twenty one); S5.2 After a preset interval Δt1, repeat step S5.1 and compare the current horizontal rotation angle θ. current Horizontal rotation angle θ with respect to the target target Horizontal deviation Δθ, current vertical rotation angle current Vertical angle with the target target Vertical deviation Δ 1, (twenty two), (twenty three); If the horizontal deviation Δθ > 0.5° or the vertical deviation Δ If 1 > 0.5°, then drive the gimbal to adjust to the horizontal or vertical target rotation angle.
[0021] The above settings, by adjusting the angles of the horizontal and vertical rotation, enable the sound unit to be precisely aligned with the human ear, ensuring that the sound waves are directionally propagated to the target location. They can also correct the gimbal posture in real time to cope with deviations caused by human ear movement or equipment vibration, and maintain the directional noise reduction effect.
[0022] Furthermore, step S6 includes: steps S6.1 to S6.2. S6.1 Update the noise parameters obtained by the receiving unit, the three-dimensional coordinates (X,Y,Z) of the human ear in the visual positioning device, and the vibration parameters of the indoor environment every preset interval time Δt2. S6.2, Collect the residual noise sound pressure level (SPL) after noise reduction using a sound acquisition device placed near the human ear. res Then compared with the noise sound pressure level (SPL) before noise reduction. original The noise reduction amount ΔSPL is obtained by comparison. (twenty four), like (25) Then re-enter S1, where SPL is the preset noise reduction threshold.
[0023] The above settings ensure timely updates to the collected environmental data, address dynamic scenarios such as noise changes and ear movement, and continuously optimize parameters through dynamic global effect feedback to ensure that the noise reduction effect consistently meets expectations. Attached Figure Description
[0024] Figure 1 This is a flowchart of the process of the present invention.
[0025] Figure 2 This is a partial view of the noise reduction system in this invention. Detailed Implementation
[0026] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0027] like Figure 1 and Figure 2 As shown, a multi-dimensional noise reduction method based on combining noise cancellation and a classifier is implemented through a noise reduction system. The noise reduction system includes an outdoor recording unit, an indoor vibration source, a pan-tilt unit positioned along the path between the outdoor and indoor areas, a visual positioning device rotating on the pan-tilt unit, and a retractable sound-emitting unit. The specific steps include: S1. Noise signals from different directions are collected by two or more outdoor sound receiving units. In this embodiment, the sound receiving unit is a microphone array. The noise signals are preprocessed to form preprocessed noise signals and noise parameters are obtained. The noise parameters include the main frequency and amplitude. Step S1 includes: steps S1.1 to S1.3. S1.1. Noise signals from different directions are collected through the radio unit, and the time domain signal of the noise signal is denoted as si(t), i=1,2,...,n,n≥3,t is time. Then, the discrete time domain data sequence si[k] is output through the radio unit, where k=1,2,...,N,N≥3,N is the number of sampling points; S1.2, Denoising using wavelet thresholding, threshold ,but (1), w j,k These are the wavelet decomposition coefficients. Denoising coefficients; These are preset parameters; Then, after filtering with a bandpass filter, a clean noise signal s is obtained. clean (t); S1.3, Processing the noise signal s clean (t) is subjected to a Discrete Fourier Transform, and then transformed to the frequency domain. (2), m=0,1,...,N 1. S[m] represents the frequency domain coefficients, j represents the imaginary unit; k represents the possible values; N represents the total number of coefficients. Noise parameters are extracted, including the dominant frequency f0, amplitude A0, and wavelength λ0, where the dominant frequency f0 is the frequency corresponding to the maximum value of |S[m]|. (3), m0 is the frequency index corresponding to the maximum amplitude, f s f is the sampling frequency. s =48kHz; (4), (5), v is the speed of sound in air under standard atmospheric pressure, v = 340 m / s.
[0028] S2. The position of the human ear is obtained in real time by a visual positioning device and classifier set up indoors, and the distance between the human ear and the visual positioning device is calculated. S3. Determine the phase compensation amount based on the distance between the human ear and the visual positioning device and the speed of sound. Determine the cancellation sound wave parameters based on the wavelength, amplitude and phase compensation amount in the noise parameters. S4. Obtain the vibration parameters of the indoor environment, and then make the vibration source set up indoors generate anti-phase vibration according to the vibration parameters; S5. Adjust the gimbal posture and the extension / retraction of the sound unit in real time based on the relationship between the position of the human ear and the position of the sound unit, and generate a cancellation sound wave for noise reduction based on the cancellation sound wave parameters. S6. Obtain the noise reduction amount by the difference between the noise sound pressure level before noise reduction and the residual noise sound pressure level after noise reduction. Then, based on whether the noise reduction amount is within the preset threshold range, readjust the noise parameters in step S1, the vibration parameters of the vibration source in step S4, and the positions of the human ear and visual positioning device in step S5.
[0029] In another embodiment, to further improve the recognition accuracy, step S5 further includes: extracting the features of the preprocessed noise signal and converting them into a quantifiable first feature vector, and then calculating the cosine similarity between the first feature vector and the second feature vector in the preset sound sample library. If the similarity exceeds a preset value, no noise reduction processing is performed; if the similarity does not exceed the preset value, noise reduction processing is performed. Specifically, this includes steps S5.1 to S5.2. S5.1, For the preprocessed noise signal s clean (t) Extract Mel-frequency cepstral coefficients (MFCCs) and generate the first eigenvector X=[x1,x2,...,x 13 ]; S5.2, Combine the first feature vector X with the second feature vector Yi=[Y] in the preset sound sample library. 1,k ,Y 2,k ,...,Y 13,kCalculate the cosine similarity S, where i is the sample index, i=1,2,... The preset sound sample library is stored in a storage module located indoors. (6), Then, it is determined whether the noise signal is exempted. If S≥T, then no cancellation sound wave is generated; otherwise, a cancellation sound wave is generated for noise reduction. T is a preset cosine similarity threshold. In this embodiment, T=0.8. For example, if the first feature vector X=[0.1,0.2,...,0.3] of the Mel frequency cepstral coefficient of the noise and the second feature vector Y1=[0.08,0.19,...,0.28] of the distress call sample, and S=0.85≥0.8 is calculated, then the noise reduction exemption is triggered, thereby avoiding the blocking of rain sounds, distress calls and other sounds that users need or are related to safety, taking into account both noise reduction effect and safety and personalized needs.
[0030] In this embodiment, the noise reduction process includes generating noise cancellation and vibration from the vibration source.
[0031] Step S4 specifically includes: steps S4.1-S4.2; S4.1. Vibration acceleration signals a(t) of the indoor environment are collected by vibration sensors installed along the length and width directions of the indoor environment, where a(k) is the value of a(t) at time k. Vibration parameters are extracted by discrete Fourier transform, including the dominant vibration frequency f. v With amplitude A v ,but (7); (8); m v This is the frequency index corresponding to the maximum value of |A[m]|. (9); S4.2 Calculate the driving voltage U(t) of the vibration source based on the vibration parameters, so that the vibration source produces vibrations with the same frequency, matching amplitude, and opposite direction as the indoor environment. (10) K is the vibration source voltage-vibration conversion coefficient. In this embodiment, the indoor environment includes indoor glass aluminum windows and doors, and the vibration source is a piezoelectric ceramic. K is determined by the type of piezoelectric ceramic, such as K=10V·s. 2 / m, while simultaneously collecting the dominant frequency f of the indoor environmental horizontal vibration acceleration a(t) a(t). v =200Hz, amplitude A v =0.5m / s 2 Then U(t) = 10 × 0.5 × sin(2π × 200t) = 5sin(400πt).
[0032] Step S2 specifically includes: steps S2.1 to S2.3; S2.1. Indoor images are acquired using a binocular camera. Then, the human ear region in the images is identified through Haar-like features and an Adaboost classifier. The pixel coordinates (u) of the human ear center in the left camera image are extracted. L ,v L ), right camera (u R ,v R Haar-like features are rectangular features based on image grayscale differences, and Adaboost classifier is a weak classifier ensemble algorithm. S2.2 Calculate the parallax d based on the pixel coordinates of the left and right cameras, and calculate the straight-line distance Z between the human ear and the camera based on the intrinsic parameters of the binocular camera. (11), p is the physical size of the pixel (p = 3μm = 3 × 10⁻⁶). 6 m), (12); Where b is the baseline of the binocular camera; In one embodiment, if the left camera pixel u L =500, right camera pixel u R =480, then d=(500) 480)×3×10 6 =6×10 5 m; Depth Z = (0.15 × 5 × 10 m) 3 ) / 6×10 5 =12.5m.
[0033] S2.3. Combining the intrinsic parameter matrix K of the stereo camera, the pixel coordinates are converted into three-dimensional coordinates (X, Y, Z) in the camera coordinate system, where... (13) u0 and v0 are the pixel coordinates of the camera's principal point (if the image center is taken, then u0=640, v0=360). (14) (15).
[0034] In one embodiment, if u L=500, u0=640, then X=(500 640)×12.5×3×10 6 / 5×10 3 = 1.05m (the negative sign indicates that the human ear is in the negative X-axis direction of the camera); Y=(360 360)×...=0m (the human ear is at the center of the camera's Y-axis).
[0035] Step S3 specifically includes: steps S3.1 to S3.2; S5.1 Generate an anti-phase cancelling sound wave based on the noise parameters, ensuring that the cancelling sound wave and the noise have the same vibration frequency, matching amplitude, and opposite phase. Let the time-domain signal of the cancelling sound wave be s. cancel (t), (16) Indicates opposite phase, Δ This is the phase compensation amount. (17) L is the distance from the sound-producing unit to the human ear, and v is the speed of sound in air under standard atmospheric pressure; If L=10m, then Δ =2π×10 / 340≈0.184rad; the canceled sound wave is s cancel (t)= 0.1×sin(2π×500t+0.184); S3.2. Based on the noise wavelength λ0, adjust the extension / retraction amount ΔL of the sound unit so that the distance L from the sound unit to the ear is an integer multiple of λ0 / 2. (18) in (19) L current L represents the current distance from the sound-producing unit to the human ear. target Let k be the target distance from the sound-producing unit to the human ear, and k be a preset positive integer, taking the closest value to L. current The value; Such as L current When λ = 9.8m and λ0 = 0.68m, then L target =14×0.68 / 2=4.76m (closest when k=14), ΔL=4.76 9.8= 5.04m, meaning the sound unit needs to be retracted by 5.04m.
[0036] Step S5 specifically includes: steps S5.1 to S5.2; S5.1 Calculate the horizontal rotation angle θ (around the Z-axis) and vertical rotation angle of the pan-tilt unit based on the three-dimensional coordinates of the human ear (X,Y,Z) and the initial coordinates of the sound-generating unit (X0,Y0,Z0). (Around the X-axis), the horizontal rotation angle θ of the gimbal adjustment is along the horizontal direction pointing towards the ear, and the vertical rotation angle of the gimbal adjustment is... To point vertically towards the ear, (20) (twenty one); If the initial coordinates of the sound-producing unit are (X0=0, Y0=0, Z0=0), the coordinates of the human ear are ( Given 1.05, 0, 12.5), then θ = arctan( 1.05 / 12.5)≈ 4.78°, meaning the gimbal rotates 4.78° to the left. =arctan(0 / 12.5)=0°, meaning the gimbal does not need to be adjusted vertically; S5.2 After a preset interval Δt1 (in this embodiment, Δt1 = 10ms), repeat step S6.1 and compare the current horizontal rotation angle θ. current Horizontal rotation angle θ with respect to the target target Deviation Δθ, current vertical rotation angle current Vertical angle with the target target deviation Δ 1, (twenty two), (twenty three); If the horizontal deviation Δθ > 0.5° or the vertical deviation Δ If 1 > 0.5°, then drive the gimbal to adjust to the horizontal or vertical target rotation angle.
[0037] Step S6 specifically includes: steps S6.1 to S6.2; S6.1 Update the noise parameters obtained by the receiving unit, the three-dimensional coordinates (X,Y,Z) of the human ear in the visual positioning device, and the vibration parameters of the indoor environment every preset interval time Δt2. S6.2, Collect the residual noise sound pressure level (SPL) after noise reduction using a sound acquisition device placed near the human ear. res Then compared with the noise sound pressure level (SPL) before noise reduction. original The noise reduction amount ΔSPL is obtained by comparison. (twenty four), like (25) Where SPL is a preset noise reduction threshold, in this embodiment, such as SPL original =65dB, SPL res =62dB, then ΔSPL=3dB<5dB, Then, the process re-enters steps S1-S6, revising the noise parameters in step S1, the vibration parameters of the vibration source in step S4, and the positions of the human ear and the visual positioning device in step S5 to achieve dynamic active noise reduction throughout the room. Specifically, this revision can be achieved by adjusting the noise reduction amount ΔSPL according to a preset relationship between the vibration parameters of the vibration source, the noise parameters, and the distance between the human ear and the visual positioning device, and then re-verifying the result.
[0038] In this embodiment, some parameters in steps S1 to S6 are explained in Table 1 below.
[0039] Table 1
[0040] In this embodiment, the rotation of the binocular camera, the angle rotation of the gimbal, and the telescopic movement of the sound unit in steps S1 to S6 are all achieved by driving the corresponding motors through a control module located indoors.
[0041] The working principle of this invention is as follows: Noise parameters are obtained by preprocessing the noise using an outdoor noise receiving unit. These parameters include the dominant frequency and amplitude. The dominant frequency is the frequency corresponding to the maximum value in the noise frequency domain, thus enabling more accurate determination of the noise type and noise cancellation. Furthermore, the phase difference is determined based on the distance between the human ear and the visual positioning device, allowing for better coupling between the human ear and the outdoor noise sound waves at the intersection point. Simultaneously, vibration parameters in the indoor environment are detected, and opposite vibrations are generated based on these parameters to achieve vibration cancellation. This prevents indoor noise from being affected by outdoor noise, further ensuring the reliability of noise reduction. After noise reduction, the difference between the noise before and after reduction is compared to determine if the noise reduction requirements are met. Further adjustments are made to the positions of the human ear and visual positioning device, indoor environmental vibration parameters, and noise parameters to ensure active and closed-loop noise reduction processing.
Claims
1. A multidimensional noise reduction method based on combining noise cancellation and a classifier, implemented through a noise reduction system, characterized in that: The noise reduction system includes an outdoor recording unit, an indoor vibration source, a pan-tilt unit positioned along the path between the outdoor and indoor areas, a visual positioning device rotating on the pan-tilt unit, and a retractable sound-emitting unit; it includes the following steps: S1. Noise signals from different directions are collected by an outdoor receiver unit, and the noise signals are preprocessed to form preprocessed noise signals and noise parameters are obtained, including the dominant frequency and amplitude. S2. The position of the human ear is obtained in real time by a visual positioning device and classifier set up indoors, and the distance between the human ear and the visual positioning device is calculated. S3. Determine the phase compensation amount based on the distance between the human ear and the visual positioning device and the speed of sound. Determine the cancellation sound wave parameters based on the dominant frequency, amplitude, and phase compensation amount in the noise parameters. S4. Obtain the vibration parameters of the indoor environment, and then make the vibration source set up indoors generate anti-phase vibration according to the vibration parameters; S5. Adjust the gimbal posture and the extension / retraction of the sound unit in real time based on the relationship between the position of the human ear and the position of the sound unit, and generate a cancellation sound wave for noise reduction based on the cancellation sound wave parameters. S6. Obtain the noise reduction amount by the difference between the noise sound pressure level before noise reduction and the residual noise sound pressure level after noise reduction. Then, based on whether the noise reduction amount is within the preset threshold range, readjust the noise parameters in step S1, the vibration parameters of the vibration source in step S4, and the human ear position and visual positioning device in step S5.
2. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S5 further includes: extracting the features of the preprocessed noise signal and converting them into a quantifiable first feature vector, and then calculating the cosine similarity between the first feature vector and the second feature vector in the preset sound sample library. If the similarity exceeds a preset value, then no noise reduction processing is performed.
3. The multidimensional noise reduction method based on combining noise cancellation and a classifier as described in claim 1, characterized in that: Step S1 includes: steps S1.1 to S1.
3. S1.
1. Noise signals from different directions are collected through the radio unit, and the time domain signal of the noise signal is denoted as si(t), i=1,2,...,n,n≥3,t is time. Then, the discrete time domain data sequence si[k] is output through the radio unit, where k=1,2,...,N,N≥3,N is the number of sampling points; S1.2, Denoising via wavelet thresholding, threshold ,but (1), w j,k These are the wavelet decomposition coefficients. j,k Denoising coefficients These are preset parameters; Then, after filtering with a bandpass filter, a clean noise signal s is obtained. clean (t); S1.3, Processing the noise signal s clean (t) is subjected to a Discrete Fourier Transform, and then transformed to the frequency domain. (2), m=0,1,...,N 1. S[m] represents the frequency domain coefficients, j represents the imaginary unit, and k represents a value; Noise parameters are extracted, including the dominant frequency f0, amplitude A0, and wavelength λ0, where the dominant frequency f0 is the frequency corresponding to the maximum value of |S[m]|. (3), m0 is the frequency index corresponding to the maximum amplitude, f s The sampling frequency; (4), (5), v is the speed of sound in air at standard atmospheric pressure.
4. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: In step S5, "extracting the preprocessed noise signal features and converting them into a quantifiable first feature vector, then calculating the cosine similarity between the first feature vector and the second feature vector in a preset sound sample library; if the similarity exceeds a preset value, then no noise reduction processing is performed," includes steps S5.1 to S5.
2. S5.1, For the preprocessed noise signal s clean (t) Extract the Mel frequency cepstral coefficients and generate the first eigenvector X=[x1,x2,...,x 13 ]; S5.2, Combine the first feature vector X with the second feature vector Yi=[Y] in the preset sound sample library. 1,k ,Y 2,k ,...,Y 13,k Calculate the cosine similarity S, where i is the sample index, i=1,2,... (6), Then, it is determined whether the noise signal is exempted. If S≥T, then no cancellation sound wave is generated; otherwise, a cancellation sound wave is generated for noise reduction. T is a preset cosine similarity threshold.
5. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S4 includes: steps S4.1-S4.
2. S4.
1. Vibration acceleration signals a(t) of the indoor environment are collected by vibration sensors installed along the length and width directions of the indoor environment, where a(k) is the value at point k in a(t). Vibration parameters are extracted by discrete Fourier transform, including the dominant vibration frequency f. v With amplitude A v ,but (7), (8), m v This is the frequency index corresponding to the maximum value of |A[m]|. (9); S4.2 Calculate the driving voltage U(t) of the vibration source based on the vibration parameters, so that the vibration source produces vibration with the same frequency, matching amplitude, and opposite direction as the glass. (10), K is the voltage-vibration conversion coefficient of the vibration source voltage.
6. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S2 includes: steps S2.1 to S2.
3. S2.
1. Indoor images are acquired using a binocular camera. Then, the human ear region in the images is identified through Haar-like features and an Adaboost classifier. The pixel coordinates (u) of the human ear center in the left camera image are extracted. L ,v L ), pixel coordinates of the right camera (u R ,v R ); S2.2 Calculate the parallax d based on the pixel coordinates of the left and right cameras, and calculate the straight-line distance Z between the human ear and the camera based on the intrinsic parameters of the binocular camera. (11), p is the pixel size, and f1 is the focal length; (12); S2.
3. Combining the intrinsic parameter matrix K of the stereo camera, the pixel coordinates are converted into three-dimensional coordinates (X, Y, Z) in the camera coordinate system, where... (13), u0 and v0 are the pixel coordinates of the camera's principal point (if the image center is taken, then u0=640, v0=360). (14), (15)。 7. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S3 includes: steps S3.1 to S3.
2. S3.1 Generate an anti-phase cancelling sound wave based on the noise parameters, ensuring that the cancelling sound wave and the noise have the same vibration frequency, matching amplitude, and opposite phase. Let the time-domain signal of the cancelling sound wave be s. cancel (t), (16), Indicates opposite phase, Δ This is the phase compensation amount. (17), L is the distance from the sound-producing unit to the human ear, and v is the speed of sound in air under standard atmospheric pressure; S3.
2. Based on the noise wavelength λ0, adjust the extension / retraction amount ΔL of the sound unit so that the distance L from the sound unit to the ear is an integer multiple of λ0 / 2. (18), in (19) L current L represents the current distance from the sound-producing unit to the human ear. target Let k be the target distance from the sound-producing unit to the human ear, and k be a preset positive integer, taking the closest value to L. current The value of .
8. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S5 includes: steps S5.1 to S5.
2. S5.1 Calculate the horizontal and vertical rotation angles of the pan-tilt unit based on the three-dimensional coordinates of the human ear (X,Y,Z) and the initial coordinates of the sound-generating unit (X0,Y0,Z0). , (20), (21); S5.2 After a preset interval Δt1, repeat step S5.1 and compare the current horizontal rotation angle θ. current Horizontal rotation angle θ with respect to the target target Horizontal deviation Δθ, current vertical rotation angle current Vertical angle with the target target Vertical deviation Δ 1, (22), (23); If the horizontal deviation Δθ > 0.5° or the vertical deviation Δ If 1 > 0.5°, then drive the gimbal to adjust to the horizontal or vertical target rotation angle.
9. The multidimensional noise reduction method based on the combination of noise cancellation and classifier as described in claim 1, characterized in that: Step S6 includes: steps S6.1 to S6.
2. S6.1 Update the noise parameters obtained by the receiving unit, the three-dimensional coordinates (X,Y,Z) of the human ear in the visual positioning device, and the vibration parameters of the indoor environment every preset interval time Δt2. S6.2, Collect the residual noise sound pressure level (SPL) after noise reduction using a sound acquisition device placed near the human ear. res Then compared with the noise sound pressure level (SPL) before noise reduction. original The noise reduction amount ΔSPL is obtained by comparison. (24), like (25), Then re-enter S1, where SPL is the preset noise reduction threshold.
Citation Information
Patent Citations
Indoor noise reduction system, method, device and equipment and readable storage medium
CN115527517A