Noise reduction method, noise reduction device and law enforcement recorder
By using a multi-microphone recording device in the law enforcement recorder, noise reduction and sound source localization are performed based on background noise feature values, and noise feature values are dynamically updated, thus solving the problem of poor recording quality with a single microphone and achieving higher quality recording results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUNDAI TECH CO LTD
- Filing Date
- 2022-10-31
- Publication Date
- 2026-05-29
Smart Images

Figure CN115798493B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a noise reduction method, a noise reduction device, and a law enforcement recorder. Background Technology
[0002] A law enforcement recorder is an evidence-gathering device worn by law enforcement officers. It integrates functions such as real-time video and audio recording, photography, audio recording, and location tracking, and can digitally record the situation on the scene during the law enforcement process.
[0003] In related technologies, law enforcement recorders use a single microphone for recording. However, in practical applications, due to the influence of various complex application environment conditions, the voice signal collected by a single microphone will be mixed with noise, resulting in poor recording quality. Summary of the Invention
[0004] This invention provides a noise reduction method, a noise reduction device, and a law enforcement recorder to improve the recording quality of the law enforcement recorder.
[0005] This invention provides a noise reduction method applied to a law enforcement recorder, the law enforcement recorder including at least two voice acquisition devices, the noise reduction method comprising:
[0006] Acquire the sampling signals collected by each of the at least two voice acquisition devices;
[0007] The sampled signal is denoised based on the background noise feature value to obtain the speech signal corresponding to each of the at least two speech acquisition devices.
[0008] Based on the voice signals corresponding to each of the at least two voice acquisition devices, the sound source is located to obtain the sound source location information;
[0009] If a change in the sound source location information is detected, the background noise feature value is updated.
[0010] A noise reduction method provided by the present invention further includes:
[0011] The speech signal is enhanced based on the sound source location information to obtain the target speech signal;
[0012] Save the target speech signal.
[0013] According to a noise reduction method provided by the present invention, updating the background noise feature values includes:
[0014] Before the moment when the location information of the sound source changes, the sampling signal segments collected by each of the at least two voice acquisition devices within a first time period;
[0015] Extract background noise from the sampled signal segment;
[0016] Audio features are extracted from the background noise to obtain the target noise feature values;
[0017] The background noise feature value is updated to the target noise feature value.
[0018] According to a noise reduction method provided by the present invention, updating the background noise feature values includes:
[0019] When a change in the sound source location information is detected, the amount of change in the sound source location information is obtained;
[0020] The noise adjustment coefficient is determined based on the change.
[0021] The background noise feature value is adjusted based on the noise adjustment coefficient.
[0022] A noise reduction method provided by the present invention further includes:
[0023] Receive selection instructions for selecting the target scene;
[0024] The target scene is determined according to the selection instruction;
[0025] Obtain the noise feature value corresponding to the target scene to obtain the background noise feature value.
[0026] A noise reduction method provided by the present invention further includes:
[0027] After each recording is started, the initial sampling signals collected by each of the at least two voice acquisition devices are obtained during the second time period after the start time.
[0028] The audio features of the initial sampled signal are extracted to obtain the background noise feature values.
[0029] According to a noise reduction method provided by the present invention, the step of extracting audio features of the initial sampled signal to obtain the background noise feature values includes:
[0030] The time-domain amplitude of the initial sampled signal is obtained, and the time-domain amplitude is corrected based on the first correction coefficient to obtain the target time-domain feature value;
[0031] And / or, perform a Fourier transform on the initial sampled signal to obtain frequency domain feature values, and correct the frequency domain feature values based on the second correction coefficient to obtain target frequency domain feature values;
[0032] The background noise feature value includes at least one of the target time-domain feature value and the target frequency-domain feature value.
[0033] A noise reduction method provided by the present invention further includes:
[0034] In response to the detection of a correction factor adjustment command, a correction factor adjustment interface is displayed, the correction factor adjustment interface including a first correction factor adjustment control and a second correction factor adjustment control;
[0035] In response to an adjustment operation on the first correction factor adjustment control, the first correction factor is adjusted;
[0036] In response to an adjustment operation on the second correction factor adjustment control, the second correction factor is adjusted.
[0037] The present invention also provides a noise reduction device for use in a law enforcement recorder, the law enforcement recorder including at least two voice acquisition devices, the noise reduction device comprising:
[0038] The acquisition module is used to acquire the sampling signals collected by each of the at least two voice acquisition devices;
[0039] The noise reduction module is used to perform noise reduction processing on the sampled signal based on background noise feature values to obtain the speech signals corresponding to each of the at least two speech acquisition devices.
[0040] The positioning module is used to locate the sound source based on the voice signals corresponding to each of the at least two voice acquisition devices, and obtain the sound source location information.
[0041] An update module is used to update the background noise feature value when a change in the sound source location information is detected.
[0042] The present invention also provides a law enforcement recorder, including a memory, a processor, at least two voice acquisition devices connected to the processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the noise reduction method described above when executing the computer program.
[0043] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the noise reduction method as described above.
[0044] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the noise reduction method as described above.
[0045] The noise reduction method, device, and law enforcement recorder provided by this invention acquire sampled signals through at least two voice acquisition devices. Then, noise reduction processing is performed on the sampled signals based on background noise feature values to obtain the corresponding voice signals for each voice acquisition device. Next, sound source localization is performed based on the voice signals corresponding to each voice acquisition device to obtain sound source location information. When the sound source location information changes, the background noise feature values are updated. In this way, by performing noise reduction processing on at least two sampled signals from the law enforcement recorder based on background noise feature values, background noise in the sampled signals can be filtered out, improving the quality of the voice signals acquired by the voice acquisition devices, thereby improving the recording effect of the law enforcement recorder. Furthermore, the background noise feature values used for noise reduction processing of the sampled signals can be dynamically updated based on the sound source location information, adapting to environmental changes and further improving recording quality. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0047] Figure 1 This is one of the flowcharts illustrating the noise reduction method provided in this embodiment of the invention;
[0048] Figure 2 This is a second schematic flowchart of the noise reduction method provided in this embodiment of the invention;
[0049] Figure 3 This is a schematic diagram of the noise reduction device provided in an embodiment of the present invention;
[0050] Figure 4 This is a schematic diagram of the structure of the law enforcement recorder provided in an embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0052] It should be noted that the serial numbers assigned to the objects described in this invention, such as "first" and "second", are only used to distinguish the objects being described and do not have any sequential or technical meaning.
[0053] The following is combined with Figures 1-2 The noise reduction method for a law enforcement recorder according to the present invention is described. This noise reduction method can be applied to a law enforcement recorder or to a noise reduction device installed in the law enforcement recorder. The noise reduction device can be implemented through software, hardware, or a combination of both.
[0054] Figure 1 An exemplary flowchart of one of the noise reduction methods provided in this embodiment of the invention is shown below. Figure 1 As shown, the noise reduction method may include the following steps 110 to 140.
[0055] Step 110: Acquire the sampling signals collected by at least two voice acquisition devices.
[0056] In this embodiment of the invention, the law enforcement recorder may include at least two voice acquisition devices, such as a microphone or a pickup unit. The law enforcement recorder can acquire sampled signals through the at least two voice acquisition devices respectively.
[0057] Step 120: Perform noise reduction processing on the sampled signal based on the background noise feature value to obtain the speech signal corresponding to each of the at least two speech acquisition devices.
[0058] Background noise feature values can characterize the features of background noise. For example, each voice acquisition device of a law enforcement recorder can have its own corresponding background noise feature value. For the sampled signal acquired by each voice acquisition device, the background noise feature value corresponding to that voice acquisition device can be used to perform noise reduction processing on the sampled signal of that channel to obtain the voice signal corresponding to each voice acquisition device.
[0059] For example, a law enforcement recorder may include two microphones, such as microphone 1 and microphone 2. These two microphones can form a pickup point at a certain angle, such as two pickup points forming a 90° angle. The law enforcement recorder can use these two microphones to acquire sampling signals within their respective pickup ranges. Assuming that the background noise characteristic value corresponding to microphone 1 is E1 and the acquired sampling signal is S1, and the background noise characteristic value corresponding to microphone 2 is E2 and the acquired sampling signal is S2, then the sampling signal S1 can be denoised using the background noise characteristic value E1 to obtain the speech signal Y1, and the sampling signal S2 can be denoised using the background noise characteristic value E2 to obtain the speech signal Y2.
[0060] For example, the initial value of the background noise feature value can be determined each time recording is started. For instance, an initial sampling signal of a preset duration (e.g., 10 seconds) can be collected after recording begins, and the audio feature value extracted from this initial sampling signal can be used as the initial background noise feature value. Alternatively, each time recording is started, an initial background noise feature value can be selected for each voice acquisition device from a pre-set set of initial background noise feature values. For example, a correspondence between application scenarios and noise feature values can be pre-established, and after recording begins, the corresponding noise feature value can be matched from this correspondence based on the selected application scenario as the initial background noise feature value.
[0061] Step 130: Based on the voice signals corresponding to at least two voice acquisition devices, perform sound source localization to obtain sound source location information.
[0062] After obtaining the voice signal from each voice acquisition device through noise reduction processing, the sound source can be located based on the voice signal from each voice acquisition device to obtain the sound source location information. This sound source location information can reflect the direction and distance of the sound source point relative to the voice acquisition device, and the direction of arrival can include azimuth and pitch angles.
[0063] For example, a law enforcement recorder includes microphone 1 and microphone 2. After noise reduction processing of the sampled signals collected by microphones 1 and 2, voice signals Y1 and Y2 are obtained. Based on the time delay between voice signals Y1 and Y2, and the distance between microphones 1 and 2, the sound source can be located using acoustic localization principles to obtain the sound source location information. It is understood that this explanation uses a law enforcement recorder with two microphones as an example. However, for law enforcement recorders with three or more microphones, the sound source location information can also be determined based on the voice signals from each microphone using acoustic localization principles.
[0064] Step 140: Update the background noise feature value if a change in the sound source location information is detected.
[0065] After obtaining the sound source location information, it is determined whether the sound source location information has changed. If the sound source location information has changed, the background noise feature value is updated. For example, the background noise feature value can be adjusted according to the amount of change in the sound source location information; or, the target noise feature value can be determined based on the sampled signal before the sound source location information changed, and the background noise feature value can be updated to the target noise feature value.
[0066] Optionally, the background noise feature values can also be updated periodically based on a set time period.
[0067] In this way, by dynamically updating the background noise feature values based on the sound source location information, or periodically updating them based on a set time period, the background noise feature values can change with the environment, reflecting the characteristics of background noise that are closer to the current environment. This makes the noise reduction processing of the sampled signal based on the background noise feature values more effective, thereby improving the quality of the recorded speech.
[0068] The noise reduction method provided by this invention involves acquiring sampling signals through at least two voice acquisition devices, then performing noise reduction processing on the sampling signals based on background noise feature values to obtain the corresponding voice signals for each voice acquisition device. Next, sound source localization is performed based on the voice signals corresponding to each voice acquisition device to obtain sound source location information. When the sound source location information changes, the background noise feature values are updated. In this way, by performing noise reduction processing on at least two sampling signals from the law enforcement recorder based on background noise feature values, background noise in the sampling signals can be filtered out, improving the quality of the voice signals acquired by the voice acquisition devices and thus improving the recording effect of the law enforcement recorder. Furthermore, the background noise feature values used for noise reduction processing of the sampling signals can be dynamically updated based on the sound source location information, adapting to environmental changes and further improving recording quality.
[0069] based on Figure 1 In one example embodiment of the noise reduction method corresponding to the embodiment, after obtaining the sound source location information, the speech signal can also be enhanced based on the sound source location information to obtain the target speech signal.
[0070] Specifically, after obtaining the sound source location information, the sampling sector in which the sound source falls within the sampling range of each voice acquisition device can be determined based on the sound source location information. The voice signal in the sampling sector can then be enhanced to obtain the target voice signal.
[0071] For example, taking a law enforcement recorder including microphone 1 and microphone 2 as an example, the sampling signal from microphone 1 is processed by noise reduction to obtain speech signal Y1, and the sampling signal from microphone 2 is processed by noise reduction to obtain speech signal Y2. Assuming that the sampling range of microphone 1 can be divided into four consecutive sampling sectors A1, A2, A3, and A4, and the sampling range of microphone 2 can be divided into three consecutive sampling sectors B1, B2, and B3, if the sound source is determined to fall in sector A4 of microphone 1 and sector B2 of microphone 2 based on the sound source location information, then when fusing speech signals Y1 and Y2, the signal corresponding to sector A4 in speech signal Y1 and the signal corresponding to sector B2 in speech signal Y2 can be enhanced. This enhancement could be achieved by multiplying by an amplification factor greater than 1, or by multiplying by an amplification factor greater than that of other sampling sectors, ultimately obtaining the desired target speech signal.
[0072] In this way, by localizing the sound source and enhancing the speech, the speech signal from the direction of the sound source can be enhanced while the speech signal from other directions is suppressed. The resulting target speech signal reflects the characteristics of the sound source more effectively, reduces the influence of background noise, and further improves the quality of the speech recorded by the law enforcement recorder.
[0073] After obtaining the target voice signal, the target voice signal is saved to realize the recording function of the law enforcement recorder.
[0074] based on Figure 1 In one example embodiment of the noise reduction method corresponding to the embodiment, when a change in the sound source location information is detected, or when the time interval for updating the background noise feature value is greater than the time interval threshold, the target noise feature value can be determined according to the update rule; and the background noise feature value can be updated to the target noise feature value.
[0075] In one optional implementation, the background noise feature value can be updated using a sampled signal of a fixed duration prior to the change in the sound source location information. Specifically, updating the background noise feature value may include: acquiring sampled signal segments collected by at least two voice acquisition devices within a first time period prior to the change in the sound source location information; extracting background noise from the sampled signal segments; performing audio feature extraction on the background noise to obtain a target noise feature value; and updating the background noise feature value to the target noise feature value.
[0076] For example, a law enforcement recorder includes microphone 1 and microphone 2. At time t1, a change in the location of the sound source is detected. At this time, the recorder can acquire the sampling signal segment ΔS1 collected by microphone 1 and the sampling signal segment ΔS2 collected by microphone 2 during a time period of T before time t1. Then, background noise ΔC1 is extracted from sampling signal segment ΔS1, and background noise ΔC2 is extracted from sampling signal segment ΔS2. The audio features of background noise ΔC1 and background noise ΔC2 are then extracted respectively to obtain the target noise feature value ΔE1 of microphone 1 and the target noise feature value ΔE2 of microphone 2. Afterwards, the background noise feature value of microphone 1 can be updated to ΔE1, and the background noise feature value of microphone 2 can be updated to ΔE2, thus updating the background noise feature values. In this way, the background noise feature values can be updated in real time according to the change in the sound source location, allowing the background noise feature values to reflect the characteristics of background noise closer to the current environment, improving the noise reduction effect of the sampled signals.
[0077] Extracting background noise from the sampled signal segment can be achieved by removing the speech signal segment obtained through noise reduction processing from the sampled signal segment. For example, for microphone 1, the speech signal ΔY1 obtained through noise reduction processing within time period T can be acquired, and the speech signal ΔY1 can be removed from the sampled signal segment ΔS1 to obtain the background noise ΔC1.
[0078] In another alternative implementation, the background noise feature value can be adjusted based on the change in the sound source location. Specifically, updating the background noise feature value may include: acquiring the change in sound source location information when a change is detected; determining a noise adjustment coefficient based on the change; and adjusting the background noise feature value based on the noise adjustment coefficient.
[0079] For example, a noise adjustment mapping table can be pre-established, storing the mapping relationship between changes in sound source location information and noise adjustment coefficients. When a sound source is detected to change from position P1 to position P2, the change in sound source location information ΔP can be calculated based on P1 and P2. Then, based on the change ΔP, the noise adjustment coefficient is matched from the noise adjustment mapping table. The target noise feature value is calculated using the pre-established functional relationship between the noise adjustment coefficient and the background noise feature value. For example, the current background noise feature value can be multiplied by the matched noise adjustment coefficient to obtain the target noise feature value. Then, the current background noise feature value is updated to the target noise feature value, thus achieving the adjustment of the background noise feature value.
[0080] In one example embodiment, the background noise feature value can also be updated actively by the user. Specifically, the noise reduction method provided in this embodiment of the invention may further include: upon detecting an update instruction to update the background noise feature value, receiving an adjustment operation on a noise feature value adjustment button; and determining a target noise feature value based on the adjustment operation.
[0081] For example, a law enforcement recorder can be equipped with a physical noise update button and a noise feature value adjustment button. The noise update button can be used to trigger the background noise feature value update function, and the noise feature value adjustment button is used to adjust the magnitude of the background noise feature value. For instance, the noise feature value adjustment button can be a rotary knob, with different rotation positions corresponding to different adjustment coefficients. When the noise update button is triggered, the user can select the adjustment coefficient using the rotary knob. The law enforcement recorder obtains the current rotation position information of the rotary knob based on the user's adjustment operation on the noise feature value adjustment button, determines the adjustment coefficient based on this rotation position information, and determines the target noise feature value based on the adjustment coefficient and the current background noise feature value. For instance, the noise feature value adjustment button can also be an adjustment control displayed on the law enforcement recorder's interface. The user can increase or decrease the adjustment coefficient using this adjustment control. The law enforcement recorder determines the adjustment coefficient based on the user's adjustment operation on the adjustment control, and determines the target noise feature value based on the adjustment coefficient and the current background noise feature value.
[0082] based on Figure 1 In one example embodiment of the noise reduction method corresponding to the embodiments, the noise reduction method may further include a step of determining an initial value for the background noise feature value. In an optional implementation, the user may set the initial value for the background noise feature value according to the specific application scenario of the law enforcement recorder, and this initial value serves as the background noise feature value when the law enforcement recorder begins recording. Specifically, the noise reduction method may further include: receiving a selection instruction for selecting a target scene; determining the target scene according to the selection instruction; and obtaining the noise feature value corresponding to the target scene to obtain the target noise feature value.
[0083] For example, before each recording session, the user can select a target scene using the target scene selection button on the body camera, based on the application scenario. This target scene selection button can be a physical button on the body camera or a target scene selection control displayed on the interface. The selectable scenes provided by the body camera may include, but are not limited to, at least one of the following: outdoor scenes, indoor scenes, noisy environment scenes, and quiet environment scenes. The body camera receives a selection command for choosing the target scene. For example, if the command indicates an indoor scene, the target scene is determined to be an indoor scene. Then, the noise characteristic value corresponding to the indoor scene can be found in a scene-to-noise characteristic value mapping table. After recording begins, the body camera can use this noise characteristic value as the background noise characteristic value to perform noise reduction processing on the collected sampled signal. In this way, appropriate background noise characteristic values can be selected according to different usage scenarios of the body camera to achieve better noise reduction processing of the sampled signal.
[0084] In another optional implementation, the law enforcement recorder can also automatically determine the background noise feature value based on the current usage environment. Specifically, the noise reduction method provided in this embodiment of the invention may further include: after each recording is started, acquiring the initial sampling signals collected by at least two voice acquisition devices within a second time period after the start time; extracting the audio features of the initial sampling signals to obtain the background noise feature value.
[0085] For example, taking a law enforcement recorder that includes two microphones, microphone 1 and microphone 2, as an example, each time recording is started, the law enforcement recorder can acquire the initial sampling signals collected by microphone 1 and microphone 2 within the initial time period t2, and obtain the initial sampling signal S1. t2 and the initial sampled signal S2 t2 Then extract the initial sampling signal S1 t2 The audio characteristics are analyzed to obtain the background noise feature value corresponding to microphone 1, and the initial sampling signal S2 is extracted. t2 The audio characteristics of microphone 1 and microphone 2 are used to obtain the background noise feature value corresponding to microphone 2. During subsequent recording, the background noise feature values of microphone 1 and microphone 2 can be dynamically updated based on noise update conditions, such as changes in the sound source location information or time intervals for updating background noise feature values exceeding a time interval threshold.
[0086] In this way, by using the audio characteristics of the sampled signal during the second time period after the law enforcement recorder starts recording as the background noise feature value, the background noise that conforms to the current application scenario can be automatically obtained. This method is highly intelligent, and the obtained background noise feature value can better reflect the characteristics of the background noise in the current environment, thereby improving the noise reduction effect of the sampled signal.
[0087] For example, the background noise feature value can be an audio feature in the time domain, an audio feature in the frequency domain, or an audio feature in both the time and frequency domains. For example, the background noise feature value can be the audio feature value of the extracted initial sampled signal, or it can be obtained by correcting the extracted audio feature value, such as by multiplying it by a corresponding correction coefficient. Specifically, extracting the audio features of the initial sampled signal to obtain the background noise feature value can include: obtaining the time domain amplitude of the initial sampled signal and correcting the time domain amplitude based on a first correction coefficient to obtain a target time domain feature value; and / or, performing a Fourier transform on the initial sampled signal to obtain a frequency domain feature value, and correcting the frequency domain feature value based on a second correction coefficient to obtain a target frequency domain feature value; wherein the background noise feature value includes at least one of the target time domain feature value and the target frequency domain feature value.
[0088] For example, the first correction factor and the second correction factor can be adjusted by the user. Specifically, the noise reduction method may further include: displaying a correction factor adjustment interface in response to detecting a correction factor adjustment command, the correction factor adjustment interface including a first correction factor adjustment control and a second correction factor adjustment control; adjusting the first correction factor in response to an adjustment operation on the first correction factor adjustment control; and adjusting the second correction factor in response to an adjustment operation on the second correction factor adjustment control.
[0089] For example, a law enforcement recorder can be equipped with an activation button for adjusting the correction coefficient. When this button is triggered, a correction coefficient adjustment interface is displayed. This interface provides a first correction coefficient adjustment control and a second correction coefficient adjustment control. The user can adjust the first correction coefficient using the first control and the second correction coefficient using the second control. For instance, the first correction coefficient adjustment control can indicate the adjustment range of the first correction coefficient, and the second correction coefficient adjustment control can indicate the adjustment range of the second correction coefficient. Thus, by adjusting the first and / or second correction coefficients, the initial value of the background noise characteristic can be adjusted. For example, the user can adjust it according to the current operating environment of the law enforcement recorder to adapt to the environment; in a noisy environment, the correction coefficient can be appropriately increased, and in a quiet environment, it can be appropriately decreased.
[0090] Based on the noise reduction methods of the above embodiments, the following example, using a law enforcement recorder including two microphones, microphone 1 and microphone 2, will be used to further illustrate the noise reduction method provided by the embodiments of the present invention.
[0091] Figure 2 This is an exemplary schematic diagram of the noise reduction method provided in an embodiment of the present invention, with reference to... Figure 2 As shown, the noise reduction method may include the following steps 201 to 210.
[0092] Step 201: Acquire the initial sampling signal for a first fixed duration.
[0093] Each time recording begins, the body camera can acquire the initial sampling signals collected by microphones 1 and 2 within the first fixed duration t2, and then convert the obtained initial sampling signal S1... t2 and the initial sampled signal S2 t2 As background noise signal.
[0094] Step 202: Extract the audio features of the initial sampled signal to obtain the background noise feature values.
[0095] For the initial sampled signal S1 t2 and the initial sampled signal S2t2 Audio features are extracted separately to obtain the background noise feature value E1 corresponding to microphone 1 and the background noise feature value E2 corresponding to microphone 2. Then, E1 and E2 are saved, for example, by using the first variable S. 01 Second variable S 02 Record the current background noise characteristic values of the two microphones, and assign E1 to the first variable S. 01 Assign E2 to the second variable S 02 After that, normal recording began.
[0096] Step 203: Acquire the sampling signals collected by each of the two microphones.
[0097] The law enforcement recorder collects sampling signals from the surrounding environment through microphone 1 and microphone 2, obtaining sampling signals S1 and S2, which are used as the raw audio signals of the law enforcement recorder.
[0098] Step 204: Perform noise reduction processing on the sampled signal to obtain the speech signal.
[0099] The law enforcement recorder can subtract the first variable S from the sampled signal S1. 01 The background noise feature values stored in the sampled signal S2 are used to obtain the noise-reduced speech signal Y1; the second variable S is subtracted from the sampled signal S2. 02 The background noise feature values stored in the data are used to obtain the noise-reduced speech signal Y2.
[0100] Step 205: Based on the speech signal, perform sound source localization and speech enhancement to obtain the target speech signal.
[0101] The law enforcement recorder can locate the sound source using acoustic localization principles based on the time delay between voice signals Y1 and Y2, as well as the distance between microphones 1 and 2, thus obtaining the sound source location information. Then, based on this location information, voice signals Y1 and Y2 are enhanced, and during this enhancement process, they are fused to obtain the target voice signal. Through sound source localization and voice enhancement, secondary noise reduction can be applied to the sampled signals acquired by the microphones, further improving the quality of the recorded speech.
[0102] Step 206: Save the target speech signal.
[0103] Step 207: Determine if the location of the sound source has changed. If it has changed, proceed to step 208; otherwise, continue with step 203.
[0104] For example, if the detected change in the sound source location information exceeds a change threshold, it can be determined that the sound source location has changed. For instance, the sound source location can be checked at set time intervals to determine if it has changed.
[0105] In an optional implementation, it can also be determined whether the time interval since the last update of the background noise feature value is greater than the time interval threshold. If the sound source location changes or the time interval since the last update of the background noise feature value is greater than the time interval threshold, then step 208 is executed; otherwise, step 203 is executed.
[0106] Step 208: Obtain the sampling signal segment of the dual microphones within the second fixed time period before the change time.
[0107] For example, if the law enforcement recorder detects a change in the location information of the sound source at time t1, it can obtain the sampling signal segment ΔS1 collected by microphone 1 and the sampling signal segment ΔS2 collected by microphone 2 during a time period of the second fixed duration T before time t1.
[0108] Step 209: Extract the audio features of the background noise in the sampled signal segment to obtain the target noise feature value.
[0109] For example, a law enforcement recorder can acquire the voice signal ΔY1 of voice signal Y1 within a time period T, and the voice signal ΔY2 of voice signal Y2 within a time period T. Then, it subtracts ΔY1 from the sampled signal segment ΔS1 and ΔY2 from the sampled signal segment ΔS2 to obtain the background noise ΔC1 collected by microphone 1 and the background noise ΔC2 collected by microphone 2 within the time period T. The audio features of the background noise ΔC1 and background noise ΔC2 can then be extracted to obtain the target noise feature value E3 corresponding to microphone 1 and the target noise feature value E4 corresponding to microphone 2.
[0110] In this example embodiment, steps 208 to 209 determine the target noise characteristic value through historical sampling signal segments. In an optional embodiment, the law enforcement recorder can also redetermine the target noise characteristic value using the method of steps 201 to 202.
[0111] Step 210: Update the background noise feature values to the target noise feature values.
[0112] After obtaining the target noise feature value E3 corresponding to microphone 1 and the target noise feature value E4 corresponding to microphone 2, the law enforcement recorder can transfer the first variable S 01 Update the value of the second variable S to E3. 02 The value is updated to E4, thus updating the background noise feature value. Then, step 203 is executed.
[0113] The noise reduction method provided in this invention can utilize two microphones to achieve the recording function of a law enforcement recorder. On one hand, during recording, noise reduction processing can be performed on the sampled signal acquired by the microphones based on background noise feature values. This can initially filter out background noise in the sampled signal and improve its quality. Then, the resulting speech signal after noise reduction can be subjected to sound source localization and speech enhancement, achieving secondary noise reduction and further improving the quality of the recorded speech. On the other hand, the background noise feature values can be updated when the sound source location changes or when the time interval between updates exceeds a threshold. This adapts to environmental changes, obtaining background noise feature values that reflect the characteristics of the current environment's background noise, improving the noise reduction effect based on background noise feature values, and thus further improving the quality of the recorded speech.
[0114] The noise reduction device provided by the present invention will be described below. The noise reduction device described below can be referred to in correspondence with the noise reduction method described above. This noise reduction device can be applied to a law enforcement recorder, which includes at least two voice acquisition devices, wherein the voice acquisition devices can be microphones or pickups, etc.
[0115] Figure 3 An exemplary schematic diagram of the noise reduction device provided in an embodiment of the present invention is shown, with reference to... Figure 3 As shown, the noise reduction device 300 may include an acquisition module 310, a noise reduction module 320, a positioning module 330, and an update module 340. Specifically: the acquisition module 310 can acquire sampling signals collected by at least two voice acquisition devices; the noise reduction module 320 can perform noise reduction processing on the sampling signals acquired by the acquisition module 310 based on background noise feature values to obtain voice signals corresponding to each of the at least two voice acquisition devices; the positioning module 330 can perform sound source localization based on the voice signals corresponding to each of the at least two voice acquisition devices obtained by the noise reduction module 320 to obtain sound source location information; and the update module 340 can update the background noise feature values when a change in the sound source location information is detected.
[0116] In one example embodiment, the noise reduction device 300 may further include: an enhancement module, configured to perform speech enhancement processing on the speech signal based on the sound source location information obtained by the positioning module 330 to obtain a target speech signal. Exemplarily, the noise reduction device 300 may further include: a storage module, configured to store the target speech signal obtained by the enhancement module.
[0117] In one example embodiment, the update module 340 may include: a first acquisition unit, configured to acquire, upon detecting a change in the sound source location information, sampled signal segments acquired by at least two voice acquisition devices within a first time period prior to the change in the sound source location information; a first extraction unit, configured to extract background noise from the sampled signal segments; a second extraction unit, configured to extract audio features from the background noise to obtain target noise feature values; and an update unit, configured to update the background noise feature values to target noise feature values.
[0118] In one example embodiment, the update module 340 may include: a second acquisition unit, configured to acquire the amount of change in the sound source location information when a change in the sound source location information is detected; a first determination unit, configured to determine a noise adjustment coefficient based on the amount of change acquired by the second acquisition unit; and a first adjustment unit, configured to adjust the background noise feature value based on the noise adjustment coefficient.
[0119] In one example embodiment, the noise reduction device 300 may further include: a receiving module, configured to receive an adjustment operation on a noise feature value adjustment button when an update instruction for updating a background noise feature value is detected; and a first determining module, configured to determine a target noise feature value based on the adjustment operation received by the receiving module, and update the background noise feature value to the target noise feature value.
[0120] In one example embodiment, the noise reduction device 300 may further include a second determining module, which may be used to: receive a selection instruction for selecting a target scene; determine the target scene according to the selection instruction; obtain noise feature values corresponding to the target scene, and obtain background noise feature values.
[0121] In one example embodiment, the noise reduction device 300 may further include an extraction module. Accordingly, the acquisition module 310 may also be used to acquire the initial sampling signals collected by at least two voice acquisition devices during a second time period after the start of recording each time recording is started; the extraction module may be used to extract the audio features of the initial sampling signals to obtain background noise feature values.
[0122] In one example embodiment, the extraction module may include a third extraction unit and / or a fourth extraction unit. The third extraction unit may be used to acquire the time-domain amplitude of the initial sampled signal and correct the time-domain amplitude based on a first correction coefficient to obtain a target time-domain feature value. The fourth extraction unit may be used to perform a Fourier transform on the initial sampled signal to obtain a frequency-domain feature value and correct the frequency-domain feature value based on a second correction coefficient to obtain a target frequency-domain feature value. The background noise feature value includes at least one of the target time-domain feature value and the target frequency-domain feature value.
[0123] In one example embodiment, the extraction module may further include: a display unit, configured to display a correction coefficient adjustment interface in response to detecting a correction coefficient adjustment command, the correction coefficient adjustment interface including a first correction coefficient adjustment control and a second correction coefficient adjustment control; a second adjustment unit, configured to adjust a first correction coefficient in response to an adjustment operation directed to the first correction coefficient adjustment control; and a third adjustment unit, configured to adjust a second correction coefficient in response to an adjustment operation directed to the second correction coefficient adjustment control.
[0124] Figure 4 An example is a schematic diagram of the structure of a law enforcement recorder, such as... Figure 4 As shown, the law enforcement recorder may include: a processor 410, and at least two voice acquisition devices 450 connected to the processor 410. Figure 4 The example uses two voice acquisition devices, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 can communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute the noise reduction methods provided in the above embodiments. These methods may include, for example, acquiring sampled signals acquired by at least two voice acquisition devices; performing noise reduction processing on the sampled signals based on background noise feature values to obtain voice signals corresponding to each of the at least two voice acquisition devices; performing sound source localization based on the voice signals corresponding to each of the at least two voice acquisition devices to obtain sound source location information; and updating the background noise feature values when a change in the sound source location information is detected.
[0125] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0126] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the noise reduction method provided in the above-described method embodiments. The method may include, for example, acquiring sampling signals acquired by at least two voice acquisition devices; performing noise reduction processing on the sampling signals based on background noise feature values to obtain voice signals corresponding to each of the at least two voice acquisition devices; performing sound source localization based on the voice signals corresponding to each of the at least two voice acquisition devices to obtain sound source location information; and updating the background noise feature values when a change in the sound source location information is detected.
[0127] In another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program is implemented to perform the noise reduction methods provided in the above-described method embodiments. The method may include, for example,: acquiring sampling signals acquired by at least two voice acquisition devices; performing noise reduction processing on the sampling signals based on background noise feature values to obtain voice signals corresponding to each of the at least two voice acquisition devices; performing sound source localization based on the voice signals corresponding to each of the at least two voice acquisition devices to obtain sound source location information; and updating the background noise feature values when a change in the sound source location information is detected.
[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0129] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A noise reduction method, characterized in that, Applied to law enforcement recorders, the law enforcement recorders include at least two voice acquisition devices, and the noise reduction method includes: Acquire the sampling signals collected by each of the at least two voice acquisition devices; The sampling signal is denoised based on the background noise feature value to obtain the speech signal corresponding to each of the at least two speech acquisition devices. Sound source localization is performed based on the speech signals corresponding to each of the at least two speech acquisition devices to obtain sound source location information; after obtaining the sound source location information, the sampling sector in which the sound source falls within the sampling range of each speech acquisition device is determined according to the sound source location information, the speech signal of the sampling sector is enhanced to obtain the target speech signal and the target speech signal is saved; wherein, the target speech signal reflects the characteristics of the sound source; If a change in the sound source location information is detected, the background noise feature value is updated.
2. The noise reduction method according to claim 1, characterized in that, Updating the background noise feature values includes: Before the moment when the location information of the sound source changes, the sampling signal segments collected by each of the at least two voice acquisition devices within a first time period; Extract background noise from the sampled signal segment; Audio features are extracted from the background noise to obtain the target noise feature values; The background noise feature value is updated to the target noise feature value.
3. The noise reduction method according to claim 1, characterized in that, Updating the background noise feature values includes: Obtain the change in the location information of the sound source; The noise adjustment coefficient is determined based on the change. The background noise feature value is adjusted based on the noise adjustment coefficient.
4. The noise reduction method according to any one of claims 1 to 3, characterized in that, Also includes: Receive selection instructions for selecting the target scene; The target scene is determined according to the selection instruction; Obtain the noise feature value corresponding to the target scene to obtain the background noise feature value.
5. The noise reduction method according to any one of claims 1 to 3, characterized in that, Also includes: After each recording is started, the initial sampling signals collected by each of the at least two voice acquisition devices are obtained during the second time period after the start time. The audio features of the initial sampled signal are extracted to obtain the background noise feature values.
6. The noise reduction method according to claim 5, characterized in that, The step of extracting the audio features of the initial sampled signal to obtain the background noise feature values includes: The time-domain amplitude of the initial sampled signal is obtained, and the time-domain amplitude is corrected based on the first correction coefficient to obtain the target time-domain feature value; And / or, perform a Fourier transform on the initial sampled signal to obtain frequency domain feature values, and correct the frequency domain feature values based on the second correction coefficient to obtain target frequency domain feature values; The background noise feature value includes at least one of the target time-domain feature value and the target frequency-domain feature value.
7. The noise reduction method according to claim 6, characterized in that, Also includes: In response to the detection of a correction factor adjustment command, a correction factor adjustment interface is displayed, the correction factor adjustment interface including a first correction factor adjustment control and a second correction factor adjustment control; In response to an adjustment operation on the first correction factor adjustment control, the first correction factor is adjusted; In response to an adjustment operation on the second correction factor adjustment control, the second correction factor is adjusted.
8. A noise reduction device, characterized in that, Applied to law enforcement recorders, the law enforcement recorders include at least two voice acquisition devices, and the noise reduction device includes: The acquisition module is used to acquire the sampling signals collected by each of the at least two voice acquisition devices; The noise reduction module is used to perform noise reduction processing on the sampled signal based on background noise feature values to obtain the speech signals corresponding to each of the at least two speech acquisition devices. A positioning module is used to locate the sound source based on the speech signals corresponding to each of the at least two speech acquisition devices to obtain sound source location information; after obtaining the sound source location information, it determines the sampling sector in which the sound source falls within the sampling range of each speech acquisition device according to the sound source location information, enhances the speech signal of the sampling sector to obtain a target speech signal, and saves the target speech signal; wherein, the target speech signal reflects the characteristics of the sound source; An update module is used to update the background noise feature value when a change in the sound source location information is detected.
9. A law enforcement recorder, characterized in that, The device includes a memory, a processor, at least two voice acquisition devices connected to the processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, it implements the noise reduction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the noise reduction method as described in any one of claims 1 to 7.