A sound signal processing method, apparatus, medium, and device

By determining the target filtering frequency band and intensity threshold of the noise signal, the original speech signal in the vehicle passenger compartment is filtered and suppressed, which solves the problem of speech recognition accuracy in noisy environments and improves the effect of speech recognition.

CN119741931BActive Publication Date: 2026-01-13VOYAH AUTOMOBILE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411742149.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2026-01-13
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In noisy environments, machines struggle to recognize speech signals from raw sound signals, leading to a decrease in speech recognition accuracy.

Method used

By acquiring the target noise type and intensity of the noise signal, and using a preset correspondence to determine the target filtering frequency band and intensity threshold, the original speech signal is filtered and suppressed to separate the user's speech signal.

Benefits of technology

It effectively reduces the impact of noise signals on speech recognition, improving the accuracy of speech recognition and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741931B_ABST
    Figure CN119741931B_ABST
Patent Text Reader

Abstract

The application provides a sound signal processing method, device, medium and equipment, the method comprises: acquiring the original voice signal collected in the passenger cabin of the vehicle, the original voice signal is a user voice signal with noise signal; acquiring the target noise type of the noise signal, determining the target filtering frequency band corresponding to the target noise type based on the preset first corresponding relationship; acquiring the target intensity of the noise signal, determining the target intensity threshold corresponding to the target intensity based on the preset second corresponding relationship; filtering the noise signal in the original voice signal based on the target filtering frequency band, and suppressing the noise signal in the original voice signal based on the target intensity threshold, to obtain the user voice signal. The application can reduce the noise signal in the original voice signal, reduce the influence of the noise signal on voice recognition, and improve the accuracy of voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer interaction, and particularly relates to a sound signal processing method and device, a medium and equipment. BACKGROUND

[0002] In the process of human-computer interaction, voice control is a relatively common voice control method. A user can issue a voice signal, and a machine identifies the voice signal and executes a corresponding function according to an instruction in the voice signal after receiving the voice signal. However, when the environment in which the machine is located is relatively noisy, the original sound signal received by the machine contains more noise, which can easily cause the machine to fail to identify the voice signal in the original sound signal. SUMMARY

[0003] Therefore, the present application provides a sound signal processing method, device, medium and equipment, which mainly aims to reduce the influence of noise signals on voice recognition and improve the accuracy of voice recognition in the process of voice recognition.

[0004] To achieve the above-mentioned purpose, the first aspect of the present application discloses a sound signal processing method, comprising:

[0005] obtaining an original voice signal collected in a passenger compartment of a vehicle, the original voice signal being a user voice signal containing noise signals;

[0006] obtaining a target noise type of the noise signals, determining a target filtering frequency band corresponding to the target noise type based on a preset first correspondence relationship;

[0007] obtaining a target intensity of the noise signals, determining a target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship;

[0008] filtering the noise signals in the original voice signal based on the target filtering frequency band, and suppressing the noise signals in the original voice signal based on the target intensity threshold, to obtain the user voice signal.

[0009] Optionally, the obtaining of the target noise type of the noise signals and the determination of the target filtering frequency band corresponding to the target noise type based on the preset first correspondence relationship comprises:

[0010] converting the original voice signal into original frequency spectrum data;

[0011] using a noise recognition model to identify a noise type of at least one noise signal in the original frequency spectrum data;

[0012] determine a target filtering frequency band corresponding to the target noise type frequency based on a preset first correspondence relationship, the target noise type being at least one of the noise types.

[0013] Optionally, the target intensity threshold value corresponding to the target intensity is determined based on a preset second correspondence relationship, and the method comprises:

[0014] an target intensity change value of the noise signal is obtained, the target intensity change value being used to represent a change of the target intensity;

[0015] a dynamic intensity threshold value is obtained by adjusting the target intensity threshold value corresponding to the target intensity based on a preset second correspondence relationship, the dynamic intensity threshold value representing a change of the target intensity threshold value, and the dynamic intensity threshold value corresponding to the target intensity change value.

[0016] Optionally, the original voice signal collected in the passenger compartment of the vehicle comprises:

[0017] a first sound signal in the vehicle cabin is collected by using a vehicle microphone array in the passenger compartment of the vehicle, the first sound signal being an original voice signal containing media sound played by a vehicle loudspeaker;

[0018] sound data for playing through the vehicle loudspeaker is obtained.

[0019] the media sound in the first sound signal is eliminated by using an echo cancellation algorithm, with reference to the sound data, to obtain the original sound signal.

[0020] Optionally, when the user voice signal is used to wake up the vehicle voice assistant, the noise signal in the original voice signal is filtered based on the target filtering frequency band, and the noise signal in the original voice signal is suppressed based on the target intensity threshold value, to obtain the user voice signal, which comprises:

[0021] a target filtering frequency band change value and a target intensity change value of the noise signal in the wake-up time period of the vehicle voice assistant are obtained, the target filtering frequency band change value being used to represent a change of the target filtering frequency band, and the target intensity change value being used to represent a change of the target intensity;

[0022] the noise signal in the original voice signal is filtered based on the target filtering frequency band change value, and the noise signal in the original voice signal is suppressed based on the target intensity change value, to obtain the user voice signal.

[0023] Optionally, when the user voice signal is used to wake up the vehicle voice assistant, the method further comprises:

[0024] The wake-up sensitivity of the in-vehicle voice assistant is adjusted according to the target intensity of the noise signal.

[0025] Optionally, after filtering the noise signal in the original speech signal based on the target filtering frequency band and suppressing the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal, the method further includes:

[0026] Determine the human voice frequency range to which the user's voice signal belongs;

[0027] The user's voice signal in the specified human voice frequency range is enhanced to obtain the enhanced user's voice signal.

[0028] The second aspect of this application discloses a sound signal processing apparatus, comprising:

[0029] The acquisition module is used to acquire the original voice signal collected in the passenger compartment of the vehicle, wherein the original voice signal is a user voice signal with noise signal;

[0030] The first determining module is used to obtain the target noise type of the noise signal and determine the target filtering frequency band corresponding to the target noise type based on a preset first correspondence relationship.

[0031] The second determining module is used to obtain the target intensity of the noise signal and determine the target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship.

[0032] The processing module is used to filter the noise signal in the original speech signal based on the target filtering frequency band, and suppress the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal.

[0033] A third aspect of this application provides an electronic device, comprising:

[0034] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform any of the methods disclosed in the first aspect.

[0035] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of the first aspect.

[0036] A fifth aspect of this application provides a vehicle in which the device as described in the second aspect or the electronic device as described in the third aspect is mounted.

[0037] In summary, according to the technical solution disclosed in this application, the problem of noise signals in the original speech signal easily causing the machine to be unable to recognize the speech signal in the original sound signal is addressed. This application provides a sound signal processing method, which first acquires the original speech signal collected in the passenger compartment of a vehicle, the original speech signal being a user speech signal containing noise; then, it acquires the target noise type of the noise signal, and determines the target filtering frequency band corresponding to the target noise type based on a preset first correspondence; then, it acquires the target intensity of the noise signal, and determines the target intensity threshold corresponding to the target intensity based on a preset second correspondence; finally, it filters the noise signal in the original speech signal based on the target filtering frequency band, and suppresses the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal. In the technical solution of this application, the original speech signal is acquired in the passenger compartment of a vehicle. When the original speech signal contains noise signals, in order to separate the user speech signal from the original speech signal, this application can determine the target noise type and target intensity of the noise signal. The target filtering frequency band corresponding to the target noise type is determined through a preset first correspondence, and the target intensity threshold corresponding to the target intensity value is determined through a preset second correspondence. The noise signal at the target filtering frequency in the original speech signal is filtered. At the same time, the noise signal in the original speech signal is suppressed by the target intensity threshold, thereby reducing the noise signal in the original speech signal, reducing the impact of the noise signal on speech recognition, and improving the accuracy of speech recognition.

[0038] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1A flowchart of a sound signal processing method provided in an embodiment of this application is shown;

[0042] Figure 2 A structural diagram of a sound signal processing device provided in an embodiment of this application is shown. Detailed Implementation

[0043] To better understand the technical solutions provided in the embodiments of this specification, the technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.

[0044] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The term "two or more" includes two or more cases.

[0045] In the process of human-computer interaction recognition, users can issue control commands in the form of voice signals. The machine can receive the voice signals, recognize the control commands from them, and execute the corresponding functions accordingly. For example, when waking up the machine, a user can speak a corresponding wake-up word. The spoken wake-up word is received by the machine as a voice signal. The machine extracts the wake-up word from the voice signal and compares it with preset wake-up words. If the comparison is successful, the machine can be woken up. However, if the voice signal received by the machine contains strong noise, the machine will have difficulty extracting the wake-up word, reducing the probability of waking up and severely impacting the user experience.

[0046] To address the aforementioned problems, this embodiment provides a sound signal processing method that can be executed in the vehicle's control system, such as... Figure 1 As shown, it includes:

[0047] Step 101: Acquire the raw voice signal collected in the passenger compartment of the vehicle. The raw voice signal is the user voice signal with noise.

[0048] The passenger compartment of a vehicle refers to the interior space of the vehicle, specifically the area where passengers and the driver reside. During the acquisition of raw voice signals, sound signals from within the passenger compartment can be collected using microphones or other audio devices. The acquired raw voice signals, in addition to the user's voice signal requiring further recognition, also include noise signals, which can interfere with the recognition of the user's voice signal. Generally speaking, all signals in the raw voice signal other than the user's voice signal can be considered noise signals. These noise signals can include: environmental noise (e.g., wind noise, tire friction, noise from other vehicles); mechanical noise (e.g., engine noise, air conditioning system noise); and electronic device noise (e.g., static from car audio systems, mobile phones, and other electronic devices).

[0049] In some embodiments, the specific process of acquiring the raw speech signal can be as follows:

[0050] The system uses an onboard microphone array in the passenger compartment of the vehicle to collect a first sound signal in the passenger compartment. The first sound signal is an original speech signal that includes media audio played by the onboard speakers. The system then acquires sound data for playback through the onboard speakers. Using an echo cancellation algorithm and referring to the sound data, the system removes the media audio from the first sound signal to obtain the original sound signal.

[0051] An array of multiple microphones installed inside the vehicle is used to capture sound within the cabin. The first sound signal collected by the in-vehicle microphone array includes all sounds within the cabin. While traveling, occupants may listen to media audio, such as navigation sounds or music, through the vehicle's speakers. During this process, the first sound signal contains the media audio played by the in-vehicle speakers. The media audio played by the in-vehicle speakers is generated based on sound data, which is the specific content played by the in-vehicle speakers, such as music files or audio files for navigation prompts. The in-vehicle speakers play the media audio by reading this music data. This sound data is extracted or recorded from the system for subsequent processing.

[0052] Echo cancellation is a signal processing technique used to remove echoes and interference signals. In this embodiment, it is used to eliminate media audio played by the vehicle's speakers from the first audio signal. Audio data acquired from the system is used as a reference signal. The echo cancellation algorithm removes the media audio portion from the first audio signal, resulting in a signal containing only the user's voice, thus eliminating the media audio played by the vehicle's speakers. For example, when the vehicle system plays navigation instructions and background music, the driver issues a voice wake-up command. The system filters out navigation audio and music signals using a preset echo cancellation algorithm, ensuring accurate recognition of the driver's voice commands.

[0053] The system employing the technical solution of this embodiment can eliminate media audio from the first sound signal obtained from the microphone when passengers listen to media audio through the vehicle's speakers, thereby reducing the impact of media audio on the recognition of the user's voice signal, enabling more accurate recognition and processing of the user's voice, and improving the quality of voice recognition.

[0054] In some embodiments, during the training process of the echo cancellation algorithm, the training data used to train the algorithm may include different media volumes, enabling the finally trained echo cancellation algorithm to cancel media sounds at different volumes. In addition, the training data used to train the algorithm may also include different vehicle noises, such as engine noise, or a mixture of engine noise and air conditioning noise. This vehicle noise can correspond to different vehicle speeds, enabling the finally trained echo cancellation algorithm to cancel noise generated by the vehicle's system at different speeds. For example, the echo cancellation algorithm can be adjusted in real time according to changes in vehicle speed and volume to ensure effective resistance to audio echoes even at high speeds.

[0055] Step 102: Obtain the target noise type of the noise signal, and determine the target filtering frequency band corresponding to the target noise type based on the preset first correspondence.

[0056] Different noise types may have different frequency characteristics and manifestations. Target noise types can be, for example: wind noise (the sound of wind blowing through a moving vehicle); mechanical noise (the sound of an engine, air conditioner, or other equipment); and environmental noise (the sounds of other vehicles on the road, pedestrians talking, etc.). After receiving the original speech signal, the noise type can be determined, and this determined noise type is the target noise type of the noise signal. After determining the target noise type, the distribution frequency band of the noise signal can be determined as the target filtering frequency band. The preset first correspondence is a pre-defined mapping table or rule that defines the relationship between different types of noise and their corresponding filtering frequency bands. By determining the target filtering frequency band, the distribution location of the noise signal in the original speech signal can be determined in detail. For example, if the target noise type is wind noise, the preset first correspondence can indicate that the wind noise is mainly concentrated in the low-frequency band (e.g., below 200Hz), then the target filtering frequency band is the frequency band below 200Hz. If the target noise type is mechanical noise, the preset first correspondence can indicate that the mechanical noise is mainly concentrated in the mid-frequency band (e.g., 500Hz to 2000Hz), then the target filtering frequency band is the 500Hz to 2000Hz band.

[0057] In some embodiments, the process of determining the target filtering frequency band can be as follows:

[0058] The original speech signal is converted into original spectrum data; a noise identification model is used to identify at least one type of noise signal in the original spectrum data; a target filtering frequency band corresponding to the target noise type frequency is determined based on a preset first correspondence, wherein the target noise type is at least one of the noise types.

[0059] Spectral data refers to the intensity information of each frequency component obtained by performing spectral analysis on the original speech signal, transforming it from the time domain to the frequency domain. Common spectral analysis methods include Fourier transform. Using spectral analysis methods (such as Fourier transform) converts the original speech signal into spectral data. This allows for a more intuitive view of the signal's distribution across different frequencies. Noise identification models can then be used to identify the signal's distribution across different frequencies, recognizing at least one type of noise signal from the original spectral data. For example, the model might identify wind noise, mechanical noise, and environmental noise. Exemplarily, noise type identification can be further combined with actual driving scenarios. For instance, when a vehicle is traveling at high speed, the detected noise is mainly high-frequency wind noise and mid-to-low-frequency road noise. After detecting the frequency band characteristics of these noises in the spectral data, the target noise type is automatically identified as "high-speed wind noise" and "road noise." When a vehicle is traveling on urban roads, the noise mainly originates from engine noise and external environmental noise, such as the sounds of other vehicles around the vehicle.

[0060] Step 103: Obtain the target intensity of the noise signal and determine the target intensity threshold corresponding to the target intensity based on the preset second correspondence.

[0061] The target intensity represents the strength value of the noise signal, which can be measured using a sound pressure level meter. The preset second correspondence is a predefined mapping between the target intensity of a noise signal and a target intensity threshold. This correspondence is typically a function or table and can be set according to different application scenarios. The target intensity threshold is used to determine whether the intensity of the noise signal has reached a standard requiring certain measures. For example, the target intensity threshold can be a level value, representing the intensity level of the noise signal. For instance, the target intensity threshold can be divided into levels 1 to 7; the higher the target intensity of the noise signal, the higher the corresponding target intensity threshold. Based on the preset second correspondence, the target intensity threshold corresponding to the current target intensity of the noise signal is found. For example, if the target intensity of the noise signal is 100dB, the preset second correspondence might specify a threshold of 2 for 100dB.

[0062] By determining the target intensity threshold corresponding to the target intensity of the noise signal, the system can better and more accurately classify the noise signal according to its noise intensity, and implement different measures based on the noise signal with different target intensity thresholds.

[0063] In some embodiments, acquiring the target intensity of the noise signal and determining a target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship specifically includes:

[0064] The target intensity change value of the noise signal is obtained, which is used to represent the change of the target intensity; the target intensity threshold corresponding to the target intensity is adjusted based on the preset second correspondence relationship to obtain the dynamic intensity threshold, which represents the change of the target intensity threshold and corresponds to the target intensity change value.

[0065] Since the target intensity represents the strength or energy level of the noise signal, the target intensity change value is used to represent how the noise signal intensity changes over time. For example, the intensity of the noise signal may increase or decrease over a certain period of time. The target intensity change value can be a numerical value representing the degree of change in noise intensity. A preset second correspondence exists between the target intensity of the noise signal and the target intensity threshold. Based on this, the preset second correspondence can be used to determine the dynamic intensity threshold corresponding to the target intensity change value. The dynamic intensity threshold represents how the target intensity threshold changes over time. The dynamic intensity threshold is adjusted in real time according to the changes in the noise signal intensity. Based on the adjusted target intensity threshold, the dynamic intensity threshold changing over time is obtained, enabling timely determination of changes in noise signal intensity and flexible response to noise signals of different intensities. In this way, the system can more effectively adapt to changes in the noise environment, improving the robustness and accuracy of speech signal processing.

[0066] Step 104: Filter the noise signal in the original speech signal based on the target filtering frequency band, and suppress the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal.

[0067] The target filtering frequency band is the frequency range where the noise signal is located. After determining the target filtering frequency band where the noise signal is located, the original speech signal in that target filtering frequency band can be filtered directly, speeding up the noise filtering process. For example, if the noise signal is mainly concentrated in the 1000Hz to 2000Hz frequency band, a filter can be used to filter the signal in this frequency band. For instance, when filtering a noise signal, an appropriate filtering strategy can be selected based on the target noise type, such as performing specific frequency band filtering for different noise types (wind noise, engine noise, road noise, etc.).

[0068] High-frequency filtering: mainly for wind noise, the system weakens or filters sounds exceeding a specific frequency.

[0069] Low-frequency filtering: For engine noise or tire friction noise, the system reduces low-frequency noise and retains mid-to-high frequency voice signals.

[0070] Besides noise filtering, technical means can be used to reduce or eliminate noise signals exceeding a target intensity threshold. If the target intensity of a noise signal exceeds the target intensity threshold, measures are taken to suppress or eliminate these noise signals. Common suppression methods include:

[0071] 1. Dynamic range compression: Reduces the dynamic range of noise signals, thereby reducing their intensity.

[0072] 2. Noise threshold: When the intensity of a noise signal exceeds a threshold, it is completely or partially removed.

[0073] 3. Adaptive filtering: Adaptive filters are used to dynamically adjust filtering parameters according to changes in the noise signal in order to suppress noise more effectively.

[0074] After filtering and suppression, a clear user voice signal is extracted from the original voice signal for subsequent processing or playback.

[0075] In some embodiments, the process of processing noise signals in the original speech signal includes:

[0076] The system obtains the target filtering frequency band change value and the target intensity change value of the noise signal during the wake-up period of the in-vehicle voice assistant. The target filtering frequency band change value is used to represent the change of the target filtering frequency band, and the target intensity change value is used to represent the change of the target intensity. The noise signal in the original speech signal is filtered based on the target filtering frequency band change value, and the noise signal in the original speech signal is suppressed based on the target intensity change value to obtain the user speech signal.

[0077] This embodiment proposes a specific application scenario: the user's voice signal generated by the sound signal processing method in this embodiment can be used to wake up the in-vehicle voice assistant. The time period during which the in-vehicle voice assistant is activated or woken up represents the wake-up time period of the in-vehicle voice assistant. This is typically the period after the in-vehicle voice assistant recognizes the user's wake-up word.

[0078] After being activated, the in-vehicle voice assistant can wait for further instructions from the passenger. During this process, the target noise type and target intensity of the noise signal in the received raw voice signal may change. For example, the dominant frequency band of the noise signal may change from a low frequency band to a high frequency band. For example, the intensity of the noise signal may first increase and then decrease. The target filter frequency band change value represents the change in frequency band. The target intensity change value represents the change in target intensity.

[0079] Based on the changes in the target filtering frequency band, this embodiment can dynamically adjust the filter parameters to effectively remove changing noise frequency bands. For example, if the main frequency band of the noise signal changes from the low frequency band to the high frequency band, the filter will adjust accordingly, so that the noise frequency filtered by the filter changes from the low frequency band to the high frequency band.

[0080] The noise suppression intensity is dynamically adjusted based on changes in the target intensity. For example, if the noise signal intensity first increases and then decreases, the noise suppression algorithm will adjust accordingly, first increasing the suppression intensity and then gradually decreasing it. After dynamic filtering and suppression processing, the user's voice signal can be extracted from the original speech signal over a certain period of time, and the noise has been effectively removed or reduced.

[0081] In this embodiment, the original speech signal is collected from the passenger compartment of the vehicle. When the original speech signal contains noise, in order to separate the user's speech signal from the original speech signal, this application can determine the target noise type and target intensity of the noise signal. A target filtering frequency band corresponding to the target noise type is determined through a preset first correspondence, and a target intensity threshold corresponding to the target intensity value is determined through a preset second correspondence. The noise signal at the target filtering frequency in the original speech signal is filtered, and simultaneously, the noise signal in the original speech signal is suppressed using the target intensity threshold, reducing the noise signal in the original speech signal, reducing the impact of the noise signal on speech recognition, and improving the accuracy of speech recognition.

[0082] In some embodiments, the sound signal processing method further includes: adjusting the wake-up sensitivity of the in-vehicle voice assistant according to the target intensity of the noise signal.

[0083] Wake-up sensitivity refers to how sensitive the in-vehicle voice assistant is to the wake-up word. High wake-up sensitivity means the system is easier to wake up, while low sensitivity means the system is harder to wake up. To achieve sensitivity matching with noise intensity, if the noise signal intensity is high, the system will increase wake-up sensitivity. This reduces the risk of missed wake-ups. If the noise signal intensity is low, the system will decrease sensitivity, thereby reducing missed wake-ups and improving user experience. For example, in high-noise environments, such as when driving on a highway, the noise inside and outside the vehicle is high. In this case, the system will increase wake-up sensitivity, allowing the voice assistant to more sensitively acquire the user's voice signal. In low-noise environments, such as in a parking lot or inside a stationary car, the noise is low. In this case, the system will decrease wake-up sensitivity, reducing the risk of false wake-ups. By being able to wake up the in-vehicle voice assistant in different noise environments, both erroneous operations and the user experience are improved.

[0084] In some embodiments, the following may also be performed: determining the human voice frequency range to which the user's voice signal belongs; enhancing the user's voice signal within the human voice frequency range to obtain the enhanced user voice signal.

[0085] Human voice frequency range: The frequency range of human speech is typically between 80 Hz and 14,000 Hz, but most useful speech information is concentrated between 300 Hz and 3,400 Hz. This range is known as the human voice frequency range. By analyzing the spectrum of a user's speech signal, the human voice frequency range in which the user's speech signal is primarily concentrated can be determined.

[0086] After determining the human voice frequency range to which the user's voice signal belongs, the signal within this range is enhanced.

[0087] The signal enhancement method includes: gain adjustment: increasing the signal strength within the frequency band. After signal enhancement processing, a clearer and higher-quality user voice signal is obtained. In this embodiment, to adapt to changes in the signal strength of the noise signal during the signal enhancement process, the gain of the user voice signal can also be adjusted according to changes in ambient noise. When the noise intensity is high, the gain of the voice signal is increased to improve the signal-to-noise ratio; when the noise intensity is low, the gain is appropriately reduced to avoid distortion caused by excessive volume.

[0088] In this way, the system can effectively improve the quality of the user's voice signal, enhance the accuracy of voice recognition, and improve the user experience.

[0089] In some embodiments, the sound signal processing method may further:

[0090] Obtain user feedback on false or missed wake-ups. Adjust noise filtering and suppression strategies, as well as sensitivity settings, based on this feedback. Train a machine learning model using long-term collected noise data to optimize filtering strategies and wake-up recognition models under different noise environments.

[0091] After the system collects enough user feedback and noise signals under different target noise types, it gradually optimizes its noise filtering and sensitivity adjustment strategies, so that the system can better adapt to individual driving habits and environmental changes over time.

[0092] This embodiment provides a structural diagram of a sound signal processing device, as shown below. Figure 2 As shown, it includes:

[0093] The acquisition module 21 is used to acquire the original voice signal collected in the passenger compartment of the vehicle, wherein the original voice signal is a user voice signal with noise signal.

[0094] The first determining module 22 is used to obtain the target noise type of the noise signal and determine the target filtering frequency band corresponding to the target noise type based on a preset first correspondence relationship.

[0095] The second determining module 23 is used to obtain the target intensity of the noise signal and determine the target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship.

[0096] The processing module 24 is used to filter the noise signal in the original speech signal based on the target filtering frequency band, and suppress the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal.

[0097] In some embodiments, the first determining module 22 is configured to:

[0098] Convert the original speech signal into original spectrum data;

[0099] Using a noise identification model, at least one type of noise signal is identified in the original spectrum data;

[0100] The target filtering frequency band corresponding to the target noise type frequency is determined based on a preset first correspondence relationship, wherein the target noise type is at least one of the noise types.

[0101] In some embodiments, the second determining module 23 is configured to:

[0102] The target intensity change value of the noise signal is obtained, and the target intensity change value is used to represent the change in the target intensity.

[0103] Based on a preset second correspondence, the target intensity threshold corresponding to the target intensity is adjusted to obtain a dynamic intensity threshold, which represents the change of the target intensity threshold and corresponds to the change value of the target intensity.

[0104] In some embodiments, the acquisition module 21 is used for:

[0105] The first sound signal inside the vehicle cabin is acquired using an onboard microphone array in the passenger compartment of the vehicle. The first sound signal is the original speech signal that includes media audio played by the onboard speakers.

[0106] Acquire sound data for playback through the vehicle-mounted speaker;

[0107] Using an echo cancellation algorithm and referring to the sound data, the media audio in the first sound signal is removed to obtain the original sound signal.

[0108] In some embodiments, the processing module 24 is configured to:

[0109] During the wake-up period of the in-vehicle voice assistant, the target filter frequency band change value and the target intensity change value of the noise signal are obtained. The target filter frequency band change value is used to represent the change of the target filter frequency band, and the target intensity change value is used to represent the change of the target intensity.

[0110] The noise signal in the original speech signal is filtered based on the target filtering frequency band change value, and the noise signal in the original speech signal is suppressed based on the target intensity change value to obtain the user speech signal.

[0111] In some embodiments, the sound signal processing apparatus further includes: an adjustment module 25, configured to:

[0112] The wake-up sensitivity of the in-vehicle voice assistant is adjusted according to the target intensity of the noise signal.

[0113] In some embodiments, the sound signal processing apparatus further includes: an enhancement module 26, configured to:

[0114] Determine the human voice frequency range to which the user's voice signal belongs;

[0115] The user's voice signal in the specified human voice frequency range is enhanced to obtain the enhanced user's voice signal.

[0116] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0117] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-readable program code.

[0118] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0119] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0120] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0121] This application also provides a computer program product, which includes computer software instructions that, when executed on a processing device, cause the processing device to perform a sound signal processing method.

[0122] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0127] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0128] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

[0129] Although preferred embodiments have been described in this specification, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of this specification. Clearly, those skilled in the art can make various alterations and modifications to this specification without departing from its spirit and scope. Thus, if such modifications and modifications fall within the scope of the claims and their equivalents, this specification also intends to include such modifications and modifications.

Claims

1. A sound signal processing method, characterized in that, include: Acquire raw speech signals collected in the passenger compartment of the vehicle, wherein the raw speech signals are user speech signals with noise signals; The target noise type of the noise signal is obtained, and the target filtering frequency band corresponding to the target noise type is determined based on a preset first correspondence relationship. The acquisition of the target noise type is combined with the actual driving scenario, which includes highway driving scenario and urban road driving scenario. Obtain the target intensity of the noise signal, and determine the target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship; The step of acquiring the target intensity of the noise signal and determining the target intensity threshold corresponding to the target intensity based on a preset second correspondence includes: The target intensity change value of the noise signal is obtained, and the target intensity change value is used to represent the change in the target intensity. Based on a preset second correspondence, the target intensity threshold corresponding to the target intensity is adjusted to obtain a dynamic intensity threshold, wherein the dynamic intensity threshold represents the change of the target intensity threshold and corresponds to the change value of the target intensity; The noise signal in the original speech signal is filtered based on the target filtering frequency band, and the noise signal in the original speech signal is suppressed based on the target intensity threshold to obtain the user speech signal; When the user voice signal is used to wake up the in-vehicle voice assistant, the noise signal in the original voice signal is filtered based on the target filtering frequency band, and the noise signal in the original voice signal is suppressed based on the target intensity threshold to obtain the user voice signal, including: During the wake-up period of the in-vehicle voice assistant, the target filter frequency band change value and the target intensity change value of the noise signal are obtained. The target filter frequency band change value is used to represent the change of the target filter frequency band, and the target intensity change value is used to represent the change of the target intensity. The noise signal in the original speech signal is filtered based on the target filtering frequency band change value, and the noise signal in the original speech signal is suppressed based on the target intensity change value to obtain the user speech signal.

2. The method according to claim 1, characterized in that, The step of acquiring the target noise type of the noise signal and determining the target filtering frequency band corresponding to the target noise type based on a preset first correspondence includes: Convert the original speech signal into original spectrum data; Using a noise identification model, at least one type of noise signal is identified in the original spectrum data; The target filtering frequency band corresponding to the target noise type frequency is determined based on a preset first correspondence relationship, wherein the target noise type is at least one of the noise types.

3. The method according to claim 1, characterized in that, The acquisition of the raw voice signal collected in the passenger compartment of the vehicle includes: The first sound signal inside the vehicle cabin is acquired using an onboard microphone array in the passenger compartment of the vehicle. The first sound signal is the original speech signal that includes media audio played by the onboard speakers. Acquire sound data for playback through the vehicle-mounted speaker; Using an echo cancellation algorithm and referring to the sound data, the media audio in the first sound signal is removed to obtain the original sound signal.

4. The method according to claim 1, characterized in that, When the user's voice signal is used to wake up the in-vehicle voice assistant, the method further includes: The wake-up sensitivity of the in-vehicle voice assistant is adjusted according to the target intensity of the noise signal.

5. The method according to claim 1, characterized in that, After filtering the noise signal in the original speech signal based on the target filtering frequency band and suppressing the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal, the method further includes: Determine the human voice frequency range to which the user's voice signal belongs; The user's voice signal in the specified human voice frequency range is enhanced to obtain the enhanced user's voice signal.

6. A sound signal processing device, characterized in that, include: The acquisition module is used to acquire the original voice signal collected in the passenger compartment of the vehicle, wherein the original voice signal is a user voice signal with noise signal; The first determining module is used to obtain the target noise type of the noise signal and determine the target filtering frequency band corresponding to the target noise type based on a preset first correspondence relationship. The acquisition of the target noise type is combined with the actual driving scenario, which includes highway driving scenario and urban road driving scenario. The second determining module is used to obtain the target intensity of the noise signal and determine the target intensity threshold corresponding to the target intensity based on a preset second correspondence relationship. The step of acquiring the target intensity of the noise signal and determining the target intensity threshold corresponding to the target intensity based on a preset second correspondence includes: The target intensity change value of the noise signal is obtained, and the target intensity change value is used to represent the change in the target intensity. Based on a preset second correspondence, the target intensity threshold corresponding to the target intensity is adjusted to obtain a dynamic intensity threshold, wherein the dynamic intensity threshold represents the change of the target intensity threshold and corresponds to the change value of the target intensity; The processing module is used to filter the noise signal in the original speech signal based on the target filtering frequency band, and suppress the noise signal in the original speech signal based on the target intensity threshold to obtain the user speech signal; When the user voice signal is used to wake up the in-vehicle voice assistant, the noise signal in the original voice signal is filtered based on the target filtering frequency band, and the noise signal in the original voice signal is suppressed based on the target intensity threshold to obtain the user voice signal, including: During the wake-up period of the in-vehicle voice assistant, the target filter frequency band change value and the target intensity change value of the noise signal are obtained. The target filter frequency band change value is used to represent the change of the target filter frequency band, and the target intensity change value is used to represent the change of the target intensity. The noise signal in the original speech signal is filtered based on the target filtering frequency band change value, and the noise signal in the original speech signal is suppressed based on the target intensity change value to obtain the user speech signal.

7. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Waking computing devices based on ambient noise

    CN108700926A

  • Voice noise reduction method, electronic equipment and storage medium

    CN115662457A

  • Speech recognition method, storage medium storing speech recognition program, and speech recognition apparatus

    US20020049587A1