A vehicle-mounted voice processing method and system
By collecting noise signals and sound signals in the vehicle, filtering and analyzing the corresponding frequency bands, the efficiency and cost reduction of on-board voice noise reduction are improved, and the problems of high workload and high cost in the existing technology are solved.
Patent Information
- Application Number
- CN202210003281.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-04
AI Technical Summary
In the prior art, the in-vehicle voice noise reduction method has a large workload and high cost, making it difficult to effectively eliminate in-vehicle noise and affect the quality of the call.
By collecting noise signals and sound signals in the car, obtaining the upper frequency limit of the noise signal and the lower frequency limit of the user's sound frequency, filtering the corresponding frequency bands in the noise signal and the sound signal, and only spectrum analysis and noise reduction processing are performed on the frequency bands below the upper frequency limit of the noise signal.
It reduces the sound signal band that requires spectrum analysis, reduces the workload and cost of voice noise reduction in vehicle on-board, and improves call quality.
Smart Images

Figure CN114495970B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a vehicle-mounted speech processing method and system. Background Art
[0002] When the vehicle is running, the noise generated by the engine operation, wind noise, tire noise, etc. will be attached to the user's call in the vehicle, resulting in noise in the sound signal collected by the sound collection device, reducing the call quality. It is necessary to eliminate the noise from the collected sound signal to obtain a speech signal and improve the call quality.
[0003] Currently, the common speech noise reduction method is as follows: perform spectrum analysis on the collected sound signal, identify the waveform characteristics and period of the noise signal, and then obtain the waveform of the noise signal. Then, generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to cancel the noise signal. Since the sound signal has multiple frequency bands, the above speech noise reduction method needs to perform spectrum analysis on each frequency band to identify the waveform of the noise signal, resulting in a large workload and high cost of spectrum analysis. Summary of the Invention
[0004] The present invention provides a vehicle-mounted speech processing method and system, which solves the technical problems of large workload and high cost of the existing vehicle-mounted speech noise reduction method.
[0005] On the one hand, the present invention provides the following technical solution:
[0006] A vehicle-mounted speech processing method, comprising:
[0007] When the vehicle is running, collect the noise signal in the vehicle under the condition of no personnel communication in the vehicle, and obtain the upper frequency limit of the noise signal;
[0008] When the vehicle is running, collect the sound signal in the vehicle;
[0009] Obtain the lower frequency limit of the user's voice frequency, and filter the frequency band of the sound signal below the lower frequency limit of the voice frequency;
[0010] Perform spectrum analysis on the frequency band of the sound signal below the upper frequency limit to obtain the waveform of the noise signal;
[0011] Generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to obtain the required speech signal.
[0012] Preferably, the upper frequency limit of the noise signal is 1000 Hz.
[0013] Preferably, the lower frequency limit of the voice frequency is 150 Hz.
[0014] Preferably, generating a waveform opposite to the waveform of the noise signal and superimposing it on the voice signal to obtain the required voice signal, and then further including:
[0015] Determining whether the voice signal is complete;
[0016] If the voice signal is incomplete, performing fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal.
[0017] Preferably, the step of, if the voice signal is incomplete, performing fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal, includes:
[0018] If the voice signal is incomplete, constructing a state network through a Hidden Markov Model and connecting the state network to the voice library;
[0019] Searching for the voice feature information that best matches the voice signal in the state network.
[0020] Preferably, the step of searching for the voice feature information that best matches the voice signal in the state network includes:
[0021] Constructing a path search algorithm;
[0022] Searching for an optimal path in the state network through the path search algorithm, and taking the voice feature information corresponding to the optimal path as the voice feature information that best matches the voice signal.
[0023] On the other hand, the present invention also provides the following technical solution:
[0024] A vehicle-mounted voice processing system, including:
[0025] A signal acquisition module, configured to collect the noise signal in the vehicle when the vehicle is running and there is no personnel communication in the vehicle, and obtain the upper frequency limit of the noise signal;
[0026] The signal acquisition module is further configured to collect the voice signal in the vehicle when the vehicle is running;
[0027] A signal filtering module, configured to obtain the lower frequency limit of the user's voice, and filter the frequency band of the voice signal below the lower frequency limit of the voice;
[0028] A signal analysis module, configured to perform spectral analysis on the frequency band of the voice signal below the upper frequency limit to obtain the waveform of the noise signal;
[0029] A signal noise reduction module, configured to generate a waveform opposite to the waveform of the noise signal and superimpose it on the voice signal to obtain the required voice signal.
[0030] Preferably, the in-vehicle voice processing system further includes:
[0031] A signal judgment module, configured to judge whether the voice signal is complete;
[0032] An information matching module, configured to perform fuzzy matching on the voice library and the voice signal if the voice signal is incomplete, and obtain the voice feature information that best matches the voice signal.
[0033] On the other hand, the present invention also provides the following technical solutions:
[0034] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned any in-vehicle voice processing method is implemented.
[0035] On the other hand, the present invention also provides the following technical solutions:
[0036] A computer-readable storage medium, which implements the above-mentioned any in-vehicle voice processing method when executed.
[0037] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0038] First, filter the frequency band of the sound signal below the lower limit of the human voice frequency, so that there is no need to perform spectrum analysis on the frequency band of the sound signal between the lower limit of the noise signal frequency and the lower limit of the voice frequency; then perform spectrum analysis on the frequency band of the sound signal below the upper limit of the noise signal frequency, so that there is no need to perform spectrum analysis on the frequency band of the sound signal between the upper limit of the noise signal frequency and the upper limit of the voice frequency, thereby greatly reducing the frequency band of the sound signal that needs spectrum analysis, and the workload and cost of in-vehicle voice noise reduction are small. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a flowchart of the in-vehicle voice processing method in the embodiment of the present invention;
[0041] Figure 2 It is a schematic structural diagram of the in-vehicle voice processing system in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In an embodiment of the present invention, by providing a vehicle-mounted voice processing method and system, the technical problem in the prior art that the workload and cost of vehicle-mounted voice noise reduction are large is solved.
[0043] To better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below in conjunction with the specification drawings and specific embodiments.
[0044] First, it should be noted that the term "and / or" appearing in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the preceding and following associated objects.
[0045] As Figure 1 shown, the vehicle-mounted voice processing method of this embodiment includes:
[0046] Step S1, when the vehicle is running, collect the noise signal in the vehicle under the condition that there is no personnel communication in the vehicle, and obtain the upper frequency limit of the noise signal;
[0047] Step S2, when the vehicle is running, collect the sound signal in the vehicle;
[0048] Step S3, obtain the lower frequency limit of the user's voice frequency, and filter the frequency band in the sound signal that is below the lower frequency limit of the voice frequency;
[0049] Step S4, perform a spectrum analysis on the frequency band in the sound signal that is below the upper frequency limit, and obtain the waveform of the noise signal;
[0050] Step S5, generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to obtain the required voice signal.
[0051] Step S1 can be regarded as multiple repetitions of an experiment to collect the noise generated when the vehicle is running. The purpose is to obtain the most comprehensive frequency range of the noise signal. Through experimental verification in this embodiment, the frequency range of specific noise generated when the vehicle is running (such as engine noise, wind noise, tire noise, etc.) is 10 - 1000 Hz, that is, the lower frequency limit of the noise signal is 10 Hz and the upper frequency limit is 1000 Hz, covering all possible noise signals generated when the vehicle is running.
[0052] Step S2 can be regarded as collecting the sound signal in the vehicle every time the user uses the vehicle after obtaining the upper frequency limit of the noise signal. If there is no one speaking in the vehicle, the sound signal is equivalent to the noise signal; if there is a call among the vehicle occupants, the sound signal includes the voice signal and the noise signal.
[0053] In step S3, through experiments in this embodiment, it is found that the frequency range of the sound emitted by a person during a voice call is 150 - 3000 Hz. Thus, the lower limit of the user's vocal frequency is 150 Hz, and the upper limit of the vocal frequency is 3000 Hz. It can be seen from this that when the vehicle is running, if there are people making a voice call in the vehicle, the frequency range of the obtained sound signal is 10 - 3000 Hz (it can be considered that 10 - 1000 Hz is the noise signal, and 1000 - 3000 Hz is the required voice signal). The overlapping frequency band of the voice signal and the noise signal is 150 - 1000 Hz. In this way, filtering the frequency band below the lower limit of the vocal frequency in the sound signal, that is, filtering the signal in the 10 - 150 Hz frequency band, the frequency range of the filtered sound signal is 150 - 3000 Hz. Of course, if there is no one talking in the vehicle, the frequency range of the filtered sound signal is 150 - 1000 Hz (all are noise signals).
[0054] In step S4, since the upper limit of the frequency of the noise signal is 1000 Hz, it can be considered that there is no noise signal in the sound signal with a frequency range of 1000 - 3000 Hz, and there is no need to eliminate the noise. Only the sound signal with a frequency range of 150 - 1000 Hz needs to be spectroscopically analyzed to obtain the waveform characteristics and period of the noise signal, and then the waveform of the noise signal with a frequency range of 150 - 1000 Hz is obtained.
[0055] In step S5, after obtaining the waveform of the noise signal with a frequency range of 150 - 1000 Hz, a waveform opposite to the waveform of the noise signal can be generated and superimposed on the sound signal to obtain the required voice signal.
[0056] As can be seen from the above, in this embodiment, the frequency band below the lower limit of the human vocal frequency in the sound signal is first filtered, so there is no need to spectroscopically analyze the frequency band between the lower limit of the noise signal frequency and the lower limit of the vocal frequency in the sound signal; then the frequency band below the upper limit of the noise signal frequency in the sound signal is spectroscopically analyzed, so there is no need to spectroscopically analyze the frequency band between the upper limit of the noise signal frequency and the upper limit of the vocal frequency in the sound signal. Thus, the frequency band of the sound signal that needs to be spectroscopically analyzed is greatly reduced, and the workload and cost of in-vehicle voice noise reduction are small.
[0057] It is easy to think that 150~3000Hz may not cover the vocal frequency ranges of all people. If there are still useful voice signals below 150Hz, step S3 will filter out some useful voice signals, resulting in an incomplete call. Regarding the problem of incomplete calls, of course, the lower limit of the vocal frequency can be appropriately reduced, such as to 100Hz, to cover the vocal frequency ranges of all people. However, in this case, the frequency band of the voice signal that needs to be analyzed by spectrum analysis becomes 100~1000Hz, resulting in an increased workload of spectrum analysis. To avoid increasing the workload of spectrum analysis, preferably after step S5 in this embodiment, the in-vehicle voice processing method of this embodiment further includes:
[0058] Determine whether the voice signal is complete; if the voice signal is incomplete, then perform fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal.
[0059] Among them, if the voice signal is complete, the call is complete and there are no voice signals that are filtered out; if the voice signal is incomplete, for example, the semantics corresponding to the obtained voice signal is "How about it, do you want to go?", the complete semantics may be "What do you think, do you want to go?", so the semantics needs to be supplemented.
[0060] The step of, if the voice signal is incomplete, then performing fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal includes: if the voice signal is incomplete, then construct a state network through a hidden Markov model and connect the state network to the voice library; find the voice feature information that best matches the voice signal in the state network. The step of finding the voice feature information that best matches the voice signal in the state network includes: constructing a path search algorithm; searching for an optimal path in the state network through the path search algorithm, and using the voice feature information corresponding to the optimal path as the voice feature information that best matches the voice signal.
[0061] Among them, a state network is constructed through a hidden Markov model. It is expanded from words and the network into a phoneme network, and then into a state network. Searching for the most matching speech feature information is actually searching for an optimal path in the state network. That is, if the probability of the speech signal corresponding to this path is the largest, it is the best fuzzy matching result. Here, the Viterbi algorithm for path search needs to be introduced to find the globally optimal path. The optimal path search is set according to the following principles: Analyze and process the speech signal to remove redundant information; Extract the key information affecting speech recognition and the feature information expressing the meaning of the language; Closely follow the feature information and identify words with the smallest unit; Identify words in the order of their respective grammars in different languages; Regard the context before and after as an auxiliary recognition condition, which is conducive to analysis and recognition; According to semantic analysis, divide paragraphs for key information, extract the recognized words and connect them, and at the same time adjust the sentence structure according to the meaning of the sentence; Combine semantics and carefully analyze the mutual connection of the context to appropriately correct the sentence being processed. Finally, the speech feature information that best matches the speech signal can be obtained as "What do you think? Do you want to go?" In this way, when filtering some useful speech signals causes the call to be incomplete, the integrity of in-vehicle voice calls can be ensured without increasing the workload of spectrum analysis.
[0062] As Figure 2 shown, this embodiment also provides an in-vehicle voice processing system, including:
[0063] A signal acquisition module, which is used to collect the noise signal in the vehicle when the vehicle is running under the condition of no in-vehicle personnel communication, and obtain the upper frequency limit of the noise signal;
[0064] The signal acquisition module is also used to collect the sound signal in the vehicle when the vehicle is running;
[0065] A signal filtering module, which is used to obtain the lower frequency limit of the user's voice frequency and filter the frequency band of the sound signal below the lower frequency limit of the voice frequency;
[0066] A signal analysis module, which is used to perform spectrum analysis on the frequency band of the sound signal below the upper frequency limit to obtain the waveform of the noise signal;
[0067] A signal noise reduction module, which is used to generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to obtain the required speech signal.
[0068] In this embodiment, the frequency bands in the voice signal below the lower limit of the human voice frequency are filtered first, so that it is not necessary to perform spectral analysis on the frequency bands in the voice signal between the lower limit of the noise signal frequency and the lower limit of the voice frequency; then, spectral analysis is performed on the frequency bands in the voice signal below the upper limit of the noise signal frequency, so that it is not necessary to perform spectral analysis on the frequency bands in the voice signal between the upper limit of the noise signal frequency and the upper limit of the voice frequency. Thus, the frequency bands of the voice signal that need spectral analysis are greatly reduced, and the workload and cost of in-vehicle voice noise reduction are small.
[0069] Furthermore, the in-vehicle voice processing system further includes:
[0070] A signal judgment module, used to judge whether the voice signal is complete;
[0071] An information matching module, used to perform fuzzy matching on the voice library and the voice signal if the voice signal is incomplete, and obtain the voice feature information that best matches the voice signal.
[0072] When this embodiment filters some useful voice signals and causes the call to be incomplete, fuzzy matching is performed on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal, which can ensure the integrity of in-vehicle voice calls without increasing the workload of spectral analysis.
[0073] Based on the same inventive concept as the in-vehicle voice processing method described above, this embodiment also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the in-vehicle voice processing methods described above.
[0074] Among them, the bus architecture (represented by the bus), the bus can include any number of interconnected buses and bridges, and the bus links various circuits including one or more processors represented by the processor and the memory represented by the memory together. The bus can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the receiver and transmitter. The receiver and transmitter can be the same element, that is, a transceiver, which provides a unit for communicating with various other devices on the transmission medium. The processor is responsible for managing the bus and general processing, and the memory can be used to store the data used by the processor when performing operations.
[0075] Since the electronic device introduced in this embodiment is the electronic device used to implement the in-vehicle voice processing method in the embodiments of the present invention, based on the in-vehicle voice processing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various forms of variation of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present invention will not be described in detail here. As long as those skilled in the art implement the electronic device used in the in-vehicle voice processing method in the embodiments of the present invention, it falls within the scope of protection of the present invention.
[0076] Based on the same inventive concept as the above in-vehicle voice processing method, the present invention also provides a computer-readable storage medium, which when executed implements any of the above in-vehicle voice processing methods.
[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0078] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or boxes. Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.
[0081] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0082] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A vehicle-mounted voice processing method, characterized in that, Including: When the vehicle is running, collect the noise signal inside the vehicle under the condition of no in-vehicle personnel communication, and obtain the upper frequency limit of the noise signal; When the vehicle is running, collect the sound signal inside the vehicle; Obtain the lower limit of the user's vocal frequency, and filter the frequency band of the sound signal below the lower limit of the vocal frequency; Perform spectrum analysis on the frequency band of the sound signal below the upper frequency limit to obtain the waveform of the noise signal; Generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to obtain the required voice signal; Judge whether the voice signal is complete; If the voice signal is incomplete, perform fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal.
2. The vehicle-mounted voice processing method according to claim 1, wherein, The upper frequency limit of the noise signal is 1000 Hz.
3. The vehicle-mounted voice processing method according to claim 1, wherein, The lower limit of the vocal frequency is 150 Hz.
4. The vehicle-mounted voice processing method according to claim 1, characterized in that, The step of if the voice signal is incomplete, perform fuzzy matching on the voice library and the voice signal to obtain the voice feature information that best matches the voice signal, including: If the voice signal is incomplete, construct a state network through a hidden Markov model, and connect the state network to the voice library; Find the voice feature information that best matches the voice signal in the state network.
5. The vehicle-mounted voice processing method according to claim 4, characterized in that, The step of finding the voice feature information that best matches the voice signal in the state network, including: Construct a path search algorithm; Search for an optimal path in the state network through the path search algorithm, and use the voice feature information corresponding to the optimal path as the voice feature information that best matches the voice signal.
6. A vehicle-mounted voice processing system, characterized in that, Including: A signal acquisition module, which is used to collect the noise signal inside the vehicle under the condition of no in-vehicle personnel communication when the vehicle is running, and obtain the upper frequency limit of the noise signal; The signal acquisition module is also used to collect the sound signal inside the vehicle when the vehicle is running; A signal filtering module, which is used to obtain the lower limit of the user's vocal frequency and filter the frequency band of the sound signal below the lower limit of the vocal frequency; A signal analysis module, which is used to perform spectrum analysis on the frequency band of the sound signal below the upper frequency limit to obtain the waveform of the noise signal; A signal noise reduction module, which is used to generate a waveform opposite to the waveform of the noise signal and superimpose it on the sound signal to obtain the required voice signal; A signal judgment module, which is used to judge whether the voice signal is complete; An information matching module, which is used to perform fuzzy matching on the voice library and the voice signal if the voice signal is incomplete, and obtain the voice feature information that best matches the voice signal.
7. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the in-vehicle voice processing method described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, When the computer-readable storage medium is executed, it implements the in-vehicle voice processing method described in any one of claims 1-5.
Citation Information
Patent Citations
HTK-based continuous speech recognition system
CN106531152A
Noise reduction method and device for vehicle-mounted environment, electronic equipment and storage medium
CN111477206A
Vehicle and sound processing system of vehicle
CN112420029A