Vehicle Audio Control Method, Device, Vehicle and Storage Medium
By using sound classification model and sound pickup strategy in the vehicle audio control system, the problem of low wake-up success rate of the on-board intelligent voice system under noise and background music interference is solved, and efficient voice wake-up in complex environments is achieved.
Patent Information
- Application Number
- CN202210884396.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-25
AI Technical Summary
The success rate of voice wake-up in traditional car intelligent voice systems is low, mainly due to noise in the car and background music interference, it is difficult for the vehicle to accurately extract wake-up words.
By collecting audio signals in the car, the audio signal classification model is used to determine the classification results of the audio signal, and appropriate sound pickup strategies are selected based on the classification results, including controlling the window status, adjusting the background music volume and turning on the audio acquisition device that matches the sound source position to improve the sound pickup effect.
It improves the vehicle's voice wake-up success rate in environments of noise and background music interference, ensuring that the vehicle can accurately extract wake-up words.
Smart Images

Figure CN115394292B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicles, and more specifically, to a vehicle audio control method, device, vehicle, and storage medium. Background Art
[0002] In the traditional automotive field, the functions of a vehicle itself are generally completed through various different buttons inside the vehicle body. Operating these different buttons during driving still poses certain risks and inconveniences. To simplify manual operations without reducing functions, in-vehicle intelligent voice systems play an important role. Through the in-vehicle intelligent voice system, users can control the vehicle by voice.
[0003] The in-vehicle intelligent voice system includes a voice wake-up module. The voice wake-up module pre-sets a wake-up word. When the user says this wake-up word, after the voice wake-up module recognizes that the voice content includes the wake-up word, the voice content after the wake-up word is used as the actual content of the operation command issued by the user. However, with the existing solutions, the voice wake-up success rate is relatively low. Summary of the Invention
[0004] In view of this, embodiments of this application propose a vehicle audio control method, device, vehicle, and storage medium.
[0005] In a first aspect, embodiments of this application provide a vehicle audio control method, the method including: collecting an audio signal inside the compartment of a target vehicle as a target audio signal; when the target audio signal meets a preset condition, determining a classification result of the target audio signal through a sound classification model as a target classification result; determining a target sound pickup strategy corresponding to the target classification result in a preset strategy set, the preset strategy set including multiple classification results and the sound pickup strategies respectively corresponding to the multiple classification results, and each sound pickup strategy in the preset strategy set is used to improve the sound pickup effect inside the compartment; controlling the target vehicle to output a corresponding action through the target sound pickup strategy.
[0006] In a second aspect, an embodiment of the present application provides a vehicle audio control device, which includes: a collection module, configured to collect an audio signal in the carriage of a target vehicle as a target audio signal; a classification module, configured to determine a classification result of the target audio signal as a target classification result through a sound classification model when the target audio signal meets a preset condition; a determination module, configured to determine a target sound pickup strategy corresponding to the target classification result from a preset strategy set, where the preset strategy set includes multiple classification results and the sound pickup strategies respectively corresponding to the multiple classification results, and each sound pickup strategy in the preset strategy set is used to improve the sound pickup effect in the carriage; and a control module, configured to control the target vehicle to output a corresponding action through the target sound pickup strategy.
[0007] Optionally, the device further includes an analysis module, configured to determine the noise suppression, clipping information, and signal-to-noise ratio corresponding to the target audio signal; obtain a noise analysis result of the target audio signal according to the signal-to-noise ratio; obtain a noise volume analysis result in the target audio signal according to the clipping information; determine a voice analysis result of the target audio signal according to the noise suppression; and determine whether the target audio signal meets the preset condition according to the noise analysis result, the voice analysis result, and the noise volume analysis result.
[0008] Optionally, the classification module is further configured to segment the target audio signal according to a silence interval in the target audio signal to obtain a plurality of audio segments; classify each of the audio segments through the sound classification model to obtain a classification result of each of the audio segments; and obtain the target classification result according to the classification results of the plurality of audio segments respectively.
[0009] Optionally, the device further includes a training module, configured to obtain training samples, where the training samples include sample audios respectively corresponding to the multiple classification results and annotation information corresponding to the sample audios; and train an initial model according to the sample audios and the annotation information to obtain the sound classification model.
[0010] Optionally, when the target classification result is that the target audio signal includes noise and a living person's voice, the control module is further configured to control the window of the target vehicle to enter the window-closing state, and determine noise reduction parameters for reducing noise of the real-time audio signal according to the collected real-time audio signal in the vehicle compartment; when the target classification result is that the target audio signal includes background music and a living person's voice, the control module is further configured to turn down the volume of the background music of the target vehicle; when the target classification result is that the target audio signal includes a living person's voice and a non-living person's voice, the control module is further configured to, if the target vehicle outputs background music, turn down the volume of the background music of the target vehicle, determine the sound source position of the target audio signal, and turn on an audio collection device matching the sound source position.
[0011] In a third aspect, an embodiment of the present application provides a vehicle, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and when the program code is run by a processor, the above method is executed.
[0013] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of the vehicle reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the vehicle executes the above method.
[0014] A vehicle audio control method, device, vehicle and storage medium provided by an embodiment of the present application determine a target sound pickup strategy through a target classification result corresponding to a target audio signal, and control the vehicle through the target sound pickup strategy, improving the sound pickup effect in the vehicle compartment of the target vehicle, enabling the target vehicle to collect audio signals with higher clarity, so that the target vehicle can accurately extract a wake-up word from the audio signals with higher clarity, and improving the voice wake-up success rate of the target vehicle. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0016] Figure 1The figure shows a schematic diagram of the implementation environment of a vehicle audio control method proposed by this application;
[0017] Figure 2 The figure shows a schematic diagram of the implementation system architecture of a vehicle audio control method proposed by this application;
[0018] Figure 3 The figure shows a flowchart of a vehicle audio control method proposed by an embodiment of this application;
[0019] Figure 4 The figure shows Figure 3 a flowchart of an implementation manner before S120 in
[0020] Figure 5 The figure shows a block diagram of a vehicle audio control device proposed by an embodiment of this application;
[0021] Figure 6 The figure shows a structural block diagram of a vehicle for executing the vehicle audio control method according to the embodiment of this application. Specific embodiments
[0022] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. According to the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.
[0023] In the following description, the terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0025] Currently, during the operation of a vehicle, there may be noises such as the surrounding environment and background music in the vehicle compartment, making the audio signals collected by the in-vehicle intelligent voice system in the vehicle compartment include more noises, resulting in the vehicle being difficult to extract the wake-up word from the audio signals including noises, the vehicle cannot be woken up in time, and the vehicle wake-up success rate is relatively low.
[0026] In view of this, the inventor provides a vehicle audio control method, apparatus, vehicle, and storage medium. By determining a target sound pickup strategy based on the target classification result corresponding to the target audio signal, and controlling the vehicle through the target sound pickup strategy, the sound pickup effect inside the carriage of the target vehicle is improved, enabling the target vehicle to collect audio signals with higher clarity. As a result, the target vehicle can accurately extract wake-up words from the audio signals with higher clarity, thereby improving the voice wake-up success rate of the target vehicle.
[0027] Please refer to Figure 1 , Figure 1 , which shows a schematic diagram of the implementation environment of a vehicle audio control method proposed in this application. In this implementation environment, it includes vehicle 100 and in-vehicle intelligent voice system 110 provided on this vehicle 100. In-vehicle intelligent voice system 110 includes in-vehicle microphone 111, processor 112, and memory 113. In-vehicle display screen 111 can be placed in the internal space of vehicle 100. The placement methods of in-vehicle microphone 111 include being embedded in the vehicle interior, suspended in the vehicle interior, being wiredly connected to the vehicle, and being wirelessly connected to the vehicle, etc. In-vehicle microphone 111 can be used for voice input by the driver or passengers.
[0028] Furthermore, vehicle 100 may include a system architecture such as Figure 2 . In the system architecture, it may include acquisition device 120, control host 130, and in-vehicle display screen 111. Among them, acquisition device 120 may include a microphone for collecting audio signals. Control host 130 is responsible for processing the collected audio signals. Specifically, control host 130 can control the vehicle to output corresponding actions or control the display of the display area of in-vehicle display screen 111 according to the voice signals collected by acquisition device 120. Information transmission and control are carried out between acquisition device 120, control host 130, and in-vehicle display screen 111 through LVDS (Low-Voltage Differential Signaling) or GMSL (Gigabit Multimedia Serial Link) signals.
[0029] Please refer to Figure 3 , Figure 3 , which shows a flowchart of a vehicle audio control method proposed in an embodiment of this application. The method can be used for Figure 1 vehicle 100 in Figure 1 (or can also be used for
[0030] in-vehicle intelligent voice system 110 in
[0031] In this embodiment, the target vehicle can be any type of vehicle, such as a sedan, an SUV, a bus, a truck, etc. The target vehicle can be equipped with an in-vehicle intelligent voice system, which includes a voice wake-up module and a voice control module. Among them, the voice wake-up module is used to turn on the voice control module according to the wake-up word spoken by the user, and the voice control module is used to extract the operation command from the voice information spoken by the user and control the target vehicle to execute the operation command.
[0032] The in-vehicle intelligent voice system can include multiple microphones, and the multiple microphones can be set at different positions of the vehicle. For example, the microphones can be set near the driver's seat of the vehicle (for collecting the voice information of the driver), near the co-driver (for collecting the voice information of the co-driver), or near the passenger seat (for collecting the voice information of the passenger).
[0033] After the target vehicle is started, the voice wake-up module of the target vehicle is started, and the audio signal in the carriage is collected through the microphone of the in-vehicle intelligent voice system as the target audio signal. The microphone of the in-vehicle intelligent voice system can collect the audio signal in real time, and the real-time collected audio signal can be used as the target audio signal, and then the target audio signal is processed in real time.
[0034] If the target audio signal includes a wake-up word, control the voice control module of the in-vehicle intelligent voice system to start. If the target audio signal does not include a wake-up word, control the voice wake-up module of the in-vehicle intelligent voice system to continue running. If no valid operation command is received within the preset duration after controlling the voice control module of the in-vehicle intelligent voice system to start, control the voice control module of the in-vehicle intelligent voice system to close, and control the voice wake-up module of the in-vehicle intelligent voice system to continue running.
[0035] The wake-up word can be a word set by the user based on needs, and can be any type of wake-up word, such as "Xiaobai", "Xiaoxin", etc. For the same vehicle, different users can set different wake-up words, and different vehicles can also have different wake-up words.
[0036] S120. When the target audio signal meets the preset conditions, determine the classification result of the target audio signal through the sound classification model as the target classification result.
[0037] In this embodiment, the preset condition can be that the target audio signal includes at least human voice. In some possible implementation manners, the preset condition can also be that the target audio signal includes human voice and noise (the noise can also be noise with a volume reaching the noise threshold). Among them, the noise threshold can be a value set based on the actual scenario and requirements, and this application does not make specific limitations.
[0038] Based on the target audio signal, it can be determined whether the target audio signal includes human voices, whether it includes noise, and whether the noise volume reaches the noise threshold. Then, it can be further determined whether the target audio signal meets the preset conditions. When the target audio signal meets the preset conditions, it indicates that the target audio signal includes human voices, and the target audio signal may include a wake word. Further analysis of the target audio signal is required, and the classification result of the target audio signal can continue to be determined through a sound classification model. When the target audio signal does not meet the preset conditions, it indicates that the target audio signal does not include human voices and does not include a wake word, so no further analysis of the target audio signal is required.
[0039] The sound classification model refers to a model for classifying the target audio signal. Through the sound classification model, the classification result of the target audio signal can be determined. The training method of the sound classification model can include: obtaining training samples, where the training samples include sample audios corresponding to multiple classification results and the annotation information corresponding to the sample audios; training an initial model based on the sample audios and the annotation information to obtain the sound classification model. Among them, the classification result of the audio signal can include at least one of the sound signal including noise, the sound signal including living human voices, the sound signal including background music, and the sound signal including non-living human voices. Living human voices refer to the voices of people in the carriage, and non-living human voices refer to the human voices played by other devices (such as radios, mobile phones, etc.).
[0040] The sample audio can refer to the audio signal used as a training sample, and the annotation information corresponding to the sample audio can refer to the manually annotated classification result of the sample audio. For example, when the sample audio includes non-living human voices, the corresponding annotation information is including non-living human voices.
[0041] The initial model used to train the human voice classification model can be a traditional decision tree model, an attention neural network model, a resnet neural network model, etc.
[0042] Input the target voice information into the trained sound classification model, and obtain the classification result output by the sound classification model as the target classification result. The possible classification results of the same target voice information may include at least one of including living human voices, including non-living human voices, including noise, and including background music. For example, a target voice information includes background music, living human voices, and non-living human voices, or a target voice information includes non-living human voices.
[0043] As an implementation, determining the classification result of the target audio signal through a sound classification model as the target classification result includes: segmenting the target audio signal according to the silent intervals in the target audio signal to obtain multiple audio segments; classifying each audio segment through the sound classification model to obtain the classification result of each audio segment; and obtaining the target classification result according to the classification results of the multiple audio segments respectively.
[0044] Under normal circumstances, background music and noise are continuous segments, while the human voice (living human voice and non-living human voice) part has short silent intervals (the silent interval can refer to the interval between each character when a person is speaking, and the character can refer to a Chinese character, word, English word, etc.). Therefore, when segmenting the target audio signal, the duration of each audio segment needs to be determined according to the actual test situation.
[0045] As an implementation, the target audio signal can be segmented to obtain multiple audio segments, each audio segment is classified through a sound classification model to obtain the classification result of each audio segment, and when the proportion of a certain classification result reaches half, this classification result is used as the target classification result of the target audio signal.
[0046] For example, the duration of the collected target audio information is 3s, and the segmentation duration is 500ms. The target audio is segmented into 6 audio segments, and the 6 audio segments are respectively input into the sound classification model to obtain the classification results. Among them, the classification result of audio segment 1 includes living human voice, background music, and noise, the classification result of audio segment 2 includes living human voice and background music, the classification result of audio segment 3 includes living human voice and noise, the classification result of audio segment 4 includes living human voice, background music, and noise, the classification result of audio segment 5 includes living human voice, background music, and non-living human voice, and the classification result of audio segment 6 includes living human voice and background music. Among them, the proportion of including living human voice is 6 / 6, reaching half, the proportion of including background music is 5 / 6, reaching half, the proportion of including noise is 1 / 2, reaching half, and the proportion of including non-living human voice is 1 / 6, not reaching half. Therefore, the target classification result is that the target audio signal includes living human voice, background music, and noise.
[0047] S130. Determine the target sound pickup strategy corresponding to the target classification result in the preset strategy set. The preset strategy set includes multiple classification results and the sound pickup strategies respectively corresponding to the multiple classification results, and each sound pickup strategy in the preset strategy set is used to improve the sound pickup effect in the carriage.
[0048] Each classification result among multiple classification results of the preset policy set includes at least one of the following: the audio signal includes noise, the audio signal includes background music, the audio signal includes a live human voice, and the audio signal includes a non-live human voice. At this time, the sound pickup strategies in the preset policy set may include the following:
[0049] 1. When the classification result is that the audio signal includes noise and a live human voice, the corresponding sound pickup strategy is to control the vehicle's windows to enter the window-closing state, and determine the noise reduction parameters for noise reduction of the real-time audio signal collected in the vehicle cabin according to the real-time audio signal collected in real time. Among them, the noise reduction parameters can be selected to be relatively strict to improve the noise reduction effect. Controlling the vehicle to enter the window-closing state reduces the influence of the external environment noise of the vehicle on the inside of the vehicle cabin, making the noise in the vehicle cabin less and the vehicle cabin quieter.
[0050] 2. When the classification result is that the audio signal includes background music and a live human voice, the corresponding sound pickup strategy is to turn down the volume of the vehicle's background music.
[0051] 3. When the classification result is that the audio signal includes a live human voice and a non-live human voice, the corresponding sound pickup strategy is to turn down the volume of the vehicle's background music if the vehicle outputs background music, and determine the sound source position of the audio signal, and turn on the audio collection device matching the sound source position. Among them, the audio collection device matching the sound source position refers to the audio collection device closest to the sound source position, so that the audio collection device can accurately collect the live human voice. The audio collection device can refer to devices such as microphones.
[0052] When the classification result is that the audio signal includes noise and a non-live human voice, in this case, even if the wake-up word is included, it is not the wake-up word directly sent by the user, but the wake-up word that appears in the voice information played by other devices, and there is no operation for the user to perform voice wake-up, so there is no special processing behavior; when the classification result is that the audio signal includes a non-live human voice and no live human voice, this situation may be that the vehicle plays an audio containing a human voice or the voice of the other party in a hands-free call, etc., and there is no operation for the user to perform voice wake-up, so there is no special processing behavior; when the classification result is that it includes background music and a non-live human voice, the voice information in this case also does not have an operation for the user to perform voice wake-up, so there is no special processing behavior.
[0053] A classification result identical to the target classification result can be determined in the preset policy set, and the sound pickup strategy corresponding to the determined classification result of the window can be determined as the target sound pickup strategy.
[0054] S130. Through the target sound pickup strategy, control the target vehicle to output corresponding actions.
[0055] Optionally, when the target classification result is that the target audio signal includes noise and human voice of a living body, controlling the target vehicle to output a corresponding action through the target pickup strategy includes: controlling the window of the target vehicle to enter a window-closing state, and determining a noise reduction parameter for reducing noise of the real-time audio signal according to the collected real-time audio signal in the carriage; when the target classification result is that the target audio signal includes background music and human voice of a living body, controlling the target vehicle to output a corresponding action through the target pickup strategy includes: turning down the volume of the background music of the target vehicle; when the target classification result is that the target audio signal includes human voice of a living body and non-living human voice, controlling the target vehicle to output a corresponding action through the target pickup strategy includes: if the target vehicle outputs background music, turning down the volume of the background music of the target vehicle, and determining the sound source position of the target audio signal, and turning on an audio collection device matching the sound source position.
[0056] When the target audio signal includes human voice, the vehicle is controlled through the target pickup strategy corresponding to the target classification result in the above different situations, the influence of external noise and background music is reduced, and at the same time, the collection accuracy of the human voice of a living body can be improved, thereby improving the pickup effect of the target vehicle and making the clarity of the audio signal collected in the carriage relatively high.
[0057] In summary, in this embodiment, the target pickup strategy is determined through the target classification result corresponding to the target audio signal, and the vehicle is controlled through the target pickup strategy, so as to improve the pickup effect in the carriage of the target vehicle, so that the target vehicle can collect an audio signal with relatively high clarity, and thus the target vehicle can accurately extract a wake-up word from the audio signal with relatively high clarity, improving the voice wake-up success rate of the target vehicle.
[0058] Please refer to Figure 4 , Figure 4 shows Figure 3 a flowchart of an implementation manner before S120 in Figure 1 The method can be used for Figure 1 the vehicle 100 in
[0059] S210. Determine the noise suppression, clipping information and signal-to-noise ratio corresponding to the target audio signal.
[0060] The signal-to-noise ratio, also known as SNR or S / N (SIGNAL-NOISE RATIO), is the ratio of the signal to the noise in a vehicle or electronic system. The signal here refers to the electronic signal from outside the device that needs to be processed by this device, and the noise refers to the irregular extra signal (or information) that does not exist in the original signal and is generated after passing through this device, and this kind of signal does not change with the change of the original signal. The higher the signal-to-noise ratio, the less noise; the lower the signal-to-noise ratio, the more noise.
[0061] The phenomenon that the amplitude of the signal waveform is too large and exceeds the linear range of the system is called clipping. Clipping is the process of limiting the amplitude of the signal to a certain fixed maximum value. Sometimes, it is also called amplitude limiting.
[0062] Voice Activity Detection (VAD), also known as speech activity detection. The purpose of voice activity detection is to identify and eliminate long silent periods from the sound signal stream to save voice channel resources without degrading the service quality. It is an important part of IP phone applications. Voice activity detection can save valuable bandwidth resources and help reduce the end-to-end delay perceived by users.
[0063] S220. Obtain the noise analysis result of the target audio signal according to the signal-to-noise ratio.
[0064] A signal-to-noise ratio threshold can be set. When the signal-to-noise ratio is lower than the signal-to-noise ratio threshold, obtain the noise analysis result of the target audio signal including noise. When the signal-to-noise ratio is not lower than the signal-to-noise ratio threshold, obtain the noise analysis result of the target audio signal without noise.
[0065] S230. Obtain the noise volume analysis result of the target audio signal according to the clipping information.
[0066] A clipping threshold can be set. The clipping threshold corresponds to the noise threshold. When the clipping information reaches the clipping threshold, it is determined that the noise of the target audio signal reaches the noise threshold, and obtain the noise volume analysis result of the target audio signal when the noise reaches the noise threshold. When the clipping information does not reach the clipping threshold, it is determined that the noise of the target audio signal does not reach the noise threshold, and obtain the noise volume analysis result of the target audio signal when the noise does not reach the noise threshold.
[0067] S240. Determine the voice analysis result of the target audio signal according to the voice activity detection.
[0068] The voice activity detection can be analyzed to determine whether the target audio signal includes voice, so as to obtain the voice analysis result that the target audio signal includes voice or the voice analysis result that the target audio signal does not include voice.
[0069] S250. Determine whether the target audio signal meets the preset condition according to the noise analysis result, the human voice analysis result, and the noise volume analysis result.
[0070] Based on the noise analysis result, the human voice analysis result, and the noise volume analysis result, determine whether the target audio signal meets the preset condition of the above embodiment. If the preset condition is met, perform S120 for maintenance. If the preset condition is not met, return to execute S110.
[0071] In this embodiment, through mute suppression, clipping information, and signal-to-noise ratio, noise analysis results, human voice analysis results, and noise volume analysis results with relatively high accuracy can be obtained. Furthermore, it can be accurately determined whether the target audio signal meets the preset condition, improving the accuracy of vehicle audio control.
[0072] Please refer to Figure 5 , Figure 5 which shows a block diagram of a vehicle audio control device proposed in an embodiment of the present application. The device 300 includes:
[0073] An acquisition module 310, configured to acquire an audio signal in the carriage of the target vehicle as the target audio signal;
[0074] A classification module 320, configured to, when the target audio signal meets the preset condition, determine a classification result of the target audio signal through a sound classification model as the target classification result;
[0075] A determination module 330, configured to determine a target sound pickup strategy corresponding to the target classification result in a preset strategy set. The preset strategy set includes multiple classification results and the respective sound pickup strategies corresponding to the multiple classification results. Each sound pickup strategy in the preset strategy set is used to improve the sound pickup effect in the carriage;
[0076] A control module 340, configured to control the target vehicle to output a corresponding action through the target sound pickup strategy.
[0077] Optionally, the device further includes an analysis module, configured to determine the mute suppression, clipping information, and signal-to-noise ratio corresponding to the target audio signal; obtain a noise analysis result of the target audio signal according to the signal-to-noise ratio; obtain a noise volume analysis result in the target audio signal according to the clipping information; determine a human voice analysis result of the target audio signal according to the mute suppression; and determine whether the target audio signal meets the preset condition according to the noise analysis result, the human voice analysis result, and the noise volume analysis result.
[0078] Optionally, the classification module 320 is further configured to segment the target audio signal according to the silence intervals in the target audio signal to obtain a plurality of audio segments; classify each of the audio segments through the voice classification model to obtain the classification result of each of the audio segments; and obtain the target classification result according to the classification results of the plurality of audio segments respectively.
[0079] Optionally, the apparatus further includes a training module, configured to obtain training samples, where the training samples include the sample audio corresponding to each of the plurality of classification results and the annotation information corresponding to the sample audio; and train an initial model according to the sample audio and the annotation information to obtain the voice classification model.
[0080] Optionally, when the target classification result is that the target audio signal includes noise and human voice of a living body, the control module 340 is further configured to control the window of the target vehicle to enter a window closing state, and determine a noise reduction parameter for reducing noise of the real-time audio signal according to the real-time audio signal collected in the vehicle compartment; when the target classification result is that the target audio signal includes background music and human voice of a living body, the control module 340 is further configured to lower the volume of the background music of the target vehicle; when the target classification result is that the target audio signal includes human voice of a living body and human voice of a non-living body, the control module 340 is further configured to, if the target vehicle outputs background music, lower the volume of the background music of the target vehicle, and determine the sound source position of the target audio signal, and turn on an audio collection device matching the sound source position.
[0081] It should be noted that the apparatus embodiments in this application correspond to the foregoing method embodiments. The specific principles in the apparatus embodiments can be referred to the content in the foregoing method embodiments, and will not be elaborated here.
[0082] Figure 6 shows a structural block diagram of a vehicle for executing the vehicle audio control method according to an embodiment of the present application. As Figure 6As shown, vehicle 1200 includes a Central Processing Unit (CPU) 1201, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 1202 or programs loaded from a storage section 1208 into a Random Access Memory (RAM) 1203, such as executing the methods in the above embodiments. In the RAM 1203, various programs and data required for system operations are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An Input / Output (I / O) interface 1205 is also connected to the bus 1204.
[0083] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.
[0084] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a Central Processing Unit (CPU) 1201, various functions defined in the system of the present application are executed.
[0085] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0087] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.
[0088] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the vehicle described in the above embodiments; or may exist alone without being assembled into the vehicle. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.
[0089] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. The processor of the vehicle reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the vehicle executes the method in any of the above embodiments.
[0090] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0091] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0092] Other embodiments of the present application will be readily conceived by those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A vehicle audio control method, characterized in that, The method includes: Collecting an audio signal inside the compartment of the target vehicle as the target audio signal; Determining the voice activity detection, clipping information, and signal-to-noise ratio corresponding to the target audio signal; Obtaining a noise analysis result of the target audio signal according to the signal-to-noise ratio; Obtaining a noise volume analysis result in the target audio signal according to the clipping information; Determining a voice analysis result of the target audio signal according to the voice activity detection; Determining whether the target audio signal meets a preset condition according to the noise analysis result, the voice analysis result, and the noise volume analysis result; When the target audio signal meets the preset condition, determining a classification result of the target audio signal through a sound classification model as the target classification result; Determining a target pickup strategy corresponding to the target classification result in a preset strategy set, where the preset strategy set includes multiple classification results and the pickup strategies respectively corresponding to the multiple classification results, and each pickup strategy in the preset strategy set is used to improve the pickup effect inside the compartment; Controlling the target vehicle to output a corresponding action through the target pickup strategy.
2. The method according to claim 1, characterized in that, The determining a classification result of the target audio signal through a sound classification model as the target classification result includes: Segmenting the target audio signal according to a silence interval in the target audio signal to obtain multiple audio segments; Classifying each audio segment through the sound classification model to obtain a classification result of each audio segment; Obtaining the target classification result according to the classification results of the multiple audio segments respectively.
3. The method according to claim 1, characterized in that, The training method of the sound classification model includes: Obtaining training samples, where the training samples include sample audios respectively corresponding to the multiple classification results and annotation information corresponding to the sample audios; Training an initial model according to the sample audios and the annotation information to obtain the sound classification model.
4. The method according to claim 1, wherein when the target classification result is that the target audio signal includes noise and a living voice, the controlling the target vehicle to output a corresponding action through the target pickup strategy includes: Controlling the windows of the target vehicle to enter a window-closing state, and determining a noise reduction parameter for noise reduction of the real-time audio signal collected inside the compartment according to the real-time audio signal collected inside the compartment; when the target classification result is that the target audio signal includes background music and a living voice, the controlling the target vehicle to output a corresponding action through the target pickup strategy includes: Lowering the volume of the background music of the target vehicle; when the target classification result is that the target audio signal includes a living voice and a non-living voice, the controlling the target vehicle to output a corresponding action through the target pickup strategy includes: If the target vehicle outputs background music, lowering the volume of the background music of the target vehicle, and determining the sound source position of the target audio signal, and turning on an audio collection device matching the sound source position.
5. The method according to claim 1, characterized in that The preset condition includes: The target audio signal includes at least a voice.
6. The method according to claim 1, characterized in that, Each of the multiple classification results includes at least one of the audio signal including noise, the audio signal including background music, the audio signal including a living human voice, and the audio signal including a non-living human voice.
7. A vehicle audio control device, characterized in that, The device includes: An acquisition module, configured to acquire an audio signal inside the compartment of a target vehicle as a target audio signal; determine noise suppression, clipping information, and signal-to-noise ratio corresponding to the target audio signal; obtain a noise analysis result of the target audio signal according to the signal-to-noise ratio; obtain a noise volume analysis result in the target audio signal according to the clipping information; determine a human voice analysis result of the target audio signal according to the noise suppression; and determine whether the target audio signal meets a preset condition according to the noise analysis result, the human voice analysis result, and the noise volume analysis result. A classification module, configured to, when the target audio signal meets the preset condition, determine a classification result of the target audio signal through a sound classification model as a target classification result. A determination module, configured to determine a target pickup strategy corresponding to the target classification result in a preset strategy set, where the preset strategy set includes multiple classification results and pickup strategies respectively corresponding to the multiple classification results, and each pickup strategy in the preset strategy set is used to improve the pickup effect inside the compartment. A control module, configured to control the target vehicle to output a corresponding action through the target pickup strategy.
8. A vehicle, characterized in that, including: One or more processors; A memory; One or more applications, where the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method and device for collecting voice signals, electronic equipment and readable storage medium
CN108156291A
In-vehicle device and sound collecting method thereof
JP2013019807A