Vehicle-mounted voice processing method and system, electronic device, and storage medium

By combining a microphone array and an audio processing module, the problem of recognition errors caused by single-microphone acquisition in in-vehicle voice systems was solved, achieving high-quality in-vehicle voice acquisition and recognition, and improving the user experience.

CN118612598BActive Publication Date: 2026-01-02CHERY NEW ENERGY AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410638707.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-22
Publication Date
2026-01-02
Estimated Expiration
2044-05-22

AI Technical Summary

Technical Problem

Existing in-vehicle voice systems use a single microphone to collect voice data, resulting in frequent voice recognition errors and failing to guarantee the voice quality and call experience for drivers and passengers.

Method used

The vehicle-mounted voice processing system, consisting of a microphone array, an audio processing module, and a communication module, collects multiple voice signals through the microphone array, performs voice processing and noise suppression through the audio processing module, and sends clear voice signals through the communication module, thereby enabling automatic selection of the voice signal of the target sound source.

Benefits of technology

It improves the quality and accuracy of in-vehicle voice acquisition and recognition, thereby enhancing the vehicle's intelligence level and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118612598B_ABST
    Figure CN118612598B_ABST
Patent Text Reader

Abstract

The application provides a vehicle-mounted voice processing method and system, electronic equipment and a storage medium, and belongs to the technical field of automobile electronics. The method comprises the following steps: collecting voice through a microphone array to obtain a plurality of first voice signals, the microphone array comprising a plurality of microphones, the plurality of first voice signals corresponding to the plurality of microphones one by one, the microphone array being connected with an audio processing module; performing voice processing on the plurality of first voice signals through the audio processing module to obtain a second voice signal, the second voice signal being a voice signal of a target sound source object, the target sound source object being a sound source object for voice communication through the microphone array, the audio processing module being connected with a communication module; and sending the second voice signal to a communication terminal of the target sound source object through the communication module. The scheme can improve the quality of vehicle-mounted voice collection, thereby improving the accuracy of vehicle-mounted voice recognition and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automotive electronics application, and in particular to a vehicle-mounted voice processing method and system, an electronic device, and a storage medium. BACKGROUND

[0002] With the rapid development of technology and the advancement of intelligent trend, the importance of vehicle-mounted voice system is increasingly prominent, bringing a more convenient, safe and efficient experience for drivers and passengers. For example, through the vehicle-mounted voice system, drivers and passengers can realize various types of operations such as vehicle component control, vehicle entertainment control, voice navigation, and telephone communication. However, most of the current vehicle-mounted voice systems use a single microphone for voice collection, and traditional vehicle-mounted voice systems often have problems such as voice recognition errors or inability to recognize voice, making it difficult to provide more accurate voice recognition support for drivers and passengers, and unable to guarantee the voice quality when drivers and passengers make vehicle calls, thus unable to guarantee the call experience of drivers and passengers. Therefore, how to build a vehicle-mounted voice system to improve the accuracy of vehicle-mounted voice recognition is a problem to be solved. SUMMARY

[0003] The embodiments of the present application provide a vehicle-mounted voice processing method and system, an electronic device, and a storage medium, which can improve the quality of vehicle-mounted voice collection and thus improve the accuracy of vehicle-mounted voice recognition. Accordingly, the present application can not only improve the intelligent level of the vehicle, but also improve the user experience and satisfaction. The technical solution is as follows:

[0004] On the one hand, a vehicle-mounted voice processing method is provided, which is applied to a vehicle-mounted voice processing system, the vehicle-mounted voice processing system comprising a microphone array, an audio processing module, and a communication module, and the method comprising:

[0005] collecting voice through the microphone array to obtain a plurality of first voice signals, the microphone array comprising a plurality of microphones, the plurality of first voice signals corresponding to the plurality of microphones one by one, and the microphone array being connected to the audio processing module;

[0006] performing voice processing on the plurality of first voice signals through the audio processing module to obtain a second voice signal, the second voice signal being a voice signal of a target sound source object, the target sound source object being a sound source object making a voice call through the microphone array, and the audio processing module being connected to the communication module;

[0007] sending the second voice signal to a communication terminal of the target sound source object through the communication module.

[0008] In some embodiments, the speech processing of the plurality of first speech signals by the audio processing module comprises:

[0009] In the case of collecting the speech of a single sound source object, the audio processing module determines the distance between the microphone corresponding to each of the plurality of first speech signals and the sound source object, respectively;

[0010] The audio processing module determines the second speech signal from the plurality of first speech signals, the distance between the microphone corresponding to the second speech signal and the sound source object being smaller than the distance between the microphone corresponding to other speech signals and the sound source object, the other speech signals being speech signals other than the second speech signal.

[0011] In some embodiments, the speech processing of the plurality of first speech signals by the audio processing module comprises:

[0012] In the case of collecting the speech of a plurality of sound source objects, the audio processing module determines the target sound source object from the plurality of sound source objects;

[0013] The audio processing module determines the distance between the microphone corresponding to each of the plurality of first speech signals and the target sound source object, respectively;

[0014] The audio processing module determines the second speech signal from the plurality of first speech signals, the distance between the microphone corresponding to the second speech signal and the target sound source object being smaller than the distance between the microphone corresponding to other speech signals and the target sound source object, the other speech signals being speech signals other than the second speech signal.

[0015] In some embodiments, the audio processing module determines the target sound source object from the plurality of sound source objects in the case of collecting the speech of a plurality of sound source objects, comprising:

[0016] The audio processing module determines the speech duration of each of the plurality of sound source objects;

[0017] The sound source object with the longest speech duration is determined as the target sound source object.

[0018] In some embodiments, the audio processing module determines the target sound source object from the plurality of sound source objects in the case of collecting the speech of a plurality of sound source objects, comprising:

[0019] The audio processing module determines the speech feature parameters of each of the plurality of sound source objects;

[0020] determining a sound source object corresponding to the voice feature parameters existing in the feature database as the target sound source object, the feature database being used to store voice feature parameters of a pre-recorded voice signal.

[0021] In some embodiments, in a case where a plurality of sound source objects are collected, the target sound source object is determined among the plurality of sound source objects by the audio processing module, including:

[0022] determining a position of each of the plurality of sound source objects by the audio processing module;

[0023] determining a sound source object at a preset position as the target sound source object.

[0024] In some embodiments, the method further includes:

[0025] performing voice processing on the second voice signal by the audio processing module based on the other voice signals, to obtain an optimized second voice signal, the optimized second voice signal being a voice signal of the target sound source object after filtering out noise signals, the noise signals including voice signals of other sound source objects, environmental noise signals, and echo signals, the other sound source objects being sound source objects other than the target sound source object among the plurality of sound source objects.

[0026] In some embodiments, the vehicle-mounted voice processing system further includes a plurality of physical buttons for controlling the microphone array, the plurality of physical buttons corresponding to the plurality of microphones one-to-one, and the method further includes:

[0027] in a case where the vehicle-mounted voice processing system is in a multi-microphone working mode, determining at least two microphones that are working, the multi-microphone working mode being used to indicate that the vehicle-mounted voice processing system supports a plurality of microphones to collect voice signals at the same time;

[0028] in response to a triggering operation on a first physical button, collecting voice signals by the at least two microphones and a microphone corresponding to the first physical button, the first physical button being a physical button corresponding to a microphone other than the at least two microphones.

[0029] In some embodiments, the vehicle-mounted voice processing system further includes a plurality of physical buttons for controlling the microphone array, the plurality of physical buttons corresponding to the plurality of microphones one-to-one, and the method further includes:

[0030] In a case that the vehicle-mounted voice processing system is in a single microphone working mode, a single microphone currently working is determined, the single microphone working mode is used to indicate that the vehicle-mounted voice processing system only supports a single microphone to collect voice;

[0031] In response to a triggering operation on any second physical key, a single voice signal is obtained by collecting voice through a microphone corresponding to the second physical key, the second physical key being a physical key corresponding to a microphone other than the single microphone; or

[0032] In response to triggering operations on at least two second physical keys at the same time, a single voice signal is obtained by collecting voice through a microphone corresponding to a second physical key with the highest priority based on priorities of the at least two second physical keys.

[0033] In some embodiments, the method further comprises:

[0034] The collected voice signals are subjected to semantic recognition by the audio processing module, the semantic recognition being used to determine whether the voice signals contain target instructions, the target instructions being used to instruct to turn on a microphone in the microphone array;

[0035] In a case that the plurality of first voice signals contain the target instructions, at least one target microphone corresponding to the target instructions is determined based on at least one microphone identifier in the target instructions, different microphones in the microphone array corresponding to different microphone identifiers;

[0036] Voice is collected through the at least one target microphone to obtain at least one voice signal.

[0037] In another aspect, a vehicle-mounted voice processing system is provided, characterized in that the system comprises a plurality of physical keys, a microphone array, an audio processing module, a communication module, a power amplifier module, a vehicle-mounted loudspeaker, and a control system;

[0038] The plurality of physical keys are used to control the microphone array, the plurality of physical keys being connected with the microphone array;

[0039] The microphone array is used to collect voice to obtain voice signals, the microphone array comprising a plurality of microphones, the microphone array being connected with the audio processing module;

[0040] The audio processing module is used to process the voice signals to obtain processed voice signals, the audio processing module being connected with the communication module;

[0041] The communication module is configured to send the processed voice signal to other devices and receive voice signals sent by other devices, and is connected with the power amplifier module.

[0042] The power amplifier module is configured to amplify the received voice signal to obtain an amplified voice signal, and is connected with the vehicle-mounted loudspeaker.

[0043] The vehicle-mounted loudspeaker is configured to play the amplified voice signal.

[0044] The control system is configured to control the plurality of modules in the system.

[0045] In another aspect, an electronic device is provided, which includes a processor and a memory configured to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to implement the vehicle-mounted voice processing method in the embodiments of the present application.

[0046] In another aspect, a computer readable storage medium is provided, which stores at least one piece of computer program, the at least one piece of computer program being loaded and executed by a processor to implement the vehicle-mounted voice processing method in the embodiments of the present application.

[0047] In another aspect, a computer program product is provided, which includes a computer program executed by a processor to implement the vehicle-mounted voice processing method in the embodiments of the present application.

[0048] The present application provides a vehicle-mounted voice processing method, which can automatically select voice signals related to a target sound source object in the process of vehicle-mounted voice communication through a vehicle-mounted voice processing system composed of a microphone array, an audio processing module, a communication module and the like. Compared with the traditional scheme of fixedly using a single microphone to collect voice signals in the vehicle, the present application can improve the quality of vehicle-mounted voice collection and thus improve the accuracy of vehicle-mounted voice recognition. Accordingly, the present application can not only improve the intelligent level of the vehicle, but also improve the user experience and satisfaction. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0050] Figure 1 is a hardware structure schematic diagram of a vehicle-mounted voice processing system provided by the embodiments of the present application;

[0051] Figure 2 is a flowchart of a vehicle-mounted voice processing method according to an embodiment of the present application;

[0052] Figure 3 is a flowchart of another vehicle-mounted voice processing method according to an embodiment of the present application;

[0053] Figure 4 is a structural schematic diagram of an electronic device according to an embodiment of the present application;

[0054] Figure 5 is a structural schematic diagram of a server according to an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the objects, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0056] In the present application, the terms "first", "second", and the like are used to distinguish the same or similar items with substantially the same function and action, and it should be understood that there is no logical or time sequence relationship between "first", "second", and "nth", and the number and execution order are not limited.

[0057] In the present application, the term "at least one" means one or more, and the term "multiple" means two or more.

[0058] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the voice signal involved in the present application is obtained under sufficient authorization.

[0059] Figure 1 is a hardware structural schematic diagram of a vehicle-mounted voice processing system according to an embodiment of the present application. Referring to Figure 1 , the hardware structure of the vehicle-mounted voice processing system includes a plurality of physical keys, a microphone array, an audio processing module, a communication module, a power amplifier module, a vehicle-mounted loudspeaker, and a control system.

[0060] For example, the plurality of physical buttons 101 are used to control the microphone array 102, and the plurality of physical buttons 101 are connected with the microphone array 102. The microphone array 102 is used to collect voice to obtain voice signals, and the microphone array 102 includes a plurality of microphones, and the microphone array 102 is connected with the audio processing module 103. Single or multiple microphones closer to the driver are selected by the plurality of physical buttons 101, voice control or system settings to collect voice input, so as to improve the effect of voice recognition and Bluetooth call. The single, double or multiple microphones are controlled by the single press, long press or double click of the plurality of physical buttons 101 to collect voice signals, so as to improve the effect of voice recognition and Bluetooth call.

[0061] The audio processing module 103 is used to perform voice processing on the voice signals to obtain processed voice signals, and the audio processing module 103 is connected with the communication module 104. The audio processing module 103 includes an audio noise reduction module 1031 and a DSP decoding module 1032, and the DSP decoding module 1032 can be directly connected with the power amplifier module 105. The audio noise reduction module 1031 compares or captures the sound in the environment, digitizes and encodes it into an electrical signal, then uses analysis and processing algorithms to separate the voice and noise, and eliminates noise and echo, so as to output single, double or multiple digital voice signals with high quality and annoying noise reduction. The voice processing is realized by the DSP decoding module 1032. The voice processing includes voice recognition, language understanding, voice synthesis, voice enhancement and voice data compression, etc. Among them, the voice recognition is to extract the feature parameters of the voice signal to be recognized in real time, and match with known voice samples, so as to determine the phoneme attribute of the voice signal to be recognized. The voice recognition method includes statistical pattern voice recognition, structure and sentence pattern voice recognition, etc. The above method can obtain important parameters such as formant frequency, pitch, voice and noise. Voice understanding is the theoretical and technical basis for people and computers to communicate in natural language. The main purpose of voice synthesis is to enable computers to speak. Through decoding, the voice signal collected by the microphone can be restored into a high-guaranteed analog audio signal output to the power amplifier module 105.

[0062] The communication module 104 is configured to send the processed voice signal to other devices and receive voice signals sent by other devices, and the communication module 104 is connected with the power amplifier module 105. The power amplifier module 105 is configured to amplify the received voice signal to obtain an amplified voice signal, and the power amplifier module 105 is connected with the vehicle-mounted loudspeaker 106. For example, the high-pitched, medium-pitched and low-pitched analog audio signals restored by the DSP decoding module 1032 are amplified by the power amplifier module 105 and output to the vehicle-mounted loudspeaker 106. The vehicle-mounted loudspeaker 106 is configured to play the amplified voice signal. That is, the high-pitched, medium-pitched and low-pitched signals amplified and restored by the power amplifier module 105 are output to the vehicle-mounted loudspeaker 106. In addition, the SOC control module 107 is configured to switch and control the voice system function and the multimedia function mode. The power supply module 108 is configured to provide stable and reliable power supply for each part of the system.

[0063] Figure 2 A flowchart of a vehicle-mounted voice processing method is provided according to an embodiment of the present application. The method is executed by a vehicle-mounted voice processing system, which includes a microphone array, an audio processing module and a communication module. Referring to Figure 2 , the method includes the following steps:

[0064] 201. Collecting voice through the microphone array to obtain a plurality of first voice signals, the microphone array including a plurality of microphones, the plurality of first voice signals corresponding to the plurality of microphones one by one, and the microphone array being connected with the audio processing module.

[0065] In the embodiment of the present application, the microphone array refers to a recording system composed of two or more microphones. For example, a microphone array composed of four microphones is used. The four microphones are arranged at the front left position, the front right position, the rear left position and the rear right position in the vehicle, respectively. By uniformly arranging a plurality of microphones, a better voice collection effect can be obtained. After four first voice signals are collected through the microphone array, the four first voice signals are sent to the audio processing module for voice processing.

[0066] 202. Voice processing the plurality of first voice signals through the audio processing module to obtain a second voice signal, the second voice signal being a voice signal of a target sound source object, and the target sound source object being a sound source object for voice communication through the microphone array, and the audio processing module being connected with the communication module.

[0067] In the embodiment of the present application, the audio processing module is used to receive the multiple voice signals sent by the microphone array and perform voice processing thereon. The voice processing includes noise reduction processing, voice recognition, voice understanding, voice synthesis, voice enhancement, voice data compression, etc., which are not limited in the embodiment of the present application. The second voice signal is used to indicate the voice of the clear target sound source object in the first voice signal. By separating the second voice signal and transmitting it to the communication terminal, the interference of other noise signals can be excluded, and the quality of the vehicle-mounted voice call can be improved.

[0068] 203. Transmitting the second voice signal to the communication terminal of the target sound source object through the communication module.

[0069] In the embodiment of the present application, the communication terminal can be wirelessly connected to the terminal where the vehicle-mounted voice processing system is located through Bluetooth. After receiving the second voice signal sent by the audio processing, the communication module transmits it to the communication terminal, and the communication terminal transmits the voice signal to the call opposite end, thereby realizing the vehicle-mounted voice call.

[0070] The embodiment of the present application provides a vehicle-mounted voice processing method. Through the vehicle-mounted voice processing system composed of a microphone array, an audio processing module, a communication module and the like, the voice signal related to the target sound source object in the process of the vehicle-mounted voice call can be automatically selected. Compared with the traditional scheme of fixedly using a single microphone to collect the voice signal in the vehicle, the present scheme can improve the quality of the vehicle-mounted voice collection, and further improve the accuracy of the vehicle-mounted voice recognition. Accordingly, the present scheme can not only improve the intelligent level of the vehicle, but also improve the user's use experience and satisfaction.

[0071] Figure 3 is a flowchart of another vehicle-mounted voice processing method provided by the embodiment of the present application. The method is executed by a vehicle-mounted voice processing system, see Figure 3 , which comprises the following steps:

[0072] 301. Collecting voice through a microphone array to obtain multiple first voice signals. The microphone array comprises multiple microphones, and the multiple first voice signals correspond to the multiple microphones one by one. The microphone array is connected with an audio processing module.

[0073] In the embodiment of the present application, the microphone array refers to a recording system composed of two or more microphones. For example, a microphone array composed of four microphones is used. The positions of the four microphones in the vehicle are left front, right front, left rear and right rear respectively. By uniformly arranging multiple microphones, a better voice collection effect can be obtained. The number, position and architecture of the multiple microphones in the microphone array are not limited in the embodiment of the present application.

[0074] It should be noted that the vehicle-mounted voice processing system can control the microphone array to switch to a multi-microphone working mode, that is, support multiple microphones to collect voice at the same time. At this time, multiple voice signals can be collected through the microphone array. The vehicle-mounted voice processing system can also control the microphone array to switch to a single-microphone working mode, that is, only support a single microphone to collect voice. At this time, only a single voice signal can be collected through the microphone array. The embodiments of the present application do not limit this.

[0075] For example, it is assumed that the vehicle-mounted voice processing system is in a multi-microphone working mode, and voice is collected through four microphones at the same time. The voice in the vehicle cabin is collected through the left front microphone, the right front microphone, the left rear microphone, and the right rear microphone to obtain four first voice signals. The four first voice signals are sent to the audio processing module for voice processing.

[0076] In the embodiments of the present application, voice is collected in the above manner to obtain multiple voice signals, which facilitates subsequent voice processing of the multiple voices. Compared with voice processing based on only a single voice, the background noise and echo interference can be more comprehensively and reliably eliminated, the voice clarity and audibility can be improved, and the accuracy and robustness of voice recognition can be improved.

[0077] In some embodiments, the driver or passenger can turn on or turn off multiple microphones in the microphone array by controlling physical buttons or voice control to meet various scene requirements. The following mode A and mode B are described by taking the driver or passenger controlling the microphone to start collecting voice as an example. It should be noted that the driver or passenger controls the microphone to turn off or controls the microphone array to switch the working mode, and the same applies here, which will not be described again.

[0078] Mode A: The driver or passenger controls the physical button. Correspondingly, the vehicle-mounted voice processing system further includes multiple physical buttons, and the multiple physical buttons are used to control the microphone array, and the multiple physical buttons correspond to the multiple microphones one by one.

[0079] Optionally, in a case where the vehicle-mounted voice processing system is in a multi-microphone working mode, the vehicle-mounted voice processing system determines at least two microphones that are working; in response to a triggering operation of the first physical button by the driver or passenger, voice is collected through the at least two microphones and the microphone corresponding to the first physical button to obtain multiple voice signals. The multi-microphone working mode is used to indicate that the vehicle-mounted voice processing system supports multiple microphones to collect voice at the same time, and the first physical button is a physical button corresponding to a microphone other than the at least two microphones.

[0080] For example, the in-vehicle voice processing system is in a multi-microphone working mode, such as the left rear microphone and the right rear microphone are working. At this time, the physical buttons corresponding to the left front microphone and the right front microphone are the first physical buttons. When the driver triggers the physical buttons corresponding to the left front microphone and the right front microphone, the in-vehicle voice processing system turns on the above two microphones. At this time, the in-vehicle voice processing system controls the original left rear microphone and the right rear microphone, and the newly added left front microphone and the right front microphone to simultaneously collect the voice in the vehicle cabin.

[0081] Optionally, in the case that the in-vehicle voice processing system is in a single-microphone working mode, the in-vehicle voice processing system determines the single microphone that is working; in response to the triggering operation of any second physical button by the driver, the voice is collected through the microphone corresponding to the second physical button to obtain a single voice signal; or in response to the triggering operation of at least two second physical buttons by the driver at the same time, based on the priority of the at least two second physical buttons, the voice is collected through the microphone corresponding to the second physical button with the highest priority to obtain a single voice signal. The single-microphone working mode is used to indicate that the in-vehicle voice processing system only supports a single microphone to collect voice, and the second physical button is the physical button corresponding to the microphone other than the single microphone.

[0082] For example, the in-vehicle voice processing system is in a single-microphone working mode, such as only the left rear microphone is working. At this time, the physical buttons corresponding to the right rear microphone, the left front microphone and the right front microphone are the second physical buttons. When the driver triggers the physical button corresponding to the left front microphone, the in-vehicle voice processing system turns off the original left rear microphone and turns on the left front microphone. At this time, the in-vehicle voice processing system only controls the newly added left front microphone to collect the voice in the vehicle cabin. Alternatively, when the driver triggers the physical buttons corresponding to the left front microphone and the right front microphone at the same time, the in-vehicle voice processing system turns off the original left rear microphone and determines the priority relationship of the left front microphone and the right front microphone, that is, the priority relationship of the physical buttons corresponding to the above two microphones. For example, when the left front microphone has a higher priority, the in-vehicle voice processing system only controls the newly added left front microphone to collect the voice in the vehicle cabin.

[0083] It should be noted that the priority relationship can be set by default by the system or by the driver, and the embodiments of the present application do not limit this. For example, the system defaults that the left front microphone closer to the main driver seat has the highest priority, and the right front microphone closer to the co-driver seat has the second highest priority.

[0084] Optionally, the in-vehicle voice processing system can be controlled by the driver. Correspondingly, the in-vehicle voice processing system can recognize the voice of the driver and control a specific microphone.

[0085] Optionally, the audio processing module is configured to perform semantic recognition on the collected voice signals, and the semantic recognition is configured to determine whether the voice signals contain target instructions, and the target instructions are configured to instruct to turn on the microphones in the microphone array; in a case where the plurality of first voice signals contain target instructions, at least one target microphone corresponding to the target instructions is determined based on at least one microphone identifier in the target instructions, and different microphones in the microphone array correspond to different microphone identifiers; and the audio processing module is configured to collect voice through the at least one target microphone to obtain at least one voice signal.

[0086] For example, the voice of the occupant contains instructions such as "turn on the left front microphone" or "turn on the driver's seat microphone and the passenger's seat microphone". The audio processing module is configured to perform semantic recognition on the voice of the occupant to determine that the voice signal contains target instructions. The audio processing module is configured to determine that the microphone identifier in the target instructions is "left front microphone" or "driver's seat microphone" and "passenger's seat microphone". The vehicle-mounted voice processing system is configured to turn on the corresponding microphone and control the microphone to collect voice in the vehicle cabin.

[0087] In the embodiments of the present application, through the above-mentioned multiple ways, the occupant can adjust the state of the multiple microphones at any time according to the needs through simple and intuitive operations, so as to meet the needs of various complex scenes. For example, in the long-distance driving scene, the driver can only turn on the microphone at the driver's seat to perform voice navigation or answer the phone, while turning off the microphones at other positions to reduce the interference of background noise. Or, in the scene of the passenger making a voice call in the vehicle, the passenger can only turn on the microphone at his own position while turning off the microphones at other positions to ensure the clarity and privacy of the voice call. It should be noted that the occupant can also use the software control system in the vehicle to access the setting options of the microphone array, and select to turn on or turn off the specific microphone to adapt to the current use scene. The embodiments of the present application will not be described again.

[0088] 302、In a case where the voice of the plurality of sound source objects is collected, the audio processing module is configured to determine a target sound source object from the plurality of sound source objects, and the target sound source object is a sound source object making a voice call through the microphone array, and the audio processing module is connected with the communication module.

[0089] In the embodiments of the present application, the sound source object usually refers to the occupant. After the audio processing module receives the plurality of first voice signals sent by the microphone array, if it is detected that the plurality of first voice signals contain the sound of the plurality of sound source objects, the vehicle-mounted voice processing module determines the target sound source object from the plurality of sound source objects, so as to facilitate subsequent reduction of the interference of the voice of other sound source objects on the voice of the target sound source object, and improve the use experience and satisfaction of the occupant.

[0090] Generally, the target sound source object is determined in any one of the following manners. Here, the vehicle to which the vehicle-mounted voice processing system belongs accesses a multi-person online conference. At this time, each occupant in the vehicle can be a sound source object. The vehicle-mounted voice processing system detects the sound in the vehicle in real time, and when multiple sound source objects emit sound, the target sound source object is determined in any one of the following manners.

[0091] Manner one: the voice duration of each sound source object in the multiple sound source objects is determined by the audio processing module; and the sound source object with the longest voice duration is determined as the target sound source object.

[0092] The voice duration can be represented by the total duration of continuous sound emission, or by the proportion of the sound emission duration within a certain duration (e.g., per second). For example, the person at the main driving position speaks in the multi-person online conference, and the person at the co-driver position emits sound briefly. The vehicle-mounted voice processing system determines the person at the main driving position as the target sound source object and preferentially retains or enhances the voice signal thereof.

[0093] Manner two: the voice feature parameters of each sound source object in the multiple sound source objects are determined by the audio processing module; and the sound source object corresponding to the voice feature parameters existing in the feature database is determined as the target sound source object. The feature database is used to store the voice feature parameters of pre-recorded voice signals.

[0094] The voice feature parameters of each sound source object are not completely the same and include voiceprint features, tone features, pitch features, and speech rate features, etc. The feature database can be automatically generated according to historical call records, or generated from voice samples pre-recorded by the participants. For example, the vehicle-mounted voice processing system identifies that the voice feature parameters of a certain sound source object in the vehicle match the voice feature parameters of a historical participant pre-recorded in the feature database, while the feature parameters of other sound source objects cannot be matched with the feature database. At this time, the sound source object is determined as the target sound source object.

[0095] Manner three: the position of each sound source object in the multiple sound source objects is determined by the audio processing module; and the sound source object at the preset position is determined as the target sound source object.

[0096] The preset position can be a specific position directly set, or a position with higher priority. For example, the positions of the sound source objects in the vehicle are usually fixed, including the main driving position, the co-driver position, and the rear seat position, etc. The priority of the above multiple positions can be pre-set by the vehicle voice processing system. For example, when the person at the main driving position and the person at the co-driver position emit sound at the same time, the person at the main driving position is determined as the target sound source object by the vehicle-mounted voice processing system because the priority of the main driving position is higher.

[0097] 303、determine distances between the microphones and the target sound source object corresponding to the plurality of first voice signals respectively by the audio processing module.

[0098] In the embodiments of the present application, after the target sound source object is determined, the vehicle-mounted voice processing system calculates the distances between the plurality of microphones in the microphone array that collect voice and the target sound source object, so as to facilitate subsequent determination of a specific voice signal for voice communication from the plurality of first voice signals based on the distances, and improve the quality of voice communication. Optionally, the vehicle-mounted voice processing system analyzes the plurality of first voice signals through a pre-configured sound source positioning algorithm. The position of the target sound source object is determined through the waveform, frequency, phase and other characteristics of the plurality of first voice signals, and then the distances between the microphones and the target sound source object are determined. Correspondingly, the phase difference, time difference or intensity difference between the plurality of first voice signals is calculated by the audio processing module, so as to determine the position and distance of the target sound source object.

[0099] 304、determine a second voice signal from the plurality of first voice signals by the audio processing module, the distance between the microphone corresponding to the second voice signal and the target sound source object is less than the distance between the microphone corresponding to other voice signals and the target sound source object, and the other voice signals are voice signals other than the second voice signal.

[0100] In the embodiments of the present application, since the second voice signal is a voice signal collected by a microphone with a shorter distance to the target sound source object, the voice signal contains less noise and distortion, and the voice signal contains higher voice clarity of the target sound source object. For example, in a vehicle-mounted conference scene or a vehicle-mounted voice control scene, by selecting a voice signal collected by a microphone with a shorter distance to the target sound source object, the vehicle-mounted voice processing system can more accurately recognize the voice content of the target sound source object, thereby providing more accurate and personalized services and improving the user experience. It should be noted that in the case of collecting the voice of a single sound source object, the sound source object is directly determined as the target sound source object, and the subsequent operations and steps 303 to 304 are the same, and will not be described here.

[0101] 305、based on the other voice signals, perform voice processing on the second voice signal by the audio processing module to obtain an optimized second voice signal, the optimized second voice signal being a voice signal of the target sound source object after filtering out noise signals, the noise signals including voice signals of other sound source objects, environmental noise signals and echo signals, and the other sound source objects being sound source objects other than the target sound source object in the plurality of sound source objects.

[0102] In the embodiment of the present application, the second voice signal not only includes the voices of multiple sound source objects, but also includes environmental noise and echo, etc. In order to further improve the quality of the second voice signal, the audio processing module distinguishes the voice signal of the target sound source object from the second voice signal and retains it, while eliminating other interference signals. The environmental noise includes wind noise, tire noise, engine noise, etc., and the echo includes the echo generated by the voice played by the loudspeaker and the echo generated by the sound source object, etc.

[0103] Optionally, the vehicle-mounted voice processing system uses a beamforming technique to enhance the signal in the direction of the target sound source object, while suppressing the noise in other directions. Alternatively, the vehicle-mounted voice processing system adopts a blind source signal separation algorithm to separate the mixed voice signals of multiple sound source objects, thereby extracting the voice signal of the target sound source object. At this time, since most of the interference signals in the second voice signal are removed, only the clear voice of the target sound source object is retained, therefore, the signal intelligibility and the intelligibility of the optimized second voice signal are both better than those of the original second voice signal, and the optimized second voice signal can also provide more accurate and reliable data support for subsequent voice processing. The intelligibility is also referred to as speech intelligibility, and the intelligibility is used to indicate the percentage of the voice signal that can be understood by the user through the sound transmission system.

[0104] 306. The second voice signal is sent to the communication terminal of the target sound source object through the communication module.

[0105] In the embodiment of the present application, the communication terminal can be wirelessly connected to the vehicle where the vehicle-mounted voice processing system is located through Bluetooth. After receiving the second voice signal sent by the audio processing, the communication terminal sends the second voice signal to the communication terminal through the communication module, and sends the voice signal to the communication terminal through the communication terminal, thereby realizing the vehicle-mounted voice communication.

[0106] For example, the communication terminal is a mobile phone. In the vehicle-mounted voice communication scenario, there is long-distance communication between the mobile phone and the communication terminal, and there is Bluetooth transmission between the mobile phone and the vehicle where the vehicle-mounted voice processing system is located. At this time, the vehicle-mounted voice processing system collects the voice through the microphone array, and then processes the voice to obtain a clear voice signal. The vehicle sends the voice signal to the mobile phone through the vehicle-mounted voice processing system, and the mobile phone sends the voice signal to the communication terminal through long-distance communication technology. Correspondingly, when the mobile phone receives the voice signal sent by the communication terminal, the mobile phone sends the voice signal to the vehicle where the vehicle-mounted voice processing system is located through Bluetooth transmission, so that the voice signal is played through the vehicle-mounted loudspeaker in the vehicle-mounted voice processing system. Through the above-mentioned manner, long-distance communication between the user in the vehicle and the user of the communication terminal can be realized in the vehicle-mounted voice communication scenario.

[0107] The embodiment of the present application provides a vehicle-mounted voice processing method, through a vehicle-mounted voice processing system composed of a microphone array, an audio processing module, a communication module and the like, a voice signal related to a target sound source object in a voice call process can be automatically selected. Compared with a single microphone used for collecting the voice signal in the vehicle in the prior art, the present application can improve the quality of the vehicle-mounted voice collection, and further improve the accuracy of the vehicle-mounted voice recognition. Accordingly, the present application can not only improve the intelligent level of the vehicle, but also improve the user experience and satisfaction.

[0108] It should be noted that the vehicle-mounted voice processing system provided by the above embodiment divides the functions of the application into different functional modules for example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the terminal is divided into different functional modules to complete all or part of the functions described above. In addition, the vehicle-mounted voice processing system and the vehicle-mounted voice processing method provided by the above embodiment belong to the same concept, and the specific implementation process is shown in the method embodiment, which will not be described here.

[0109] Figure 4 Fig. 4 is a structural schematic diagram of an electronic device according to the embodiment of the present application. The electronic device 400 can be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a notebook computer or a desktop computer. The electronic device 400 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal or other names.

[0110] Generally, the electronic device 400 includes a processor 401 and a memory 402.

[0111] The processor 401 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 401 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 401 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 401 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 401 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.

[0112] The memory 402 can include one or more computer-readable storage media, which can be non-transitory. The memory 402 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one computer program for being executed by the processor 401 to implement the vehicle-mounted voice processing method provided by the method embodiments in the present application.

[0113] In some embodiments, the electronic device 400 can also optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface 403 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 403 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, and a power supply 408.

[0114] The peripheral interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 401 and the memory 402. In some embodiments, the processor 401, the memory 402 and the peripheral interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 401, the memory 402 and the peripheral interface 403 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.

[0115] The radio frequency circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic wave signals. The radio frequency circuit 404 communicates with a communication network and other communication devices through electromagnetic wave signals. The radio frequency circuit 404 converts electrical signals into electromagnetic wave signals for transmission, or converts received electromagnetic wave signals into electrical signals. In some embodiments, the radio frequency circuit 404 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 404 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 4G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 404 can also include NFC (Near Field Communication) related circuitry, and the present application is not limited in this regard.

[0116] The display screen 405 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 405 is a touch display screen, the display screen 405 is further configured to capture touch signals on or above the surface of the display screen 405. The touch signals can be input to the processor 401 as control signals for processing. In this case, the display screen 405 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 405 can be one, disposed on the front panel of the electronic device 400; in other embodiments, the display screen 405 can be at least two, respectively disposed on different surfaces of the electronic device 400 or in a folding design; in other embodiments, the display screen 405 can be a flexible display screen, disposed on a curved surface or a folding surface of the electronic device 400. Even, the display screen 405 can also be disposed in an irregular shape, i.e., a special-shaped screen. The display screen 405 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.

[0117] The camera assembly 406 is configured to capture images or videos. In some embodiments, the camera assembly 406 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 406 can further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0118] The audio circuit 407 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into an electrical signal input to the processor 401 for processing, or input to the radio frequency circuit 404 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the electronic device 400. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker can be a conventional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can it convert electrical signals into sound waves that humans can hear, but it can also convert electrical signals into sound waves that humans cannot hear for ranging purposes. In some embodiments, the audio circuit 407 can also include a headphone jack.

[0119] The power supply 408 is used to supply power to various components in the electronic device 400. The power supply 408 can be alternating current, direct current, disposable batteries, or rechargeable batteries. When the power supply 408 includes rechargeable batteries, the rechargeable batteries can support wired charging or wireless charging. The rechargeable batteries can also be used to support fast charging technology.

[0120] In some embodiments, the electronic device 400 also includes one or more sensors 409. The one or more sensors 409 include, but are not limited to, an acceleration sensor 44, a gyroscope sensor 411, a pressure sensor 410, an optical sensor 413, and a proximity sensor 414.

[0121] The acceleration sensor 44 can detect the acceleration in three coordinate axes of the coordinate system established by the electronic device 400. For example, the acceleration sensor 44 can be used to detect the components of the gravitational acceleration in three coordinate axes. The processor 401 can control the display screen 405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 44. The acceleration sensor 44 can also be used for game or user motion data collection.

[0122] The gyroscope sensor 411 can detect the body orientation and rotation angle of the electronic device 400. The gyroscope sensor 411 can work with the acceleration sensor 44 to collect 3D actions of the user on the electronic device 400. The processor 401 can realize the following functions according to the data collected by the gyroscope sensor 411: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.

[0123] The pressure sensor 410 can be disposed on the side bezel of the electronic device 400 and / or on the lower layer of the display screen 405. When the pressure sensor 410 is disposed on the side bezel of the electronic device 400, it can detect the user's grip signal on the electronic device 400, and the processor 401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 410. When the pressure sensor 410 is disposed on the lower layer of the display screen 405, the processor 401 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0124] Optical sensor 413 is used to collect ambient light intensity. In one embodiment, processor 401 can control the display brightness of display screen 405 based on the ambient light intensity collected by optical sensor 413. Optionally, when the ambient light intensity is high, the display brightness of display screen 405 is increased; when the ambient light intensity is low, the display brightness of display screen 405 is decreased. In another embodiment, processor 401 can also dynamically adjust the shooting parameters of camera assembly 409 based on the ambient light intensity collected by optical sensor 413.

[0125] A proximity sensor 414, also known as a distance sensor, is installed on the front panel of the electronic device 400. The proximity sensor 414 is used to detect the distance between the user and the front of the electronic device 400. In one embodiment, when the proximity sensor 414 detects that the distance between the user and the front of the electronic device 400 is gradually decreasing, the processor 401 controls the display screen 405 to switch from a screen-on state to a screen-off state; when the proximity sensor 414 detects that the distance between the user and the front of the electronic device 400 is gradually increasing, the processor 401 controls the display screen 405 to switch from a screen-off state to a screen-on state.

[0126] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on the electronic device 400, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0127] Figure 5is a structural schematic diagram of a server provided by an embodiment of the present application. The server 150 can be quite different in configuration or performance, and can include one or more processors (Central Processing Units, CPUs) 151 and one or more memories 152. The memory 152 stores at least one computer program, which is loaded and executed by the processor 151 to implement the vehicle-mounted voice processing method provided by each of the above-mentioned method embodiments. Of course, the server can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for implementing device functions, and the like, so as to perform input and output. The server can also include other components for implementing device functions, which are not described here.

[0128] The embodiment of the present application also provides a computer readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to implement the vehicle-mounted voice processing method in the above-mentioned embodiments. For example, the computer readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0129] The embodiment of the present application also provides a computer program product, which includes a computer program. The computer program is executed by a processor to implement the vehicle-mounted voice processing method in the embodiment of the present application.

[0130] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a Read-Only Memory, a magnetic disk or an optical disk, and the like.

[0131] The above-mentioned is only an optional embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, and the like made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A vehicle-mounted voice processing method, characterized by, The application is applied to a vehicle-mounted voice processing system, the vehicle-mounted voice processing system comprises a microphone array, a plurality of physical buttons, an audio processing module and a communication module, the microphone array comprises a plurality of microphones, the plurality of physical buttons are used for controlling the microphone array, the plurality of physical buttons correspond to the plurality of microphones one by one, and the method comprises the following steps: Collecting voice through the microphone array to obtain a plurality of first voice signals, the plurality of first voice signals correspond to the plurality of microphones one by one, the microphone array is connected with the audio processing module; Performing voice processing on the plurality of first voice signals through the audio processing module to obtain a second voice signal, the second voice signal is a voice signal of a target sound source object, the target sound source object is a sound source object for voice communication through the microphone array, the audio processing module is connected with the communication module; Sending the second voice signal to a communication terminal of the target sound source object through the communication module; In a case where the vehicle-mounted voice processing system is in a multi-microphone working mode, determining at least two microphones that are working, the multi-microphone working mode is used for indicating that the vehicle-mounted voice processing system supports a plurality of microphones to collect voice at the same time; in response to a trigger operation on a first physical button, collecting voice through the at least two microphones and a microphone corresponding to the first physical button to obtain a plurality of voice signals, the first physical button is a physical button corresponding to a microphone other than the at least two microphones; In a case where the vehicle-mounted voice processing system is in a single-microphone working mode, determining a single microphone that is working, the single-microphone working mode is used for indicating that the vehicle-mounted voice processing system only supports a single microphone to collect voice; in response to a trigger operation on any second physical button, collecting voice through a microphone corresponding to the second physical button to obtain a single voice signal, the second physical button is a physical button corresponding to a microphone other than the single microphone; or, in response to trigger operations on at least two second physical buttons at the same time, collecting voice through a microphone corresponding to a second physical button with the highest priority based on priorities of the at least two second physical buttons to obtain a single voice signal; Performing semantic recognition on the collected voice signals through the audio processing module, the semantic recognition is used for determining whether the voice signals contain a target instruction, the target instruction is used for indicating to turn on a microphone in the microphone array; in a case where the plurality of first voice signals contain the target instruction, determining at least one target microphone corresponding to the target instruction based on at least one microphone identifier in the target instruction, different microphones in the microphone array correspond to different microphone identifiers; collecting voice through the at least one target microphone to obtain at least one voice signal.

2. The method of claim 1, wherein, The method comprises the following steps: Performing voice processing on the plurality of first voice signals through the audio processing module to obtain a second voice signal, the second voice signal is a voice signal of a target sound source object, the target sound source object is a sound source object for voice communication through the microphone array, the audio processing module is connected with the communication module; In the case of collecting the voice of a single sound source object, the audio processing module is used to determine the distance between the microphone corresponding to each of the first voice signals and the sound source object; The audio processing module is used to determine the second voice signal from the first voice signals, the distance between the microphone corresponding to the second voice signal and the sound source object being smaller than the distance between the microphone corresponding to other voice signals and the sound source object, the other voice signals being voice signals other than the second voice signal.

3. The method of claim 1, wherein, The audio processing module is used to process the first voice signals to obtain the second voice signal, including: In the case of collecting the voice of multiple sound source objects, the audio processing module is used to determine the target sound source object from the multiple sound source objects; The audio processing module is used to determine the distance between the microphone corresponding to each of the first voice signals and the target sound source object; The audio processing module is used to determine the second voice signal from the first voice signals, the distance between the microphone corresponding to the second voice signal and the target sound source object being smaller than the distance between the microphone corresponding to other voice signals and the target sound source object, the other voice signals being voice signals other than the second voice signal.

4. The method of claim 3, wherein, The audio processing module is used to determine the target sound source object from the multiple sound source objects, including: The audio processing module is used to determine the voice duration of each of the multiple sound source objects; The sound source object with the longest voice duration is determined as the target sound source object.

5. The method of claim 3, wherein, The audio processing module is used to determine the target sound source object from the multiple sound source objects, including: The audio processing module is used to determine the voice feature parameters of each of the multiple sound source objects; The sound source object corresponding to the voice feature parameters existing in the feature database is determined as the target sound source object, the feature database being used to store the voice feature parameters of pre-recorded voice signals.

6. The method of claim 3, wherein, The audio processing module is used to determine the target sound source object from the multiple sound source objects, including: The audio processing module is used to determine the position of each of the multiple sound source objects; The sound source object at the preset position is determined as the target sound source object.

7. The method of claim 3, wherein, The method further includes: Based on the other voice signals, the audio processing module is used to process the second voice signal to obtain an optimized second voice signal, the optimized second voice signal being the voice signal of the target sound source object after filtering out noise signals, the noise signals including the voice signals of other sound source objects, environmental noise signals, and echo signals, the other sound source objects being sound source objects other than the target sound source object among the multiple sound source objects.

8. An in-vehicle speech processing system, characterized by comprising: The system is used for executing the vehicle-mounted voice processing method as claimed in any one of claims 1 to 7, and comprises a plurality of physical buttons, a microphone array, an audio processing module, a communication module, a power amplifier module, a vehicle-mounted loudspeaker and a control system. The plurality of physical buttons are used for controlling the microphone array, and the plurality of physical buttons are connected with the microphone array. The microphone array is used for collecting voice to obtain a voice signal, and comprises a plurality of microphones, and is connected with the audio processing module. The audio processing module is used for performing voice processing on the voice signal to obtain a processed voice signal, and is connected with the communication module. The communication module is used for sending the processed voice signal to other devices and receiving voice signals sent by other devices, and is connected with the power amplifier module. The power amplifier module is used for amplifying the received voice signal to obtain an amplified voice signal, and is connected with the vehicle-mounted loudspeaker. The vehicle-mounted loudspeaker is used for playing the amplified voice signal. The control system is used for controlling a plurality of modules in the system.

9. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory is used for storing at least one computer program, the at least one computer program is loaded and executed by the processor to execute the vehicle-mounted voice processing method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium is used for storing at least one computer program, and the at least one computer program is used for executing the vehicle-mounted voice processing method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice processing method and device, computer readable storage medium and electronic equipment

    CN114598963A

  • Voice signal processing method and device

    WO2016183791A1