Voice processing method of vehicle, vehicle and computer readable storage medium
By monitoring the vehicle seat status and dynamically adjusting the sensitivity of the audio acquisition device or shutting down the audio input channel, the problem of false wake-up and false recognition of unmanned seats in smart cars has been solved, achieving more efficient voice recognition and smoother human-computer interaction.
Patent Information
- Application Number
- CN202511268226.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-02
AI Technical Summary
Existing intelligent vehicle voice recognition systems struggle to distinguish between valid speech and noise in the complex environment inside the vehicle, resulting in high rates of false wake-up and false recognition, especially when the microphone still attempts to recognize ambient audio signals even when no one is in the seat.
By monitoring the occupancy status of vehicle seats, the sensitivity threshold of the audio acquisition device is dynamically adjusted or the audio input channel is turned off. An audio signal suppression strategy is implemented for unoccupied seats, and speech recognition processing is performed only for seats with passengers. The text recognition results are then displayed in the display area.
It effectively reduces the false wake-up rate, improves the accuracy and efficiency of speech recognition, reduces computing resources and power consumption, and enhances the user experience.
Smart Images

Figure CN121053993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to vehicle technology and voice processing technology, specifically to a voice processing method for a vehicle, a vehicle, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of intelligent vehicle technology, voice recognition has become an important tool for improving driving experience and safety. Intelligent vehicles are equipped with multiple microphones to accommodate the voice interaction needs of passengers in different seats. However, the complex environment inside the vehicle, such as background noise, air conditioning noise, and engine noise, poses a challenge to the performance of the voice recognition system.
[0003] Currently, voice recognition systems in smart cars generally adopt a full-vehicle mode to capture voice commands from all seats. However, this approach fails to effectively distinguish between valid speech and pure noise. In particular, when some seats are not occupied, the microphones in these seats will still receive and attempt to recognize non-directly related audio signals in the environment, resulting in a high rate of false wake-ups and false recognitions.
[0004] There is currently no good solution to the above problems. Summary of the Invention
[0005] This application provides a voice processing method for a vehicle, a vehicle, and a computer-readable storage medium to at least solve the technical problems of false wake-up and false recognition in related technologies.
[0006] According to one aspect of the embodiments of this application, a voice processing method for a vehicle is provided, comprising: monitoring the seating status of the vehicle and obtaining a monitoring result, wherein the monitoring result is used to reflect whether there is a passenger in any seat in the vehicle; in response to the absence of a passenger in a first seat in the vehicle, adjusting a first audio acquisition device based on an audio signal suppression strategy, wherein the audio signal suppression strategy is used to increase the sensitivity threshold of the first audio acquisition device or to close the audio input channel of the first audio acquisition device, the first audio acquisition device being used to acquire audio signals in the area where the first seat is located; in response to the presence of a passenger in a second seat in the vehicle, performing speech recognition processing on the first audio signal acquired by the second audio acquisition device to obtain a text recognition result, wherein the second audio acquisition device is used to acquire audio signals in the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal; and displaying the text recognition result in a display area of the vehicle.
[0007] Furthermore, adjusting the first audio acquisition device based on the audio signal suppression strategy includes: in response to a one-to-one configuration relationship between the first audio acquisition device and the first seat, increasing the sensitivity threshold of the first audio acquisition device or turning off the first audio acquisition device, wherein the one-to-one configuration relationship indicates that the first audio acquisition device acquires the audio signal of the area where the first seat is located alone; or, in response to a one-to-many configuration relationship between the first audio acquisition device and the seats in the vehicle, determining the second audio signal acquired by the first audio acquisition device and suppressing the second audio signal, wherein the one-to-many configuration relationship indicates that the first audio acquisition device acquires the audio signal of the area where the first seat is located simultaneously, as well as the audio signals of the areas where the other seats besides the first seat are located, and the second audio signal is the audio signal of the area where the first seat is located.
[0008] Furthermore, the suppression processing of the second audio signal includes: reducing the signal strength of the second audio signal; and / or, performing pruning processing on the second audio signal.
[0009] Furthermore, the method also includes: in response to detecting a person sitting in the first seat, lowering the sensitivity threshold of the first audio acquisition device, or opening the audio input channel of the first audio acquisition device.
[0010] Further, the speech recognition processing of the first audio signal acquired by the second audio acquisition device to obtain the text recognition result includes: performing audio preprocessing on the first audio signal to obtain a third audio signal; extracting audio features from the third audio signal to obtain audio features; recognizing the audio features based on a recognition algorithm to obtain a recognition result, wherein the recognition result is used to reflect whether the third audio signal is a valid voice command containing user intent; in response to the third audio signal being a valid voice command containing user intent, performing text conversion processing on the audio features to obtain the text recognition result.
[0011] Furthermore, the audio features are processed into text to obtain the character recognition result, including: processing the audio features into text to obtain a first character recognition result; performing text correction on the first character recognition result to obtain a second character recognition result; and adjusting the second character recognition result based on the context to obtain the final character recognition result.
[0012] Furthermore, the method also includes: calculating the confidence level of the recognition result; if the confidence level is lower than the confidence level threshold, adjusting the recognition algorithm; or, evaluating the recognition result based on historical voice commands within a preset time period to obtain an evaluation result; and adjusting the recognition algorithm in response to the evaluation result indicating that there is a misrecognition.
[0013] Furthermore, monitoring the seating status of a vehicle includes: monitoring the seating status of a vehicle based on monitoring equipment, wherein the monitoring equipment includes at least one of the following: pressure sensing equipment, infrared sensing equipment, and visual monitoring equipment.
[0014] According to another aspect of the embodiments of this application, a voice processing device for a vehicle is also provided, comprising: a monitoring module for monitoring the seating status of the vehicle and obtaining a monitoring result, wherein the monitoring result is used to reflect whether there is a passenger in any seat in the vehicle; an adjustment module for adjusting a first audio acquisition device based on an audio signal suppression strategy in response to the absence of a passenger in a first seat in the vehicle, wherein the audio signal suppression strategy is used to increase the sensitivity threshold of the first audio acquisition device or to close the audio input channel of the first audio acquisition device, the first audio acquisition device being used to acquire audio signals in the area where the first seat is located; a processing module for performing speech recognition processing on the first audio signal acquired by the second audio acquisition device in response to the presence of a passenger in a second seat in the vehicle, obtaining a text recognition result, wherein the second audio acquisition device is used to acquire audio signals in the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal; and a display module for displaying the text recognition result in a display area of the vehicle.
[0015] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0016] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0018] In this embodiment, by monitoring the seating status of the vehicle, a monitoring result is obtained to determine whether any seat in the vehicle is occupied. Then, for the first seat where no one is occupying, an audio acquisition device is adjusted based on an audio signal suppression strategy. This strategy either increases the sensitivity threshold of the first audio acquisition device or disables its audio input channel. The first audio acquisition device is used to acquire the audio signal of the area where the first seat is located. For the second seat where someone is occupying, the first audio signal acquired by the second audio acquisition device undergoes speech recognition processing to obtain a text recognition result. The second audio acquisition device is used to acquire the audio signal of the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal. Finally, the text recognition result is displayed in the vehicle's display area. This achieves the goal of improving speech recognition accuracy by dynamically adjusting the working status of microphones in different areas of the vehicle through real-time monitoring of the occupancy status of vehicle seats. For unoccupied seats, the false wake-up caused by background noise is suppressed by increasing the sensitivity threshold of the microphone or directly shutting down the audio input channel, thus effectively solving the problem of false wake-up. For occupied seats, the audio signals from these areas are processed centrally, reducing the false recognition rate, thereby solving the technical problems of false wake-up and false recognition in related speech recognition technologies. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a flowchart of a vehicle voice processing method according to one embodiment of this application;
[0021] Figure 2 This is a flowchart of speech processing according to one embodiment of this application;
[0022] Figure 3 This is a structural block diagram of a vehicle voice processing device according to one embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] For ease of understanding, some concepts related to the embodiments of this application are illustrated below for reference.
[0026] Pruning: This usually refers to an optimization technique for data processing flow or algorithms, which aims to improve algorithm efficiency and save computing resources and energy consumption by removing unnecessary computational branches or reducing data complexity.
[0027] Mel Frequency Cepstral Coefficients (MFCC) is a signal feature extraction method widely used in speech recognition and audio processing. MFCC mimics the human ear's perception of different frequencies of sound, converting audio signals into a series of numerical features. These features can effectively represent the spectral characteristics of speech, especially excelling in distinguishing different speech information.
[0028] Hidden Markov Model (HMM): A statistical model commonly used in fields such as speech recognition and natural language processing to describe a Markov process with unknown parameters.
[0029] Cars with intelligent interactive systems typically have multiple microphones for voice recognition. Voice source localization modes are often available for different seats, such as the driver, front passenger, and rear seats, to accommodate the voice interaction needs of passengers in different seats. To meet the voice needs of different seats, car owners usually choose the full vehicle mode for voice source localization.
[0030] There may be various background noises inside a vehicle, such as talking, making phone calls, listening to music, air conditioning sounds, engine sounds, wind sounds, etc. If there are no people in the vehicle seats, then the audio signals from these seats (such as the sound collected by the microphone) may be invalid or pure noise.
[0031] Currently, all voice recognition systems in the industry suffer from false wake-up and false recognition issues. Because these cannot be completely avoided, the user experience is poor. The industry average for false wake-up is 6 times per hour in dynamic mode and 12 times per hour in silent mode. False wake-up issues are frequently reported in after-sales customer service complaints. For example, a voice message saying "Hey there" might suddenly trigger a false wake-up when the passenger is not in the front seat, or the driver might wake up and receive a response from the passenger. These seemingly occasional false wake-up issues lead to user complaints.
[0032] This application aims to propose a method and vehicle solution to reduce the false voice wake-up rate, thereby reducing the voice misrecognition rate, improving user experience, and reducing user complaints.
[0033] According to an embodiment of this application, a method embodiment for voice processing of a vehicle is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] This embodiment provides a voice processing method for vehicles. Figure 1 This is a flowchart of a vehicle voice processing method according to one embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:
[0035] Step S10: Monitor the seating status of the vehicle and obtain monitoring results, wherein the monitoring results are used to reflect whether there is a person sitting in any seat in the vehicle.
[0036] In this embodiment of the application, the monitoring result is obtained by monitoring whether there is a person sitting in the vehicle seat. The monitoring result is used to reflect whether there is a person sitting in any seat in the vehicle. That is, the monitoring result will clearly indicate the status of each seat, that is, whether someone is sitting or no one is sitting.
[0037] For example, the monitoring results could be obtained by monitoring the seating status of the vehicle through an occupant monitoring module. The occupant monitoring module monitors the seating status of the vehicle in real time using technologies such as pressure sensing, infrared sensing, or visual monitoring.
[0038] It can be seen that by monitoring the seating status of the vehicle, the system can ensure that it understands the distribution of occupants inside the vehicle, providing a basis for subsequent audio signal processing. It can also avoid processing audio signals from unoccupied seats, saving computing resources and improving processing speed.
[0039] Step S12: In response to the absence of a passenger in the first seat of the vehicle, the first audio acquisition device is adjusted based on an audio signal suppression strategy. The audio signal suppression strategy is used to increase the sensitivity threshold of the first audio acquisition device or to close the audio input channel of the first audio acquisition device. The first audio acquisition device is used to acquire the audio signal of the area where the first seat is located.
[0040] In this embodiment of the application, the first seat refers to the seat that is shown as unoccupied in the monitoring results, that is, the seat in the vehicle that is not occupied.
[0041] The first audio acquisition device may be a microphone or microphone array installed in or near the first seating area to capture audio signals from the first seating area.
[0042] Audio signal suppression strategies reduce the response of the first audio acquisition device to background noise by either increasing the audio reception threshold or directly shutting down the input channel. In other words, the audio signal suppression strategy is used to either increase the sensitivity threshold of the first audio acquisition device or shut down its audio input channel.
[0043] The sensitivity threshold refers to the minimum sound intensity at which a microphone begins to capture sound. Typically, a microphone's sensitivity threshold determines the range and minimum volume of sound signals it can capture. By default, to ensure the pickup of weak speech signals, the microphone's sensitivity threshold is often set low. Therefore, when the system decides to suppress a microphone in an unoccupied seating area, it increases the microphone's sensitivity threshold. This means that unless the sound intensity is significantly higher than the original threshold, the microphone will not capture any signal. This effectively filters out background noise and other irrelevant audio input, as these sounds often do not reach the new, higher threshold.
[0044] An audio input channel refers to the path or interface through which the sound signal collected by the microphone is transmitted to the speech processing system. Each microphone or microphone array has its own input channel, ensuring that the sound signal can be processed independently. For areas where no one is sitting, the system directly shuts down the audio input channels of the microphones in that area. This means that even if the microphone receives an audio signal, this signal will not be transmitted to the speech processing system for processing, thus ensuring that the microphones in unoccupied seats are completely unaffected by background noise and do not participate in the subsequent misidentification analysis process.
[0045] It can be seen that for unoccupied seats in vehicles, adjusting the audio acquisition devices corresponding to unoccupied seats through audio signal suppression strategies, that is, by suppressing or turning off the audio input of unoccupied seats, can significantly reduce the false wake-up rate caused by background noise, reduce unnecessary audio signal processing, and improve the overall operating efficiency of the system.
[0046] Step S14: In response to the presence of a passenger in the second seat of the vehicle, the first audio signal collected by the second audio acquisition device is processed for speech recognition to obtain a text recognition result. The second audio acquisition device is used to collect audio signals in the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal.
[0047] In this embodiment of the application, the second seat refers to the seat that is shown to be occupied in the monitoring results.
[0048] The second audio acquisition device may be a microphone or microphone array installed in or near the second seating area to capture audio signals from the second seating area, such as capturing voice commands from passengers sitting in the second seat.
[0049] The first audio signal is the audio signal captured by the second audio acquisition device.
[0050] The first audio signal is processed by a speech processing system to convert the speech content in the first audio signal into text, thus obtaining the text recognition result. That is, the text recognition result is the textual representation of the speech content in the first audio signal.
[0051] It can be seen that for occupied seats in the vehicle, only the audio signal of occupied seats is processed to ensure the accuracy and effectiveness of the recognition results. Furthermore, it can quickly respond to passenger commands and provide a smooth and natural interactive experience.
[0052] Step S16: Display the text recognition results in the vehicle's display area.
[0053] In this embodiment, the display area can be a central control touch screen, instrument panel display, head-up display (HUD), or other areas used to display information in the vehicle, and is not limited here.
[0054] As can be seen, displaying the text recognition results in the display area allows drivers and passengers to intuitively see the recognition results of voice commands, thereby increasing the transparency and trust in the interaction. Furthermore, displaying the results helps users confirm whether the command has been correctly recognized, reducing operational errors.
[0055] In summary, this application monitors the seating status of a vehicle to determine whether any seat in the vehicle is occupied. Then, for the first seat where no one is present, an audio signal suppression strategy is applied to the first audio acquisition device. This strategy either increases the sensitivity threshold of the first audio acquisition device or disables its audio input channel. The first audio acquisition device then acquires the audio signal from the area where the first seat is located. For the second seat where someone is present, speech recognition processing is performed on the first audio signal acquired by the second audio acquisition device to obtain a text recognition result. The second audio acquisition device acquires the audio signal from the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal. Finally, the text recognition result is displayed in the vehicle's display area.
[0056] As can be seen, this application, by monitoring the occupancy status of vehicle seats, enables the voice processing system to intelligently identify which areas of voice signals are valid input, thereby avoiding processing invalid signals from unoccupied seats and improving the targeting and efficiency of recognition. For unoccupied seat areas, this application employs suppression strategies such as increasing the sensitivity threshold or disabling the audio input channel, effectively reducing false wake-ups caused by background noise and other non-occupant voice signals, significantly lowering the false wake-up rate. Furthermore, this application only performs feature extraction and recognition processing on audio signals from occupied seat areas, reducing the proportion of noise signals processed overall, thereby reducing the possibility of misrecognition and improving the accuracy of recognition results. This allows the voice processing system to respond accurately and quickly when an occupant issues a command, providing high-quality text display or voice feedback, enhancing the naturalness and fluency of human-computer interaction, and improving the user experience of in-vehicle voice recognition. Therefore, this application avoids processing a large number of invalid signals, saving computing resources and power consumption, allowing the system to operate more efficiently, while also extending vehicle battery life or reducing energy consumption.
[0057] The above steps of this application involve monitoring the seating status of a vehicle to determine whether any seat in the vehicle is occupied. Then, for the first seat where no one is present, an audio signal suppression strategy is applied to the first audio acquisition device. This strategy either increases the sensitivity threshold of the first audio acquisition device or disables its audio input channel. The first audio acquisition device is used to acquire the audio signal of the area where the first seat is located. For the second seat where someone is present, speech recognition processing is performed on the first audio signal acquired by the second audio acquisition device to obtain a text recognition result. The second audio acquisition device is used to acquire the audio signal of the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal. Finally, the text recognition result is displayed in the vehicle's display area. This achieves the goal of improving speech recognition accuracy by dynamically adjusting the working status of microphones in different areas of the vehicle through real-time monitoring of the occupancy status of vehicle seats. For unoccupied seats, the false wake-up caused by background noise is suppressed by increasing the sensitivity threshold of the microphone or directly shutting down the audio input channel, thus effectively solving the problem of false wake-up. For occupied seats, the audio signals from these areas are processed centrally, reducing the false recognition rate, thereby solving the technical problems of false wake-up and false recognition in related speech recognition technologies.
[0058] Optionally, in step S12, adjusting the first audio acquisition device based on the audio signal suppression strategy may include the following steps:
[0059] Step S121: In response to the one-to-one configuration relationship between the first audio acquisition device and the first seat, increase the sensitivity threshold of the first audio acquisition device or turn off the first audio acquisition device. The one-to-one configuration relationship indicates that the first audio acquisition device acquires the audio signal of the area where the first seat is located.
[0060] Alternatively, in step S122, in response to the one-to-many configuration relationship between the first audio acquisition device and the vehicle seats, the second audio signal acquired by the first audio acquisition device is determined, and the second audio signal is suppressed. The one-to-many configuration relationship indicates that the first audio acquisition device simultaneously acquires the audio signal of the area where the first seat is located, as well as the audio signals of the areas where the other seats are located, and the second audio signal is the audio signal of the area where the first seat is located.
[0061] In this embodiment of the application, when adjusting the first audio acquisition device based on the audio signal suppression strategy, in response to the one-to-one configuration relationship between the first audio acquisition device and the first seat, the sensitivity threshold of the first audio acquisition device is increased, or the first audio acquisition device is turned off.
[0062] In this context, a one-to-one configuration relationship indicates that the first audio acquisition device exclusively acquires the audio signal of the area where the first seat is located. For example, this could be a scenario where four microphones (MICs) correspond to four frequency zones. Essentially, a one-to-one configuration relationship refers to the direct association between audio acquisition devices (such as microphones) and seats within the vehicle; each microphone is responsible for acquiring the audio signal of a single seat area. Typically, in microphone array designs, specific positioning and directionality ensure that each microphone primarily captures the sound of its corresponding seat area.
[0063] If the audio acquisition device and the seat are configured one-to-one, the system will automatically increase the sensitivity threshold of the microphone in the corresponding seat area when no one is sitting there. Therefore, the microphone will only start to capture the signal when the sound intensity reaches or exceeds the set higher threshold, which helps to reduce false wake-ups caused by weak background noise or distant non-target sounds.
[0064] As can be seen, when the system detects that the first seat is unoccupied and there is a one-to-one configuration between the first audio acquisition device and the first seat, the system will adjust the sensitivity threshold of the microphone to make it more stringent in capturing sound, or directly turn off the microphone to ensure that the voice recognition system is not triggered by noise in the unoccupied area.
[0065] Therefore, by increasing the sensitivity threshold, the impact of background noise on speech recognition can be significantly reduced, lowering the false wake-up rate. Furthermore, disabling microphones in unoccupied areas reduces unnecessary power consumption and computing resource usage, improving overall system energy efficiency. This allows the system to focus more on processing audio signals from occupied areas, further enhancing the effectiveness and accuracy of recognition.
[0066] Alternatively, in response to the one-to-many configuration relationship between the first audio acquisition device and the vehicle seats, the second audio signal acquired by the first audio acquisition device is determined, and the second audio signal is suppressed.
[0067] In this context, a one-to-many configuration means that the first audio acquisition device simultaneously acquires the audio signal from the area where the first seat is located, as well as the audio signals from the areas where all other seats are located, such as 4MICs corresponding to the five-tone register. This can be understood as one audio acquisition device being responsible for acquiring audio signals from multiple seat areas. In this configuration, the microphones typically have a wider capture range, or the sound is analyzed from multiple directions using software algorithms.
[0068] The second audio signal is the audio signal from the area where the first seat is located, that is, the audio signal from the unoccupied area of the first seat acquired by the first audio acquisition device. In a one-to-many configuration, the second audio signal may be mixed with signals from other seat areas.
[0069] As can be seen, when the microphones in the vehicle are configured in a one-to-many relationship with multiple seating areas, the system will suppress the second audio signal collected by the microphone from the first unoccupied seating area. This suppression may be achieved through software algorithms during the signal preprocessing stage, such as frequency domain filtering, noise reduction, or specific signal separation techniques, to ensure that the signal from the unoccupied area is not included in the speech recognition system for analysis.
[0070] Therefore, by identifying and suppressing audio signals from unoccupied seating areas, these signals can be prevented from mixing with valid speech signals, thereby reducing misrecognition. Even in a one-to-many configuration, this application's system can intelligently adjust signal processing strategies by analyzing occupant distribution in real time to adapt to different usage scenarios and ensure the accuracy of speech recognition.
[0071] Optionally, in step S122, suppressing the second audio signal may include the following steps:
[0072] Step S121: Reduce the signal strength of the second audio signal.
[0073] And / or, in step S122, the second audio signal is subjected to pruning processing.
[0074] In this embodiment, when suppressing the second audio signal, the signal strength of the second audio signal can be reduced. It can be seen that the signal strength of the second audio signal can be reduced during suppression processing, which can be achieved by adjusting the gain or using dynamic range compression technology. The purpose is to weaken these second audio signals, reducing their influence in subsequent processing stages, thereby reducing background noise interference and the possibility of false triggering. Thus, by reducing the audio signal strength in unoccupied areas, background noise is effectively shielded, and the overall signal quality is improved.
[0075] And / or, pruning processing can be applied to the second audio signal. It can be seen that while suppressing the second audio signal, pruning processing can also be performed. For example, subsequent feature extraction and pattern recognition operations can be stopped, or only simple analysis can be performed to maintain the basic operation of the system. Thus, by stopping further complex processing of the second audio signal, computational resources and energy consumption are saved. Furthermore, pruning processing reduces the system's processing load, enabling the speech recognition algorithm to recognize and respond to occupant voice commands more quickly.
[0076] Optionally, the method may further include the following execution steps:
[0077] Step S18: In response to detecting that there is a person sitting in the first seat, lower the sensitivity threshold of the first audio acquisition device or open the audio input channel of the first audio acquisition device.
[0078] In this embodiment of the application, if a person is detected sitting in the first seat that was originally unoccupied during subsequent monitoring, the sensitivity threshold of the first audio acquisition device is lowered, or the audio input channel of the first audio acquisition device is opened.
[0079] As can be seen, when the system detects a passenger in the previously unoccupied first seat, it can automatically lower the microphone's sensitivity threshold, making it more responsive to passenger voice signals and thus capturing clearer and more accurate voice information. Alternatively, if the audio input of the first audio acquisition device was previously disabled due to vacancy, the system will re-enable the audio input function of the first audio acquisition device when a passenger is detected, opening the audio input channel to ensure that the passenger's voice commands can be captured and processed by the system.
[0080] Therefore, by adjusting the sensitivity threshold or activating the audio input channel, the system can more effectively capture passengers' voice commands, reduce the chance of misrecognition, and improve the overall accuracy of recognition. Furthermore, the system dynamically adjusts the working state of the audio acquisition device, effectively avoiding waste of resources in unoccupied areas and improving the energy efficiency and response speed of the entire voice recognition system.
[0081] Optionally, in step S14, performing speech recognition processing on the first audio signal acquired by the second audio acquisition device to obtain the text recognition result may include the following steps:
[0082] Step S141: Perform audio preprocessing on the first audio signal to obtain the third audio signal.
[0083] Step S142: Extract audio features from the third audio signal to obtain audio features.
[0084] Step S143: Based on the recognition algorithm, the audio features are identified to obtain the recognition result, wherein the recognition result is used to reflect whether the third audio signal is valid and contains a voice command with the user's intent.
[0085] Step S144: In response to the third audio signal being a valid voice command containing the user's intent, the audio features are processed into text to obtain the text recognition result.
[0086] In this embodiment of the application, when the first audio signal acquired by the second audio acquisition device is processed for speech recognition to obtain the text recognition result, the first audio signal can be preprocessed to obtain the third audio signal. The audio preprocessing is used to improve the audio signal quality and remove interference, including operations such as noise reduction, filtering, and gain adjustment, to ensure that the audio signal is clearer and easier to recognize in subsequent processing.
[0087] It can be seen that by preprocessing the first audio signal, the aim is to reduce background noise, optimize the audio signal quality, and make it more suitable for subsequent feature extraction and recognition processing.
[0088] Then, audio features are extracted from the third audio signal to obtain audio features. Audio feature extraction involves extracting a series of parameters describing the characteristics of sound from the audio signal, such as spectral features, energy features, Mel-frequency cepstral coefficients (MFCCs), pitch period, formants, and other features. The extracted audio features must reflect the basic attributes of speech to facilitate understanding and matching by the recognition algorithm.
[0089] As can be seen, by extracting key audio features from the third audio signal—that is, converting the third audio signal into a set of numerical features—the recognition algorithm can analyze and identify the signal based on these features. Thus, by extracting audio features, the complex audio signal is simplified into easily processed numerical features, providing a foundation for algorithmic recognition.
[0090] Then, the audio features are identified using a recognition algorithm to obtain the recognition result. This recognition algorithm can be a deep learning-based neural network model, a Hidden Markov Model (HMM), or other type of pattern recognition algorithm, used to determine whether the audio features correspond to valid speech commands.
[0091] The recognition results are used to reflect whether the third audio signal is valid and contains the user's intended voice command.
[0092] As can be seen, by analyzing the extracted audio features using a recognition algorithm, it is determined whether the third audio signal contains valid voice commands, thus obtaining the recognition result. The recognition result indicates whether the third audio signal contains voice commands expressing the occupant's intention. Therefore, by analyzing the signal through the algorithm, the validity of the signal is determined, avoiding unnecessary processing of non-voice signals, enhancing the intelligence and effectiveness of the system, and enabling accurate recognition of user intentions.
[0093] Finally, in response to the third audio signal being a valid voice command containing user intent, text conversion processing is performed on the audio features to obtain the text recognition result. This text conversion process transforms the speech signal into text, i.e., speech-to-text (STT). STT is used to convert the recognized speech command into a text command that can be understood and executed by the system.
[0094] As can be seen, if the third audio signal is confirmed to contain a valid voice command, then the third audio signal is converted into text, transforming the speech features in the third audio signal into human-readable text recognition results. This allows the in-vehicle system to understand and execute the occupant's specific commands, completing the corresponding functions or services. Thus, the system can understand and respond to the user's voice commands, promoting the fluency and intelligence of human-computer interaction.
[0095] Optionally, in step S144, performing text conversion processing on the audio features to obtain the text recognition result may include the following steps:
[0096] Step S1441: Perform text conversion processing on the audio features to obtain the first character recognition result.
[0097] Step S1442: Perform text correction on the first character recognition result to obtain the second character recognition result.
[0098] Step S1443: Adjust the second character recognition result based on the context content to obtain the character recognition result.
[0099] In this embodiment of the application, when performing text conversion processing on audio features to obtain text recognition results, the audio features can be processed to obtain a first text recognition result. This first text recognition result is the text output obtained from the initial text conversion processing of the audio features; it may not have been corrected or optimized, i.e., it is a preliminary textual description directly decoded from the audio features.
[0100] As can be seen, this application converts audio features into text through an algorithm, that is, transforms the occupant's voice into a preliminary text description. For example, an Hidden Markov Model (HMM) or a deep learning model can be used to analyze and interpret the audio features. Thus, the transformation of sound signals into text information is achieved, providing a foundation for subsequent processing.
[0101] Then, text correction is performed on the first character recognition result to obtain the second character recognition result. The text correction is a process of checking the grammar, spelling, and semantics of the preliminary character recognition result (i.e., the first character recognition result), aiming to improve the accuracy and readability of the text.
[0102] The second character recognition result is the corrected text output, which is usually more accurate and standardized, and closer to the original meaning of the passenger.
[0103] As can be seen, this application corrects the initial first character recognition result by using methods such as grammar checking, spelling correction, and homonym identification to eliminate errors and ambiguities that may arise from the initial recognition, resulting in a more accurate second character recognition result. This improves the accuracy and consistency of character recognition, reduces misunderstandings caused by minor errors in the initial recognition, and ensures that the occupant's intentions are accurately conveyed to the vehicle system.
[0104] Finally, the second text recognition result is adjusted based on the context to obtain the final text recognition result. The context refers to the surrounding information of the current voice command, including but not limited to the already recognized text, previous voice commands from the occupant, and the current state of the vehicle.
[0105] The text recognition result is the final text expression after all processing steps have been optimized, which can accurately reflect the occupant's intentions and needs.
[0106] As can be seen, this application further adjusts the recognition result based on the contextual information of the second character recognition result to ensure its logic, coherence, and practicality. For example, the system may need to consider the occupant's previous voice commands to understand the meaning of the current command, or judge the urgency of the occupant's needs based on the current state of the vehicle. Therefore, the final character recognition result is more closely aligned with the actual situation, improving the practicality of the recognition and the occupant's satisfaction, while also enhancing the system's ability to understand complex commands and reducing execution errors.
[0107] Optionally, the method may further include the following execution steps:
[0108] Step S145: Calculate the confidence level of the recognition result. If the confidence level is lower than the confidence level threshold, adjust the recognition algorithm.
[0109] Alternatively, in step S146, the recognition results are evaluated based on historical voice commands within a preset time period to obtain the evaluation results.
[0110] Step S147: In response to the evaluation result indicating that there is a misidentification, the recognition algorithm is adjusted.
[0111] In this embodiment, the confidence level of the recognition result can also be calculated. If the confidence level is lower than the confidence level threshold, the recognition algorithm is adjusted. The confidence level of the recognition result is an evaluation index of the correctness of a certain recognition result by the speech recognition system. It is usually a value between 0 and 1, with a higher value indicating higher accuracy.
[0112] A confidence threshold is a pre-defined criterion used to distinguish between high-confidence and low-confidence recognition results. When the confidence level of a recognition result falls below this threshold, the result is considered potentially problematic and corrective measures need to be taken.
[0113] As can be seen, this application calculates the confidence level of each recognition result and evaluates its correctness. If the confidence level of a recognition is found to be lower than a predetermined threshold, it indicates that there is uncertainty or error in the recognition process. At this time, the recognition algorithm will be adjusted to reduce the occurrence of such errors in the future. Thus, by monitoring the confidence level, this application can promptly identify potential problems in the recognition process, react quickly, adjust algorithm parameters, and improve future recognition accuracy.
[0114] Alternatively, the recognition results can be evaluated based on historical voice commands within a preset time period. If the evaluation indicates misrecognition, the recognition algorithm can be adjusted. The system records all voice commands issued by passengers within a certain period, serving as the basis for evaluating the current recognition result—that is, historical voice commands within the preset time period. The length of the preset time period can be set according to specific usage scenarios and needs; it should not be too short to provide sufficient information, nor too long to introduce excessive irrelevant data.
[0115] The evaluation result is an assessment of the accuracy of the current recognition result derived from the analysis of historical voice commands. The evaluation result reflects the consistency of the recognition result with past occupant behavior and whether it conforms to normal behavioral patterns.
[0116] As can be seen, this application evaluates the reasonableness of the current recognition result based on historical voice command records over a period of time. By comparing the similarity between the current voice command and past voice commands, and by understanding the passenger's habits, it determines whether the current recognition result is likely a misrecognition. If the evaluation result shows that the current recognition result is a misrecognition, the system will immediately adjust the recognition algorithm and optimize its parameters or logic to improve its ability to process similar voice commands in the future.
[0117] Therefore, by using historical voice commands as an evaluation benchmark, the system can better understand the needs and habits of passengers, improve the accuracy of voice command judgment, and reduce misrecognition caused by accidental factors. Furthermore, through continuous evaluation and algorithm adjustments, the system can continuously improve itself, gradually reducing the occurrence of misrecognition and enhancing the overall performance of speech recognition.
[0118] Optionally, in step S10, monitoring the seating status of the vehicle may include the following steps:
[0119] Step S101: Monitor the seat occupancy status of the vehicle based on monitoring equipment, wherein the monitoring equipment includes at least one of the following: pressure sensing equipment, infrared sensing equipment, and visual monitoring equipment.
[0120] In this embodiment of the application, when monitoring the seating status of a vehicle, the seating status can be monitored using monitoring equipment. This monitoring equipment is used to detect in real time whether there are occupants in the vehicle seats. This equipment can sense the presence of occupants and convert it into electronic signals, which are then sent to the vehicle's onboard system.
[0121] The monitoring equipment includes at least one of the following: pressure sensing devices, infrared sensing devices, and visual monitoring devices. Pressure sensing devices determine whether a seat is occupied by measuring changes in pressure on the seat surface and are suitable for most seat materials. Common pressure sensing devices include pressure sensors and barometric pressure sensors.
[0122] Infrared sensing devices use infrared light to detect the presence of objects or people in a seating area. They typically determine the presence of occupants by detecting the infrared radiation emitted by a human body. Infrared sensing devices are highly sensitive, have a fast response time, and are suitable for monitoring in dynamic environments.
[0123] Visual monitoring devices analyze whether there are occupants in seats using image recognition technology based on images captured by cameras. These devices can capture more details about occupants, such as posture and facial features, but their accuracy may be affected by lighting conditions.
[0124] As can be seen, the vehicle described in this application collects information in real time about whether the seats are occupied through monitoring devices installed under or around the seats, such as pressure sensors, infrared sensors, or visual monitoring devices (e.g., cameras). The monitoring devices convert physical signals (pressure, infrared radiation, images) into electronic signals, which are then analyzed and judged by the onboard system to ultimately determine the occupancy status of the seats.
[0125] Therefore, by monitoring the seating status of vehicles using monitoring equipment, it is possible to accurately detect whether there are passengers sitting in the seats.
[0126] In summary, this application proposes a method and vehicle for reducing voice misrecognition and false wake-up rates, including the following components: an audio acquisition module, an audio preprocessing module, an occupant monitoring module, an audio feature extraction module, and a voice recognition system. When the user sets the sound source localization to the whole vehicle mode and all seats are occupied, all microphones in the vehicle can normally receive in-vehicle voice audio information for recognition and response. When only one or a few seats are occupied, the vehicle senses the unoccupied seats through the occupant monitoring module and reduces the probability of false wake-up recognition caused by sounds such as people talking, making phone calls, listening to music, air conditioning, engine noise, and wind noise inside or outside the vehicle. That is, this application suppresses or reduces the voice audio signal of unoccupied seats by adding an occupant monitoring module, preventing them from being recognized by voice, thereby reducing the overall vehicle recognition probability. The suppression or reduction is achieved by the noise reduction module of the audio preprocessing module.
[0127] The audio acquisition module's function is to continuously collect audio signals from inside the vehicle using a microphone or other audio acquisition devices. These audio signals include various sound information from inside the vehicle, such as passengers' conversations and ambient noise.
[0128] The audio preprocessing module's function is to preprocess the acquired raw audio signal to improve audio quality and reduce interference. Common preprocessing operations include noise reduction and filtering. Noise reduction removes background noise, making the audio signal clearer. Filtering removes unwanted frequency components, highlighting useful audio information.
[0129] The function of the occupant detection module is to prevent false wake-ups by disabling the sound source localization function for seats without passengers. The occupant detection module can be implemented in various ways, such as pressure sensing modules, infrared sensing modules, and visual monitoring modules, to detect whether there is someone in the seat.
[0130] The function of the audio feature extraction module is to extract representative features from the preprocessed audio signal for subsequent pattern recognition and classification. Common audio features include spectral features, energy features, and Mel-frequency cepstral coefficients (MFCCs). These features reflect different properties of the audio signal, such as frequency distribution and energy variation.
[0131] Figure 2 This is a flowchart of speech processing according to one embodiment of this application, such as... Figure 2As shown, the key steps of the speech recognition in this application include: speech signal acquisition, preprocessing, feature extraction, acoustic modeling, language modeling, dictionary and decoding, and post-processing. In the speech recognition process of this application, sound is converted into an analog signal using devices such as a microphone, and then converted into a digital signal by an analog-to-digital converter for computer processing. Occupant monitoring signals are added during computer processing to suppress or reduce the speech audio signals from unoccupied seats, preventing them from being recognized and thus reducing the overall vehicle recognition probability. Specifically, it includes the following steps:
[0132] Step 1: Voice Signal Acquisition and Preprocessing: The vehicle's audio acquisition module uses microphones and other equipment to convert sound into analog signals, which are then converted into digital signals by an analog-to-digital converter. Signal conditioning and sampling are then performed to prepare the acquired digital voice signal for subsequent processing. This initial processing includes filtering, noise reduction, and gain adjustment to improve the quality of the voice signal. This module helps reduce background noise and interference, making the voice clearer.
[0133] Step 2, Occupant Monitoring Module: When the occupant monitoring module detects an empty seat using sensors or cameras, the system automatically shuts down the audio signal input channel corresponding to that seat by increasing the sound threshold or using software suppression through the voice recognition function of that seat. When the system detects that there is someone in that seat, it automatically turns on the audio input channel of that seat.
[0134] In the "No One Sitting" audio processing, when the system determines that no one is sitting, this branch is responsible for special processing of the audio signal. Specifically, it can choose to stop further complex processing or reduce the intensity of processing to save computing resources and energy consumption. For example, it can stop subsequent feature extraction and pattern recognition operations, or only perform simple analysis to maintain the basic operation of the system.
[0135] When audio processing detects that someone is sitting, the audio signal is processed according to the normal procedure. First, audio features are extracted, then pattern recognition and classification are performed, and finally, misidentification is identified and adjusted, before the final recognition result is output.
[0136] Step 3, Audio Feature Extraction Module: When processing audio with someone sitting, the audio feature extraction module extracts the time-domain features of the speech signal, such as short-time energy and short-time zero-crossing rate. Short-time energy reflects the energy changes of the speech signal, while the short-time zero-crossing rate reflects the number of times the speech signal crosses the zero axis.
[0137] Frequency domain analysis can be performed using methods such as the Fast Fourier Transform (FFT) to transform the speech signal from the time domain to the frequency domain, obtaining spectral information. Common characteristic parameters include Mel-frequency cepstral coefficients (MFCC) and linear predictive cepstral coefficients (LPCC), which can better reflect the spectral characteristics of speech.
[0138] In addition, features such as pitch period and formants can be extracted. The pitch period refers to the period of frequency of vocal cord vibration, and the formants are regions of concentrated energy in the speech signal. The pitch period and formants are also important for characterizing the properties of speech.
[0139] Step four, speech recognition system: misidentification judgment and adjustment, which involves evaluating and judging the recognition and classification results to determine whether misidentification exists. If misidentification is found, adjustments and corrections are made according to the specific circumstances. For example, setting thresholds or referring to historical data can be used to determine whether it is a misidentification and to take appropriate action to improve recognition accuracy.
[0140] Following this, post-processing will be performed, including necessary corrections to the generated text results, such as correcting spelling errors and adjusting capitalization, to improve recognition accuracy and readability. Furthermore, based on semantic analysis and contextual relationships, the recognition results will be further optimized and adjusted to better align with actual language expression habits.
[0141] Finally, the output module delivers the final recognition results to the user or other relevant systems. The output results may include information such as the audio signal classification and confidence level, allowing the user to understand the in-vehicle audio situation.
[0142] Furthermore, the collected audio data, recognition results, and other relevant information can be stored and recorded through the storage and recording module. This data can be used for subsequent analysis, research, and system improvement to continuously enhance the system's reliability and performance.
[0143] This application targets unoccupied seats and can suppress or reduce the audio signal at each stage. For example, the audio acquisition module does not acquire audio, the audio preprocessing module does not preprocess, the audio feature extraction module does not extract features, and a suppression term is added to each stage of the speech recognition system (such as the automatic sound source localization logic of the vehicle system).
[0144] Compared with traditional speech recognition solutions, this application reduces false wake-up and false recognition rates by using an occupant monitoring system to identify unoccupied seats, raising the sound threshold, or suppressing the speech recognition function of those seats through software. This has the following beneficial effects:
[0145] Beneficial effects: (1) Improved signal quality, including: reduced noise, i.e., after turning off audio signals when no one is sitting, the amount of audio data processed by the system is reduced, and the sources of noise are also reduced accordingly. Improved signal-to-noise ratio, i.e., the ratio of effective signals (such as the voice of the driver or passenger) to noise is increased, making it easier for the system to identify effective signals.
[0146] Beneficial effects (2) Reduced false recognition rate, including: reduced false triggering, that is, after noise is reduced, the probability of the system misidentifying noise as the target signal is reduced. Improved recognition accuracy, that is, the system can focus more on processing effective audio signals, thereby improving the accuracy of speech recognition or sound monitoring.
[0147] Beneficial effects (3) System optimization, including: dynamic adjustment, that is, the system can dynamically adjust the audio signal processing strategy according to the seat occupancy, further optimizing performance. Resource saving, that is, turning off invalid signals can also save computing resources, enabling the system to process valid signals more efficiently.
[0148] Therefore, by disabling audio signals when no one is in the vehicle, this application reduces noise interference, improves the signal-to-noise ratio of the effective signal, thereby lowering the false recognition rate and enhancing overall performance and user experience. This method has significant practical implications for applications such as in-vehicle voice recognition and call systems.
[0149] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0150] According to an embodiment of this application, a device embodiment for a vehicle voice processing apparatus is provided. It should be noted that the apparatus can be used to execute the above-described vehicle voice processing method.
[0151] Figure 3 This is a structural block diagram of a vehicle voice processing device according to one embodiment of this application, such as... Figure 3As shown, taking a vehicle's voice processing device 300 as an example, the device includes: a monitoring module 301, used to monitor the seating status of the vehicle and obtain monitoring results, wherein the monitoring results reflect whether there is a person sitting in any seat in the vehicle; an adjustment module 302, used to adjust a first audio acquisition device based on an audio signal suppression strategy in response to the absence of a person sitting in the first seat in the vehicle, wherein the audio signal suppression strategy is used to increase the sensitivity threshold of the first audio acquisition device or to close the audio input channel of the first audio acquisition device, wherein the first audio acquisition device is used to acquire audio signals in the area where the first seat is located; a processing module 303, used to perform speech recognition processing on the first audio signal acquired by the second audio acquisition device in response to the presence of a person sitting in the second seat in the vehicle, and obtain text recognition results, wherein the second audio acquisition device is used to acquire audio signals in the area where the second seat is located, and the text recognition results are the textual representation of the speech content in the first audio signal; and a display module 304, used to display the text recognition results in the display area of the vehicle.
[0152] Embodiments of this application also provide a vehicle, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods described in various embodiments of this application when it runs.
[0153] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0154] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0155] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0160] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A voice processing method for vehicles, characterized in that, The method includes: The vehicle's seating status is monitored to obtain monitoring results, wherein the monitoring results are used to reflect whether there is a person sitting in any seat in the vehicle; In response to the absence of a passenger in the first seat of the vehicle, the first audio acquisition device is adjusted based on an audio signal suppression strategy. The audio signal suppression strategy is used to either increase the sensitivity threshold of the first audio acquisition device or close the audio input channel of the first audio acquisition device. The first audio acquisition device is used to acquire audio signals in the area where the first seat is located. In response to the presence of a passenger in the second seat of the vehicle, the first audio signal collected by the second audio acquisition device is processed for speech recognition to obtain a text recognition result. The second audio acquisition device is used to collect audio signals in the area where the second seat is located, and the text recognition result is a textual representation of the speech content in the first audio signal. The text recognition result is displayed in the display area of the vehicle.
2. The method according to claim 1, characterized in that, The adjustment of the first audio acquisition device based on the audio signal suppression strategy includes: In response to the one-to-one configuration between the first audio acquisition device and the first seat, the sensitivity threshold of the first audio acquisition device is increased, or the first audio acquisition device is turned off, wherein the one-to-one configuration indicates that the first audio acquisition device acquires audio signals only from the area where the first seat is located; or... In response to the one-to-many configuration relationship between the first audio acquisition device and the seats in the vehicle, the second audio signal acquired by the first audio acquisition device is determined, and the second audio signal is suppressed. The one-to-many configuration relationship indicates that the first audio acquisition device simultaneously acquires the audio signal of the area where the first seat is located, as well as the audio signals of the areas where the other seats are located, and the second audio signal is the audio signal of the area where the first seat is located.
3. The method according to claim 2, characterized in that, The suppression processing of the second audio signal includes: Reduce the signal strength of the second audio signal; and / or, The second audio signal is subjected to pruning processing.
4. The method according to claim 1, characterized in that, The method further includes: In response to detecting that there is a person sitting in the first seat, the sensitivity threshold of the first audio acquisition device is lowered, or the audio input channel of the first audio acquisition device is opened.
5. The method according to claim 1, characterized in that, The process of performing speech recognition processing on the first audio signal acquired by the second audio acquisition device to obtain text recognition results includes: The first audio signal is preprocessed to obtain the third audio signal; Audio features are extracted from the third audio signal to obtain audio features; The audio features are identified based on the recognition algorithm to obtain a recognition result, wherein the recognition result is used to reflect whether the third audio signal is a valid voice command containing the user's intent; In response to the third audio signal being a valid voice command containing user intent, the audio features are processed for text conversion to obtain the text recognition result.
6. The method according to claim 5, characterized in that, The text conversion processing of the audio features to obtain the text recognition result includes: The audio features are subjected to text conversion processing to obtain the first character recognition result; The first character recognition result is corrected to obtain the second character recognition result; The second character recognition result is adjusted based on the context to obtain the character recognition result.
7. The method according to claim 5, characterized in that, The method further includes: Calculate the confidence level of the recognition result; if the confidence level is lower than a confidence threshold, adjust the recognition algorithm; or, The recognition results are evaluated based on historical voice commands within a preset time period to obtain an evaluation result; In response to the evaluation result indicating that the identification result has a misidentification, the identification algorithm is adjusted.
8. The method according to any one of claims 1-7, characterized in that, The seat occupancy status of the monitored vehicle includes: The vehicle's seating status is monitored using monitoring equipment, wherein the monitoring equipment includes at least one of the following: a pressure sensing device, an infrared sensing device, and a visual monitoring device.
9. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the voice processing method for a vehicle as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the voice processing method for a vehicle as described in any one of claims 1 to 8 when run on a computer or processor.