Method, apparatus, electronic device and readable storage medium for speech recognition

By obtaining image information according to changes in doors or gears in new energy vehicles, determining seat information and selecting adaptive voice recognition algorithms, and turning off power supply without microphones, the problem of large power consumption of voice recognition in new energy vehicles is solved, and the energy-saving voice recognition effect is achieved.

CN117116268BActive Publication Date: 2025-07-11CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311041475.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-07-11
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

There is a problem of large power consumption in existing new energy vehicles.

Method used

By obtaining image information in the vehicle when the door switch status or gear information changes, determining the seat information of the person riding, and selecting a voice recognition algorithm adapted to the distribution of personnel for recognition based on the corresponding relationship between the preset riding position and the voice recognition algorithm, turning off the sound pickup microphone of the seat that is not in the passenger, and controlling its power supply.

Benefits of technology

It realizes that while ensuring the quality of voice recognition, it saves power, avoids the increase in energy consumption caused by unreasonable use of voice recognition algorithms, and improves battery usage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117116268B_ABST
    Figure CN117116268B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of automobiles, and provides a method, an apparatus, an electronic device and a readable storage medium for voice recognition. The method includes: obtaining image information inside a target vehicle when the door switch state or gear information of the target vehicle changes; determining target seat information where a person inside the target vehicle is sitting according to the image information; determining a target voice recognition algorithm corresponding to the target seat information according to the correspondence relationship between the preset riding position information and the voice recognition algorithm; and recognizing the voice inside the target vehicle according to the target voice recognition algorithm. The embodiments of the present application solve the problem of wasting electric vehicle power in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of automobiles, and particularly to a method, device, electronic device and readable storage medium for speech recognition. Background Art

[0002] With the development of new energy vehicle technology, speech recognition technology has become an essential function in the automotive field. In practical applications, speech recognition technology can simplify the cumbersome steps of traditional mechanical operations and greatly improve driving safety.

[0003] In new energy vehicles, generally multiple pickup microphones are installed inside the vehicle. The sounds emitted by the occupants inside the vehicle are collected through all the pickup microphones, and the speech recognition system processes the sounds received by the pickup microphones in real time in the background for speech recognition. However, this speech recognition method has the problem of relatively high power consumption. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, device, electronic device and readable storage medium for speech recognition to solve the problem of relatively high power consumption in the speech recognition method in related technology vehicles.

[0005] In the first aspect of the embodiments of this application, a method for speech recognition is provided, including:

[0006] When the door switch state or gear information of the target vehicle changes, obtain the image information inside the target vehicle;

[0007] According to the image information, determine the target seat information of the person sitting inside the target vehicle;

[0008] According to the correspondence between the preset seating position information and the speech recognition algorithm, determine the target speech recognition algorithm corresponding to the target seat information;

[0009] Recognize the speech inside the target vehicle according to the target speech recognition algorithm.

[0010] In the second aspect of the embodiments of this application, a device for speech recognition is provided, including:

[0011] An acquisition module, configured to obtain the image information inside the target vehicle when the door switch state or gear information of the target vehicle changes;

[0012] A first determination module, configured to determine the target seat information of the person sitting inside the target vehicle according to the image information;

[0013] A second determination module, configured to determine a target speech recognition algorithm corresponding to the target seat information according to the correspondence between pre-set riding position information and a speech recognition algorithm;

[0014] A speech recognition module, configured to recognize speech in the target vehicle according to the target speech recognition algorithm.

[0015] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0016] In a fourth aspect of the embodiments of the present application, a readable storage medium is provided. The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0017] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0018] When the door switch state or gear information of the target vehicle changes, image information inside the target vehicle is acquired, realizing that image information is acquired each time the door is opened or the gear information changes, and determining the personnel distribution inside the vehicle through the image information, ensuring the accuracy of the determined personnel distribution; in addition, according to the image information, the target seat information occupied by the personnel inside the target vehicle is determined, and according to the correspondence between the pre-set riding position information and the speech recognition algorithm, the target speech recognition algorithm corresponding to the target seat information is determined, so that the adopted target speech recognition algorithm is adapted to the personnel distribution, realizing saving speech recognition power while ensuring speech recognition quality, and avoiding the problem of excessive power consumption caused by using the same speech recognition algorithm regardless of the number of people inside the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of a method for speech recognition provided by an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of the module operation of a method for speech recognition provided by an embodiment of the present application;

[0022] Figure 3It is a schematic flowchart of another speech recognition method provided by an embodiment of the present application;

[0023] Figure 4 It is an internal structure diagram of a speech recognition module provided by an embodiment of the present application;

[0024] Figure 5 It is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application;

[0025] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the related objects before and after.

[0028] In addition, it should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusively, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.

[0029] Next, the speech recognition method, device, electronic device, and readable storage medium of the embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flowchart of a speech recognition method provided by an embodiment of the present application. AsFigure 1 As shown, the method includes:

[0031] Step 101, when the door switch state or gear information of the target vehicle changes, obtain the image information inside the target vehicle.

[0032] The change in the door switch state of the target vehicle includes the door switching from the closed state to the open state, or from the open state to the closed state.

[0033] The change in the gear information of the target vehicle includes the gear of the target vehicle changing from the parking gear to the forward gear, or from the neutral gear to the forward gear, etc.

[0034] The door switch state or gear information of the target vehicle can be monitored in real time through the Controller Area Network (CAN).

[0035] The image information can be an overall panoramic view inside the target vehicle or an image of the seat part inside the target vehicle, and the image information can be obtained through a camera installed inside the target vehicle.

[0036] By obtaining the image information inside the target vehicle when the door switch state or gear information of the target vehicle changes, it is possible to determine whether the personnel on the seats inside the target vehicle have changed, including from having personnel to no personnel on the seat and from having personnel to no personnel on the seat, achieving obtaining the image information once every time the door is opened or the gear information changes, ensuring the real-time nature of the obtained image information.

[0037] Step 102, based on the image information, determine the target seat information of the personnel inside the target vehicle.

[0038] Specifically, the target seat information indicates the seat on which the personnel inside the target vehicle are sitting, including the driver's seat, the co-driver's seat, and the left rear seat, the right rear seat, the middle rear seat, etc. inside the target vehicle.

[0039] Determining the target seat information of the personnel inside the target vehicle through the image information, that is, determining the personnel distribution inside the vehicle, ensures the accuracy of the determined personnel distribution.

[0040] It should be noted that in this embodiment, the target seat information of the personnel inside the target vehicle can also be determined through an infrared sensor installed inside the target vehicle or through a pressure sensor installed on the seat of the target vehicle, and specific limitations are not provided here.

[0041] Step 103, according to the correspondence relationship between the pre-set riding position information and the speech recognition algorithm, determine the target speech recognition algorithm corresponding to the target seat information.

[0042] Specifically, the speech recognition algorithm can convert the audio signal into text form or other forms, aiming to recognize and understand the speech features extracted from the speech signal, so as to convert it into the corresponding text representation.

[0043] The speech recognition algorithm can include single-microphone speech algorithms, dual-microphone speech algorithms, four-microphone speech algorithms, etc. In addition, the speech recognition algorithm can include Hidden Markov Models (HMMs), Deep Neural Networks (DNNs), or Convolutional Temporal Deep Neural Networks (CTC-DNN) recognition algorithms, which are not specifically limited here.

[0044] A corresponding relationship is preset between the seating position information and the speech recognition algorithm. For example, as an example, if the seating position of the person is the driver's seat, the speech recognition algorithm corresponding to the driver's seat can be used for speech recognition; in this way, if other speech recognition algorithms are used, the sound data collected by the pick-up microphones corresponding to the other positions of the target vehicle is far less than the sound data collected by the pick-up microphone corresponding to the driver's seat. Such sound data does not improve the speech recognition effect and increases the occupancy rate of the Central Processing Unit (CPU), resulting in an increase in the energy consumption of the entire target vehicle. Therefore, by using the speech recognition algorithm corresponding to the driver's seat for speech recognition, not only the speech recognition effect is guaranteed, but also the energy consumption of the vehicle is reduced.

[0045] In this way, in this embodiment, by using the target speech recognition algorithm corresponding to the target seat information, the adopted target speech recognition algorithm is adapted to the personnel distribution situation, achieving the goal of saving speech recognition power while ensuring the speech recognition quality, and avoiding the problem of increased processor energy consumption of the vehicle caused by the unreasonable use of the speech recognition algorithm.

[0046] Step 104, recognize the speech in the target vehicle according to the target speech recognition algorithm.

[0047] Specifically, the speech in the target vehicle is recognized by the target speech recognition algorithm. Since the adopted target speech recognition algorithm is adapted to the personnel distribution situation, the speech recognition power is saved while ensuring the speech recognition quality.

[0048] It should be noted that when recognizing the speech in the target vehicle, the speech signal can be preprocessed, such as noise reduction, enhancement, and equalization, etc., to improve the accuracy and reliability of speech recognition.

[0049] In this way, in this embodiment, an image is obtained each time the door opening state or the gear position information changes, and the personnel distribution in the vehicle is determined through the image information, ensuring the accuracy of the determined personnel distribution. In addition, according to the image information, the target seat information on which the personnel in the target vehicle are sitting is determined, and according to the corresponding relationship between the preset sitting position information and the speech recognition algorithm, the target speech recognition algorithm corresponding to the target seat information is determined, so that the target speech recognition algorithm adopted is adapted to the personnel distribution, achieving the purpose of saving speech recognition power while ensuring the speech recognition quality, and avoiding the problem of excessive power consumption caused by using the same speech recognition algorithm regardless of the number of personnel in the vehicle.

[0050] In some embodiments, after obtaining the image information of the target vehicle, when it is determined according to the image information that there is no person on a seat in the target vehicle, the pick-up microphone corresponding to the seat is controlled to be turned off; and the power supply of the pick-up microphone is controlled to be cut off.

[0051] Specifically, if it is determined according to the image information that there is no person sitting on a seat, the pick-up microphone corresponding to the seat is turned off and the power supply of the pick-up microphone is cut off.

[0052] For example, as an example, assume that the target vehicle is equipped with a pick-up microphone at each of the driver's seat and the co-driver's seat. If it is determined according to the image information that there is a person sitting on the driver's seat and no person sitting on the co-driver's seat, the pick-up microphone corresponding to the co-driver's seat is controlled to be turned off and the power supply of the pick-up microphone is cut off.

[0053] The seat of the vehicle can be any seat of the target vehicle, and this embodiment does not make specific limitations.

[0054] By turning off the pick-up microphone corresponding to the seat where no one is sitting and cutting off the power supply, it is realized to determine the pick-up microphone that needs to be turned off according to the personnel distribution, reducing the power-consuming devices of the target vehicle, thereby achieving the purpose of power saving and solving the problem of waste of the target vehicle's power by turning on all the pick-up microphones.

[0055] In some embodiments, before determining the target seat information on which the personnel in the target vehicle are sitting according to the image information, it further includes:

[0056] According to the image information, detect whether there is a face image within the space range corresponding to each seat;

[0057] If there is no such face image within the space range corresponding to the seat, it is determined that there is no person sitting on the seat; if there is such face image within the space range corresponding to the seat, multiple face images corresponding to different moments of the seat are obtained, and in the case where it is determined that the multiple face images are different, it is determined that there is a person sitting on the seat.

[0058] Specifically, it can be determined whether there is a face feature through a trained image recognition model to determine whether there is a face image within the space range corresponding to the seat. If there is no face image within the space range corresponding to the seat, it can be directly determined that there is no person sitting on the seat.

[0059] In addition, in order to ensure that the person on the seat is not a model or a photo, when it is determined that there is a face image within the space range corresponding to the seat according to the image information, multiple face images corresponding to different moments of the seat can be obtained, and whether the face image is a model or a photo can be detected through the multiple face images corresponding to different moments of the same seat. For example, if it is determined that there are changes in face features such as the blinking degree and the opening and closing degree of the corners of the mouth of the multiple face images, it is determined that the person on the seat is not a model or a photo, but a real person.

[0060] In this embodiment, according to the image information, it is detected whether there is a face image within the space range corresponding to each seat to determine whether there is a person sitting on the seat, and whether there are action changes in the face image is detected through multiple face images corresponding to different moments of the same seat, so as to determine that the person on the seat is not a model or a photo, realizing accurate identification of whether there is a person on the seat and avoiding the situation of misidentification caused by a model or a photo on the seat.

[0061] In addition, this embodiment can also control the activation of the sound pickup microphone corresponding to the target seat information according to the target seat information. Specifically, in some embodiments, there is at least one sound pickup microphone corresponding to the front row seats of the target vehicle, and there are at least two sound pickup microphones corresponding to the rear row seats of the target vehicle:

[0062] After determining the target seat information of the person sitting in the target vehicle according to the image information, the following at least one is further included:

[0063] First, in the case where the target seat information is only the driver's seat, control to activate only the sound pickup microphone corresponding to the driver's seat; where the sound pickup microphone corresponding to the driver's seat is the sound pickup microphone closest to the driver's seat among the at least one sound pickup microphones corresponding to the front row seats.

[0064] Specifically, when the image information indicates that the target seat information is only the driver's seat, the sound pickup microphone corresponding to the driver's seat can be controlled to be activated.

[0065] For example, if the target vehicle is only equipped with two pickup microphones and both are located in the front row, then control to turn on the pickup microphone closest to the driver's seat; if the two pickup microphones are configured with one in the front row and one in the back row of the target vehicle, then control to turn on the pickup microphone corresponding to the front row position; if the target vehicle is equipped with four pickup microphones, then control to turn on the pickup microphone closest to the driver's seat.

[0066] This control only turns on the pickup microphone corresponding to the driver's seat, ensuring that the pickup microphone can accurately pick up the sounds inside the vehicle and avoiding the problem of wasting power by turning on all pickup microphones.

[0067] Second, when the target seat information is the driver's seat and the passenger seat, then control to only turn on all the pickup microphones corresponding to the front row seats.

[0068] Specifically, when the image information indicates that the target seat information is the driver's seat and the passenger seat, it can control to turn on all the pickup microphones corresponding to the front row seats.

[0069] For example, as an example, if the target vehicle is only equipped with two pickup microphones and both are located in the front row, then control to turn on the two pickup microphones corresponding to the front row seats; if the two pickup microphones are configured with one in the front row and one in the back row of the target vehicle, then control to turn on the pickup microphone corresponding to the front row position; if the target vehicle is equipped with four pickup microphones, then control to turn on the two pickup microphones corresponding to the front row seats.

[0070] Third, when the target seat information is the driver's seat and one of the rear row seats, then control to turn on all the pickup microphones corresponding to the front row seats and the pickup microphone corresponding to the rear row seat on one side.

[0071] Specifically, when the image information indicates that the target seat information is the driver's seat and one of the rear row seats, it can control to turn on all the pickup microphones corresponding to the front row seats and the pickup microphone corresponding to the rear row seat on one side.

[0072] For example, as an example, if the target vehicle is only equipped with two pickup microphones and both are located in the front row, then control to turn on the two pickup microphones corresponding to the front row seats; if the two pickup microphones are configured with one in the front row and one in the back row of the target vehicle, then control to turn on all the pickup microphones; if the target vehicle is equipped with four pickup microphones, then control to turn on the two pickup microphones corresponding to the front row seats and the microphone corresponding to the seat of the rear row passenger.

[0073] Fourthly, when the target seat information is the driver's seat and the two side seats in the back row, control is performed to turn on all the pick-up microphones corresponding to the front row seats and all the pick-up microphones corresponding to the back row seats.

[0074] Specifically, when the image information indicates that the target seat information is the driver's seat and the two side seats in the back row, control can be performed to turn on all the pick-up microphones corresponding to the front row seats and all the pick-up microphones corresponding to the back row seats.

[0075] For example, as an example, if the target vehicle is only equipped with two pick-up microphones, control is performed to turn on the two pick-up microphones in the target vehicle; if the target vehicle is only equipped with four pick-up microphones, control is performed to turn on the four pick-up microphones in the target vehicle.

[0076] In this embodiment, by controlling the activation of the pick-up microphones corresponding to the target seat information according to the target seat information, the reasonable use of the pick-up microphones in the target vehicle is achieved. While ensuring the accurate pickup of sounds inside the vehicle, the power consumption caused by turning on all the pick-up microphones in the target vehicle is avoided.

[0077] In addition, in some embodiments, the voice recognition algorithm includes a single-microphone voice algorithm, a dual-microphone voice algorithm, and a four-microphone voice algorithm;

[0078] Determining the target voice recognition algorithm corresponding to the target seat information according to the correspondence relationship between the pre-set seating position information and the voice recognition algorithm includes at least one of the following:

[0079] First, according to the correspondence relationship, when the target seat information is only the driver's seat, determine that the target voice recognition algorithm is the single-microphone voice algorithm.

[0080] Specifically, the single-microphone voice algorithm refers to an algorithm for recognizing voice signals collected by a single pick-up microphone, mainly relying on the audio information collected by a single pick-up microphone for voice analysis and recognition.

[0081] If the current image information indicates that the target seat information is only the driver's seat, it can be determined that the target voice recognition algorithm is the single-microphone voice algorithm, that is, control is performed to use the single-microphone voice algorithm to perform voice recognition on the collected sound data.

[0082] Second, according to the correspondence relationship, when the target seat information is the driver's seat and the co-driver's seat, determine that the target voice recognition algorithm is the dual-microphone voice algorithm.

[0083] Specifically, the dual-microphone voice algorithm refers to an algorithm that performs recognition based on voice signals collected by two pick-up microphones. The dual-microphone voice recognition algorithm utilizes the audio signals collected by two pick-up microphones, and through information such as the time difference and sound intensity difference between the pick-up microphones, provides more sound source localization and noise reduction capabilities.

[0084] If the current image information indicates that the target seat information is the driver's seat and the front passenger seat, then the target voice recognition algorithm can be determined as the dual-microphone voice algorithm, that is, control the use of the dual-microphone voice algorithm to perform voice recognition on the collected sound data.

[0085] Thirdly, according to the corresponding relationship, when the target seat information is the driver's seat and at least one side seat in the back row, determine that the target voice recognition algorithm is the four-microphone voice algorithm.

[0086] Specifically, the four-microphone voice algorithm refers to an algorithm that performs recognition based on voice signals collected by four pick-up microphones. The four-microphone voice algorithm can more accurately locate the sound source position and provide stronger noise reduction and echo cancellation capabilities through the audio signals collected by four pick-up microphones, thereby improving the performance and accuracy of voice recognition.

[0087] If the current image information indicates that the target seat information is the driver's seat and at least one side seat in the back row, then determine that the target voice recognition algorithm is the four-microphone voice algorithm, that is, the four-microphone voice algorithm can be used to perform voice recognition on the collected sound data.

[0088] It should be noted that compared with the dual-microphone voice algorithm and the four-microphone voice algorithm, the single-microphone voice algorithm has a low computational load and the least power consumption. The four-microphone voice recognition algorithm has a high computational load and the most power consumption.

[0089] For example, taking the number of requests per second (Queries Per Second, QPS) as an indicator, QPS represents the number of voice recognition requests that the system can process per second in voice recognition. This indicator is usually used to measure the processing ability and performance of a voice recognition system. The single-microphone voice algorithm usually has no request indicators to process. The dual-microphone voice algorithm needs to process 1-3 request indicators, and the four-microphone voice algorithm needs to process 4-10 request indicators. It can be seen that the single-microphone voice algorithm has the lowest computational load, the four-microphone voice algorithm has the highest computational load, and the computational load of the dual-microphone voice algorithm is between that of the single-microphone voice algorithm and the four-microphone voice algorithm.

[0090] In this way, by determining the target voice recognition algorithm in the above manner, this embodiment realizes the adoption of the corresponding voice recognition algorithm according to the number of people and position information in the target vehicle, realizes the reduction of CPU operation, and thus achieves the purpose of power saving.

[0091] In addition, in some embodiments, each pickup microphone in the target vehicle corresponds to a piece of pickup microphone data; the step of recognizing the voice in the target vehicle according to the target voice recognition algorithm includes:

[0092] Obtain the voice data in the target vehicle, where the voice data includes first data and second data, the first data is the data picked up by the pickup microphones in the on state, and the second data is zero corresponding to the pickup microphones in the off state;

[0093] According to the voice data, recognize the voice in the target vehicle through the target voice recognition algorithm.

[0094] Specifically, the pickup microphone data is the sound signal data collected by the pickup microphone. When the vehicle enables the voice recognition function, the pickup microphone will collect the voice input in the vehicle and convert it into a digital signal for voice recognition processing. For example, the driver's instructions or the conversation content of the passengers. By analyzing and processing these pickup microphone data, the system can implement functions such as voice command control, phone interaction, and navigation instructions.

[0095] When the pickup microphone is turned off, if no processing is performed, the voice recognition unit may receive some environmental noise or other non-voice signals, resulting in incorrect recognition by the voice recognition unit. Therefore, the voice data corresponding to the pickup microphones in the off state can be set to zero; that is, the first data is the data picked up by the pickup microphones in the on state, and the second data is zero corresponding to the pickup microphones in the off state.

[0096] In this embodiment, the voice data includes first data and second data. The first data is the data picked up by the pickup microphones in the on state, and the second data is zero corresponding to the pickup microphones in the off state, which realizes setting the data of the pickup microphones corresponding to the seats without passengers to 0, thereby simulating the voice input data as a mute state, effectively reducing the error rate of system recognition, and avoiding affecting the voice recognition function.

[0097] In addition, in some embodiments, after recognizing the voice in the target vehicle according to the target voice recognition algorithm, a feedback message input by the user can also be received, where the feedback message is used to indicate the user's satisfaction with the voice recognition result; according to the satisfaction, update the correspondence between the seating position information and the voice recognition algorithm.

[0098] Specifically, after a user interacts with the voice interaction system in the vehicle, the user can input feedback information, which is used to indicate the user's satisfaction with the voice recognition result. For example, if the user is not satisfied with the inaccurate voice recognition result, the user can input dissatisfaction.

[0099] After receiving the feedback message, the vehicle can update the correspondence between the sitting position information and the voice recognition algorithm according to the satisfaction, so that the updated voice recognition algorithm can more accurately recognize the voice in the vehicle and improve the accuracy of voice recognition.

[0100] In this way, by updating the correspondence between the sitting position information and the voice recognition algorithm according to the satisfaction, the voice recognition algorithm is updated according to the user's needs, and the accuracy of voice recognition is improved.

[0101] Next, in combination with Figure 2 , the module working process of a voice recognition method according to an embodiment of the present application will be described, as Figure 2 shown:

[0102] First, the body signal monitoring module will detect the door switch state or gear information of the target vehicle in real time. When it detects that the door switch state or gear information of the target vehicle changes, it will send the detection signal to the image recognition module.

[0103] Then, after receiving the signal, the image recognition module will recognize the image information obtained through the camera installed in the vehicle and send the recognition result to the voice recognition module and the system driving module.

[0104] Finally, the system driving module will control the voice recognition module to analyze and process the sound data collected by the pickup microphone hardware using the voice algorithm corresponding to the image information according to the recognition result of the image recognition module. At the same time, the system driving module will control to turn off the corresponding pickup microphone and set the pickup microphone data to 0 according to the recognition result of the image recognition module.

[0105] Next, in combination with Figure 3 , a voice recognition method according to an embodiment of the present application will be described, as Figure 3 shown in the book. The method includes:

[0106] First, the door switch state or gear information of the target vehicle is monitored in real time through the CAN bus. When the door switch state or gear information of the target vehicle changes, the image information inside the target vehicle is obtained. According to the image information, it is detected whether there is a face image within the space range corresponding to each seat to determine the target seat information where the person is sitting. It should be noted that in order to ensure the accuracy of image information recognition, various sitting postures and facial features of the person can be extracted through an image recognition algorithm for training to obtain an image recognition model for recognizing the image information to improve the accuracy of image recognition.

[0107] Then, it is determined whether there are two or four pick-up microphones installed inside the target vehicle.

[0108] If there are only two pick-up microphones installed in the target vehicle, it is detected whether there is a person in the co-pilot seat. When the image information indicates that there are people in both the driver's seat and the co-pilot seat within the space range of the target vehicle, at least one pick-up microphone corresponding to the front row seats is controlled to be turned on, and the remaining pick-up microphones are turned off; for example, if the target vehicle is only equipped with two pick-up microphones and both are located in the front row, then the two pick-up microphones corresponding to the front row seats are controlled to be turned on, and the remaining pick-up microphones are turned off; if the configuration of the two pick-up microphones is that there is one pick-up microphone in each of the front and rear rows of the target vehicle, then the pick-up microphone corresponding to the front row position is controlled to be turned on, and the remaining pick-up microphones are turned off; if the target vehicle is equipped with four pick-up microphones, then the two pick-up microphones corresponding to the front row seats are controlled to be turned on, the remaining pick-up microphones are turned off, and the voice recognition controller is controlled to perform voice recognition using the dual-microphone voice algorithm.

[0109] When the image information indicates that there is only a person in the driver's seat within the space range of the target vehicle, the pick-up microphone corresponding to the driver's seat is controlled to be turned on, and the remaining pick-up microphones are turned off. For example, if the target vehicle is only equipped with two pick-up microphones and both are located in the front row, then the pick-up microphone closest to the driver's seat is controlled to be turned on; if the configuration of the two pick-up microphones is that there is one pick-up microphone in each of the front and rear rows of the target vehicle, then the pick-up microphone corresponding to the front row position is controlled to be turned on, and the remaining pick-up microphones are turned off; if the target vehicle is equipped with four pick-up microphones, then the pick-up microphone closest to the driver's seat is controlled to be turned on, the remaining pick-up microphones are turned off, and the voice recognition controller is controlled to perform voice recognition using the single-microphone voice algorithm.

[0110] Finally, the data of the turned-off pick-up microphones is set to 0, so as to simulate the voice input data as a mute state, effectively reducing the error rate of system recognition and avoiding affecting the voice recognition function.

[0111] Next, in combination with Figure 4 , the internal structure diagram of a voice recognition module according to an embodiment of the present application will be described. As shown in Figure 4As shown, the structure includes:

[0112] The speech recognition module internally includes six parts: acquiring audio data, voice activity detection, noise reduction, natural language understanding, automatic speech recognition, and sound source localization.

[0113] Acquiring audio data means obtaining real-time audio input from an input source (such as a pickup microphone, audio file, etc.) for subsequent speech recognition processing.

[0114] Voice Activity Detection (VAD) is the process of identifying active speech segments (i.e., time periods containing valid speech) and non-speech segments (i.e., time periods not containing valid speech) in a speech signal. In speech recognition or speech processing, for long audio inputs, it is usually necessary to identify the speech parts for subsequent processing while ignoring background noise or silent parts. This requires voice activity detection to determine when there is speech activity and when there is not. By accurately determining the time periods of speech activity, speech signals can be more effectively extracted and processed, improving the recognition accuracy and quality of the speech parts. At the same time, it can also help eliminate noise, reduce the amount of output data, and save computing resources.

[0115] Noise reduction is the process of reducing or removing the noise components in an audio signal. In speech recognition and audio processing, noise refers to unwanted environmental sounds or other non-speech sound components that are irrelevant to the signal of interest. Noise can be generated from various sources, such as background noise, electrical noise, wind noise, etc. These noise reduction methods can be used alone or in combination, and the specific methods selected and applied depend on factors such as the noise environment, application requirements, and performance requirements. Through noise reduction processing, the quality, clarity, and intelligibility of the audio signal can be effectively improved, and the performance and user experience of speech-related applications can be enhanced.

[0116] Natural Language Understanding (NLU) is the ability to enable a computer to understand and interpret natural language. It involves converting natural language text or speech input into a form that the computer can understand and process to extract information such as intentions, entities, and relationships in the text or speech. The goal of natural language understanding is to enable the computer to understand and process natural language input in a way similar to humans. Through effective natural language understanding technology, the computer can better understand and respond to the user's language needs, achieving a more intelligent and user-friendly interaction experience.

[0117] Automatic Speech Recognition (ASR) refers to the automated process of converting human speech input into text representation. It is a technology that analyzes and decodes sound signals and converts them into corresponding text.

[0118] Sound source localization refers to the process of determining the location of the sound source in an audio signal. Its goal is to determine the direction and location of the sound by analyzing features such as the time delay, intensity, and spectrum of the sound signal.

[0119] Any combination of the above optional technical solutions can be used to form optional embodiments of this application, and will not be elaborated one by one here.

[0120] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the process of the embodiments of this application.

[0121] The following is an embodiment of the device of this application, which can be used to execute the method embodiment of this application. For details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.

[0122] Figure 5 This is a voice recognition device provided by an embodiment of this application. As Figure 5 shown, the device includes:

[0123] An acquisition module 501, configured to acquire image information inside the target vehicle when the door switch state or gear information of the target vehicle changes;

[0124] A first determination module 502, configured to determine target seat information of the person sitting in the target vehicle according to the image information;

[0125] A second determination module 503, configured to determine a target speech recognition algorithm corresponding to the target seat information according to the correspondence between the pre-set sitting position information and the speech recognition algorithm;

[0126] A speech recognition module 504, configured to recognize the speech inside the target vehicle according to the target speech recognition algorithm.

[0127] In some embodiments, the first determination module is further configured to, when it is determined according to the image information that there is no person on a seat inside the target vehicle, control to turn off the pick-up microphone corresponding to the seat; and control to cut off the power supply of the pick-up microphone.

[0128] In some embodiments, the first determination module is further configured to detect whether there is a face image within the space range corresponding to each seat according to the image information; if there is no face image within the space range corresponding to the seat, it is determined that there is no person sitting on the seat; if there is a face image within the space range corresponding to the seat, multiple face images corresponding to different moments of the seat are obtained, and when it is determined that the multiple face images are different, it is determined that there is a person sitting on the seat.

[0129] In some embodiments, there is at least one sound pickup microphone corresponding to the front row seats of the target vehicle, and at least two sound pickup microphones corresponding to the rear row seats of the target vehicle:

[0130] The first determination module is specifically configured to perform at least one of the following:

[0131] When the target seat information is only the driver's seat, control to only turn on the sound pickup microphone corresponding to the driver's seat; the sound pickup microphone corresponding to the driver's seat is the sound pickup microphone closest to the driver's seat among the at least one sound pickup microphones corresponding to the front row seats; when the target seat information is the driver's seat and the co-driver's seat, control to only turn on all the sound pickup microphones corresponding to the front row seats; when the target seat information is the driver's seat and a rear row side seat, control to turn on all the sound pickup microphones corresponding to the front row seats and the sound pickup microphone corresponding to the rear row side seat; when the target seat information is the driver's seat and both rear row seats, control to turn on all the sound pickup microphones corresponding to the front row seats and all the sound pickup microphones corresponding to the rear row seats.

[0132] In some embodiments, the voice recognition algorithm includes a single-microphone voice algorithm, a dual-microphone voice algorithm, and a four-microphone voice algorithm; the second determination module is specifically configured to perform at least one of the following:

[0133] According to the corresponding relationship, when the target seat information is only the driver's seat, determine that the target voice recognition algorithm is the single-microphone voice algorithm; according to the corresponding relationship, when the target seat information is the driver's seat and the co-driver's seat, determine that the target voice recognition algorithm is the dual-microphone voice algorithm; according to the corresponding relationship, when the target seat information is the driver's seat and at least one side seat of the rear seat, determine that the target voice recognition algorithm is the four-microphone voice algorithm.

[0134] In some embodiments, each sound pickup microphone in the target vehicle corresponds to a sound pickup microphone data;

[0135] The speech recognition module is specifically configured to obtain the speech data in the target vehicle, where the speech data includes first data and second data. The first data is the data picked up by the pickup microphone in the on state, and the second data is zero corresponding to the pickup microphone in the off state. According to the speech data, the speech in the target vehicle is recognized by the target speech recognition algorithm.

[0136] In some embodiments, the second determination module is further configured to receive a feedback message input by the user, where the feedback message is used to indicate the user's satisfaction with the speech recognition result. According to the satisfaction degree, the corresponding relationship between the riding position information and the speech recognition algorithm is updated.

[0137] The device provided in the embodiments of the present application can implement all the method steps of the above method embodiments and achieve the same technical effects, which will not be described in detail here.

[0138] Figure 6 It is a schematic diagram of the electronic device 6 provided in the embodiments of the present application. As Figure 6 shown, the electronic device 6 in this embodiment includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, the steps in the above various method embodiments are implemented. Alternatively, when the processor 601 executes the computer program 603, the functions of each module / unit in the above device embodiments are implemented.

[0139] The electronic device 6 may be a desktop computer, a notebook, a palm computer, a cloud server and other electronic devices. The electronic device 6 may include, but is not limited to, the processor 601 and the memory 602. Those skilled in the art can understand that Figure 6 is only an example of the electronic device 6, and does not constitute a limitation on the electronic device 6. It may include more or fewer components than those shown in the figure, or different components.

[0140] The processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0141] The memory 602 can be an internal storage unit of the electronic device 6, for example, the hard disk or memory of the electronic device 6. The memory 602 can also be an external storage device of the electronic device 6, for example, a plug-in hard disk equipped on the electronic device 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 602 can also include both the internal storage unit of the electronic device 6 and the external storage device. The memory 602 is used to store computer programs and other programs and data required by the electronic device.

[0142] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0143] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0144] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for speech recognition, characterized in that, Including: When the door switch state or gear information of the target vehicle changes, obtaining image information inside the target vehicle; Determining target seat information where the person inside the target vehicle is sitting according to the image information; Determining a target speech recognition algorithm corresponding to the target seat information according to the pre-set correspondence between the sitting position information and the speech recognition algorithm; Recognizing the speech inside the target vehicle according to the target speech recognition algorithm; The speech recognition algorithm includes a single-microphone speech algorithm, a dual-microphone speech algorithm, and a four-microphone speech algorithm; determining the target speech recognition algorithm corresponding to the target seat information according to the pre-set correspondence between the sitting position information and the speech recognition algorithm includes at least one of the following: According to the correspondence, when the target seat information is only the driver's seat, determining the target speech recognition algorithm as the single-microphone speech algorithm; According to the correspondence, when the target seat information is the driver's seat and the front passenger seat, determining the target speech recognition algorithm as the dual-microphone speech algorithm; According to the correspondence, when the target seat information is the driver's seat and at least one side seat in the back row, determining the target speech recognition algorithm as the four-microphone speech algorithm; Recognizing the speech inside the target vehicle according to the target speech recognition algorithm includes: Obtaining speech data in the target vehicle, where the speech data includes first data and second data, the first data is the data picked up by the turned-on pickup microphone, and the second data is zero corresponding to the turned-off pickup microphone; Recognizing the speech inside the target vehicle according to the speech data through the target speech recognition algorithm.

2. The method for speech recognition according to claim 1, characterized in that After obtaining the image information inside the target vehicle, it further includes: When it is determined according to the image information that there is no person on a seat inside the target vehicle, controlling to turn off the pickup microphone corresponding to the seat; and, Controlling to cut off the power supply of the pickup microphone.

3. The method for speech recognition according to claim 1, wherein Before determining the target seat information where the person inside the target vehicle is sitting according to the image information, it further includes: Detecting whether there is a face image within the space range corresponding to each seat according to the image information; If there is no such face image within the space range corresponding to the seat, determining that there is no person sitting on the seat; If there is such face image within the space range corresponding to the seat, obtaining multiple face images corresponding to the seat at different times, and when it is determined that there are differences among the multiple face images, determining that there is a person sitting on the seat.

4. The method for speech recognition according to claim 1, wherein At least one pickup microphone corresponds to the front row seats of the target vehicle, and at least two pickup microphones correspond to the back row seats of the target vehicle: After determining the target seat information where the person inside the target vehicle is sitting according to the image information, it further includes at least one of the following: When the target seat information is only the driver's seat, control to only turn on the pick-up microphone corresponding to the driver's seat; the pick-up microphone corresponding to the driver's seat is the pick-up microphone closest to the driver's seat among at least one pick-up microphone corresponding to the front row seats; When the target seat information is the driver's seat and the co-driver's seat, then control to only turn on all the pick-up microphones corresponding to the front row seats; When the target seat information is the driver's seat and a rear row side seat, then control to turn on all the pick-up microphones corresponding to the front row seats and the pick-up microphone corresponding to the rear row side seat; When the target seat information is the driver's seat and both rear row seats, then control to turn on all the pick-up microphones corresponding to the front row seats and all the pick-up microphones corresponding to the rear row seats.

5. The method for speech recognition according to claim 4, wherein After recognizing the voice in the target vehicle according to the target voice recognition algorithm, it further includes: Receiving a feedback message input by the user, where the feedback message is used to indicate the user's satisfaction with the voice recognition result; Updating the corresponding relationship between the seating position information and the voice recognition algorithm according to the satisfaction.

6. A device for speech recognition, characterized in that, It includes: An acquisition module, configured to acquire the image information in the target vehicle when the door switch state or gear information of the target vehicle changes; A first determination module, configured to determine the target seat information of the person sitting in the target vehicle according to the image information; A second determination module, configured to determine the target voice recognition algorithm corresponding to the target seat information according to the pre-set corresponding relationship between the seating position information and the voice recognition algorithm; A voice recognition module, configured to recognize the voice in the target vehicle according to the target voice recognition algorithm; The voice recognition algorithm includes a single microphone voice algorithm, a dual microphone voice algorithm, and a four microphone voice algorithm; the second determination module is specifically configured to perform at least one of the following: According to the corresponding relationship, when the target seat information is only the driver's seat, determine that the target voice recognition algorithm is the single microphone voice algorithm; according to the corresponding relationship, when the target seat information is the driver's seat and the co-driver's seat, determine that the target voice recognition algorithm is the dual microphone voice algorithm; according to the corresponding relationship, when the target seat information is the driver's seat and at least one side seat in the back seat, determine that the target voice recognition algorithm is the four microphone voice algorithm; Each pick-up microphone in the target vehicle corresponds to a pick-up microphone data; the voice recognition module is specifically configured to acquire the voice data in the target vehicle, where the voice data includes first data and second data, the first data is the data picked up by the pick-up microphone in the on state, and the second data is zero corresponding to the pick-up microphone in the off state; according to the voice data, recognize the voice in the target vehicle through the target voice recognition algorithm.

7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the steps of the method according to any one of claims 1 to 5 when executing the computer program.

8. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Vehicle-mounted voice acquisition method and device, vehicle and medium

    CN116741173A