Vehicle-mounted voice recognition method, device, equipment, storage medium and program product
By combining the software application information and interface images on the vehicle's infotainment screen with the in-vehicle voice recognition system, and dynamically adjusting the voice recognition parameters and vocabulary, the problem of insufficient recognition accuracy in the in-vehicle environment is solved, achieving more efficient human-computer interaction and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VOYAH AUTOMOBILE TECH CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-19
AI Technical Summary
Existing in-vehicle voice recognition systems struggle to achieve high-precision recognition in complex in-vehicle environments, resulting in high misrecognition rates of user commands and poor interactive experiences. This is mainly because static word libraries cannot dynamically adapt to the contextual requirements of different applications.
By acquiring software application information and current interface images from the vehicle's infotainment screen, the speech recognition parameters and vocabulary are dynamically adjusted. Combined with multimodal information fusion technology, the speech recognition model is optimized to adapt to multi-user, multi-application, and multi-noise environments.
It significantly improves the accuracy and scene adaptability of speech recognition, enhances the ability to resolve ambiguous references, homophones and scene-related commands, realizes more accurate and natural human-computer interaction, and improves the reliability of voice control and user experience.
Smart Images

Figure CN122067518A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular to an in-vehicle speech recognition method, device, equipment, storage medium, and program product. Background Technology
[0002] In-vehicle voice recognition technology is a key function for human-machine interaction in modern smart car cockpits. It allows users to conveniently control vehicle functions through voice commands, which is of great significance for improving driving safety.
[0003] In existing technologies, in-vehicle voice recognition systems typically rely on speech signal processing and natural language processing techniques to parse and semantically understand user-input speech data. When processing user-input speech data, they usually depend solely on the speech data itself, employing static pre-set dictionaries and general speech recognition models for recognition.
[0004] However, the in-vehicle environment is highly dynamic and complex. The diversity of in-vehicle applications leads to a wide semantic range of voice commands. Static word libraries cannot dynamically adapt to the contextual requirements of different applications, resulting in insufficient accuracy and scenario adaptability of current speech recognition when understanding user interactions with complex screen interfaces. Summary of the Invention
[0005] This application provides an in-vehicle voice recognition method, device, equipment, storage medium, and program product to improve the accuracy and scene adaptability of voice recognition.
[0006] In a first aspect, embodiments of this application provide an in-vehicle voice recognition method, the method comprising:
[0007] In response to voice data input by the target user, the system acquires software application information and the current interface image from the target vehicle infotainment screen; wherein, the target vehicle infotainment screen is the vehicle infotainment screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of the in-vehicle applications on the target vehicle infotainment screen.
[0008] Based on the software application information and the current interface image, the voice data is enhanced for recognition to obtain a voice recognition result.
[0009] In one possible implementation, the step of enhancing the recognition of the voice data based on the software application information and the current interface image to obtain a voice recognition result includes:
[0010] Based on the software application information and the current interface image, adjust the speech recognition parameters; wherein, the speech recognition parameters are parameters used to optimize the preset speech recognition model;
[0011] The preset speech recognition model is configured based on the adjusted speech recognition parameters, and the configured speech recognition model is used to recognize the speech data to obtain the speech recognition result.
[0012] In one possible implementation, adjusting the speech recognition parameters based on the software application information and the current interface image includes:
[0013] The current interface image is semantically analyzed using a preset image recognition model to extract key information from the current interface image; wherein, the key information is text information related to the functions of the currently running in-vehicle application.
[0014] The speech recognition parameters are adjusted based on the software application information and the key information.
[0015] In one possible implementation, adjusting the speech recognition parameters based on the software application information and the key information includes:
[0016] Determine target word domains in a preset word library that are related to the software application information and the key information; wherein, the target word domains include at least one preset word;
[0017] Adjust the semantic weight and / or semantic priority of the target word domain to be higher than that of other word domains; wherein, the other word domains are word domains in the preset lexicon other than the target word domain.
[0018] In one possible implementation, before acquiring the software application information and current interface image from the target vehicle infotainment screen, the method further includes:
[0019] The usage status of each vehicle infotainment screen in the car cabin is obtained, and the target vehicle infotainment screen is determined based on the usage status of each vehicle infotainment screen.
[0020] In one possible implementation, determining the target vehicle infotainment screen based on the usage status of each of the vehicle infotainment screens includes:
[0021] If only one in-vehicle infotainment screen is in use in the vehicle cabin, then the in-vehicle infotainment screen in use is determined as the target in-vehicle infotainment screen.
[0022] If multiple in-vehicle screens in the vehicle cabin are in use, the in-vehicle screen associated with the location of the audio source of the voice data is identified as the target in-vehicle screen.
[0023] In one possible implementation, the method further includes:
[0024] Based on the voice recognition results, the corresponding operation instructions are executed on the target vehicle screen.
[0025] In one possible implementation, the method further includes, prior to responding to voice data input by the target user:
[0026] Voice data input from multiple users is collected through a microphone array distributed throughout the car cabin;
[0027] Based on the time difference and phase difference of the voice data received by each microphone, the voice data of each user is separated.
[0028] Secondly, embodiments of this application provide an in-vehicle voice recognition device, the device comprising:
[0029] The first processing unit is configured to respond to voice data input by the target user and acquire software application information and current interface image from the target vehicle infotainment screen; wherein, the target vehicle infotainment screen is the vehicle infotainment screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of in-vehicle applications in the target vehicle infotainment screen.
[0030] The second processing unit is used to enhance the recognition of the voice data based on the software application information and the current interface image to obtain the voice recognition result.
[0031] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0032] The memory stores computer-executed instructions;
[0033] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0034] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0035] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0036] The vehicle-mounted voice recognition method, device, equipment, storage medium, and program products provided in this application enhance the recognition of user voice data by introducing software application information and real-time interface images from the vehicle's infotainment screen as contextual information. This effectively solves the technical problems of low recognition accuracy and large intent comprehension deviation caused by the lack of visual context in traditional vehicle-mounted voice systems. This multimodal information fusion mechanism enables the system to dynamically understand the current application functions and interface elements, significantly improving the ability to resolve ambiguous references, homophones, and scene-related instructions. As a result, more accurate and natural human-computer interaction is achieved in complex vehicle environments, greatly improving the reliability of voice control and user experience. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0038] Figure 1 A flowchart illustrating an in-vehicle voice recognition method provided in an embodiment of this application;
[0039] Figure 2 A flowchart illustrating another in-vehicle voice recognition method provided in an embodiment of this application;
[0040] Figure 3 This is a schematic diagram of the structure of an in-vehicle voice recognition device provided in an embodiment of this application;
[0041] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0042] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0043] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0044] With the rapid development of smart cockpit technology, in-vehicle voice interaction systems have become one of the core functions for enhancing the driving experience. However, the in-vehicle environment is highly dynamic and complex, and the diversity of in-vehicle applications (such as map navigation, music and video, and contact lists) results in a wide semantic range for voice commands. Static word libraries cannot dynamically adapt to the contextual needs of different applications. Existing technologies struggle to achieve high-precision voice recognition in complex noise, multi-user scenarios, and dynamic contexts, leading to high user command misrecognition rates and poor interactive experiences.
[0045] To address the aforementioned technical issues, this application provides an in-vehicle voice recognition method. By dynamically sensing software application information and current interface images on the vehicle's infotainment screen, a multi-dimensional voice recognition enhancement mechanism is constructed, thereby improving the accuracy and adaptability of voice recognition in complex scenarios. This application breaks through the limitations of traditional static lexicons and fixed models, enabling adaptive optimization for multi-user, multi-application, and multi-noise environments.
[0046] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0047] It should be noted that the executing entity of the vehicle voice recognition method provided in this application embodiment can be an vehicle voice recognition device. This device can be deployed on the vehicle's intelligent cockpit system, vehicle controller, or other electronic devices. It can implement the method of this application through software, hardware, or a combination of software and hardware. This application embodiment does not impose any limitations. This application embodiment uses an vehicle voice recognition device as an example for detailed description.
[0048] Figure 1 This is a flowchart illustrating an in-vehicle voice recognition method provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the in-vehicle voice recognition method may include:
[0049] S101. In response to the voice data input by the target user, obtain the software application information and the current interface image on the target vehicle infotainment screen; wherein, the target vehicle infotainment screen is the vehicle infotainment screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of the in-vehicle application on the target vehicle infotainment screen.
[0050] For example, the target user refers to the user who issues voice data. When multiple users issue voice commands simultaneously, the target user refers to the user corresponding to the target vehicle infotainment screen. The voice data input by the user refers to the user's voice commands collected by the in-vehicle microphone array, such as "play song XX" or "navigate to location XXX". The microphone array can be combined with sound source localization technology to determine the location of the user issuing the command.
[0051] The target in-vehicle infotainment screen is the screen in the car cabin that corresponds to the voice data input by the target user. In a cabin with multiple in-vehicle infotainment screens (such as the driver's instrument panel screen, central control screen, passenger entertainment screen, rear entertainment screen, etc.), the target screen can be determined based on the usage status of each screen and the target user's voice position, etc., and this application embodiment does not impose any limitations. For example, based on the sound source localization result, the screen directly facing the user or the closest screen is determined as the target in-vehicle infotainment screen.
[0052] Software application information includes relevant information about all currently running and non-running in-vehicle applications displayed on the target vehicle's infotainment screen, such as application identification information (e.g., application name), application type, and running status. For example, for a music application, its status might be "Playing song 'XX'"; for a navigation application, its status might include the destination information of the currently planned route.
[0053] The current interface image refers to the complete view of the target vehicle screen captured in real time through a screenshot or image acquisition module. This image contains all visual elements, such as icons, text, buttons, lists, and pictures, providing the most direct and richest contextual information.
[0054] When the target user issues a voice command, the voice data of the user can be collected through the in-vehicle microphone array. Then, the in-vehicle voice recognition device can respond to the voice data input by the target user, obtain the software application information on the target vehicle screen in real time through the in-vehicle operating system interface (such as Android's ActivityManager), and call the screen capture API (such as PixelCopy) to capture the current interface image of the target vehicle screen, so as to provide key input for subsequent voice recognition enhancement.
[0055] This application embodiment simultaneously collects structured software application information and unstructured visual interface information, laying a solid data foundation for subsequent deep semantic understanding and overcoming the one-sidedness of relying solely on speech recognition.
[0056] S102. Based on the software application information and the current interface image, enhance the recognition of the speech data to obtain the speech recognition result.
[0057] For example, embodiments of this application aim to enhance speech data recognition based on software application information and the current interface image to obtain more accurate speech recognition results. For instance, an in-vehicle speech recognition device can upload the acquired software application information, the current interface image, and the speech data input by the target user to a local or cloud server speech recognition model for processing. The speech recognition model then enhances the speech data recognition, resulting in a more accurate speech recognition result. Alternatively, the in-vehicle speech recognition device can utilize a local image recognition library (such as an Optical Character Recognition (OCR) engine) to first parse the text of the current interface image and extract key information strongly related to the current application (such as POI name, song title, etc.) to provide contextual enhancement for the speech recognition model. The speech recognition model then performs enhanced recognition, resulting in a more accurate speech recognition result.
[0058] Optionally, in one possible embodiment, enhancing the speech recognition based on software application information and the current interface image to obtain a speech recognition result may include:
[0059] S1. Adjust the speech recognition parameters based on the software application information and the current interface image; where the speech recognition parameters are the parameters used to optimize the preset speech recognition model;
[0060] S2. Configure a preset speech recognition model based on the adjusted speech recognition parameters, and use the configured speech recognition model to recognize the speech data to obtain the speech recognition result.
[0061] For example, the preset speech recognition model refers to the pre-trained core model used to convert speech data into text. It is usually a complex model based on deep learning, which may include components such as acoustic models and language models. It can be deployed locally or on a cloud server. This application embodiment does not impose any restrictions.
[0062] Speech recognition parameters are configurable strategy parameters that affect the recognition behavior and results of a speech recognition model during recognition. These parameters determine which words or paths the model tends to select during the decoding process. Specifically, they may include the semantic weights and semantic priorities of each word domain in the preset vocabulary, as well as other acoustic model adaptation parameters of the speech recognition model, such as fine-tuning the input processing of the acoustic model according to the current environment (e.g., the in-car noise characteristics when the application is playing music). This application does not impose any limitations on these parameters.
[0063] In-vehicle voice recognition devices can first analyze the acquired software application information and current interface images. For example, the current interaction area can be clearly identified from the software application information. For instance, if the current foreground application is identified as "navigation software," then the core interaction area can be considered to be "map" and "navigation"; if it is "music software," then the core interaction area can be considered to be "music" and "playback control," etc. By performing optical character recognition (OCR) and UI element recognition on the current interface image, key information on the interface (such as keywords and optional operations) can be extracted. For example, in a music application interface, the OCR engine may recognize the currently playing song / dance title, singer, and other text, while UI recognition may detect the presence of "favorite" and "pause" buttons, etc. The in-vehicle voice recognition device can adjust the voice recognition parameters based on information recognized from software application information and current interface images. The adjusted voice recognition parameters can then be configured into a preset voice recognition model. The user's original voice data is then sent to this specifically configured model for recognition. Since the model's recognition search space and preferences have been optimized, the recognition process is more efficient and accurate, resulting in more accurate voice recognition results.
[0064] This optional embodiment transforms screen context information into internal configuration parameters of the speech recognition model, thereby realizing a front-end, deeply integrated enhanced recognition mechanism. This fundamentally improves the intelligence level of in-vehicle voice interaction and enhances the accuracy and scene adaptability of the model from the source of the recognition process.
[0065] Optionally, in one possible embodiment, step S1, adjusting the speech recognition parameters based on software application information and the current interface image, may include:
[0066] S11. Perform semantic analysis on the current interface image using a preset image recognition model to extract key information from the current interface image; where the key information is text information related to the functions of the currently running in-vehicle application.
[0067] S12. Adjust the speech recognition parameters based on software application information and key information.
[0068] For example, the preset image recognition model is a pre-trained computer vision model specifically designed for analyzing and understanding the in-vehicle infotainment screen interface. This could be a deep learning-based text detection model, a rule-based optical character recognition model, or similar. The key information consists of text information related to the functions of the currently running in-vehicle application. It is a semantically refined collection of text information strongly relevant to the current application's functions; it is no longer a simple list of all the text on the original image, but rather a filtered and categorized version.
[0069] In this embodiment, the in-vehicle voice recognition device can first input the current interface image into a preset image recognition model. The preset image recognition model performs semantic analysis on the current interface image to extract key information from it. Then, based on software application information and key information, the voice recognition parameters are adjusted. For example, a dynamic language model can be constructed, which generates a dynamic, small-scale domain language model or vocabulary based on software application information and key information. During subsequent decoding by the voice recognition model, the output weight of this dynamic language model is increased.
[0070] This embodiment introduces a preset image recognition model to perform deep semantic analysis on the interface, which can make implicit information in the user operation scenario (such as navigation destination, music playlist, etc.) explicit, transform ambiguous visual information into precise key information, and then dynamically optimize the core parameters of the speech recognition model accordingly, forming a closed-loop optimization system from vision to speech, which significantly improves the accuracy and intelligence of in-vehicle voice interaction.
[0071] Optionally, in one possible embodiment, step S12, adjusting the speech recognition parameters based on software application information and key information, may include:
[0072] S121. Determine the target word domain in the preset word library that is related to software application information and key information; wherein, the target word domain includes at least one preset word;
[0073] S122. Adjust the semantic weight and / or semantic priority of the target word domain to be higher than that of other word domains; wherein, other word domains are word domains in the preset lexicon other than the target word domain.
[0074] For example, a preset vocabulary refers to a large set of words built into the speech recognition system that covers all words of common languages (such as Chinese and English words and phrases). This vocabulary is the foundation upon which the speech recognition model can output various texts. A word domain refers to a set of sub-vocabularies in the preset vocabulary that are divided according to specific themes, scenarios, or functions. For example, a "navigation word domain" may include words such as "navigate to," "starting point," "ending point," "zoom in," "zoom out," "highway priority," and "avoid congestion"; a "music word domain" may include words such as "play," "pause," "favorite," and "next song." A target word domain is one or more highly relevant word domains dynamically selected from numerous word domains based on the current state of the vehicle's infotainment screen. It is a subset of the preset vocabulary and is considered to be the set of words that the user is most likely to say at this moment.
[0075] In this embodiment, software application information and key information extracted from the current interface image can be fused together. Then, based on the fused information, a matching query is performed in a preset word library to determine a target word domain related to the software application information and key information in this scenario. For example, if the software application information determines that the currently running application is a music app, and the extracted key information includes a "favorite" button, the song title "XX Sunny Day", and the singer's name "Zhou XX", then through association matching, not only will the basic word domain strongly related to music be locked, but also a deep filter and expansion will be performed based on the key information, taking into account words directly or indirectly related to "favorite", "XX Sunny Day", and "Zhou XX". Finally, a target word domain customized for this scenario is formed, the content of which may include: play, pause, favorite, unfavorite, next song, Zhou XX, XX Sunny Day, etc.
[0076] In the decision-making process of speech recognition models (especially at the language model level), each word or phrase has a basic statistical probability, namely its "semantic weight." Words with higher weights are more likely to be selected as the final result by the model when their acoustic features are similar. Semantic priority refers to the search strategy in the decoding process, which prioritizes the model to search and expand within the "target word domain."
[0077] The vehicle-mounted voice recognition device can send instructions to the voice recognition model through an application programming interface or a configuration interface to temporarily increase the semantic weight and / or semantic priority of all words in the "target word domain" so that they are higher than the semantic weight and / or semantic priority of words in other word domains.
[0078] This embodiment effectively solves the recognition error problem caused by homophones and near-homophones by increasing the weight of relevant words, and greatly improves the accuracy of the first recognition. By focusing the decoding search space on the target word domain, the number of candidate paths that the model needs to calculate can be reduced, thereby reducing the computational overhead, speeding up the response, and improving the recognition efficiency.
[0079] For example, when a user inputs voice data, the in-vehicle voice recognition device can obtain the name of the current foreground application (such as "map navigation") on the target vehicle screen through the operating system interface and take a screenshot of the screen to obtain the current interface image (such as the navigation interface). Then, it uses the Optical Character Recognition (OCR) engine to parse the screenshot, extract key information related to the current foreground application (such as the destination "XX Bund"), and upload the application name and key information to the cloud server. The cloud server then classifies the recognition enhancements based on the application information, such as POI enhancement for map applications, media enhancement for music and video applications, and address book enhancement for phone applications, thereby improving the accuracy of voice recognition.
[0080] The in-vehicle voice recognition method provided in this application enhances the recognition of user voice data by introducing software application information and real-time interface images from the vehicle screen as contextual information. This effectively solves the technical problems of low recognition accuracy and large intention comprehension deviation caused by the lack of visual context in traditional in-vehicle voice systems. This multimodal information fusion mechanism enables the system to dynamically understand the current application functions and interface elements, significantly improving the ability to resolve ambiguous references, homophones, and scene-related instructions. As a result, it achieves more accurate and natural human-computer interaction in complex in-vehicle environments, greatly improving the reliability of voice control and user experience.
[0081] Figure 2 This is a flowchart illustrating another in-vehicle voice recognition method provided in an embodiment of this application. Figure 2 As shown, in this embodiment... Figure 1 Based on the embodiments, a detailed description is provided for the scenario where multiple users simultaneously input voice data and the car cabin includes multiple in-vehicle infotainment screens. The method may include:
[0082] S201. Collect voice data input from multiple users through a microphone array distributed in the car cabin.
[0083] For example, a microphone array refers to multiple microphones arranged in a specific geometry within a car cabin. It is not a simple recording device, but a collaborative acoustic signal acquisition system. Its core value lies in its ability to spatially locate sound sources by comparing the signals received by different microphones. When one or more users (such as the driver and front passenger) in the car cabin issue voice commands simultaneously or sequentially, the microphones distributed in different locations synchronously acquire the mixed audio streams, obtaining voice data input by multiple users.
[0084] S202. Based on the time difference and phase difference of the voice data received by each microphone, separate the voice data of each user.
[0085] For example, since sound waves travel different distances to different microphones, the time it takes for the same sound signal to reach each microphone will be slightly different, resulting in a time difference. The waveform phase will also change accordingly, resulting in a phase difference. The time difference and phase difference are key features for sound source localization and separation.
[0086] By employing signal processing algorithms such as blind source separation and beamforming, the voice data of each user can be separated. Specifically, by adjusting the phase and weight of each microphone signal, one or more "steering pickup beams" can be formed, acting like a virtual "directional microphone" that amplifies sound from only a specific direction (such as the driver's mouth) while suppressing noise and interference from other directions, thereby achieving the separation of user voice data.
[0087] Optionally, the separated speech can be authenticated by combining pre-stored user voiceprint features, so that not only the speech can be separated, but also the user's identity can be identified. This application embodiment does not impose any limitations.
[0088] S203. Obtain the usage status of each vehicle infotainment screen in the car cabin, and determine the target vehicle infotainment screen based on the usage status of each vehicle infotainment screen.
[0089] For example, in this embodiment of the application, when multiple users simultaneously input voice in an in-vehicle environment, the location information of the user's voice input can be obtained through sound source recognition technology, and the sound regions can be divided in combination with the screen location information. Voice commands can then be segmented for recognition, thereby improving the accuracy of voice recognition. Specifically, after separating the voice data of each user, voice recognition processing is performed separately for each user's voice data.
[0090] For example, for each user's voice data, the usage status of each in-vehicle infotainment screen in the car cabin is first obtained, and then the target in-vehicle infotainment screen for each user is determined based on preset decision logic. For example, the target in-vehicle infotainment screen for the driver's seat user corresponds to the driver's screen, the target in-vehicle infotainment screen for the passenger seat user corresponds to the passenger's screen, and the target in-vehicle infotainment screen for the rear seat user corresponds to the ceiling-mounted screen, etc.
[0091] In a multi-screen cockpit, explicitly binding a user's voice commands to a specific screen and accurately identifying the interaction object can effectively reduce the risk of misoperation and improve the accuracy of command execution.
[0092] Optionally, determining the target vehicle infotainment screen based on the usage status of each vehicle infotainment screen may include: if only one vehicle infotainment screen in the vehicle cabin is in use, then that vehicle infotainment screen in use is determined as the target vehicle infotainment screen; if multiple vehicle infotainment screens in the vehicle cabin are in use, then the vehicle infotainment screen associated with the location of the voice data source is determined as the target vehicle infotainment screen.
[0093] For example, in a typical scenario where multiple users can operate different screens, combining the usage status of each screen can clearly bind the user's voice commands to the screen they are operating, thereby reducing the possibility of accidental operation between the user and the screen. Specifically, by querying the usage status of all in-vehicle screens in the cabin, if only one screen is found to be in use, then each user can only control that screen, so it can be directly locked as the target in-vehicle screen without initiating complex sound source localization analysis; however, if multiple screens (such as the central control screen and the passenger screen) are detected to be in use at the same time, then sound source localization analysis needs to be initiated to obtain the sound source location of the target user's voice data, and compare it with the layout information of the screens in the cabin, determining the screen that best matches in space as the target in-vehicle screen.
[0094] This embodiment determines the target vehicle infotainment screen based on the usage status of each vehicle infotainment screen, providing a clear, efficient, and reliable decision-making logic. In typical multi-user scenarios where multiple users can operate their own screens, it can reduce the possibility of users accidentally operating the screen.
[0095] S204. Obtain software application information and current interface image from the target vehicle's infotainment screen.
[0096] The target vehicle infotainment screen is the screen in the car cabin that corresponds to the voice data, and the software application information includes relevant information about the in-vehicle applications on the target vehicle infotainment screen.
[0097] It should be noted that the specific implementation of step S204 can be referred to the description of step S101, and will not be repeated here.
[0098] S205. Based on the software application information and the current interface image, enhance the recognition of the speech data to obtain the speech recognition result.
[0099] It should be noted that the specific implementation of step S205 can be referred to the description of step S102, and will not be repeated here.
[0100] S206. Based on the voice recognition results, execute the corresponding operation command on the target vehicle screen.
[0101] For example, after obtaining the speech recognition result, the speech recognition result can be further converted into one or more specific, executable operation instructions. Then, through the operating system interface or application protocol, the instruction is sent to a specific application on the target vehicle screen to drive it to complete the corresponding operation, thereby realizing end-to-end automated interaction from "voice input" to "screen operation" and providing a seamless user experience.
[0102] It should be noted that when using the vehicle-mounted voice recognition method of this application in practical applications, some or all of the above steps may be included, and the embodiments of this application do not impose any limitations.
[0103] The in-vehicle voice recognition method provided in this application solves the core problems of ambiguous command attribution and inaccurate intent understanding in complex scenarios by integrating sound source localization, screen status perception, and multimodal information. Specifically, by deploying a microphone array and combining it with sound source separation technology, the method effectively distinguishes the voice commands of different users, solving the voice interference problem in multi-passenger scenarios. By intelligently analyzing the usage status of each vehicle screen and associating it with sound source localization, the method accurately determines the target vehicle screen for the user's intended operation, avoiding misoperation in multi-screen environments. By integrating the software application information of the target vehicle screen with real-time interface images to enhance the recognition of voice data, the method fully utilizes visual context information to significantly improve the accuracy and robustness of voice recognition. Finally, it forms a complete closed loop from voice input to precise screen control, greatly improving the accuracy, scenario adaptability, and user experience of in-vehicle voice interaction.
[0104] Figure 3 This is a schematic diagram of the structure of an in-vehicle voice recognition device provided in an embodiment of this application, as shown below. Figure 3 As shown, the in-vehicle voice recognition device 30 provided in this embodiment includes: a first processing unit 301 and a second processing unit 302.
[0105] The first processing unit 301 is used to respond to the voice data input by the target user and obtain the software application information and the current interface image on the target vehicle screen; wherein, the target vehicle screen is the vehicle screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of the vehicle application on the target vehicle screen.
[0106] The second processing unit 302 is used to enhance the recognition of speech data based on software application information and the current interface image to obtain speech recognition results.
[0107] In one possible implementation, the second processing unit 302 is specifically used for:
[0108] Adjust the speech recognition parameters based on the software application information and the current interface image; the speech recognition parameters are used to optimize the preset speech recognition model.
[0109] Based on the adjusted speech recognition parameters, a preset speech recognition model is configured, and the configured speech recognition model is used to recognize speech data to obtain speech recognition results.
[0110] In one possible implementation, the second processing unit 302 is specifically used for:
[0111] The current interface image is semantically analyzed by a preset image recognition model to extract key information from the current interface image; among which, the key information is text information related to the functions of the currently running in-vehicle application.
[0112] Adjust the speech recognition parameters based on software application information and key information.
[0113] In one possible implementation, the second processing unit 302 is specifically used for:
[0114] Identify target word domains in a pre-defined word library that are related to software application information and key information; wherein, the target word domain includes at least one pre-defined word;
[0115] Adjust the semantic weight and / or semantic priority of the target word domain to be higher than that of other word domains; where other word domains are word domains in the preset lexicon other than the target word domain.
[0116] In one possible implementation, before acquiring the software application information and current interface image from the target vehicle infotainment screen, the first processing unit 301 is further configured to:
[0117] The system obtains the usage status of each in-vehicle infotainment screen in the car cabin and determines the target in-vehicle infotainment screen based on the usage status of each screen.
[0118] In one possible implementation, the first processing unit 301 is specifically used for:
[0119] If only one in-vehicle infotainment screen is in use in the car cabin, then that in-use screen is identified as the target screen.
[0120] If multiple in-vehicle screens are in use in the car cabin, the screen associated with the location of the voice data source will be identified as the target screen.
[0121] In one possible implementation, the second processing unit 302 is further configured to:
[0122] Based on the voice recognition results, execute the corresponding operation commands on the target vehicle's infotainment screen.
[0123] In one possible implementation, the first processing unit 301, before responding to the voice data input by the target user, is further configured to:
[0124] Voice data input from multiple users is collected through a microphone array distributed throughout the car cabin;
[0125] Based on the time difference and phase difference of the voice data received by each microphone, the voice data of each user is separated.
[0126] The apparatus provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0127] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. These modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. Furthermore, they can be stored as program code in the device's memory, and the data processing modules can be called and executed by a specific processing element. The implementation of other modules is similar. These modules can be fully or partially integrated together, or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0128] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.
[0129] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.
[0130] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0131] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0132] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0133] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0134] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0135] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0136] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0137] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0138] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0141] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0143] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A vehicle-mounted voice recognition method, characterized in that, The method includes: In response to voice data input by the target user, the system acquires software application information and the current interface image from the target vehicle infotainment screen; wherein, the target vehicle infotainment screen is the vehicle infotainment screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of the in-vehicle applications on the target vehicle infotainment screen. Based on the software application information and the current interface image, the voice data is enhanced for recognition to obtain a voice recognition result.
2. The method according to claim 1, characterized in that, The step of enhancing the recognition of the voice data based on the software application information and the current interface image to obtain a voice recognition result includes: Based on the software application information and the current interface image, adjust the speech recognition parameters; wherein, the speech recognition parameters are parameters used to optimize the preset speech recognition model; The preset speech recognition model is configured based on the adjusted speech recognition parameters, and the configured speech recognition model is used to recognize the speech data to obtain the speech recognition result.
3. The method according to claim 2, characterized in that, The step of adjusting the speech recognition parameters based on the software application information and the current interface image includes: The current interface image is semantically analyzed using a preset image recognition model to extract key information from the current interface image; wherein, the key information is text information related to the functions of the currently running in-vehicle application. The speech recognition parameters are adjusted based on the software application information and the key information.
4. The method according to claim 3, characterized in that, The step of adjusting the speech recognition parameters based on the software application information and the key information includes: Determine target word domains in a preset word library that are related to the software application information and the key information; wherein, the target word domains include at least one preset word; Adjust the semantic weight and / or semantic priority of the target word domain to be higher than that of other word domains; wherein, the other word domains are word domains in the preset lexicon other than the target word domain.
5. The method according to claim 1, characterized in that, Before acquiring the software application information and current interface image from the target vehicle infotainment screen, the method further includes: The usage status of each vehicle infotainment screen in the car cabin is obtained, and the target vehicle infotainment screen is determined based on the usage status of each vehicle infotainment screen.
6. The method according to claim 5, characterized in that, The step of determining the target vehicle infotainment screen based on the usage status of each of the vehicle infotainment screens includes: If only one in-vehicle infotainment screen is in use in the vehicle cabin, then the in-vehicle infotainment screen in use is determined as the target in-vehicle infotainment screen. If multiple in-vehicle screens in the vehicle cabin are in use, the in-vehicle screen associated with the location of the audio source of the voice data is identified as the target in-vehicle screen.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Based on the voice recognition results, the corresponding operation instructions are executed on the target vehicle screen.
8. The method according to any one of claims 1-6, characterized in that, Prior to responding to voice data input by the target user, the method further includes: Voice data input from multiple users is collected through a microphone array distributed throughout the car cabin; Based on the time difference and phase difference of the voice data received by each microphone, the voice data of each user is separated.
9. A vehicle-mounted voice recognition device, characterized in that, The device includes: The first processing unit is configured to respond to voice data input by the target user and acquire software application information and current interface image from the target vehicle infotainment screen; wherein, the target vehicle infotainment screen is the vehicle infotainment screen in the car cabin corresponding to the voice data, and the software application information includes relevant information of in-vehicle applications in the target vehicle infotainment screen. The second processing unit is used to enhance the recognition of the voice data based on the software application information and the current interface image to obtain the voice recognition result.
10. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.