Speech recognition method of vehicle, vehicle and storage medium

The vehicle voice recognition method, which utilizes multi-dimensional feature extraction and intent recognition, solves the problem of low accuracy in vehicle voice recognition systems, achieving higher recognition accuracy and user experience, especially improving safety and stability in multi-passenger environments.

CN121963726APending Publication Date: 2026-05-01GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU AUTOMOBILE GROUP CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing vehicle voice recognition systems have low accuracy and struggle to accurately recognize user commands, leading to misoperation and insufficient security.

Method used

By extracting multi-dimensional features and recognizing intent, including voiceprint, acoustic and keyword features, a comprehensive user speech recognition model is constructed to determine whether the user intends to exit the voice interaction system, and to control the voice interaction system based on the intent recognition results.

Benefits of technology

It improves the accuracy and response speed of speech recognition, reduces misoperation, and enhances user experience and safety, especially significantly improving the accuracy of speech recognition and system stability in multi-passenger environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963726A_ABST
    Figure CN121963726A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice recognition method of a vehicle, the vehicle and a storage medium, and the method comprises the steps: responding to a received wake-up instruction of a voice interaction system in the vehicle, and collecting voice information outputted by a user in the vehicle; performing multi-dimensional feature extraction on the voice information to obtain voice feature information of the user, the voice feature information being used for representing feature information associated with a target intention, and the target intention being used for representing that the user has an intention of exiting the voice interaction system; based on the voice feature information, intention recognition is carried out on the user, an intention recognition result is obtained, and the intention recognition result is used for representing whether the user has a target intention or not; and controlling the voice interaction system based on the intention recognition result. According to the invention, the technical problem of low voice recognition accuracy of the vehicle in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and more particularly to a voice recognition method for vehicles, a vehicle, and a storage medium. Background Technology

[0002] With the booming development of the intelligent vehicle industry, in-vehicle voice recognition systems have become increasingly important due to their convenience and improved driving safety. Existing technologies mainly rely on silence detection and semantic rejection strategies to determine whether the user is giving a command. However, silence detection and semantic rejection strategies often lack flexibility, making it difficult to accurately recognize user commands, which in turn leads to low voice recognition accuracy in vehicles using related technologies. Summary of the Invention

[0003] This application provides a vehicle voice recognition method, a vehicle, and a storage medium, aiming to improve the technical problem of low accuracy in vehicle voice recognition in related technologies.

[0004] According to one aspect of the embodiments of this application, a voice recognition method for a vehicle is provided, comprising: in response to receiving a wake-up command of a voice interaction system in the vehicle, acquiring voice information output by a user in the vehicle; performing multi-dimensional feature extraction on the voice information to obtain voice feature information of the user, wherein the voice feature information is used to represent feature information associated with a target intent, and the target intent is used to indicate that the user has the intention to exit the voice interaction system; performing intent recognition on the user based on the voice feature information to obtain an intent recognition result, wherein the intent recognition result is used to indicate whether the user has the target intent; and controlling the voice interaction system based on the intent recognition result.

[0005] By extracting multi-dimensional features and recognizing intent from the voice information collected after responding to the wake-up command, the system effectively distinguishes the user's true intent, especially the intent to exit the voice interaction system. This improves the accuracy and response speed of the voice recognition system, reduces erroneous operations, and enhances user experience and security.

[0006] Furthermore, multi-dimensional feature extraction is performed on the speech information to obtain the user's speech feature information, including: extracting voiceprint features from the speech information to obtain the user's voice feature information, wherein the voice feature information is used to represent the user's category information; extracting acoustic features from the speech information to obtain the user's acoustic feature information, wherein the acoustic feature information is used to represent the user's speech output type; extracting keyword features from the speech information to obtain the user's keyword feature information; and determining the speech feature information based on the voice feature information, acoustic feature information, and keyword feature information.

[0007] By combining multi-dimensional feature extraction of voiceprint features, acoustic features, and keyword features, a comprehensive user speech recognition model was constructed. This enables the system not only to identify the user's identity type but also to understand the nature and specific needs of the speech output, significantly enhancing the system's intelligent recognition capabilities and the accuracy of user intent understanding.

[0008] Furthermore, based on voice feature information, user intent is identified to obtain intent identification results, including: determining whether the user is a preset type based on the sound feature information in the voice feature information; if the user is a preset type, determining that the intent identification result is that the user has the intent to exit the voice interaction system; if the user is not a preset type, the user intent is identified based on other feature information in the voice feature information besides the sound feature information to obtain intent identification results.

[0009] By utilizing voiceprint data from voice feature information, the system can quickly determine whether a user belongs to a preset type. If a user belongs to a preset type, the system will directly determine that the user's intention is to exit voice interaction, thus avoiding the risk of accidental operation by children and enhancing the stability and security of the system in specific user scenarios.

[0010] Furthermore, based on other features in the speech feature information besides sound features, the user's intent is identified to obtain an intent identification result, including: determining whether the user is in a preset interaction scenario based on the acoustic features in the speech feature information, wherein the preset interaction scenario is used to represent the scenario in which the user interacts with other users via voice; if the user is in the preset interaction scenario, the intent identification result is determined to be that the user has the intent to exit the voice interaction system; if the user is not in the preset interaction scenario, the user's intent is identified based on other features in the speech feature information besides sound features and acoustic features to obtain an intent identification result.

[0011] By analyzing acoustic feature information, the system can determine whether the user is in a preset interaction scenario. When such a scenario is identified, the system will automatically recognize the user's intention to exit, avoiding false triggering of the vehicle's voice system under non-control intention conditions and improving the voice recognition experience in multi-passenger environments.

[0012] Furthermore, user intent is identified based on other features besides voice features in the voice feature information to obtain intent identification results, including: determining whether the user has a need for voice interaction based on keyword features in the voice feature information; if the user does not have a need for voice interaction, determining the intent identification result as the user has the intent to exit the voice interaction system; if the user has a need for voice interaction, determining the intent identification result as the user does not have the intent to exit the voice interaction system.

[0013] Keyword feature extraction enables the system to accurately understand whether the user has made an actual voice control request. If the system determines that the user is not giving a specific control command but is instead chatting or asking a question, it will identify the user's intention to exit the voice interaction. This intelligent judgment avoids meaningless waiting and responses, improving system efficiency and user satisfaction.

[0014] Furthermore, based on voice feature information, user intent is identified to obtain intent identification results, including: determining whether the user is a preset type based on sound feature information in the voice feature information; determining whether the user is in a preset interaction scenario based on acoustic feature information in the voice feature information; determining whether the user has a voice interaction need based on keyword feature information in the voice feature information; if the user is a preset type, or the user is in a preset interaction scenario, or the user does not have a voice interaction need, the intent identification result is determined to be that the user has the intent to exit the voice interaction system for voice interaction; if the user is not a preset type, and the user is not in a preset interaction scenario, and the user has a voice interaction need, the intent identification result is determined to be that the user does not have the intent to exit the voice interaction system for voice interaction.

[0015] By comprehensively considering voiceprint, acoustic, and keyword features, an intelligent recognition logic is formed. The system remains active only when the user does not meet the preset type, is not in the preset interaction scenario, and has a clear need for voice interaction. Otherwise, the system will exit the voice interaction mode, greatly reducing false triggers and improving the system's intelligence and controllability.

[0016] Furthermore, based on the intent recognition result, the voice interaction system is controlled, including: if the intent recognition result indicates that the user does not have the intent to exit the voice interaction system, the voice information is task-based based on voice feature information to obtain the target task, and the voice interaction system is controlled based on the target task; if the intent recognition result indicates that the user has the intent to exit the voice interaction system, the user exits the voice interaction system.

[0017] Based on accurate intent recognition results, the system can intelligently determine whether to continue executing the task. If the user's intent is clear, the system will execute the corresponding control action; otherwise, it will exit voice interaction, avoiding resource waste and unnecessary operations, making the system more efficient and user interaction safer.

[0018] Furthermore, controlling the voice interaction system based on the target task includes: retrieving the target control parameters corresponding to the target task from multiple preset control parameters, wherein different preset control parameters are used to control the voice interaction system to achieve different tasks.

[0019] By calling the control parameters corresponding to the recognition task, the system can flexibly and efficiently execute user commands, improving the accuracy of task execution and the intelligence level of system response, and ensuring the reliability and smoothness of voice control in different task scenarios.

[0020] Furthermore, the method also includes: in response to receiving preset voice information, obtaining the current configuration information of the voice interaction system, wherein the current configuration information is used to represent the configuration information of the interaction mode of the voice interaction system; generating voice interaction information based on the current configuration information and outputting the voice interaction information; in response to receiving voice feedback information of the voice interaction information, updating the current configuration information based on the voice feedback information to obtain the target configuration information of the voice interaction system.

[0021] By receiving voice feedback from users, the system can automatically update its configuration information. This self-learning and adaptive capability enables the system to continuously evolve, better adapt to the needs of different users and scenarios, and improve the overall voice interaction quality and the system's user-friendliness.

[0022] According to another aspect of the embodiments of this application, a voice recognition device for a vehicle is provided, comprising: a collection module, configured to collect voice information output by a user in the vehicle in response to receiving a wake-up command from a voice interaction system in the vehicle; a feature extraction module, configured to perform multi-dimensional feature extraction on the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent feature information associated with a target intent, and the target intent is used to indicate that the user has the intention to exit the voice interaction system; an intent recognition module, configured to perform intent recognition on the user based on the voice feature information to obtain an intent recognition result, wherein the intent recognition result is used to indicate whether the user has the target intent; and a control module, configured to control the voice interaction system based on the intent recognition result.

[0023] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory for storing a computer program; and a processor for executing the program stored in the memory, wherein the program executes the voice recognition method of the vehicle when it runs.

[0024] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to execute the above-mentioned vehicle voice recognition method. Attached Figure Description

[0025] Figure 1 This is a flowchart of a vehicle voice recognition method according to an embodiment of this application;

[0026] Figure 2This is a flowchart of a vehicle voice recognition method according to an embodiment of this application;

[0027] Figure 3 This is a schematic diagram of a vehicle voice recognition device according to an embodiment of this application;

[0028] Figure 4 This is a schematic diagram of a vehicle provided in one embodiment of this application. Detailed Implementation

[0029] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0030] Terminology Explanation

[0031] Retrieval-Augmented Generation (RAG): RAG is a model that combines retrieval and generation techniques, primarily used for handling open-domain dialogue or information generation tasks. In applications such as intelligent voice assistants and customer service robots, RAG models can enhance the accuracy and richness of generated responses by retrieving a large amount of relevant documents or knowledge base information, thereby providing higher-quality dialogue interaction or content generation services.

[0032] An intelligent agent is a software entity possessing autonomy, responsiveness, and sociality, capable of perceiving its environment and taking actions to achieve specific goals. Intelligent agents can independently perform tasks, interact with the external environment, and communicate and collaborate with other agents or users. In this application, "Agent" specifically refers to the intelligent component in an in-vehicle voice assistant system, which can understand and respond to the user's voice commands, while providing guidance and solutions on how to use the voice system more effectively, enhancing the interactivity and convenience between the user and the voice assistant.

[0033] This application provides a vehicle voice recognition method, comprising: in response to receiving a wake-up command from a voice interaction system in the vehicle, collecting voice information output by a user in the vehicle; extracting multi-dimensional features from the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent feature information associated with a target intent, and the target intent is used to indicate that the user has the intention to exit the voice interaction system; based on the voice feature information, performing intent recognition on the user to obtain an intent recognition result, wherein the intent recognition result is used to indicate whether the user has the target intent; and controlling the voice interaction system based on the intent recognition result.

[0034] The vehicle voice recognition method provided in this application achieves the following technical effects:

[0035] This application responds immediately upon receiving a wake-up command from the vehicle's voice interaction system, collecting real-time voice information output by the user inside the vehicle. This step ensures the system can promptly capture all user voice input, providing foundational data for subsequent multi-dimensional feature analysis. Next, multi-dimensional feature extraction is performed on the voice information to obtain voice feature information associated with the target intent. This series of feature extractions comprehensively depicts the user's current voice state and context, providing rich evidence for subsequent intent recognition. Compared to traditional single-dimensional feature analysis, the introduction of multi-dimensional features significantly enhances the system's ability to understand user intent. Then, based on the extracted voice feature information, user intent is recognized to determine whether the user has expressed an intent to exit the voice interaction system, thus obtaining the intent recognition result. Finally, the voice interaction system is controlled based on the obtained intent recognition result. This application, by comprehensively extracting multi-dimensional features and intelligently recognizing user voice after wake-up, achieves the technical objective of intelligently distinguishing user target intents, significantly improving voice recognition accuracy and user interaction experience, thereby solving the technical problem of low voice recognition accuracy in vehicles in related technologies.

[0036] This application provides a vehicle voice recognition method. Figure 1 This is a flowchart of a vehicle voice recognition method according to an embodiment of this application. Please refer to it. Figure 1 This includes the following steps:

[0037] S102: In response to receiving a wake-up command from the voice interaction system in the vehicle, collect the voice information output by the user in the vehicle.

[0038] The aforementioned vehicle may refer to a vehicle equipped with an advanced intelligent voice interaction system. The type of vehicle may include, but is not limited to, electric vehicles, hybrid vehicles, and gasoline vehicles, with the specific vehicle type to be determined based on the actual situation. The vehicle can serve as the actual carrier for the voice interaction system in the voice recognition method of this application, with the user emitting voice information from within the vehicle.

[0039] The aforementioned voice interaction system refers to an intelligent system capable of receiving and parsing human voice commands, and then executing corresponding operations based on these commands. A voice interaction system may include, but is not limited to, a speech recognition module, a semantic understanding module, an acoustic model, a language model, and a response execution module. Specifically, the speech recognition module is responsible for converting speech into text; the semantic understanding module analyzes the meaning of the text; the acoustic and language models are used for recognizing sound and language features; and the response execution module executes the corresponding function based on the understood instructions. The specific voice interaction system needs to be determined based on the actual design. A voice interaction system allows drivers and passengers to control various vehicle functions, such as navigation, music playback, and temperature adjustment, via voice without manual operation, thereby improving driving safety and convenience.

[0040] The users mentioned above refer to drivers or passengers in the vehicle, who are the primary users of the voice interaction system. Users send commands to the system via voice, expecting quick and accurate responses to meet various needs during the driving process.

[0041] User types need to be determined according to the classification method, which may include, but is not limited to, the following: classification by role (users can be drivers or passengers); classification by age (adults or children, especially children aged 0-6); classification by language habits (users can communicate using Mandarin or other dialects). Specific user types and classification methods need to be determined based on the actual situation and are not limited here.

[0042] The aforementioned voice information can refer to verbal commands or dialogue transmitted by the user to the voice interaction system through a microphone or other recording device. Voice information serves as the basic input for the voice interaction system to function, and the system responds by recognizing and understanding this information.

[0043] The aforementioned wake-up command can refer to a specific voice or physical signal used to activate the voice interaction system, bringing it from standby to a responsive state. Wake-up commands can activate the voice interaction system, preventing the system from wasting computing resources in useless sound environments while ensuring user privacy and security. It enables the system to distinguish between background noise and genuine user commands, thus accurately responding to user needs.

[0044] The types of wake-up commands can include, but are not limited to, the following:

[0045] The first type is a voice wake-up word, such as a phrase like "Hello, Xiao A". Users need to say these specific wake-up words before issuing any other voice commands.

[0046] The second method is physical wake-up, which activates the voice system through a physical button in the vehicle or a specific icon on the touch screen. This is more intuitive and suitable for emergencies or when voice wake-up is not applicable.

[0047] The third method is gesture wake-up, where the system can recognize specific gestures to activate the voice assistant, providing extra convenience by eliminating the need to search for physical buttons while driving.

[0048] The fourth type is context-based wake-up, where the system automatically wakes up based on changes in the vehicle's internal and external environment, such as when the vehicle starts, arrives at a specific location, or wakes up at a specific time, in order to proactively provide services or information.

[0049] The fifth type is continuous wake-up: After a successful wake-up, the system automatically enters a continuous listening state. Users can issue commands continuously without using the wake word again until the system automatically exits the listening mode according to specific logic or the user manually closes it.

[0050] The above wake-up command types are for illustrative purposes only; the specific wake-up command needs to be determined based on the actual situation.

[0051] In one alternative embodiment, in response to receiving a wake-up command from the vehicle's voice interaction system, the voice interaction system immediately becomes active. At this time, the system can employ audio capture technology to collect the user's voice information through an in-vehicle microphone network. These microphones are carefully positioned to cover all seating areas within the vehicle, ensuring that the voice information of both the driver and passengers is clearly captured.

[0052] S104: Perform multi-dimensional feature extraction on the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent feature information associated with the target intent, and the target intent is used to represent the user's intention to exit the voice interaction system.

[0053] The aforementioned multi-dimensional feature extraction refers to the process of extracting a series of feature information representing the speech content, sound characteristics, context, and user identity from the user's output speech information using various algorithms and techniques. Multi-dimensional feature extraction helps voice interaction systems more comprehensively understand user intent, distinguish various types of speech input, and thus make more accurate responses. During the process of exiting the voice interaction system, this feature information can be used to determine whether the user has expressed a willingness to end the conversation, ensuring that the system does not misidentify or delay exit.

[0054] The aforementioned voice feature information refers to a data set extracted from user speech that reflects speech content, sound quality, user identity, and emotional state. Voice feature information may include, but is not limited to, acoustic features, linguistic features, dialogue context, emotional features, and voiceprint features; the specific voice feature information needs to be determined based on actual needs. In voice interaction systems, voice feature information can be used to identify and understand specific user commands or needs, and can also be used to determine whether the user's intent is related to exiting the voice system.

[0055] The aforementioned target intent can refer to a specific action or need that a user wishes to perform, conveyed through a voice interaction system. Identifying the target intent can help the system determine how to respond to the user's instructions, thereby meeting the user's needs or resolving the problem.

[0056] In one optional embodiment, by performing multi-dimensional feature extraction on the voice information, voice feature information can be captured and analyzed to identify the user's target intent, especially the intent to exit. Through these mechanisms, the system can understand the user's intent in a timely and accurate manner, effectively reducing false triggers and improving the overall efficiency and user experience of voice interaction.

[0057] S106: Based on voice feature information, perform intent recognition on the user to obtain intent recognition results, wherein the intent recognition results are used to indicate whether the user has a target intent.

[0058] The aforementioned intent recognition refers to understanding the specific actions or needs a user wants to perform, based on voice feature information extracted from the user's speech, through machine learning models or rule engines. Within the framework of voice interaction, intent recognition is a crucial step in converting sound wave signals into executable commands for the system. It enables the system to correctly respond to user instructions, avoiding misunderstandings or misoperations. Especially when determining whether a user intends to exit the current interaction, accurate intent recognition can significantly improve user experience and system efficiency.

[0059] The aforementioned intent recognition result refers to the conclusion about the user's intent derived by the system through algorithms based on the user's voice feature information. The intent recognition result may include, but is not limited to, whether the user has a target intent or not. The intent recognition result directly guides the system's subsequent behavior. If the system recognizes that the user has a target intent to exit the voice interaction system, it will immediately take measures to end the listening mode; otherwise, it will remain active and prepare to receive more instructions.

[0060] In one optional embodiment, by using intent recognition based on voice feature information, an intent recognition result that reflects the user's true intent can be obtained. This can accurately parse the user's speech, especially when the user expresses the intent to exit the voice interaction system. The system can immediately generate the corresponding recognition result, thereby adjusting its working state in a timely manner, ensuring that the user's needs are responded to appropriately and quickly, and enhancing the security and user-friendliness of the voice interaction system.

[0061] S108: Control the voice interaction system based on the intent recognition results.

[0062] In one optional embodiment, based on the intent recognition result, the voice interaction system will take targeted actions to respond to the user's instructions or needs. If the system parses that the user has not expressed an intention to exit the voice interaction, it further parses the user's intent to meet the user's specific needs. If the system parses that the user has expressed an intention to exit the voice interaction, it terminates the current continuous listening state, shuts down the voice recognition module, and ends the dialogue with the user, avoiding misoperation and resource waste, thereby improving overall interaction efficiency and user experience. This process reflects the flexible and accurate response mechanism of the intelligent voice system, ensuring that every user interaction is both safe and efficient. In this way, the voice interaction system can not only meet the diverse needs of users, but also exit in a timely manner without interfering with the user's normal activities, demonstrating a high level of intelligent and humanized design.

[0063] This application responds immediately upon receiving a wake-up command from the vehicle's voice interaction system, collecting real-time voice information output by the user inside the vehicle. This step ensures the system can promptly capture all user voice input, providing foundational data for subsequent multi-dimensional feature analysis. Next, multi-dimensional feature extraction is performed on the voice information to obtain voice feature information associated with the target intent. This series of feature extractions comprehensively depicts the user's current voice state and context, providing rich evidence for subsequent intent recognition. Compared to traditional single-dimensional feature analysis, the introduction of multi-dimensional features significantly enhances the system's ability to understand user intent. Then, based on the extracted voice feature information, user intent is recognized to determine whether the user has expressed an intent to exit the voice interaction system, thus obtaining the intent recognition result. Finally, the voice interaction system is controlled based on the obtained intent recognition result. This application, by comprehensively extracting multi-dimensional features and intelligently recognizing user voice after wake-up, achieves the technical objective of intelligently distinguishing user target intents, significantly improving voice recognition accuracy and user interaction experience, thereby solving the technical problem of low voice recognition accuracy in vehicles in related technologies.

[0064] Optionally, S104 includes: extracting voiceprint features from the speech information to obtain the user's voice feature information, wherein the voice feature information is used to represent the user's category information; extracting acoustic features from the speech information to obtain the user's acoustic feature information, wherein the acoustic feature information is used to represent the user's speech output type; extracting keyword features from the speech information to obtain the user's keyword feature information; and determining speech feature information based on the voice feature information, acoustic feature information, and keyword feature information.

[0065] The aforementioned voiceprint feature extraction refers to the process of creating a voiceprint template by analyzing the unique attributes of a user's voice. Voiceprint feature extraction can be used to distinguish different speakers, ensuring that the source of voice commands or information is the intended user.

[0066] The aforementioned voice feature information refers to the output result of the voiceprint feature extraction process. It includes the biological and physical attributes of the user's voice and can be used for subsequent user category information determination. Based on the voice feature information, the system can classify users into adults, children, authorized users, etc., in order to take appropriate interaction strategies. Voice feature information can be used to help the system identify and verify the speaker's identity, ensuring that only authorized users can use the voice interaction system. In this application, it is used to identify children's voiceprints for dynamic masking to protect driving safety.

[0067] The aforementioned acoustic feature extraction refers to the analysis of the physical properties of speech signals. Acoustic feature extraction can be used to determine the type of user speech output, that is, to analyze whether the speech is a formal instruction, casual conversation, singing, or other types of speech activity. This is crucial for improving semantic rejection models and reducing false triggers.

[0068] The aforementioned acoustic feature information refers to data extracted from speech signals that can quantify the physical properties of sound. Acoustic feature information can include, but is not limited to, time-domain features, frequency-domain features, and deep learning-based features. These features can be used individually or in combination to enhance the performance of acoustic models. In speech recognition and voice interaction systems, acoustic feature information helps the system understand the nature of the current speech, such as whether it is normal conversation, a command, emotional expression, or background noise. It helps improve the accuracy of speech recognition and the ability to reject non-target speech, especially in intelligent vehicle environments, where it can distinguish the acoustic feature differences between casual conversation, singing, and formal commands, effectively reducing false triggers and improving user experience and safety.

[0069] The aforementioned voice output types can refer to the different natures or classifications represented by the voice content when a user interacts with the system via voice, reflecting the user's interaction mode or purpose with the system.

[0070] The aforementioned keyword feature extraction refers to the process of identifying and extracting a series of words or phrases with key meanings from user speech.

[0071] The aforementioned keyword feature information refers to keywords used to describe user intent. Keyword feature information may include, but is not limited to, keywords such as "navigate to", "play music", "close", and "exit". These are direct manifestations of user intent and can guide the system to perform appropriate processing.

[0072] In one optional embodiment, firstly, voiceprint features are extracted from the voice information to obtain the user's voice feature information. This voice feature information represents the user's category information, such as the difference between a child's voiceprint and an adult's voiceprint. This effectively identifies and dynamically blocks children's voices, preventing accidental voice triggering caused by children's unconscious conversations during driving, thereby improving driving safety. Secondly, acoustic features are extracted from the voice information to obtain the user's acoustic feature information. This acoustic feature information distinguishes the type of the user's voice output, such as whether it is casual conversation, singing, or issuing commands. Through training an acoustic model, the system can accurately judge the user's dialogue content during continuous listening, rejecting non-command voice and reducing accidental triggering. Thirdly, keyword features are extracted from the voice information to obtain keyword feature information, which is used to improve semantic rejection. By refining the domain classification of voice commands, the system can more accurately identify user intent and avoid misidentification of casual conversation content. Finally, voice feature information is comprehensively determined based on voice feature information, acoustic feature information, and keyword feature information. Through multi-dimensional feature judgment, the accuracy of voice recognition is improved, the accidental trigger rate is reduced, and a smoother and safer voice interaction experience is provided to the user.

[0073] Optionally, S106 includes: determining whether the user is a preset type based on the sound feature information in the speech feature information; if the user is a preset type, determining that the intent recognition result is that the user has the intent to exit the voice interaction system; if the user is not a preset type, performing intent recognition on the user based on other feature information in the speech feature information besides the sound feature information, and obtaining the intent recognition result.

[0074] The aforementioned preset types refer to user categories that the system uses to determine whether a user is interacting with belongs to a specific category, based on specific criteria or predefined user classifications. In intelligent voice interaction systems, preset types can include, but are not limited to, any predefined user attributes, such as child, adult, authorized user, and unauthorized user. Preset types can be categorized in several ways, including but not limited to: by age (children (0-6 years, over 6 years old), adults, etc.); by authorization status (authorized users, unauthorized users); by emotional state (happy, sad, angry, etc.); and by language type (Mandarin-speaking users, dialect users, etc.). Specific preset types need to be determined based on the classification method and actual needs. Determining preset types allows the system to adopt different interaction strategies based on the user's specific identity, especially in terms of security and personalized services, providing more accurate and tailored responses. For example, in an intelligent vehicle environment, recognizing a child user can trigger additional safety measures to avoid risks caused by accidental operation by children.

[0075] The aforementioned other feature information refers to various types of voice data collected and analyzed by the system during voice interaction, in addition to the voice feature information required for specific type judgments. Analysis of other feature information helps enhance the accuracy and intelligence of voice recognition.

[0076] In one optional embodiment, firstly, based on the voice feature information in the voice feature information, it is determined whether the user belongs to a preset type, such as a child's voiceprint or specific dialect features. If the user does belong to the preset type, such as if the voiceprint recognition result indicates that the user is a child, the system will determine, according to preset rules, that the user's intent recognition result indicates an intention to exit the interaction with the voice interaction system, thereby avoiding the execution of vehicle control commands that may endanger safety or reducing false triggers. Conversely, if the user does not belong to the preset type, the system further performs user intent recognition based on other feature information, such as semantic analysis of the voice content. Through this technical solution, the system can more accurately determine whether the user is interacting with the voice interaction system, effectively distinguishing between casual conversation and formal commands between the user and the passenger, thereby reducing the false trigger rate and improving the user experience and safety of the in-vehicle voice assistant.

[0077] Optionally, the user's intent is identified based on other features in the speech feature information besides sound feature information to obtain an intent identification result, including: determining whether the user is in a preset interaction scenario based on acoustic feature information in the speech feature information, wherein the preset interaction scenario is used to represent a scenario in which the user interacts with other users via voice; if the user is in the preset interaction scenario, the intent identification result is determined to be that the user has the intent to exit the voice interaction system; if the user is not in the preset interaction scenario, the user's intent is identified based on other features in the speech feature information besides sound feature information and acoustic feature information to obtain an intent identification result.

[0078] The aforementioned preset interaction scenarios refer to a series of voice interaction situations predefined by the system based on common user behaviors and environmental conditions. Defining these preset interaction scenarios helps the system make more reasonable judgments in specific situations. For example, in a multi-person car ride scenario, the system can recognize casual conversations and distinguish them from issuing commands to the system, thereby reducing false triggers and improving user experience.

[0079] The aforementioned other feature information refers to data points collected from user voice input that can assist in intent recognition, in addition to sound feature information and acoustic feature information. This other feature information enriches the system's input dimensions, enabling intent recognition to not only rely on the physical characteristics of sound but also consider the semantic and emotional context of the user's expression. This more comprehensive dataset helps the system to more accurately understand the user's intent and make appropriate responses.

[0080] In one optional embodiment, if the acoustic features in the voice feature information indicate that the user is in a preset interaction scenario, such as conversing with other passengers in the vehicle, the system can recognize that the user does not actually intend to engage in a conversation with the voice interaction system. This allows the system to promptly end the voice listening state, avoiding interference with the user's normal communication and effectively reducing the chance of accidental touches. The effectiveness of this strategy lies in its ability to intelligently distinguish between casual conversation between the user and passengers and the user's commands to the voice assistant, thereby significantly improving the user experience while ensuring the voice assistant's response efficiency. When the user is not in a preset interaction scenario, the system further utilizes other features in the voice feature information, such as speech rate and tone, to determine the user's true intention. This process also helps reduce accidental touches.

[0081] Optionally, user intent is identified based on other features besides voice features in the voice feature information to obtain intent identification results, including: determining whether the user has a need for voice interaction based on keyword features in the voice feature information; if the user does not have a need for voice interaction, determining the intent identification result as the user has the intent to exit the voice interaction system; if the user has a need for voice interaction, determining the intent identification result as the user does not have the intent to exit the voice interaction system.

[0082] The aforementioned voice interaction needs refer to users' behavioral tendencies to communicate, control, or query via voice with an intelligent voice system. Identifying whether a user has a voice interaction need is one of the core functions of a voice recognition system. It determines whether the system should activate the voice recognition module to respond to user input, and how to respond. By accurately determining user needs, the system can improve interaction efficiency, reduce misrecognition, and simultaneously enhance user satisfaction and experience.

[0083] In one optional embodiment, the system determines whether the user has a need to continue voice interaction with the voice interaction system based on keyword feature information in the voice feature information. If the keyword feature information indicates that the user is expressing non-command-like content, such as casual conversation or complaints, it indicates that the user does not have a need for voice interaction. The system will determine that the user intends to exit the current voice interaction and will terminate the voice monitoring mode in a timely manner to prevent accidental activation and improve driving safety and user experience. Conversely, if the keyword feature information indicates that the user still has a need for subsequent operations, such as setting navigation, calling contacts, or controlling in-vehicle devices, it indicates that the user has a need for voice interaction. The system will then determine that the user wants to maintain the connection with the voice interaction system, avoiding premature termination of voice monitoring and interruption of user operations, and ensuring smooth interaction. Through this accurate intent recognition based on keyword feature information, the voice interaction system can more sensitively capture the user's true intentions and effectively reduce the probability of accidental triggering due to misunderstanding of user intentions.

[0084] Optionally, S106 includes: determining whether the user is a preset type based on the sound feature information in the speech feature information; determining whether the user is in a preset interaction scenario based on the acoustic feature information in the speech feature information; determining whether the user has a voice interaction need based on the keyword feature information in the speech feature information; if the user is a preset type, or the user is in a preset interaction scenario, or the user does not have a voice interaction need, determining the intent recognition result as the user having the intent to exit the voice interaction system for voice interaction; if the user is not a preset type, and the user is not in a preset interaction scenario, and the user has a voice interaction need, determining the intent recognition result as the user does not have the intent to exit the voice interaction system for voice interaction.

[0085] In one optional embodiment, the system first determines whether the user belongs to a specific age group, such as children (0-6 years old), or a group with specific language characteristics, such as people who speak a dialect, based on voice feature information. This determines whether the user belongs to a preset type, thereby reducing erroneous voice command activations for specific user types. Next, based on acoustic feature information, the system identifies whether the user is in a preset interaction scenario, such as a multi-person conversation in the vehicle or the user talking to themselves. This step effectively identifies whether the user is communicating with someone outside the voice interaction system. Furthermore, keyword feature information is used to determine whether the user has expressed a specific voice interaction need, such as "navigate to the nearest gas station" or "play my favorite music," to confirm whether the user intends to engage in a conversation with the voice interaction system. If any of the following conditions are met—that is, the user is a preset type of child or dialect speaker, the user is in a multi-person conversation scenario, or the user has not expressed a clear need for the voice interaction system—the system infers that the user intends to withdraw from the conversation with the voice interaction system. This recognition mechanism significantly reduces erroneous voice responses, improving driving safety and user experience. Conversely, if a user does not belong to a preset type, is not in a preset interaction scenario, and expresses a need to interact with the voice interaction system, the system determines that the user does not intend to terminate the conversation, thus maintaining the continuity and efficiency of voice interaction. This refined intent recognition strategy not only enhances the system's adaptability and reduces the false recognition rate, but also ensures that users receive timely responses when they need the voice assistant and quickly exit when they do not, achieving a balance between intelligence and humanization, by setting reasonable thresholds and logical judgments.

[0086] Optionally, S108 includes: if the intent recognition result indicates that the user does not have the intent to exit the voice interaction system, performing task recognition on the voice information based on voice feature information to obtain a target task, and controlling the voice interaction system based on the target task; if the intent recognition result indicates that the user has the intent to exit the voice interaction system, exiting the voice interaction system.

[0087] Task recognition, as described above, refers to the process by which an intelligent voice system analyzes user voice commands to determine the specific type or purpose of the user's request. Based on voice feature information, especially keyword features and semantic understanding, it parses the user's intent and categorizes it into predefined task types. Task recognition can include, but is not limited to, control tasks, query tasks, and navigation tasks; the specific task type needs to be determined based on actual needs. Task recognition is a core step in an intelligent voice interaction system's response to user requests. It determines how the system should interpret user commands and what actions or information needs to be taken. Through accurate task recognition, the system can quickly and accurately meet user needs, improving interaction efficiency and user experience.

[0088] The aforementioned target task can refer to the specific instruction or need that the user wants the system to execute, as determined during the task recognition process. It is a specific operation or information query, derived by the system after parsing the user's voice command. Once the target task is recognized, the voice interaction system can respond to the user in a targeted manner, performing the corresponding operation or providing the required information. Determining the target task serves as a bridge connecting user needs and system actions, ensuring the accuracy and effectiveness of voice interaction.

[0089] In one optional embodiment, when the system detects that the user does not intend to exit the voice interaction system, it further performs task recognition on the user's voice based on the acquired voice feature information. This process analyzes the user's voice content through a deep learning model, determines the user's target task based on context and specific keywords, and whether it's adjusting the air conditioning temperature, playing music, or navigating to a destination, the system can respond accurately and execute the corresponding operation to meet the user's immediate needs. Conversely, if the system detects that the user intends to exit the voice interaction system, such as by using colloquial commands like "exit" or "turn off," or by analyzing voiceprints and tone of voice to determine that the user clearly expresses a desire to end the conversation, the system will immediately terminate the voice listening state and exit the voice interaction mode to avoid accidental touches and disturbances, while improving the system's intelligence and security. This task recognition and exit mechanism based on user intent and voice features not only enhances the user experience of the voice interaction system but also effectively reduces the risk of misoperation, ensuring the safety of drivers and passengers.

[0090] Optionally, controlling the voice interaction system based on the target task includes: retrieving the target control parameters corresponding to the target task from a plurality of preset control parameters, wherein different preset control parameters are used to control the voice interaction system to achieve different tasks.

[0091] The aforementioned preset control parameters refer to a series of pre-set operation instructions or configuration rules in an intelligent voice system to achieve different tasks. These preset control parameters may include, but are not limited to, device control parameters, service execution parameters, and information query parameters; the specific preset control parameters need to be determined based on the task type. The diversity of preset control parameters ensures that the voice interaction system can flexibly respond to various user commands, providing accurate and timely responses whether controlling devices, executing services, or querying information (such as checking the weather or reading messages).

[0092] The aforementioned target control parameters refer to the specific instructions or settings retrieved by the intelligent voice system from multiple preset control parameters based on the user's voice commands and intent recognition results, used to implement the user's specified tasks. Target control parameters ensure that the system can execute user intents accurately and without error. Whether it's a simple device operation or a complex multi-step service, the system can respond according to the requirements of the target parameters, achieving seamless communication between the user and the vehicle's intelligent system.

[0093] In one optional embodiment, when controlling the voice interaction system based on a target task, target control parameters corresponding to the target task are retrieved from multiple preset control parameters. The core of this control mechanism lies in adjusting the response strategy of the voice interaction system using specific control parameters according to different types of voice tasks, such as navigation, telephone calls, and vehicle control. For example, for complex tasks requiring multiple rounds of dialogue, such as setting a navigation destination, the system automatically activates a longer continuous listening window to ensure that the user can continuously issue commands without repeated wake-up calls, improving task execution efficiency. For simple commands, such as adjusting the air conditioning temperature, the system immediately exits the continuous listening state after the command is executed to avoid accidental activation. This dynamically adjusted strategy enables the voice interaction system to more intelligently adapt to user needs in different scenarios, not only reducing the probability of accidental triggering but also improving the user experience and safety for drivers and passengers, achieving refined management and efficient operation of voice control in intelligent vehicle applications.

[0094] Optionally, the method further includes: in response to receiving preset voice information, obtaining current configuration information of the voice interaction system, wherein the current configuration information is used to represent the configuration information of the interaction mode of the voice interaction system; generating voice interaction information based on the current configuration information and outputting the voice interaction information; in response to receiving voice feedback information of the voice interaction information, updating the current configuration information based on the voice feedback information to obtain target configuration information of the voice interaction system.

[0095] The aforementioned preset voice information refers to specific voice commands or sets of commands issued by the user to the voice interaction system, which have been programmed and recognized by the system. Preset voice information can simplify user operations and improve the efficiency and accuracy of human-computer interaction. It ensures that even in noisy environments, users can control the system with clear and easily identifiable commands, without having to remember complex command sequences or text input.

[0096] The aforementioned current configuration information refers to the set of parameters stored in the voice interaction system regarding the system's interaction methods and operating status. This current configuration information guides the real-time operation of the voice interaction system, ensuring that the system can respond to user commands according to the currently set interaction methods. By adjusting these configurations, the system can adapt to the needs and preferences of different users, providing more personalized services.

[0097] The aforementioned voice interaction information refers to voice feedback or prompts generated by the system based on current configuration information, used to communicate with the user. Voice interaction information can enhance communication between the user and the system, ensuring the user understands the system's response or status, while also providing guidance and feedback to improve the user experience.

[0098] The aforementioned voice feedback information refers to the user's response to the voice interaction information output by the system. This can include confirmation, correction, rejection, or request for more information regarding the system's operation. Voice feedback information is an important source for system learning and self-improvement. It helps the system understand whether the user's true intentions have been correctly recognized and whether the system's response meets the user's needs, thereby continuously improving its interaction strategies.

[0099] The aforementioned target configuration information refers to the updated configuration that the voice interaction system is prepared to adopt after correction based on user voice feedback. It reflects user satisfaction with the current interaction method and their personalized needs, guiding the system's subsequent operation. Updating the target configuration information is crucial for the system to dynamically adapt to user needs. Through user feedback, the system can adjust its configuration to provide an interactive experience that better suits user preferences.

[0100] In one optional embodiment, in response to receiving preset voice information, the system acquires the current configuration information of the voice interaction system, which reflects the interaction mode settings of the voice interaction system. Based on this configuration information, the system generates and outputs corresponding voice interaction information to guide the user or provide feedback on the current status. When the system receives feedback from the user on the voice interaction information, it can automatically update the current configuration information based on this feedback to form target configuration information, thereby improving the subsequent voice interaction experience. This dynamic adjustment mechanism can not only adapt to different user needs and usage habits, but also continuously learn and improve in practical applications, reducing false triggers and improving the accuracy of voice recognition, ensuring safety and convenience during driving. By continuously monitoring and adjusting voice interaction parameters, such as recognition thresholds, dialect support, and multi-turn dialogue initiation conditions, the system can more intelligently identify the user's true intentions, reduce erroneous voice responses, and improve the overall user experience.

[0101] In one optional embodiment, a vehicle voice recognition method includes the following steps: dynamic masking of children's voiceprints; rejection based on human-to-human dialogue features; rejection based on dialect feature detection; improvement of semantic rejection model; reduction of the learning threshold for "exiting voice"; differentiation of dialogue exit mechanisms based on task type, and improvement of multi-turn dialogue exit mechanisms; construction of a voice agent that can autonomously answer questions asked by users regarding voice usage.

[0102] The system includes dynamic voiceprint masking for children. Specifically, it determines whether a user's voiceprint belongs to a child aged 0-6 or older than 6. During a 15-second continuous voice listening period, the system detects external sounds and valid intentions. If the detected voiceprint is for a child under 6 years old, voice control of driving-related functions such as windows and the trunk is disabled. The system also informs drivers that opening windows and the tailgate during driving could easily cause accidents. Furthermore, it prevents children from accidentally activating voice commands during the 15-second continuous listening period, thus avoiding potential accidents caused by casual conversation.

[0103] Human-to-human dialogue feature rejection: Because the vocal characteristics of a person singing are different from those of a person speaking, and the vocal characteristics of a person reporting in a formal setting are different from those of a person chatting casually; the vocal characteristics of a person are different in different scenarios. By collecting voices in different scenarios and labeling their features, and training the acoustic model in the recognition process, when a user makes a sound, the acoustic model outputs whether the content of the current user's dialogue is casual conversation or a voice command. During the continuous listening for 15 seconds, even if the user makes a sound, if the acoustic model judges that the user's vocal characteristics are casual conversation rather than human-computer dialogue, then the user command is rejected, reducing accidental touches.

[0104] Dialect Feature Detection and Rejection: Since the pronunciation of words differs between languages, features are collected and labeled using sounds from different languages. These features are then used to train a language model during the recognition process. When a user speaks, the language model outputs whether the current conversation is in a dialect or standard Mandarin. This dialect rejection avoids forced conversion of dialects into standard Mandarin, leading to misrecognition of the speech.

[0105] Improved semantic rejection model: All voice commands are subdivided and categorized into different domains. The user's speech is semantically understood to determine its domain. If the command is determined to be from the domain of navigation, telephone, or vehicle control, it is executed; if it is determined to be casual conversation or meaningless dialogue, it is rejected to avoid false recognition.

[0106] Lowering the learning curve for "exiting voice": Currently, there are usually two exit mechanisms for voice commands: ① exiting voice via steering wheel buttons, ② exiting voice by issuing a voice command. However, some users are unfamiliar with both of these methods. Therefore, this application proposes to generalize the voice "exit" command set, adding colloquial expressions such as "get lost," "exit," "step back," "close," "voice is annoying," "why is he still exiting," "why is he still talking," "shut up," "I don't want to talk to you anymore," etc., to help users quickly end the voice listening time and reduce accidental touches.

[0107] Based on task type, the dialogue exit mechanism is differentiated, and the multi-turn dialogue exit mechanism has been improved: Users do not need to initiate multi-turn dialogues every time they issue a command; whether or not a multi-turn dialogue is initiated depends on whether the need is met, or for voice, whether the task is completed. Voice tasks are categorized into tasks requiring multi-turn dialogues such as "navigation" and "phone calls," and tasks requiring single-turn dialogues such as "checking weather," "music," "videos," and "air conditioning." For example, in a multi-turn task like navigation, the task is considered complete when route planning is initiated. Voice listening automatically starts during address lookup, address selection, and route selection to ensure efficient task completion via voice. Voice automatically exits after task completion to avoid accidental activation. For a single-turn task like air conditioning, voice listening stops automatically after the first command is issued, avoiding accidental activation. Simultaneously, a wake-up-free function is activated. Users can use this function for other needs, reducing the probability of accidental activation. The wake-up-free context association lasts for 15 seconds, a duration derived from the common 15-second conversational pause before the dialogue ends.

[0108] A voice intelligence agent is built, capable of autonomously answering user inquiries about voice usage. Through Retrieval-Augmented Generation (RAG), users can ask voice commands for solutions when encountering problems while driving. A fault assistant provides answers based on the user's questions and the actual state of the vehicle's functions. The values ​​used in the above process are for illustrative purposes only; specific values ​​should be determined based on actual needs and are not limited here.

[0109] In one alternative embodiment, Figure 2 This is a flowchart of a vehicle voice recognition method according to an embodiment of this application, such as... Figure 2 As shown, the process begins with receiving the user's voice signal (e.g., "Why does the voice keep accidentally wake up?").

[0110] Next, determine if wake-up-free is enabled. If it is, determine whether the Smart Listening menu is currently used only by the driver or by the entire vehicle. If it is used by the entire vehicle, you can turn off voice wake-up here to resolve the issue of false voice wake-up. You can then wake up the voice using the steering wheel buttons, and the troubleshooting is complete. If it is only used by the driver, you can turn off Smart Listening to reduce false wake-ups. If you are still not satisfied, you can choose to turn off voice wake-up to resolve the issue.

[0111] If the Smart Listening menu is missing or not set, further investigation can be conducted to determine if multiple wake-up zones are enabled. If so, the Smart Listening selection range can be set to "Driver Only" to reduce false wake-ups and meet the driver's needs. Alternatively, the wake-up zone can be set to "Driver Only" to reduce false wake-ups and meet the driver's needs. The troubleshooting is now complete. If not, Smart Listening can be disabled to reduce false wake-ups. If the problem persists, disabling voice wake-up will resolve the issue, and the troubleshooting is now complete.

[0112] If not (wake-up-free is off), then determine the wake-up sound zone and whether multiple sound zones are enabled. If not, you can turn off voice wake-up here to solve the problem of accidental voice wake-up. You can then wake up the voice using the steering wheel buttons. If yes, you can set the wake-up sound zone to driver only to reduce accidental voice wake-up and meet the needs of the car owner.

[0113] The above process embodies a user-centric design philosophy, not only addressing the common pain point of accidental voice wake-up but also fully considering the personalized needs of car owners in different scenarios. By dynamically adjusting the usage scope of Smart Listening and the settings of the wake-up sound zone, it effectively reduces the accidental wake-up rate and significantly improves the user experience. Furthermore, the introduction of steering wheel buttons as an alternative wake-up method further enhances the system's flexibility and security, ensuring that users have a reliable voice control channel under any circumstances. It also reduces over-reliance on voice recognition technology, balancing technological advancements with practical needs.

[0114] This application embodiment also provides a vehicle voice recognition device 30. Figure 3 This is a schematic diagram of a vehicle voice recognition device according to an embodiment of this application. Please refer to it. Figure 3 It includes: acquisition module 302, feature extraction module 304, intent recognition module 306, and control module 308.

[0115] The acquisition module 302 is used to acquire voice information output by the user in the vehicle in response to receiving a wake-up command from the voice interaction system in the vehicle; the feature extraction module 304 is used to extract multi-dimensional features from the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent feature information associated with the target intent, and the target intent is used to indicate that the user has the intention to exit the voice interaction system; the intent recognition module 306 is used to recognize the user's intent based on the voice feature information to obtain the intent recognition result, wherein the intent recognition result is used to indicate whether the user has the target intent; and the control module 308 is used to control the voice interaction system based on the intent recognition result.

[0116] Optionally, the feature extraction module is used to extract voiceprint features from the speech information to obtain the user's voice feature information, wherein the voice feature information is used to represent the user's category information; to extract acoustic features from the speech information to obtain the user's acoustic feature information, wherein the acoustic feature information is used to represent the user's speech output type; to extract keyword features from the speech information to obtain the user's keyword feature information; and to determine the speech feature information based on the voice feature information, acoustic feature information, and keyword feature information.

[0117] Optionally, the intent recognition module is used to determine whether the user is a preset type based on the sound feature information in the voice feature information; if the user is a preset type, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system; if the user is not a preset type, the intent recognition is performed on the user based on other feature information in the voice feature information besides the sound feature information to obtain the intent recognition result.

[0118] Optionally, the intent recognition module is further configured to determine whether the user is in a preset interaction scenario based on the acoustic feature information in the speech feature information, wherein the preset interaction scenario represents a scenario in which the user interacts with other users via voice; if the user is in the preset interaction scenario, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system; if the user is not in the preset interaction scenario, the intent recognition is performed on the user based on other feature information in the speech feature information besides sound feature information and acoustic feature information, to obtain the intent recognition result.

[0119] Optionally, the intent recognition module is also used to determine whether the user has a need for voice interaction based on the keyword feature information in the voice feature information; if the user does not have a need for voice interaction, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system; if the user has a need for voice interaction, the intent recognition result is determined to be that the user does not have the intent to exit the voice interaction system.

[0120] Optionally, the intent recognition module is further configured to determine whether the user is a preset type based on the voice feature information in the voice feature information; determine whether the user is in a preset interaction scenario based on the acoustic feature information in the voice feature information; determine whether the user has a voice interaction need based on the keyword feature information in the voice feature information; if the user is a preset type, or the user is in a preset interaction scenario, or the user does not have a voice interaction need, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system for voice interaction; if the user is not a preset type, and the user is not in a preset interaction scenario, and the user has a voice interaction need, the intent recognition result is determined to be that the user does not have the intent to exit the voice interaction system for voice interaction.

[0121] Optionally, the control module is configured to, if the intent recognition result indicates that the user does not have the intent to exit the voice interaction system, perform task recognition on the voice information based on the voice feature information to obtain the target task, and control the voice interaction system based on the target task; and if the intent recognition result indicates that the user has the intent to exit the voice interaction system, exit the voice interaction system.

[0122] Optionally, the control module is also used to retrieve the target control parameters corresponding to the target task from a plurality of preset control parameters, wherein different preset control parameters are used to control the voice interaction system to achieve different tasks.

[0123] Optionally, the device is further configured to, in response to receiving preset voice information, obtain current configuration information of the voice interaction system, wherein the current configuration information is used to represent the configuration information of the interaction mode of the voice interaction system; generate voice interaction information based on the current configuration information and output the voice interaction information; and, in response to receiving voice feedback information of the voice interaction information, update the current configuration information based on the voice feedback information to obtain the target configuration information of the voice interaction system.

[0124] This application also provides a vehicle 40, please refer to... Figure 4 It includes a processor 410 and a memory 420, wherein the memory 410 is used to store computer programs; the processor 420 is used to execute the programs stored in the memory 410 to implement the vehicle voice recognition method described in any embodiment of this application.

[0125] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the vehicle voice recognition method described in any embodiment of this application.

[0126] In this application, "multiple" refers to two or more.

[0127] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a physical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0128] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0129] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0130] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if the method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if the method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.

[0131] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice recognition method for vehicles, characterized in that, include: In response to receiving a wake-up command from the voice interaction system in the vehicle, the system collects voice information output by the user inside the vehicle. Multi-dimensional feature extraction is performed on the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent feature information associated with the target intent, and the target intent is used to indicate that the user has the intention to exit the voice interaction system; Based on the voice feature information, the user's intent is identified to obtain an intent identification result, wherein the intent identification result is used to indicate whether the user has the target intent; Based on the intent recognition results, the voice interaction system is controlled.

2. The method according to claim 1, characterized in that, Multi-dimensional feature extraction is performed on the voice information to obtain the user's voice feature information, including: Voiceprint features are extracted from the voice information to obtain the user's voice feature information, wherein the voice feature information is used to represent the user's category information; Acoustic features are extracted from the speech information to obtain the user's acoustic feature information, wherein the acoustic feature information is used to represent the user's speech output type; Keyword feature extraction is performed on the voice information to obtain the user's keyword feature information; The speech feature information is determined based on the sound feature information, the acoustic feature information, and the keyword feature information.

3. The method according to claim 1 or 2, characterized in that, Based on the voice feature information, the user's intent is identified to obtain the intent identification result, including: Based on the voice feature information, determine whether the user is a preset type; If the user is of the preset type, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system. If the user is not the preset type, the user's intent is identified based on other feature information in the voice feature information besides the voice feature information, and the intent identification result is obtained.

4. The method according to claim 3, characterized in that, Based on other features in the speech feature information besides the voice feature information, the user's intent is identified to obtain the intent identification result, including: Based on the acoustic feature information in the voice feature information, it is determined whether the user is in a preset interaction scenario, wherein the preset interaction scenario is used to represent the scenario in which the user interacts with other users via voice. If the user is in the preset interaction scenario, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system. If the user is not in the preset interaction scenario, the user's intent is identified based on other feature information in the voice feature information besides the sound feature information and the acoustic feature information, and the intent identification result is obtained.

5. The method according to claim 4, characterized in that, Based on other features in the speech feature information besides the voice feature information, the user's intent is identified to obtain the intent identification result, including: Based on the keyword feature information in the voice feature information, it is determined whether the user has a voice interaction need; If the user does not have the voice interaction requirement, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system. If the user has the voice interaction requirement, the intent recognition result is determined to be that the user does not have the intent to exit the voice interaction system.

6. The method according to any one of claims 1 or 2, characterized in that, Based on the voice feature information, the user's intent is identified to obtain the intent identification result, including: Based on the voice feature information, determine whether the user is a preset type; Based on the acoustic feature information in the voice feature information, determine whether the user is in a preset interaction scenario; Based on the keyword feature information in the voice feature information, it is determined whether the user has a voice interaction need; If the user is the preset type, or the user is in the preset interaction scenario, or the user does not have the voice interaction requirement, the intent recognition result is determined to be that the user has the intent to exit the voice interaction system. If the user is not the preset type, and the user is not in the preset interaction scenario, and the user has the voice interaction requirement, the intent recognition result is determined to be that the user does not have the intent to exit the voice interaction system.

7. The method according to any one of claims 1-6, characterized in that, Based on the intent recognition result, the voice interaction system is controlled, including: If the intent recognition result indicates that the user does not have the intent to exit the voice interaction system, the voice information is used for task recognition based on the voice feature information to obtain the target task, and the voice interaction system is controlled based on the target task. If the intent recognition result indicates that the user intends to exit the voice interaction system for voice interaction, then the user exits the voice interaction system.

8. The method according to claim 7, characterized in that, Controlling the voice interaction system based on the target task includes: The target control parameters corresponding to the target task are retrieved from a plurality of preset control parameters, wherein different preset control parameters are used to control the voice interaction system to achieve different tasks.

9. The method according to claim 1, characterized in that, The method further includes: In response to receiving preset voice information, the system obtains the current configuration information of the voice interaction system, wherein the current configuration information is used to represent the configuration information of the interaction mode of the voice interaction system; Based on the current configuration information, generate voice interaction information and output the voice interaction information; In response to the voice feedback information received from the voice interaction information, the current configuration information is updated based on the voice feedback information to obtain the target configuration information of the voice interaction system.

10. A vehicle, characterized in that, Including processor and memory, among which, Memory, used to store computer programs; A processor for executing a program stored in memory to implement the method described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 9.