Speech recognition optimization method and device, vehicle and storage medium

By using multimodal sensors to collect data at the vehicle end to optimize the in-vehicle speech recognition model, the problem that cloud-based unified training cannot achieve real-time personalized learning in existing technologies is solved, thereby improving the accuracy of real-time speech recognition and user experience.

CN121789663APending Publication Date: 2026-04-03SAIC GM WULING AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing in-vehicle voice systems rely on unified training in the cloud, which cannot achieve real-time personalized learning. This makes it difficult to correct semantic understanding errors in a timely manner, resulting in a poor user experience.

Method used

By collecting multimodal teaching data based on speech recognition confidence and multimodal sensors at the vehicle end, the speech recognition model is optimized in real time, including immediate and delayed learning mechanisms, to acquire multimodal teaching data and perform iterative optimization.

Benefits of technology

It enables real-time personalized learning of in-vehicle voice recognition models, improving voice recognition accuracy and user experience, and balancing driving safety with learning needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789663A_ABST
    Figure CN121789663A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition optimization method and device, a vehicle and a storage medium, and relates to the technical field of intelligent cabin man-machine interaction, and the method comprises the steps: carrying out the intention recognition of an interaction instruction inputted by a user through a voice recognition model, and obtaining an intention analysis result; performing confidence evaluation on the intention analysis result, and generating learning prompt information when an evaluation result meets a preset learning trigger condition; acquiring multi-modal teaching data acquired based on the learning prompt information; and performing iterative optimization on the speech recognition model according to the multi-modal teaching data. The method does not depend on cloud strong interaction to carry out voice recognition optimization, and directly realizes continuous optimization of the vehicle-mounted voice recognition model at the vehicle end based on the voice recognition confidence and the acquired multi-modal teaching data, so that the voice recognition accuracy can be timely improved, and the user experience is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent cockpit human-computer interaction technology, and in particular to a voice recognition optimization method, device, vehicle, and storage medium. Background Technology

[0002] Current in-vehicle voice systems face significant challenges in handling personalized or complex user commands. Due to the limitations of natural language understanding models, these systems frequently exhibit semantic comprehension errors, leading to the inability to execute voice commands correctly.

[0003] However, the speech recognition function of existing in-vehicle large models relies solely on unified training and batch updates in the cloud, which cannot achieve real-time personalized learning and response. This makes it difficult to correct erroneous semantic understanding in a timely manner, resulting in a poor user experience.

[0004] Therefore, existing in-vehicle voice recognition systems suffer from problems such as reliance on unified cloud training, inability to learn in real time with personalization, and poor user experience, which urgently need to be addressed. Summary of the Invention

[0005] The main purpose of this application is to provide a speech recognition optimization method, device, vehicle, and storage medium, which aims to solve the technical problems of existing in-vehicle speech recognition systems that rely on unified cloud training, cannot learn in real time with personalization, and have poor user experience.

[0006] To achieve the above objectives, this application proposes a speech recognition optimization method, which includes: The intention of the user's interactive commands is identified by a speech recognition model, and the intention parsing result is obtained. The confidence level of the intent parsing result is evaluated, and a learning prompt message is generated when the evaluation result meets the preset learning triggering conditions. Acquire multimodal teaching data based on the learning prompt information; The speech recognition model is iteratively optimized based on the multimodal teaching data.

[0007] In one embodiment, the step of acquiring multimodal teaching data based on the learning prompt information includes: Output the learning prompt information; Receive response instructions from the user based on the output learning prompts; Multimodal teaching data is obtained based on the type of the response instruction, which includes immediate learning instructions or delayed learning instructions.

[0008] In one embodiment, the step of acquiring multimodal teaching data according to the type of the response instruction includes: When the type of the response instruction is an immediate learning instruction, a multi-turn interactive dialogue is initiated. The interactive dialogue is used to guide the user to confirm the instruction text, intent category, and key semantic slot information of the interactive instruction. Obtain multimodal auxiliary information associated with the interaction command collected based on the multi-round interactive session, wherein the multimodal auxiliary information includes at least one of interactive image data, gesture trajectory data, touch coordinate data, or environmental sensing data; The instruction text, the intent category, the key semantic slot information, and the multimodal structure information are subjected to preset structured processing to generate multimodal teaching data.

[0009] In one embodiment, the step of obtaining multimodal teaching data according to the type of the response instruction further includes: When the type of the response instruction is a delayed learning instruction, the context information corresponding to the interaction instruction is stored in a preset cache. Acquire vehicle operating status parameters and occupant detection signals; When the vehicle operating status parameters and the occupant detection signal meet the preset trigger conditions, the context information is retrieved from the preset buffer. The learning prompt information is output again based on the context information until the type of the response instruction corresponding to the learning prompt information is an immediate learning instruction.

[0010] In one embodiment, the step of iteratively optimizing the speech recognition model based on the multimodal teaching data includes: Skill items are generated based on the multimodal teaching data, and the skill items are stored in their own learning record database according to time sequence; Perform a pre-defined association analysis on the skill items to identify target sharing users; Send the sharing suggestions corresponding to the skill entries to the target sharing users; If the target sharing user confirms the sharing information based on the sharing suggestion, the skill item will be sent to the learning record database of the vehicle equipment corresponding to the target sharing user.

[0011] In one embodiment, the step of performing a preset association analysis on the skill items to determine the target sharing user includes: A skill universality analysis was performed on the skill items to obtain a universality index; The skill items are analyzed for closeness based on current user relationship data to obtain a relationship closeness index. The current user relationship data includes explicit user data and implicit behavioral data. A shared recommendation score is generated based on the generality index and the relationship tightness index, and target shared users are determined according to a preset recommendation threshold and the shared recommendation score.

[0012] In one embodiment, after generating skill items based on the multimodal teaching data and storing the skill items in its own learning record database according to time sequence, the method further includes: Upon receiving a user's error reporting command, determine the fault skill item corresponding to the error reporting command; Delete the multimodal teaching data corresponding to the fault technology item in its own learning record database, and re-acquire the optimized teaching data corresponding to the fault skill item; The speech recognition model is iteratively optimized based on the optimized teaching data.

[0013] Furthermore, to achieve the above objectives, this application also proposes a speech recognition optimization device, which includes: The instruction analysis module is used to identify the intent of the user's interactive instructions through a speech recognition model and obtain the intent parsing results. The learning trigger module is used to evaluate the confidence level of the intent parsing result and generate learning prompt information when the evaluation result meets the preset learning trigger conditions. A multimodal data acquisition module is used to acquire multimodal teaching data based on the learning prompt information. The model optimization module is used to iteratively optimize the speech recognition model based on the multimodal teaching data.

[0014] In addition, to achieve the above objectives, this application also proposes a vehicle, the device including: a memory, a processor, and a speech recognition optimization program stored in the memory and executable on the processor, the speech recognition optimization program being configured to implement the steps of the speech recognition optimization method as described above.

[0015] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing the speech recognition optimization method is stored, and the program implementing the speech recognition optimization method is executed by a processor to implement the steps of the speech recognition optimization method as described above.

[0016] This application provides a speech recognition optimization method, device, vehicle, and storage medium. The method includes: performing intent recognition on user-input interactive commands using a speech recognition model to obtain intent parsing results; evaluating the confidence level of the intent parsing results, and generating learning prompt information when the evaluation results meet preset learning trigger conditions; acquiring multimodal teaching data collected based on the learning prompt information; and iteratively optimizing the speech recognition model based on the multimodal teaching data. This application does not rely on strong cloud-based interaction for speech recognition optimization. Instead, it directly performs real-time personalized learning on the vehicle side based on speech recognition confidence level and multimodal teaching data collected by local multimodal sensors in the vehicle, achieving continuous optimization of the in-vehicle speech recognition model. This not only improves speech recognition accuracy in a timely manner but also enhances the user experience. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a first flowchart illustrating the first embodiment of the speech recognition optimization method of this application; Figure 2 This is a second flowchart illustrating the first embodiment of the speech recognition optimization method of this application; Figure 3 This is a schematic diagram of the first process of the second embodiment of the speech recognition optimization method of this application; Figure 4 This is a schematic diagram of the second process of the second embodiment of the speech recognition optimization method of this application; Figure 5 A simplified flowchart of the speech recognition optimization method of this application is provided; Figure 6 This is a schematic diagram of the module structure of the speech recognition optimization device according to an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the speech recognition optimization method in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of this application is as follows: to perform intent recognition on the interactive commands input by the user through a speech recognition model to obtain intent parsing results; to evaluate the confidence of the intent parsing results and generate learning prompt information when the evaluation results meet the preset learning trigger conditions; to acquire multimodal teaching data collected based on the learning prompt information; and to iteratively optimize the speech recognition model based on the multimodal teaching data.

[0024] Current in-vehicle voice systems face significant challenges in handling personalized or complex user commands. Due to the limitations of natural language understanding models, systems often fail to understand commands or have insufficient confidence, resulting in incorrect execution of voice commands. Furthermore, existing optimization solutions for in-vehicle voice systems, which primarily rely on unified cloud training and batch updates, cannot achieve real-time personalized learning and response, making it difficult to correct semantic misunderstandings in a timely manner.

[0025] To address the aforementioned issues, this application does not rely on strong cloud-based interaction for speech recognition optimization. Instead, it performs real-time personalized learning on the vehicle side based on speech recognition confidence and multimodal teaching data collected by local multimodal sensors in the vehicle. This enables continuous optimization of the in-vehicle speech recognition model, which not only improves speech recognition accuracy in a timely manner but also enhances the user experience.

[0026] It should be noted that the executing entity in this embodiment can be a voice recognition optimization system, or an in-vehicle terminal with data processing, network communication, and program execution functions, such as an in-vehicle computer, or a vehicle capable of performing the above functions. This embodiment does not specifically limit the specific implementation. The following description uses an in-vehicle terminal as the executing entity to illustrate this embodiment and the following embodiments.

[0027] Based on this, embodiments of this application provide a speech recognition optimization method, referring to... Figure 1 , Figure 1 This is a first flowchart illustrating the first embodiment of the speech recognition optimization method of this application.

[0028] In this embodiment, the speech recognition optimization method is applied to a vehicle, and the method includes steps S10 to S40: Step S10: The user's input interaction command is identified by a speech recognition model to obtain the intent parsing result; Step S20: Calculate the confidence level of the intent parsing result, and generate a learning prompt message when the evaluation result meets the preset learning triggering conditions; It is easy to understand that, in this embodiment, the user-inputted interactive commands can be control requests or information query requests issued by the user to the vehicle via voice, such as commands for air conditioning control, navigation settings, music playback, etc. The aforementioned voice recognition model can refer to the Automatic Speech Recognition (ASR) module and Natural Language Understanding (NLU) module integrated into the in-vehicle terminal, which can be used to convert user-input voice into text and parse the intent. Specifically, ASR can be used to convert interactive commands into text data, while NLU is responsible for parsing the semantic information in the text data and outputting the intent and key slots corresponding to the interactive commands.

[0029] Therefore, the intent parsing results described above can be structured data output by the NLU module, which may include intent categories (such as air conditioning control, navigation, and weather inquiries), key semantic slots (slots refer to key information input by the user), and slot values ​​(specific parameters of the key information, such as "temperature value"). For example, when the user inputs the interaction command as the voice command "set the air conditioning temperature to 23 degrees," the ASR module can first convert the user's voice input "set the air conditioning temperature to 23 degrees" into text, and the NLU module then parses out the intent of "air conditioning control," the slot for "temperature," and the slot value of "23℃." Alternatively, when the interaction command is "go to Beijing tomorrow," the intent could be "navigation," and the slot and slot value could be "time = tomorrow" and "destination = Beijing."

[0030] Furthermore, the aforementioned intent parsing results may also include the parsing confidence score output by the NLU module. In this case, the confidence assessment can be based on a preset confidence threshold (which can be set by technicians according to the actual scenario, such as 0.7). For example, in this embodiment, the aforementioned preset learning trigger condition can be set to a confidence score lower than the preset confidence threshold (such as lower than 80%) or complete parsing failure (i.e., the NLU module cannot parse).

[0031] In the specific implementation, if the NLU module outputs an intent parsing result and its confidence level is higher than a preset threshold, the function corresponding to the interactive instruction is successfully executed; if the NLU module outputs an intent parsing result but its confidence level is lower than the preset threshold, it is judged as "low confidence", triggering the subsequent learning optimization mechanism and generating learning prompt information; if the NLU module cannot parse the instruction at all, that is, the intent parsing result is that it cannot be mapped to any known intent or slot, it is judged as "explicit failure", directly triggering the learning optimization mechanism and generating learning prompt information.

[0032] Step S30: Obtain multimodal teaching data based on the learning prompt information; It is understandable that the aforementioned learning prompts may be prompts used to ask users whether they wish to participate in model teaching, and may be output in a multimodal manner, such as UI animations on the vehicle's central control screen and voice assistant broadcasts.

[0033] Furthermore, existing technologies lack flexible user-participatory teaching mechanisms, failing to adapt to the needs of users in different scenarios (such as while driving or during leisure time), resulting in low user engagement and insufficient data collection efficiency. To address this issue and ensure driving safety, in this embodiment, the aforementioned multimodal teaching data collection process can be divided into an immediate collection process and a delayed collection process. Therefore, in a feasible implementation, referring to... Figure 2 , Figure 2 This is a second flowchart illustrating the first embodiment of the speech recognition optimization method of this application. In this embodiment, step S30 may include steps A1 to A3: Step A1: Output the learning prompt information; Step A2: Receive the response instruction from the user based on the output learning prompt information; Step A3: Obtain multimodal teaching data according to the type of the response instruction, wherein the type of the response instruction includes immediate learning instructions or delayed learning instructions.

[0034] It is easy to understand that after the learning optimization mechanism is triggered and learning prompts are generated, the in-vehicle terminal can display slight animations on the edge of the central control screen, such as flashing icons and gradient halos; at the same time, the voice assistant will output learning prompts in a multimodal format, asking the user whether to learn immediately.

[0035] The aforementioned response commands can be user actions in response to learning prompts, including immediate learning commands (users choose to participate in teaching immediately) and delayed learning commands (users choose to participate in teaching later). The feedback methods for these commands can be voice or touch. Voice commands correspond to users replying with voice commands, such as "learn now" or "we'll talk about it later." Touch commands correspond to users directly clicking on different corresponding option buttons displayed on the central control screen.

[0036] Therefore, in this embodiment, the user can provide a response command corresponding to "instant learning" or "delayed learning" via voice command or touch-screen selection. The vehicle terminal receives the response command in real time and identifies its type. If the response command is an instant learning command, the vehicle terminal can directly enter the multimodal teaching data acquisition process; if it is a delayed learning command, the vehicle terminal can first store the relevant context information and restart data acquisition after the preset trigger conditions are met, ultimately completing the acquisition of multimodal teaching data.

[0037] Therefore, in the first feasible implementation, step A3 may include steps A31 to A33: Step A31: If the type of the response instruction is an immediate learning instruction, initiate a multi-turn interactive dialogue. The interactive dialogue is used to guide the user to confirm the instruction text, intent category, and key semantic slot information of the interactive instruction. Step A32: Obtain multimodal auxiliary information associated with the interaction command collected based on the multi-round interactive session. The multimodal auxiliary information includes at least one of interactive image data, gesture trajectory data, touch coordinate data, or environmental sensing data. It's important to understand that the aforementioned multi-turn interactive dialogue refers to multiple question-and-answer interactions between the in-vehicle terminal and the user via a voice assistant. Its core purpose is to further clarify key information about the command based on the interactive command text output by the ASR and the intent parsing results output by the NLU module. Therefore, the logic of this multi-turn interactive dialogue mainly revolves around the command text, intent category, and key semantic slot information of the interactive command.

[0038] The aforementioned multimodal assistance information can be non-textual data collected by multimodal sensors installed in the vehicle to enhance command recognition. This data may include interactive image data collected by in-vehicle cameras, gesture trajectory data collected by gesture sensors, touch coordinate data collected by the central control screen, and environmental sensing data such as temperature / humidity collected by environmental sensors. It is easy to understand that all privacy-related content in this multimodal assistance information is based on the user's consent.

[0039] Step A33: Perform preset structured processing on the instruction text, the intent category, the key semantic slot information, and the multimodal structure information to generate multimodal teaching data.

[0040] It is easy to understand that the aforementioned pre-defined structured processing can refer to the standardization of various types of data collected by the vehicle terminal, including timestamp alignment (associating voice text, gesture data, and image data in chronological order), semantic annotation (clarifying intent labels, slot names, and values), and format unification (converting to a dataset format that the model can recognize).

[0041] Therefore, in the specific implementation, suppose user A, while driving the vehicle, first issues the voice command "Set the air conditioning temperature to 23 degrees Celsius, but not too dry." The ASR module successfully converts the speech to text, but the NLU module cannot fully parse the newly added condition "but not too dry," and thus judges it as "low confidence." At this point, after the user provides feedback on the instant learning command, the in-vehicle terminal automatically initiates a multi-turn interactive dialogue: First, the system asks the user about the intent category of their command (e.g., "Do you want to control the air conditioning?"). Then, it confirms the accuracy of the command text (e.g., "You said 'set the air conditioning to 23 degrees but don't make it too dry,' right?"). Finally, it guides the user to supplement key semantic information (e.g., by asking the voice assistant, "What specific function does 'don't make it too dry' refer to?", and user A replies, "When the air conditioning is turned on, also turn on the humidity control function," the supplemented information would then be "humidity control"). Accordingly, the in-vehicle terminal can add a "humidity control" slot under the "air conditioning control" intent and set its value to "on".

[0042] Then, for complex commands requiring physical references, gestures, or location information (such as "open the passenger glove compartment" or "adjust the driver's seat backrest angle"), the in-vehicle terminal can provide further guidance. This involves automatically activating the in-vehicle camera, gesture sensor, touchscreen, and environmental sensors upon user permission to collect corresponding multimodal auxiliary information, including images, gesture trajectories, touch coordinates, and environmental data. For example, the voice assistant can ask, "Do you need a gesture to reinforce this command?" If user A replies "yes," the voice assistant can further guide user A to make a gesture (A) in front of the air vent. Simultaneously, the in-vehicle camera can record a video clip of this gesture, thus collecting the aforementioned interactive image data.

[0043] Finally, the in-vehicle terminal performs pre-defined structured processing, including timestamp alignment, semantic annotation, and format unification, on the confirmed command text, intent category, key semantic slot information, and collected multimodal auxiliary information, ultimately generating complete multimodal teaching data. For example, corresponding to the above process, the command text "set the air conditioner to 23 degrees but not too dry," the final confirmed semantic structure (intent: air conditioner control; slot: temperature = 23℃, humidity hold = on), and the associated gesture video data can be processed with pre-defined structured processing such as timestamp alignment, structured annotation, and encryption to generate complete multimodal teaching data. If, three days later, user A again utters the command "set the air conditioner to 23 degrees but not too dry," the in-vehicle terminal can successfully recognize and execute it: setting the air conditioner temperature to 23℃ and simultaneously activating the humidity hold function.

[0044] Therefore, it is difficult to accurately capture key information of complex or personalized instructions (such as implicit functional requirements and gesture associations) through single text data, resulting in insufficient training sample quality and limited recognition accuracy. This embodiment can ensure the accuracy of core instruction information through multi-turn interactive dialogue, enrich the sample dimensions by combining multimodal auxiliary information, and ensure data standardization through structured processing, thereby significantly improving the quality of training samples. This enables the model to accurately recognize complex and personalized instructions and further optimize the speech recognition effect.

[0045] Furthermore, since providing learning instructions during driving can distract users and affect driving safety, and existing technologies lack a delayed learning mechanism, some users abandon the learning process due to the inconvenience of immediate participation, resulting in incomplete data collection. To avoid this problem, in the second feasible implementation, step A3 in this embodiment may include steps A34-A37: Step A34: If the type of the response instruction is a delayed learning instruction, store the context information corresponding to the interaction instruction in a preset cache. It is easy to understand that the aforementioned context information can be complete data related to the current interaction command, including the original speech audio, ASR-converted text, preliminary NLU parsing results (if any), and current vehicle scene information (such as driving status, gear position, and ambient temperature). The preset cache area can be a secure local storage area of ​​the vehicle. In this embodiment, it can be used to encrypt and store the context information of the delayed learning task, and has data encryption and anti-tampering functions.

[0046] Step A35: Obtain vehicle operating status parameters and occupant detection signals; Step A36: When the vehicle operating status parameters and the occupant detection signal meet the preset trigger conditions, retrieve the context information from the preset buffer. Step A37: Output the learning prompt information again based on the context information until the type of the response instruction corresponding to the learning prompt information is an immediate learning instruction.

[0047] It is easy to understand that the aforementioned vehicle operating status parameters may include parameters such as vehicle gear position (P or D, etc.), ignition status (started or off), and driving speed. The occupant detection signal can be a signal collected by devices such as in-vehicle seat sensors and cameras to determine whether the occupants are present; for example, a signal from the driver's seat sensor indicates that the driver is present.

[0048] At this time, the above-mentioned preset triggering conditions can represent the specific operating scenario of the pre-set start-up delay learning prompt. For example, in this embodiment, it may include vehicle operating status parameters and passenger detection signals corresponding to the vehicle being idle, such as the end of the current ignition cycle (vehicle is turned off), the next vehicle start and in P gear (parked state), or there are people in the driver's seat but the vehicle is not moving.

[0049] At this time, when the vehicle terminal detects that the preset trigger conditions are met, it can retrieve the stored context information from the preset cache area and output the learning prompt information again through UI animation and voice broadcast until the user gives feedback on the immediate learning command and then executes the multimodal data acquisition process shown in steps A31 to A33 above.

[0050] Therefore, this embodiment can achieve delayed learning by caching the context information of interactive instructions and setting precise preset trigger conditions, avoiding interference with users during driving, balancing learning needs and driving safety; and ensuring that no potential teaching data is missed, improving the completeness of data collection, and providing more sample support for model optimization.

[0051] The above are merely two feasible implementations of step A3 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step A3. Therefore, this embodiment can provide a multimodal prompting method to ensure user perception, while supporting both immediate and delayed learning options to adapt to different user scenarios and enhance user participation; it accurately acquires data based on response command types, ensuring the effective collection of multimodal teaching data and laying the foundation for model optimization.

[0052] Step S40: Iteratively optimize the speech recognition model based on the multimodal teaching data.

[0053] It is easy to understand that the vehicle terminal can perform structured processing such as local anonymization and encryption on multimodal teaching data to generate multimodal training samples, and upload the multimodal training samples to the cloud for incremental training of the speech recognition models of the current vehicle and other vehicles, updating the model parameters to improve the ability to recognize user-inputted voice commands.

[0054] In this embodiment, without relying on strong cloud-based interaction or a large amount of pre-labeled data, continuous optimization and personalized customization of the model can be achieved solely by utilizing the vehicle's existing multimodal sensors and in-vehicle infotainment system. Furthermore, it responds to semantic understanding errors in real time and corrects the model promptly through multimodal data, significantly improving the accuracy of speech recognition and user experience.

[0055] Furthermore, through the collaborative design of "multimodal learning prompts" and "delayed teaching," this embodiment enables the system to guide users to participate in teaching in a low-intrusive manner at appropriate times, balancing the conflict between immediate learning and driving safety and user attention.

[0056] This embodiment provides a speech recognition optimization method, which includes: performing intent recognition on user-input interactive commands using a speech recognition model to obtain intent parsing results; evaluating the confidence level of the intent parsing results, and generating learning prompt information when the evaluation results meet preset learning trigger conditions; outputting the learning prompt information; receiving response commands from the user based on the output learning prompt information; acquiring multimodal teaching data according to the type of response command, which includes immediate learning commands or delayed learning commands. When the response command is an immediate learning command, a multi-turn interactive dialogue is initiated, which guides the user to confirm the command text, intent category, and key semantic slot information of the interactive command; acquiring multimodal auxiliary information associated with the interactive command based on the multi-turn interactive dialogue, which includes at least one of interactive image data, gesture trajectory data, touch coordinate data, or environmental sensing data; and performing preset structured processing on the command text, intent category, key semantic slot information, and multimodal structure information to generate multimodal teaching data. When the response command is a delayed learning command, the context information corresponding to the interaction command is stored in a preset cache; vehicle operating status parameters and occupant detection signals are acquired; when the vehicle operating status parameters and occupant detection signals meet preset trigger conditions, the context information is retrieved from the preset cache; learning prompts are output again based on the context information until the response command corresponding to the learning prompt is an immediate learning command. The speech recognition model is iteratively optimized based on multimodal teaching data. This embodiment does not rely on strong cloud interaction or a large amount of pre-labeled data, but only utilizes the vehicle's existing multimodal sensors and vehicle system to achieve continuous optimization and personalized customization of the model; and responds to semantic understanding errors in real time, correcting the model in a timely manner through multimodal data, significantly improving the accuracy of speech recognition and user experience. In addition, through the collaborative design of "multimodal learning prompts" and "delayed teaching," this embodiment can guide users to participate in teaching in a low-intrusive manner at appropriate times, balancing the contradiction between immediate learning and driving safety and user attention.

[0057] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.

[0058] It is important to understand that, in addition to the problem that the rich multimodal sensor data in the in-vehicle environment has not been fully utilized for intent understanding, existing solutions also have limitations such as the isolation of learning outcomes among different family members, which prevents the realization of the synergistic effect of knowledge sharing and seriously restricts the further improvement of the intelligent cockpit interaction experience.

[0059] Therefore, based on the first embodiment, please refer to Figure 3 , Figure 3This is a first flowchart illustrating the second embodiment of the speech recognition optimization method of this application. In this embodiment, after step S40, the speech recognition optimization method further includes steps B1 to B4: Step B1: Generate skill items based on the multimodal teaching data, and store the skill items in their own learning record database according to time sequence; It should be noted that after the model iterative optimization is completed, the skill items generated by the system can be standardized data units formed by structuring multimodal teaching data. These units may include instruction text, semantic structure (intent category and key semantic slot information), associated multimodal tags (such as image tags and gesture tags), and timestamps. The aforementioned learning record database can be a data storage module in the vehicle that stores skill items. It can organize data according to time sequence (the order in which skills are generated) and support operations such as querying, calling, and deleting skill items. Externally, this learning record database can be presented as a persistent application entry, "Evolution Log," on the vehicle's display screen, showcasing new skill items learned within a specific time period in a timeline or card stream format.

[0060] Step B2: Perform a preset association analysis on the skill items to determine the target sharing users; It is easy to understand that in order to effectively push newly learned skill items to the sharing users who need them, this embodiment can analyze the universality of skill items and the relationship between users, that is, perform the above-mentioned preset association analysis, thereby filtering out target sharing users who may need to learn new skill items.

[0061] In one feasible implementation, refer to Figure 4 , Figure 4 This is a second flowchart illustrating the second embodiment of the speech recognition optimization method of this application. In this embodiment, step B2 may include steps B21 to B23: Step B21: Perform a skill universality analysis on the skill items to obtain a universality index; It should be understood that the aforementioned generality index can be a parameter measuring the degree to which a skill item is applicable to other users. In this embodiment, the generality index can be calculated by weighting parameters from four dimensions: usage frequency, scenario generality, semantic generalization ability, and learning cost. The weights of each parameter can be pre-configured by technical personnel, such as setting the usage frequency weight to 0.3, scenario generality to 0.25, semantic generalization ability to 0.25, and learning cost to 0.2. Specific values ​​are not limited in this embodiment.

[0062] It should be noted that the above usage frequency refers to the frequency with which users call the skills corresponding to each skill item through the "evolution log" after they are stored in the learning record database; the scenario universality can be determined by whether the vehicle function associated with the skill item is a commonly used function of most users (the specific proportion can be set according to the actual situation); the semantic generalization ability can be determined by analyzing the structure and keywords of the instruction text corresponding to the skill item through natural language processing technology to determine whether it is easy for other users to understand and express naturally; the learning cost can be determined by evaluating the time and step complexity of teaching the skill corresponding to the skill item.

[0063] Step B22: Perform a closeness analysis on the skill items based on the current user relationship data to obtain a relationship closeness index. The current user relationship data includes explicit user data and implicit behavioral data. Understandably, the aforementioned user relationship data can be parameters used to analyze the degree of correlation between newly generated skill entries and potential users. This data may include explicit user data, such as family account groups (pre-defined in the vehicle terminal, users bound to the same vehicle or family cloud account), frequently used contacts (determined by synchronizing the user's mobile phone address book or vehicle call records), and implicit behavioral data, such as function usage overlap (i.e., the similarity of vehicle function usage between different users, determined through user profile analysis), and vehicle sharing mode (determined through vehicle account switching records). It is readily understood that all of the aforementioned user relationship data is obtained with the user's consent. Therefore, in this embodiment, user identity can be pre-distinguished through voiceprint recognition and vehicle account login status, and each user's teaching records, preference settings, and learning samples need to be independently and encryptedly stored in their personal cloud archive.

[0064] At this point, the aforementioned relationship density index can be a score of the degree of association between users calculated based on user relationship data. It can be generated by summing the acquired explicit user data and implicit behavioral data after assigning preset weights to them by the vehicle terminal.

[0065] Step B23: Generate a shared recommendation score based on the universality index and the relationship tightness index, and determine the target shared users according to the preset recommendation threshold and the shared recommendation score.

[0066] It should be noted that the in-vehicle terminal can generate a score for determining whether to push a sharing suggestion to a user by weighting and summing the universality index and the relationship density index according to a preset fusion ratio (e.g., universality index accounts for 50% and relationship density index accounts for 50%). This score is the aforementioned sharing recommendation score. The preset recommendation threshold can be a critical value (e.g., 0.6) for determining whether to trigger a sharing suggestion, and is set by technicians according to the actual scenario. Therefore, after the in-vehicle terminal determines the sharing recommendation score for a user based on the preset fusion ratio, if the score is higher than the preset recommendation threshold, the user can be determined to be a target sharing user.

[0067] Therefore, this embodiment can accurately calculate the sharing recommendation score through multi-dimensional generality analysis and user relationship analysis, thereby achieving precise screening of target sharing users; and effectively avoiding invalid sharing pushes, improving user acceptance, while ensuring that skills are shared with users who truly need them, thus improving sharing efficiency and resource utilization.

[0068] Step B3: Send the sharing suggestion corresponding to the skill item to the target sharing user; Step B4: If the target sharing user confirms the sharing information based on the sharing suggestion, the skill item is sent to the learning record database of the vehicle equipment corresponding to the target sharing user.

[0069] It's easy to understand that the aforementioned sharing suggestions can be push notifications containing the skill name and function description of the skill item, which can be sent to the target sharing user via the vehicle's infotainment system or the cloud. For example, user A's in-vehicle terminal can push a sharing suggestion to the target sharing user, i.e., user B's in-vehicle terminal, such as "User A has added a new skill, <Air Conditioning Temperature Adjustment and Humidification>, do you want to learn it?"

[0070] If the target user confirms the sharing information via voice or touch feedback, the in-vehicle terminal can further simplify the skill items, generating a structured data packet that retains the semantic and multimodal frameworks. After encrypting the structured data packet, it is synchronized to the learning record database of all authorized vehicle devices corresponding to the target user via the vehicle's message center or the cloud, thus achieving skill sharing. If the target user refuses to share, the structured data packet corresponding to the skill item will not be shared with that user.

[0071] Therefore, this embodiment can achieve cross-user and cross-device sharing of personalized learning outcomes through a skill item storage and sharing mechanism, avoiding repetitive teaching and improving user efficiency; at the same time, it expands the application scope of skills, allowing more users to benefit from personalized optimization and enhancing the collaborative interactive experience of the product.

[0072] It should be noted that the key semantic slot information contained in the skill items shared across devices is a logical label, not an absolutely unique hardware instruction. If the vehicle configuration being synchronized is outdated, it may lack the corresponding hardware function or software interface for the technology item to be synchronized. In this case, the vehicle's fault tolerance mechanism or degradation processing mechanism can search for the most semantically similar available function in the local function library for association. Ultimately, in practice, this may lead to errors in function association. Therefore, to avoid this problem, in another feasible implementation, in this embodiment, step B1 may be followed by steps C1 to C3: Step C1: Upon receiving a user input error command, determine the fault skill item corresponding to the error command; Step C2: Delete the multimodal teaching data corresponding to the fault technology item in the learning record database of the user, and re-acquire the optimized teaching data corresponding to the fault skill item; Step C3: Iteratively optimize the speech recognition model based on the optimized teaching data.

[0073] It is important to understand that each skill card in the "Evolution Log," a permanent application entry on the in-vehicle terminal display screen, provides both an "Try" button and a "Report Error" button. When a user clicks the "Try" button, the skill process can be simulated and the result verified. When the user clicks the "Report Error" button, the in-vehicle terminal can roll back the skill data and automatically trigger the teaching mode to guide the user to make corrections, thus forming a closed loop of "teaching -> verification -> correction" for voice recognition optimization.

[0074] Therefore, the aforementioned error reporting command can be a command triggered by the user through the "Report Error" button in the "Evolution Log" application when an error is found in the execution of a skill item, used to report a skill malfunction. Correspondingly, the aforementioned malfunctioning skill item is the skill item whose execution result does not meet expectations, corresponding to the user-triggered error reporting command (e.g., the command is associated with a vehicle function error). In this case, the aforementioned optimized teaching data can be corrected multimodal teaching data collected after multiple rounds of interactive dialogue for the malfunctioning skill item, used to correct the model parameters of the speech recognition model.

[0075] For example, when user A synchronizes the skill item "Air conditioning temperature adjustment and humidification" corresponding to the voice command "Set the air conditioning to 23 degrees but not too dry" to user B, if user B discovers during verification that the command is incorrectly associated with "seat ventilation" on their vehicle (due to differences in vehicle configuration), user B can click the "Report Error" button on the skill card. The in-vehicle terminal can immediately roll back the skill data and automatically re-trigger a multi-round interactive dialogue for the command "Set the air conditioning to 23 degrees but not too dry," guiding user B to make corrections. Then, user B can upload the corrected sample back to the cloud, thus completing an optimization iteration for a specific vehicle model.

[0076] Therefore, this embodiment can construct a closed-loop mechanism of "teaching-verification-correction" so that after the user triggers the error command, it can quickly respond and re-collect data to optimize the speech recognition model and correct the erroneous skill items in a timely manner; thereby achieving effective adaptation of technical items to the configuration differences of different car models, ensuring the accuracy and adaptability of model recognition, and continuously improving the user experience.

[0077] In summary, this embodiment provides users with transparent management and error correction capabilities for their vehicle's "learned skills" by constructing an "evolution log" and a "trusted verification closed loop"; and enables the secure and efficient flow and sharing of personalized knowledge within the user's social circle.

[0078] This embodiment discloses a method for generating skill items based on multimodal teaching data and storing these skill items in a time-series learning record database. The method performs a skill universality analysis on the skill items to obtain a universality index. It then performs a relationship closeness analysis on the skill items based on current user relationship data, including explicit and implicit user behavior data, to obtain a relationship closeness index. A sharing recommendation score is generated based on the universality index and the relationship closeness index, and target sharing users are determined according to a preset recommendation threshold and the sharing recommendation score. Sharing suggestions corresponding to the skill items are sent to the target sharing users. If the target sharing user confirms the sharing information based on the sharing suggestions, the skill items are sent to the learning record database of the vehicle equipment corresponding to the target sharing user. This embodiment enables the secure and efficient flow and sharing of personalized knowledge within a user's social circle.

[0079] This embodiment also discloses that upon receiving a user's error input command, the system determines the fault skill item corresponding to the error command; deletes the multimodal teaching data corresponding to the fault skill item from its own learning record database, and re-acquires the optimized teaching data corresponding to the fault skill item; and iteratively optimizes the speech recognition model based on the optimized teaching data. Therefore, this embodiment provides users with transparent management and error correction capabilities for their vehicle's "learned skills" by constructing an "evolution log" and a "trusted verification closed loop."

[0080] For example, to help understand the technical concept or principle of the speech recognition optimization method combined with Embodiments 1 and 2 above, please refer to Figure 5 , Figure 5 A simplified flowchart of the speech recognition optimization method of this application is provided below: 1) Speech recognition: The vehicle-mounted terminal continuously receives interactive commands input by the user, that is... Figure 5Voice commands are converted into text by the ASR (Automatic Speech Recognition) module, and then parsed and identified by the NLU (Natural Language Understanding) module to determine the intent category, key semantic slots, and slot values ​​of the commands, thus forming the intent parsing result.

[0081] Then, the confidence level of the intent parsing result output by NLU is evaluated. If the confidence level of the parsing result output by NLU is higher than the preset threshold, it is considered high confidence level, and the function corresponding to the voice command is executed directly. If the confidence level of the NLU output is lower than a preset threshold, or if the NLU module cannot map to any known intent at all (i.e., parsing fails), the learning optimization mechanism is triggered to optimize the multimodal data.

[0082] 2) Multimodal teaching data collection: First, attract the user's attention by outputting multimodal learning prompts, guide the user to respond to the instructions, and provide voice or touch methods for the user to choose whether to learn immediately; If the user chooses to learn immediately, they will enter the multimodal teaching module, which guides them through multi-turn interactive dialogue and multimodal sensor input of the semantic structure and multimodal information of the defined instructions, thereby collecting multimodal teaching data including text, images, gestures, etc. Then, the recognized instruction text, intent category, key semantic slot information, and the collected multimodal teaching data are timestamped and standardized to generate structured training samples, which are then encrypted and uploaded to the cloud. The cloud can use these training samples to incrementally train the speech recognition model, complete the iterative optimization of the model, and then synchronize the optimized model to the vehicle. If the user chooses to delay learning, the context information of the interaction command is saved, and the multimodal learning prompt is triggered again when the preset triggering conditions are met.

[0083] 3) Trusted verification: This application can display skill items arranged in a time sequence by setting an evolution log in the vehicle display screen. The evolution log can be equipped with a verification button for users to simulate the function of the skill item; and an error reporting button for users to report errors in the skill item, so as to roll back the multimodal teaching data corresponding to the faulty skill item.

[0084] Then, users are guided to collect optimized teaching data corresponding to the fault technology items, so as to perform iterative optimization of the model for specific vehicle types based on the optimized teaching data.

[0085] 4) Collaborative learning and knowledge synchronization: After verifying the effectiveness of the simulation, skill items can be shared across devices, enabling personalized storage and cross-user knowledge sharing.

[0086] Specifically, this application can analyze the skill universality and closeness of newly generated skill entries, determine target sharing users based on the generated universality index and closeness index, and push sharing suggestions to target sharing users. Newly generated skill entries can only be pushed to all authorized devices of the target user after the target user confirms skill sharing.

[0087] In summary, the vehicle-mounted large model self-learning method provided in this application can automatically identify user commands that fail to be parsed or have low confidence, actively guide users to participate in teaching, comprehensively utilize multimodal data to achieve incremental model learning, continuously improve speech understanding and execution capabilities, and at the same time ensure the accuracy and generalization of learning effects through verification closed loop and collaborative sharing mechanisms.

[0088] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the speech recognition optimization method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0089] This application also provides a speech recognition optimization device, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the module structure of the speech recognition optimization device according to an embodiment of this application. In this embodiment, the speech recognition optimization device includes: The instruction analysis module 601 is used to perform intent recognition on the interactive instructions input by the user through a speech recognition model and obtain intent parsing results; The learning trigger module 602 is used to evaluate the confidence level of the intent parsing result and generate learning prompt information when the evaluation result meets the preset learning trigger conditions. The multimodal data acquisition module 603 is used to acquire multimodal teaching data based on the learning prompt information. The model optimization module 604 is used to iteratively optimize the speech recognition model based on the multimodal teaching data.

[0090] Optionally, in this embodiment, the multimodal data acquisition module 603 is further configured to output the learning prompt information; receive response instructions from the user based on the output learning prompt information; and acquire multimodal teaching data according to the type of the response instructions, wherein the type of the response instructions includes immediate learning instructions or delayed learning instructions.

[0091] Optionally, in this embodiment, the multimodal data acquisition module 603 is further configured to: initiate a multi-turn interactive dialogue when the type of the response instruction is an immediate learning instruction; guide the user to confirm the instruction text, intent category, and key semantic slot information of the interactive instruction; acquire multimodal auxiliary information associated with the interactive instruction based on the multi-turn interactive dialogue; the multimodal auxiliary information includes at least one of interactive image data, gesture trajectory data, touch coordinate data, or environmental sensing data; and perform preset structured processing on the instruction text, intent category, key semantic slot information, and multimodal structure information to generate multimodal teaching data.

[0092] Optionally, in this embodiment, the multimodal data acquisition module 603 is further configured to: store the context information corresponding to the interaction instruction in a preset buffer when the type of the response instruction is a delayed learning instruction; acquire vehicle operating status parameters and driver / passenger detection signals; retrieve the context information from the preset buffer when the vehicle operating status parameters and the driver / passenger detection signals meet preset trigger conditions; and output the learning prompt information again based on the context information until the type of the response instruction corresponding to the learning prompt information is an immediate learning instruction.

[0093] Optionally, in this embodiment, the model optimization module 604 is further configured to generate skill items based on the multimodal teaching data, and store the skill items in its own learning record database according to the time sequence; perform a preset association analysis on the skill items to determine the target sharing user; send the sharing suggestion corresponding to the skill item to the target sharing user; and, if the target sharing user confirms the sharing information based on the sharing suggestion, send the skill item to the learning record database of the vehicle equipment corresponding to the target sharing user.

[0094] Optionally, in this embodiment, the model optimization module 604 is further configured to perform skill universality analysis on the skill items to obtain a universality index; perform closeness analysis on the skill items based on current user relationship data to obtain a relationship closeness index, wherein the current user relationship data includes explicit user data and implicit behavioral data; generate a shared recommendation score based on the universality index and the relationship closeness index; and determine the target shared user based on a preset recommendation threshold and the shared recommendation score.

[0095] Optionally, in this embodiment, the model optimization module 604 is further configured to, upon receiving a user input error command, determine the fault skill item corresponding to the error command; delete the multimodal teaching data corresponding to the fault skill item in its own learning record database, and re-acquire the optimized teaching data corresponding to the fault skill item; and iteratively optimize the speech recognition model based on the optimized teaching data.

[0096] The speech recognition optimization device provided in this application, employing the speech recognition optimization method in the above embodiments, can solve the technical problem of speech recognition optimization. Compared with the prior art, the beneficial effects of the speech recognition optimization device provided in this application are the same as those of the speech recognition optimization method provided in the above embodiments, and other technical features in the speech recognition optimization device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0097] This application provides a vehicle, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the speech recognition optimization method in Embodiment 1 above.

[0098] The following is for reference. Figure 7 It shows a structural diagram of a vehicle suitable for implementing the embodiments of this application, that is, a device structural diagram of the hardware operating environment involved in the speech recognition optimization method in the embodiments of this application. Figure 7 The vehicle shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0099] like Figure 7As shown, the vehicle may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for vehicle operation. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the vehicle to communicate wirelessly or wiredly with other devices to exchange data. Although vehicles with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0100] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment disclosed in this application includes a speech recognition optimization program product, which includes a speech recognition optimization program carried on a computer-readable medium, the speech recognition optimization program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the speech recognition optimization program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the speech recognition optimization program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0101] The vehicle provided in this application, employing the speech recognition optimization method described in the above embodiments, can solve the technical problem of speech recognition optimization. Compared with the prior art, the beneficial effects of the vehicle provided in this application are the same as those of the speech recognition optimization method provided in the above embodiments, and other technical features of the vehicle are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0102] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0103] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0104] This application provides a storage medium having computer-readable program instructions (i.e., a speech recognition optimization program) stored thereon, which are used to execute the speech recognition optimization method in the above embodiments.

[0105] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of the storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0106] The aforementioned storage medium may be included in the vehicle or may exist independently without being installed in the vehicle.

[0107] The aforementioned storage medium carries one or more programs, which, when executed by the vehicle, enable the vehicle to optimize voice recognition.

[0108] Speech recognition optimization program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and speech recognition optimization program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0110] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0111] The readable storage medium provided in this application is a storage medium that stores computer-readable program instructions (i.e., a speech recognition optimization program) for executing the above-described speech recognition optimization method, and is capable of solving the technical problem of speech recognition optimization. Compared with the prior art, the beneficial effects of the storage medium provided in this application are the same as the beneficial effects of the speech recognition optimization method provided in the above embodiments, and will not be repeated here.

[0112] The above are only some embodiments of this application and do not limit the scope of the solution of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.

Claims

1. A speech recognition optimization method, characterized in that, The method is applied to a vehicle, and the method includes: The intention of the user's interactive commands is identified by a speech recognition model, and the intention parsing result is obtained. The confidence level of the intent parsing result is evaluated, and a learning prompt message is generated when the evaluation result meets the preset learning triggering conditions. Acquire multimodal teaching data based on the learning prompt information; The speech recognition model is iteratively optimized based on the multimodal teaching data.

2. The speech recognition optimization method as described in claim 1, characterized in that, The step of acquiring multimodal teaching data based on the learning prompt information includes: Output the learning prompt information; Receive response instructions from the user based on the output learning prompts; Multimodal teaching data is obtained based on the type of the response instruction, which includes immediate learning instructions or delayed learning instructions.

3. The speech recognition optimization method as described in claim 2, characterized in that, The step of obtaining multimodal teaching data according to the type of the response instruction includes: When the type of the response instruction is an immediate learning instruction, a multi-turn interactive dialogue is initiated. The interactive dialogue is used to guide the user to confirm the instruction text, intent category, and key semantic slot information of the interactive instruction. Obtain multimodal auxiliary information associated with the interaction command collected based on the multi-round interactive session, wherein the multimodal auxiliary information includes at least one of interactive image data, gesture trajectory data, touch coordinate data, or environmental sensing data; The instruction text, the intent category, the key semantic slot information, and the multimodal structure information are subjected to preset structured processing to generate multimodal teaching data.

4. The speech recognition optimization method as described in claim 3, characterized in that, The step of obtaining multimodal teaching data according to the type of the response instruction further includes: When the type of the response instruction is a delayed learning instruction, the context information corresponding to the interaction instruction is stored in a preset cache. Acquire vehicle operating status parameters and occupant detection signals; When the vehicle operating status parameters and the occupant detection signal meet the preset trigger conditions, the context information is retrieved from the preset buffer. The learning prompt information is output again based on the context information until the type of the response instruction corresponding to the learning prompt information is an immediate learning instruction.

5. The speech recognition optimization method as described in claim 1, characterized in that, The step of iteratively optimizing the speech recognition model based on the multimodal teaching data includes: Skill items are generated based on the multimodal teaching data, and the skill items are stored in their own learning record database according to time sequence; Perform a pre-defined association analysis on the skill items to identify target sharing users; Send the sharing suggestions corresponding to the skill entries to the target sharing users; If the target sharing user confirms the sharing information based on the sharing suggestion, the skill item will be sent to the learning record database of the vehicle equipment corresponding to the target sharing user.

6. The speech recognition optimization method as described in claim 5, characterized in that, The step of performing a pre-defined association analysis on the skill items to determine the target sharing users includes: A skill universality analysis was performed on the skill items to obtain a universality index; The skill items are analyzed for closeness based on current user relationship data to obtain a relationship closeness index. The current user relationship data includes explicit user data and implicit behavioral data. A shared recommendation score is generated based on the generality index and the relationship tightness index, and target shared users are determined according to a preset recommendation threshold and the shared recommendation score.

7. The speech recognition optimization method as described in claim 5, characterized in that, After generating skill items based on the multimodal teaching data and storing the skill items in its own learning record database according to time series, the process further includes: Upon receiving a user's error reporting command, determine the fault skill item corresponding to the error reporting command; Delete the multimodal teaching data corresponding to the fault technology item in its own learning record database, and re-acquire the optimized teaching data corresponding to the fault skill item; The speech recognition model is iteratively optimized based on the optimized teaching data.

8. A speech recognition optimization device, characterized in that, The voice recognition optimization device is applied to a vehicle, and the device includes: The instruction analysis module is used to identify the intent of the user's interactive instructions through a speech recognition model and obtain the intent parsing results. The learning trigger module is used to evaluate the confidence level of the intent parsing result and generate learning prompt information when the evaluation result meets the preset learning trigger conditions. A multimodal data acquisition module is used to acquire multimodal teaching data based on the learning prompt information. The model optimization module is used to iteratively optimize the speech recognition model based on the multimodal teaching data.

9. A vehicle, characterized in that, The vehicle includes: a memory, a processor, and a speech recognition optimization program stored in the memory and executable on the processor, the speech recognition optimization program being configured to implement the steps of the speech recognition optimization method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a speech recognition optimization program, which, when executed by a processor, implements the steps of the speech recognition optimization method as described in any one of claims 1 to 7.