Voice instruction recognition method, system and device based on intention prediction and medium
By predicting the intent of voice commands and matching voice databases, the problem of inaccurate recognition of existing vehicle voice recognition systems in driving scenarios is solved, and the recognition accuracy and user experience are improved.
Patent Information
- Application Number
- CN202510117350.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
The voice text recognized by the existing vehicle voice recognition system in driving scenarios may not be applicable, affecting the recognition accuracy and user experience.
By obtaining the user's real-time voice information and historical operation behavior data, combining the current vehicle status and environment information, predicting the voice command intent, and finding candidate voice commands in the preset voice command library to perform semantic and voice matching to obtain the target voice command.
It improves the accuracy of voice command recognition, enhances the user's driving experience, and avoids the problem of directly using the recognized voice text as voice commands.
Smart Images

Figure CN119943034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle control technology, and in particular to a method, system, device and medium for voice command recognition based on intention prediction. Background Art
[0002] At present, in order to improve the user's driving experience, most vehicles are equipped with voice recognition systems for recognizing user voice commands. However, the voice recognition models used by existing in-vehicle voice recognition systems are mostly trained based on conversation texts in general scenarios and their social general semantics, resulting in the recognized voice texts not being suitable for voice commands in driving scenarios, nor meeting the user's original operating intentions, affecting the accuracy of voice command recognition and the user's driving experience. Summary of the invention
[0003] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0004] To this end, an object of an embodiment of the present invention is to provide a voice command recognition method based on intention prediction, which improves the accuracy of voice command recognition and the user's driving experience.
[0005] Another object of an embodiment of the present invention is to provide a voice command recognition system based on intention prediction.
[0006] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0007] In a first aspect, an embodiment of the present invention provides a method for voice command recognition based on intention prediction, comprising the following steps:
[0008] Acquire real-time voice information of the user, and obtain a first voice text according to the real-time voice information recognition;
[0009] Acquire historical operation behavior data, historical voice text, current vehicle status and current environment information of the user in the previous period, and predict multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information and the current environment information;
[0010] According to the voice instruction intention, a plurality of corresponding candidate voice instructions are searched in a preset voice instruction library, and the first voice text is matched according to the candidate voice instructions to obtain a target voice instruction.
[0011] Further, in one embodiment of the present invention, the acquiring of the user's real-time voice information and obtaining the first voice text according to the real-time voice information recognition specifically includes:
[0012] Acquiring the real-time voice information through an audio acquisition device;
[0013] The real-time voice information is input into a preset voice recognition model to obtain the first voice text.
[0014] Further, in one embodiment of the present invention, the acquisition of the user's historical operation behavior data, historical voice text, current vehicle status and current environmental information in the previous period, and prediction of multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information and the current environmental information specifically includes:
[0015] Acquire the user's historical operation records in the previous period through the cockpit system, determine the historical operation behavior type and the corresponding operation time according to the historical operation records, and then construct the historical operation behavior data according to the historical operation behavior type and the operation time;
[0016] Acquire a second voice text obtained by recognizing the user's voice in the previous period, and construct the historical voice text according to the second voice text and the corresponding voice acquisition time;
[0017] Acquire the current state of each operable component of the vehicle through the vehicle body controller, and construct the current vehicle state information according to each operable component and the corresponding current state;
[0018] Acquire the in-vehicle environment information and the out-vehicle environment information respectively through sensors arranged in the vehicle and outside the vehicle, and construct the current environment information according to the in-vehicle environment information and the out-vehicle environment information;
[0019] The historical operation behavior data, the historical voice text, the current vehicle state information and the current environment information are input into a pre-trained voice command intention prediction model to obtain the voice command intention.
[0020] Furthermore, in one embodiment of the present invention, the voice instruction intention prediction model is trained by the following steps:
[0021] Obtaining the operating behavior sample data and the voice and text sample data of the test personnel in the first period and the vehicle state sample data and the environment sample data of the test vehicle in the second period, and generating training sample data according to the operating behavior sample data, the voice and text sample data, the vehicle state sample data and the environment sample data;
[0022] Acquire voice command sample data of the tester in the third period, and determine the voice command intention label of the training sample data according to the voice command sample data;
[0023] Inputting the training sample data into a pre-built deep learning neural network to obtain a voice command intention prediction result;
[0024] A loss value is determined according to the voice command intention prediction result and the voice command intention label, and the parameters of the deep learning neural network are updated according to the loss value to obtain a trained voice command intention prediction model.
[0025] Further, in one embodiment of the present invention, the step of searching for a plurality of corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, matching the first voice text according to the candidate voice instructions, and obtaining a target voice instruction specifically includes:
[0026] Determining a candidate instruction type according to the voice instruction intention;
[0027] Using the candidate instruction type as a key field to perform traversal query in the voice instruction library to obtain the corresponding candidate voice instruction;
[0028] The semantic similarity between each of the candidate voice instructions and the first voice text is calculated, and when there is a semantic similarity greater than or equal to a preset first threshold, the candidate voice instruction with the highest semantic similarity is determined as the target voice instruction.
[0029] Furthermore, in one embodiment of the present invention, the voice instruction recognition method further comprises the following steps:
[0030] When the semantic similarities are all less than the first threshold, the voice similarities between each of the candidate voice instructions and the first voice text are calculated, and the candidate voice instruction with the highest voice similarity is determined as the target voice instruction.
[0031] Furthermore, in one embodiment of the present invention, the voice instruction recognition method further includes the step of constructing and updating the voice instruction library, which specifically includes:
[0032] Acquire a software operation vocabulary built into the cockpit system, and generate a first voice command according to the software operation vocabulary;
[0033] Obtaining a vehicle user manual, performing word segmentation processing on the vehicle user manual to obtain a plurality of operation objects and a plurality of corresponding operation types, and constructing a second voice instruction according to the operation objects and the operation types;
[0034] Obtaining the custom password data uploaded by the user, and generating a third voice command according to the custom password data;
[0035] Building the voice instruction library according to the first voice instruction, the second voice instruction and the third voice instruction;
[0036] The number of successful matches of each voice instruction in the voice instruction library in a historical period is obtained, the matching priority of each voice instruction is updated according to the number of successful matches, and the voice instructions that have not been successfully matched are deleted.
[0037] In a second aspect, an embodiment of the present invention provides a voice command recognition system based on intention prediction, comprising:
[0038] A speech recognition module, used to obtain real-time speech information of a user, and obtain a first speech text according to the real-time speech information recognition;
[0039] A voice command intention prediction module is used to obtain the user's historical operation behavior data, historical voice text, current vehicle status and current environmental information in the previous period, and predict multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information and the current environmental information;
[0040] The voice instruction matching module is used to search for a plurality of corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, and match the first voice text according to the candidate voice instructions to obtain a target voice instruction.
[0041] In a third aspect, an embodiment of the present invention provides a voice command recognition device based on intention prediction, comprising:
[0042] at least one processor;
[0043] at least one memory for storing at least one program;
[0044] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned voice instruction recognition method based on intention prediction.
[0045] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to execute the above-mentioned method of voice command recognition based on intent prediction when executed by the processor.
[0046] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:
[0047] The embodiment of the present invention obtains the real-time voice information of the user, obtains the first voice text according to the real-time voice information recognition, obtains the historical operation behavior data, historical voice text, current vehicle status and current environmental information of the user in the previous period, predicts multiple voice command intentions according to the historical operation behavior data, historical voice text, current vehicle status information and current environmental information, searches for multiple corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, matches the first voice text according to the candidate voice instructions, and obtains the target voice instruction. After recognizing the voice text of the user, the embodiment of the present invention predicts multiple voice command intentions according to the historical operation behavior data, historical voice text, current vehicle status information and current environmental information, searches for multiple candidate voice instructions in a preset voice instruction library based on the voice instruction intention, so that the candidate voice instructions and the recognized voice text can be matched based on semantics and voice to obtain the target voice instruction, avoiding the problem of directly using the recognized voice text as a voice instruction, which is not suitable for driving scenarios and does not meet the original operation intention of the user, and improves the accuracy of voice instruction recognition and the user's driving experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 A flowchart of a method for voice command recognition based on intention prediction provided by an embodiment of the present invention;
[0050] Figure 2 A structural block diagram of a voice command recognition system based on intention prediction provided by an embodiment of the present invention;
[0051] Figure 3 A structural block diagram of a voice command recognition device based on intention prediction provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0053] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.
[0054] Reference Figure 1 , an embodiment of the present invention provides a method for voice command recognition based on intention prediction, which specifically includes the following steps:
[0055] S101, obtaining real-time voice information of a user, and obtaining a first voice text according to the real-time voice information recognition;
[0056] S102, obtaining historical operation behavior data, historical voice text, current vehicle status and current environment information of the user in the previous period, and predicting multiple voice command intentions based on the historical operation behavior data, historical voice text, current vehicle status information and current environment information;
[0057] S103: searching for a plurality of corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, matching the first voice text according to the candidate voice instructions, and obtaining a target voice instruction.
[0058] After recognizing the user's voice text, the embodiment of the present invention predicts multiple voice command intentions based on historical operation behavior data, historical voice texts, current vehicle status information and current environmental information, and searches for multiple candidate voice commands in a preset voice command library based on the voice command intentions, so that the candidate voice commands and the recognized voice text can be matched based on semantics and voice to obtain the target voice command, avoiding the problem of directly using the recognized voice text as a voice command that is not suitable for driving scenarios and does not meet the user's original operation intentions, thereby improving the accuracy of voice command recognition and the user's driving experience.
[0059] As an optional implementation, real-time voice information of the user is obtained, and a first voice text is obtained according to the real-time voice information recognition, which specifically includes:
[0060] S1011, obtaining real-time voice information through an audio collection device;
[0061] S1012: Input the real-time voice information into a preset voice recognition model to obtain a first voice text.
[0062] Specifically, the user's real-time voice information is obtained through an audio collection device set in the cockpit, and the real-time voice information is input into a preset voice recognition model to obtain the corresponding first voice text. It should be noted that the voice recognition model used here can be a general voice recognition model disclosed in the prior art. The embodiment of the present invention does not retrain the voice recognition model, but predicts the user's voice command intention in subsequent steps, and performs targeted matching on the recognized voice text in combination with the predicted voice command intention, thereby achieving fine-tuning of the voice command recognition result.
[0063] As an optional implementation, historical operation behavior data, historical voice text, current vehicle status and current environment information of the user in the previous period are obtained, and multiple voice command intentions are predicted based on the historical operation behavior data, historical voice text, current vehicle status information and current environment information, which specifically include:
[0064] S1021. Obtaining the user's historical operation records in the previous period through the cockpit system, determining the historical operation behavior type and the corresponding operation time according to the historical operation records, and then constructing the historical operation behavior data according to the historical operation behavior type and the operation time;
[0065] S1022, obtaining a second voice text obtained by recognizing the user's voice in the previous period, and constructing a historical voice text according to the second voice text and the corresponding voice acquisition time;
[0066] S1023, obtaining the current state of each operable component of the vehicle through the vehicle body controller, and constructing current vehicle state information according to each operable component and the corresponding current state;
[0067] S1024, obtaining in-vehicle environment information and out-vehicle environment information respectively through sensors disposed in the vehicle and outside the vehicle, and constructing current environment information according to the in-vehicle environment information and out-vehicle environment information;
[0068] S1025. Input historical operation behavior data, historical voice text, current vehicle status information, and current environment information into a pre-trained voice command intention prediction model to obtain the voice command intention.
[0069] Specifically, data points are embedded in the vehicle controller to record the user's operation behavior data. The data format is {operation, time}, such as {navigation home, 2024-12-01 12:12}, based on the historical operation records of the previous period, determine the historical operation behavior type and the corresponding operation time, so as to obtain the historical operation behavior data; obtain multiple voice texts recognized in the previous period, sort the voice texts in combination with the voice acquisition time, and obtain the historical voice texts; obtain the vehicle status data through the body controller, and the status data is defined as {component name, functional status}, such as {window, open}; obtain the in-vehicle environment information and the out-vehicle environment information through the temperature sensors and light sensors set inside and outside the vehicle, so as to obtain the current environment information; input the obtained historical operation behavior data, historical voice texts, current vehicle status information and current environment information into the pre-trained voice command intention prediction model, and the voice command intention can be predicted. For example, when the owner goes to work at a fixed time every morning, he will turn on the air conditioner after starting the music player. Then, within the corresponding time, if the system has executed the operation of starting the music player, the current air conditioner status is not turned on, the current temperature in the car is 30 degrees, and the historical voice text contains the instruction to start the music player, the voice command intention to turn on the air conditioner will be predicted with the greatest probability.
[0070] As an optional implementation, the voice command intention prediction model is trained by the following steps:
[0071] S201, obtaining the operating behavior sample data and voice text sample data of the test personnel in the first period and the vehicle state sample data and environment sample data of the test vehicle in the second period, and generating training sample data according to the operating behavior sample data, voice text sample data, vehicle state sample data and environment sample data;
[0072] S202, obtaining voice command sample data of the tester in the third period, and determining the voice command intention label of the training sample data according to the voice command sample data;
[0073] S203, inputting the training sample data into a pre-built deep learning neural network to obtain a voice command intention prediction result;
[0074] S204. Determine a loss value based on the voice command intention prediction result and the voice command intention label, and update the parameters of the deep learning neural network based on the loss value to obtain a trained voice command intention prediction model.
[0075] Specifically, after inputting the training sample data into the initialized deep learning neural network, the prediction result of the model output, i.e., the prediction result of the voice command intention, can be obtained. The accuracy of the model prediction can be evaluated based on the voice command intention prediction result and the aforementioned voice command intention label, thereby updating the parameters of the model. For the voice command intention prediction model, the accuracy of the model prediction result can be measured by the loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of a single training data and the prediction result of the model for the training data. In actual training, a training data set has a lot of training data, so the cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the model. For general machine learning models, based on the aforementioned cost function, plus the regularization term that measures the complexity of the model, it can be used as the objective function of the training. Based on this objective function, the loss value of the entire training data set can be calculated. There are many types of commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross entropy loss function, etc., which can all be used as loss functions of machine learning models, which will not be elaborated here one by one. In an embodiment of the present invention, any one of the loss functions can be selected to determine the loss value of training. Based on the loss value of the training, the back propagation algorithm is used to update the parameters of the model, and a trained voice command intention prediction model can be obtained after several rounds of iteration. The specific number of iterations can be set in advance, or the training is considered to be completed when the test set meets the accuracy requirements.
[0076] As an optional implementation, according to the voice instruction intention, a plurality of corresponding candidate voice instructions are searched in a preset voice instruction library, and the first voice text is matched according to the candidate voice instructions to obtain the target voice instruction, which specifically includes:
[0077] S1031, determining a candidate instruction type according to the voice instruction intention;
[0078] S1032, using the candidate instruction type as a key field to perform a traversal query in the voice instruction library to obtain a corresponding candidate voice instruction;
[0079] S1033, calculating the semantic similarity between each candidate voice instruction and the first voice text, and when there is a semantic similarity greater than or equal to a preset first threshold, determining the candidate voice instruction with the highest semantic similarity as the target voice instruction.
[0080] Specifically, the candidate command type is determined according to the recognized voice command intention. For example, if the recognized voice command intention is to turn on the air conditioner, the corresponding candidate command type is the in-vehicle environment adjustment command. The candidate command type is then used as a key field to traverse and query in the preset voice command library to obtain multiple corresponding candidate voice commands, such as turning on the air conditioner, starting the fragrance system, starting the in-vehicle ventilation function, etc. After obtaining the candidate voice commands, the semantic similarity between each candidate voice command and the first voice text recognized previously is determined. If there is at least one candidate voice command whose semantic similarity with the first voice text is greater than or equal to a preset first threshold (such as 0.8), the candidate voice command with the highest semantic similarity is used as the target voice command.
[0081] As an optional implementation, the voice command recognition method further includes the following steps:
[0082] S1034: When the semantic similarities are all less than the first threshold, calculate the voice similarity between each candidate voice instruction and the first voice text, and determine the candidate voice instruction with the highest voice similarity as the target voice instruction.
[0083] Specifically, when the semantic similarity between all candidate voice commands and the first voice text is less than a first threshold, it is possible that the vocabulary in the driving scenario is recognized as a homonym in the general scenario during voice recognition. Therefore, the voice similarity between each candidate voice command and the first voice text can be calculated, and the candidate voice command with the highest voice similarity can be used as the target voice command.
[0084] As an optional implementation, the voice command recognition method further includes the step of constructing and updating a voice command library, which specifically includes:
[0085] S301, obtaining a software operation vocabulary built into the cockpit system, and generating a first voice command according to the software operation vocabulary;
[0086] S302, obtaining a vehicle user manual, performing word segmentation processing on the vehicle user manual to obtain a plurality of operation objects and a plurality of corresponding operation types, and constructing a second voice instruction according to the operation objects and the operation types;
[0087] S303, obtaining the user-defined password data uploaded by the user, and generating a third voice command according to the user-defined password data;
[0088] S304, constructing a voice instruction library according to the first voice instruction, the second voice instruction and the third voice instruction;
[0089] S305: Obtain the number of successful matches of each voice command in the voice command library within the historical period, update the matching priority of each voice command according to the number of successful matches, and delete the voice commands that have not been successfully matched.
[0090] Specifically, a first voice command is generated according to the operation vocabulary on the UI of the cockpit software; according to the vehicle user manual, a plurality of operation objects and corresponding operation types are obtained by using NLP technology for word segmentation, and a second voice command is constructed according to the operation object and the operation type; the user can upload a customized document through the mobile phone APP, extract the words in the document by using the OFFICE control, and then use the NLP model to segment the document to generate a third voice command of a customized password; a voice command library is constructed according to the first voice command, the second voice command and the third voice command. During use, the number of successful matches of each voice command is recorded. When the user expresses the demand for the last time, the owner does not reply again after the vehicle responds, and the user does not manually perform vehicle control within the specified time, it is considered to be a successfully matched voice command, otherwise it is considered to be unsuccessful. The matching priority of each voice command is updated according to the number of successful matches. The matching priority can be used for the order of traversing the query voice command library, and can also be used as a coefficient to fine-tune the semantic similarity and voice similarity between the candidate voice command and the first voice text, and delete the voice command that has not been successfully matched.
[0091] The above is a description of the method steps of the embodiment of the present invention. It can be understood that after the embodiment of the present invention recognizes the user's voice text, it predicts multiple voice command intentions based on historical operation behavior data, historical voice text, current vehicle status information and current environmental information, and searches for multiple candidate voice commands in a preset voice command library based on the voice command intention, so that the candidate voice commands and the recognized voice text can be matched based on semantics and voice to obtain the target voice command, avoiding the problem of directly using the recognized voice text as a voice command, which is not suitable for driving scenarios and does not meet the user's original operation intention, and improving the accuracy of voice command recognition and the user's driving experience.
[0092] Reference Figure 2 , an embodiment of the present invention provides a voice command recognition system based on intention prediction, comprising:
[0093] A speech recognition module, used to obtain real-time speech information of a user, and obtain a first speech text according to the real-time speech information recognition;
[0094] The voice command intention prediction module is used to obtain the user's historical operation behavior data, historical voice text, current vehicle status and current environmental information in the previous period, and predict multiple voice command intentions based on the historical operation behavior data, historical voice text, current vehicle status information and current environmental information;
[0095] The voice instruction matching module is used to search for multiple corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, match the first voice text according to the candidate voice instructions, and obtain the target voice instruction.
[0096] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0097] Reference Figure 3 , an embodiment of the present invention provides a voice command recognition device based on intention prediction, comprising:
[0098] at least one processor;
[0099] at least one memory for storing at least one program;
[0100] When the at least one program is executed by the at least one processor, the at least one processor implements the voice command recognition method based on intention prediction.
[0101] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0102] An embodiment of the present invention also provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned voice command recognition method based on intention prediction.
[0103] A computer-readable storage medium according to an embodiment of the present invention can execute a voice command recognition method based on intent prediction provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0104] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.
[0105] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0106] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0107] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0108] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0109] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.
[0110] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0111] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0112] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0113] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for voice command recognition based on intention prediction, characterized in that: The following steps are involved: Acquire real-time voice information of the user, and obtain a first voice text according to the real-time voice information recognition; Acquire historical operation behavior data, historical voice text, current vehicle status and current environment information of the user in the previous period, and predict multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information and the current environment information; According to the voice instruction intention, a plurality of corresponding candidate voice instructions are searched in a preset voice instruction library, and the first voice text is matched according to the candidate voice instructions to obtain a target voice instruction.
2. The method for voice command recognition based on intention prediction according to claim 1, characterized in that: The step of acquiring the real-time voice information of the user and obtaining the first voice text according to the real-time voice information recognition specifically includes: Acquiring the real-time voice information through an audio acquisition device; The real-time voice information is input into a preset voice recognition model to obtain the first voice text.
3. The method for voice command recognition based on intention prediction according to claim 1, characterized in that: The acquiring of the user's historical operation behavior data, historical voice text, current vehicle status, and current environment information in the previous period, and predicting multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information, and the current environment information, specifically includes: Acquire the user's historical operation records in the previous period through the cockpit system, determine the historical operation behavior type and the corresponding operation time according to the historical operation records, and then construct the historical operation behavior data according to the historical operation behavior type and the operation time; Acquire a second voice text obtained by recognizing the user's voice in the previous period, and construct the historical voice text according to the second voice text and the corresponding voice acquisition time; Acquire the current state of each operable component of the vehicle through the vehicle body controller, and construct the current vehicle state information according to each operable component and the corresponding current state; Acquire the in-vehicle environment information and the out-vehicle environment information respectively through sensors arranged in the vehicle and outside the vehicle, and construct the current environment information according to the in-vehicle environment information and the out-vehicle environment information; The historical operation behavior data, the historical voice text, the current vehicle state information and the current environment information are input into a pre-trained voice command intention prediction model to obtain the voice command intention.
4. The method for voice command recognition based on intention prediction according to claim 3, characterized in that: The voice command intention prediction model is trained by the following steps: Obtaining the operating behavior sample data and the voice and text sample data of the test personnel in the first period and the vehicle state sample data and the environment sample data of the test vehicle in the second period, and generating training sample data according to the operating behavior sample data, the voice and text sample data, the vehicle state sample data and the environment sample data; Acquire voice command sample data of the tester in the third period, and determine the voice command intention label of the training sample data according to the voice command sample data; Inputting the training sample data into a pre-built deep learning neural network to obtain a voice command intention prediction result; A loss value is determined according to the voice command intention prediction result and the voice command intention label, and the parameters of the deep learning neural network are updated according to the loss value to obtain a trained voice command intention prediction model.
5. The method for voice command recognition based on intention prediction according to claim 1, characterized in that: The step of searching for a plurality of corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, and matching the first voice text according to the candidate voice instructions to obtain a target voice instruction specifically includes: Determining a candidate instruction type according to the voice instruction intention; Using the candidate instruction type as a key field to perform traversal query in the voice instruction library to obtain the corresponding candidate voice instruction; The semantic similarity between each of the candidate voice instructions and the first voice text is calculated, and when there is a semantic similarity greater than or equal to a preset first threshold, the candidate voice instruction with the highest semantic similarity is determined as the target voice instruction.
6. The method for voice command recognition based on intention prediction according to claim 5, characterized in that: The voice command recognition method further comprises the following steps: When the semantic similarities are all less than the first threshold, the voice similarities between each of the candidate voice instructions and the first voice text are calculated, and the candidate voice instruction with the highest voice similarity is determined as the target voice instruction.
7. The method for voice command recognition based on intention prediction according to claim 5, characterized in that: The voice instruction recognition method further includes the step of constructing and updating the voice instruction library, which specifically includes: Acquire a software operation vocabulary built into the cockpit system, and generate a first voice command according to the software operation vocabulary; Obtaining a vehicle user manual, performing word segmentation processing on the vehicle user manual to obtain a plurality of operation objects and a plurality of corresponding operation types, and constructing a second voice instruction according to the operation objects and the operation types; Obtaining the custom password data uploaded by the user, and generating a third voice command according to the custom password data; Building the voice instruction library according to the first voice instruction, the second voice instruction and the third voice instruction; The number of successful matches of each voice instruction in the voice instruction library in a historical period is obtained, the matching priority of each voice instruction is updated according to the number of successful matches, and the voice instructions that have not been successfully matched are deleted.
8. A voice command recognition system based on intention prediction, characterized in that: include: A speech recognition module, used to obtain real-time speech information of a user, and obtain a first speech text according to the real-time speech information recognition; A voice command intention prediction module is used to obtain the user's historical operation behavior data, historical voice text, current vehicle status and current environmental information in the previous period, and predict multiple voice command intentions based on the historical operation behavior data, the historical voice text, the current vehicle status information and the current environmental information; The voice instruction matching module is used to search for a plurality of corresponding candidate voice instructions in a preset voice instruction library according to the voice instruction intention, and match the first voice text according to the candidate voice instructions to obtain a target voice instruction.
9. A voice command recognition device based on intention prediction, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a voice command recognition method based on intention prediction as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to execute a voice command recognition method based on intention prediction as described in any one of claims 1 to 7 when executed by the processor.
Citation Information
Cited By
Searching method and system for vehicle-mounted system and storage medium
CN120196802A
Voice interaction method and system based on 3D virtualization
CN120388565A
A voice interaction method and system based on 3D virtualization
CN120388565B
Voice interaction method and apparatus, and XR device
CN120544579A
Voice interruption processing method and device, electronic equipment and storage medium
CN120748401A