Voice interaction method, vehicle, and storage medium
By combining the model and rule engine in the vehicle for slot extraction and sentence segmentation, the problem of automatic execution of customized solutions for network abnormal download systems is solved, improving user experience and efficiency.
Patent Information
- Application Number
- CN202111629752.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-12-28
AI Technical Summary
When customizing private customization solutions during user driving, the existing technology has the risk of long data links and stability, especially in the case of poor network signals or no network, which leads to poor user experience.
The slot extraction is performed using a solution combining the first model and the second model in the vehicle, combining the distillation model and the rule engine to reduce the chip usage while ensuring the accuracy of the extraction and analysis process. Customized results are generated through slot extraction and sentence segmentation to realize the automatic execution of the vehicle-mounted system in the abnormal network state.
In the abnormal situation of network, users can trigger the on-board system to automatically complete the customization plan through voice requests, improving the efficiency of user experience and customization results generation, and providing timely feedback.
Smart Images

Figure CN114399994B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction technology, and in particular to a voice interaction method, a vehicle, and a storage medium. Background Art
[0002] Currently, when a user needs to complete a customized plan while driving, he or she must first make voice customization on the mobile phone or vehicle App, and then upload it to the vehicle through the server. The overall data link is long and there are risks to stability.
[0003] Furthermore, if users use their mobile phones to customize, they cannot easily customize while driving. In particular, when the network signal is poor or unavailable, the vehicle cannot respond to the user's voice commands, making it difficult to quickly complete personalized customization, resulting in a poor user experience. Summary of the Invention
[0004] The present invention provides a voice interaction method, a vehicle and a storage medium.
[0005] The present invention provides a voice interaction method. The voice interaction method includes: in a state of abnormal network connection, performing voice recognition on a user's voice request for voice operation of an in-vehicle system function customization using natural language in a vehicle to obtain a recognized text, wherein the customization enables the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met; performing slot extraction on the recognized text; performing sentence segmentation on the recognized text based on the results of the slot extraction; and generating a customization result based on the content of the sentence segmentation.
[0006] In this way, the voice interaction method of the present invention can enable the user to trigger the vehicle system to automatically complete the execution operation corresponding to the customized solution through a voice request when the network connection is abnormal.
[0007] The slot extraction of the recognized text includes: performing a first slot extraction process on the recognized text using a first model in the vehicle to obtain a first extraction result; performing a second slot extraction process on the recognized text using a second model in the vehicle to obtain a second extraction result; and fusing the first extraction result and the second extraction result to obtain a slot extraction word.
[0008] In this way, the voice interaction method of the present invention takes into account the computing power limitations of existing vehicle chips and adopts a solution that combines the first model and the second model to complete slot extraction. On the premise of reducing the utilization rate of the vehicle-side chip, it can ensure the accuracy of the subsequent private customized extraction and analysis process in the event of network anomalies.
[0009] The method includes: training a preset model using training data; and distilling the trained preset model to obtain the first model.
[0010] In this way, a first model can be obtained, and thus a first slot extraction process can be performed on the recognized text using the first model to obtain a first extraction result.
[0011] The method includes: comparing the prediction results of the trained preset model and the first model; and obtaining the second model according to the comparison results, the standard statement supported by the customization, and the generalized statement corresponding to the standard statement.
[0012] In this way, a second model can be obtained, and thus a second slot extraction process can be performed on the recognized text using the second model to obtain a second extraction result.
[0013] The slot extraction of the recognized text further includes: using the standard statement of the keyword and the generalized statement of the keyword to link the slot extraction word with the standard statement of the keyword.
[0014] In this way, by linking the slot extraction words to the standard expressions, the generated slot extraction words are more consistent with the standard expressions of the keywords in the standard customization scheme of the vehicle system function.
[0015] The sentence segmentation of the recognized text according to the result of the slot extraction includes: segmenting the recognized text according to the order of appearance of the keywords and the expression habits of the user to obtain trigger conditions and execution instructions.
[0016] In this way, the present invention can directly extract keyword information from the original recognized sentence, and use identifiers to divide trigger conditions and execution instructions, so as to directly extract personalized structured data, simplifying the pre-work of sentence division of traditional speech recognition solutions and improving the efficiency of customized result generation.
[0017] The generating of the customized result according to the content of the sentence segmentation includes: generating the customized result according to the trigger condition and the execution instruction.
[0018] In this way, the vehicle-mounted system can complete the corresponding execution operation according to the customization result.
[0019] The voice interaction method includes: feeding back the customization result to the user.
[0020] In this way, the interactive method of the present invention can trigger the vehicle to execute the customized plan according to voice in the event of network abnormality, and the user can obtain relevant feedback and know the customization results in time, thereby improving the user experience.
[0021] The present invention further provides a vehicle comprising a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the voice interaction method described in any one of the above embodiments is implemented.
[0022] In this way, the vehicle of the present invention can realize that the user can trigger the vehicle system to automatically complete the execution operation corresponding to the customized plan through voice request when the network connection is abnormal.
[0023] The present invention also provides a non-volatile computer-readable storage medium for a computer program. When the computer program is executed by one or more processors, the voice interaction method described in any one of the above embodiments is implemented.
[0024] In this way, the storage medium of the present invention can enable the user to trigger the vehicle-mounted system to automatically complete the execution operation corresponding to the customized solution through a voice request when the network connection is abnormal.
[0025] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments with reference to the following drawings, in which:
[0027] Figure 1 It is a flowchart of the voice interaction method of the present invention;
[0028] Figure 2 It is a structural diagram of the voice interaction device of the present invention;
[0029] Figure 3 It is a flowchart of the voice interaction method of the present invention;
[0030] Figure 4 It is a flowchart of the voice interaction method of the present invention;
[0031] Figure 5 It is a flowchart of the voice interaction method of the present invention;
[0032] Figure 6 It is a flowchart of the voice interaction method of the present invention;
[0033] Figure 7 It is a structural diagram of the voice interaction device of the present invention;
[0034] Figure 8 It is a flowchart of the voice interaction method of the present invention;
[0035] Figure 9 It is a structural diagram of the voice interaction device of the present invention;
[0036] Figure 10 It is a flowchart of the voice interaction method of the present invention;
[0037] Figure 11 It is a flowchart of the voice interaction method of the present invention;
[0038] Figure 12 It is a flowchart of the voice interaction method of the present invention;
[0039] Figure 13 It is a flowchart of the voice interaction method of the present invention;
[0040] Figure 14 It is a flowchart of the voice interaction method of the present invention;
[0041] Figure 15 It is a structural diagram of the voice interaction device of the present invention;
[0042] Figure 16 It is a schematic structural diagram of a vehicle of the present invention;
[0043] Figure 17 It is a schematic structural diagram of the computer-readable storage medium of the present invention. DETAILED DESCRIPTION
[0044] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present invention, and should not be understood as limiting the embodiments of the present invention.
[0045] See also Figure 1 The present invention provides a voice interaction method. The voice interaction method includes:
[0046] 01: When the network connection is abnormal, the user's voice request for customized in-vehicle system functions using natural language is recognized in the vehicle and the recognized text is obtained. The customization can enable the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met;
[0047] 03: Extract slots from the recognized text;
[0048] 05: Sentence segmentation of the recognized text based on the results of slot extraction;
[0049] 07: Generate customized results based on the content of sentence segmentation.
[0050] See also Figure 2 The present invention further provides a voice interaction device 10. The voice interaction device 10 includes: a voice recognition module 11, a slot extraction module 13, a sentence segmentation module 15 and a result generation module 17.
[0051] Step 01 can be implemented by the speech recognition module 11, step 03 can be implemented by the slot extraction module 13, step 05 can be implemented by the sentence segmentation module 15, and step 07 can be implemented by the result generation module 17. That is, the speech recognition module 11 is used to perform speech recognition on a user's voice request for voice operation of an in-vehicle system function customization using natural language in a vehicle when the network connection is abnormal, to obtain a recognized text. The customization can enable the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met; the slot extraction module 13 is used to extract slots from the recognized text; the sentence segmentation module 15 is used to perform sentence segmentation on the recognized text based on the results of the slot extraction; and the result generation module 17 is used to generate a customized result based on the content of the sentence segmentation.
[0052] Abnormal network connections include no network or poor network conditions. Users use natural language to complete voice operation to customize the functions of the in-vehicle system. Customization can enable the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met. Meeting the preset conditions means that the user's voice request conforms to the voice content and format in the customization plan of the in-vehicle system function. The voice content includes one or more trigger conditions and one or more trigger methods. In other words, users can trigger the in-vehicle system function to automatically execute the customization plan through voice requests, that is, automatically execute multiple functions of the preset conditions.
[0053] First, when the network connection is abnormal, voice requests for customized in-vehicle system functions, performed in the vehicle using natural language, are recognized and converted into text. That is, after the user wakes up the vehicle with a voice request, the vehicle can perform local automatic speech recognition (ASR) to convert the user's speech into text. For example, if the user's voice request is "Every day from 10:00 to 11:00, when I get in the car and say I'm going to work, all the windows are closed, music is playing, the heater is on, and I adjust my chair," the corresponding recognized text is "Every day from 10:00 to 11:00, when I get in the car and say I'm going to work, all the windows are closed, music is playing, the heater is on, and I adjust my chair."
[0054] Then, slot extraction is performed on the recognized text. This involves extracting the keywords from the user's customized voice request. Keywords contain information such as trigger conditions or execution instructions. For example, for the recognized text "Every day from 10:00 to 11:00, I get in the car and say I'm going to work. I close all the windows, play music, turn on the heater, and adjust the chair," slot extraction yields "Time 1: Every day, Time 2: 10:00, Time 3: 11:00, Action: Get in the car, Voice content: I'm going to work, Execution instructions: Close all the windows, play music, turn on the heater, and adjust the chair."
[0055] Next, the recognized text is segmented based on the slot extraction results. Sentence segmentation can be performed based on the order in which the extracted keywords appear. During the segmentation process, user expression habits, such as imperative and inverted sentences, can be taken into account to achieve sentence segmentation that suits user habits.
[0056] Finally, the customized results are generated based on the sentence segmentation content. The customized results can then be fed back to the user in the form of voice and card on the vehicle's large screen, allowing the customization process to be completed on the vehicle side even when there is no network or poor network connection.
[0057] In this way, the voice interaction method of the present invention can enable the user to trigger the vehicle system to automatically complete the execution operation corresponding to the customized solution through a voice request when the network connection is abnormal.
[0058] See also Figure 3 , step 03 includes:
[0059] 031: Using the first model in the vehicle to perform a first slot extraction process on the recognition text to obtain a first extraction result;
[0060] 032: Using the second model in the vehicle to perform a second slot extraction process on the recognized text to obtain a second extraction result;
[0061] 033: Fusion of the first extraction result and the second extraction result to obtain the slot extraction word.
[0062] Please combine Figure 2 Steps 031, 032, and 033 can be implemented by the slot extraction module 13. That is, the slot extraction module 13 is configured to perform a first slot extraction process on the recognized text using the first model in the vehicle to obtain a first extraction result; perform a second slot extraction process on the recognized text using the second model in the vehicle to obtain a second extraction result; and fuse the first extraction result and the second extraction result to obtain a slot extraction word.
[0063] The first model includes a distillation model. The second model includes a rule engine. The rule engine can be used to compensate for the overall impact of the decline in speech inference performance after model distillation, ensuring the accuracy of subsequent customized extraction and analysis processes.
[0064] First, the distillation model in the vehicle is used to extract the first slot of the recognized text to obtain the first extraction result, and then the rule engine in the vehicle is used to extract the second slot of the recognized text to obtain the second extraction result. For example, when the recognized text is "Every day from 10 to 11 o'clock, after I get in the car, I say I am going to work, close all the windows, play music, turn on the heater, and adjust the chair", the first extraction result and the second extraction result are as follows: Figure 4 shown.
[0065] Then, the first extraction result and the second extraction result are combined to obtain the slot extraction word, and the obtained slot extraction word is also as follows Figure 5 shown.
[0066] In this way, the voice interaction method of the present invention takes into account the computing power limitations of existing vehicle chips and adopts a solution that combines an offline rule engine and a distillation model to complete slot extraction. While reducing the utilization rate of the chip on the vehicle side, it can ensure the accuracy of the subsequent personalized extraction and analysis process in the event of network anomalies.
[0067] See also Figure 6 , methods include:
[0068] 02: Use training data to train the preset model;
[0069] 04: Distill the trained preset model to obtain the first model.
[0070] See also Figure 7 , the voice interaction device 10 includes a first model creation module 12.
[0071] Steps 02 and 04 can be implemented by the first model creation module 12. That is, the first model creation module 12 is used to train a preset model using training data; and distill the trained preset model to obtain the first model.
[0072] The creation process of the first model is:
[0073] First, the customized text data is manually labeled using the BOI method to mark out keyword tags such as trigger conditions and execution instructions.
[0074] Then, the default model is the standard BERT model. Using the standard BERT model for data training, we can obtain the upper limit of the model's capabilities. By training the pre-trained model, we can effectively utilize a large amount of expected semantic information and obtain sufficient generalization.
[0075] Finally, the standard BERT model is distilled using the public model distillation method to reach the carrying capacity of the vehicle-side chip and obtain the first model.
[0076] In this way, a first model can be obtained, and thus a first slot extraction process can be performed on the recognized text using the first model to obtain a first extraction result.
[0077] See also Figure 8 , methods include:
[0078] 06: Compare the prediction results of the trained preset model and the first model;
[0079] 08: Obtain the second model based on the comparison results, the customized supported standard statements and the generalized statements corresponding to the standard statements.
[0080] See also Figure 9 , the voice interaction device 10 includes a second model creation module 16.
[0081] Steps 06 and 08 can be implemented by the second model creation module 16. That is, the second model creation module 16 is used to compare the prediction results of the trained preset model and the first model; and obtain the second model based on the comparison results, the customized supported standard statements, and the generalized statements corresponding to the standard statements.
[0082] The creation process of the second model is: compare the prediction results of the trained preset model with those of the first model, and filter out the keyword labels and individual high-frequency words that are inaccurately predicted by the first model.
[0083] The second model is obtained by combining the inaccurately predicted keyword labels and individual high-frequency words filtered out by the first model, and using the standard statements supported by private customization and the generalized statements corresponding to the standard statements.
[0084] In this way, a second model can be obtained, and thus a second slot extraction process can be performed on the recognized text using the second model to obtain a second extraction result.
[0085] See also Figure 10 , step 03 also includes:
[0086] 034: Use the standard terms and generalized terms of keywords to link the slot extraction words with the standard terms of keywords.
[0087] Please combine Figure 2 Step 034 can be implemented by the slot extraction module 13. That is, the slot extraction module 13 is used to link the slot extraction word with the standard statement of the keyword by using the standard statement of the keyword and the generalized statement of the keyword.
[0088] Standard expressions and generalized expressions of keywords. For example, the standard expression corresponding to the generalized expression "turn off" of the keyword is "close", and the standard expression corresponding to the generalized expression "adjust it" is "adjust".
[0089] The process of linking the slot extraction words with the standard expressions of keywords is the process of aligning the slot extraction words with the standard expressions of keywords, and it is also the Figure 5 normalization process shown below. For example, the standard expression linked by the slot extraction word "when turning off all windows" is "close all windows". The standard expression linked by the slot extraction word "adjust the chair" is "adjust the seat".
[0090] In this way, by linking the slot extraction words to the standard expressions, the generated slot extraction words are more in line with the standard expressions of keywords in the standard customization scheme of the in-vehicle system functions.
[0091] Please refer to Figure 11 , step 05 includes:
[0092] 051: Split the recognized text into trigger conditions and execution instructions according to the appearance order of keywords and the user's expression habits.
[0093] Please refer to Figure 2 , step 051 can be implemented by the sentence splitting module 15. That is, the sentence splitting module 15 is used to split the recognized text into trigger conditions and execution instructions according to the appearance order of keywords and the user's expression habits.
[0094] After keyword extraction, use the identifying words for separating the components of the sentence. For example, the identifying words include "to", "after", "say", and "when", etc., and use the appearance order of keywords to perform sentence splitting. During the splitting process, consider the user's expression habits, including whether the user is accustomed to using imperative sentences or inverted sentences for expression, so as to obtain a sentence splitting result that conforms to the user's habits, and obtain trigger conditions and execution instructions, as Figure 12 shown below.
[0095] In this way, the present invention can directly extract keyword information from the recognized original sentence, use the identifying words to divide the trigger conditions and execution instructions, so as to directly extract the personalized structured data, simplify the pre-work of sentence division in the traditional speech recognition scheme, and improve the efficiency of generating the customization result.
[0096] Please refer to Figure 13 , step 07 includes:
[0097] 071: Generate a customization result according to the trigger conditions and execution instructions.
[0098] ]>Please refer to Figure 2Step 071 can be implemented by the result generation module 17, that is, the result generation module 17 is used to generate customized results according to the trigger conditions and execution instructions.
[0099] Finally, customized results are generated based on the trigger conditions and execution instructions. Figure 12 As shown, by recognizing the text "Every day from 10 to 11 o'clock, after I get in the car, I say I'm going to work, close all the windows, play music, turn on the heater, and adjust the chair", the trigger conditions include "time: every day, time: 10 o'clock, identification word: arrive, time: 11 o'clock, action: get in the car, identification word: say, voice content: when I'm going to work (voice condition)", and the execution instructions obtained include "execution instruction 1: close all windows, execution instruction 2: play music, execution instruction 3: turn on the air conditioner, temperature 26, execution instruction 4: adjust the seat".
[0100] In this way, the vehicle-mounted system can complete the corresponding execution operation according to the customization result.
[0101] See also Figure 14 , the voice interaction methods include:
[0102] 09: Feedback customization results to users.
[0103] See also Figure 15 , the voice interaction device 10 includes a feedback module 19 .
[0104] Step 09 can be implemented by the feedback module 19. That is, the feedback module 19 is used to feed back the customization result to the user.
[0105] The customized results are fed back to the user in the form of voice playback feedback and text display on the vehicle display screen. The feedback can also be in other ways, which are not limited here.
[0106] In this way, the interactive method of the present invention can trigger the vehicle to execute the customized plan according to voice in the event of network abnormality, and the user can obtain relevant feedback and know the customization results in time, thereby improving the user experience.
[0107] See also Figure 16 The present invention further provides a vehicle 20. The vehicle 20 includes a processor 21 and a memory 22. The memory 22 stores a computer program 221. When the computer program 221 is executed by the processor 21, the voice interaction method in any of the above embodiments is implemented.
[0108] In this way, the vehicle of the present invention can realize that the user can trigger the vehicle system to automatically complete the execution operation corresponding to the customized plan through voice request when the network connection is abnormal.
[0109] See also Figure 17The present invention further provides a non-volatile computer-readable storage medium 30 for a computer program. When the computer program 31 is executed by one or more processors 40, the voice interaction method of any of the above embodiments is implemented.
[0110] For example, when the computer program 31 is executed by the processor 40, the following steps of the voice interaction method are implemented:
[0111] 01: When the network connection is abnormal, the user's voice request for customized in-vehicle system functions using natural language is recognized in the vehicle and the recognized text is obtained. The customization can enable the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met;
[0112] 03: Extract slots from the recognized text;
[0113] 05: Sentence segmentation of the recognized text based on the results of slot extraction;
[0114] 07: Generate customized results based on the content of sentence segmentation.
[0115] It is understood that the computer program 31 includes computer program code. The computer program code may be in source code form, object code form, executable file, or some intermediate form. Computer-readable storage media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media.
[0116] In this way, the storage medium 30 of the present invention can enable the user to trigger the vehicle-mounted system to automatically complete the execution operation corresponding to the customized solution through a voice request when the network connection is abnormal.
Claims
1. A voice interaction method, characterized in that: include: In the event of an abnormal network connection, a user's voice request for customized in-vehicle system functions using natural language is recognized in the vehicle to obtain recognized text, wherein the customization enables the in-vehicle system to automatically execute multiple functions corresponding to the preset conditions when the preset conditions are met; Performing slot extraction on the recognized text; Segmenting the recognized text according to the result of the slot extraction; Generate customized results based on the content of the sentence segmentation; The extracting slots from the recognized text includes: performing a first slot extraction process on the recognized text using a first model in the vehicle to obtain a first extraction result, wherein the first model includes a distillation model; performing a second slot extraction process on the recognized text using a second model in the vehicle to obtain a second extraction result, wherein the second model includes a rule engine; The first extraction result and the second extraction result are combined to obtain a slot extraction word.
2. The voice interaction method according to claim 1, characterized in that: The method comprises: Use training data to train the preset model; The trained preset model is distilled to obtain the first model.
3. The voice interaction method according to claim 2, characterized in that: The method comprises: Comparing the prediction results of the trained preset model and the first model; The second model is obtained according to the comparison result, the standard statement supported by the customization, and the generalized statement corresponding to the standard statement.
4. The voice interaction method according to claim 1, wherein: The slot extraction of the recognized text further includes: The slot extraction word is linked to the standard statement of the keyword by using the standard statement of the keyword and the generalized statement of the keyword.
5. The voice interaction method according to claim 1, wherein: The sentence segmentation of the recognized text according to the result of the slot extraction includes: The recognition text is segmented into sentences according to the order of appearance of keywords and the expression habits of users to obtain trigger conditions and execution instructions.
6. The voice interaction method according to claim 5, characterized in that: Generating customized results according to the sentence segmentation content includes: The customized result is generated according to the trigger condition and the execution instruction.
7. The voice interaction method according to claim 1, wherein: The voice interaction method includes: Feedback the customization result to the user.
8. A vehicle, characterized in that: The vehicle includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the voice interaction method according to any one of claims 1 to 7 is implemented.
9. A non-volatile computer-readable storage medium containing a computer program, characterized in that: When the computer program is executed by one or more processors, the voice interaction method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Conference record generation method based on voice recognition, device and storage medium
CN110335612A
Voice control method, terminal and computer storage medium
CN110706705A
Text extraction model training and text extraction method and device
CN113204616A