Voice interaction method, device, server, and readable storage medium

By rewritten and recognized user intentions in multiple rounds of voice requests in the voice interaction system of smart cars, the problem of difficulty in identifying users' continuous adjustment intentions is solved, and accurate scale adjustment and improved user experience is achieved.

CN114299931BActive Publication Date: 2025-05-30GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111574477.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-05-30
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

In the voice interaction system of smart cars, it is difficult for users to accurately identify the user's intentions when performing multiple rounds of voice interaction, especially in scenarios where the user wants to make continuous adjustments, the system cannot accurately recognize the current round of voice requests.

Method used

By receiving voice requests forwarded by the vehicle, the previous round of voice requests is read and querying in the cache engine. If the missed cache is used, the voice request of the current round is rewritten by using the voice request of the previous round, and intent recognition is performed to complete the voice interaction.

Benefits of technology

It realizes that in multiple rounds of voice request scenarios, accurately identify users' intentions, meet users' precise scale adjustment needs for vehicle components, and improves the accuracy and user experience of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299931B_ABST
    Figure CN114299931B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice interaction method, its device, server, and readable storage medium. The voice interaction method includes: receiving a current round of voice requests for adjusting preset functions of a vehicle forwarded by the vehicle, where the preset functions refer to functions for simulating the operation of vehicle components for scale adjustment; reading the previous round of voice requests for adjusting the preset functions of the vehicle; performing a cache query in a cache engine based on the current round of voice requests and the previous round of voice requests; in the case where the result of the cache query fails to find the corresponding cache, using the previous round of voice requests to rewrite the current round of voice requests; performing intent recognition on the rewritten current round of voice requests; and completing voice interaction according to the result of the intent recognition. The present invention combines two rounds of voice requests and uses a combination of a high-frequency cache engine and intent recognition to identify the intent of the voice requests, achieving accurate identification of user intent under multiple rounds of voice requests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice technology, and particularly relates to a voice interaction method, an apparatus, a server, and a readable storage medium therefor. Background Art

[0002] Currently, in the scenario of intelligent vehicles, voice interaction can be applied to control vehicle components by users, such as "open the window", "turn up the volume", etc. However, for scenarios where users hope to make continuous adjustments, in the voice scenario, it is manifested as multi-round interaction. After the previous round of voice interaction, users naturally omit some content in each subsequent round of conversation. For example, in the following conversation between the user and the voice assistant Xiaop:[[]]

[0003] User: What's the weather like today?

[0004] Xiaop: It's sunny in Guangzhou today, 26-30°.

[0005] User: How about Shanghai?

[0006] In multi-round conversations, like in the first example above, the literal meaning of what the user asks is about Shanghai, but actually the user wants to know the weather in Shanghai. Omitting some content conforms to the habit of human conversations. However, this may cause the vehicle's in-vehicle system to not accurately recognize certain rounds of voice requests, or prompt that it doesn't understand.

[0007] Furthermore, if the user needs to adjust the volume, the user can operate the mechanical knob for adjusting the car volume on the vehicle to rotate the mechanical knob to the desired volume. However, if using voice to adjust the volume, it can only be turned up or down. In the following second example:[[]]

[0008] User: Turn up the volume

[0009] Xiaop: The volume has been turned up

[0010] User: Louder, louder, louder

[0011] As can be seen from the second example, the current vehicle's in-vehicle system cannot accurately recognize "louder, louder, louder" in the current round, or prompts that it doesn't understand. Such a situation cannot meet the user's demand for precise scale continuous adjustment like a mechanical knob. Summary of the Invention

[0012] Embodiments of the present invention provide a voice interaction method, an apparatus, a server, and a readable storage medium therefor.

[0013] An embodiment of the present invention provides a voice interaction method. The voice interaction method includes: receiving a current round of voice requests for adjusting preset functions of a vehicle forwarded by the vehicle, where the preset functions refer to functions for simulating scale adjustment of operations on vehicle components; reading the previous round of voice requests for adjusting the preset functions of the vehicle; performing a cache query in a cache engine according to the current round of voice requests and the previous round of voice requests; in the case where the result of the cache query fails to find a corresponding cache, using the previous round of voice requests to rewrite the current round of voice requests; performing intent recognition on the rewritten current round of voice requests; and completing voice interaction according to the result of the intent recognition.

[0014] In this way, after receiving a voice request from the user for scale adjustment of vehicle components, the voice interaction method of the present invention can read the previous round of voice requests, combine the two rounds of voice requests to query whether the cache is hit, and in the case where the corresponding cache cannot be found, use the previous round of voice requests to rewrite the current round of voice requests, so that the rewritten voice request can be recognized by the in-vehicle system of the vehicle as the corresponding intent, and then realize scale adjustment of vehicle components in a voice interaction manner according to the result of the intent recognition. The intention of the voice request is recognized by combining a high-frequency cache engine and intent recognition, and the accurate recognition of the user's intention is realized under multiple rounds of voice requests.

[0015] The voice interaction method includes: adding adjacent two rounds of voice requests with an occurrence frequency greater than a preset frequency to the cache engine.

[0016] In this way, the cache of the cache engine of the present invention is composed of adjacent two rounds of voice requests with an occurrence frequency greater than the preset frequency, realizing the statistics of high-frequency set voice requests.

[0017] The voice interaction method includes: establishing a mapping relationship between the current round of voice requests and preset intents.

[0018] In this way, after establishing the mapping relationship between the current round of voice requests and preset intents in the present invention, each preset intent is associated with the corresponding adjacent two rounds of voice requests, so as to determine the intent corresponding to the voice request in the cache engine query.

[0019] The voice interaction method includes: in the case where the result of the cache query finds a corresponding cache, determining the preset intent corresponding to the current round of voice requests as the target intent according to the mapping relationship to complete voice interaction.

[0020] Thus, in the present invention, the voice request of the current round and the voice request of the adjacent previous round are queried in the cache engine. According to the established mapping relationship, the target intention of the voice request of the current round can be directly determined, so that the voice interaction can be completed according to the determined target intention corresponding to the voice request of the current round.

[0021] The rewriting of the voice request of the current round by using the voice request of the previous round includes: training a rewriting model through training data for rewriting, where the training data for rewriting includes voice requests of two adjacent rounds; and rewriting the voice request of the current round by using the voice request of the previous round and the rewriting model.

[0022] Thus, in the present invention, through machine learning, a rewriting model is trained from voice requests of two adjacent rounds, so that the voice request of the current round can be rewritten according to the voice request of the previous round and the rewriting model, and the rewritten voice request can be recognized by the vehicle-mounted system of the vehicle for the corresponding intention.

[0023] The intention recognition of the rewritten voice request of the current round includes: training an intention recognition model through intention training data, where the intention training data is related to vehicle components that can be adjusted in scale and the scale adjustment range of the vehicle components; and recognizing the intention of the rewritten voice request of the current round by using the intention recognition model.

[0024] Thus, in the present invention, through machine learning, an intention recognition model is trained from the training data corresponding to vehicle components that can be adjusted in scale and the scale adjustment range of the vehicle components, and then the intention of the rewritten voice request is recognized to achieve accurate recognition of the user's intention.

[0025] The completion of the voice interaction according to the result of the intention recognition includes: obtaining the intention discrimination probability corresponding to each preset intention of the result of the intention recognition; and determining, as the target intention corresponding to the voice request of the current round to complete the voice interaction, one of the preset intentions whose intention discrimination probability is greater than the probability threshold.

[0026] Thus, the voice interaction method of the present invention can obtain the intention discrimination probability corresponding to each preset intention of the result of the intention recognition, and determine, as the target intention corresponding to the voice request, one of the preset intentions whose intention discrimination probability is greater than the probability threshold, so as to recognize the intention of the user to accurately adjust the vehicle components.

[0027] The preset intents include at least one of the following: increasing the volume, decreasing the volume, increasing the air volume, decreasing the air volume, raising the temperature, lowering the temperature, zooming in on the map, zooming out on the map, brightening the screen, dimming the screen, swiping up the screen, swiping down the screen, brightening the instrument panel, dimming the instrument panel, brightening the ambient light, dimming the ambient light, moving the seat forward, moving the seat backward, raising the seat, lowering the seat, moving the seat backrest forward, moving the seat backrest backward, raising the window, and lowering the window.

[0028] In this way, setting multiple preset intents can further lay a foundation for identifying the voice interaction intent of the user.

[0029] The voice interaction method includes: when the intent discrimination probability of each of the preset intents is not greater than a probability threshold, determining that the intent of the voice request in the current round is a non-scale adjustment intent.

[0030] In this way, when the intent discrimination probability of each preset intent is not greater than the probability threshold, determining that the voice request is a non-scale adjustment intent can exclude voice requests with non-scale adjustment intents.

[0031] The present invention also provides a voice interaction device. The voice interaction device includes: a receiving instruction module, a reading instruction module, a query module, a rewriting module, an intent recognition module, and an interaction module. The receiving instruction module is used to receive the voice request of the current round for adjusting the preset functions of the vehicle forwarded by the vehicle, and the preset functions refer to the functions of simulating the operation of vehicle components for scale adjustment; the reading instruction module is used to read the voice request of the previous round for adjusting the preset functions of the vehicle; the query module is used to perform a cache query in the cache engine according to the voice request of the current round and the voice request of the previous round; the rewriting module is used to rewrite the voice request of the current round with the voice request of the previous round when the result of the cache query fails to query the corresponding cache; the intent recognition module is used to perform intent recognition on the rewritten voice request of the current round; and the interaction module is used to complete voice interaction according to the result of the intent recognition.

[0032] In this way, after receiving the voice request from the user for scale adjustment of vehicle components, the voice interaction device of the present invention can, by reading the voice request of the previous round, combine the two rounds of voice requests to query whether the cache is hit. When the corresponding cache cannot be queried, the voice request of the current round is rewritten with the voice request of the previous round so that the rewritten voice request can be recognized by the in-vehicle system of the vehicle for the corresponding intent, and then the scale adjustment of vehicle components is realized in a voice interaction manner according to the intent recognition result. The intent of the voice request is recognized by combining the high-frequency cache engine with intent recognition, and the accurate recognition of the user's intent is realized under multiple rounds of voice requests.

[0033] The present invention provides a server. The server includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the voice interaction method described in any of the above embodiments is implemented.

[0034] In this way, when the server of the present invention executes the computer program through the processor, after receiving a voice request from the user for adjusting the scale of vehicle parts, by reading the previous voice request and combining the two rounds of voice requests to query whether the cache is hit, in the case where the corresponding cache cannot be found, the previous voice request is used to rewrite the current round of voice request so that the rewritten voice request can be recognized by the vehicle's in-vehicle system for the corresponding intention, and then the scale of the vehicle parts is adjusted in a voice interaction manner according to the intention recognition result. The intention of the voice request is recognized by combining the high-frequency cache engine and intention recognition, and the accurate recognition of the user's intention is realized under multiple rounds of voice requests.

[0035] An embodiment of the present invention also provides a non-volatile computer-readable storage medium containing a computer program. When the computer program is executed by one or more processors, the voice interaction method described in any of the above embodiments is implemented.

[0036] In this way, when the computer program stored in the readable storage medium of the present invention is executed by the processor, after receiving a voice request from the user for adjusting the scale of vehicle parts, by reading the previous voice request and combining the two rounds of voice requests to query whether the cache is hit, in the case where the corresponding cache cannot be found, the previous voice request is used to rewrite the current round of voice request so that the rewritten voice request can be recognized by the vehicle's in-vehicle system for the corresponding intention, and then the scale of the vehicle parts is adjusted in a voice interaction manner according to the intention recognition result. The intention of the voice request is recognized by combining the high-frequency cache engine and intention recognition, and the accurate recognition of the user's intention is realized under multiple rounds of voice requests.

[0037] Additional aspects and advantages of the embodiments of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0039] Figure 1 is a schematic flowchart of the voice interaction method of the present invention;

[0040] Figure 2 is a schematic structural diagram of the voice interaction device of the present invention;

[0041] Figure 3 It is a schematic flow diagram of the voice interaction method of the present invention;

[0042] Figure 4 It is a schematic structural diagram of the voice interaction device of the present invention;

[0043] Figure 5 It is a schematic flow diagram of the voice interaction method of the present invention;

[0044] Figure 6 It is a schematic flow diagram of the voice interaction method of the present invention;

[0045] Figure 7 It is a schematic flow diagram of the voice interaction method of the present invention;

[0046] Figure 8 It is a schematic structural diagram of the voice interaction device of the present invention;

[0047] Figure 9 It is a schematic flow diagram of the voice interaction method of the present invention;

[0048] Figure 10 It is a schematic structural diagram of the voice interaction device of the present invention;

[0049] Figure 11 It is a schematic flow diagram of the voice interaction method of the present invention;

[0050] Figure 12 It is a schematic structural diagram of the interaction module in the voice interaction device of the present invention;

[0051] Figure 13 It is a schematic structural diagram of the server of the present invention;

[0052] Figure 14 It is a schematic structural diagram of the computer-readable storage medium of the present invention. Detailed Embodiments

[0053] The following describes in detail the embodiments of the present invention. Examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the embodiments of the present invention and should not be construed as limiting the embodiments of the present invention.

[0054] Currently, in the case of a vehicle's voice interaction system receiving multiple rounds of voice requests from a user, for example, when the user's first-round voice request is "brighten the screen" and the second-round voice request uses a concise voice request like "bright bright bright", the voice interaction system cannot accurately recognize from the user's voice request that the user's second-round requirement is to increase the screen brightness by 3 graduations, and cannot correctly issue a vehicle control instruction to accurately increase the screen brightness by the three brightness levels required by the user, resulting in a poor user experience.

[0055] To solve the above problems, please refer to Figure 1 , the present invention provides a voice interaction method. The voice interaction method includes:

[0056] 01. Receiving the current-round voice request for adjusting a preset function of the vehicle forwarded by the vehicle, where the preset function refers to a function of simulating the operation of vehicle components for graduation adjustment;

[0057] 02. Reading the previous-round voice request for adjusting the preset function of the vehicle;

[0058] 03. Conducting a cache query in a cache engine based on the current-round voice request and the previous-round voice request;

[0059] 04. In the case where the result of the cache query fails to find the corresponding cache, using the previous-round voice request to rewrite the current-round voice request;

[0060] 05. Conducting intent recognition on the rewritten current-round voice request;

[0061] 06. Completing the voice interaction according to the result of the intent recognition.

[0062] Please refer to Figure 2 , the present invention also provides a voice interaction device 10. The voice interaction device 10 includes: a receiving instruction module 11, a reading instruction module 12, a query module 13, a rewriting module 14, an intent recognition module 15, and an interaction module 16.

[0063] Step 01 can be implemented by the receiving instruction module 11, step 02 can be implemented by the reading instruction module 12, step 03 can be implemented by the query module 13, step 04 can be implemented by the rewriting module 14, step 05 can be implemented by the intention recognition module 15, and step 06 can be implemented by the interaction module 16. That is to say, the receiving instruction module 11 can be used to receive the voice request of the current round for adjusting the preset functions of the vehicle forwarded by the vehicle, and the preset function refers to the function of simulating the operation of vehicle components for scale adjustment; the reading instruction module 12 can be used to read the voice request of the previous round for adjusting the preset functions of the vehicle; the query module 13 can be used to perform cache query in the cache engine according to the voice request of the current round and the voice request of the previous round; the rewriting module 14 can be used to rewrite the voice request of the current round with the voice request of the previous round in the case that the result of the cache query fails to find the corresponding cache; the intention recognition module 15 can be used to recognize the intention of the rewritten voice request of the current round; and the interaction module 16 can be used to complete the voice interaction according to the result of the intention recognition.

[0064] In the process where the user uses voice interaction to simulate the scale adjustment of vehicle components, the corresponding voice requests may include but are not limited to "Make the screen brighter", "Make the volume louder", "Move the seat backward". Among them, the preset function refers to the function of completing scale adjustment through vehicle components, and the vehicle components may refer to physical components such as mechanical knobs or buttons, which are components that can be adjusted in scale. Currently in intelligent vehicles, for scenarios where users hope to perform continuous adjustment, in the voice scenario, it is manifested as multi-round interaction. For example, if the user's previous voice request was "Make the volume louder", after the system volume is increased, the user issues the current voice request "Make it a little lower". At this time, the system will not recognize the current voice request or prompt that it doesn't understand, unable to meet the user's need for precise scale continuous adjustment similar to a mechanical knob.

[0065] The present invention can, after receiving the user's voice request for the vehicle preset function, by reading the previous voice request, combine the two rounds of voice requests to query whether the cache is hit. In the case that the corresponding cache cannot be found, use the previous voice request to rewrite the current voice request so that the rewritten voice request can be recognized by the system with the corresponding intention. Thus, after the intention of the rewritten current voice request is recognized, the user's intention can be accurately recognized, and then a control instruction can be issued according to the intention recognition result to control the corresponding vehicle components to complete the voice interaction. Use the method of combining a high-frequency cache engine with intention recognition to recognize the intention of voice requests, and accurately recognize the user's intention of simulating the operation of vehicle components through voice interaction to achieve scale adjustment under multi-round voice requests.

[0066] It should be noted that after receiving the user's voice request for the current round of the vehicle's preset function, the received voice request for the current round is subjected to speech recognition to obtain the speech recognition text for the current round for subsequent processing. For example, when performing speech recognition on the user's voice request for the current round "Make the screen brighter, brighter, brighter", the recognized text for the current round obtained is "Make the screen brighter, brighter, brighter".

[0067] In actual situations, it may be restricted by vehicle hardware, or due to network instability, the user's colloquial or dialectal expressions, etc., resulting in the text instructions after ASR recognition not being clear and accurate enough. The received voice request for the current round can be preprocessed. The preprocessing includes correcting some conventional text errors, such as correcting "Volume deeper, deeper, deeper, deeper, deeper" to "Volume increase, increase, increase, increase, increase", and removing some meaningless words, such as "ah", "please", etc.

[0068] Please combine Figure 3 , before step 01, the voice interaction method may include:

[0069] 011, determine the control range and non-control range of vehicle components.

[0070] Please combine Figure 4 , the voice interaction device 10 further includes a first determination module 111.

[0071] Step 011 can be implemented by the first determination module 111. That is to say, the first determination module 111 can be used to determine the control range and non-control range of vehicle components.

[0072] It can be understood that not all functions of the vehicle can be, are able to, or need to be adjusted with precise graduations. For example, the movement of the seat in various directions can be adjusted through vehicle components. While for the door, there are no vehicle components such as knobs or buttons to achieve graduated adjustment, and it is usually only opened and closed through the door handle. Therefore, seat adjustment belongs to the control range of vehicle components, while door adjustment belongs to the non-control range of vehicle components.

[0073] Obtain the information of vehicle components, based on the information of vehicle components, determine the hardware that can be adjusted with graduations through the components, determine it as the control range of vehicle components, and determine the hardware that cannot be adjusted through vehicle components as the non-control range.

[0074] First, identify the vehicle components that can be adjusted by scale, such as: "volume knob", "screen brightness button", "air conditioner air volume knob / button", "seat adjustment knob / button", etc. Further, determine that the control scope of the vehicle components can include: in-vehicle audio, the screen inside the vehicle, the vehicle air conditioner, the vehicle seat, the ambient light inside the vehicle, the vehicle exterior lights, or the windows, etc. The non-control scope of the vehicle components can include: doors, rearview mirrors, trunks, etc.

[0075] During the subsequent voice interaction process, a voice prompt can be given when the voice request is for the non-control scope of the vehicle components.

[0076] In this way, by collecting vehicle component information and confirming the functions that can be adjusted by scale through the components, the control scope of the vehicle components is determined, that is, the control scope that can be adjusted by scale through voice interaction.

[0077] The voice interaction method further includes:

[0078] 012. Determine the adjustable range of the vehicle components.

[0079] The voice interaction device 10 further includes a second determination module 112.

[0080] Step 012 can be implemented by the second determination module 112. That is to say, the second determination module 112 is used to determine the adjustable range of the vehicle components.

[0081] It can be understood that after determining the control scope and non-control scope of the vehicle components, it is necessary to determine the adjustable range for each vehicle component in the control scope. The adjustable range of the vehicle components corresponds to the scale range adjusted by operating the vehicle components. For different vehicle components, the adjustable range can be gears or ranges. For example, if the screen brightness button is continuously pressed 5 times in succession, the screen brightness is adjusted to the maximum brightness in 1 to 5 brightness levels in sequence, then the adjustable range of the screen brightness button is 1 to 5 levels. Another example is that the total scale value of the knob for adjusting the seat forward and backward is 90, then the adjustable range of the seat adjustment knob is scale values 1 to 90.

[0082] The voice interaction method further includes:

[0083] 013. Map the control scope and the adjustable range to a preset intention.

[0084] The voice interaction device 100 further includes a mapping module 113.

[0085] Step 013 can be implemented by the mapping module 113. That is to say, the mapping module 113 can be used to map the control scope and the adjustable range to a preset intention.

[0086] Map the control range of vehicle components and the adjustable range of each vehicle component to an intention system that can be understood by the intention recognition model. For each object in the control range of the vehicle component and the corresponding adjustable range of the vehicle component, a corresponding preset intention is formulated. For example, system_volume_up represents the preset intention "increase volume" and system_volume_down represents the preset intention "decrease volume", and it includes all adjustable range expressions. For example, "increase volume greatly" corresponds to system_volume_up of the preset intention, and "increase volume very greatly" also corresponds to this intention. In this way, a specific intention mapping system is formulated for the component control range and the adjustable range of vehicle components.

[0087] The preset intentions may include at least one of the following: increase volume, decrease volume, increase air volume, decrease air volume, increase temperature, decrease temperature, zoom in on the map, zoom out on the map, brighten the screen, dim the screen, swipe up the screen, swipe down the screen, brighten the instrument panel, dim the instrument panel, brighten the ambient light, dim the ambient light, move the seat forward, move the seat backward, raise the seat, lower the seat, tilt the backrest forward, tilt the backrest backward, raise the window, and lower the window.

[0088] In this way, setting multiple preset intentions can further lay a foundation for identifying the voice interaction intentions of users. Different intentions are identified based on the voice requests with concise words provided by users, so as to achieve the corresponding target intentions.

[0089] Please combine Figure 5 , the voice interaction method includes:

[0090] 014. Add adjacent two rounds of voice requests with an appearance frequency greater than the preset frequency to the cache engine.

[0091] Step 014 can be implemented by the query module 13. That is to say, the query module 13 can be used to add adjacent two rounds of voice requests with an appearance frequency greater than the preset frequency to the cache engine.

[0092] In this way, the cache of the cache engine of the present invention is composed of adjacent two rounds of voice requests with an appearance frequency greater than the preset frequency, realizing the statistics of high-frequency set voice requests.

[0093] First, the server can collect the historical voice requests of the user for a period of time with the user's permission. The voice requests collected here need to include at least two rounds of voice requests. Among them, it is expected to collect more than 10,000 historical voice requests.

[0094] Secondly, the server can perform a simple screening on the collected historical voice requests to filter out voice requests with obviously unclear semantics, as well as some short voice requests that only contain modal particles, such as "ah", "oh", etc., leaving voice requests with clear semantics and specific purposes, such as "Navigate to the company", "Help me turn on the air conditioner", "Search for nearby hospitals", "Play songs by singer A", "What's the weather like today", etc.; and remove single-round voice requests during the screening.

[0095] The server can perform high-frequency statistics on the screened voice requests to count the occurrence frequencies of adjacent two-round voice requests. Among them, count the number of times that adjacent two-round voice requests appear as unique values. When the number of occurrences is greater than a certain number, it can be considered that the occurrence frequency of the corresponding adjacent two-round voice requests is greater than the preset frequency.

[0096] For example, in the case where the previous-round voice request is "Louder, louder" and the current-round voice request is "Softer", if the number of occurrences in the screened voice requests exceeds the predetermined number, then the adjacent two-round voice requests of "Louder, louder" and "Softer" can be added to the cache engine.

[0097] The voice interaction method further includes:

[0098] 015. Establish a mapping relationship between the current-round voice request and a preset intention.

[0099] Step 015 can be implemented by the query module 13. That is to say, the query module 13 can be used to establish a mapping relationship between the current-round voice request and a preset intention.

[0100] In this way, after the present invention establishes a mapping relationship between the current-round voice request and a preset intention, each preset intention is associated with the corresponding adjacent two-round voice requests, so as to determine the intention corresponding to the voice request in the query of the cache engine.

[0101] It should be understood that the previous-round voice request and the current-round voice request are two adjacent rounds. Among them, after determining the mapping relationship between the current-round voice request and the preset intention in the adjacent two-round voice requests of the cache engine, it can be determined whether the combination of the previous-round voice request and the current-round voice request belongs to a high-frequency set instruction, and whether the target intention of the current-round voice request can be determined according to the preset intention corresponding to the high-frequency set instruction.

[0102] For example, if the previous-round voice request is "Louder, louder" and the current-round voice request is "Softer", then in the adjacent two-round voice requests of "Louder, louder" and "Softer", the preset intention corresponding to the current-round voice request "Softer" is "Turn down the volume".

[0103] Please combine with Figure 6 , the voice interaction method includes:

[0104] 07. When the result of the cache query is that the corresponding cache is found, determine the preset intention corresponding to the current round of voice request as the target intention according to the mapping relationship to complete the voice interaction.

[0105] Step 07 can be implemented by the interaction module 16. That is to say, the interaction module 16 can be used to determine the preset intention corresponding to the current round of voice request as the target intention according to the mapping relationship when the result of the cache query is that the corresponding cache is found, so as to complete the voice interaction.

[0106] In the case that the present invention finds adjacent two rounds of voice requests corresponding to the current round of voice request and the previous round of voice request in the cache engine, according to the established mapping relationship between the current round of voice request and the preset intention, the target intention corresponding to the current round of voice request can be directly determined, so that the voice interaction can be completed according to the determined target intention corresponding to the current round of voice request.

[0107] For example, the previous round of voice request is "Make it louder, make it louder", and the current round of voice request is "Make it quieter". If the adjacent two rounds of voice requests cached in the cache engine are "Make it louder, make it louder" and "Make it quieter", and the preset intention corresponding to "Make it quieter" is "Turn down the volume", then the target intention of the current round of voice request can be directly determined as the queried preset intention "Turn down the volume", so that the operation of vehicle parts can be simulated through voice interaction according to the intention of "Turn down the volume", and the accurate recognition of the user's intention can be realized under multiple rounds of voice requests.

[0108] Please combine with Figure 7 , step 04 includes:

[0109] 041. Train a rewriting model through rewriting training data, and the rewriting training data includes adjacent two rounds of voice requests;

[0110] 042. Rewrite the current round of voice request by using the previous round of voice request and the rewriting model.

[0111] Please combine with Figure 8 , the voice interaction device 10 includes a rewriting training module 114.

[0112] Step 041 can be used to be implemented by the rewriting training module 114, and step 042 can be implemented by the rewriting module 14. That is to say, the rewriting training module 114 can be used to train a rewriting model through rewriting training data. The rewriting module 14 can be used to rewrite the current round of voice request by using the previous round of voice request and the rewriting model.

[0113] In the present invention, a rewriting model is trained by means of machine learning from adjacent two rounds of voice requests, so that the current round of voice request can be rewritten according to the previous round of voice request and the rewriting model, and the rewritten voice request can be recognized by the vehicle-mounted system of the vehicle to identify the corresponding intention. Among them, for the rewriting model, BERT (Bidirectional Encoder Representation from Transformers) and sequence annotation can be used for model training to obtain a trained rewriting model.

[0114] Among them, the rewritten data can be obtained by annotating adjacent two rounds of voice requests in the above-screened voice requests. The current round of voice request in the adjacent two rounds of voice requests can be rewritten and annotated manually. For example, if the previous round of voice request is "Louder, louder", and the current round of voice request is "Softer", then the current round of voice request can be rewritten and annotated as "Volume a little softer". In this way, the annotated adjacent two rounds of voice requests are input into the established rewriting model. During the training process, the rewriting model can learn how to rewrite the current round of voice request before annotation into the current round of voice request after annotation through adjacent two rounds of voice requests by means of feature extraction.

[0115] During the training process, the adjacent two rounds of voice requests in the annotated voice requests are divided into a rewriting training set and a rewriting verification set, and the division ratio can be set according to requirements and is not limited here. For example, the rewriting training set is 80% and the rewriting verification set is 20%. For the established rewriting model, at least part of the data in the rewriting training set is first used to train the rewriting model, and then at least part of the data in the rewriting verification set is used to verify the accuracy of the trained rewriting model. In the case that the accuracy of the rewriting verification does not reach the rewriting accuracy threshold, the rewriting model is trained again through at least another part of the data in the rewriting training set, and the accuracy of the rewritten model after the retraining is verified again by using at least another part of the data in the rewriting verification set. In this way, the process of training and rewriting verification is repeated until the accuracy of the rewriting verification reaches the rewriting accuracy threshold, and it can be considered that the rewriting model has reached the standard and the training of the rewriting model is completed.

[0116] It should be noted that each data in the rewriting training set and the rewriting verification set is only used once. In the case that the rewriting model fails to reach the training standard after traversing all the data in the rewriting training set and the rewriting verification set, more voice requests can be collected again with the permission of the user, so as to screen and annotate to obtain more rewritten training data to train the rewriting model, so as to ensure that the rewriting model can accurately rewrite the voice request.

[0117] Please combine Figure 9 with step 05 including:

[0118] 051, An intent recognition model is trained using intent training data, where the intent training data is related to vehicle components that can be adjusted in scale and the scale adjustment range of the vehicle components.

[0119] 052, Use the intent recognition model to perform intent recognition on the rewritten voice request of the current turn.

[0120] Please combine Figure 10 , The voice interaction device 10 includes an intent training module 115.

[0121] Step 051 can be implemented by the intent training module 115, and step 052 can be implemented by the intent recognition module 15. That is to say, the intent training module 115 can be used to train an intent recognition model using intent training data. The intent recognition module 15 can be used to perform intent recognition on the rewritten voice request of the current turn using the intent recognition model.

[0122] In the present invention, through machine learning, an intent recognition model is trained using training data corresponding to vehicle components that can be adjusted in scale and the scale adjustment range of the vehicle components, and then intent recognition is performed on the rewritten voice request of the current turn to achieve accurate recognition of user intent. Among them, model training can utilize models such as BERT, ALBERT, XLNet, RoBERTa, etc.

[0123] Among them, the intent training data is related to vehicle components that can be adjusted in scale and the scale adjustment range of the components. Vehicle components refer to components on intelligent vehicles that can be adjusted in scale, such as: "volume knob", "screen brightness button", "air conditioner air volume knob / button", "seat adjustment knob / button", etc. The adjustable range of the vehicle component corresponds to the scale range adjusted by operating the vehicle component. For different vehicle components, the adjustable range can be a gear or a range.

[0124] Among them, the intent training data can be obtained by annotating the previous voice request in the adjacent two rounds of voice requests among the above-screened voice requests. The previous voice request in the adjacent two rounds of voice requests can be manually annotated for intent. It can be understood that the previous voice request should include content related to the intent that the user needs to adjust. For example, if the previous voice request is "make the volume louder, louder", and the user needs to adjust the volume up 2 times, at this time, the intent corresponding to the previous voice request can be manually annotated as "increase the volume". In this way, the annotated previous voice request is given to the established intent recognition model. During the training process, the intent recognition model can learn how to recognize the target intent that the user wants to achieve through the input voice request by feature extraction.

[0125] During the training process, the labeled voice requests of the previous round can be divided into an intent training set and an intent verification set. The division ratio can be set according to requirements and is not limited here. For example, the intent training set is 80% and the intent verification set is 20%. For the established intent recognition model, at least part of the data in the intent training set is first used to train the intent recognition model, and then at least part of the data in the intent verification set is used to verify the accuracy of the trained intent recognition model. In the case where the accuracy of the intent verification does not reach the intent accuracy threshold, the intent recognition model is trained again with at least another part of the data in the intent training set, and the accuracy of the intent recognition model after the re-training is verified again with another part of the data in the intent verification set. The process of training and intent verification is repeated in this way until the accuracy of the intent verification reaches the intent accuracy threshold. It can be considered that the intent recognition model has reached the standard and the training of the intent recognition model is completed.

[0126] It should be noted that each piece of data in the intent training set and the intent verification set is only used once. In the case where the intent recognition model fails to reach the training standard after traversing all the data in the intent training set and the intent verification set, more voice requests can be collected again with the user's permission, so as to screen and label more intent training data to train the intent recognition model, so as to ensure that the intent recognition model can accurately recognize the intent corresponding to the input voice request.

[0127] It can be understood that the training of the above rewriting model and intent recognition model can be carried out offline. After deploying the offline-trained rewriting model and intent recognition model to the server, the server can use the voice request rewriting model of the previous round to rewrite the current round of voice requests after receiving the current round of voice requests, and use the intent recognition model to recognize the intent of the rewritten current round of voice requests.

[0128] Please combine Figure 11 with step 06, which includes:

[0129] 061, obtaining the intent discrimination probability corresponding to each preset intent for the result of intent recognition;

[0130] 062, determining a preset intent with an intent discrimination probability greater than the probability threshold as the target intent corresponding to the current round of voice requests to complete the voice interaction.

[0131] Please combine Figure 12 with the interaction module 16 including an obtaining unit 161 and an intent determination unit 162.

[0132] Step 061 can be implemented by the acquisition unit 161, and step 062 can be implemented by the intention determination unit 162. That is to say, the acquisition unit 161 can be used to obtain the intention discrimination probability corresponding to each preset intention of the intention recognition result. The intention determination unit 162 can be used to determine a preset intention with an intention discrimination probability greater than the probability threshold as the target intention corresponding to the voice request to complete the voice interaction.

[0133] According to the recognition results of each preset intention category corresponding to multiple categories of preset intentions, the intention recognition module 15 can give the intention discrimination probabilities matching each preset intention, and then multiple intention discrimination probabilities can be obtained. If the probability threshold is 0.9, and the intention discrimination probability of the recognition result of a certain category of preset intention exceeds 0.9, then the server considers that the preset intention of this category is the target intention of the current user's voice request. The probability threshold can also be other values. The probability threshold can be a default set value or can be set by the user according to needs, and there is no limitation here.

[0134] In this way, the voice interaction method of the present invention can obtain the intention discrimination probability corresponding to each preset intention of the intention recognition result, and determine a preset intention with an intention discrimination probability greater than the probability threshold as the target intention corresponding to the voice request, so as to meet the requirement of accurately identifying the user's intention to adjust vehicle components.

[0135] The voice interaction method includes:

[0136] 063. When the intention discrimination probabilities of all preset intentions are not greater than the probability threshold, determine that the intention of the current round of voice request is a non-scale adjustment intention.

[0137] Step 063 can be implemented by the intention determination unit 162. That is to say, the intention determination unit 162 can be used to determine that the voice request is a non-scale adjustment intention when the intention discrimination probabilities of all preset intentions are not greater than the probability threshold.

[0138] For example, when the intention discrimination probabilities obtained according to each category of preset intentions are not greater than the probability threshold, that is, the probability of the intention recognition result of the user obtained according to the voice request matching each category of preset intention is relatively low and is lower than the probability threshold. For example, the probability threshold can be 0.9, then it is determined that the voice request is a non-scale adjustment intention. The non-scale adjustment intention refers to the user's intention for components that cannot be adjusted by a knob or button with a scale. For example, the voice request input by the user is "Open the car door", because the car door is not a component adjusted by a knob or button with a scale, so the voice request "Open the car door" is a non-scale adjustment intention.

[0139] Thus, when the intention discrimination probabilities of all preset intentions are not greater than the probability threshold, it is determined that the voice request is a non-scale adjustment intention, and voice requests with non-scale adjustment intentions can be excluded.

[0140] Please refer to Figure 13 , the present invention also provides a server 20. The server 20 includes a processor 21 and a memory 22. A computer program 221 is stored on the memory 22. When the computer program 221 is executed by the processor 21, the voice interaction method described in any one of the above embodiments is implemented.

[0141] The server of the present invention can execute the computer program 221 through the processor 21. After receiving a voice request from the user for a preset function of the vehicle, by reading the previous voice request and combining the two rounds of voice requests to query whether there is a hit in the cache. When the corresponding cache cannot be found, the previous voice request is used to rewrite the current voice request so that the rewritten voice request can be recognized by the system for the corresponding intention. Thus, after the intention recognition of the rewritten current voice request, the user's intention can be accurately recognized, and then a control instruction can be issued according to the intention recognition result to control the corresponding vehicle components to complete the voice interaction. The intention of the voice request is recognized by combining the high-frequency cache engine and the intention recognition, so as to accurately recognize the user's intention of simulating the operation of the vehicle components through voice interaction to achieve scale adjustment under multiple rounds of voice requests.

[0142] Please refer to Figure 14 , the present invention also provides a non-volatile computer-readable storage medium 30 containing a computer program 31. When the computer program 31 is executed by one or more processors 40, the voice interaction method described in any of the above embodiments is implemented.

[0143] For example, when the computer program 31 is executed by the processor 40, the following steps of the data processing method are implemented:

[0144] 01. Receive the current round of voice request for adjusting the preset function of the vehicle forwarded by the vehicle. The preset function refers to the function of simulating the operation of the vehicle components for scale adjustment;

[0145] 02. Read the previous round of voice request for adjusting the preset function of the vehicle;

[0146] 03. Perform a cache query in the cache engine according to the current round of voice request and the previous round of voice request;

[0147] 04. When the result of the cache query fails to find the corresponding cache, rewrite the current round of voice request using the previous round of voice request;

[0148] 05. Perform intent recognition on the current round of rewritten voice requests;

[0149] 06. Complete voice interaction according to the results of intent recognition.

[0150] Understandably, a computer program includes computer program code. The computer program code can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution media, etc.

[0151] When the computer program 31 stored in the computer-readable storage medium 30 of the present invention is executed by the processor 40, after receiving the user's voice request for the preset function of the vehicle, by reading the previous round of voice requests, it combines the two rounds of voice requests to query whether there is a cache hit. In the case where no corresponding cache is found, it uses the previous round of voice requests to rewrite the current round of voice requests so that the rewritten voice requests can be recognized by the system for corresponding intents. Thus, after performing intent recognition on the rewritten current round of voice requests, the user's intent can be accurately recognized, and then control instructions can be issued according to the results of intent recognition to control the corresponding vehicle components to complete voice interaction. The method of combining a high-frequency cache engine with intent recognition is used to recognize the intent of voice requests, and under multiple rounds of voice requests, accurately recognize the user's intent to simulate the operation of vehicle components through voice interaction to achieve scale adjustment.

Claims

1. A voice interaction method, characterized in that, it includes: Receiving the voice request of the current round for adjusting the preset functions of the vehicle, where the preset functions refer to the functions of simulating the operation of vehicle components for scale adjustment; Reading the voice request of the previous round for adjusting the preset functions of the vehicle; Performing a cache query in the cache engine according to the voice request of the current round and the voice request of the previous round; In the case where the result of the cache query fails to find the corresponding cache, using the voice request of the previous round to rewrite the voice request of the current round; Performing intent recognition on the rewritten voice request of the current round; Completing voice interaction according to the result of the intent recognition; The voice interaction method includes: Adding adjacent two rounds of voice requests with a frequency of occurrence greater than the preset frequency to the cache engine.

2. The voice interaction method according to claim 1, characterized in that, the voice interaction method includes: Establishing a mapping relationship between the voice request of the current round and the preset intent.

3. The voice interaction method according to claim 2, characterized in that, the voice interaction method includes: In the case where the result of the cache query finds the corresponding cache, determining the preset intent corresponding to the voice request of the current round as the target intent according to the mapping relationship to complete the voice interaction.

4. The voice interaction method according to claim 1, characterized in that, the using the voice request of the previous round to rewrite the voice request of the current round includes: Training a rewriting model through rewritten training data, where the rewritten training data includes adjacent two rounds of voice requests; Using the voice request of the previous round and the rewriting model to rewrite the voice request of the current round.

5. The voice interaction method according to claim 1, characterized in that, the performing intent recognition on the rewritten voice request of the current round includes: Training an intent recognition model through intent training data, where the intent training data is related to the vehicle components that can be adjusted by scale and the scale adjustment range of the vehicle components; Using the intent recognition model to perform intent recognition on the rewritten voice request of the current round.

6. The voice interaction method according to claim 5, characterized in that, the completing voice interaction according to the result of the intent recognition includes: Obtaining the intent discrimination probability of each preset intent corresponding to the result of the intent recognition; Determining one of the preset intents with the intent discrimination probability greater than the probability threshold as the target intent corresponding to the voice request of the current round to complete the voice interaction.

7. The voice interaction method according to claim 6, characterized in that, the preset intents include at least one of: volume up, volume down, air volume up, air volume down, temperature up, temperature down, map zoom in, map zoom out, screen brightness up, screen brightness down, screen swipe up, screen swipe down, instrument brightness up, instrument brightness down, ambient light brightness up, ambient light brightness down, seat forward, seat backward, seat up, seat down, backrest forward, backrest backward, window up and window down.

8. The voice interaction method according to claim 6, wherein, the voice interaction method includes: when the intention discrimination probabilities of all the preset intentions are not greater than a probability threshold, determining that the intention of the current round of voice request is a non-scale adjustment intention.

9. A voice interaction device, wherein, the voice interaction device includes: a receiving instruction module, which is configured to receive the current round of voice request for adjusting a preset function of the vehicle forwarded by the vehicle, and the preset function refers to a function of simulating scale adjustment of operations on vehicle components; a reading instruction module, which is configured to read the previous round of voice request for adjusting the preset function of the vehicle; a query module, which is configured to perform cache query in a cache engine according to the current round of voice request and the previous round of voice request; a rewriting module, which is configured to rewrite the current round of voice request by using the previous round of voice request when the result of the cache query fails to find a corresponding cache; an intention recognition module, which is configured to recognize the intention of the rewritten current round of voice request; an interaction module, which is configured to complete voice interaction according to the result of the intention recognition; the query module is further configured to add adjacent two rounds of voice requests with an occurrence frequency greater than a preset frequency to the cache engine.

10. A server, wherein, the server includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the voice interaction method according to any one of claims 1-8 is implemented.

11. A non-volatile computer-readable storage medium containing a computer program, wherein, when the computer program is executed by one or more processors, the voice interaction method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Voice control method, server, voice control system and readable storage medium

    CN112581955A

  • Semantic recognition method and device

    CN113806470A