A Voice Interaction Method, Device, Electronic Device, Storage Medium, and Vehicle
By using the local side to convert and understand the text of voice commands when the vehicle network status is abnormal, the problem of decreasing recognition speed caused by network status is solved and the user experience is improved.
Patent Information
- Application Number
- CN202310179444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-02-27
AI Technical Summary
In the prior art, when the vehicle network status is abnormal, the recognition speed of voice commands on the server side decreases, resulting in a decrease in user experience.
When the vehicle network status is abnormal, the local side will perform text conversion and semantic understanding of voice commands, obtain target text information and intentions, and select appropriate policies for identification when the network returns to normal.
Improve the recognition speed and overall efficiency of voice commands and improve the user experience.
Smart Images

Figure CN116168693B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of human-computer interaction technologies, and in particular, to a voice interaction method, apparatus, electronic device, storage medium, and vehicle.
Background Art
[0002] With the popularization of intelligent voice cockpits, users can interact with vehicles through voice to control the cockpit. After a user utters a voice command, the voice command uttered by the user is usually first recognized into a text command through Automatic Speech Recognition (ASR), and then the text command is sent to a Natural Language Understanding (NLU) model to recognize the user's intention in the text command, and then the command uttered by the user is responded to.
[0003] In the prior art, both ASR and NLU are on the server side, and the recognition speed of voice commands issued by users is usually greatly affected by the network speed. Once the vehicle network is not smooth or there is no network, the recognition speed of ASR and NLU located on the server side will decrease or even ASR and NLU are in an unresponsive state, resulting in the inability to quickly recognize voice commands issued by users, thereby reducing the user experience.
Summary of the Invention
[0004] The embodiments of the present application provide a voice interaction method, apparatus, electronic device, and storage medium, which can still provide voice interaction services for users in the case of abnormal network status, thereby improving the user experience.
[0005] In a first aspect, the embodiments of the present application provide a voice interaction method, and the method includes:
[0006] Receiving a voice command issued by a user;
[0007] Obtaining current first network quality information;
[0008] If the network quality characterized by the first network quality information is in an abnormal state, converting the voice command into text to obtain target text information; and, performing semantic understanding on the target text information to obtain a target intention;
[0009] Performing a target action corresponding to the target intention.
[0010] In the embodiments of the present application, after receiving a voice command issued by a user, if the network quality characterized by the current first network quality information of the vehicle is in an abnormal state, it can be considered that the current network state cannot support text conversion and semantic understanding on the server side. Therefore, the local side first performs text conversion on the voice command issued by the user to obtain target text information corresponding to the voice command, then performs semantic understanding on the target text information to obtain a target intention, and corresponding target actions can be executed according to the target intention. Compared with the prior art in which text conversion and semantic understanding of the voice command issued by the user can only be performed on the server side, this method can perform text conversion and semantic understanding on the voice command on the local side when the network is abnormal, reducing the dependence on the network in the process of recognizing the voice command, improving the recognition speed of the voice command as a whole, and thus enhancing the user experience.
[0011] Optionally, after obtaining the current first network quality information of the vehicle, the method further includes:
[0012] If the network quality characterized by the first network quality information is in a normal state, send the voice command to the server;
[0013] Receive the target text information corresponding to the voice command;
[0014] Determine whether the target text information belongs to a local text information set, where the local text information set is all text information corresponding to services in which the vehicle can perform voice interaction with the user without relying on the network;
[0015] If it is determined that the target text information belongs to the local text information set, then based on the target text information, obtain the corresponding target intention, and the target intention is used to execute the target action.
[0016] In the embodiment of the present application, when the network quality characterized by the current first network quality information of the vehicle is in a normal state, the voice combination instruction is sent to the server side for text conversion. After receiving the target text information generated by the server side for text conversion, in the process of obtaining the target intention, some of the obtained target intentions need to be obtained through the query of the server side, while obtaining the target intention corresponding to the local text information does not require network query. By judging whether the target text information belongs to the local text information set, it can be distinguished whether the target intention corresponding to the target text information is obtained through the query of the server side. If the target text information belongs to the local text information set, the semantic understanding of the target text information can be completed at the local side, and the semantic understanding at the local side does not depend on the network. The semantic understanding of the target text information belonging to the local text information set at the local side is faster than that at the server side, which also improves the efficiency of the overall recognition process of the voice instruction to a certain extent, thereby improving the user experience.
[0017] Optionally, after judging whether the target text information belongs to the local text information set, the method further includes:
[0018] If it is determined that the target text information does not belong to the local text information set, send the target text information to the server;
[0019] Receive the target intention corresponding to the target text information.
[0020] In the embodiment of the present application, since the text conversion and semantic understanding of the voice instruction at the server side are higher in recognition rate and faster in recognition speed than those at the local side, when the network quality characterized by the current first network quality information of the vehicle is in a normal state, the voice combination instruction is sent to the server side for text conversion. After receiving the target text information generated by the server side for text conversion, in the process of obtaining the target intention, some of the obtained target intentions need to be obtained through the query of the server side, while obtaining the target intention corresponding to the local text information does not require network query. By judging whether the target text information belongs to the local text information set, it can be distinguished whether the target intention corresponding to the target text information is obtained through the query of the server side. If the local text information does not belong to the local text information set, the target text information can only obtain the corresponding target intention through semantic understanding at the server side, which not only better meets the user's voice interaction needs, but also improves the efficiency of the overall recognition process of the voice instruction to a certain extent, thereby improving the user experience.
[0021] Optionally, performing semantic understanding on the target text information to obtain a target intention includes:
[0022] Obtain the current second network quality information;
[0023] If the network quality represented by the second network quality information is in a normal state, determine whether the target text information belongs to the local text information set, where the local text information set is all the text information corresponding to the services in which the vehicle can perform voice interaction with the user without relying on the network;
[0024] If it is determined that the target text information belongs to the local text information set, perform semantic understanding on the target text information to obtain the corresponding target intent.
[0025] In the embodiments of the present application, if the network quality represented by the obtained first network quality information is in an abnormal state, it can be considered that the current network state cannot support text conversion on the server side. Therefore, the local side first performs text conversion on the voice command issued by the user to obtain the target text information corresponding to the voice command. After obtaining the target text information, the current network quality may return to a normal state. Then, re-obtain the quality parameter of the second network in the vehicle. If the network quality represented by the second network quality information is in a normal state, it is possible to select whether to perform semantic understanding on the target text information on the local side or the server side by determining whether the target text information belongs to the local text information set. If the target text information belongs to the local text information set, semantic understanding can be completed on the local side. When the network quality quickly recovers, a more reasonable semantic understanding strategy can be adopted to improve the efficiency of the semantic understanding process, and to a certain extent, improve the overall recognition efficiency of the voice command and enhance the user experience.
[0026] Optionally, after determining whether the target text information belongs to the local text information set, the method further includes:
[0027] If it is determined that the target text information does not belong to the local text information set, send the target text information to the server;
[0028] Receive the target intent corresponding to the target text information.
[0029] In the embodiment of the present application, if the network quality represented by the acquired first network quality information is in an abnormal state, it can be considered that the current network state cannot support text conversion on the server side. Therefore, the local side first performs text conversion on the voice command issued by the user to obtain the target text information corresponding to the voice command. After obtaining the target text information, the current network quality may be restored to a normal state, and then the quality parameters of the vehicle's current second network are re-acquired. If the network quality represented by the second network quality information is in a normal state, by judging whether the target text information belongs to the local text information set, it can be distinguished whether the target intent corresponding to the target text information is obtained through server-side query. If the target local information does not belong to the local text information set, the target text information can only be semantically understood on the server side to obtain the corresponding target intent. This not only better meets the user's voice interaction needs, but also judges the network quality at that time before performing text conversion and semantic understanding, and can promptly meet the user's voice interaction needs when the network quickly returns to normal, and to a certain extent, it also improves the efficiency of the overall recognition process of voice commands, thereby improving the user experience.
[0030] Optionally, when executing the target action corresponding to the target intention, the method further includes:
[0031] The target action and the target sentence corresponding to the target action are announced to the user by voice.
[0032] In an embodiment of the present application, after the target intent is acquired, while the target action corresponding to the target intent is executed, a preset target sentence corresponding to the target action is broadcast to the user, thereby creating a good voice interaction scene for the user and improving the user experience.
[0033] In a second aspect, an embodiment of the present application provides a voice interaction device, the device comprising:
[0034] A first receiving unit, used to receive a voice command issued by a user;
[0035] A first acquisition unit, acquiring current first network quality information;
[0036] a conversion unit, configured to convert the voice command into text to obtain target text information when the network quality represented by the first network quality information is in an abnormal state;
[0037] An understanding unit, used to perform semantic understanding on the target text information to obtain the target intention;
[0038] An execution unit is used to execute the target action corresponding to the target intention.
[0039] Optionally, the device further comprises:
[0040] A first sending unit, configured to send the voice command to a server when the network quality characterized by the first network quality information is in a normal state;
[0041] The first receiving unit is further configured to receive the target text information corresponding to the voice command;
[0042] A first determining unit, configured to determine whether the target text information belongs to a local text information set, where the local text information set is all text information corresponding to services for which the vehicle can perform voice interaction with a user without relying on a network;
[0043] The understanding unit is further configured to, when it is determined that the target text information belongs to the local text information set, obtain a corresponding target intent based on the target text information, where the target intent is used to perform a target action.
[0044] Optionally, the first sending unit is further configured to send the target text information to the server when it is determined that the target text information does not belong to the local text information set;
[0045] The first receiving unit is further configured to receive the target intent corresponding to the target text information.
[0046] Optionally, the understanding unit includes:
[0047] A second obtaining unit, configured to obtain current second network quality information;
[0048] A second determining unit, configured to determine whether the target text information belongs to the local text information set if the network quality characterized by the second network quality information is in a normal state, where the local text information set is all text information corresponding to services for which the vehicle can perform voice interaction with a user without relying on a network;
[0049] A voice understanding unit, configured to perform semantic understanding on the target text information to obtain a corresponding target intent if it is determined that the target text information belongs to the local text information set.
[0050] Optionally, the understanding unit further includes:
[0051] A second sending unit, configured to send the target text information to the server when it is determined that the target text information does not belong to the local text information set;
[0052] A second receiving unit, configured to receive the target intent corresponding to the target text information.
[0053] Optionally, when performing the target action corresponding to the target intention, the execution unit is further configured to:
[0054] Voice broadcast the target statement corresponding to the target action to the user.
[0055] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a processor and a memory. When the processor executes a computer program stored in the memory, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.
[0056] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.
[0057] In a fifth aspect, an embodiment of the present invention provides a vehicle, including: the electronic device provided in the third aspect embodiment of the present invention.
Description of the Drawings
[0058] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present application;
[0060] Figure 2 It is a schematic flowchart of a process for obtaining a target intention provided by an embodiment of the present application;
[0061] Figure 3 It is a schematic flowchart of a process for obtaining a target intention provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present application;
[0063] Figure 5 It is a schematic flowchart of a process for obtaining a target intention provided by an embodiment of the present application;
[0064] Figure 6 It is a schematic flowchart of a process for obtaining a target intention provided by an embodiment of the present application;
[0065] Figure 7 It is a schematic flowchart of a voice broadcast process provided by an embodiment of the present application;
[0066] Figure 8 Schematic flowchart of a voice interaction method provided by an embodiment of the present application;
[0067] Figure 9 Schematic structural diagram of a voice interaction device provided by an embodiment of the present application;
[0068] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0069] Figure 11 Schematic structural diagram of a vehicle provided by an embodiment of the present application.
Specific embodiments
[0070] For a better understanding of the technical solutions of this specification, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0071] It should be clear that the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts belong to the scope protected by this specification.
[0072] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit this specification. The singular forms of "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0073] It has been found by the inventors of the present application that when a user conducts voice interaction with a vehicle, the voice commands issued by the user are all recognized on the server side. However, the recognition of the voice commands issued by the user on the server side is greatly affected by the current network state of the vehicle. When the current network state of the vehicle is abnormal, the server cannot recognize the voice commands issued by the user, thus affecting the user experience.
[0074] In view of this, the embodiments of the present application provide a voice interaction method. In this method, the voice commands issued by the user can be recognized locally when the vehicle network state is abnormal, so that a voice interaction service can still be provided to the user when the network state is abnormal, thereby improving the user experience.
[0075] The technical solutions provided by the embodiments of the present application will be introduced below with reference to the accompanying drawings. Please refer to Figure 1 , the embodiments of the present application provide a voice interaction method. This method is applied to a vehicle, and the process of this method is described as follows:
[0076] Step 101: Receive a voice command issued by the user.
[0077] When a user conducts a voice interaction with a vehicle, the voice command issued by the user represents the action that the user wants the vehicle to perform. For example, if the voice command issued by the user is "Open the sunroof", the action performed by the vehicle after recognition is to open the sunroof; another example is that if the voice command issued by the user is "What's the weather today", the action performed by the vehicle after recognition is to display the weather forecast information for today to the user. Therefore, in the embodiments of the present application, when a user conducts a voice interaction with a vehicle and wants the vehicle to perform a specific action, it is necessary to first obtain the voice command issued by the user.
[0078] Step 102: Obtain the current first network quality information.
[0079] Considering that the network quality of the vehicle has a great impact on the voice interaction between the user and the vehicle, therefore, by obtaining the current first network quality information, the first network quality information can reflect the state of the vehicle network quality and provide a judgment basis for determining whether the voice command should be recognized on the server side or the local side.
[0080] For example, the first network quality information can be network speed, transmission delay, bandwidth, etc., which can be set according to requirements and are not limited herein.
[0081] Step 103: If the network quality characterized by the first network quality information is in an abnormal state, perform text conversion on the voice command to obtain target text information.
[0082] Step 104: Perform semantic understanding on the target text information to obtain a target intention.
[0083] In the embodiments of the present application, a module capable of recognizing voice commands is set on the local side. First, the first network quality information is judged. If the network quality characterized by the first network quality information is in an abnormal state, and if the first network quality information is network speed, when the network speed is lower than the set threshold, it can be judged that the network quality is in an abnormal state. (Here, when the first network quality information is different, the method of determining that the network characterized by the first network quality information is abnormal is also different, and no specific limitation is made here, and it can be set according to different requirements). The ASR and NLU in the server side require network support and cannot respond to the voice commands issued by the user, but the ASR and NLU on the local side are not affected by the network quality. Therefore, the voice commands issued by the user can be completed by the ASR and NLU on the local side. The ASR on the local side performs text conversion on the voice command to obtain target text information, and then the NLU on the local side performs semantic understanding on the target text information to further obtain the target intention. Therefore, when the vehicle network quality is abnormal, recognizing the voice command on the local side can reduce the dependence on the network to a certain extent, improve the recognition efficiency of the vehicle for voice commands as a whole, and thus improve the user experience.
[0084] Figure 2 This is a schematic flowchart of a process for obtaining a target intent provided in an embodiment of the present application. As a possible implementation, when performing step 104, the specific process of obtaining the target intent can be achieved through steps 1041 - 1043:
[0085] Step 1041: Obtain the current second network quality information.
[0086] Step 1042: If the network quality represented by the second network quality information is in a normal state, determine whether the target text information belongs to the local text information set. The local text information set is all the text information corresponding to services in which the vehicle can interact with the user by voice without relying on the network.
[0087] Step 1043: If it is determined that the target text information belongs to the local text information set, perform semantic understanding on the target text information to obtain the corresponding target intent.
[0088] In an embodiment of the present application, when the network quality represented by the first network quality information is in an abnormal state, after converting the voice command into text at the local end and obtaining the corresponding target text information, the current second network quality information can be obtained again. The second network quality information can better represent the network quality at this time. When the network quality represented by the second network quality information is in a normal state, it can be considered that the network quality has recovered from the abnormal state to the normal state. By determining whether the target text information belongs to the local text information set, it is judged whether the target text information is suitable for semantic understanding at the local end. If the target text information belongs to the local text information set, semantic understanding is performed on the target text information at the local end to obtain the corresponding target intent. The network quality is judged before both text conversion and semantic understanding. By detecting the network quality more frequently, more efficient semantic understanding strategies can be adopted when the network quality quickly recovers, which can improve the efficiency of the overall recognition process of the voice command to a certain extent and enhance the user experience.
[0089] Figure 3 This is a schematic flowchart of a process for obtaining a target intent provided in an embodiment of the application. As a possible implementation, after performing step 1042, the specific process of obtaining the target intent can be achieved through steps 1044 - 1045:
[0090] Step 1044: If it is determined that the target text information does not belong to the local text information set, send the target text information to the server.
[0091] Step 1045: Receive the target intent corresponding to the target text information.
[0092] In the embodiments of the present application, when the network quality characterized by the first network quality information is in an abnormal state, after converting the voice command into text at the local end and obtaining the target text information, the current second network quality information is obtained again. The second network quality information can better characterize the network quality at this time. When the network quality characterized by the second network quality information is in a normal state, it can be considered that the network quality has recovered from abnormal to normal at this time. By determining whether the target text information belongs to the local text information set, it can be distinguished whether the target intention corresponding to the target text information is obtained through querying the server end, so as to determine whether the target text information is suitable for semantic understanding at the local end. If the target text information does not belong to the local text information set, the target text information can only obtain the corresponding target intention through semantic understanding at the server end. This not only meets the user's voice interaction needs, but also judges the network quality at that time before converting the voice command into text and performing semantic understanding. By detecting the network quality more frequently, it can timely meet the user's voice interaction needs when the network quickly returns to normal, improving the efficiency of the semantic understanding process, and to a certain extent, also improving the efficiency of the overall recognition process of the voice command, enhancing the user experience.
[0093] Figure 4 FIG. is a schematic flowchart of a voice interaction method provided in the embodiments of the present application. After step 103 is executed, step 104 can also be implemented by sub-steps 1041-1045. Among them, after step 1041 is executed, if the network quality characterized by the second network quality information obtained in step 1041 is in a normal state, step 1042 is executed. If it is determined in step 1042 that the target text information belongs to the local text information set, it jumps to step 1043. If it is determined in step 1042 that the target text information does not belong to the local text information set, it jumps to steps 1044-1045.
[0094] Step 105: Execute the target action corresponding to the target intention.
[0095] In the embodiments of the present application, each obtained target intention corresponds to a corresponding target action. After identifying the target intention in the voice command, the target action corresponding to the target intention is executed, so as to truly complete the user's needs.
[0096] In some embodiments, when the network quality characterized by the first network quality state is in a normal state, in order to improve the text conversion efficiency of the voice command, after converting the voice command into text at the server end and obtaining the target text information, before performing semantic understanding, it is determined whether the target text information is suitable for semantic understanding locally. If the target text information is suitable for semantic understanding locally, semantic understanding is preferentially performed locally.
[0097] Figure 5A schematic flowchart of a process for obtaining a target intention provided by an embodiment of the present application. As a possible implementation manner, after performing step 102, the specific process of obtaining the target intention can be implemented through steps 201-204:
[0098] Step 201: If the network quality characterized by the first network quality information is in a normal state, send a voice command to the server.
[0099] Step 202: Receive the target text information corresponding to the voice command.
[0100] Step 203: Determine whether the target text information belongs to the local text information set, where the local text information set is all the text information corresponding to services in which the vehicle can interact with the user without relying on the network.
[0101] Step 204: If it is determined that the target text information belongs to the local text information set, then based on the target text information, obtain the corresponding target intention, and the target intention is used to perform the target action.
[0102] In the embodiment of the present application, when the network quality characterized by the first network quality information is in a normal state, for example, if the first network quality information is the network speed, when the network speed is greater than the set threshold, it can be considered that the network quality is in a normal state. To improve the text conversion efficiency of the voice command, the voice command is converted into text on the server side to obtain the corresponding target text information. Then, it is determined whether the generated target text information belongs to the local text information set. By determining whether the target text information belongs to the local text information set, it can be distinguished whether the target intention corresponding to the target text information is obtained through querying on the server side. If the target text information belongs to the local text information set, the semantic understanding of the target text information on the local side is relatively faster than that on the server side. When the network quality is normal, the server side is used to convert the voice command into text, which improves the text conversion efficiency. When the target text information belongs to the local text information set, the semantic understanding of the target text information is performed on the local side, which improves the semantic understanding efficiency and can also reduce the dependence on the network to a certain extent during the semantic understanding process. In summary, it can be considered that this embodiment improves the recognition speed of the voice command as a whole, thereby improving the user experience.
[0103] In some embodiments, when the network quality characterized by the first network quality state is in a normal state, to improve the text conversion efficiency of the voice command, before performing semantic understanding, it is determined whether the target text information is suitable for semantic understanding locally. If the target text information is not suitable for semantic understanding locally, the semantic understanding of the target text information is performed on the server side.
[0104] Figure 6A flowchart for obtaining a target intent provided by an embodiment of the present application. As a possible implementation, after performing step 203, the specific process of obtaining the target intent can also be achieved through steps 205-206:
[0105] Step 205: If it is determined that the target text information does not belong to the local text information set, send the target text information to the server.
[0106] Step 206: Receive the target intent corresponding to the target text information.
[0107] In the embodiment of the present application, when the network quality characterized by the first network quality information is in a normal state, in order to improve the text conversion efficiency of voice commands, the voice commands are converted into text on the server side to obtain the target text information. By determining whether the generated target text information belongs to the local text information set, it can be distinguished whether the target intent corresponding to the target text information is obtained through query on the server side. If the target text information does not belong to the local text information set, the target text information can only obtain the corresponding target intent through semantic understanding on the server side, which better meets the user's voice interaction needs. In summary, it can be considered that this embodiment improves the recognition efficiency of voice commands as a whole, thereby improving the user experience.
[0108] Furthermore, it is known that the services corresponding to the text information in the local text information set are services that can interact with users without relying on the network. For example, voice commands such as turning on the air conditioner, turning on the radio, and opening the sunroof can be executed locally, and such text information belongs to the local text information set. However, some target text information requires querying the server to obtain the target intent. For example, voice commands such as displaying weather information and displaying the road conditions ahead require querying the server to obtain accurate information. Therefore, such text information does not belong to the local text information set.
[0109] Since only the network quality information is obtained once in the above embodiment during voice interaction, sometimes the network quality anomaly only lasts for a very short time. It is possible that the network quality is abnormal when obtaining the first network quality information, but the network quality has returned to normal after a very short time. Therefore, in the embodiment of the present application, the network quality information is obtained before converting the voice command into text and performing semantic understanding, and a more appropriate voice command recognition strategy is adopted according to the network quality information at different times.
[0110] Figure 7 A voice broadcast flowchart provided by an embodiment of the application. As a possible implementation, the embodiment of the present application can also perform voice broadcast to the user, and the specific broadcast process can be achieved through step 301:
[0111] Step 301: Verbally announce the target statement corresponding to the target action to the user.
[0112] In the embodiments of the present application, when performing the target action corresponding to the target intent, the target statement corresponding to the target action is simultaneously announced to the user. The statements corresponding to each target action can be set with different target statements according to different requirements, and no specific limitation is made here. Verbally announcing to the user can create a good voice interaction scenario for the user and further enhance the user experience.
[0113] Figure 8 For the flowchart of a voice interaction method provided in the embodiments of the present application, please refer to Figure 8 ., after receiving the voice command issued by the user, the voice command will be converted into text and semantically understood to obtain the target intent. The vehicle performs corresponding actions based on the obtained target intent. The process of obtaining the target intent includes all the above situations, which will be briefly described below:
[0114] Step 101: Receive the voice command issued by the user.
[0115] Step 102: Obtain the current first network quality information.
[0116] When the network quality characterized by the first network quality information is in an abnormal state, jump to Step 103 - 104. When the network quality characterized by the first network quality information is in a normal state, jump to Step 201 - 203. If it is determined in Step 203 that the target text information belongs to the local text information set, jump to Step 204; if it is determined in Step 203 that the target text information does not belong to the local text information set, jump to Step 205 - 206.
[0117] After obtaining the target intent corresponding to the target text information in Steps 104, 206, and 204, execute Step 105, and simultaneously execute Step 301 while executing Step 105.
[0118] Please refer to Figure 9 ., based on the same inventive concept, the embodiments of the present application also provide a voice interaction device, which includes: a first receiving unit 401, a first obtaining unit 402, a conversion unit 403, an understanding unit 404, and an execution unit 405.
[0119] The first receiving unit 401 is configured to receive the voice command issued by the user;
[0120] The first obtaining unit 402 obtains the current first network quality information;
[0121] The conversion unit 403 is configured to perform text conversion on the voice command to obtain the target text information when the network quality characterized by the first network quality information is in an abnormal state;
[0122] An understanding unit 404 for semantically understanding the target text information to obtain a target intention;
[0123] A unit 405 for performing a target action corresponding to the target intention.
[0124] Optionally, the device further includes:
[0125] A first sending unit for sending a voice command to the server when the network quality characterized by the first network quality information is in a normal state;
[0126] The first receiving unit 401 is further configured to receive the target text information corresponding to the voice command;
[0127] A first determination unit for determining whether the target text information belongs to a local text information set, where the local text information set is all the text information corresponding to services in which the vehicle can interact with the user by voice without relying on the network;
[0128] The understanding unit 404 is further configured to, when it is determined that the target text information belongs to the local text information set, obtain a corresponding target intention based on the target text information, and the target intention is used to perform a target action.
[0129] Optionally, the first sending unit is further configured to send the target text information to the server when it is determined that the target text information does not belong to the local text information set;
[0130] The first receiving unit 401 is further configured to receive the target intention corresponding to the target text information.
[0131] Optionally, the understanding unit 404 includes:
[0132] A second obtaining unit for obtaining the current second network quality information;
[0133] A second determination unit for determining whether the target text information belongs to the local text information set if the network quality characterized by the second network quality information is in a normal state, where the local text information set is all the text information corresponding to services in which the vehicle can interact with the user by voice without relying on the network;
[0134] A voice understanding unit for semantically understanding the target text information to obtain a corresponding target intention if it is determined that the target text information belongs to the local text information set.
[0135] Optionally, the understanding unit 404 further includes:
[0136] A second sending unit for sending the target text information to the server when it is determined that the target text information does not belong to the local text information set;
[0137] A second receiving unit, configured to receive a target intention corresponding to target text information.
[0138] Optionally, when performing a target action corresponding to the target intention, the execution unit 405 is further configured to:
[0139] Voice broadcast a target statement corresponding to the target action to the user.
[0140] Please refer to Figure 10 , this application embodiment provides an electronic device 100, which includes at least one processor 501. The processor 501 is configured to execute a computer program stored in a memory, so as to implement the steps of the voice interaction method provided in this application embodiment as Figure 1 shown.
[0141] Optionally, the processor 501 may specifically be a central processing unit or a specific ASIC, and may be one or more integrated circuits for controlling program execution.
[0142] Optionally, the electronic device 100 may further include a memory 502 connected to at least one processor 501. The memory 502 may include a ROM, a RAM, and a disk memory. The memory 502 is used to store data required during the operation of the processor 501, that is, instructions that can be executed by at least one processor 501 are stored. The at least one processor 501 executes the instructions stored in the memory 502 to execute the method as Figure 1 shown. Among them, the number of the memories 502 is one or more. Among them, the memory 502 is shown together in the figure, but it should be noted that the memory 502 is not an essential functional module, so it is shown by a dashed line in Figure 10 .
[0143] Among them, the physical devices corresponding to the first receiving unit 401, the first obtaining unit 402, the conversion unit 403, the understanding unit 404, and the execution unit 405 may all be the foregoing processor 501. The electronic device may be used to execute the method provided in the Figures 1 - 8 shown embodiment. Therefore, regarding the functions that can be realized by each functional module in the electronic device, reference may be made to the corresponding descriptions in the Figures 1 - 8 shown embodiment, which will not be elaborated here.
[0144] This application embodiment further provides a computer storage medium. Among them, the computer storage medium stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the method as Figures 1 - 8 described.
[0145] Please refer to Figure 11 , this application embodiment further provides a vehicle 200, and the vehicle 200 includes as Figure 9The electronic device 100 shown.
[0146] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of protection of this specification.
Claims
1. A voice interaction method, characterized in that, Applied to a vehicle, the method comprises: Receive voice commands from users; Obtain current first network quality information; If the network quality represented by the first network quality information is in an abnormal state, the voice command is converted into text locally to obtain target text information; and semantic understanding of the target text information is performed to obtain a target intent; Execute the target action corresponding to the target intention; Perform semantic understanding on the target text information to obtain the target intention, including: Obtain the current second network quality information again; If the network quality represented by the second network quality information is normal, determining whether the target text information belongs to a local text information set, where the local text information set is all text information corresponding to a service for which the vehicle can perform voice interaction with a user without relying on a network; If it is determined that the target text information belongs to the local text information set, semantic understanding is performed locally on the target text information to obtain the corresponding target intent; If it is determined that the target text information does not belong to the local text information set, the target text information is sent to a server; and the target intent corresponding to the target text information is received.
2. The method according to claim 1, characterized in that, After obtaining the current first network quality information of the vehicle, the method further includes: If the network quality represented by the first network quality information is normal, sending the voice command to the server; Receiving the target text information corresponding to the voice command; Determining whether the target text information belongs to a local text information set, where the local text information set is all text information corresponding to a service for which the vehicle can perform voice interaction with a user without relying on a network; If it is determined that the target text information belongs to the local text information set, the corresponding target intent is obtained based on the target text information, and the target intent is used to execute the target action.
3. The method according to claim 2, wherein The method further comprises: If it is determined that the target text information does not belong to the local text information set, sending the target text information to the server; Receive the target intent corresponding to the target text information.
4. The method according to claim 1, wherein When executing the target action corresponding to the target intention, the method further includes: The target action and the target sentence corresponding to the target action are announced to the user by voice.
5. A voice interaction device, characterized in that, Applied to a vehicle, the device comprises: A first receiving unit, used to receive a voice command issued by a user; A first acquisition unit, acquiring current first network quality information; a conversion unit, configured to convert the voice command into text locally to obtain target text information when the network quality represented by the first network quality information is in an abnormal state; An understanding unit, used to perform semantic understanding on the target text information to obtain the target intention; An execution unit, used for executing a target action corresponding to the target intention; The understanding unit comprises: A second acquisition unit, used to acquire the current second network quality information again; A second determination unit, configured to determine whether the target text information belongs to a local text information set if the network quality characterized by the second network quality information is in a normal state, where the local text information set is all text information corresponding to services in which the vehicle can perform voice interaction with the user without relying on the network; A voice understanding unit, configured to perform semantic understanding on the target text information locally to obtain the corresponding target intent if it is determined that the target text information belongs to the local text information set; The understanding unit further includes: A second sending unit, configured to send the target text information to the server when it is determined that the target text information does not belong to the local text information set; A second receiving unit, configured to receive the target intent corresponding to the target text information.
6. An electronic device, characterized in that, The electronic device includes at least one processor and a memory connected to the at least one processor, and the at least one processor is configured to implement the steps of the method according to any one of claims 1-4 when executing a computer program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the steps of the method according to any one of claims 1-4.
8. A vehicle, characterized in that, An electronic device including the electronic device according to claim 6.
Citation Information
Patent Citations
Conversation marking method and device, aggregation server and storage medium
CN107894972A
Voice processing method and device for vehicle-mounted equipment, equipment and storage medium
CN112509585A