Voice interaction method and device, computer equipment and storage medium
By setting visible and invisible controls with the same triggering events in the on-board smart cockpit system and setting hot word attributes on the invisible control, the problem of cumbersome logic when implementing multiple voice functions is solved, and the simplification of functions and the accuracy of operation is achieved.
Patent Information
- Application Number
- CN202311583262.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-24
AI Technical Summary
In the existing vehicle-mounted smart cockpit system, when the voice assistant function needs to implement multiple voice functions, it needs to repeatedly change the hot word attributes of the control in the code logic, resulting in cumbersome logic and error-prone.
By setting visible controls and invisible controls that have the same triggering event after being triggered, the hot word attributes are not set for visible controls, but the hot word attributes are set for invisible controls. The visible control is used to respond to manual click triggers, while the invisible control is used to respond to visible and speakable voice input triggers, thus simplifying multiple functions.
By separating the controls that carry hot words from the controls visible to the interface, the implementation logic of multiple functions is simplified, the cumbersome and error risks of repeatedly setting hot words attributes in the code logic are avoided, and the accuracy and flexibility of operations are improved.
Smart Images

Figure CN120048254A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer program technology, and in particular to a voice interaction method, apparatus, computer device and storage medium. Background Art
[0002] At present, most of the in-vehicle smart cockpit systems on the market are equipped with voice assistant functions. Among them, the visible-is-speaking technology refers to the interface elements that can be seen by the user in the in-vehicle system interface. Through voice wake-up and voice input, the corresponding UI controls can be directly controlled and hit, thereby realizing the technology of controlling the interface operation. This technology registers and scans the text on the interface through the Android accessibility service, and when the user wakes up with voice, the voice text entered by the user is matched with the UI controls with the registered related hot words on the current interface, so as to accurately respond to the specified control control.
[0003] However, since each control can only set one hot word attribute to implement a voice function, when multiple voice functions need to be implemented through one control, it is necessary to add processing logic to the code logic to repeatedly change the hot word attributes of the control to implement multiple voice functions, which makes the logic cumbersome and error-prone. Summary of the invention
[0004] Based on this, it is necessary to provide a voice interaction method, device, computer equipment, storage medium and computer program product that simplifies the logic of multiple function implementation in order to solve the above technical problems.
[0005] In a first aspect, the present application provides a voice interaction method, which is applied to a voice interaction interface, wherein there are matching visible controls and invisible controls in the voice interaction interface, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot words; the method comprises:
[0006] Get the target hot words carried in the voice interaction command;
[0007] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0008] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0009] In one of the embodiments, when a visible control is used to implement multiple functions, the number of functions implemented by the visible control is the same as the number of matching invisible controls, and the hot word attribute of each invisible control that matches the visible control corresponds to a function, and the invisible controls that match the visible controls are superimposed on the voice interaction interface.
[0010] In one of the embodiments, invisible controls are disabled for touch events.
[0011] In one embodiment, the implementation process of disabling touch events for invisible controls includes:
[0012] For the processing method corresponding to the invisible control for processing touch events, the return value of the processing method is set to false.
[0013] In one embodiment, the method further comprises:
[0014] In response to the voice interaction instruction, the corresponding animation effect is displayed based on the position of the invisible control in the voice interaction interface.
[0015] In one of the embodiments, the display size of the invisible control is set to 0 to 2 pixel units, and the type of the invisible control is one of the following control types, including a picture control, a text control, and a button control.
[0016] In a second aspect, the present application also provides a voice interaction device, the device comprising:
[0017] An acquisition module is used to acquire target hot words carried in voice interaction instructions;
[0018] A matching module, used to match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0019] The execution module is used to execute the corresponding voice interaction process based on the trigger event of the target invisible control after being triggered by the hot word.
[0020] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, applied to a voice interaction interface, the voice interaction interface having matching visible controls and invisible controls, the visible controls not being set with hot word attributes, the triggering event of the visible controls after being touched being the same as the triggering event of the matching invisible controls after being triggered by the hot words; when the processor executes the computer program, the following steps are implemented:
[0021] Get the target hot words carried in the voice interaction command;
[0022] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0023] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0024] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which is applied to a voice interaction interface, wherein there are matching visible controls and invisible controls in the voice interaction interface, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot words; when the computer program is executed by a processor, the following steps are implemented:
[0025] Get the target hot words carried in the voice interaction command;
[0026] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0027] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0028] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, which is applied to a voice interaction interface, in which there are matching visible controls and invisible controls, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot words; when the computer program is executed by a processor, the following steps are implemented:
[0029] Get the target hot words carried in the voice interaction command;
[0030] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0031] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0032] The above-mentioned voice interaction method, device, computer equipment, storage medium and computer program product, by setting visible controls and invisible controls that have the same trigger event after being triggered, do not set the hot word attribute for the visible control, and set the hot word attribute for the invisible control. Since the control carrying the hot word can be separated from the control visible in the interface, when multiple functions are realized for the visible control, multiple matching invisible controls can be set for the visible control, and the hot word attribute of each invisible control can be set separately to realize the corresponding function, thereby improving the cumbersome and error-prone situation caused by repeatedly setting the hot word attribute through code logic.
[0033] In addition, since the controls that carry hot words can be separated from the visible controls on the interface, the visible controls are dedicated to responding to manual click triggers, and the invisible controls are dedicated to responding to visible and audible voice input triggers, so that the source of the control response trigger can be distinguished in real time to meet diverse needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of an implementation scenario of a voice interaction method in an embodiment;
[0035] Figure 2 is a flowchart of a voice interaction method in one embodiment;
[0036] Figure 3 is a flowchart of a voice interaction method in another embodiment;
[0037] Figure 4 A schematic diagram of implementing the functions of a UI control and a View control in one embodiment;
[0038] Figure 5 is a structural block diagram of a voice interaction device in one embodiment;
[0039] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0041] It is understood that the terms "first", "second", etc. used in this application can be used in this article to describe various professional terms, but unless otherwise specified, these professional terms are not limited by these terms. These terms are only used to distinguish one professional term from another professional term. For example, without departing from the scope of this application, the third preset threshold and the fourth preset threshold can be the same or different.
[0042] At present, most of the in-vehicle smart cockpit systems on the market are equipped with voice assistant functions. Among them, the visible-is-speaking technology refers to the interface elements that can be seen by the user in the in-vehicle system interface. Through voice wake-up and voice input, the corresponding UI controls can be directly controlled and hit, thereby realizing the technology of controlling the interface operation. This technology registers and scans the text on the interface through the Android accessibility service, and when the user wakes up with voice, the voice text entered by the user is matched with the UI controls with the registered related hot words on the current interface, so as to accurately respond to the specified control control.
[0043] In the related art, generally, in the Android layout file, the voice visible command that the UI control needs to support is directly written into the control's android:contentDescription attribute. In this way, when the voice accessibility server scans the interface, it can match the user hot words that the control needs to respond to through the control's android:contentDescription attribute, for example:
[0044] <Button < / button>
[0045] android:id="@+id / pause_button"
[0046] android:src=”@drawable / pause”
[0047] android:contentDescription="Pause" / >
[0048] That is, you can configure this button control so that when the user inputs the "pause" hot word by voice, the control will respond, thereby pausing the music.
[0049] In the above process, there may be multiple defects:
[0050] (1) For controls with multiple functions, it is impossible to distinguish the attribute settings of hot words. For example, the pause button mentioned above needs to respond to the "play" command when pausing. During playback, this button needs to respond to the "pause" command. However, in the layout, only one instruction related to android:contentDescription can be set for this button, resulting in the need to repeatedly reset the hot word attributes of contentDescription dynamically in the code logic according to whether the current audio is playing or paused, resulting in cumbersome logic and prone to errors.
[0051] (2) It is impossible to distinguish between the user's manual click response and the user's visible and spoken voice input response. Since hot words are set, after the voice accessibility service scans and matches successfully, it will directly trigger a click event in response to the hot word. The user's hand click trigger will also trigger the same click event, which makes it impossible to distinguish the triggering method of the click event in the interface, and it is impossible to design targeted interactions for the two triggering methods. For example, if you want to give the user voice TTS feedback in the interaction when the voice triggers a hot word, but there is no feedback when the hand click triggers, you cannot distinguish them.
[0052] (3) It is impossible to accurately locate the trigger position of the visible and speakable animation. Currently, when adding animation to the visible and speakable control response, such as a water ripple animation, it is generally achieved on the voice accessibility service scanning end by extracting and matching the center coordinate position of the hot word control. Generally, the controls on the interface will not implement the animation. One reason is that the above-mentioned defect (2) cannot distinguish whether it is a voice response or a manual click response. The second reason is that each app application has to perform the animation, which is a lot of repetitive work. However, the scanning end can only define an extraction rule in advance to determine the location of the animation, such as extracting the center position. However, based on the user's interactive experience, not all interface designs need to respond to the center position, and thus cannot meet the differentiated animation trigger positions.
[0053] In the related art, the android:contentDescription attribute of the control can also be defined more complexly using a JSON string, so that the response mode of the control can be expanded to be more diverse, for example:
[0054] <Button < / button>
[0055] android:id="@+id / pause_button"
[0056] android:src=”@drawable / pause”
[0057] android:contentDescription="{"func":"open","value":["Collection","Collection Music","My Collection"]}" / >
[0058] In the above code, by specifying the func field as open, the specified hot word can be generalized and expanded to "Open xxx", and the value field can have three hot words. Therefore, the button can respond to three types of voice input, namely "Open Favorites", "Open Favorite Music" and "Open My Favorites". The above process can make the setting of hot words more flexible through the preset JSON rules, and can support multiple hot word responses.
[0059] However, the above process still has the defect mentioned in (3), that is, even if multiple hot word responses can be achieved, it is still impossible to accurately locate the visible and audible dynamic effect trigger position.
[0060] In view of the above problems, the present application provides a voice interaction method, which can be applied between a terminal and a server, and can be specifically applied to: Figure 1 In the application environment shown. Among them, the terminal 102 can obtain the target hot word carried by the voice interaction instruction. Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the hot word. Based on the trigger event of the target invisible control after being triggered by the hot word, a voice interaction request is sent to the server 104, and the server 104 returns the feedback result to the terminal, so that the terminal 102 can execute the corresponding voice interaction process based on the feedback result. Among them, the terminal 102 can be a mobile terminal, a personal computer, a smart TV, a tablet or an s car terminal, etc., and the server 104 can be implemented by a server cluster composed of multiple servers, and the embodiment of the present application does not make specific limitations on this.
[0061] In some embodiments, see Figure 2 , provides a voice interaction method. This method is applied to Figure 1 The terminal 102 in the example is used for explanation. Accordingly, the method can be applied to a voice interaction interface, in which there are matching visible controls and invisible controls, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot words.
[0062] Among them, the voice interaction interface can be a mobile terminal interface or a vehicle terminal interface, and the embodiments of the present application do not specifically limit this. Visible controls refer to controls that users can perceive in the interface through vision, and invisible controls refer to controls that users cannot perceive in the interface through vision. Visible controls are mainly provided to users for hand click responses, while invisible controls are mainly provided to users for voice input responses. Both can correspond to the same trigger event after being triggered. Hot word attributes can be set through contentDescription. In the embodiments of the present application, in order to make the visible control not respond to hot words, the hot word attributes may not be set for the visible control, but the hot word attributes may be set for the invisible control.
[0063] For ease of understanding, the embodiment of the present application takes the vehicle-machine interaction scenario and the voice interaction interface as an example for explanation. Therefore, the method includes the following steps:
[0064] 202. Obtain the target hot word carried by the voice interaction command.
[0065] The voice interaction command can be triggered by the user. After the vehicle terminal obtains the voice interaction command, it can perform voice recognition to obtain the target hot words.
[0066] 204. Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word.
[0067] Specifically, there may be multiple visible controls in the voice interaction interface. For each visible control, at least one matching invisible control may be set, which is not specifically limited in the embodiments of the present application. Among them, "matched" means that the visible control and the invisible control correspond to the same trigger event after being triggered respectively. By matching the target hot word with the hot word attributes of the invisible control on the voice interaction interface through text or direct voice matching, it can be determined which invisible control on the voice interaction interface matches the target hot word.
[0068] 206. Based on the triggering event of the target invisible control after being triggered by the hot word, execute the corresponding voice interaction process.
[0069] Among them, the triggering event of the target invisible control after being triggered by the hot word can be an onClick event, which is not specifically limited in the embodiment of the present application. The corresponding voice interaction process can be realized through the execution logic corresponding to the triggering event.
[0070] The above-mentioned voice interaction method, by setting visible controls and invisible controls that have the same trigger event after being triggered, does not set the hot word attribute for the visible control, and sets the hot word attribute for the invisible control. Since the control carrying the hot word can be separated from the visible control of the interface, when multiple functions are realized for the visible control, multiple matching invisible controls can be set for the visible control, and the hot word attribute of each invisible control can be set separately to realize the corresponding function, thereby improving the cumbersome and error-prone situation caused by repeatedly setting the hot word attribute through code logic.
[0071] In addition, since the controls that carry hot words can be separated from the visible controls on the interface, the visible controls are dedicated to responding to manual click triggers, and the invisible controls are dedicated to responding to visible and audible voice input triggers, so that the source of the control response trigger can be distinguished in real time to meet diverse needs.
[0072] In some embodiments, when a visible control is used to implement multiple functions, the number of functions implemented by the visible control is the same as the number of matching invisible controls, and the hot word attribute of each invisible control that matches the visible control corresponds to a function, and the invisible controls that match the visible controls are superimposed on the voice interaction interface.
[0073] As can be seen from the previous content, each visible control sets multiple matching invisible controls, each invisible control can correspond to a function to be implemented, and the corresponding hot word attributes of each invisible control can be set based on the function to be implemented. In addition, since invisible controls do not need to respond to manual click triggers, in order to save control layout space on the voice interaction interface, invisible controls can be superimposed on the voice interaction interface.
[0074] In the above embodiment, since the control carrying the hot word can be separated from the visible control of the interface, when implementing multiple functions for the visible control, multiple matching invisible controls can be set for the visible control, and the hot word properties of each invisible control can be set separately to achieve the corresponding functions, thereby improving the cumbersome and error-prone situation caused by repeatedly setting the hot word properties through code logic.
[0075] In some embodiments, invisible controls are disabled for touch events.
[0076] In actual implementation, the invisible control may choose to disable touch events or not disable touch events. In the case where touch events are not disabled, the invisible control may be triggered by mistake.
[0077] In the above embodiment, since the touch events of the invisible controls can be disabled, the visible controls can be used exclusively to respond to manual click triggers, and the invisible controls can be used exclusively to respond to visible, i.e., spoken voice input triggers, thereby reducing the possibility of the invisible controls being accidentally triggered and improving the accuracy of the operation.
[0078] In some embodiments, the implementation process of disabling touch events for invisible controls includes:
[0079] For the processing method corresponding to the invisible control for processing touch events, the return value of the processing method is set to false.
[0080] Specifically, the method for processing a touch event may be onTouchEvent. In actual implementation, the onTouchEvent event may be rewritten so that the method returns false. The onTouchEvent method returns true when the event has been completely processed and no other callback method is expected to process the event again, otherwise it returns false.
[0081] In the above embodiment, since the return value of the processing method for processing touch events can be set to false, the visible control can be dedicated to responding to manual click triggers, and the invisible control can be dedicated to responding to visible and spoken voice input triggers, so as to reduce the possibility of the invisible control being accidentally triggered and improve the operation accuracy.
[0082] In some embodiments, the method further comprises:
[0083] In response to the voice interaction instruction, the corresponding animation effect is displayed based on the position of the invisible control in the voice interaction interface.
[0084] The position of each invisible control can be set according to the requirements of the dynamic effect. When the dynamic effect is subsequently realized, the dynamic effect triggering position can be determined based on the position of each invisible control according to the corresponding extraction rule of each invisible control, so as to display the dynamic effect at the corresponding position.
[0085] In the above embodiment, since multiple matching invisible controls can be set for the visible controls, and the position of each invisible control can be set according to the motion effect requirements, instead of extracting the motion effect trigger position based only on the position of the visible control during the voice accessibility service scanning process as before, differentiated motion effect trigger position requirements can be met.
[0086] In some embodiments, the display size of the invisible control is set to 0 to 2 pixel units, and the type of the invisible control is one of the following control types, including a picture control, a text control, and a button control.
[0087] Among them, the display size of the invisible control can be set to 0 to 2 pixel units. In actual implementation, the invisible control can also be set to other display sizes, and the display size makes it invisible to the user. In addition, the type of the invisible control can be a control that can set hot word attributes, and the embodiment of the present application does not specifically limit this.
[0088] In the above embodiment, since the display size can be set to make the control invisible, the interference caused by the control carried by the hot word to the voice interaction interface can be reduced. In addition, since there can be multiple types of controls as invisible controls, the versatility can be improved.
[0089] For ease of understanding, the method flow provided in the embodiment of the present application is described by taking the vehicle system using an Android-related system as an example. Figure 3 , the specific implementation can be shown in the following process steps:
[0090] 302. Add a 0px or 1px View control to the Android layout file, set the android:contentDescription of the View control to the hot word that needs to be responded to, and set the View control to the position where the animation needs to be displayed in the voice interaction interface.
[0091] In addition to using the View control, you can also use other sub-controls under the View class, such as TextView, ImageView, and Button, etc. Since barrier-free scanning does not filter the display size of the View, even if the display size is 0px or 1px, it will not affect the response of the hot word.
[0092] 304. Disable the touch event of the View control, and set the trigger source to be triggered only from the voice hot word so as not to affect the user's manual click.
[0093] The onTouchEvent method of the View control may be rewritten and the method may return false so that the View control is triggered only by a voice hot word.
[0094] 306. The onClick event of the View control and the onClick event of the UI control matching the View control are set as response events consistent with the hot word function, and android:contentDescription is not set for the UI control matching the View control.
[0095] In the case of multiple functions, you can add several 0px View controls and set the hot word attributes of different functions respectively. It should be noted that in actual implementation, you can uniformly define a custom control to encapsulate the function of the hot word button, so that it becomes a universal custom control and the logic does not need to be rewritten. The above process can be referred to Figure 4 ,exist Figure 4 In the , users can trigger functions by manually clicking and speaking when visible. The 0px-sized View1 to View3 controls can be bound to the layout and placement positions between the visible UI controls on the page. Through the above three controls, different hot word responses can be triggered to achieve different voice functions.
[0096] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0097] Based on the same inventive concept, the embodiment of the present application also provides a voice interaction device for implementing the voice interaction method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more voice interaction device embodiments provided below can refer to the limitations on the voice interaction method above, and will not be repeated here.
[0098] In one embodiment, Figure 5 As shown, a voice interaction device is provided, which is applied to a voice interaction interface. There are matching visible controls and invisible controls in the voice interaction interface. The visible controls are not set with hot word attributes. The triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot words. The device includes: an acquisition module 502, a matching module 504 and an execution module 506, wherein:
[0099] An acquisition module 502 is used to acquire a target hot word carried in a voice interaction instruction;
[0100] A matching module 504 is used to match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0101] The execution module 506 is used to execute the corresponding voice interaction process based on the trigger event of the target invisible control after being triggered by the hot word.
[0102] In some embodiments, when a visible control is used to implement multiple functions, the number of functions implemented by the visible control is the same as the number of matching invisible controls, and the hot word attribute of each invisible control that matches the visible control corresponds to a function, and the invisible controls that match the visible controls are superimposed on the voice interaction interface.
[0103] In some embodiments, invisible controls are disabled for touch events.
[0104] In some embodiments, the execution module 506 is further used to set the return value of the processing method corresponding to the invisible control for processing the touch event to false.
[0105] In some embodiments, the execution module 506 is further used to display corresponding animation effects in response to the voice interaction instruction based on the position of the invisible control in the voice interaction interface.
[0106] In some embodiments, the display size of the invisible control is set to 0 to 2 pixel units, and the type of the invisible control is one of the following control types, including a picture control, a text control, and a button control.
[0107] The above-mentioned voice interaction device, by setting visible controls and invisible controls that have the same triggering event after being triggered, does not set the hot word attribute for the visible controls, but sets the hot word attribute for the invisible controls. Since the controls carrying the hot words can be separated from the controls visible in the interface, when multiple functions are implemented for the visible controls, multiple matching invisible controls can be set for the visible controls, and the hot word attributes of each invisible control can be set separately to implement the corresponding functions, thereby improving the cumbersome and error-prone situation caused by repeatedly setting the hot word attributes through code logic.
[0108] In addition, since the controls that carry hot words can be separated from the visible controls on the interface, the visible controls are dedicated to responding to manual click triggers, and the invisible controls are dedicated to responding to visible and audible voice input triggers, so that the source of the control response trigger can be distinguished in real time to meet diverse needs.
[0109] For the specific definition of the voice interaction device, please refer to the definition of the voice interaction method above, which will not be repeated here. Each module in the above-mentioned voice interaction device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0110] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be realized through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a voice interaction method is realized. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0111] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0112] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory and applied to a voice interaction interface, wherein there are matching visible controls and invisible controls in the voice interaction interface, the visible controls are not set with hot word attributes, and a triggering event of the visible controls after being touched is the same as a triggering event of the matching invisible controls after being triggered by a hot word; when the processor executes the computer program, the following steps are implemented:
[0113] Get the target hot words carried in the voice interaction command;
[0114] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0115] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which is applied to a voice interaction interface, in which there are matching visible controls and invisible controls, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot word; when the computer program is executed by a processor, the following steps are implemented:
[0117] Get the target hot words carried in the voice interaction command;
[0118] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0119] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0120] In one embodiment, a computer program product is provided, including a computer program, applied to a voice interaction interface, wherein there are matching visible controls and invisible controls in the voice interaction interface, the visible controls are not set with hot word attributes, and the triggering event of the visible controls after being touched is the same as the triggering event of the matching invisible controls after being triggered by the hot word; when the computer program is executed by a processor, the following steps are implemented:
[0121] Get the target hot words carried in the voice interaction command;
[0122] Match the target hot word with the hot word attribute of the invisible control on the voice interaction interface, and determine the target invisible control that matches the target hot word;
[0123] Based on the trigger event of the target invisible control after being triggered by the hot word, the corresponding voice interaction process is executed.
[0124] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0125] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0126] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A voice interaction method, characterized in that, applied to a voice interaction interface, in which there are matching visible controls and invisible controls. The visible controls do not have hot word attributes, and the trigger event of the visible controls after being touched is the same as the trigger event of the matching invisible controls after being triggered by a hot word; the method includes: Obtaining a target hot word carried by a voice interaction instruction; Matching the target hot word with the hot word attributes of the invisible controls on the voice interaction interface to determine a target invisible control that matches the target hot word; Based on the trigger event of the target invisible control after being triggered by a hot word, executing a corresponding voice interaction process.
2. The method according to claim 1, characterized in that, when the visible control is used to implement multiple functions, the number of functions implemented by the visible control is the same as the number of the matching invisible controls. The hot word attribute of each invisible control matching the visible control corresponds to one function, and the invisible controls matching the visible control are stacked and placed on the voice interaction interface.
3. The method according to claim 1, characterized in that, the touch event of the invisible control is disabled.
4. The method according to claim 3, characterized in that, the implementation process of disabling the touch event of the invisible control includes: For the processing method corresponding to the invisible control for processing touch events, setting the return value of the processing method to false.
5. The method according to claim 1, characterized in that, the method further includes: In response to the voice interaction instruction, based on the position of the invisible control in the voice interaction interface, displaying a corresponding dynamic effect.
6. The method according to any one of claims 1 to 5, characterized in that, the display size of the invisible control is set to 0 to 2 pixel units, and the type of the invisible control is one of the following control types, and the following control types include picture controls, text controls, and button controls.
7. A voice interaction device, characterized in that, the device includes: An obtaining module, configured to obtain a target hot word carried by a voice interaction instruction; A matching module, configured to match the target hot word with the hot word attributes of the invisible controls on the voice interaction interface to determine a target invisible control that matches the target hot word; An execution module, configured to execute a corresponding voice interaction process based on the trigger event of the target invisible control after being triggered by a hot word.
8. A computer device, including a memory and a processor, the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, including a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and system for setting shortcut keys
CN107562482A
Display device and control identification method
CN114282544A
Method and device for controlling speaking window and electronic equipment
CN115641843A
Voice interaction method and voice interaction device
CN116364080A
Method and system for generating and controlling composite user interface control
US20170185422A1
Cited By
Knowable window control method, control device, electronic equipment and vehicle
CN121070303A