Interaction method and device based on VGUI, equipment, storage medium and product
By automatically sorting out and registering the initial entries of the target page in the voice user interface, the problem of high entry registration cost in the prior art is solved, and high-quality human-computer interaction is achieved.
Patent Information
- Application Number
- CN202411960020.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the entry registration process for each element in the Voice Graphical User Interface requires developers to manually sort out, resulting in a high registration cost.
A VGUI-based interactive method is provided, which automatically sorts out the initial entries that can be registered in the target page by responding to page update events, and registers these entries to the voice link to realize automated entries registration.
It reduces the registration cost of the initial entry, improves the quality of human-computer interaction, and realizes the visible and easy-to-speak function of each first control in the target page.
Smart Images

Figure CN119993141A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a VGUI-based interaction method, device, equipment, storage medium and product. Background Art
[0002] With the development of voice interaction technology, seeing is speaking has become a development trend. Voice Graphical User Interface (Voice Graphical User Interface) is a way to implement multimodal interaction, combining voice recognition and graphical user interface, allowing users to control various elements on the graphical user interface through voice commands, without manual touch, press and other physical operations, to achieve the visible and speakable function of various elements of the graphical user interface.
[0003] In order to realize the "see and speak" function, the entries of each element need to be registered in the voice interaction system. In the related art, the entry registration process of each element requires developers to sort out, which increases the registration cost of the entry. Summary of the invention
[0004] In order to overcome the problems in the related art, the present disclosure provides an interactive method, device, equipment, storage medium and product based on VGUI, so that when the page is updated to the target page, the initial entries in the target page that can be registered with VGUI can be automatically sorted out, which is beneficial to reducing the registration cost of the initial entries; and the initial entries of the first control are registered to the voice link, which can realize the visible and audible function of each first control in the target page, thereby improving the quality of human-computer interaction.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a VGUI-based interaction method, comprising:
[0006] In response to updating to a target page based on a page update event, obtaining first controls associated with page elements in the target page, and registering initial entries of each of the first controls to a voice link;
[0007] In response to detecting a voice instruction, determining a target word matching the voice instruction from each of the initial words registered to the voice link, and determining a first control corresponding to the target word as a target control;
[0008] Based on the voice instruction, a control function supported by the target control is executed.
[0009] In some embodiments, registering the initial entry of each of the first controls to the voice link includes:
[0010] In response to identifying the first control, determining a type of the first control based on a control function of the first control;
[0011] A mapping relationship between the initial terms and the types of the first controls is established, and based on the mapping relationship, the initial terms and the types of the respective first controls are registered to the voice link.
[0012] In some embodiments, in response to updating to a target page based on a page update event, obtaining a first control associated with a page element in the target page includes:
[0013] In response to updating to the target page, traversing the page elements in the target page to determine candidate controls associated with the page elements;
[0014] Determine the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control;
[0015] When the proportion is greater than a preset threshold, the candidate control is determined as the first control.
[0016] In some embodiments, the method further comprises:
[0017] Generalizing the type registered to the voice link to generate an execution statement;
[0018] generating candidate terms based on the initial terms registered to the voice link and the execution statement;
[0019] In response to detecting a voice instruction, determining a target word matching the voice instruction from each of the initial words registered to the voice link comprises:
[0020] In response to detecting the voice instruction, converting the voice instruction into text content;
[0021] The target term is determined from the candidate terms based on the similarity between the candidate terms and the text content.
[0022] In some embodiments, executing a control function supported by the target control based on the voice instruction includes:
[0023] In the case where the voice instruction includes a task to be performed, determining a current working state of the target control;
[0024] In the case that the current working state of the target control does not match the task to be executed, a target function corresponding to the task to be executed is determined from control functions supported by the target control, and the target function is executed.
[0025] In some embodiments, the method further comprises:
[0026] When the current working state of the target control matches the task to be executed, outputting prompt information;
[0027] The prompt information is used to prompt that the current working status is the same as the working status indicated by the task to be executed.
[0028] In some embodiments, executing a control function supported by the target control based on the voice instruction includes:
[0029] In the case that the voice instruction does not include the task to be executed, executing a preset function supported by the target control;
[0030] The preset function indicates generating a selection event for the target control.
[0031] In some embodiments, determining the first control corresponding to the target entry as the target control includes:
[0032] In response to determining one of the target terms, determining a first type from among the types based on the mapping relationship;
[0033] A first control of the first type in the voice link is determined as the target control.
[0034] In some embodiments, the method further comprises:
[0035] In response to determining at least two of the target terms, determining the second type of each of the target terms from each of the types based on the mapping relationship;
[0036] Based on each of the second types, the target control is determined from first controls of the second type.
[0037] In some embodiments, determining the target control from a first control whose control type is the second type includes:
[0038] In response to the second types being the same, determining the first control with the highest priority indicated by the preset execution policy as the target control; wherein the priorities indicated by the preset execution policy for different first controls are different;
[0039] In response to the second types being different, a first control whose function matches the task to be performed is determined as the target control.
[0040] According to a second aspect of an embodiment of the present disclosure, there is provided a VGUI-based interactive device, comprising:
[0041] A registration module, configured to, in response to updating to a target page based on a page update event, obtain first controls associated with page elements in the target page, and register initial entries of each of the first controls to a voice link;
[0042] A hit module, configured to, in response to detecting a voice instruction, determine a target entry matching the voice instruction from each of the initial entries registered to the voice link, and determine a first control corresponding to the target entry as a target control;
[0043] An execution module is configured to execute a control function supported by the target control based on the voice instruction.
[0044] In some embodiments, the registration module is specifically configured as follows:
[0045] In response to identifying the first control, determining a type of the first control based on a control function of the first control;
[0046] A mapping relationship between the initial terms and the types of the first controls is established, and based on the mapping relationship, the initial terms and the types of the respective first controls are registered to the voice link.
[0047] In some embodiments, the registration module is specifically configured as follows:
[0048] In response to updating to the target page, traversing the page elements in the target page to determine candidate controls associated with the page elements;
[0049] Determine the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control;
[0050] When the proportion is greater than a preset threshold, the candidate control is determined as the first control.
[0051] In some embodiments, the apparatus further comprises:
[0052] A generalization module, configured to perform generalization processing on the type registered to the voice link and generate an execution statement;
[0053] generating candidate terms based on the initial terms registered to the voice link and the execution statement;
[0054] The hit module is further configured as:
[0055] In response to detecting the voice instruction, converting the voice instruction into text content;
[0056] The target term is determined from the candidate terms based on the similarity between the candidate terms and the text content.
[0057] In some embodiments, the execution module is specifically configured as follows:
[0058] In the case where the voice instruction includes a task to be performed, determining a current working state of the target control;
[0059] In the case that the current working state of the target control does not match the task to be executed, a target function corresponding to the task to be executed is determined from control functions supported by the target control, and the target function is executed.
[0060] In some embodiments, the apparatus further comprises:
[0061] A prompt module, configured to output prompt information when the current working state of the target control matches the task to be executed;
[0062] The prompt information is used to prompt that the current working status is the same as the working status indicated by the task to be executed.
[0063] In some embodiments, the execution module is further configured to:
[0064] In the case that the voice instruction does not include the task to be executed, executing a preset function supported by the target control;
[0065] The preset function indicates generating a selection event for the target control.
[0066] In some embodiments, the hit module is further configured to:
[0067] In response to determining one of the target terms, determining a first type from among the types based on the mapping relationship;
[0068] A first control of the first type in the voice link is determined as the target control.
[0069] In some embodiments, the apparatus further comprises:
[0070] a selection module configured to, in response to determining at least two of the target terms, respectively determine a second type of each of the target terms from each of the types based on the mapping relationship;
[0071] Based on each of the second types, the target control is determined from first controls of the second type.
[0072] In some embodiments, the selection module is specifically configured as follows:
[0073] In response to the second types being the same, determining the first control with the highest priority indicated by the preset execution policy as the target control; wherein the priorities indicated by the preset execution policy for different first controls are different;
[0074] In response to the second types being different, a first control whose function matches the task to be performed is determined as the target control.
[0075] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0076] processor;
[0077] Memory for storing computer programs or instructions;
[0078] The processor executes a computer program or instruction to implement the steps of any one of the VGUI-based interaction methods in the first aspect.
[0079] According to a fourth aspect of an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, including:
[0080] When the computer program or instruction in the storage medium is executed by the processor, the steps in any one of the VGUI-based interaction methods in the first aspect described above are implemented.
[0081] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of any one of the VGUI-based interaction methods in the above-mentioned first aspect are implemented.
[0082] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:
[0083] In an embodiment of the present disclosure, in response to updating a target page based on a page update event, a first control associated with a page element in the target page is obtained, and an initial entry of each of the first controls is registered to a voice link; in response to detecting a voice instruction, a target entry matching the voice instruction is determined from the initial entries registered to the voice link, and the first control corresponding to the target entry is determined as the target control; based on the voice instruction, a control function supported by the target control is executed.
[0084] On the one hand, when the page is updated to the target page, the initial entries in the target page that can be registered for VGUI can be automatically sorted out, which reduces the processing burden of developers and helps to reduce the registration cost of the initial entries; on the other hand, registering the initial entries of the first control to the voice link can realize the visible and audible function of each first control in the target page, thereby improving the quality of human-computer interaction.
[0085] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0087] Figure 1 This is a schematic diagram of a VGUI-based interactive method provided by an embodiment of the present disclosure. Figure 1 ;
[0088] Figure 2 is a classification diagram of a control type provided by an embodiment of the present disclosure;
[0089] Figure 3 This is a schematic diagram of a VGUI-based interactive method provided by an embodiment of the present disclosure. Figure 2 ;
[0090] Figure 4 This is a schematic diagram of a VGUI-based interactive method provided by an embodiment of the present disclosure. Figure 3 ;
[0091] Figure 5 is a block diagram of a VGUI-based interactive device provided by an embodiment of the present disclosure;
[0092] Figure 6 It is a structural block diagram of an electronic device 600 provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0093] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.
[0094] Figure 1 This is a schematic diagram of a VGUI-based interactive method provided by an embodiment of the present disclosure. Figure 1 ,like Figure 1 As shown, the interactive method mainly includes the following steps:
[0095] In step 101, in response to updating to a target page based on a page update event, first controls associated with page elements in the target page are obtained, and initial entries of each first control are registered to a voice link;
[0096] In step 102, in response to detecting a voice instruction, a target word matching the voice instruction is determined from each initial word registered to the voice link, and a first control corresponding to the target word is determined as a target control;
[0097] In step 103, based on the voice instruction, a control function supported by the target control is executed.
[0098] It should be noted that the VGUI-based interaction method proposed in the present disclosure can be applied to electronic devices. Here, the electronic device may include: a terminal device, for example, a mobile terminal or a fixed terminal. Among them, the mobile terminal may include: a mobile phone, a tablet computer, a laptop computer, a wearable electronic device and other devices. The fixed terminal may include: a desktop computer, a smart TV, a vehicle-mounted device, etc. In other embodiments, the interaction method may also be applied to an application installed on an electronic device.
[0099] In other embodiments, the VGUI-based interaction method in the disclosed embodiment may be configured in a VGUI-based interaction device, which may be provided in an electronic device, and the disclosed embodiment does not limit this. It should be noted that the execution subject of the disclosed embodiment may be a central processing unit (CPU) in an electronic device in terms of hardware, and may be a related background service in an electronic device in terms of software, and this is not limited.
[0100] Here, a page update event refers to an event related to a change in the content or state of a page, which can be triggered by a user operation or by the internal logic of the system, and is used to implement dynamic updates and interactions of the page. In some embodiments, when a page update event is triggered by a user operation, the page update event can be a click event, such as a user clicking on a page element to implement a page jump; it can also be an input event, such as a user entering content in an input box to implement real-time search of the page. In other embodiments, when a page update event is triggered by the internal logic of the system, the page update event can be a page life cycle event, such as a page has a life cycle, including stages such as creation, mounting, updating or uninstalling, and when the page is in the update stage, the page is dynamically updated; it can also be a page visibility event, such as in order to optimize the user experience and reduce resource loss, when the page visibility event is triggered, it can pause or resume video playback, pause or continue data requests, etc., to implement dynamic updates of the page.
[0101] In the disclosed embodiment, the target page is a new page displayed after the page update event is triggered. The target page can refer to a page of a third-party application, a page of a system tool and application built into the electronic device, or a page of a browser, etc., which is not limited in the disclosed embodiment.
[0102] Here, page elements refer to the various components that make up a web page or application interface, which are used to define the structure, function, appearance, and interaction method of the page.
[0103] In some embodiments, the page elements may be text elements for conveying information to the user, such as titles, links, or pictures; form elements for receiving user input information, such as text boxes or function menus; or interactive elements that allow users to interact, such as applications.
[0104] Here, the first control associated with the page element can be understood as an interactive control for triggering and executing an interactive event associated with the page element. The first control has the characteristics of being operable and viewable on the target page.
[0105] One page element may be associated with one or more first controls. For example, the page element is an application, and the first control is a click control for entering the program interface of the application, a selection control for authorizing the corresponding permissions of the application's functions, and / or a sliding control for executing the preset functions of the application, etc. For another example, the page element is an input box, and the first control is a click control for clicking the input box, and / or a check box selection control for executing multiple preset functions of the text box.
[0106] It needs to be explained that in order to realize the visible and speakable function of the first control, the initial entry of the first control can be registered to the voice link, so that an index or label for voice recognition is established for the first control, so that when a voice command is detected, the target entry corresponding to the voice command can be matched.
[0107] In some embodiments, the voice link includes at least: a first receiving module for receiving voice instructions and the initial entry of the first control registered to the voice link; a processing module for recognizing the voice instruction and preprocessing the voice instruction; a matching module for determining the target entry that matches the voice instruction from the initial entries registered to the voice link; a transmission module for sending the voice instruction to the cloud, and the cloud recognizes the voice instruction, or sending the initial entry registered to the voice link to the cloud, and the cloud performs generalization processing on the initial entry; a second receiving module for receiving the processing result returned by the cloud. Among them, the processing module includes at least a voice recognition unit, a voiceprint recognition unit, a natural speech processing unit, and a speech synthesis unit.
[0108] Here, the initial entry is an identifier of the first control, which is used to characterize what kind of control the first control is or what function the first control implements.
[0109] In some embodiments, when the target page is a page of a system tool or application built into the electronic device, a first association relationship between the text label of the page element in the target page and the text label of the associated first control, as well as a second association relationship between each first control and the text label of the first control are pre-established. When the first control is identified, the text label of the first control is determined based on the first association relationship and the second association relationship, and the text label is determined as the initial entry.
[0110] In other embodiments, when the target page is an application interface of a third-party application, the first control is a third-party control, and the text label corresponding to the third-party control can be searched to determine the initial entry of the third-party control. For example, if the number of text labels is 1, the text label is determined as the initial entry of the third-party control; if the number of text labels is greater than 1, the coordinate information of the text label and the coordinate information of the third-party control are obtained, the target text label is selected from each text label, and the target text label is determined as the initial entry of the third-party control. When the initial entry of the third-party control is determined, a corresponding relationship is established between the initial entry of the third-party control and the third-party control, which is saved in the memory of the electronic device, and the initial entry of the third-party control is sent to the voice link, thereby completing the registration of the entry of the third-party control to the voice link. In this way, the visible and audible function of the third-party control can be realized, and the intelligence of human-computer interaction can be improved.
[0111] Here, when there are multiple first controls, each first control can be registered to the voice link individually according to the recognition order, or can be registered to the voice link in batches, which is not limited in this embodiment of the present disclosure.
[0112] It can be understood that, when a target entry matching a voice instruction is determined from various initial entries, the first control corresponding to the target entry can be determined as the target control, and based on the voice instruction, the control function supported by the target control can be executed.
[0113] In some embodiments, after obtaining the user's voice command, the voice command can be recognized by using technologies such as Automatic Speech Recognition (ASR) to obtain a recognition result. The recognition result is matched with the initial entry in the voice link. When the corresponding target entry is matched, it means that the voice command hits the target entry. The first control corresponding to the target entry can be viewed and the first control is determined as the target control. Finally, according to the voice command, the control function supported by the target control is executed.
[0114] In some embodiments, in order to improve the accuracy of determining the target entry, when a voice instruction is detected, the voice instruction can be converted into text content, and the text content can be preprocessed, such as filtering, word segmentation, etc.; then the key information is extracted from the preprocessed text content; then the similarity between the key information and each initial entry registered in the voice link is determined; and then the initial entry with the greatest similarity is determined as the target entry that matches the voice instruction. Here, there are many ways to determine the similarity between the key information and the initial entry, such as cosine similarity, Jaccard similarity, or word vector similarity.
[0115] Here, the control functions supported by the target control are determined according to the type and properties of the target control. For example, if the target control is a button, the control function is to trigger a click event; if the target control is an input box, the control function is to input text content.
[0116] In some embodiments, in actual applications, there will be entries with the same text associated with different first controls. For example, a sliding control supports sliding to play songs, and a switch control supports clicking to play songs. When the voice command is "play songs", it is determined that there are two target entries that match the voice command. At this time, the first control with the highest priority indicated by the preset execution policy can be determined as the target control, wherein the preset execution policy indicates different priorities for different first controls.
[0117] In an embodiment of the present disclosure, in response to updating a target page based on a page update event, a first control associated with a page element in the target page is obtained, and an initial entry of each of the first controls is registered to a voice link; in response to detecting a voice instruction, a target entry matching the voice instruction is determined from the initial entries registered to the voice link, and the first control corresponding to the target entry is determined as the target control; based on the voice instruction, a control function supported by the target control is executed.
[0118] On the one hand, when the page is updated to the target page, the initial entries in the target page that can be registered for VGUI can be automatically sorted out, which reduces the processing burden of developers and helps to reduce the registration cost of the initial entries; on the other hand, registering the initial entries of the first control to the voice link can realize the visible and audible function of each first control in the target page, thereby improving the quality of human-computer interaction.
[0119] In some embodiments, registering the initial entry of each first control to the voice link includes:
[0120] In response to identifying the first control, determining a type of the first control based on a control function of the first control;
[0121] A mapping relationship between the initial terms and types of the first controls is established, and based on the mapping relationship, the initial terms and types of each first control are registered to the voice link.
[0122] It needs to be explained that, taking into account the differences in users' language habits, the initial terms registered to the voice link can be generalized so that the generalized initial terms match the language habits of different users, so that when matching with voice commands, the target terms matching the voice commands can be accurately located, thereby improving the accuracy of the term hits.
[0123] Since the control function of the first control is related to the user's intention, the type of the first control determines the control function. After generalizing the type of the first control, it can quickly identify and respond to the user's voice instructions for the first control. Therefore, the type of the first control and the initial entry can be registered to the voice link at the same time.
[0124] Here, the type of the first control determines the control function, and different types of first controls have different control functions. For example, the control function of a button control is to trigger an event or perform an operation, such as clicking a button to submit a form or open a new window; the control function of a text box control is to input and display text data, such as inputting information or editing information in a text box; the control function of a list box and a combo box is to display and select list data of a certain list.
[0125] In some embodiments, different control type sets are preset by summarizing the control functions of different controls; after identifying the type of the first control, the type of the first control is matched with the type tags of the different control type sets; when the control type set in which the first control is located is determined, the first control carries the type tag of the corresponding control type set. Here, the first controls in the same control type set carry the same type tag, and the first controls in different control type sets carry different type tags.
[0126] For example, Figure 2 is a classification diagram of a control type provided by an embodiment of the present disclosure, such as Figure 2 As shown, six control type sets are preset, and each control type set includes different sub-sets, and the control functions indicated by different type sets are different. Specifically, the control type sets include: a switchable control set, a checkable control set, a radioable control set, a selectable control set, a slidable control set, and a touchable control set.
[0127] Taking the checkable control set and the selectable control set as examples, the first control in the checkable control set is a control with a check function, such as a checkbox; and the first control in the selectable control set is a control with a switch function, such as a multiswitch control.
[0128] In addition, since the initial entry and type of the first control are registered to the voice link at the same time, after determining the target entry that matches the voice command, the first control corresponding to the target entry can be quickly determined based on the mapping relationship, thereby improving the efficiency of human-computer interaction.
[0129] It is understandable that in order to improve the registration accuracy of the type and initial term of the first control, a mapping relationship between the initial term and type of the first control can be established first, and based on the mapping relationship, the initial term and type can be registered to the voice link at the same time.
[0130] In some embodiments, a mapping table or database is pre-created to store the initial entry, type and mapping relationship between the first control; and a registration interface is set to receive the initial entry and type of the first control; and then a registration logic is set to store the received initial entry and type of the first control in the registration table of the voice link, so as to realize efficient retrieval and matching of the initial entry and type of the first control.
[0131] In some embodiments, after determining the type tag carried by the first control, a mapping relationship between the initial entry and the type tag of the first control is established, and based on the mapping relationship, the initial entry and type tag of each first control are registered to the voice link.
[0132] In the disclosed embodiment, when the type of the first control is determined, a mapping relationship between the initial terms and the type of the first control is established, and based on the mapping relationship, the initial terms and the type of each first control are registered to the voice link. In this way, by classifying the first controls and registering the type and the initial terms to the voice link at the same time, the generalization of the initial terms can be achieved, thereby not only improving the registration accuracy of the term registration process, but also improving the processing efficiency of the term hit process.
[0133] In some embodiments, in response to updating to a target page based on a page update event, obtaining a first control associated with a page element in the target page includes:
[0134] In response to updating to a target page, traversing page elements in the target page to determine candidate controls associated with the page elements;
[0135] Determine the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control;
[0136] When the proportion is greater than a preset threshold, the candidate control is determined as the first control.
[0137] It should be noted that the operable area of the first control is related to the control function of executing the first control. When the proportion of the operable area of the first control in the display area of the first control is unreasonable, even if the target control is identified, invalid execution or execution failure may occur when executing the control function of the first control, thereby reducing the accuracy of human-computer interaction.
[0138] Therefore, in an embodiment of the present disclosure, when an update to a target page is detected, page elements in the target page are traversed, candidate controls associated with the page elements are determined, and the candidate controls are screened to obtain a first control, so that the first control registered to the voice link has an operable area that can execute the control function, thereby improving the quality of human-computer interaction.
[0139] In some embodiments, when there are multiple page elements, each page element can be traversed in turn according to the execution order until the first control associated with each page element is identified.
[0140] It can be understood that in order to improve the accuracy of determining the first control, a preset threshold can be set in advance. When determining the proportion of the operable area indicated by the control attributes of the candidate control in the display area of the candidate control, the proportion will be compared with the preset threshold; and the candidate control whose proportion is greater than the preset threshold will be determined as the first control.
[0141] Here, the preset threshold can be set arbitrarily according to requirements, such as 1 / 2 or 2 / 3, etc., and the embodiments of the present disclosure are not limited to this.
[0142] Exemplarily, the preset threshold is 1 / 2, the initial entry of the candidate control is a playlist, and the type of the candidate control is a sliding control. When the sliding area indicated by the control attribute of the playlist accounts for 1 / 3 of the playlist, it is determined that the playlist is not the first control, and there is no need to register the playlist to the voice link.
[0143] In the disclosed embodiment, the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control is determined, and the candidate control corresponding to the proportion greater than the preset threshold is determined as the first control. In this way, each candidate control can be filtered through the proportion and the preset threshold, thereby reducing the registration of invalid controls and improving the registration accuracy.
[0144] In some embodiments, the method further comprises:
[0145] Generalize the types registered to the voice link and generate execution statements;
[0146] Generating candidate terms based on the initial terms and execution statements registered to the voice link;
[0147] In response to detecting a voice command, determining a target word matching the voice command from each initial word registered to the voice link, including:
[0148] In response to detecting the voice command, converting the voice command into text content;
[0149] Based on the similarity between each candidate term and the text content, a target term is determined from each candidate term.
[0150] It should be noted that in order to improve the accuracy of term hits, the types and terms registered to the voice link can be generalized to obtain candidate terms that match user habits, thereby improving the accuracy of determining target terms that match voice commands and thus improving the response capabilities to voice commands.
[0151] Since the initial entry represents the identifier of the first control, in order to improve the accuracy of generating candidate entries, the type of the first control can be generalized to obtain an execution statement; then the execution statement is associated with the initial entry to obtain a candidate entry that matches the user's habits.
[0152] Here, generalizing the type of the first control includes but is not limited to expanding the control function corresponding to the type, replacing the control function corresponding to the type with synonyms, or generalizing the usage scenario corresponding to the type, generalizing the intent, etc., and the embodiments of the present disclosure are not limited to this.
[0153] Exemplarily, the type of the first control is a sliding control, and the registered type is a sliding control. The sliding control can be generalized to obtain execution statements, such as sliding to make the volume louder, sliding to make the volume lower, sliding to close, or sliding to open, etc.; and the initial entry is a playlist. Based on the initial entry and the execution statement, multiple candidate entries can be generated, such as sliding the playlist to make the volume louder, sliding the playlist to make the volume lower, sliding to close the playlist, or sliding to open the playlist, etc.
[0154] It is understandable that the voice recognition results of voice commands are inaccurate due to factors such as noise, accent or speech speed changes in voice signals. After converting voice commands into text content, it is possible to identify and process ambiguous information in voice commands and improve the semantic recognition accuracy of voice commands.
[0155] At the same time, after converting the voice command into text content, it is convenient to compare the vocabulary, sentence structure and semantics between the text content and the candidate entries, and obtain the similarity between the text content and the candidate entries, thereby improving the accuracy of identifying the target entry.
[0156] In some embodiments, the text content is converted into a single character required for a candidate entry by means of an edit distance (Levenshtein Distance), and the similarity between the text content and the candidate entry can be determined based on the number of edit operations, wherein the edit operation at least includes inserting a character, deleting a character, or replacing a character. When the number of edit operations is small, the smaller the edit distance is, the higher the similarity between the text content and the candidate entry is; and when the number of edit operations is large, the larger the edit distance is, the lower the similarity between the text content and the candidate entry is.
[0157] In other embodiments, the text content and the candidate terms are converted into TF-IDF vectors by means of term frequency-inverse document frequency (TF-IDF) vectors, and then the cosine value of the angle between the two TF-IDF vectors in the vector space is calculated to further determine the similarity between the text content and the candidate terms. For example, when the cosine value is closer to 1, it means that the directions of the two vectors are closer, that is, the similarity between the text content and the candidate terms is higher; and when the cosine value is closer to 0, it means that the directions of the two vectors are perpendicular, that is, the similarity between the text content and the candidate terms is lower.
[0158] In some embodiments, after obtaining the similarities between the text content and each candidate term, the similarities may be sorted, and the candidate term with the greatest similarity among the similarities may be determined as the target term that matches the voice instruction.
[0159] In the disclosed embodiment, the types registered to the voice link are generalized to generate execution statements; and based on the initial terms registered to the voice link and the execution statements, candidate terms are generated to generalize the initial terms so that the generalized candidate terms can match the user's habits; then, when a voice command is detected, the voice command is converted into text content; based on the similarity between each candidate term and the text content, a target term is determined from each candidate term, thereby improving the accuracy of the term hit.
[0160] In some embodiments, based on the voice command, executing a control function supported by the target control includes:
[0161] In the case where the voice command includes a task to be performed, determining the current working state of the target control;
[0162] In the case that the current working state of the target control does not match the task to be executed, the target function corresponding to the task to be executed is determined from the control functions supported by the target control, and the target function is executed.
[0163] Here, the voice command includes tasks to be performed, which refers to specific operations or commands that are recognized and prepared to be executed by the electronic device through voice input; for example, playing the next song, or adjusting the air conditioning temperature to 25 degrees, etc.
[0164] In some embodiments, in the vehicle-computer interaction scenario, the voice command is "turn on Bluetooth", that is, the user's intention is to turn on Bluetooth; if Bluetooth is currently on, and the switch function supported by the Bluetooth control is executed, Bluetooth will be switched from on to off, making the execution result inconsistent with the user's intention, thereby reducing the accuracy of the interaction.
[0165] Therefore, by recognizing the matching degree between the task to be executed and the current working state of the target control, and then executing the voice command, the situation of erroneous operation can be reduced and the quality of human-computer interaction can be improved.
[0166] It can be understood that when the current working state of the target control does not match the task to be executed, it is determined that the current working state of the target control does not match the user's intention. The target function corresponding to the task to be executed can be determined from the control functions supported by the target control, and the target function can be executed, thereby improving the accuracy and efficiency of the interaction.
[0167] Exemplarily, the voice command is "close the song currently playing in the playlist", and the current working status of the playlist is playing songs. It is determined that the task to be executed included in the voice command does not match the current working status of the playlist, so that the pause function can be selected from the supported functions in the playlist and executed, thereby improving the accuracy of human-computer interaction.
[0168] In the embodiment of the present disclosure, when the voice instruction includes a task to be executed, the current working state of the target control is determined; when the current working state of the target control does not match the task to be executed, the target function corresponding to the task to be executed is determined from the control functions supported by the target control, and the target function is executed, thereby improving the accuracy of executing the voice instruction and improving the quality of human-computer interaction.
[0169] In some embodiments, the method further comprises:
[0170] When the current working state of the target control matches the task to be executed, output prompt information;
[0171] The prompt information is used to prompt that the current working status is the same as the working status indicated by the task to be executed.
[0172] It is understandable that when the current working status of the target control matches the task to be executed, there is no need to execute voice commands; in order to improve the user experience, instant feedback can be provided to the user by outputting prompt information, prompting that the current working status of the target control is the same as the working status indicated by the task to be executed.
[0173] Here, the prompt information can be output in the form of voice. For example, when Bluetooth is on and the task to be executed is to turn on Bluetooth, a voice prompt instruction of "Bluetooth is currently on" is output through text synthesis (Text-to-Speech, TTS) technology; it can also be output in the form of text, such as outputting a pop-up window of "Bluetooth is currently on", which is not limited to the embodiments of the present disclosure.
[0174] In the disclosed embodiment, when the current working state of the target control matches the task to be executed, prompt information is output to provide instant feedback to the user, thereby improving the quality of human-computer interaction.
[0175] In some embodiments, based on the voice command, executing a control function supported by the target control includes:
[0176] When the voice command does not include the task to be performed, the preset function supported by the target control is executed;
[0177] The preset function indicates that a selection event is generated for the target control.
[0178] Here, the voice instruction does not include the task to be performed means that the voice instruction lacks a clear indication of the specific task or operation that the electronic device should perform. For example, the voice instruction is "WIFI", and there is no clear indication of the control instruction for WIFI, such as turning on WIFI, turning off WIFI, or connecting WIFI. In this case, the electronic device cannot recognize the task to be performed.
[0179] It is understandable that when the voice command does not include the task to be performed, if the electronic device directly refuses to respond or requires the user to re-enter, it will reduce the user's experience and is not conducive to improving the intelligence of human-computer interaction.
[0180] Therefore, in the embodiment of the present disclosure, when the voice instruction does not include the task to be executed, the preset function supported by the target control is executed, so that a selection event is generated in the target control, and the user's attention is directed to the target control. While prompting the user to modify or expand the operation on the target control through other voice instructions, instant feedback can be output for the current voice instruction, thereby improving the user's experience.
[0181] In the embodiment of the present disclosure, when the voice command does not include the task to be executed, the preset function supported by the target control can be executed to generate a selection event for the target control, thereby providing feedback for the voice command and improving the intelligence of human-computer interaction.
[0182] In some embodiments, determining the first control corresponding to the target entry as the target control includes:
[0183] In response to determining a target term, determining a first type from the types based on the mapping relationship;
[0184] A first control of a first type in the voice link is determined as a target control.
[0185] It is understandable that, since the initial entry and type of the first control are simultaneously registered to the voice link when the initial entry is registered, after a target entry is determined, the first type corresponding to the target entry can be accurately determined from various types based on the mapping relationship.
[0186] In the disclosed embodiment, when a target term is determined, the first type corresponding to the target term can be accurately located based on the mapping relationship; then the first control of the first type is determined as the target control, thereby improving the accuracy of determining the target control.
[0187] In some embodiments, the method further comprises:
[0188] In response to determining at least two target terms, determining the second type of each target term from each type based on the mapping relationship;
[0189] Based on each second type, a target control is determined from among first controls of the second type.
[0190] In some embodiments, if the voice command is "open that" without clear indication of what is to be opened, for example, a music application, a file, or a device light, then the initial entries with the text "open" all belong to the target entry. At this time, there will be multiple target entries, and there will also be multiple first controls corresponding to the target entries.
[0191] It is understandable that, when at least two target terms are determined, the second type corresponding to each target term can be determined from each type based on the mapping relationship, and the first control of the second type in the voice link can be determined. The type of the control reflects the control function of the control. By determining the type of the control, the search range of the target control can be narrowed down, and the first control that does not meet the requirements can be excluded.
[0192] In some embodiments, when the second types are the same, any first control from the first controls of the second types can be selected as the target control; and when the second types are different, the first control among the first controls of the second types whose control function matches the task to be performed can be determined as the target control.
[0193] In the disclosed embodiment, when multiple target terms are determined, the target control is determined from the first controls of each second type by the second type corresponding to each target term. In this way, even if there are duplicate terms, the target control can be selected, thereby improving the accuracy of voice interaction.
[0194] In some embodiments, determining a target control from among first controls whose control type is the second type includes:
[0195] In response to the second types being the same, determining the first control with the highest priority indicated by the preset execution policy as the target control; wherein the preset execution policy indicates different priorities for different first controls;
[0196] In response to the second types being different, a first control whose function matches the task to be performed is determined as a target control.
[0197] It should be noted that when there are multiple target entries, in order to improve the accuracy of determining the target control, two determination methods for the target control are formulated. When the second types are the same, it is determined that the control functions of the first controls of the second types are similar, and the matching degree between the first controls and the user's intention is similar. By determining the priority of the first controls of the second types, the matching degree between the first controls and the user's intention can be predicted, and it can be ensured that the first control that meets the user's intention is selected to execute the voice command.
[0198] Exemplarily, the first controls corresponding to the high beam and the wipers both support three control positions, such as sliding to open, sliding to close, and sliding to automatic. It is determined that the control types of the first controls corresponding to the high beam and the wipers are the same, both being sliding controls.
[0199] In some embodiments, the priority indicated by the preset execution strategy is related to the position at which each first control is displayed on the screen. For example, in a vehicle-computer interaction scenario, the first control corresponding to the high beam is displayed in the upper left corner of the screen, and the first control corresponding to the wiper is displayed in the upper right corner of the screen. Then the priority indicated by the preset execution strategy for the first control corresponding to the high beam is higher than the priority indicated by the preset execution strategy for the first control corresponding to the wiper.
[0200] In other embodiments, the priority indicated by the preset execution policy is related to the environment in which the electronic device is located. For example, when the environment in which the electronic device is located is a dim environment, the priority indicated by the preset execution policy for the first control corresponding to the high beam is higher than the priority indicated by the preset execution policy for the first control corresponding to the wiper. For another example, when the environment in which the electronic device is located is a rainy environment, the priority indicated by the preset execution policy for the first control corresponding to the wiper is higher than the priority indicated by the preset execution policy for the first control corresponding to the high beam.
[0201] It is understandable that when the second types are different, the control functions of the first controls of the second types are different, and the degree of match between the target control determined by the priority indicated by the preset execution strategy and the user's intention cannot be determined. Therefore, in order to improve the accuracy of human-computer interaction, the first control whose control function matches the task to be executed can be determined as the target control.
[0202] For example, the task to be executed in the voice command is "turn the page", and the target terms matching the voice command may be "click on a category of news" and "slide to the next page". Since the user intends to switch the current page to the previous page or the next page, and the function of the control to realize page turning is sliding, the sliding control can be determined as the target control.
[0203] In the embodiment of the present disclosure, when the second types are the same, the first control with the highest priority indicated by the preset execution strategy is determined as the target control; and when the second types are different, the first control whose control function matches the task to be executed is determined as the target control. In this way, by determining the second types, the target control is determined in different ways, thereby improving the accuracy of determining the target control.
[0204] Figure 3 A schematic diagram of a VGUI-based interactive method according to an embodiment of the present disclosure Figure 2 ,like Figure 3 As shown, the interaction method includes at least the following steps:
[0205] In step 301, the target page is updated based on a page update event.
[0206] In step 302, the cache is cleared.
[0207] In step 303, the page elements in the target page are traversed to determine candidate controls associated with the page elements.
[0208] In step 304, it is determined whether the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control is greater than a preset threshold.
[0209] In some embodiments, when it is determined that the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control is greater than a preset threshold, step 305 is executed.
[0210] In other embodiments, when it is determined that the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control is less than or equal to a preset threshold, step 311 is performed.
[0211] In step 305, it is determined whether the candidate control is interactive.
[0212] In some embodiments, if it is determined that the candidate control is interactive, step 306 is performed.
[0213] In other embodiments, when it is determined that the candidate control is not interactive, step 307 is performed.
[0214] In step 306, the candidate control is determined to be the first control, and a mapping relationship between the initial entry and the type of the first control is established.
[0215] In step 307, it is determined whether the candidate control is operable.
[0216] In some embodiments, if it is determined that the candidate control is operable, step 308 is performed.
[0217] In other embodiments, when it is determined that the candidate control is not operable, step 303 is performed.
[0218] In step 308, it is determined whether it is a candidate control binding view.
[0219] In some embodiments, if it is determined to be a candidate control binding view, step 306 is performed.
[0220] In other embodiments, when it is determined that no view is bound to the candidate control, step 309 is executed.
[0221] In step 309, it is determined whether the candidate control has an initial entry.
[0222] In some embodiments, if it is determined that the candidate control has an initial entry, step 306 is performed.
[0223] In other embodiments, when it is determined that the candidate control has an initial entry, step 311 is performed.
[0224] In step 310, based on the mapping relationship, the initial entry and type of the first control are registered to the voice link.
[0225] In step 311, registration is completed.
[0226] In the embodiments of the present disclosure, on the one hand, when a page is updated to a target page, the initial entries in the target page that can be registered with the VGUI can be automatically sorted out, thereby reducing the processing burden on developers and helping to reduce the registration cost of the initial entries; on the other hand, registering the initial entries of the first controls to the voice link can realize the visible-and-speakable function of each first control in the target page, thereby improving the quality of human-computer interaction.
[0227] Figure 4 This is a schematic diagram of a VGUI-based interactive method provided by an embodiment of the present disclosure. Figure 3 ,like Figure 4 As shown, the interaction method includes at least the following steps:
[0228] In step 401 , in response to detecting a voice instruction, the voice instruction is converted into text content.
[0229] In step 402, the text content is preprocessed to generate target content.
[0230] In step 403, the target content is obtained through the callback function.
[0231] In some embodiments, when a mapping relationship between the initial terms and types of the first controls is established, the initial terms and types of each first control are registered to the voice link based on the mapping relationship. Here, the initial terms and types of the first controls can be registered in a storage module of the electronic device, or the initial terms and types of the first controls can be sent to the cloud by the electronic device and registered in the cloud.
[0232] In some embodiments, the text content is preprocessed to generate the target content, which can be executed by the electronic device or by the cloud. Here, the xiaoyund module can receive the target content returned by the electronic device or by the cloud through the callback function.
[0233] In step 404 , it is determined whether the target content includes a task to be performed.
[0234] In some embodiments, when it is determined that the target content includes a task to be performed, step 405 is performed.
[0235] In other embodiments, when it is determined that the target content does not include a task to be performed, step 410 is performed.
[0236] In step 405 , it is determined whether there are at least two target terms matching the target content.
[0237] In some embodiments, when it is determined that there are at least two target terms matching the target content, step 406 is performed.
[0238] In some other embodiments, when it is determined that there is only one target term matching the target content, step 407 is executed.
[0239] In step 406, the second type of each target term is determined from each type based on the mapping relationship.
[0240] In step 407, a first type is determined from various types based on the mapping relationship, and a first control of the first type in the voice link is determined as a target control.
[0241] In step 408, based on each second type, a target control is determined from among the first controls of the second type.
[0242] In step 409, based on the voice instruction, the control function supported by the target control is executed.
[0243] In step 410, a preset function supported by the target control is executed; wherein the preset function indicates generating a selection event for the target control.
[0244] In the disclosed embodiments, by preprocessing the voice commands to obtain the target content, the semantic recognition accuracy of the voice commands is improved; by determining whether the target content includes the task to be executed, the target control is determined in different ways, thereby improving the accuracy of executing the voice commands and improving the quality of human-computer interaction.
[0245] Figure 5 is a block diagram of a VGUI-based interactive device provided by an embodiment of the present disclosure, such as Figure 5 As shown, the interactive device 500 includes:
[0246] A registration module 501 is configured to, in response to updating to a target page based on a page update event, obtain first controls associated with page elements in the target page, and register initial entries of each of the first controls to a voice link;
[0247] The hit module 502 is configured to, in response to detecting a voice instruction, determine a target word matching the voice instruction from each of the initial words registered to the voice link, and determine a first control corresponding to the target word as a target control;
[0248] The execution module 503 is configured to execute the control function supported by the target control based on the voice instruction.
[0249] In some embodiments, the registration module 501 is specifically configured as follows:
[0250] In response to identifying the first control, determining a type of the first control based on a control function of the first control;
[0251] A mapping relationship between the initial terms and the types of the first controls is established, and based on the mapping relationship, the initial terms and the types of the respective first controls are registered to the voice link.
[0252] In some embodiments, the registration module 501 is specifically configured as follows:
[0253] In response to updating to the target page, traversing the page elements in the target page to determine candidate controls associated with the page elements;
[0254] Determine the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control;
[0255] When the proportion is greater than a preset threshold, the candidate control is determined as the first control.
[0256] In some embodiments, the apparatus 500 further includes:
[0257] A generalization module, configured to perform generalization processing on the type registered to the voice link and generate an execution statement;
[0258] generating candidate terms based on the initial terms registered to the voice link and the execution statement;
[0259] The hit module 502 is further configured to:
[0260] In response to detecting the voice instruction, converting the voice instruction into text content;
[0261] The target term is determined from the candidate terms based on the similarity between the candidate terms and the text content.
[0262] In some embodiments, the execution module 503 is specifically configured as follows:
[0263] In the case where the voice instruction includes a task to be performed, determining a current working state of the target control;
[0264] In the case that the current working state of the target control does not match the task to be executed, a target function corresponding to the task to be executed is determined from control functions supported by the target control, and the target function is executed.
[0265] In some embodiments, the apparatus 500 further includes:
[0266] A prompt module, configured to output prompt information when the current working state of the target control matches the task to be executed;
[0267] The prompt information is used to prompt that the current working status is the same as the working status indicated by the task to be executed.
[0268] In some embodiments, the execution module 503 is further configured to:
[0269] In the case that the voice instruction does not include the task to be executed, executing a preset function supported by the target control;
[0270] The preset function indicates generating a selection event for the target control.
[0271] In some embodiments, the hit module 502 is further configured to:
[0272] In response to determining one of the target terms, determining a first type from among the types based on the mapping relationship;
[0273] A first control of the first type in the voice link is determined as the target control.
[0274] In some embodiments, the apparatus 500 further includes:
[0275] a selection module configured to, in response to determining at least two of the target terms, respectively determine a second type of each of the target terms from each of the types based on the mapping relationship;
[0276] Based on each of the second types, the target control is determined from first controls of the second type.
[0277] In some embodiments, the selection module is specifically configured as follows:
[0278] In response to the second types being the same, determining the first control with the highest priority indicated by the preset execution policy as the target control; wherein the priorities indicated by the preset execution policy for different first controls are different;
[0279] In response to the second types being different, a first control whose function matches the task to be performed is determined as the target control.
[0280] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0281] Figure 66 is a block diagram of an electronic device 600 provided by an embodiment of the present disclosure. For example, the electronic device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0282] Reference Figure 6 , the electronic device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output (I / O) interface 612 , a sensor component 614 , and a communication component 616 .
[0283] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with at least one of display, phone calls, data communications, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0284] The memory 604 is configured to store various types of data to support operations on the electronic device 600. Examples of such data include at least one of the following: instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, and videos. The memory 604 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.
[0285] The power supply component 606 provides power to various components of the electronic device 600. The power supply component 606 may include at least one of the following: a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 600.
[0286] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0287] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), and when the electronic device 600 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 604 or sent via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.
[0288] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0289] The sensor assembly 614 includes one or more sensors for providing various aspects of status assessment for the electronic device 600. For example, the sensor assembly 614 can detect the open / closed state of the electronic device 600, the relative positioning of the components, such as the display and keypad of the electronic device 600, and the sensor assembly 614 can also detect the position change of the electronic device 600 or a component in the electronic device 600, the presence or absence of contact between the user and the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and the temperature change of the electronic device 600. The sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or a charge coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include, but is not limited to, at least one of the following: an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, and a temperature sensor.
[0290] The communication component 616 is configured to facilitate wired or wireless communication between the electronic device 600 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as Wi-Fi, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, UWB technology, Bluetooth (BT) technology and other technologies.
[0291] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components.
[0292] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including executable instructions or a computer program, which can be executed by a processor 620 of an electronic device 600 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0293] A non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform any one of the VGUI-based interaction methods described above in the embodiments of the present disclosure. For example, the interaction method includes:
[0294] In response to updating to a target page based on a page update event, obtaining first controls associated with page elements in the target page, and registering initial entries of each first control to a voice link;
[0295] In response to detecting a voice instruction, determining a target word matching the voice instruction from each initial word registered to the voice link, and determining a first control corresponding to the target word as a target control;
[0296] Based on the voice command, execute the control functions supported by the target control.
[0297] The embodiment of the present disclosure provides a computer program product, which includes: a computer program or an executable instruction, which is stored in a computer-readable storage medium. The processor of the computer device reads the computer program or executable instruction from the computer-readable storage medium, and the processor executes the computer program or executable instruction, so that the computer device executes any one of the above-mentioned VGUI-based interaction methods of the embodiment of the present disclosure. After considering the specification and practicing the invention disclosed here, those skilled in the art will easily think of other embodiments of the present disclosure. The present disclosure is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0298] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A VGUI-based interactive method, characterized in that: The method comprises: In response to updating to a target page based on a page update event, obtaining first controls associated with page elements in the target page, and registering initial entries of each of the first controls to a voice link; In response to detecting a voice instruction, determining a target word matching the voice instruction from each of the initial words registered to the voice link, and determining a first control corresponding to the target word as a target control; Based on the voice instruction, a control function supported by the target control is executed.
2. The method according to claim 1, characterized in that The registering the initial entries of each of the first controls to the voice link includes: In response to identifying the first control, determining a type of the first control based on a control function of the first control; A mapping relationship between the initial terms and the types of the first controls is established, and based on the mapping relationship, the initial terms and the types of the respective first controls are registered to the voice link.
3. The method according to claim 1, characterized in that The step of updating the target page based on the page update event and acquiring a first control associated with a page element in the target page includes: In response to updating to the target page, traversing the page elements in the target page to determine candidate controls associated with the page elements; Determine the proportion of the operable area indicated by the control attribute of the candidate control in the display area of the candidate control; When the proportion is greater than a preset threshold, the candidate control is determined as the first control.
4. The method according to claim 2, characterized in that: The method further comprises: Generalizing the type registered to the voice link to generate an execution statement; generating candidate terms based on the initial terms registered to the voice link and the execution statement; In response to detecting a voice instruction, determining a target word matching the voice instruction from each of the initial words registered to the voice link comprises: In response to detecting the voice instruction, converting the voice instruction into text content; The target term is determined from the candidate terms based on the similarity between the candidate terms and the text content.
5. The method according to claim 2 or 4, characterized in that: The executing, based on the voice instruction, a control function supported by the target control includes: In the case where the voice instruction includes a task to be performed, determining a current working state of the target control; In the case that the current working state of the target control does not match the task to be executed, a target function corresponding to the task to be executed is determined from control functions supported by the target control, and the target function is executed.
6. The method according to claim 5, characterized in that The method further comprises: When the current working state of the target control matches the task to be executed, outputting prompt information; The prompt information is used to prompt that the current working status is the same as the working status indicated by the task to be executed.
7. The method according to claim 5, characterized in that The executing, based on the voice instruction, a control function supported by the target control includes: In the case that the voice instruction does not include the task to be executed, executing a preset function supported by the target control; The preset function indicates generating a selection event for the target control.
8. The method according to claim 5, characterized in that The step of determining the first control corresponding to the target entry as the target control includes: In response to determining one of the target terms, determining a first type from among the types based on the mapping relationship; A first control of the first type in the voice link is determined as the target control.
9. The method according to claim 5, characterized in that The method further comprises: In response to determining at least two of the target terms, determining the second type of each of the target terms from each of the types based on the mapping relationship; Based on each of the second types, the target control is determined from first controls of the second type.
10. The method according to claim 9, characterized in that The determining the target control from a first control whose control type is the second type includes: In response to the second types being the same, determining the first control with the highest priority indicated by the preset execution policy as the target control; wherein the priorities indicated by the preset execution policy for different first controls are different; In response to the second types being different, a first control whose function matches the task to be performed is determined as the target control.
11. An interactive device based on VGUI, characterized in that: The device comprises: A registration module, configured to, in response to updating to a target page based on a page update event, obtain first controls associated with page elements in the target page, and register initial entries of each of the first controls to a voice link; A hit module, configured to, in response to detecting a voice instruction, determine a target entry matching the voice instruction from each of the initial entries registered to the voice link, and determine a first control corresponding to the target entry as a target control; An execution module is configured to execute a control function supported by the target control based on the voice instruction.
12. An electronic device, characterized in that: include: processor; Memory for storing computer programs or instructions; The processor executes the computer program or instructions to implement the steps of the method of any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing a computer program or instruction, characterized in that: When the computer program or instructions in the storage medium are executed by a processor, the steps of the method of any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program or instructions, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.