Voice control method and device, program product, electronic equipment and storage medium
By obtaining interactive nodes and rendering identification information from the virtual document object model of the terminal device, the problem that the voice interaction function of the terminal device cannot be used for general interface operations is solved, achieving efficient and accurate voice control and improving the user experience.
Patent Information
- Application Number
- CN202410705432.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-02
AI Technical Summary
In existing technologies, the voice interaction function of terminal devices cannot be used for general interface operations, resulting in a poor user experience, as well as low image recognition efficiency and high energy consumption.
By obtaining interactive nodes from the application's Virtual Document Object Model (VDom), rendering identification information within the interface, and using voice control commands to execute operations, the application scenarios of voice interaction functions are expanded.
It enables efficient and accurate voice control of terminal devices in system applications and third-party application interfaces, reduces energy consumption, and improves user experience. In particular, it significantly enhances the application scenarios and user experience of voice interaction in scenarios where touch operation is inconvenient.
Smart Images

Figure CN121053977A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of voice control technology, specifically to a voice control method, device, program product, electronic device, and storage medium. Background Technology
[0002] In recent years, the functions of terminal devices such as smartphones and tablets have become increasingly rich, and their performance has also improved significantly. For example, some terminal devices are now equipped with voice interaction functions, meaning that users can control the terminal device to perform various functions and operations by inputting voice commands. However, in related technologies, the voice interaction function of terminal devices can only be used for specific functions and cannot be used for general interface operations, resulting in many scenarios where voice control cannot be implemented. Summary of the Invention
[0003] To overcome the problems existing in the related technologies, this disclosure provides a voice control method, device, program product, electronic device and storage medium to solve the defects in the related technologies.
[0004] According to a first aspect of the present disclosure, a voice control method is provided, the method comprising:
[0005] Obtain at least one interactive node corresponding to the current interface of the application from the Virtual Document Object Model (VDom) of the application, and render the identification information of the interactive node within the current interface. The application runs on an application framework, the interactive node corresponds to an interface element within the current interface, and the identification information of the interactive node is used to mark the interface element corresponding to the interactive node.
[0006] In response to receiving a voice control command input by the user for the identification information of any interactive node in the current interface, the operation content of the voice control command is executed for the interactive node.
[0007] In one possible embodiment of this disclosure, the method further includes:
[0008] The application's description code is parsed using VDom to obtain VDom node data;
[0009] A VDom model is constructed based on the VDom node data, wherein the VDom model contains multiple VDom nodes distributed in a tree structure, and each VDom node corresponds to a UI element of the application.
[0010] Interactive identification is performed on the VDom node data and / or the VDom model to obtain the interactive nodes contained in the VDom model, and the interactive nodes are labeled in the VDom model.
[0011] In one possible embodiment of this disclosure, the method further includes:
[0012] Interaction identification is performed on the VDom node data and / or the VDom model to obtain the interaction type of the interactive nodes contained in the VDom model;
[0013] In the VDom model, the interaction type of interactive nodes is labeled.
[0014] In one possible embodiment of this disclosure, the method further includes:
[0015] Based on the VDom model, identification information is determined for the interactive nodes corresponding to each interface of the application.
[0016] In one possible embodiment of this disclosure, the method further includes:
[0017] The application's interface is rendered based on the VDom model.
[0018] In one possible embodiment of this disclosure, rendering the identification information of the interactive node within the current interface includes:
[0019] Add an identifier layer on top of the current interface, and render the identifier information of the interactive node at the position corresponding to the interactive node within the identifier layer; or,
[0020] The identification information of the interactive node is rendered at the position corresponding to the interactive node within the current interface.
[0021] In one possible embodiment of this disclosure, the method further includes:
[0022] In response to the application launch, the system monitors the speech recognition results of the system's voice service and performs intent recognition on the speech recognition results to determine whether the speech recognition results contain a wake-up command.
[0023] The step of obtaining the interactive nodes contained in the current interface of the application from the application's Virtual Document Object Model (VDom) and rendering the identification information of each interactive node within the current interface includes:
[0024] In response to the speech recognition result containing a wake-up command, the interactive nodes contained in the current interface of the application are obtained in the application's VDom model, and the identification information of each interactive node is rendered in the current interface.
[0025] In one possible embodiment of this disclosure, after rendering the identification information of each interactive node within the current interface, the method further includes:
[0026] The system monitors the speech recognition results of the system's speech service and performs intent recognition on the speech recognition results to obtain voice control commands.
[0027] In one possible embodiment of this disclosure, obtaining the interactive nodes contained in the current interface of the application from the application's Virtual Document Object Model (VDom) and rendering the identification information of each interactive node within the current interface includes:
[0028] In response to receiving a wake-up command from the system intent module, the system obtains the interactive nodes contained in the current interface of the application in the application's VDom model, and renders the identification information of each interactive node in the current interface. The wake-up command is obtained by the system intent module through intent recognition of the speech recognition results of the system speech service.
[0029] In one possible embodiment of this disclosure, after rendering the identification information of each interactive node within the current interface, the method further includes:
[0030] The system receives voice control commands sent by the system intent module, wherein the voice control commands are obtained by the system intent module through intent recognition of the voice recognition results of the system voice service.
[0031] In one possible embodiment of this disclosure, the operation of executing the voice control command for any of the interactive nodes includes:
[0032] The operation content in response to the voice control command includes native events, and the method interface based on the native events simulates the native events for any of the interactive nodes;
[0033] The operation content in response to the voice control command includes component events, and the component events are simulated for any interactive node based on the component interface of the interactive node.
[0034] According to a second aspect of the present disclosure, a voice control device is provided, the device comprising:
[0035] The interaction discovery module is used to obtain at least one interactive node contained in the current interface of the application in the application's VDom node model, and to render the identification information of the interactive node in the current interface.
[0036] The voice control module is used to respond to a voice control command input by the user for the identification information of any interactive node in the current interface, and to execute the operation content of the voice control command for any interactive node.
[0037] In one possible embodiment of this disclosure, the apparatus further includes a model building module for:
[0038] The application's description code is parsed using VDom to obtain VDom node data;
[0039] A VDom model is constructed based on the VDom node data, wherein the VDom model contains multiple VDom nodes distributed in a tree structure, and each VDom node corresponds to a UI element of the application.
[0040] Interactive identification is performed on the VDom node data and / or the VDom model to obtain the interactive nodes contained in the VDom model, and the interactive nodes are labeled in the VDom model.
[0041] In one possible embodiment of this disclosure, the apparatus further includes a type labeling module for:
[0042] Interaction identification is performed on the VDom node data and / or the VDom model to obtain the interaction type of the interactive nodes contained in the VDom model;
[0043] In the VDom model, the interaction type of interactive nodes is labeled.
[0044] In one possible embodiment of this disclosure, the apparatus further includes an identifier management module for:
[0045] Based on the VDom model, identification information is determined for the interactive nodes corresponding to each interface of the application.
[0046] In one possible embodiment of this disclosure, the device further includes an interface module for:
[0047] The application's interface is rendered based on the VDom model.
[0048] In one possible embodiment of this disclosure, the interaction discovery module is used to, when rendering the identification information of the interaction node within the current interface, perform the following:
[0049] Add an identifier layer on top of the current interface, and render the identifier information of the interactive node at the position corresponding to the interactive node within the identifier layer; or,
[0050] The identification information of the interactive node is rendered at the position corresponding to the interactive node within the current interface.
[0051] In one possible embodiment of this disclosure, the apparatus further includes a listening module for:
[0052] In response to the application launch, the system monitors the speech recognition results of the system's voice service and performs intent recognition on the speech recognition results to determine whether the speech recognition results contain a wake-up command.
[0053] The interactive discovery module is used for:
[0054] In response to the speech recognition result containing a wake-up command, the interactive nodes contained in the current interface of the application are obtained in the application's VDom model, and the identification information of each interactive node is rendered in the current interface.
[0055] In one possible embodiment of this disclosure, the monitoring module is further configured to:
[0056] After rendering the identification information of each interactive node within the current interface, the speech recognition results of the system's speech service are monitored, and intent recognition is performed on the speech recognition results to obtain voice control commands.
[0057] In one possible embodiment of this disclosure, the interaction discovery module is used to:
[0058] In response to receiving a wake-up command from the system intent module, the system obtains the interactive nodes contained in the current interface of the application in the application's VDom model, and renders the identification information of each interactive node in the current interface. The wake-up command is obtained by the system intent module through intent recognition of the speech recognition results of the system speech service.
[0059] In one possible embodiment of this disclosure, the voice control module is used for:
[0060] After rendering the identification information of each interactive node within the current interface, the system receives voice control commands sent by the system intent module. These voice control commands are obtained by the system intent module through intent recognition of the speech recognition results of the system voice service.
[0061] In one possible embodiment of this disclosure, when the voice control module executes the operation content of the voice control command for any of the interactive nodes, it is used to:
[0062] The operation content in response to the voice control command includes native events, and the method interface based on the native events simulates the native events for any of the interactive nodes;
[0063] The operation content in response to the voice control command includes component events, and the component events are simulated for any interactive node based on the component interface of the interactive node.
[0064] According to a third aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0065] According to a fourth aspect of the present disclosure, an electronic device is provided, the electronic device including a memory and a processor, the memory being configured to store computer instructions executable on the processor, and the processor being configured to implement the method of the first aspect when executing the computer instructions.
[0066] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0067] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0068] The voice control method provided in this disclosure can obtain at least one interactive node corresponding to the current interface of the application from the VDom model of the application running on the application framework. That is, the interactive node corresponding to at least one interactive interface element in the current interface, and render the identification information of the interactive node in the current interface to mark the interface element corresponding to the interactive node. This can prompt the user which interface elements in the current interface are interactive, allowing the user to input voice control commands by referring to these interface elements using the identification information. Furthermore, after receiving the user's voice command for the identification information of any interactive node in the current interface, the operation content of the voice control command can be executed for that interactive node. This method controls the application to perform general interface operations through voice control commands, expanding the application scenarios and user experience of voice interaction functions. Moreover, this method uses the VDom model on the application framework to obtain the interactive interface elements in the interface, which is very accurate and efficient. In particular, compared with obtaining interactive interface elements by image recognition of the interface, it can reduce energy consumption and improve efficiency and accuracy. Attached Figure Description
[0069] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0070] Figure 1 This is a schematic diagram illustrating the structure of a cross-platform application framework according to an exemplary embodiment of this disclosure;
[0071] Figure 2 This is a flowchart illustrating a voice control method according to an exemplary embodiment of this disclosure;
[0072] Figure 3 This is a schematic diagram of a voice control process illustrated in an exemplary embodiment of this disclosure;
[0073] Figure 4 This is a schematic diagram of a voice control process illustrated in another exemplary embodiment of this disclosure;
[0074] Figure 5 This is a schematic diagram of a voice control scenario illustrated in another exemplary embodiment of this disclosure;
[0075] Figure 6 This is a schematic diagram of the structure of a voice control device shown in an exemplary embodiment of the present disclosure;
[0076] Figure 7 This is a structural block diagram of an electronic device illustrated in an exemplary embodiment of the present disclosure. Detailed Implementation
[0077] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0078] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0079] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0080] In recent years, the functions of terminal devices such as smartphones and tablets have become increasingly richer, and their performance has also improved significantly. For example, some terminal devices are now equipped with voice interaction functions, meaning users can control the terminal device to perform various functions and operations by inputting voice commands. Voice interaction functions involve many scenarios; broadly speaking, they need to be applicable to specific functions and general interface operations. However, in related technologies, voice interaction functions perform better when applied to specific functions, but perform worse when applied to general interface operations, resulting in a generally mediocre user experience.
[0081] Voice interaction functions for specific purposes can be voice assistants configured on terminal devices. These voice assistants are pre-configured with a number of specific functions, especially those specific to system applications on the terminal device, such as launching an application, setting an alarm for a specific time, or turning silent mode on or off.
[0082] Voice interaction functionality oriented towards user interface (VIE) allows users to interact with controls within an application's interface via voice commands (e.g., touch input) when the application is displayed on a terminal device. However, current VIE-oriented VIE technologies require application adaptation or image recognition of the interface to identify controls, resulting in low efficiency, high computational costs, and inaccurate recognition results.
[0083] Based on this, in a first aspect, at least one embodiment of this disclosure provides a voice control method that enables a terminal device to implement voice interaction functions for generalized interface operations, particularly for the interfaces of system applications or third-party applications; moreover, this method does not require application function adaptation, especially not third-party applications; furthermore, the control recognition of this method is accurate, efficient, and has low power consumption.
[0084] This method can be applied to the application framework level of terminal devices, such as cross-platform application frameworks. Cross-platform application frameworks are installed on different terminals such as smartphones, wearable devices, in-vehicle devices, televisions, and PCs along with the operating system. Applications installed on these terminals only need to be adapted to the cross-platform application framework, without the need to develop separate versions for different terminals, thereby improving the compatibility of applications and reducing redundant application development.
[0085] Cross-platform application frameworks may include the following: Figure 1The application framework and native platform are shown. The application framework includes a template interpreter, which parses the description code of the application running on the cross-platform application framework to obtain VDom (Virtual Document Object Model) node data. The native platform includes a communication module, a VUI (Voice User Interface) control module, a rendering module, and a VUI event triggering module. The communication module is used for data transmission between the application framework and the native platform. The VUI control module is used for interactive recognition of VDom node data and VDom model, querying the interactive nodes corresponding to each interface of the application, and controlling the rendering module to enable or disable the rendering of the identification information of interactive nodes. The rendering module is used to render the identification information of interactive nodes. The VUI event triggering module is used to trigger native events, component events, and OCR services.
[0086] Optional, as shown in the appendix Figure 1 As shown, the native platform includes the system voice service, which can collect ambient voice and perform ASR (Automatic Speech Recognition) recognition, and send the obtained speech recognition result (i.e., the speech recognition result) to the voice control module. The cross-platform application framework also includes a voice control module and an intent recognition module. The voice control module interfaces with the system voice service, and after receiving the speech recognition result from the system voice service, controls the intent recognition module to perform intent recognition on the speech recognition result. When the intent recognition result displays a wake-up command within the speech recognition result, it notifies the VUI to control the rendering module to enable or disable the rendering of the interactive node's identifier information. When the intent recognition result displays a voice control command within the speech recognition result, it sends the voice control command to the VUI event to trigger the execution of the operation. The voice control module and the intent recognition module can be located within the application framework or the native platform.
[0087] Optionally, the native platform includes the system voice service and the system intent module. The system voice service can collect ambient voice and perform ASR (Automatic Speech Recognition) recognition, and send the obtained speech recognition result (i.e., the speech recognition result) to the system intent module for intent recognition. The system intent module then sends the intent recognition result to the voice control module. The cross-platform application framework also includes a voice control module, which interfaces with the system intent module. Upon receiving the intent recognition result from the system intent module, the voice control module responds by displaying a wake-up command within the speech recognition result and notifying the VUI control to start the rendering module to initiate the rendering of the interactive node's identification information. It also responds by sending a voice control command to the VUI event to trigger the execution of the operation. The voice control module can be located within the application framework or the native platform.
[0088] Please refer to the appendix. Figure 2 The example illustrates the flow of a voice control method, including steps S201 to S202.
[0089] In step S201, at least one interactive node corresponding to the current interface of the application is obtained from the application's Virtual Document Object Model (VDom), and the identification information of the interactive node is rendered within the current interface. The application runs on an application framework, the interactive node corresponds to an interface element within the current interface, and the identification information of the interactive node is used to label the interface element corresponding to the interactive node.
[0090] Hereinafter, the Virtual Document Object Model (VDom) will be referred to simply as the VDom model. A cross-platform application framework can run one or more applications, such as Quick Apps; each application within the cross-platform application framework has its own VDom model. The VDom model guides the cross-platform application framework in laying out and rendering the application's interface.
[0091] For example, the application's VDom model can be pre-built as follows:
[0092] First, the application's description code is parsed using VDom to obtain VDom node data.
[0093] For example, the template interpreter within the application framework parses the application's description code (DSL, Domain-Specific Language) and sends the parsed VDOM node data to the native platform. The description code of an application running on the application framework needs to conform to the framework's rules, such as template development. Therefore, the template interpreter within the application framework can parse VDom node data from the application's description code, that is, discover the VDom node corresponding to each UI element of the application, as well as the organizational relationships between these UI elements.
[0094] Next, a VDom model is constructed based on the VDom node data. The VDom model contains multiple VDom nodes distributed in a tree structure, and each VDom node corresponds to a UI element of the application.
[0095] For example, a tree-like topology can be built on the VDom node data using the native platform to construct a tree-like VDom model.
[0096] Finally, interactive identification is performed on the VDom node data and / or the VDom model to obtain the interactive nodes contained in the VDom model, and the interactive nodes are labeled in the VDom model.
[0097] For example, the native platform's VUI control module can be used to perform interaction recognition on each VDom node in the VDom model in at least one of the following optional dimensions:
[0098] Node type: Whether the VDom node is a special component with its own triggering event. If so, it is an interactive node; otherwise, it is a non-interactive node.
[0099] Registered events: Whether the VDom node has events registered by the developer for calling back certain functions. If so, it is an interactive node; otherwise, it is not an interactive node.
[0100] Parent Node Type: Whether the parent node of a VDom node in the VDom model affects whether its child nodes have interactive properties. If so, it is an interactive node; otherwise, it is a non-interactive node. For example, the TabBar component will make its child components switch tabs when clicked.
[0101] Additional attributes: Whether a VDom node has additional attributes that give it extra interactive behavior when it is used as a component. If so, it is an interactive node; otherwise, it is not an interactive node. For example, an Input component of type button has button behavior.
[0102] This example demonstrates the discovery and registration of interactive nodes, allowing us to determine whether a corresponding UI element in the application has interactive functionality by checking if a VDom node in the VDom model is marked as an interactive node. For instance, the native platform's VUI control module can be used to perform the discovery and registration of interactive nodes.
[0103] Preferably, the VDom node data and / or the VDom model can also be interactively identified to obtain the interaction type of the interactive nodes contained in the VDom model; and the interaction type of the interactive nodes can be labeled in the VDom model. For example, when using the VUI control module for interaction identification, it can simultaneously identify whether the VDom node is an interactive node and the type of the interactive node. For example, the interaction type can be click, swipe, text input, etc.
[0104] This preferred example labels the interaction type of the interactive node, so that it can be determined whether the voice control command is a valid command in subsequent voice control, that is, whether the operation content in the command matches the interaction type of the interactive node to which the command is targeted.
[0105] Preferably, identification information can also be determined for the interactive nodes corresponding to each interface of the application based on the VDom model.
[0106] This preferred example implements the identification management of interactive nodes, that is, assigning identification information to the interactive nodes corresponding to each interface as a unit. For example, for each interface in the application that contains interface elements with interactive functions, a label is assigned to at least one interactive node corresponding to the interface to distinguish different interactive nodes within the interface. For example, the management function of interactive nodes is implemented using the VUI control module of the native platform.
[0107] Preferably, the application's interface can also be rendered based on the VDom model. This preferred example represents the basic functionality of the VDom; when the application's interface switches, the native platform's rendering module can render the switched interface.
[0108] As another example, the identification information of the interactive node can be rendered within the current interface in any of the following optional methods:
[0109] Option 1: Add an identifier layer on top of the current interface, and render the identifier information of the interactive node at the position corresponding to the interactive node within the identifier layer. The position corresponding to the interactive node can be a location within the identifier layer used to cover the corresponding interface element.
[0110] This rendering method can be called top-level identifier rendering. The rendering of the identifier is independent of the rendering of the corresponding UI elements of the interactive nodes. It is rendered on an independent layer according to the position of the UI elements. Because the identifier rendering process is independent, this rendering method requires minimal modification to the existing framework, and the cost of updating the identifier is low, without the need to redraw the corresponding UI elements.
[0111] Option 2: Render the identifier information of the interactive node at the location corresponding to the interactive node within the current interface. The location corresponding to the interactive node can be the location of the interface element corresponding to the interactive node.
[0112] This rendering method can be called element-bound rendering, where the identifier is rendered based on the local coordinate system of the interface element when the corresponding interface element of the interactive node is rendered. Because the identifier is bound to the element in terms of hierarchy and position, when the interface element moves, the identifier moves synchronously with it, resulting in a smooth visual performance; when the element is obscured by an element with a higher hierarchy, the element identifier can also be correctly obscured, which meets general expectations.
[0113] It should be understood that this example can utilize the native platform's rendering module to complete the various rendering methods mentioned above. The specific rendering process can employ methods such as on-time rendering checks or additional rendering functions.
[0114] The method of checking during rendering refers to modifying the common rendering method so that when the interface is rendered, it checks whether the interface corresponds to an interactive node. If it does, the identification information of the interactive node is rendered.
[0115] The method of attaching rendering functions refers to calling the interface of the attached rendering function of the interface element, adding the rendering logic of the corresponding interactive node's identification information to the above-mentioned attached rendering function, so that this logic is executed when rendering the interface element to complete the rendering of the identification information.
[0116] In this example, the identification information of the interactive nodes is rendered within the current interface to mark the interface elements corresponding to the interactive nodes. This can prompt the user which interface elements in the current interface are interactive, allowing the user to use the identification information to refer to these interface elements and input voice control commands, thus facilitating voice control.
[0117] In step S202, in response to receiving a voice control command input by the user for the identification information of any interactive node in the current interface, the operation content of the voice control command is executed for any interactive node.
[0118] The voice control command includes at least two parts: the target identification information and the operation content. For example, if a voice control command is "perform a click operation on the interactive node with label 1", then the target identification information is label 1 and the operation content is a click operation.
[0119] For example, the operation content may include native events. The native platform provides method interfaces to simulate such events. When a trigger is needed, the corresponding method interface can be called. For example, a scrollable list on Android can be completed by calling the native scrollBy method.
[0120] In this step, when the operation content of the voice control command is executed for any of the interactive nodes, the operation content of the voice control command may include a native event, and the method interface based on the native event may simulate the native event for any of the interactive nodes.
[0121] For another example, the operation content can include component events. These events are triggered similarly to native events, but their behavior is component-dependent, requiring the component itself to implement a specific interface to expose the triggering method. For instance, the swiping behavior of a carousel component (Swiper) is not simply the movement of component content, but rather switching between content pages. In this case, the native platform may not provide relevant method interfaces to simulate this operation. Therefore, the VDom node corresponding to the Swiper component needs to implement an interface that defines a general swiping method. In the implementation of the Swiper component, the internal pages are switched according to the swiping direction.
[0122] In this step, when the operation content of the voice control command is executed for any of the interactive nodes, the operation content of the voice control command may include component events, and the component events may be simulated for any of the interactive nodes based on the component interface of the interactive node.
[0123] The voice control method provided in this embodiment can obtain at least one interactive node corresponding to the current interface of the application in the VDom model of the application running on the application framework, that is, the interactive node corresponding to at least one interactive interface element in the current interface, and render the identification information of the interactive node in the current interface to mark the interface element corresponding to the interactive node. This can prompt the user which interface elements in the current interface are interactive, so that the user can use the identification information to refer to these interface elements and input voice control commands. Furthermore, after receiving the user's voice command for the identification information of any interactive node in the current interface, the operation content of the voice control command can be executed for the interactive node. This method controls applications to perform generalized interface operations through voice control commands, expanding the application scenarios and user experience of voice interaction. Moreover, this method utilizes the VDom model on the application framework to obtain interactive interface elements within the interface, which is highly accurate and efficient. In particular, compared to obtaining interactive interface elements through image recognition, it can reduce energy consumption and improve efficiency and accuracy. Furthermore, this method does not require the application implementing the interface operation to undergo functional adaptation or separate code development. As long as it can run on a cross-platform application framework, it can be directly controlled by the cross-platform application framework for interface-oriented voice operation.
[0124] This voice control method can significantly improve the user experience, especially in scenarios where touch operation is inconvenient, such as smart speakers and televisions, or when users are unable to operate the terminal device (e.g., people with disabilities, or normal users who are performing other tasks such as laundry or cooking).
[0125] It should be understood that if the native platform has on-device AI (Artificial Intelligence) capabilities, it can introduce the system's OCR service to recognize the text on the interface, thereby clarifying the content of the interface element corresponding to each interaction node. Then, the user only needs to speak the text on the interface element they want to interact with, and the application framework can find the corresponding interface element and trigger the operation event on it. This solution requires the introduced OCR service to be able to recognize the text location and determine the interface element to trigger the click event based on the text location; it also requires the native platform to be able to simulate user click behavior based on absolute screen coordinates.
[0126] In some embodiments of this disclosure, the voice control module interfaces with the system voice service within the native platform. The voice control module then makes the following interface requirements to the system voice service:
[0127] Enable continuous listening: Enabling continuous listening allows users to wake up the voice service within the application.
[0128] Obtain speech recognition results: Extract VUI commands from the text obtained from the speech recognition results.
[0129] Enable continuous dialogue: In continuous dialogue mode, users can have a more seamless interactive experience.
[0130] Dialogue state change callback: Based on the open / closed state of the dialogue, instruct the VUI control to show / hide the identification information of the interactive node.
[0131] Based on the above content in this embodiment, the method may further include: in response to the application startup, monitoring the speech recognition result of the system voice service, and performing intent recognition on the speech recognition result to determine whether the speech recognition result contains a wake-up command.
[0132] For example, the voice control module sends the voice recognition result to the intent module for intent recognition to determine whether the voice recognition result contains a wake-up command.
[0133] Based on the above content in this embodiment, the appendix Figure 2 Step S201 in the illustrated process method may include: in response to the speech recognition result containing a wake-up command, obtaining the interactive nodes contained in the current interface of the application in the application's VDom model, and rendering the identification information of each interactive node in the current interface.
[0134] For example, when the voice control module determines whether the voice recognition result contains a wake-up command, it can send a notification to the native platform's VUI control module so that the VUI control module can control the rendering module to render the identification information of each interactive node corresponding to the current interface according to step S201.
[0135] Based on the above content in this embodiment, the appendix Figure 2 The process method shown may further include, after step S201: listening to the speech recognition result of the system's speech service and performing intent recognition on the speech recognition result to obtain a voice control command.
[0136] For example, the voice control module sends the voice recognition result to the intent module for intent recognition to determine whether the voice recognition result contains a voice control command. If the voice recognition result contains a voice control command, the module recognizes that a voice control command has been received and determines the interactive node targeted by the voice control command and the operation content indicated based on the intent recognition result.
[0137] The appendix can be obtained by combining the above content described in this embodiment. Figure 3 The voice control flow is shown below. The identifier information in this flow is rendered in the upper right corner of the interface elements, hence it will be referred to as a superscript.
[0138] First, the voice control module within the application framework recognizes that after the application starts, it will continuously listen to the system voice service within the native platform.
[0139] Next, the user wakes up the voice service by using a wake-up command (such as a specific wake-up word). The wake-up command is then recognized by the system voice service using ASR. The resulting speech recognition is sent to the voice control module via an activation event callback. The voice control module then notifies the native platform's VUI control module to enable badge rendering. The VUI control module updates the badge and controls the rendering module to render the badge.
[0140] Finally, the user inputs a voice command. The voice command is then recognized by the system's voice service using ASR (Automatic Speech Recognition) and the resulting voice recognition result is sent to the voice control module. The voice control module then responds to the voice command, for example, by performing intent recognition on the voice recognition result and controlling the VUI event triggering module to perform operations based on the intent recognition result.
[0141] In some embodiments of this disclosure, the voice control module interfaces with the system intent module within the native platform. The voice control module then makes the following interface requirements to the system intent module:
[0142] Dialogue state change callback: Based on the open / closed state of the dialogue, instruct the VUI control to show / hide the badges of interactive nodes.
[0143] Based on the above content in this embodiment, the appendix Figure 2 Step S201 in the illustrated process method may include: in response to receiving a wake-up command sent by the system intent module, obtaining the interactive nodes contained in the current interface of the application in the VDom model of the application, and rendering the identification information of each interactive node in the current interface, wherein the wake-up command is obtained by the system intent module performing intent recognition on the speech recognition result of the system speech service.
[0144] For example, the system intent module performs intent recognition on the speech recognition results of the system voice service, and notifies the voice control module when it determines that the speech recognition structure contains a wake-up command; then, the voice control module can send a notification to the native platform's VUI control module so that the VUI control module controls the rendering module to render the identification information of each interactive node corresponding to the current interface according to step S201.
[0145] Based on the above content in this embodiment, the appendix Figure 2The process method shown may further include, after step S201: receiving a voice control command sent by the system intent module, wherein the voice control command is obtained by the system intent module performing intent recognition on the voice recognition result of the system voice service.
[0146] For example, when the system intent module determines that the speech recognition result of the system speech service contains a voice control command, it sends the voice control command to the voice control module so that the voice control module can determine the interactive node targeted by the voice control command and the operation content indicated.
[0147] The appendix can be obtained by combining the above content described in this embodiment. Figure 4 The voice control flow is shown below. The identifier information in this flow is rendered in the upper right corner of the interface elements, hence it will be referred to as a superscript.
[0148] First, the application starts.
[0149] Next, the user wakes up the voice service using a wake-up command (such as a specific wake-up word). The wake-up command is then recognized by the system voice service using ASR (Automatic Speech Recognition). The system intent module then performs intent recognition on the speech recognition result of the system voice service to determine whether the speech recognition result contains a wake-up command. If it does, it will be sent to the voice control module via an activation event callback. The voice control module then notifies the native platform's VUI control module to enable badge rendering. The VUI control module updates the badge and controls the rendering module to render the badge.
[0150] Finally, the user inputs a voice command. The voice command is then recognized by the system's voice service using ASR (Automatic Speech Recognition). The resulting voice recognition result is sent to the system intent module for intent recognition. The system intent module then sends the intent recognition result, which is determined to be a voice command, to the voice control module. The voice control module then responds to the voice command, for example, by controlling the VUI event triggering module to perform an operation based on the aforementioned intent recognition result.
[0151] Next, we will combine the appendix Figure 5 The specific scenario shown will be used to provide a more detailed and specific introduction to the process of this method.
[0152] Please refer to the appendix. Figure 5 In scenario a, the terminal device is playing a video. At this moment, a WeChat notification pops up on the screen: XXX's message "Do you have time to have dinner today?" If the user is unable to operate the device manually at this time, they can say, "XX (the name of the voice assistant), open WeChat." The voice assistant on the terminal device can then perform the specific function of opening WeChat, thereby switching the terminal device's interface to the WeChat chat list.
[0153] Please refer to the appendix. Figure 5In step b, the user can continue by saying, "Student xx, open the first one," where "Student xx" is the wake-up command, "first one" is the label of the interactive node, and "open" is the interface-oriented operation. The cross-platform application framework of the terminal device will then respond to the wake-up command by rendering the label corresponding to each interactive node on the interface, and in response to the voice command, switch the terminal device's interface to the appropriate location. Figure 5 The chat box shown in C; and the cross-platform application framework of the terminal device will render the corresponding label of each interactive node on the interface.
[0154] Please refer to the appendix. Figure 5 In step c, the user can continue by saying "The first one to enter 'OK'", where "first one" is the label of the interaction node and "enter 'OK'" is the interface-oriented operation. The cross-platform application framework on the terminal device will then respond to this voice command by entering "OK" in the text box corresponding to label 1, causing the terminal device's interface to display the corresponding text. Figure 5 The chat box shown in d.
[0155] Please refer to the appendix. Figure 5 In step d, the user can continue by saying "Click the second one," where "second one" is the label of the interactive node, and "click" is the interface-oriented operation. The cross-platform application framework on the terminal device will then respond to this voice command by sending the "OK" text entered in the text box as the dialogue content, causing the terminal device's interface to display the corresponding text. Figure 5 The chat box shown in the middle (e).
[0156] Please refer to the appendix. Figure 5 In the middle of the e, the user did not attach Figure 5 If you input a voice command on the interface shown in the image, the terminal device will switch the interface to the attached screen. Figure 5 The desktop shown in f is now accessible. At this point, the user can continue by saying, "XX (the name of the voice assistant), open XX video." The voice assistant on the terminal device can then perform the specific function of opening XX video (the application that initially played the video), allowing the terminal device to continue playing the video within XX video.
[0157] According to a second aspect of the embodiments of this disclosure, a voice control device is provided; please refer to the appendix. Figure 6 The device includes:
[0158] The interaction discovery module 601 is used to obtain at least one interactive node contained in the current interface of the application in the VDOM node model of the application, and to render the identification information of the interactive node in the current interface.
[0159] The voice control module 602 is used to respond to a voice control command input by a user for the identification information of any interactive node in the current interface, and to execute the operation content of the voice control command for any interactive node.
[0160] In one possible embodiment of this disclosure, the apparatus further includes a model building module for:
[0161] The application's description code is parsed using VDom to obtain VDom node data;
[0162] A VDom model is constructed based on the VDom node data, wherein the VDom model contains multiple VDom nodes distributed in a tree structure, and each VDom node corresponds to a UI element of the application.
[0163] Interactive identification is performed on the VDom node data and / or the VDom model to obtain the interactive nodes contained in the VDom model, and the interactive nodes are labeled in the VDom model.
[0164] In one possible embodiment of this disclosure, the apparatus further includes a type labeling module for:
[0165] Interaction identification is performed on the VDom node data and / or the VDom model to obtain the interaction type of the interactive nodes contained in the VDom model;
[0166] In the VDom model, the interaction type of interactive nodes is labeled.
[0167] In one possible embodiment of this disclosure, the apparatus further includes an identifier management module for:
[0168] Based on the VDom model, identification information is determined for the interactive nodes corresponding to each interface of the application.
[0169] In one possible embodiment of this disclosure, the device further includes an interface module for:
[0170] The application's interface is rendered based on the VDom model.
[0171] In one possible embodiment of this disclosure, the interaction discovery module is used to, when rendering the identification information of the interaction node within the current interface, perform the following:
[0172] Add an identifier layer on top of the current interface, and render the identifier information of the interactive node at the position corresponding to the interactive node within the identifier layer; or,
[0173] The identification information of the interactive node is rendered at the position corresponding to the interactive node within the current interface.
[0174] In one possible embodiment of this disclosure, the apparatus further includes a listening module for:
[0175] In response to the application launch, the system monitors the speech recognition results of the system's voice service and performs intent recognition on the speech recognition results to determine whether the speech recognition results contain a wake-up command.
[0176] The interactive discovery module is used for:
[0177] In response to the speech recognition result containing a wake-up command, the interactive nodes contained in the current interface of the application are obtained in the application's VDom model, and the identification information of each interactive node is rendered in the current interface.
[0178] In one possible embodiment of this disclosure, the monitoring module is further configured to:
[0179] After rendering the identification information of each interactive node within the current interface, the speech recognition results of the system's speech service are monitored, and intent recognition is performed on the speech recognition results to obtain voice control commands.
[0180] In one possible embodiment of this disclosure, the interaction discovery module is used to:
[0181] In response to receiving a wake-up command from the system intent module, the system obtains the interactive nodes contained in the current interface of the application in the application's VDom model, and renders the identification information of each interactive node in the current interface. The wake-up command is obtained by the system intent module through intent recognition of the speech recognition results of the system speech service.
[0182] In one possible embodiment of this disclosure, the voice control module is used for:
[0183] After rendering the identification information of each interactive node within the current interface, the system receives voice control commands sent by the system intent module. These voice control commands are obtained by the system intent module through intent recognition of the speech recognition results of the system voice service.
[0184] In one possible embodiment of this disclosure, when the voice control module executes the operation content of the voice control command for any of the interactive nodes, it is used to:
[0185] The operation content in response to the voice control command includes native events, and the method interface based on the native events simulates the native events for any of the interactive nodes;
[0186] The operation content in response to the voice control command includes component events, and the component events are simulated for any interactive node based on the component interface of the interactive node.
[0187] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments of the method in the first aspect, and will not be elaborated upon here.
[0188] According to a third aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0189] According to the fourth aspect of the embodiments of this disclosure, please refer to the appendix. Figure 7 The diagram illustrates, for example, a block diagram of an electronic device. For instance, device 700 could be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0190] Reference Figure 7 The device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0191] Processing component 702 typically controls the overall operation of device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.
[0192] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0193] The power supply component 706 provides power to the various components of the device 700. The power supply component 706 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 700.
[0194] Multimedia component 708 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0195] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.
[0196] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0197] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or a component of device 700, the presence or absence of user contact with device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0198] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0199] In an exemplary embodiment, device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the voice control method of the aforementioned electronic device.
[0200] Fifthly, in exemplary embodiments, this disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory 704 including instructions, which can be executed by a processor 720 of device 700 to complete the voice control method of the electronic device. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0201] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0202] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A voice control method, characterized in that, The method includes: Obtain at least one interactive node corresponding to the current interface of the application from the Virtual Document Object Model (VDom) of the application, and render the identification information of the interactive node within the current interface. The application runs on an application framework, the interactive node corresponds to an interface element within the current interface, and the identification information of the interactive node is used to mark the interface element corresponding to the interactive node. In response to receiving a voice control command input by the user for the identification information of any interactive node in the current interface, the operation content of the voice control command is executed for the interactive node.
2. The voice control method according to claim 1, characterized in that, The method further includes: The application's description code is parsed using VDom to obtain VDom node data; A VDom model is constructed based on the VDom node data, wherein the VDom model contains multiple VDom nodes distributed in a tree structure, and each VDom node corresponds to a UI element of the application. Interactive identification is performed on the VDom node data and / or the VDom model to obtain the interactive nodes contained in the VDom model, and the interactive nodes are labeled in the VDom model.
3. The voice control method according to claim 2, characterized in that, The method further includes: Interaction identification is performed on the VDom node data and / or the VDom model to obtain the interaction type of the interactive nodes contained in the VDom model; In the VDom model, the interaction type of interactive nodes is labeled.
4. The voice control method according to claim 2, characterized in that, The method further includes: Based on the VDom model, identification information is determined for the interactive nodes corresponding to each interface of the application.
5. The voice control method according to claim 2, characterized in that, The method further includes: The application's interface is rendered based on the VDom model.
6. The voice control method according to claim 1, characterized in that, The step of rendering the identification information of the interactive node within the current interface includes: Add an identifier layer on top of the current interface, and render the identifier information of the interactive node at the position corresponding to the interactive node within the identifier layer; or, The identification information of the interactive node is rendered at the position corresponding to the interactive node within the current interface.
7. The voice control method according to claim 1, characterized in that, The method further includes: In response to the application launch, the system monitors the speech recognition results of the system's voice service and performs intent recognition on the speech recognition results to determine whether the speech recognition results contain a wake-up command. The step of obtaining the interactive nodes contained in the current interface of the application from the application's Virtual Document Object Model (VDom) and rendering the identification information of each interactive node within the current interface includes: In response to the speech recognition result containing a wake-up command, the interactive nodes contained in the current interface of the application are obtained in the application's VDom model, and the identification information of each interactive node is rendered in the current interface.
8. The voice control method according to claim 7, characterized in that, After rendering the identification information of each interactive node within the current interface, the method further includes: The system monitors the speech recognition results of the system's speech service and performs intent recognition on the speech recognition results to obtain voice control commands.
9. The voice control method according to claim 1, characterized in that, The step of obtaining the interactive nodes contained in the current interface of the application from the application's Virtual Document Object Model (VDom) and rendering the identification information of each interactive node within the current interface includes: In response to receiving a wake-up command from the system intent module, the system obtains the interactive nodes contained in the current interface of the application in the application's VDom model, and renders the identification information of each interactive node in the current interface. The wake-up command is obtained by the system intent module through intent recognition of the speech recognition results of the system speech service.
10. The voice control method according to claim 9, characterized in that, After rendering the identification information of each interactive node within the current interface, the method further includes: The system receives voice control commands sent by the system intent module, wherein the voice control commands are obtained by the system intent module through intent recognition of the voice recognition results of the system voice service.
11. The voice control method according to claim 1, characterized in that, The operation content for executing the voice control command for any of the interactive nodes includes: The operation content in response to the voice control command includes native events, and the method interface based on the native events simulates the native events for any of the interactive nodes; The operation content in response to the voice control command includes component events, and the component events are simulated for any interactive node based on the component interface of the interactive node.
12. A voice control device, characterized in that, The device includes: The interaction discovery module is used to obtain at least one interactive node contained in the current interface of the application in the application's VDom node model, and to render the identification information of the interactive node in the current interface. The voice control module is used to respond to a voice control command input by the user for the identification information of any interactive node in the current interface, and to execute the operation content of the voice control command for any interactive node.
13. A computer program product comprising a computer program / instructions, characterized in that, When the computational program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 11.
14. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store computer instructions that can be executed on the processor, and the processor being used to implement the method of any one of claims 1 to 11 when executing the computer instructions.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method of any one of claims 1 to 11.