Methods and apparatuses for interaction control, device, medium and program product
By leveraging voice interaction between terminal devices and wearable devices, and utilizing machine learning models and digital assistants, the problem of limited interactive functions in wearable devices has been solved, enabling rich interactive control and efficient function execution, thereby enhancing the user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2026-03-26
AI Technical Summary
Existing wearable devices have limited interactive functions, making it difficult to meet users' needs in various scenarios, and the interactive experience is not rich or convenient enough.
By using voice interaction between terminal devices and wearable devices, machine learning models and digital assistants are utilized to determine and execute target functions based on the current interaction state and voice information, and to provide responses, thereby enabling functional control and information feedback of the terminal devices.
It expands the usage scenarios of terminal devices, improves the convenience and speed of user interaction, meets users' needs in work, study and other aspects, and enhances the interactive experience.
Smart Images

Figure CN2025094973_26032026_PF_FP_ABST
Abstract
Description
Method, apparatus, device, medium and program product for interactive control
[0001] The present application claims priority to the Chinese patent application No. 202411311076.6, filed on September 19, 2024, entitled “Method, apparatus, device, medium and program product for interactive control”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device, computer-readable storage medium and computer program product for interactive control. BACKGROUND
[0003] A wearable device is a portable device that is directly worn on a user or integrated into a user's clothes or accessories. A wearable device is not only a hardware device, but also a device that realizes diversified functions through software support and data interaction and cloud interaction. A wearable device will bring great changes to people's life and perception. People expect that the interaction with a wearable device can handle more functions, so that the wearable device can play a greater role. SUMMARY
[0004] In a first aspect of the present disclosure, a method for interactive control is provided. The method is implemented at a terminal device, comprising: receiving voice information of a user from a wearable device connected to the terminal device; in response to receiving the voice information, determining a target function to be executed at the terminal device based at least on a current interaction state at the terminal device; executing the target function based at least on the voice information; and providing a reply to the voice information at the terminal device, the reply indicating an execution result of the target function.
[0005] In a second aspect of the present disclosure, a method for interactive control is provided. The method is implemented at a wearable device, comprising: collecting voice information of a user at the wearable device; and sending the voice information to a terminal device connected to the wearable device, so that the terminal device determines a target function to be executed at the terminal device based at least on a current interaction state at the terminal device and the voice information.
[0006] In a third aspect of the present disclosure, an apparatus for interaction control is provided. The apparatus includes a receiving module configured to receive voice information of a user from a wearable device connected with a terminal device; a determining module configured to determine, in response to receiving the voice information, a target function to be executed at the terminal device based at least on a current interaction state at the terminal device; an executing module configured to execute the target function based at least on the voice information; and a replying module configured to provide a reply to the voice information at the terminal device, the reply indicating an execution result of the target function.
[0007] In a fourth aspect of the present disclosure, an apparatus for interaction control is provided. The apparatus includes a collecting module configured to collect voice information of a user at a wearable device; and a sending module configured to send the voice information to a terminal device connected with the wearable device, so as to cause the terminal device to determine a target function to be executed at the terminal device based at least on a current interaction state at the terminal device and the voice information.
[0008] In a fifth aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the electronic device to perform the method of the first aspect.
[0009] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon computer-executable instructions that, when executed by a processor, implement the method of the first aspect.
[0010] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0011] It should be understood that the description in this section is not intended to define key or essential features of embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following detailed description in conjunction with the accompanying drawings. In the drawings, like reference numerals denote like elements, in which:
[0013] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0014] FIG. 2 shows a flowchart of a process for interaction control according to some embodiments of the present disclosure;
[0015] FIG. 3 shows a flowchart of a process for interaction control, according to some embodiments of the present disclosure;
[0016] FIG. 4 shows a schematic structural block diagram of an apparatus for interaction control, according to some embodiments of the present disclosure;
[0017] FIG. 5 shows a schematic structural block diagram of an apparatus for interaction control, according to some embodiments of the present disclosure; and
[0018] FIG. 6 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It will be appreciated that the drawings of the present disclosure and the embodiments are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.
[0020] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of", i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0021] In this document, unless explicitly stated otherwise, performing a step "in response to" an event does not mean that the step is performed immediately after the event, but can include one or more intermediate steps.
[0022] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained by appropriate means, wherein the relevant user can include any type of right subject, such as an individual, an enterprise or a group.
[0024] For example, in response to receiving an active request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using information of the relevant user, so that the relevant user can autonomously select whether to provide information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.
[0025] As an optional but non-limiting implementation manner, in response to receiving an active request of a relevant user, the manner of sending a prompt information to the relevant user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0026] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0027] As used herein, the term "model" can learn an association between respective inputs and outputs from training data, such that a corresponding output can be generated for a given input after training is completed. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network", or a "learning network", which terms are used interchangeably herein.
[0028] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.
[0029] Generally, machine learning can include three stages, namely a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to derive consistent inferences from the training data that satisfy an intended objective. Through training, the model can be considered to have learned an association (also referred to as a mapping) from input to output from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to determine whether the model is able to provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the trained parameter values to determine corresponding outputs.
[0030] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, a terminal device 110 has a digital assistant 120 installed therein. The digital assistant 120 can assist a user 140 in processing tasks. The digital assistant 120 can have a conversation and task processing capability with the user.
[0031] In the environment 100, the user 140 can implement an interactive operation by wearing at least one wearable device 130. In some embodiments, the wearable device 130 can include a ring, a watch, a bracelet, a handle, a glove, a finger sleeve, glasses, a brooch, and the like, which can be worn on various parts of the human body. In addition, the wearable device 130 can also be referred to as a wearable interactive device.
[0032] In some embodiments, the user 140 can wake up the digital assistant 120 through the wearable device 130, and input a voice command to the wearable device 130, thereby implementing an interaction with the digital assistant 120. In such embodiments, the wearable device 130 is equipped with a sound receiving device, such as a microphone.
[0033] In some embodiments, the user 140 can interact with the digital assistant 120 via the terminal device 110 and / or an attached device of the terminal device 110. Such an attached device can include, for example, an audio output device (e.g., a headset) for outputting a voice reply corresponding to a voice command. In some embodiments, the terminal device 110 can also present a text reply corresponding to the voice command through a user interface. In some embodiments, the voice reply can also be provided to the wearable device 130 for playing by the wearable device 130. In such a case, the wearable device 130 can be equipped with an audio output device, such as a loudspeaker.
[0034] In some embodiments, the digital assistant 120 can utilize machine learning models (which can include one or more machine learning models) to support the control of the terminal device 110 by the user 140. For example, the digital assistant 120 can utilize one or more machine learning models to provide a question-answering service to the user 140. It should be appreciated that the machine learning models can be different types of models.
[0035] In some embodiments, the terminal device 110 communicates with the service-end device 130 to implement the provision of services of the digital assistant 120 and / or other services at the terminal device 110. The service-end device 130 can invoke a machine learning model to support the interaction between the digital assistant 120 and the user 140 based on the output of the machine learning model. For example, upon receiving a voice command of the user to the digital assistant, the terminal device 110 can request the service-end device 130 to determine or assist in determining a voice reply to the voice command. The service-end device 130 can send the voice reply or related information for determining the voice reply to the terminal device 110.
[0036] The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).
[0037] The service-end device 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like. The service-end device 130 can be implemented based on a cloud environment, for example.
[0038] It should be appreciated that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0039] As mentioned previously, it is expected that the interaction with the wearable device can handle more functions, making the wearable device play a greater role, so as to meet the needs of the user in various scenarios and provide a better interaction experience for the user.
[0040] Embodiments of the present disclosure provide a scheme for interaction control. In the scheme, voice information of a user is received from a wearable device connected with a terminal device. In response to receiving the voice information, a target function to be executed at the terminal device is determined based at least on a current interaction state at the terminal device. The target function is executed based at least on the voice information. A reply to the voice information is provided at the terminal device, the reply indicating an execution result of the target function.
[0041] In this way, the interaction of the user and the wearable device can be combined with the interaction of the user and the terminal device to support the user to conveniently control the execution of various functions through voice interaction. This can expand the use scenarios of the terminal device and fully meet the user's needs for working, learning, and the like using the terminal device. By enabling the user to interact with the wearable device in a voice manner, the user's interaction convenience and speed are effectively improved.
[0042] Some example embodiments of the present disclosure will be described in detail below with reference to examples of the accompanying drawings.
[0043] FIG. 2 shows a flowchart of a process 200 for interaction control according to some embodiments of the present disclosure. For ease of discussion, the process 200 will be described with reference to the environment 100 of FIG. 1. The process 200 can be implemented in the terminal device 110 of FIG. 1.
[0044] At block 210, the terminal device 110 receives voice information of the user 140 from the wearable device 130 connected therewith. In some embodiments, the wearable device 130 can include a ring. In other embodiments, the wearable device 130 can also include a watch, a bracelet, a handle, a glove, a finger sleeve, glasses, a brooch, and the like, which can be worn on various parts of the human body. It should be understood that such wearable devices 130 can belong to smart devices. For example, the ring of the embodiments of the present disclosure can be a smart ring.
[0045] In some embodiments, the communication connection manner of the terminal device 110 and the wearable device 130 includes, but is not limited to, a wireless communication manner such as Bluetooth connection. In some embodiments, the wearable device 130 can be configured with a microphone, a speaker, and the like audio devices, thereby facilitating the provision of voice interaction services for the user 140.
[0046] FIG. 3 shows a flowchart of a process 300 for interaction control according to some embodiments of the present disclosure. For ease of discussion, the process 300 will be described with reference to the environment 100 of FIG. 1. The process 300 can be implemented in the wearable device 130 of FIG. 1.
[0047] At block 310, voice information of the user 140 is collected at the wearable device 130. The voice information of the user 140 can be collected by configuring a microphone for the wearable device 130. The voice information can include, for example, a question, a recording, a voice call, etc. of the user 140. The voice information can also include a relevant voice of the user 140 for waking up the digital assistant 120 at the terminal device 110. Such a relevant voice for waking up the digital assistant 120 can include, for example, a wake-up word, a wake-up sentence, etc. without limitation.
[0048] At block 320, the wearable device 130 sends the voice information to the terminal device 110 connected thereto. Accordingly, the terminal device 110 receives the voice information. The voice information is sent by the wearable device 130 to the terminal device 110 connected thereto, so that the terminal device 110 determines a target function to be performed at the terminal device 110 based at least on a current interaction state at the terminal device 110 and the voice information. The specific manner of determining the target function to be performed at the terminal device 110 will be described in detail from the side of the terminal device 110.
[0049] Referring back to FIG. 2, at block 220, if the terminal device 110 receives the voice information, a target function to be performed at the terminal device 110 is determined based at least on a current interaction state at the terminal device 110. The current interaction state at the terminal device 110 can include, for example, but not limited to, an interaction state of the digital assistant 120, an interaction interface presented by the terminal device 110, an interaction state of an application running at the terminal device 110, etc. In addition, the current interaction state at the terminal device 110 can also include an empty state, i.e., the terminal device 110 is not currently in an interaction state. Specific embodiments of determining a corresponding target function to be performed at the terminal device 110 based on different current interaction states at the terminal device 110 will be described in detail below.
[0050] At block 230, the terminal device 110 performs the target function based at least on the voice information. In some embodiments, when performing the target function based at least on the voice information, the terminal device 110 can perform the target function based on the current interaction state at the terminal device 110 and the voice information. In other words, when the current interaction state at the terminal device 110 is an empty state, the target function can be performed based on the voice information. The empty state can include, for example, that no interaction content is displayed on the interface of the terminal device 110, that no application is running at the terminal device 110, that no configuration is made at the terminal device 110, etc. It should be understood that any appropriate case in which the terminal device 110 is not currently in an interaction state can belong to the empty state. In addition, when the current interaction state at the terminal device 110 is not an empty state, the target function can be performed based on the current interaction state and the voice information.
[0051] At block 240, a reply to the voice information is provided at the terminal device 110, the reply indicating an execution result of the target function. In some embodiments, the reply corresponding text content can be presented at the terminal device 110 through a floating window. In some embodiments, alternatively or additionally, the reply corresponding voice content can be played via the terminal device 110 or an attached device of the terminal device 110. It should be appreciated that the reply to the voice information can also be in any other suitable manner, which is not limited herein.
[0052] In some embodiments, the floating window can be a system tool of the terminal device 110 or a tool of the digital assistant 120. When the digital assistant 120 is used, the floating window can be a movable window floating on the interface of the terminal device 110 to present the reply corresponding text content in the window.
[0053] In some embodiments, the terminal device 110 can be configured with a speaker capable of playing the reply corresponding voice content. In some embodiments, the attached device of the terminal device 110 can include, but is not limited to, an audio output device such as a headset or other speaker, etc. Such an audio output device can be attached to the terminal device 110 in a wired or wireless manner.
[0054] In some embodiments, it can be determined whether the digital assistant 120 at the terminal device 110 is woken up, and the terminal device 110 determines to execute at least one function of the digital assistant 120 in the case that the digital assistant 120 is determined to be woken up. There can be various ways to wake up the digital assistant 120, which will be discussed in detail below.
[0055] In some examples, physical wake-up, breath wake-up, combination of physical wake-up and breath wake-up, somatosensory wake-up, etc. can be employed to detect whether to collect voice information for the digital assistant 120 at the wearable device 130. As an example, the physical wake-up can employ a button on the wearable device 130 to wake up the digital assistant 120 at the terminal device 110. For example, the button can be used to control the microphone switch of the wearable device 130. When it is detected that the user long-presses the button, the microphone is turned on, and at the same time the digital assistant 120 is woken up. In this case, the user 140 can hold the button to speak. The button can employ a physical button, a touch button, a pressure-sensitive button, etc.
[0056] As another example, breath wake-up can be that user 140 puts wearable device 130 to the mouth, wearable device 130 can open the microphone to receive sound in case of detecting human voice at the end side. Cloud VAD (Voice activity detection) can determine whether user 140 finishes speaking. In this case, no wake-up word can be needed to wake up digital assistant 120. Breath wake-up can be implemented through multiple processes, such as: a) two-stage wake-up, sound detection through microphone hardware (detect whether sound is recorded, and start VAD detection after recording); human voice detection through end-side VAD, and start recording. b) three-stage wake-up: recognize lifting hand action through sensor, and then open the microphone; sound detection through microphone hardware; human voice detection through end-side VAD, and start recording. The implementation process here is only an example, and more implementation ways can be designed in practice.
[0057] As yet another example, breath wake-up can be combined with physical wake-up as described above. In this way, breath or key press can individually wake up digital assistant 120, and can be set by the user.
[0058] As another example, body sense wake-up can be to wake up digital assistant 120 by detecting a predetermined gesture of a part of user 140's body. If wearable device 130 is a ring, the predetermined gesture can be that the finger wearing the ring and the thumb pinch twice, and digital assistant 120 can be woken up. It should be noted that in order to protect user privacy, this way can be set to be supported only when terminal device 110 is in an active state (such as an unlocked state, a use state, etc.).
[0059] In some embodiments, digital assistant 120 can have a prompt after being successfully woken up. As an example, a breathing light can be set on wearable device 130. Through different states of the breathing light, such as on-off, flashing, etc., different states of the current voice interaction can be indicated. For example, when the microphone is opened, the breathing light can be always on. When digital assistant 120 is feeding back voice content, the breathing light can flash at a low frequency. The breathing light can also have other conventional functions, such as power display, pairing state display of wearable device 130 and terminal device 110, etc. As another example, different sound prompts can also be set for wearable device 130, so as to respectively indicate the corresponding voice interaction states. As yet another example, a vibration device can also be set for wearable device 130, so as to respectively indicate the corresponding voice interaction states through different vibration prompts. It should be understood that these prompt ways can also be combined.
[0060] Through the above-mentioned example wake-up ways and prompt ways, user can be informed in time about the situation of digital assistant being woken up, and the current voice interaction situation, which helps to improve user experience.
[0061] In some embodiments, if it is detected that the digital assistant 120 is woken up, based on the current active interaction interface presented at the terminal device 110 and the received voice information, the terminal device 110 can determine at least one function of the digital assistant 120 to be executed. In such embodiments, after the digital assistant 120 is woken up, the digital assistant 120 can be used in combination with the current active interaction interface presented at the terminal device 110. In this way, the terminal device 110, by detecting the current active interaction interface presented thereby, helps to determine the interaction intention of the user 140 and further determines at least one function to be executed in combination with the plurality of preset functions of the digital assistant 120.
[0062] In some embodiments, in determining at least one function of the digital assistant 120 to be executed, the terminal device 110 can determine whether the current active interaction interface thereof is a user interface of an application. In such embodiments, the application can be any application in the terminal device 110. It should be understood that such an application can be a system application of the terminal device 110 or an application downloaded by the user 140 at the terminal device 110, etc., without limitation herein. The type of the application can include, for example, an application of an online document, an application of an offline document, an application in which a webpage is located, a video playing application, a social application, etc. If the application is in an active state, the terminal device 110 can present a corresponding user interface through the application.
[0063] In some embodiments, if it is detected that the user interface of the application is not presented at the terminal device 110, the terminal device 110 can determine a first question-and-answer function of the digital assistant 120 to be executed, which can be configured to determine a reply based on the voice information.
[0064] In such embodiments, after the digital assistant 120 is woken up, the user 140 can directly ask the digital assistant 120 of the terminal device 110 by speaking to the wearable device 130 to obtain a corresponding reply. In addition, if the terminal device 110 has an attached audio output device, the corresponding reply can be directly broadcasted through the attached device. If the terminal device 110 does not have an attached device, the corresponding reply can be displayed by the digital assistant 120 on the terminal device 110 and / or played via a built-in speaker of the terminal device 110.
[0065] Taking a specific application scenario as an example, for example, if the user 140 wants to search for content, find information, etc., the user 140 can not manually open the digital assistant 120, but directly ask the wearable device 130, such as "what functions does XX product have", "please help find articles related to XXX", etc. In this way, the user can directly ask in this way to quickly obtain the required content, improving the interactivity.
[0066] In some embodiments, if it is detected that a user interface of an application is rendered at the terminal device 110 and it is detected that at least part of the text in the user interface is selected, the terminal device 110 can determine that one or more processing functions of the digital assistant 120 on the selected at least part of the text are to be performed.
[0067] In such embodiments, different terminal devices 110 can correspond to different ways of selecting text. For example, if the terminal device 110 is a personal computer, it can be determined that the corresponding processing functions of the digital assistant 120 are to be performed by detecting whether part of the text in the user interface is selected via a mouse connected to the terminal device 110 or whether part of the text is selected via a touchpad of the terminal device 110. For example, if the terminal device 110 is a mobile phone, it can be determined that the corresponding processing functions of the digital assistant 120 are to be performed by detecting whether part of the text in the user interface of the touch screen is selected via a touch swipe. It should be understood that the way in which at least part of the text in the user interface of the terminal device 110 is selected can be determined according to actual conditions, which is not limited herein.
[0068] In some embodiments, for at least part of the text selected in the user interface, the processing functions of the digital assistant 120 to be performed can be further determined according to different states of the text. For example, the processing functions of the digital assistant 120 to be performed can be determined according to whether at least part of the text is in an editable state. It should be understood that the processing functions of the digital assistant 120 to be performed can be further determined according to other states of at least part of the text.
[0069] In some embodiments, in determining that one or more processing functions of the digital assistant 120 on the selected text are to be performed, if it is detected that at least part of the text in the user interface is editable, the terminal device 110 can determine that an editing function of the digital assistant 120 on the at least part of the text is to be performed.
[0070] In such embodiments, for detecting the editable state of the selected at least part of the text, it includes but is not limited to detecting an editable control (such as an input box), an editable region, an editable page, an editable window, etc. of the application in which the at least part of the text is located, so as to determine the editable state of the at least part of the text. For example, it can be detected whether the selected at least part of the text is in an input box or whether it is in an editable region of a document application, etc. For the selected editable at least part of the text, it can be determined that a corresponding editing function, such as rewriting, polishing, translation, etc. is to be performed, which is not limited herein.
[0071] For example, when the user 140 is writing or inputting text content, the user 140 can select the text by using a mouse or a touchpad, and then can speak to the wearable device 130 to ask the digital assistant 120 to "polish this paragraph to make it more beautiful", or to ask the digital assistant 120 to perform other editing functions. In this case, the digital assistant 120 can polish the selected text (or perform other editing functions) according to the user's request. For another example, when the user 140 is writing code, the user 140 can select the code by using a mouse or a touchpad, and then can speak to the wearable device 130 to ask the digital assistant 120 to "rewrite this code", or to ask the digital assistant 120 to perform other functions. In this case, the digital assistant 120 can rewrite the selected code (or perform other functions) according to the user's request.
[0072] In some embodiments, if it is detected that at least part of the text in the user interface is not editable, the terminal device 110 can determine to perform a third question-answering function of the digital assistant 120 in combination with the at least part of the text, and the third question-answering function can be configured to determine a reply based on the voice information and the selected at least part of the text.
[0073] In such embodiments, the non-editable state of the selected at least part of the text can be detected in various manners, including but not limited to detecting whether the at least part of the text is in a non-editable area, a non-editable page, a non-editable panel, a non-editable window, or the like of the application, so as to determine the non-editable state of the at least part of the text. For the at least part of the text in the non-editable state, the third question-answering function of the digital assistant 120 can be performed on the at least part of the text. For example, the user 140 can ask a question in combination with the at least part of the text by speaking, so as to obtain a corresponding reply. In other embodiments, other functions of the digital assistant 120 can also be performed on the at least part of the text in the non-editable state.
[0074] For example, in the scenario of browsing a webpage (which can belong to a non-editable page) or reading code, the user 140 can select the text by using a mouse or a touchpad, and then can speak to the wearable device 130 to ask a question, such as "explain the meaning of this word", or "introduce this product", so as to obtain a reply to the question in combination with the selected text by using the digital assistant 120. After obtaining the reply, the user 140 can further ask a question to the wearable device 130, such as "make a sentence with this word", or "what are the advantages of this product compared with similar products on the market", so as to obtain a corresponding reply.
[0075] In some embodiments, after detecting that at least part of the text is selected, the terminal device 110 can present a floating window on the interface of the terminal device 110, and the floating window can display the names of the corresponding processing functions of the digital assistant 120, so that the user 140 can directly select a processing function required by the user 140, and the digital assistant 120 can assist in performing the processing function selected by the user 140.
[0076] By making the digital assistant execute an editing function or a third question and answer function in combination with the selected at least part of the text, the user's editing or question and answer of specific content in the interface can be satisfied, the use scenarios of the wearable device and the digital assistant are enriched, and greater convenience is brought to the user's work, study, and the like.
[0077] In some embodiments, if it is detected that the user interface of the application is presented at the terminal device 110 and no selection of the content in the user interface is detected, the terminal device 110 can determine to execute a second question and answer function of the digital assistant 120 in combination with the content in the user interface, and the second question and answer function can be configured to determine a reply based on the voice information and the content presented in the user interface.
[0078] In such embodiments, the terminal device 110 can determine the second question and answer function of the digital assistant 120 to be executed based on the content of the user interface presented by the currently running application. Such a user interface may, for example, include a picture frame of a video application, a document page of a document application, a web page of a browser, and the like. In some embodiments, such an application can be an application supported by the digital assistant 120. The user 140 can speak to the wearable device 130 based on such a user interface to ask a question, so as to obtain a corresponding reply.
[0079] Taking specific application scenarios as examples, for example, the user 140 watches a release conference of a certain mobile phone on a video application, and can speak to the wearable device 130, saying "jump to the part of the camera of the mobile phone", at which time the digital assistant 120 can jump the release conference video to the corresponding part. For another example, the user 140 can inquire the video content, such as asking "summarize the shooting function of the mobile phone", so as to obtain a corresponding reply. Taking a document application as an example, the reading function of the document application can be combined as a voice input channel, for example, the user 140 can speak to the wearable device 130, asking "summarize the innovation of this paper", so as to obtain a corresponding reply.
[0080] In some embodiments, in determining the target function to be executed at the terminal device 110, if it is detected that the digital assistant 120 is not woken up and the current active interaction interface at the terminal device 110 is an input box, the terminal device 110 can determine to execute a voice-to-text function, and the voice-to-text function can be configured to convert voice information into text and input into the input box.
[0081] In such embodiments, in the case where the digital assistant 120 is not woken up, if it is detected that the input box is focused, for example, the cursor is in the input box, the terminal device 110 can determine that the voice-to-text function is to be performed. In some embodiments, the voice-to-text function can include a voice information filtering function, which can be configured to filter redundant information in the voice information, so that the text corresponding to the redundant information is not input in the input box.
[0082] For example, in a specific application scenario, for example, when the user 140 communicates in the form of text through the terminal device 110 in an office environment, if the communication content is typed in the input box, it can be inefficient and inconvenient. If the user 140 uses the voice input method installed in the terminal device 110 to input text into the input box, the user 140 needs to speak loudly so that the voice information can be collected by the terminal device 110, which can be inconvenient for the user 140 in an office environment, and can cause the user 140 to refuse to use this way. In such a scenario, the user 140 can move the wearable device 110 close to the mouth, so as to speak softly or use a whisper to present the corresponding voice content in the input box of the terminal device 110. In addition, some redundant information in the speech content of the user 140, such as "um" and "ah", which have little practical meaning, can be filtered out. In this way, the converted text content can be more formal and more readable.
[0083] The way used in such embodiments can effectively save the communication time of the user and ensure the quality of the communication content conveyed by the user, thereby improving the user experience.
[0084] In some embodiments, in determining the target function to be performed at the terminal device 110, if it is detected that the digital assistant 120 is not woken up and the terminal device 110 joins a voice call session, the terminal device 110 can determine that the voice call function is to be performed, which can be configured to input voice information as voice input of the voice call session.
[0085] In such embodiments, in the case where the digital assistant 120 is not woken up, if it is detected that the terminal device 110 joins a voice call session, for example, an online meeting, or a voice call of a chat application, a video call containing a voice call, etc., the terminal device 110 can determine that the voice call function is to be performed. In this way, the user 140 can speak to the wearable device 110 as voice input of the voice call session, thereby realizing voice communication.
[0086] With specific application scenarios as examples, for example, when the user 140 participates in an online meeting through the terminal device 110 in an office environment, according to the traditional way, the user 140 may need to speak loudly to ensure that the voice information is collected by the terminal device 110 or its attached device (such as a headset), so as to ensure smooth communication. This will cause inconvenience to the user 140 in the office environment, and may affect the online communication of the user 140. In such a scenario, the user 140 can participate in the meeting by moving the wearable device 110 to the vicinity of the mouth as the default microphone device and speaking softly. Therefore, with the aid of the wearable device 130, the manner of the present embodiment can make the user conveniently make voice calls in various environments, thereby helping the user to improve work and study efficiency, and effectively improving the user experience.
[0087] In some embodiments, when determining the target function to be performed at the terminal device 110, if it is detected that the digital assistant 120 is not woken up and there is no currently active interaction interface at the terminal device 110, the terminal device 110 can determine to perform a control function of the terminal device 110, which can be configured to determine at least one control operation of the terminal device 110 based on the voice information.
[0088] In such embodiments, the objects of the control operation in the terminal device 110 can include but are not limited to files, applications, system configurations, configurations of attached devices, etc. The control operations for these objects can include but are not limited to operations such as opening, closing, querying, searching, setting, etc.
[0089] In some embodiments, the at least one control operation can include one or more of the following: one or more control operations of the operating system of the terminal device 110, or one or more custom instructions, each of which can indicate a corresponding trigger word and include one or more control operations of the terminal device 110.
[0090] In such embodiments, the one or more control operations of the operating system of the terminal device 110 may, for example, include control to open a file, open an application, open a do-not-disturb mode, etc. It should be understood that such one or more control operations can belong to the operation at the basic system level.
[0091] With specific application scenarios as examples, for example, when a user wants to search for a file from the terminal device 110, if the storage path of the file is forgotten or it is troublesome to find the file, the user can speak to the wearable device 130 to find the file. For example, the user 140 can say to the wearable device 130 “help me find a file, which is about the courseware of XXX”, so that the courseware can be presented on the interface of the terminal device 110. For another example, the user 140 can speak to the wearable device 110 to request to open an application, such as directly saying “open XXX”, so that the page after the application is opened can be presented on the interface of the terminal device 110.
[0092] In addition, for a custom instruction, a custom trigger word can be set, and based on the trigger word, a single control operation or a series of control operations on the terminal device 110 can be triggered. Such a custom instruction can also be referred to as a shortcut instruction. The user 140 can speak such a trigger word to the wearable device 130, so that the terminal device 110 performs one or more corresponding control operations.
[0093] With specific application scenarios as examples, for example, the user 140 can speak the trigger word “start work” to the wearable device 130, and the terminal device 110 can perform a series of control operations based on the trigger word, such as including: opening the do-not-disturb mode, opening a certain enterprise application, starting to play light music, closing a certain chat application, etc. For another example, the user 140 can speak the trigger word “start reading news” to the wearable device 130, and the terminal device 110 can perform a series of control operations based on the trigger word, such as including: creating a new desktop, a browser simultaneously opening multiple news sources, etc.
[0094] Through the control function of the terminal device, the user can more conveniently and quickly control the content in the terminal device, thereby helping the terminal device to better serve the user.
[0095] In some embodiments, when determining the target function to be executed at the terminal device 110, based on the matching degrees of the current interaction state and the plurality of functions respectively and the priorities of the plurality of functions respectively, the terminal device 110 can determine the target function from the plurality of functions.
[0096] In such embodiments, for each function discussed above, if the current interaction state of the terminal device 110 can match multiple functions, the target function to be executed can be determined according to the priority of each of the multiple functions. In some embodiments, the function with the highest priority among the multiple functions to which the current interaction state can match can be determined as the target function in order of the priority of the multiple functions from high to low. The multiple functions to which the current interaction state can match have been discussed in detail above and will not be repeated here. The priority of each function can be pre-set according to the actual use of the terminal device 110, so as to determine the target function faster according to the set priority order when matching.
[0097] For example, the function corresponding to the case where the digital assistant 120 of the terminal device 110 needs to be woken up can have a first priority. The speech-to-text function can have a second priority. The voice call function can have a third priority. The control function of the terminal device 110 can have a fourth priority. It should be understood that such priority arrangement is only an example and is not intended to be any limitation.
[0098] Through embodiments of the present disclosure, the user can interact with the terminal device through the interaction of the user with the wearable device, thereby expanding the use scenario of the terminal device and fully meeting the user's needs for working, learning, and the like using the terminal device. By enabling the user to interact with the wearable device in a voice manner, the user's interaction convenience and speed are effectively improved. In addition, by determining the target function to be executed at the terminal device based on the current interaction state at the terminal device, a high-quality reply to the voice information can be obtained, and the user experience is effectively improved.
[0099] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 4 shows a schematic structural block diagram of an apparatus 400 for interaction control according to some embodiments of the present disclosure. The apparatus 400 may, for example, be implemented in or included in the terminal device 110. Each module / component in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.
[0100] As shown, the apparatus 400 can include a receiving module 410 configured to receive voice information of a user from a wearable device connected with a terminal device; a determining module 420 configured to, in response to receiving the voice information, determine a target function to be executed at the terminal device based at least on a current interaction state at the terminal device; an executing module 430 configured to execute the target function based at least on the voice information; and a replying module 440 configured to provide a reply to the voice information at the terminal device, the reply indicating an execution result of the target function.
[0101] In some embodiments, the determining module 420 is further configured to determine the target function from the plurality of functions based on the matching degree of the current interaction state with each of the plurality of functions and the priority of each of the plurality of functions.
[0102] In some embodiments, the performing module 430 is further configured to perform the target function based on the current interaction state at the terminal device and the voice information.
[0103] In some embodiments, the determining module 420 is further configured to determine whether a digital assistant at the terminal device is woken up; and in response to detecting that the digital assistant is woken up, determine the at least one function of the digital assistant to be performed based on a current active interaction interface presented at the terminal device.
[0104] In some embodiments, the determining module 420 is further configured to determine whether a current active interaction interface at the terminal device is a user interface of an application; in response to detecting that the user interface of the application is not presented at the terminal device, determine a first question-and-answer function of the digital assistant to be performed, the first question-and-answer function being configured to determine a reply based on the voice information; in response to detecting that the user interface of the application is presented at the terminal device and detecting that at least part of text in the user interface is selected, determine one or more processing functions of the digital assistant on the at least part of text selected to be performed; and in response to detecting that the user interface of the application is presented at the terminal device and detecting that no selection on content in the user interface, determine a second question-and-answer function of the digital assistant in combination with the content in the user interface to be performed, the second question-and-answer function being configured to determine a reply based on the voice information and the content presented in the user interface.
[0105] In some embodiments, the determining module 420 is further configured to, in response to detecting that the at least part of text in the user interface is editable, determine an editing function of the digital assistant on the at least part of text to be performed; and in response to detecting that the at least part of text in the user interface is not editable, determine a third question-and-answer function of the digital assistant in combination with the at least part of text to be performed, the third question-and-answer function being configured to determine a reply based on the voice information and the at least part of text selected.
[0106] In some embodiments, the determining module 420 is further configured to, in response to detecting that the digital assistant is not woken up and that a current active interaction interface at the terminal device is an input box, determine a voice-to-text function to be performed, the voice-to-text function being configured to convert the voice information into text and input into the input box.
[0107] In some embodiments, the determining module 420 is further configured to, in response to detecting that the digital assistant is not woken up and that the terminal device joins a voice call session, determine a voice call function to be performed, the voice call function being configured to input the voice information as voice of the voice call session.
[0108] In some embodiments, the determining module 420 is further configured to determine, in response to detecting that the digital assistant is not woken up and there is no current active interaction interface at the terminal device, a control function to be performed on the terminal device, the control function being configured to determine at least one control operation on the terminal device based on the voice information.
[0109] In some embodiments, the at least one control operation comprises one or more of the following: one or more control operations on an operating system of the terminal device, or one or more custom instructions, each of which indicates one or more control operations on the terminal device configured to have a corresponding trigger word and include.
[0110] In some embodiments, the replying module 440 is further configured to present the reply to the corresponding text content through a floating window at the terminal device, and play the reply to the corresponding voice content via the terminal device or an attached device of the terminal device.
[0111] In some embodiments, the wearable device comprises a ring.
[0112] FIG. 5 shows a schematic structural block diagram of an apparatus 500 for interaction control, according to some embodiments of the present disclosure. The apparatus 500 may, for example, be implemented in or included in the wearable device 130. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0113] As shown, the apparatus 500 can include a collecting module 510 configured to collect voice information of a user at the wearable device, and a sending module 520 configured to send the voice information to a terminal device connected with the wearable device, to enable the terminal device to determine a target function to be performed at the terminal device based at least on a current interaction state at the terminal device and the voice information.
[0114] The units and / or modules included in the apparatus 400, the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition or as an alternative, part or all of the units and / or modules in the apparatus 400, the apparatus 500 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0115] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 illustrated in FIG. 6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 illustrated in FIG. 6 can include or be implemented as the terminal device 110 or the wearable device 130 of FIG. 1, or the apparatus 400 of FIG. 4 or the apparatus 500 of FIG. 5.
[0116] As illustrated in FIG. 6, the electronic device 600 is in the form of a general electronic device. Components of the electronic device 600 can include, but are not limited to, one or more processors 610 or processing units, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 620. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve parallel processing capability of the electronic device 600.
[0117] The electronic device 600 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 600 and includes both volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable media and can include a machine-readable medium, such as a flash drive, a disk drive, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 600.
[0118] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive and disk drive interface can be provided for reading from or writing to a removable, non-volatile optical disk (e.g., a "CD-ROM" or "DVD"). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0119] The communication unit 640 enables communication through communication media with other electronic devices. Additionally, the functionality of the components of the electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. As such, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0120] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 640, as necessary, with one or more devices that enable a user to interact with the electronic device 600, or with any device (e.g., a network card, a modem, etc.) that enables the electronic device 600 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0121] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0122] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0123] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0124] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0125] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions (s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in some cases, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0126] implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The herein described subject matter is to be considered in all its novel implications and applications and can be practiced in a variety of environments. It is also to be understood that such features are only examples and that many modifications, additions, or omissions can be made without departing from the scope of the claimed disclosure.
Claims
1. An interactive control method implemented at a terminal device, comprising: receiving speech information of a user from a wearable device connected with the terminal device; in response to receiving the speech information, determining a target function to be executed at the terminal device based at least on a current interaction state at the terminal device; executing the target function based at least on the speech information; and providing a reply to the speech information at the terminal device, the reply indicating an execution result of the target function.
2. The method of claim 1, wherein determining the target function to be executed at the terminal device comprises: determining the target function from a plurality of functions based on a matching degree of the current interaction state with each of the plurality of functions and a priority of each of the plurality of functions.
3. The method of claim 1, wherein executing the target function based at least on the speech information comprises: executing the target function based on the current interaction state at the terminal device and the speech information.
4. The method of claim 1, wherein determining the target function to be executed at the terminal device comprises: determining whether a digital assistant at the terminal device is woken up; and in response to detecting that the digital assistant is woken up, determining at least one function of the digital assistant to be executed based on a current active interaction interface presented at the terminal device.
5. The method of claim 4, wherein determining the at least one function of the digital assistant to be executed comprises: determining whether the current active interaction interface at the terminal device is a user interface of an application; in response to detecting that no user interface of an application is presented at the terminal device, determining a first question-answering function of the digital assistant to be executed, the first question-answering function being configured to determine the reply based on the speech information; in response to detecting that a user interface of an application is presented at the terminal device and detecting that at least a portion of text in the user interface is selected, determining one or more processing functions of the digital assistant on the selected at least portion of text to be executed; in response to detecting that a user interface of an application is presented at the terminal device and detecting no selection on content in the user interface, determining a second question-answering function of the digital assistant in combination with the content in the user interface to be executed, the second question-answering function being configured to determine the reply based on the speech information and the content presented in the user interface.
6. The method of claim 5, wherein determining the one or more processing functions of the digital assistant on the selected text to be executed comprises: in response to detecting that the at least portion of text in the user interface is editable, determining an editing function of the digital assistant on the at least portion of text to be executed; in response to detecting that the at least portion of text in the user interface is not editable, determining a third question-answering function of the digital assistant in combination with the at least portion of text to be executed, the third question-answering function being configured to determine the reply based on the speech information and the selected at least portion of text. 7.The method of claim 4, wherein determining the target function to be performed at the terminal device further comprises: in response to detecting that the digital assistant is not woken up and a current active interaction interface at the terminal device is an input box, determining to perform a speech-to-text function configured to convert the speech information into text and input into the input box. 8.The method of claim 4, wherein determining the target function to be performed at the terminal device comprises: in response to detecting that the digital assistant is not woken up and the terminal device joins a voice call session, determining to perform a voice call function configured to input the speech information as voice input of the voice call session. 9.The method of claim 4, wherein determining the target function to be performed at the terminal device comprises: in response to detecting that the digital assistant is not woken up and there is no current active interaction interface at the terminal device, determining to perform a control function of the terminal device configured to determine at least one control operation of the terminal device based on the speech information. 10.The method of claim 9, wherein the at least one control operation comprises one or more of: one or more control operations of an operating system of the terminal device, or one or more custom instructions, each custom instruction indicating a function configured to have a corresponding trigger word and comprising one or more control operations of the terminal device. 11.The method of claim 1, wherein providing a reply to the speech information at the terminal device comprises at least one of: presenting text content corresponding to the reply at the terminal device through a pop-up window, and playing speech content corresponding to the reply via the terminal device or an attached device of the terminal device. 12.The method of claim 1, wherein the wearable device comprises a ring. 13.An interaction control method implemented at a wearable device, comprising: capturing speech information of a user at the wearable device; and sending the speech information to a terminal device connected with the wearable device to make the terminal device determine a target function to be performed at the terminal device based on at least a current interaction state at the terminal device and the speech information. 14.An apparatus for interaction control, comprising: a receiving module configured to receive speech information of a user from a wearable device connected with a terminal device; a determining module configured to determine a target function to be performed at the terminal device based on at least a current interaction state at the terminal device in response to receiving the speech information; an executing module configured to perform the target function based on at least the speech information; and a reply module configured to provide a reply to the speech information at the terminal device, the reply indicating an execution result of the target function. 15.An apparatus for interaction control, comprising: a capturing module configured to capture speech information of a user at a wearable device; and a sending module configured to send the speech information to a terminal device connected with the wearable device to make the terminal device determine a target function to be performed at the terminal device based on at least a current interaction state at the terminal device and the speech information. The sending module is configured to send the voice information to a terminal device connected with the wearable device, so that the terminal device determines a target function to be executed at the terminal device based on at least a current interaction state at the terminal device and the voice information.
16. An electronic device, comprising: at least one processor; and at least one memory that is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-13.
17. A computer-readable storage medium having computer-executable instructions stored therein, the computer-executable instructions being executable by a processor to implement the method according to any one of claims 1-13.
18. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-13.
Citation Information
Patent Citations
Voice control method and device, electronic equipment and storage medium
CN111968640A
Voice recognition method, wearable device, and system
CN112334977A
Voice control method, wearable equipment and terminal
CN112420035A
Interaction method and device, wearable equipment and storage medium
CN118426592A
Wearable Earpiece Voice Command Control System and Method
US20170110124A1