Voice interaction method, interaction device, electronic device and storage medium

By collecting voice signals through an interactive device worn on the finger and sending them to the terminal device, the problem of low efficiency in noisy environments in traditional voice interaction is solved, and an efficient and flexible voice interaction experience is achieved.

CN121641004APending Publication Date: 2026-03-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Traditional voice interaction methods are inefficient in noisy environments, making it difficult to effectively wake up the voice assistant and affecting the user experience.

Method used

By using an interactive device worn on the user's finger, an audio acquisition unit collects voice signals and sends them to a terminal device to generate a second voice signal. It supports multiple wake-up methods such as button presses, audio, and motion sensing, improving the flexibility and efficiency of voice interaction.

Benefits of technology

It enables efficient voice interaction activation in noisy environments, reduces the probability of accidental operation, supports more voice interaction scenarios, and improves the efficiency and user experience of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641004A_ABST
    Figure CN121641004A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a voice interaction method, interaction equipment, electronic equipment and a storage medium. The method comprises the steps that in response to a preset operation for an interaction device, an audio collection unit of the interaction device is controlled to start collecting a first voice signal, and the interaction device is suitable for being worn on a finger of a user; acquiring a first voice signal acquired by an audio acquisition unit; and sending the first voice signal to a terminal device connected with the interaction device, so that the terminal device generates a second voice signal by using the first voice signal. In this way, according to the embodiment of the invention, voice interaction can be carried out through the interaction device suitable for being worn on the finger, so that the interaction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, interactive devices, electronic devices, and storage media for voice interaction. Background Technology

[0002] With the development of computer technology, voice interaction technology has become one of the important methods of human-computer interaction. For example, users can interact with terminal devices through voice to obtain feedback from the terminal devices.

[0003] Traditional voice interaction scenarios typically involve direct interaction between the user and the terminal device. This approach has significant limitations and affects the efficiency of voice interaction. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for voice interaction is provided. The method includes: in response to a preset operation on an interactive device, controlling an audio acquisition unit of the interactive device to begin acquiring a first voice signal, the interactive device being adapted to be worn on a user's finger; acquiring the first voice signal acquired by the audio acquisition unit; and sending the first voice signal to a terminal device connected to the interactive device, so that the terminal device can generate a second voice signal using the first voice signal.

[0005] In a second aspect of this disclosure, an interactive device suitable for wearing on a user's finger is provided. The interactive device includes: an interaction module configured to acquire preset operations performed by a user on the interactive device; a control module configured to, in response to the preset operations performed on the interactive device connected to a terminal device, control an audio acquisition unit of the interactive device to begin acquiring a first voice signal; an acquisition module configured to acquire the first voice signal acquired by the audio acquisition unit; and a transmission module configured to transmit the first voice signal to a terminal device connected to the interactive device, so that the terminal device can generate a second voice signal using the first voice signal.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0011] Figure 2 A flowchart of a voice interaction process according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A flowchart illustrating an example interaction process according to some embodiments of this disclosure is shown;

[0013] Figure 4 A schematic structural block diagram of an example interactive device according to some embodiments of the present disclosure is shown; and

[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0020] As mentioned above, voice interaction is an important interactive capability. For example, some terminal devices can deploy voice assistants to interact with users via voice. Traditional solutions typically involve waking up the voice assistant with voice commands. However, this wake-up method is not suitable for certain voice interaction scenarios (e.g., noisy public places), thus affecting the user's voice interaction experience.

[0021] Embodiments of this disclosure propose a voice interaction scheme. According to this scheme, the interactive device can respond to a preset operation on the interactive device by controlling the audio acquisition unit of the interactive device to start acquiring a first voice signal, wherein the interactive device is adapted to be worn on the user's finger.

[0022] Furthermore, the interactive device can acquire the first voice signal collected by the audio acquisition unit and send the first voice signal to a terminal device connected to the interactive device, so that the terminal device can use the first voice signal to generate a second voice signal.

[0023] In this way, embodiments of the present disclosure enable voice interaction via an interactive device suitable for wearing on a finger. Specifically, embodiments of the present disclosure can activate the voice interaction capability of a terminal device by detecting preset operations on the interactive device. Furthermore, by utilizing the audio acquisition unit on the interactive device to acquire voice signals and provide them to the terminal device, embodiments of the present disclosure can support more voice interaction scenarios and improve the efficiency of voice interaction.

[0024] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0025] Example Environment

[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include terminal device 110.

[0027] In this example environment 100, terminal device 110 may run an application 120 that supports voice interaction. Application 120 may include, but is not limited to, voice assistant applications, digital assistant applications, etc. User 140 may interact with application 120 via terminal device 110 and / or its attached devices.

[0028] like Figure 1 As shown, terminal device 110 can also be connected to interactive device 160. Interactive device 160 can be worn on the finger of user 140. As an example, interactive device 160 can be in the form of a ring or other suitable device.

[0029] In some embodiments, the interactive device 160 may include an audio acquisition unit, such as a microphone or microphone array. After detecting a preset operation by the user 140, the audio acquisition device of the interactive device 160 may begin to acquire the user's voice signal for voice interaction with the terminal device 110.

[0030] As will be detailed below, user 140 can activate the voice processing capability of terminal device 110 by interacting with interactive device 160. The specific process of voice interaction will be described in detail below. Figure 2 Detailed description.

[0031] exist Figure 1 In environment 100, if application 120 is active, terminal device 110 can also present interface 150 for supporting user interface interaction through application 120. As an example, interface 150 may be an interactive interface with a digital assistant, for example, to demonstrate the interaction process between the user and the digital assistant. This disclosure is not intended to limit the specific form of interface 150.

[0032] In some embodiments, terminal device 110 communicates with server 130 to provide services to application 120. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0033] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 supporting virtual scenarios in terminal device 110.

[0034] A communication connection can be established between server 130 and terminal device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and terminal device 110 can achieve signaling interaction through the communication connection between them.

[0035] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0036] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0037] Example voice interaction process

[0038] The voice interaction process according to embodiments of the present disclosure will now be described with reference to the accompanying drawings. Figure 2 A flowchart 200 of a voice interaction process according to some embodiments of the present disclosure is shown.

[0039] like Figure 2 As shown, at 205, the interactive device 160 can detect a preset operation for the interactive device 160 to control the audio acquisition unit in the interactive device 160 to start acquiring the first voice signal.

[0040] In some embodiments, the interactive device 160 may be equipped with buttons. The buttons may include physical buttons or pressure-sensitive buttons. The interactive device 160 may, for example, detect interaction operations on the buttons to detect preset interactions on the interactive device 160.

[0041] For example, the interactive device 160 can detect the user 140 pressing a physical button. Alternatively, the interactive device 160 can also detect the pressing of a pressure-sensitive button using a pressure-sensitive haptic sensor.

[0042] In some embodiments, such a press may include appropriate pressing operations such as single click, double click, or long press. In some embodiments, the specific button operation used to trigger the audio acquisition unit of the interactive device 160 may be determined based on the configuration of the user 140. As an example, the user may configure the specific button operation type via the interface 150 provided by the application 120.

[0043] By setting a button for waking up voice interaction on the interactive device 160, embodiments of this disclosure can enable users to achieve more efficient voice wake-up through an interactive device worn on their fingers, and reduce the probability of accidental operation.

[0044] In some embodiments, the interactive device 160 may also use an audio acquisition unit to detect whether a preset audio input has been received, so as to detect a preset interaction for the interactive device 160.

[0045] In some embodiments, the interactive device 160 may utilize an audio acquisition unit (e.g., a microphone) to perform sound detection. When a sound with a volume exceeding a threshold is detected, the interactive device 160 may determine that a preset audio input has been received and may utilize the audio acquisition unit to begin recording.

[0046] In other embodiments, the preset audio input may also include preset voice content, such as a wake-up word. The interactive device 160 may use an audio acquisition unit to receive the voice content output by the user 140, and the interactive device 160 or the terminal device 110 may determine whether the voice content matches the preset voice content, such as whether it includes a preset wake-up word.

[0047] When the preset voice content is detected, the interactive device 160 can determine that the preset audio input has been received and can start recording using the audio acquisition unit.

[0048] In other examples, after detecting a preset audio input, the interactive device 160 can further detect whether human voice content has been received. Specifically, the received voice signal can be processed by the interactive device 160 or the terminal device 110 to determine whether human voice content has been received. After detecting human voice content, the interactive device 160 can start recording using the audio acquisition unit, instead of starting recording immediately after receiving the preset audio input. In this way, embodiments of this disclosure can further provide the effectiveness of audio acquisition.

[0049] In some embodiments, the interactive device 160 may also detect preset interactions for the interactive device 160 by detecting preset haptic movements associated with the interactive device 160.

[0050] Specifically, such preset motion-sensing actions may include, for example, actions associated with the user's hand. In some embodiments, the interactive device 160 or the terminal device 110 may detect the user's wrist movements using motion sensors (e.g., accelerometers, gyroscopes, etc.) within the interactive device 160. For example, if the interactive device 160 detects that a user wearing the interactive device 160 has performed a wrist-raising motion, the interactive device 160 may begin recording using an audio acquisition unit.

[0051] In some embodiments, the interactive device 160 or the terminal device 110 can detect the user 140's hand gestures using motion sensors (e.g., accelerometers, gyroscopes, etc.) within the interactive device 160. For example, if the interactive device 160 detects that a user wearing the interactive device 160 has performed a two-finger pinch gesture, the interactive device 160 can start recording using an audio acquisition unit.

[0052] In some embodiments, such preset motion-sensing actions can also be determined based on configuration operations performed by user 140. As an example, user 140 can configure the motion-sensing action via interface 150 provided by terminal device 110. For example, user 140 can select a provided preset motion-sensing action, or can collect a custom motion-sensing action.

[0053] In some embodiments, the interactive device 160 may also activate the audio acquisition unit based on a combination of various methods to acquire the first voice signal.

[0054] As an example, the interactive device 160 can first detect a preset motion-sensing action associated with the interactive device 160, such as the wrist-raising action or gesture mentioned above. Further, after detecting the preset motion-sensing action, the interactive device 160 can activate the audio acquisition unit (e.g., a microphone) at the interactive device 160.

[0055] Furthermore, the interactive device 160 can use the activated audio acquisition unit to detect a preset audio input in order to determine whether to use the audio acquisition unit to acquire the first speech signal.

[0056] The above describes a combination of motion-activated wake-up and audio-activated wake-up. In another example, the interactive device 160 may also support a combination of button-activated wake-up and audio-activated wake-up. For example, the interactive device 160 may activate the audio acquisition unit after detecting a button press and further detect a preset audio input to determine whether to use the audio acquisition unit to acquire a first voice signal.

[0057] The above describes different wake-up methods for activating voice interaction. In some embodiments, the interactive device 160 may support multiple wake-up methods simultaneously. Additionally, the wake-up methods supported by the interactive device 160 (e.g., the operation type of the interactive operation, etc.) may also be determined based on the configuration operation of the user 140.

[0058] As an example, user 140 can configure support for button wake-up and motion wake-up via interface 150 provided by terminal device 110. In this case, when a button press or a preset motion action is detected, interactive device 160 can correspondingly control the audio acquisition unit to acquire a first voice signal.

[0059] Continue to refer to Figure 2 At 210, the interactive device 160 can use the audio acquisition unit to acquire a first voice signal. In some embodiments, the interactive device 160 can respond to the above wake-up method to turn on the audio acquisition unit and start recording.

[0060] In other embodiments, after the audio acquisition unit is turned on, the interactive device 160 or the terminal device 110 may also perform human voice detection on the audio signal acquired by the audio acquisition unit, and if it is determined that there is human voice content, start recording to acquire the first voice signal.

[0061] In some embodiments, taking a button-activated wake-up method as an example, the interactive device 160 can also detect the button's pressed state. While the button remains pressed, the interactive device 160 can control the audio acquisition device to begin acquiring a first voice signal.

[0062] For example, user 140 can trigger the audio acquisition unit to start recording by pressing and holding a button on interactive device 160. During the period the button is pressed, the audio acquisition unit continues to record, and after the pressing state ends, the audio acquisition unit stops recording, thereby completing the acquisition of the first voice signal.

[0063] As an example, under different triggering methods, the interactive device 160 can determine the end of the first voice signal acquisition in different ways. For example, the interactive device 160 can determine the completion of the first voice signal acquisition based on the end of the button press state.

[0064] As another example, the interactive device 160 or the terminal device 110 can detect whether human voice content is continuously received for a preset duration (e.g., 3 seconds). If not, the interactive device 160 can determine that the acquisition of the first voice signal has been completed.

[0065] As another example, the interactive device 160 can also detect preset haptic movements to determine the completion of the acquisition of the first voice signal. For example, when waking up the voice interaction by raising the wrist, the interactive device 160 can detect the haptic movement of the user lowering their wrist, and upon detecting this haptic movement, complete the acquisition of the first voice signal.

[0066] In some embodiments, the interactive device 160 may also include a notification unit. Taking a ring as an example, the notification unit may include a signal light disposed on the ring. The signal light may be disposed on the front of the ring, for example, a breathing light. Alternatively, the signal light may be disposed on the side of the ring, for example, a ring light.

[0067] In some embodiments, in response to the audio acquisition unit starting to acquire a first voice signal, the interactive device 160 may use an alert unit to present a first alert signal. For example, during audio acquisition unit activation or recording, an indicator light may remain constantly on to indicate that audio acquisition is currently in progress.

[0068] Continue to refer to Figure 2 After the first voice signal is acquired, at 215, the interactive device 160 can send the first voice signal to the terminal device 110.

[0069] Furthermore, at 220, terminal device 110 can use the received first voice signal to generate a second voice signal.

[0070] In some embodiments, the terminal device 110 may process the first voice signal using a locally deployed application or model to generate a second voice signal. For example, the terminal device 110 may have a locally deployed voice processing model to generate a second voice signal based on the first voice signal as a response to the first voice signal.

[0071] In some embodiments, such as Figure 2 As shown, terminal device 110 can also generate a second voice signal based on communication with server 130. As an example, server 130 can provide voice interaction services associated with application 140.

[0072] At 225, after receiving the first voice signal, the terminal device 110 can send the first voice signal to the server 130. The server 130 can, for example, generate a corresponding second voice signal based on the first voice signal.

[0073] As an example, server 130 may convert the first speech signal into first text content and provide it to a model associated with application 140 to generate second text content. Further, server 130 may generate a corresponding second speech signal based on the second text content.

[0074] It should be understood that such a model can include any suitable generative model, such as a language model. Alternatively, the model can also directly process the first speech signal to generate a second speech signal, for example.

[0075] At 230, terminal device 110 can obtain the generated second voice signal from server 130. In some embodiments, the second voice signal can be played by terminal device 110 or audio device 202 connected to terminal device 110. In some embodiments, audio device 202 may include an external audio device connected to terminal device 110 and having an audio playback unit, such as headphones, speakers, etc.

[0076] In some embodiments, although the interactive device 160 and the audio device 202 are in Figure 2 The audio device 202 can also be an interactive device 160 with audio playback units (e.g., speakers or speaker arrays) deployed, as shown in two separate boxes.

[0077] Specifically, such as Figure 2As shown, further, at 235, terminal device 110 can control the playback of the second voice signal. For example, at 240, terminal device 110 can send the second voice signal to audio device 202 to play the second voice signal using audio device 202. As another example, at 245, terminal device 110 can send the second voice signal to interactive device 160 to play the second voice signal using the audio playback unit of interactive device 160.

[0078] In some embodiments, when the terminal device 110 is connected to multiple devices with audio playback capabilities, the terminal device 110 can select a specific device for playback according to a preset strategy. This disclosure is not intended to limit the specific selection strategy.

[0079] In some embodiments, during the playback of the second voice signal, the reminder unit at the interactive device may present a second reminder signal. For example, a signal light may be flashing.

[0080] In some embodiments, during the playback of the second voice signal, the interactive device 160 may also detect a specific operation performed on the interactive device 160. In response to detecting the specific operation, the interactive device 160 may control the second voice signal to stop playing.

[0081] For example, this specific operation may include interactive operation of the buttons on the interactive device 160. For example, while playing a second voice signal using the terminal device 110, an external audio device, or the interactive device 160, if a press of a button on the interactive device 160 by the user 140 is received, the playback of the second voice signal can be stopped.

[0082] Based on the voice interaction process described above, embodiments of this disclosure enable voice interaction through an interactive device suitable for wearing on a finger. Specifically, embodiments of this disclosure can activate the voice interaction capability of a terminal device by detecting preset operations performed on the interactive device. Furthermore, by utilizing the audio acquisition unit on the interactive device to acquire voice signals and provide them to the terminal device, embodiments of this disclosure can support more voice interaction scenarios and improve the efficiency of voice interaction.

[0083] Example process

[0084] Figure 3 A flowchart of an example interaction process 300 according to some embodiments of the present disclosure is shown. Process 300 can be implemented at the interaction device 160. Reference is made below. Figure 1 To describe process 300.

[0085] As shown in the figure, in box 310, the interactive device 160 responds to a preset operation for the interactive device 160, controlling the audio acquisition unit of the interactive device 160 to start acquiring the first voice signal, and the interactive device 160 is suitable for wearing on the user's finger.

[0086] In frame 320, interactive device 160 acquires a first speech signal acquired by audio acquisition unit.

[0087] In box 330, the interactive device 160 sends a first voice signal to the terminal device 110 connected to the interactive device 160, so that the terminal device 110 can use the first voice signal to generate a second voice signal.

[0088] In some embodiments, the second voice signal can be played by the terminal device 110 or a target audio device connected to the terminal device 110.

[0089] In some embodiments, the interactive device 160 is equipped with an audio playback unit, the target audio device includes the interactive device 160 equipped with the audio playback unit, and the second voice signal can be played by the interactive device 160.

[0090] In some embodiments, the terminal device 110 is connected to an external audio device, the target audio device includes an external audio device, and the second voice signal can be played by the external audio device.

[0091] In some embodiments, the preset operation includes an interactive operation on a button on an interactive device.

[0092] In some embodiments, controlling the audio acquisition unit of the interactive device to start acquiring the first voice signal includes: in response to the button being held down, controlling the audio acquisition device to start acquiring the first voice signal.

[0093] In some embodiments, sending a first voice signal to a terminal device connected to an interactive device includes: sending a first voice signal to the terminal device in response to the end of a pressing state.

[0094] In some embodiments, process 300 further includes: during the playback of the second voice signal, in response to receiving an interactive operation on a key, controlling the playback of the second voice signal to stop.

[0095] In some embodiments, the preset operation includes a preset audio input for the audio acquisition unit.

[0096] In some embodiments, controlling the audio acquisition unit of the interactive device to start acquiring a first voice signal includes: detecting whether the audio acquisition unit receives human voice input in response to detecting a preset audio input; and controlling the audio acquisition unit to start acquiring the first voice signal in response to the audio acquisition unit receiving human voice input.

[0097] In some embodiments, controlling the audio acquisition unit of the interactive device to start acquiring a first voice signal includes: responding to the detection of a preset audio input, the preset audio containing preset voice content; and controlling the audio acquisition unit of the interactive device to start acquiring the first voice signal.

[0098] In some embodiments, process 300 further includes: activating an audio acquisition unit at the interactive device 160 in response to detecting a preset haptic action associated with the interactive device 160; and using the audio acquisition unit to detect whether a preset audio input has been received.

[0099] In some embodiments, the preset operation includes a preset motion-sensing action associated with the interactive device 160.

[0100] In some embodiments, preset motion-sensing operations include wrist-raising actions and gesture actions.

[0101] In some embodiments, process 300 further includes: in response to the audio acquisition unit starting to acquire the first voice signal, presenting the first reminder signal using the reminder unit at the interactive device 160.

[0102] In some embodiments, process 300 further includes presenting a second reminder signal using a reminder unit at the interactive device 160 during the playback of the second voice signal.

[0103] In some embodiments, the operation type of the preset operation is determined based on the user's configured operation.

[0104] Example devices and equipment

[0105] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example interactive device 160 according to certain embodiments of the present disclosure is shown. The interactive device 160 may be implemented as or included in the interactive device 160. Various modules / components in the interactive device 160 may be implemented by hardware, software, firmware, or any combination thereof.

[0106] like Figure 4 As shown, the interactive device 160 includes an interactive module 410 configured to collect preset operations performed by a user on the interactive device 160; a control module 420 configured to control the audio acquisition unit of the interactive device to start collecting a first voice signal in response to the preset operation performed on the interactive device connected to the terminal device; an acquisition module 430 configured to acquire the first voice signal collected by the audio acquisition unit; and a sending module 440 configured to send the first voice signal to the terminal device 110 connected to the interactive device 160, so that the terminal device 110 can generate a second voice signal using the first voice signal.

[0107] In some embodiments, the second voice signal can be played by the terminal device 110 or a target audio device connected to the terminal device 110.

[0108] In some embodiments, the interactive device 160 is equipped with an audio playback unit, the target audio device includes the interactive device 160 equipped with the audio playback unit, and the second voice signal can be played by the interactive device 160.

[0109] In some embodiments, the terminal device 110 is connected to an external audio device, the target audio device includes an external audio device, and the second voice signal can be played by the external audio device.

[0110] In some embodiments, the preset operation includes an interactive operation on a button on an interactive device.

[0111] In some embodiments, the interactive device is configured to control the audio acquisition device to begin acquiring a first voice signal in response to a button being held down.

[0112] In some embodiments, the interactive device is further configured to send a first voice signal to the terminal device in response to the end of the pressing state.

[0113] In some embodiments, the interactive device 160 further includes a stop module configured to control the second voice signal to stop playing in response to receiving an interactive operation on a button during the playback of the second voice signal.

[0114] In some embodiments, the preset operation includes a preset audio input for the audio acquisition unit.

[0115] In some embodiments, the control module 420 is further configured to: detect whether the audio acquisition unit receives human voice input in response to detecting a preset audio input; and control the audio acquisition unit to start acquiring a first voice signal in response to the audio acquisition unit receiving human voice input.

[0116] In some embodiments, the control module 420 is further configured to: respond to detecting a preset audio input, the preset audio containing preset voice content; and control the audio acquisition unit of the interactive device to start acquiring a first voice signal.

[0117] In some embodiments, the control module 420 is further configured to: activate the audio acquisition unit at the interactive device 160 in response to detecting a preset haptic action associated with the interactive device 160; and use the audio acquisition unit to detect whether a preset audio input has been received.

[0118] In some embodiments, the preset operation includes a preset motion-sensing action associated with the interactive device 160.

[0119] In some embodiments, preset motion-sensing operations include wrist-raising actions and gesture actions.

[0120] In some embodiments, the interactive device 160 further includes a reminder module configured to: in response to the audio acquisition unit starting to acquire a first voice signal, present a first reminder signal using the reminder unit at the interactive device 160.

[0121] In some embodiments, the reminder module is further configured to present a second reminder signal using a reminder unit at the interactive device 160 during the playback of the second voice signal.

[0122] In some embodiments, the operation type of the preset operation is determined based on the user's configured operation.

[0123] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 can be used for at least some steps of the voice interaction process described above.

[0124] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0125] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0126] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0127] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0128] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0129] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0130] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0131] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0132] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0134] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1.A voice interaction method applied to an interaction device, the method comprising: controlling an audio capturing unit of the interaction device to start capturing a first voice signal in response to a preset operation on the interaction device, the interaction device being adapted to be worn on a finger of a user; obtaining the first voice signal captured by the audio capturing unit; and sending the first voice signal to a terminal device connected to the interaction device, for the terminal device to generate a second voice signal using the first voice signal. 2.The method of claim 1, wherein the second voice signal is capable of being played by the terminal device or a target audio device connected to the terminal device. 3.The method of claim 2, wherein the interaction device is deployed with an audio playing unit, the target audio device comprises the interaction device deployed with the audio playing unit, and the second voice signal is capable of being played by the interaction device. 4.The method of claim 2, wherein the terminal device is connected to an external audio device, the target audio device comprises the external audio device, and the second voice signal is capable of being played by the external audio device. 5.The method of claim 1 or 2, wherein the preset operation comprises an interaction operation on a button on the interaction device. 6.The method of claim 5, wherein the controlling the audio capturing unit of the interaction device to start capturing the first voice signal comprises: controlling the audio capturing unit to start capturing the first voice signal in response to the button being kept in a pressed state. 7.The method of claim 6, wherein the sending the first voice signal to the terminal device connected to the interaction device comprises: sending the first voice signal to the terminal device in response to the pressed state ending. 8.The method of claim 5, further comprising: controlling the second voice signal to stop playing in response to receiving an interaction operation on the button during playing of the second voice signal. 9.The method of claim 1, wherein the preset operation comprises a preset audio input to the audio capturing unit. 10.The method of claim 9, wherein the controlling the audio capturing unit of the interaction device to start capturing the first voice signal comprises: detecting whether the audio capturing unit receives a human voice input in response to detecting the preset audio input; and controlling the audio capturing unit to start capturing the first voice signal in response to the audio capturing unit receiving the human voice input. 11.The method of claim 9, wherein the controlling the audio capturing unit of the interaction device to start capturing the first voice signal comprises: controlling the audio capturing unit to start capturing the first voice signal in response to detecting the preset audio input, the preset audio containing preset voice content. 12.The method of claim 9, wherein the controlling the audio capturing unit of the interaction device to start capturing the first voice signal comprises: starting the audio capturing unit at the interaction device in response to detecting a preset somatosensory action associated with the interaction device. ​ ​ ​ ​ and detecting, by the audio acquisition unit, whether the preset audio input is received. 13.The method of claim 1, wherein the preset operation comprises a preset somatosensory action associated with the interaction device. 14.The method of claim 12 or 13, wherein the preset somatosensory operation comprises a wrist-lifting action, a gesture action. 15.The method of claim 1, wherein the interaction device is further configured to present, by a reminder unit at the interaction device, a first reminder signal in response to the audio acquisition unit starting to acquire the first voice signal. 16.The method of claim 15, further comprising: presenting, by the reminder unit at the interaction device, a second reminder signal during the playing of the second voice signal. 17.The method of claim 1, wherein a type of the preset operation is determined based on a configuration operation of a user. 18.An interaction device adapted to be worn on a finger of a user, comprising: an interaction module configured to acquire a preset operation of a user on the interaction device; a control module configured to control an audio acquisition unit of the interaction device to start to acquire a first voice signal in response to the preset operation on the interaction device connected to a terminal device; an acquisition module configured to acquire the first voice signal acquired by the audio acquisition unit; and a sending module configured to send the first voice signal to a terminal device connected with the interaction device, for the terminal device to generate a second voice signal by using the first voice signal. 19.The interaction device of claim 18, further comprising: a receiving module configured to receive the second voice signal sent from the terminal device, the second voice signal being generated by the terminal device based on the first voice signal; an audio playing module configured to play the second voice signal at the interaction device. 20.An electronic device comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 17. 21.A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 17. ​

Citation Information

Patent Citations

  • Voice control realizing method and system based on intelligent wearable equipment

    CN105096951A

  • Voice interaction method for wearable electronic equipment, and wearable electronic equipment

    CN107481721A

  • Voice interaction wake-up electronic equipment and method based on microphone signal and medium

    CN110111776A

  • Voice response method and device, equipment and storage medium

    CN112201230A

  • Wearable intelligent sound box and intelligent sound box system

    CN209017223U