Voice interaction method, apparatus, and related devices

Augmented reality devices enhance flexibility and versatility by performing voice interaction with diverse target devices, including those without a connected terminal, through type determination and adaptive response modes, addressing the limitation of interacting with only compatible terminals.

JP2026502430AActive Publication Date: 2026-01-23HANGZHOU LINGBAN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025534437
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-12
Filing Date
2024-01-12
Publication Date
2026-01-23
Estimated Expiration
2044-01-12

AI Technical Summary

Technical Problem

Augmented reality devices are limited to interacting with only compatible terminal devices, lacking flexibility and multifunctionality, which restricts their usage scenarios.

Method used

The augmented reality device can perform voice interaction with different types of target devices, including those without a connected terminal, by determining the type of target device and employing appropriate response modes, utilizing voice keyword detection technology to enable voice interaction even when not connected, and converting commands as needed.

Benefits of technology

This enhances the flexibility and versatility of augmented reality devices, allowing them to maintain voice interaction capabilities across various devices, expanding their usage scenarios and improving user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502430000001_ABST
    Figure 2026502430000001_ABST
Patent Text Reader

Abstract

The present invention provides a voice interaction method, apparatus, and related device, which includes: acquiring a user's voice command through a voice collection unit of an augmented reality device; determining a type of a target device connected to the augmented reality device; determining a target response mode for the voice command as a first response mode if the target device is a first type; determining a target response mode for the voice command as a second response mode if the target device is a second type or if the target device is not detected; and responding to the voice command using the target response mode, thereby providing technical support for normal voice interaction even when no target device is connected to the augmented reality device or a different target device is connected.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-Incorporation of Related Applications) This application claims priority from a Chinese patent application bearing application number 202310096785.6, filed with the China Intellectual Property Office on January 12, 2023, the entire contents of which are incorporated herein by reference.

[0002] The present invention relates to the field of voice interaction, and in particular to a voice interaction method, apparatus and related device. [Background technology]

[0003] Augmented reality devices are typically used to display or collect sensor data, and in most cases, they need to be connected to a compatible (compatible) terminal device in order to be used. DISCLOSURE OF THE INVENTION

[0004] The present application provides a voice interaction method, apparatus, and related devices, which solve at least the above technical problems existing in the prior art.

[0005] In a first aspect, the present application provides a voice interaction method, the method comprising: acquiring a voice command of the user by a voice collection unit of the augmented reality device; determining a type of target device connected to the augmented reality device; If the target device is of a first type, determining a target response mode to the voice command as a first response mode; determining a target response mode to the voice command as a second response mode if the target device is of a second type or if no target device is detected; Responding to the voice command using the target response mode.

[0006] In the above configuration, responding to the voice command using the target response mode includes: If the target response mode is a first response mode, sending the voice command to the target device to respond; If the target response mode is a second response mode, determining whether the augmented reality device or the target device responds based on the type of the voice command.

[0007] In the configuration, determining which of the augmented reality device and the target device will respond based on the type of voice command includes: if the voice command is a first type of command, the augmented reality device responds to the voice command; If the voice command is a second type of command, sending the voice command to the target device for response.

[0008] In the above configuration, when the voice command is a second type command, transmitting the voice command to the target device to cause the target device to respond includes: If the voice command is a second type command, convert the voice command into a target type command and send the target type command to the target device for response.

[0009] In the configuration, the sound collection unit is used to acquire target sound data, and the target sound data includes a sound command; The method further comprises: determining whether a voice command is present in the target voice data, thereby obtaining the voice command; and / or transmitting the target audio data to the target device.

[0010] In the above configuration, determining whether a voice command is present in the target voice data to obtain the voice command; and / or transmitting the target audio data to the target device performing noise reduction processing on the target voice data to obtain noise-reduced target voice data; determining whether a voice command is present in the noise-reduced target voice data, thereby obtaining the voice command; and / or transmitting the noise-reduced target audio data to the target device.

[0011] In the above configuration, determining whether a voice command is present in the target voice data after noise reduction to acquire the voice command; and / or transmitting the noise-reduced target audio data to the target device, acquiring first target sound data and second target sound data based on the noise-reduced target sound data; determining whether a voice command is present in the first target voice data to obtain the voice command; and / or transmitting the second target audio data to the target device. Includes:

[0012] In a second aspect, the present application provides a voice interaction device, the device including: an acquisition unit, a first determination unit, a second determination unit, a third determination unit and a response unit; The acquisition unit acquires a voice command of a user through a voice collection unit of an augmented reality device; the first determining unit is configured to determine a type of a target device connected to the augmented reality device; the second determination unit determines a target response mode to the voice command as a first response mode when the target device is a first type; the third determination unit determines a target response mode to the voice command as a second response mode when the target device is of a second type or when no target device is detected; The response unit responds to the voice command using the target response mode.

[0013] In a third aspect, the present application provides an augmented reality device, the augmented reality device comprising at least the voice interaction apparatus described in the present application.

[0014] In a fourth aspect, the present application provides an electronic device, comprising at least one processor and a memory communicatively connected to the at least one processor, the memory storing commands executable by the at least one processor, the commands being executed by the at least one processor to enable the at least one processor to perform a method described in the present application.

[0015] The present application includes obtaining a user's voice command through a voice collecting unit of an augmented reality device, determining the type of a target device connected to the augmented reality device, determining a target response mode for the voice command as a first response mode if the target device is a first type, determining a target response mode for the voice command as a second response mode if the target device is a second type or if the target device is not detected, and responding to the voice command using the target response mode, thereby providing technical support for normal voice interaction even when no target device is connected to the augmented reality device or a different target device is connected.

[0016] It should be understood as follows: The contents described in this section are not intended to identify the key points or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will be easily understood from the following specification. [Brief explanation of the drawings]

[0017] These and other objects, features, and advantages of exemplary embodiments of the present application will become more readily apparent from the following detailed description taken in conjunction with the drawings in which several embodiments of the present application are illustrated by way of example and not of limitation, and in which: In the drawings, like or corresponding reference numerals indicate like or corresponding parts.

[0018] [Figure 1] FIG. 1 is a schematic diagram of the implementation flow of the voice interaction method of the embodiment of the present application. [Figure 2] FIG. 2 is a schematic diagram of the realization flow of different target response modes in an embodiment of the present application. [Figure 3] FIG. 3 is a schematic diagram of the implementation flow of different target response modes in an embodiment of the present application. [Figure 4] FIG. 4 is a schematic diagram showing the data flow of an augmented reality device terminal in an embodiment of the present application. [Figure 5] FIG. 5 is a schematic diagram of the configuration of a voice interaction device in an embodiment of the present application. [Figure 6] FIG. 6 is a schematic diagram of the configuration of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] In order to make the objectives, features, and advantages of the present application clearer and easier to understand, the following provides a clear and complete description of the technical configurations in the embodiments of the present application, in conjunction with the drawings in the embodiments of the present application. Needless to say, the described embodiments are only a part of the embodiments of the present application, and are not all of the embodiments. Based on the embodiments of the present application, any other embodiments that can be obtained by a person skilled in the art without any creative effort are all within the scope of protection of the present application.

[0020] In the related art, an augmented reality device can only communicate with a terminal device that is compatible with it, and the flexibility and multi-functionality of the augmented reality device cannot be realized.

[0021] It can be understood as follows: Augmented reality (AR) devices are currently the mainstream wearable device. Intelligent voice interaction, as the mainstream interaction method for AR devices, allows for hands-free, simple, and quick input or control of AR devices. Considering factors such as appearance, wearability, power consumption, and heat generation, AR devices are typically used only as displays (projections) and sensor data collection functions (e.g., images, audio, inertial measurement units, etc.) rather than as complex computing units. Based on this, AR devices can only be connected to compatible terminal devices via wired or wireless methods and transmit collected sensor data to the terminal device for algorithmic calculations. If AR devices could successfully perform voice interaction even when not connected to a terminal device or connected to another terminal device, the functionality of AR devices would undoubtedly be enhanced. This lays the foundation for the widespread application of AR devices in everyday life.

[0022] The technical configuration of the embodiments of the present application relates to a voice interaction configuration. The augmented reality device can perform voice interaction with different types of target devices, and can also perform voice interaction without being connected to a target device, demonstrating the multifunctionality and flexibility of the augmented reality device. Depending on the acquired voice command and the type of target device connected to the augmented reality device, the augmented reality device can respond to the voice interaction using different target response modes. Technical support is provided to enable normal voice interaction even when the augmented reality device is not connected to a target device or is connected to a different target device, thereby expanding the usage scenarios of the augmented reality device.

[0023] The voice interaction method of the embodiment of the present application will be described in detail below.

[0024] An embodiment of the present application provides a voice interaction method. As shown in Figure 1, the method includes:

[0025] S101: Acquire a user's voice command through a voice collection unit of the augmented reality device. In this step, the augmented reality device is an electronic device capable of AR interaction. The augmented reality device can be a smart wearable device, including but not limited to smart glasses and smart watches. In this application, the augmented reality device will be described using separated AR glasses as an example.

[0026] The augmented reality device includes a voice collection unit, for example a microphone, and in this step, obtains the user's voice commands by collecting the voice commands uttered by the user to the augmented reality device through the microphone.

[0027] It can be understood as follows: The microphone (Mic) includes a microphone array sensor (MicArray) for capturing a user's voice commands issued to the augmented reality device.

[0028] In practical applications, a user can issue voice commands to an augmented reality device even when the augmented reality device is not in use, for example, when the screen of the augmented reality device is turned off, the user can issue a voice command such as "Turn on the screen" or "Turn off the power."

[0029] Furthermore, the user can also issue voice commands to the augmented reality device when the augmented reality device is being used, for example, when the augmented reality device is being used to play a movie or play audio such as a song, etc. That is, when the augmented reality device is being used to output multimedia data, the user's voice commands are captured by the voice collection unit of the augmented reality device.

[0030] It can be understood as follows: An augmented reality device typically outputs multimedia data, such as images and audio, during the AR interaction process. When the augmented reality device and a target device are connected, a video in the target device can be projected and output on the augmented reality device, for example, a movie in the target device can be projected and output on the augmented reality device. In such a case, the multimedia data can also refer to a video in the target device that can be projected by the augmented reality device. Furthermore, audio in the target device can also be output via the augmented reality device, for example, a voice call or a voicemail can be made through the target device. In such a case, the multimedia data can also refer to audio in the target device that can be output by the augmented reality device.

[0031] That is, when connected to a target device, the augmented reality device of the present application serves as an alternative target device, serving as a device that outputs video and audio. Such alternatives are primarily considered for the following reasons: In some applications, using the target device as a device that outputs video and audio is far less convenient and results in a lower output effect than using the augmented reality device as a device that outputs video and audio. For example, in projection applications, an augmented reality device can be used to project video projectable in the target device, providing an immersive experience to the wearer of the augmented reality device. Furthermore, in situations where it is inconvenient to answer a call from the target device in a crowded viewing environment, such as during rush hour on the subway, the user can answer the call by wearing split AR glasses. The split AR glasses can be incorporated into regular eyeglasses worn by the wearer. Using the split AR glasses to answer a call can avoid situations where the wearer is unable to take their cell phone out of their pocket in a crowded environment.

[0032] If desired, the user can issue voice commands to the augmented reality device, which specifically uses a microphone to collect the voice commands.

[0033] For example, if the augmented reality device is split AR glasses, and the current scene is a scene in which a user is playing audio or video using the split AR glasses, the multimedia data output from the split AR glasses is the audio and video information being played by the user. When the user issues a voice command such as "Turn up the volume," the microphone can collect the voice command issued by the user to obtain the voice command for audio and video playback. In response to the voice command, the volume of the currently playing audio and video is increased.

[0034] S102: Determine the type of target device connected to the augmented reality device.

[0035] In this step, the target device can be any device capable of voice interaction with the augmented reality device, such as a smartphone, a computer, a personal digital assistant, etc. The target device connected to the augmented reality device in this application can be a different type of terminal. For example, the target device is a self-developed terminal, i.e., a terminal compatible with the augmented reality device. It should be understood as follows: After an augmented reality device is manufactured by a manufacturer, there is usually a self-developed terminal manufactured by the manufacturer that is compatible with the augmented reality device. The self-developed terminal can be understood as a normal-sized terminal without a display screen or with a relatively small display screen but with computing capabilities. By connecting the augmented reality device with a compatible self-developed terminal, the intelligent voice interaction function of the augmented reality device can be realized.

[0036] The target device can be another type of terminal, such as a third-party mobile phone or a third-party computer. In this application, the intelligent voice interaction function of the augmented reality device is realized by connecting the augmented reality device to another type of terminal. For example, if the augmented reality device is split AR glasses, the split AR glasses can not only connect to a terminal compatible with the split AR glasses, but also to a third-party terminal. In this application, whether the split AR glasses are connected to a terminal compatible with the split AR glasses or a third-party terminal, the split AR glasses can perform intelligent voice interaction with the device connected to the split AR glasses.

[0037] In practical applications, voice interaction is the mainstream interaction method for augmented reality devices, but it is usually limited to in-house developed devices. That is, if a user wants to use an augmented reality device, they must purchase a in-house developed device that is compatible with that augmented reality device in order to properly use voice interaction. In such cases, on the one hand, if a user already owns a mobile phone, they may not purchase an additional compatible device due to cost and portability considerations. On the other hand, when an augmented reality device is connected to a personal computer (PC), the PC is a computing unit and does not connect to other devices. In these two cases, the voice interaction function of the augmented reality device becomes unavailable, greatly limiting the use of voice interaction in AR glasses.

[0038] In this application, the augmented reality device can be connected to different types of terminals and can realize the voice interaction functions of the augmented reality device by determining the type of target device connected to the augmented reality device, while the augmented reality device can realize basic voice interaction functions, such as adjusting the volume and brightness, through its own voice command control application without connecting to any type of terminal.

[0039] In practical applications, the augmented reality device can be connected to different types of terminals or not connected to any terminal at all, which allows users to purchase only the augmented reality device without purchasing a compatible proprietary terminal. Users can realize the voice interaction function of the augmented reality device by directly connecting the augmented reality device to their mobile phone or computer.

[0040] S103: If the target device is of a first type, determine a target response mode to the voice command as a first response mode.

[0041] In this application, the augmented reality device can be connected or coupled to different types of terminals, and the response mode to the voice command will be different depending on the type of target device connected to the augmented reality device.

[0042] In this application, the target device type is divided into a first type and a second type. The first type refers to a terminal that has a built-in voice keyword detection technology service, such as the self-developed terminal. The second type refers to a terminal that does not have a built-in voice keyword detection technology service, such as the third-party mobile phone or computer.

[0043] For example, if the target device connected to the augmented reality device is the in-house developed terminal, the corresponding target response mode for the voice command may be Mode A (first response mode). If the target device connected to the augmented reality device is another type of terminal (e.g., a third-party mobile phone, a computer, etc.), the corresponding target response mode for the voice command may be Mode B (second response mode).

[0044] In implementation, the augmented reality device identifies whether a target device is connected to it. If no device is connected, the target response mode for the voice command is determined to be the second response mode. If a target device is connected to the augmented reality device, the device acquires an identifier of the connected device and determines whether the connected device is of a first type or a second type based on the identifier. If the identifier of the connected device is identifier A and identifier A represents a device of the first type, the device is determined to be of the first type. If the identifier of the connected device is identifier B and identifier B represents a device of the second type, the device is determined to be of the second type.

[0045] In this application, two response modes are preset based on whether the terminal connected to the augmented reality device is a terminal compatible (adaptable) with the augmented reality device or a third-party terminal not compatible with the augmented reality device. One of the response modes is used when the terminal connected to the augmented reality device is a terminal compatible with the augmented reality device. The other response mode is used when the terminal connected to the augmented reality device is a third-party terminal. Two different types of terminals and the modes used by each type of terminal are preset in a correspondence relationship. During implementation, based on the type of terminal connected to the augmented reality device, the corresponding mode of the terminal of that type is searched from the correspondence relationship and used as the target response mode for responding to a voice command.

[0046] S104: If the target device is of a second type or if no target device is detected, determine the target response mode to the voice command as a second response mode.

[0047] Since the second type of target device is a terminal that does not include a voice keyword detection technology service, the augmented reality device of the present application utilizes a voice keyword detection technology service to enable normal voice interaction when connected to the second type of target device, even when the augmented reality device is connected to the second type of target device. In addition, since the augmented reality device of the present application is equipped with a voice keyword detection technology service, even when the augmented reality device is not connected to the target device, i.e., the target device is not detected, the augmented reality device can use the corresponding target response mode to perform basic voice interactions, such as adjusting brightness and volume.

[0048] In this application, by arranging a low-power voice keyword detection technology service in an augmented reality device, it is possible to perform identification of voice commands with the help of the low-power voice keyword detection technology service without significantly increasing the power consumption or heat generation of the augmented reality device.

[0049] By implementing different target response modes for different types of terminals, the augmented reality device can maintain voice interaction capabilities even when connected to different types of terminals.

[0050] S105: Respond to the voice command using the target response mode.

[0051] Depending on the type of the connected target device, different target response modes are used to respond to the voice command. That is, different target response modes are used to analyze what type of command the voice command is, and then processing is performed for this type of command. For example, if the voice command is analyzed as a voice command for "increase volume," the output volume of the augmented reality device is increased. If the voice command is analyzed as a voice command for "decrease screen brightness," the brightness of the screen of the augmented reality device is decreased. In this way, an operation corresponding to the voice command is realized.

[0052] In the configuration shown in S101 to S105, the augmented reality device can not only communicate with terminals compatible with the augmented reality device and third-party terminals, but also perform voice interaction even when no terminal is connected, which demonstrates the flexibility and versatility of the augmented reality device.

[0053] Furthermore, in this application, a target response mode is determined based on the type of target device connected to the augmented reality device, and a response to a voice command is realized using the target response mode. Different target response modes can be used to respond to voice commands depending on the type of target device connected to the augmented reality device. This allows the augmented reality device to maintain voice interaction functionality even when connected to different types of terminals. This provides technical support for the augmented reality device to normally perform voice interaction even when connected to different target devices.

[0054] In one optional configuration, responding to the voice command using the target response mode includes:

[0055] If the target response mode is the first response mode, the voice command is sent to the target device to cause it to respond.

[0056] If the target response mode is the second response mode, it is determined whether the augmented reality device or the target device responds based on the type of the voice command.

[0057] The type of voice command indicates whether the voice command is a command type to which the augmented reality device responds or a command type to which the target device responds.

[0058] As shown in FIG. 2 , in this application, the responding entity is mainly divided into two types: a target device and an augmented reality device. The augmented reality device mainly includes a voice keyword detection technology service and a voice command control application. If the target response mode is the first response mode, i.e., if the target response mode is determined to be the first response mode based on the type of the target device connected to the augmented reality device being a first type, such as a proprietary terminal, the voice command is sent to the operating system of the target device for full voice control. If the target response mode is the second response mode, i.e., if the target response mode is determined to be the second response mode based on the type of the target device connected to the augmented reality device being a second type, such as a third-party smartphone or computer, or if the target device is not detected, the voice command is converted into an event ID (Event ID) and notified to the voice command control application in the augmented reality device. The voice command control application determines whether the command is a control command for the augmented reality device itself. If the command is a control command for the augmented reality device itself, the voice command control application responds to the voice command, such as adjusting the volume or brightness. For example, if the voice command is a control command for the target device, the event ID (Event Id) is converted into a universal keyboard (USB Keyboard, Universal Serial Bus Keyboard) protocol ID, the USB Keyboard Id is sent to the target device, and the target device responds to the voice command.

[0059] Specifically, as shown in FIG. 3, when the voice keyword detection technology service of the augmented reality device is running, the voice keyword detection technology service determines whether the target device connected to the augmented reality device is a first type or a second type. Different target response modes are used depending on the type of the connected target device. If the target response mode is the first response mode, i.e., if the target response mode is determined to be the first response mode based on the fact that the type of the target device connected to the augmented reality device is a first type (e.g., a self-developed device), the voice keyword detection technology service in the augmented reality device goes into a dormant state, and voice commands are handed over to the operating system of the target device for full voice control. For example, assuming that the target device connected to the augmented reality device is a self-developed device, since the self-developed device is equipped with the voice keyword detection technology service, there is usually a set of predefined voice commands on the self-developed device. For example, a volume up command adjusts the volume of multimedia data, a brightness up command adjusts the display brightness of the screen, and a mode switch command adjusts the display mode of multimedia data (e.g., from normal mode to 3D mode, or from 3D mode to normal mode, etc.).

[0060] When a user issues a corresponding voice command, the augmented reality device determines that the connected target device is a proprietary device and sends the collected voice command to the operating system of the target device for full voice control. A voice keyword detection technology service running on the proprietary device passes the voice command to the proprietary device. The proprietary device, specifically its application processing unit, responds to the voice command and performs a control action corresponding to the voice command, such as adjusting the volume, adjusting the brightness, or switching modes.

[0061] If the voice command received by the augmented reality device is a command (voice adjustment) for adjusting the voice of the multimedia data output by the augmented reality device, and if the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, the voice adjustment command is passed on to the proprietary terminal, and the proprietary terminal adjusts the voice of the multimedia data output by the augmented reality device.

[0062] When the voice command acquired by the augmented reality device is a command for adjusting the display brightness of multimedia data output by the augmented reality device, if the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, the command for adjusting the display brightness is passed to the proprietary terminal, which then adjusts the display brightness of the screen of the augmented reality device.

[0063] If the voice command acquired by the augmented reality device is a command to adjust the display mode of multimedia data output by the augmented reality device, for example, if the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, the command to adjust the display brightness is passed to the proprietary terminal, which then adjusts the display mode of the augmented reality device.

[0064] If the target response mode is the second response mode, i.e., if the target response mode is determined to be the second response mode based on the type of the target device connected to the augmented reality device being a second type, such as a third-party mobile phone or computer, or the target device not being detected, the voice keyword detection technology service of the augmented reality device remains operational.

[0065] In one optional configuration, when the target response mode is a second response mode, determining which of the augmented reality device and the target device will respond based on the type of voice command includes:

[0066] If the voice command is a first type of command, the augmented reality device responds to the voice command.

[0067] If the voice command is a second type of command, the voice command is sent to the target device for response.

[0068] In the present application, when the target response mode is the second response mode, the voice commands include two types: the first type of commands are commands to which the augmented reality device can respond, such as the volume adjustment command, display brightness adjustment command, and display mode adjustment command; and the second type of commands are commands to which the augmented reality device cannot respond but the target device can respond, such as the "back" command, "confirm" command, and "main menu" command.

[0069] In the present application, when the target response mode is the second response mode, the augmented reality device responds to the voice command if the voice command is a command to which the augmented reality device can respond, such as a volume adjustment command, a display brightness adjustment command, a display mode adjustment command, etc. If the voice command is a "back" command, a "confirm" command, a "main menu" command, etc., the voice command is passed on to the target device, and the target device responds to the voice command.

[0070] In view of this, in the present application, in the second response mode, different response entities respond to voice commands according to different types of voice commands, thereby avoiding the problem of excessive load caused by only one of the response entities responding to all voice commands. In the present application, different types of voice commands are responded to by different response entities, which is easy to implement in engineering and highly feasible.

[0071] In one optional configuration, if the voice command is a second type of command, sending the voice command to the target device to respond includes:

[0072] If the voice command is a second type command, convert the voice command into a target type command and send the target type command to the target device for response.

[0073] In this application, for example, if a voice command is a command that the augmented reality device cannot respond to but the target device can respond to, the voice command needs to be converted into a type that the target device can recognize or respond to, thereby realizing a response to the voice command.

[0074] For example, assuming that the target device connected to the augmented reality device is a third-party mobile phone or computer, when a voice keyword detection technology service in the augmented reality device identifies a voice command, it converts the voice command into an event ID (Event ID) and notifies the voice command control application in the augmented reality device. Upon receiving the command's Event ID (Event ID), the voice command control application compares the command's Event ID (Event ID) with numbers corresponding to predefined command functions and provides corresponding feedback based on the predefined command functions represented by the corresponding numbers. The predefined command functions are command functions represented by different numbers in a predefined command set, such as number 1 for "increase volume," number 2 for "increase brightness," and number 3 for "switch modes."

[0075] If the command corresponding to the command event ID (Event Id) is a first-type command, i.e., if the command event ID (Event Id) corresponds to a number corresponding to a command function predefined in the augmented reality device, i.e., if the command corresponding to the command event ID (Event Id) is a control-related command to which the augmented reality device can respond, such as a command related to the hardware of the augmented reality device, such as volume adjustment, brightness adjustment, display switching, or mode switching, the voice command control application executes the corresponding control operation directly via the system application programming interface (API). If the command corresponding to the command event ID (Event Id) is a second-type command, i.e., if the command event ID (Event Id) corresponds to a number corresponding to a command function predefined in the target device, i.e., if the command corresponding to the command event ID is a command requiring a response from the target device, the command event ID (Event Id) is converted into a USB keyboard protocol ID and sent to the connected target device. The target device responds to the USB keyboard protocol ID and executes an operation on the multimedia data output from the augmented reality device. For example, a "Back" command may be defined as the "F1" key on the keyboard of the target device. When a user presses the "F1" key, the multimedia data output by the augmented reality device returns to the previous step. Alternatively, a "Main Menu" command may be defined as the "F2" key on the keyboard of the target device. When a user presses the "F2" key, the augmented reality device pops up a main menu interface. Alternatively, a "Confirm" command may be defined as the "Enter" key on the keyboard of the target device. When a user presses the "Enter" key, a confirmation operation can be performed.

[0076] The above-described process performed by the augmented reality device in the present application can be realized by a high-performance dedicated chip installed in the augmented reality device. The use of a high-performance dedicated chip improves computational efficiency and enables rapid response to voice commands. Furthermore, the present application utilizes a low-power voice keyword detection technology service, which has significant advantages, such as low computational power and low energy consumption. By incorporating AR functionality into the augmented reality device without significantly increasing power consumption, the controllability of the augmented reality device itself is significantly improved. Furthermore, voice commands are customizable and highly scalable, overcoming the lack of key operations. By using a generic USB keyboard protocol, the augmented reality device converts the identified voice command into a target type that the target device can recognize or respond to. The target device's response to the voice command can be seen as a user operating a key (e.g., the "F1" key on the target device's keyboard, the "F2" key on the target device's keyboard, and the "Enter" key on the target device's keyboard) to respond to the voice command. This configuration is equivalent to using the augmented reality device as a control peripheral for the target device, providing the target device with voice interaction capabilities at low cost. This will improve the product competitiveness of augmented reality devices in practical applications.

[0077] In one optional configuration, the audio collection unit is adapted to acquire target audio data, the target audio data including audio commands.

[0078] The method further includes:

[0079] Obtaining the voice command by determining whether the voice command is present in the target voice data.

[0080] and / or transmitting the target audio data to the target device.

[0081] In this configuration, the augmented reality device further comprises a keyword detection unit and a control application unit. The keyword detection unit includes a voice keyword detection technology service for determining the type of connected target device and identifying voice commands. The control application unit includes a voice command control application for determining the type of command and sending different types of commands to different response entities. The target voice data includes the voice command and / or voice data generated when the augmented reality device outputs multimedia data.

[0082] It can be understood as follows: In this application, a voice command refers to a command input to an augmented reality device in the form of voice, and is a type of command data. The target voice data collected by the microphone array sensor may be command data. Furthermore, the voice data acquired by the microphone array sensor may be voice data generated when multimedia data is output. For example, when answering a phone call through an augmented reality device, the voice data collected by the microphone array sensor may be the content of the call. Alternatively, when projecting a movie through an augmented reality device, the voice data collected by the microphone array sensor may be the audio content of the movie.

[0083] The augmented reality device collects target voice data via a microphone array sensor, and identifies whether the target voice data includes voice commands and / or voice data generated when the augmented reality device outputs multimedia data. If it is identified that the target voice data includes both voice commands and voice data generated when the augmented reality device outputs multimedia data, it may separate the two types of voice data and perform different processes on the two types of voice data.

[0084] When the target voice data includes a voice command, the augmented reality device uses the keyword detection unit to determine different types of target devices and use different response modes accordingly. Specifically, when the target device is a first type, the target device responds to the voice command. When the target device is a second type or the target device is not detected, the keyword detection unit identifies the voice command, and the control application unit determines the type of the command and selects different response entities to respond to the voice command accordingly.

[0085] If the target audio data includes audio data generated when the augmented reality device outputs multimedia data, the audio data can be sent to the target device, and the target device can perform processing on the audio data, such as recording and word recognition.

[0086] In the above configuration, transmitting the collected target audio data to the augmented reality device and / or target device facilitates engineering implementation and ensures the target audio data can be used correctly by the augmented reality device and / or target device.

[0087] In one optional configuration, obtaining the voice command by determining whether a voice command is present in the target voice data; And / or, transmitting the target audio data to the target device includes:

[0088] Noise reduction processing is performed on the target voice data to obtain noise-reduced target voice data.

[0089] The voice command is acquired by determining whether or not a voice command is present in the target voice data after noise reduction.

[0090] And / or, the noise-reduced target audio data is transmitted to the target device.

[0091] In this configuration, the augmented reality device further includes a noise reduction unit, which performs noise reduction processing on the target voice data to obtain noise-reduced target voice data. The augmented reality device collects target voice data through a microphone array sensor, and the collected target voice data is subjected to noise reduction processing by the noise reduction unit to obtain noise-reduced target voice data. The noise-reduced target voice data is sent to the augmented reality device and / or the target device. If the target voice data includes a voice command, the augmented reality device uses a keyword detection unit to determine different target device types and use different response modes accordingly. If the target device type is a first type, the target device responds to the voice command. If the target device type is a second type or the target device is not detected, the keyword detection unit identifies the voice command, and the control application unit determines different command types and selects different response entities to respond to the voice command.

[0092] If the target audio data includes audio data generated when the augmented reality device outputs multimedia data, the noise-reduced audio data is sent to the target device, and the target device can perform processing such as recording and word classification on the noise-reduced audio data.

[0093] By performing noise reduction processing on the collected target audio data, the noise-reduced target audio data becomes audio data that contains no noise or less noise, thereby improving the quality of the audio data sent to the augmented reality device and / or target device, thereby achieving an accurate response to the target audio data.

[0094] In one optional configuration, obtaining the voice command by determining whether a voice command is present in the noise-reduced target voice data; And / or, transmitting the noise-reduced target audio data to the target device includes:

[0095] acquiring first target sound data and second target sound data based on the noise-reduced target sound data; determining whether a voice command is present in the first target voice data to obtain the voice command; and / or transmitting the second target audio data to the target device.

[0096] In this configuration, the augmented reality device further includes a copy unit that copies the noise-reduced target audio data to obtain two pieces of noise-reduced target audio data and then splits the two pieces of noise-reduced target audio data to obtain first and second target audio data. As shown in FIG. 4 , the augmented reality device sends the acquired microphone audio data (Audio In) to a specific noise reduction unit for directional noise cancellation to obtain audio data (Clean Audio) that is free of noise or contains relatively little noise. The audio data that is free of noise or contains relatively little noise is then passed through the copy unit to obtain the first and second target audio data. This can be understood as follows: Both the first and second target audio data are the same audio data as the clean audio that is free of noise or contains relatively little noise after noise removal. The second target audio data is directly sent to the Audio Out of the target device via a wired or wireless hardware connection, and the first target audio data is processed by the augmented reality device. Furthermore, if a voice command exists in the target voice data, the augmented reality device acquires the voice command and uses a keyword detection unit to determine different target device types and accordingly uses different response modes. Specifically, if the target device is a first type, the target device responds to the voice command. If the target device is a second type or if the target device is not detected, the keyword detection unit identifies the voice command and uses a control application unit to determine different command types and accordingly select different response entities to respond to the voice command.

[0097] The noise-reduced target voice data is duplicated and split to obtain first and second target voice data. The first target voice data is processed by the augmented reality device to perform voice command identification and to have different response entities respond. The second target voice data is sent to the target device so that the collected target voice data can be used by other applications on the target device. For example, when a user makes a call, the collected target voice data is used by a call application on the target device as the call content. The noise-reduced target voice data is duplicated and split to ensure that the augmented reality device and applications on the target device can obtain the required voice data without interfering with each other.

[0098] The voice interaction device provided in the embodiment of the present application includes, as shown in FIG. 5, an acquisition unit 501, a first determination unit 502, a second determination unit 503, a third determination unit 504 and a response unit 505.

[0099] The acquisition unit 501 is for acquiring the user's voice commands through the voice collection unit of the augmented reality device.

[0100] The first determining unit 502 is for determining the type of the target device connected to the augmented reality device.

[0101] The second determining unit 503 is for determining, if the target device is of a first type, the target response mode to the voice command as a first response mode.

[0102] The third determining unit 504 is for determining the target response mode to the voice command as a second response mode if the target device is of a second type or if no target device is detected.

[0103] The response unit 505 responds to the voice command using the target response mode.

[0104] In one optional configuration, the response unit 505 passes the voice command to the target device to cause the target device to respond if the target response mode is a first response mode, and determines whether the augmented reality device or the target device responds based on the type of the voice command if the target response mode is a second response mode.

[0105] In one optional configuration, the response unit 505 is configured to instruct the augmented reality device to respond to the voice command if the voice command is a command of a first type, and to send the voice command to the target device to respond if the voice command is a command of a second type.

[0106] In one optional configuration, if the voice command is a second type command, the response unit 505 converts the voice command into a target type command and sends the target type command to the target device for response.

[0107] In one optional configuration, the sound collection unit is configured to collect target sound data, the target sound data including sound commands.

[0108] The apparatus further includes:

[0109] The voice data unit obtains a voice command by determining whether the target voice data includes a voice command, and / or transmits the target voice data to the target device. In one optional configuration, the device further includes a noise reduction unit and a voice data unit, the noise reduction unit being adapted to perform noise reduction processing on the target voice data to obtain noise-reduced target voice data, and the voice data unit being adapted to determine whether the noise-reduced target voice data includes a voice command to obtain a voice command and / or to transmit the noise-reduced target voice data to the target device.

[0110] In one optional configuration, the device further includes a replica division unit that obtains first target voice data and second target voice data based on the noise-reduced target voice data, and correspondingly, the voice data unit further obtains a voice command by determining whether the first target voice data includes a voice command and / or transmits the second target voice data to the target device.

[0111] It should be noted that the principle of the problem solved by the voice interaction device in the embodiment of the present application is similar to that of the voice interaction method, so the implementation process, implementation principle and beneficial effects of the device can all refer to the implementation process, implementation principle and beneficial effects of the method, and redundant descriptions will not be repeated.

[0112] An embodiment of the present application provides an augmented reality device, which includes at least the voice interaction apparatus described in the present application.

[0113] According to an embodiment of the present application, the present application further provides an electronic device.

[0114] 6 is a block schematic diagram illustrating an exemplary electronic device 600 that may be used to practice embodiments of the present application. The electronic device may refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants (PDAs), servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants (PDAs), cell phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the practice of the present application as described and / or claimed herein.

[0115] 6, the electronic device 600 includes an arithmetic unit 601, which performs various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data necessary for the operation of the electronic device 600. The arithmetic unit 601, the ROM 602, and the RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0116] A number of components in the electronic device 600 are connected to an I / O interface 605, including input units 606 such as a keyboard, a mouse, etc., output units 607 such as various types of displays, speakers, etc., storage units 608 such as a disk, an optical disk, etc., and communication units 609 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 enables the electronic device 600 to exchange information / data with other devices via a computer network, for example the Internet, and / or various telecommunication networks.

[0117] The computing unit 601 can be various general-purpose and / or special-purpose process assemblies with processing and computing capabilities. Examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processing unit (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, some or all of the computer program can be loaded and / or installed into the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, it can perform one or more steps of the voice interaction method described above. Alternatively, in other embodiments, the computing unit 601 is arranged to implement voice interaction methods in any other manner (eg, via firmware).

[0118] Various embodiments of the systems and techniques described herein above may be realized in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include one or more computer software programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special purpose or a general purpose programmable processor, and that can receive data and commands from, and send data and commands to, a storage system, at least one input device, and at least one output device.

[0119] The program code used to implement the methods of the present application can be written in any combination of one or more programming languages. These program codes are provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code can run entirely on the machine, partially on the machine, as a stand-alone software package on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] It should be understood that various forms of the flow shown in the preceding paragraph may be used, and the order of steps may be changed, steps may be added, or steps may be deleted. For example, the steps described in this application may be performed in parallel, sequentially, or in a different order. As long as the desired results of the technical configuration disclosed in this application are achieved, the present application is not limited thereto.

[0121] Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly specifying the number of technical features shown. Thus, a feature qualified as "first" or "second" explicitly or implicitly includes at least one of the feature. In the specification of this application, the term "plurality" means two or more than two, unless otherwise defined.

[0122] The above description is merely a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application are included in the scope of protection of the present application. Therefore, the scope of protection of the present application is determined by the scope of protection of the claims.

Claims

1. 1. A voice interaction method, comprising: acquiring a voice command of the user by a voice collection unit of the augmented reality device; determining a type of target device connected to the augmented reality device; If the target device is of a first type, determining a target response mode to the voice command as a first response mode; determining a target response mode to the voice command as a second response mode if the target device is of a second type or if no target device is detected; Responding to the voice command using the target response mode; The voice interaction method, comprising:

2. Responding to the voice command using the target response mode includes: If the target response mode is a first response mode, sending the voice command to the target device to respond; if the target response mode is a second response mode, determining whether the augmented reality device or the target device will respond based on the type of the voice command; 2. The method of claim 1, comprising:

3. determining whether the augmented reality device or the target device will respond based on the type of voice command; if the voice command is a first type of command, the augmented reality device responds to the voice command; If the voice command is a second type command, transmitting the voice command to the target device and causing the target device to respond.

3. The method of claim 2, comprising:

4. If the voice command is a second type command, transmitting the voice command to the target device to respond to the voice command includes:

4. The method of claim 3, wherein if the voice command is a second type command, the voice command is converted into a target type command and the target type command is sent to the target device for response.

5. the sound collection unit is used to acquire target sound data, the target sound data including a sound command; The method further comprises: determining whether a voice command is present in the target voice data, thereby obtaining the voice command; and / or transmitting the target audio data to the target device.

2. The method of claim 1, comprising:

6. determining whether a voice command is present in the target voice data, thereby obtaining the voice command; and / or transmitting the target audio data to the target device performing noise reduction processing on the target voice data to obtain noise-reduced target voice data; determining whether a voice command is present in the noise-reduced target voice data, thereby obtaining the voice command; and / or transmitting the noise-reduced target audio data to the target device; 6. The method of claim 5, comprising:

7. determining whether a voice command is present in the noise-reduced target voice data, thereby obtaining the voice command; and / or transmitting the noise-reduced target audio data to the target device, acquiring first target sound data and second target sound data based on the noise-reduced target sound data; determining whether a voice command is present in the first target voice data to obtain the voice command; and / or transmitting the second target audio data to the target device.

7. The method of claim 6, comprising:

8. A voice interaction device, comprising: an acquisition unit, a first determination unit, a second determination unit, a third determination unit and a response unit; The acquisition unit acquires a voice command of a user through a voice collection unit of an augmented reality device; the first determining unit is configured to determine a type of a target device connected to the augmented reality device; the second determination unit determines a target response mode to the voice command as a first response mode when the target device is a first type; the third determination unit determines a target response mode to the voice command as a second response mode when the target device is of a second type or when no target device is detected; The voice interaction device, wherein the response unit responds to the voice command using the target response mode.

9. An augmented reality device comprising at least the voice interaction device according to claim 8.

10. 1. An electronic device comprising: at least one processor; and a memory communicatively coupled to the at least one processor, wherein: the memory stores instructions executable by the at least one processor; The electronic device, characterized in that the commands, when executed by the at least one processor, enable the at least one processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech dialogue method, apparatus and system

    JP2019204074A

  • Electronic device, server and method of controlling the same

    JP2020028129A

  • Method, system and apparatus for providing a composite graphical assistant interface for controlling connected devices

    JP2021523456A

  • Systems and methods for augmented reality

    JP2022542363A

  • Operational command boundaries

    US20220358914A1