Voice interaction method, apparatus, and related devices

By determining the type of connected device and using appropriate response modes, augmented reality devices can perform voice interaction effectively, addressing the limitations of existing systems and enhancing their versatility.

JP7838876B2Active Publication Date: 2026-04-01HANGZHOU LINGBAN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Augmented reality devices are limited to communicating only with compatible terminal devices, lacking flexibility and versatility in voice interaction.

Method used

The augmented reality device determines the type of connected target device and employs different response modes to handle voice commands, enabling interaction even when not connected to a terminal or connected to different types of terminals.

Benefits of technology

This approach enhances the versatility and flexibility of augmented reality devices, allowing voice interaction regardless of the connected device type, expanding their application range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838876000001
    Figure 0007838876000001
  • Figure 0007838876000002
    Figure 0007838876000002
  • Figure 0007838876000003
    Figure 0007838876000003
Patent Text Reader

Abstract

The present invention provides a voice interaction method, apparatus, and related device, which includes: acquiring a user's voice command through a voice collection unit of an augmented reality device; determining a type of a target device connected to the augmented reality device; determining a target response mode for the voice command as a first response mode if the target device is a first type; determining a target response mode for the voice command as a second response mode if the target device is a second type or if the target device is not detected; and responding to the voice command using the target response mode, thereby providing technical support for normal voice interaction even when no target device is connected to the augmented reality device or a different target device is connected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross-Reference to Related Applications) This application claims the priority of the Chinese patent application with application number 202310096785.6, filed with the China National Intellectual Property Administration on January 12, 2023, and the entire content thereof is incorporated herein by reference.

[0002] The present invention relates to the field of voice interaction, and particularly to a voice interaction method, apparatus, and related devices.

Background Art

[0003] Generally, an extended reality device is used to display or collect sensor data. In most cases, it is necessary to connect to a terminal device that is compatible (conforms) with the extended reality device for use.

Disclosure of the Invention

[0004] This application provides a voice interaction method, apparatus, and related devices, which at least solve the above technical problems existing in the prior art.

[0005] In a first aspect, this application provides a voice interaction method. The method includes: obtaining a user's voice command by a voice collection unit of an extended reality device; determining the type of a target device connected to the extended reality device; when the target device is of a first type, determining a target response mode for the voice command as a first response mode; when the target device is of a second type, or when no target device is detected, determining a target response mode for the voice command as a second response mode; responding to the voice command using the target response mode.

[0006] In the above configuration, responding to the voice command using the target response mode is: If the target response mode is the first response mode, the voice command is sent to the target device to cause it to respond. If the target response mode is a second response mode, the process includes determining which of the augmented reality device and the target device will respond based on the type of voice command.

[0007] In the above configuration, determining which of the augmented reality device and the target device will respond based on the type of voice command is: If the voice command is of a first type, the augmented reality device responds to the voice command. If the voice command is of a second type, the method includes sending the voice command to the target device to elicit a response.

[0008] In the above configuration, if the voice command is a second type of command, sending the voice command to the target device to elicit a response means that If the voice command is a second type of command, the voice command is converted into a target type command, and the target type command is sent to the target device to elicit a response.

[0009] In the above configuration, the voice acquisition unit is used to acquire target voice data, and the target voice data includes voice commands. The above method further, A voice command exists in the aforementioned target voice data. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or, including transmitting the target audio data to the target device.

[0010] In the above configuration, a voice command exists in the target voice data. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or, transmitting the target audio data to the target device is The process involves applying noise reduction processing to the target audio data and obtaining the target audio data after noise reduction. Voice commands are present in the target audio data after noise reduction. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or, transmit the noise-reduced target audio data to the target device.

[0011] In the above configuration, a voice command exists in the target voice data after noise reduction. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or, transmitting the noise-reduced target audio data to the target device is Based on the noise-reduced target audio data, the first target audio data and the second target audio data are acquired. A voice command exists in the aforementioned first target voice data. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or, including transmitting the second target audio data to the target device.

[0012] In a second embodiment, this application provides a voice interaction device, the device comprising an acquisition unit, a first determination unit, a second determination unit, a third determination unit, and a response unit. The aforementioned acquisition unit acquires the user's voice commands using the voice collection unit of the augmented reality device. The first determination unit determines the type of target device connected to the augmented reality device, The second determination unit determines that the target response mode to the voice command is the first response mode when the target device is of the first type. The third determination unit determines that the target response mode to the voice command is the second response mode if the target device is of the second type or if no target device is detected. The response unit responds to the voice command using the target response mode.

[0013] In a third embodiment, this application provides an augmented reality device, the augmented reality device comprising at least the voice interaction device described in this application.

[0014] In a fourth embodiment, the present application provides an electronic device comprising at least one processor and a memory communicably connected to the at least one processor, wherein the memory stores commands executable by the at least one processor, and the commands are executed by the at least one processor so that the at least one processor can perform the method described herein.

[0015] In this application, obtaining a user's voice command by a voice collection unit of an extended reality device, determining the type of a target device connected to the extended reality device, and when the target device is of a first type, determining a target response mode for the voice command as a first response mode; when the target device is of a second type, or when no target device is detected, determining a target response mode for the voice command as a second response mode; and responding to the voice command using the target response mode. Thereby, technical support is provided for normal voice interaction even when no target device is connected to the extended reality device or when different target devices are connected.

[0016] It should be understood as follows. The content described in this part is not intended to identify the points or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become more understandable from the following specification.

Brief Description of Drawings

[0017] By reading the detailed description of the following specification while referring to the drawings, the above and other objects, features, and advantages of the exemplary embodiments of this application will become more understandable. In the drawings, some embodiments of this application are shown illustratively and non - limitatively, and are not intended to limit the present invention. Among them, In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts.

[0018] [Figure 1] FIG. 1 is a schematic diagram of the implementation flow of the voice interaction method according to the embodiment of this application. [Figure 2] FIG. 2 is a schematic diagram of the implementation flow of different target response modes according to the embodiment of this application. [Figure 3] FIG. 3 is a second schematic diagram of the implementation flow of different target response modes according to the embodiment of this application. [Figure 4] Figure 4 is a schematic diagram showing the data flow of an augmented reality device terminal according to an embodiment of this application. [Figure 5] Figure 5 is a schematic diagram of the configuration structure of the voice interaction device in an embodiment of this application. [Figure 6] Figure 6 is a schematic diagram of the configuration structure of an electronic device in an embodiment of the present invention. [Modes for carrying out the invention]

[0019] To make the purpose, features, and advantages of this application clearer and easier to understand, the technical configurations in the embodiments of this application will be clearly and completely described below, along with the drawings in the embodiments of this application. Needless to say, the embodiments described are only a part of the embodiments of this application, not all of them. All other embodiments that a person skilled in the art can obtain based on the embodiments of this application without any creative work are all within the scope of protection of this application.

[0020] In related technologies, augmented reality devices can only communicate with compatible terminal devices, which prevents them from achieving the flexibility and versatility required for augmented reality devices.

[0021] The following can be understood: Augmented reality (AR) devices are currently the mainstream wearable devices. Intelligent voice interaction, as the mainstream method of interaction for augmented reality devices, allows for hands-free, easy, and rapid input or control of the augmented reality device. Considering factors such as appearance, wearability, power consumption, and heat generation, augmented reality devices are typically used only as displays (projections) and sensor data acquisition functions (e.g., image, sound, inertial measurement units, etc.), rather than as complex computing units. Based on this, augmented reality devices can only connect to compatible terminal devices via wired or wireless methods, and transmit the collected sensor data to the terminal device for algorithmic calculations. The functionality of augmented reality devices is undoubtedly enhanced if they can perform voice interaction successfully even when not connected to a terminal device or when connected to another terminal device. This provides a foundation for the wide application of augmented reality devices in daily life.

[0022] The technical configuration of the embodiment of this application relates to the configuration of voice interaction. The augmented reality device can perform voice interaction with different types of target devices, as well as voice interaction when not connected to a target device, demonstrating the versatility and flexibility of the augmented reality device. Based on the acquired voice command and the type of target device connected to the augmented reality device, it can respond to voice interaction using different target response modes. Technical support is provided to enable successful voice interaction even when the augmented reality device is not connected to a target device, or when connected to a different target device, thereby expanding the range of applications for the augmented reality device.

[0023] The voice interaction method of the embodiment of this application will be described in detail below.

[0024] Embodiments of this application provide a voice interaction method. As shown in Figure 1, the method includes the following:

[0025] S101: The voice acquisition unit of the augmented reality device acquires the user's voice commands. In this step, the augmented reality device is an electronic device capable of AR interaction. The augmented reality device can be a smart wearable device, including but not limited to smart glasses and smartwatches. In this application, the augmented reality device will be described using separated AR glasses as an example.

[0026] The augmented reality device includes a voice acquisition unit, such as a microphone. In this step, the user's voice commands are obtained by collecting the voice commands that the user makes to the augmented reality device via the microphone.

[0027] It can be understood as follows: A microphone includes a microphone array sensor (MicArray) for acquiring user voice commands uttered to an augmented reality device.

[0028] In practical applications, users can issue voice commands to augmented reality devices even when they are not in use. For example, when the augmented reality device's screen is off, they can issue voice commands such as "Turn on the screen" or "Turn off the power."

[0029] Furthermore, users can issue voice commands to the augmented reality device while it is in use, for example, when playing a movie or audio such as music. In other words, when the augmented reality device is outputting multimedia data, the user's voice commands are captured by the augmented reality device's voice acquisition unit.

[0030] It can be understood as follows: Augmented reality devices typically output multimedia data, such as images and audio, during the process of AR interaction. When an augmented reality device and a target device are connected, videos from the target device can be projected onto the augmented reality device and output; for example, a movie from the target device can be projected onto the augmented reality device and output. In such cases, multimedia data can also refer to videos from the target device that can be projected by the augmented reality device. Furthermore, audio from the target device can also be output via the augmented reality device; for example, the target device can be used to make voice calls or listen to voicemails. In such cases, multimedia data can also refer to audio from the target device that can be output by the augmented reality device.

[0031] In other words, when connected to a target device, the augmented reality device of this application becomes a video and audio output device, acting as an alternative target device. Such alternatives mainly include the following: In some application scenarios, using a target device as a video and audio output device is far less convenient and less effective than using an augmented reality device as a video and audio output device. For example, in projection applications, an immersive experience can be provided to the wearer of the augmented reality device by projecting projectable videos from the target device using the augmented reality device. Furthermore, in situations where it is inconvenient to answer calls from a target device, such as in crowded viewing environments, for example, during rush hour on the subway, the user can answer calls by wearing split AR glasses. Split AR glasses can be integrated into the wearer's regular eyeglasses. By answering calls using split AR glasses, the situation of not being able to take a mobile phone out of one's pocket in a crowded environment can be avoided.

[0032] Users can issue voice commands to the augmented reality device as needed. Specifically, the augmented reality device collects voice commands using a microphone.

[0033] For example, if the augmented reality device is split-frame AR glasses, and the current scene is one in which the user is playing audio or video using the split-frame AR glasses, then the multimedia data output from the split-frame AR glasses is the audio / video information being played by the user. When the user issues a voice command such as "turn up the volume," the microphone collects this voice command from the user, thereby obtaining a voice command for audio / video playback. In response to the voice command, the volume of the currently playing audio / video is increased.

[0034] S102: Determine the type of target device connected to the augmented reality device.

[0035] In this step, the target device can be any device capable of voice interaction with the augmented reality device. Examples include smartphones, computers, and personal digital assistants. The target device connected to the augmented reality device in this application can be of a different type. For example, the target device may be a self-developed terminal, i.e., a terminal compatible with the augmented reality device. This should be understood as follows: After the augmented reality device is manufactured by a manufacturer, there is usually a self-developed terminal manufactured by that manufacturer that is compatible with the augmented reality device. This self-developed terminal can be understood as a standard-sized terminal with no display screen, or a relatively small display screen, but with processing power. By connecting a self-developed terminal compatible with the augmented reality device, the intelligent voice interaction function of the augmented reality device is realized.

[0036] The target device can be other types of terminals, such as a third-party mobile phone or a third-party computer. In this application, the intelligent voice interaction function of the augmented reality device is realized by connecting the augmented reality device with other types of terminals. For example, if the augmented reality device is a split AR glasses, the split AR glasses can connect not only to a terminal that is compatible with it, but also to a third-party terminal. In this application, whether the split AR glasses are connected to a terminal that is compatible with it or to a third-party terminal, the split AR glasses can perform intelligent voice interaction with the device to which they are connected.

[0037] In practical applications, voice interaction is the dominant interaction method for augmented reality devices, and is typically limited to proprietary devices. That is, if a user wants to use an augmented reality device, they need to purchase a proprietary device compatible with that device in order to use voice interaction correctly. In such cases, on the one hand, if the user already owns a mobile phone, they may not purchase an additional compatible device due to cost and portability considerations. On the other hand, when connecting an augmented reality device to a personal computer (PC), the PC is a processing unit and does not connect to other devices. In these two scenarios, the voice interaction function of the augmented reality device becomes unavailable, severely limiting the applications of voice interaction in AR glasses.

[0038] In this application, the augmented reality device can connect to different types of terminals, and by determining the type of target device connected to the augmented reality device, the augmented reality device can realize voice interaction functions. The augmented reality device can realize basic voice interaction functions, such as volume adjustment and brightness adjustment, using its own voice command control application without connecting to any type of terminal.

[0039] In practical applications, augmented reality devices can connect to different types of terminals or remain unconnected to any terminal. This allows users to purchase only the augmented reality device without having to buy compatible, proprietary terminals. Users can utilize the voice interaction features of the augmented reality device by connecting it directly to their mobile phone or computer.

[0040] S103: If the target device is of the first type, the target response mode to the voice command is determined to be the first response mode.

[0041] In this application, the augmented reality device can be connected to or linked to different types of terminals. The response mode to voice commands also differs depending on the type of target device to which the augmented reality device is connected.

[0042] In this application, there are two types of target devices. The first type refers to a terminal that has an internal voice keyword detection technology service, such as the in-house developed terminal mentioned above. The second type refers to a terminal that does not have an internal voice keyword detection technology service, such as the third-party mobile phone or computer mentioned above.

[0043] For example, assuming that the target device connected to the augmented reality device is the in-house developed terminal, the target response mode corresponding to the voice command can be Mode A (first response mode). If the target device connected to the augmented reality device is another type of terminal (e.g., a third-party mobile phone, computer, etc.), the target response mode corresponding to the voice command can be Mode B (second response mode).

[0044] During execution, the augmented reality device identifies whether a target device is connected to it. If no device is connected, the target response mode for the voice command is determined to be the second response mode. If a target device is connected to the augmented reality device, the device obtains the identifier of the connected device and determines whether the connected device is of the first or second type based on that identifier. If the identifier of the connected device is identifier A, and identifier A represents a device of the first type, the connected device is determined to be of the first type. If the identifier of the connected device is identifier B, and identifier B represents a device of the second type, the connected device is determined to be of the second type.

[0045] In this application, two response modes are pre-configured based on whether the terminal connected to the augmented reality device is a terminal compatible with the augmented reality device or a third-party terminal not compatible with the augmented reality device. One of the response modes is used when the terminal connected to the augmented reality device is a terminal compatible with the augmented reality device. The other response mode is used when the terminal connected to the augmented reality device is a third-party terminal. Two different types of terminals and the modes used for each type of terminal are pre-configured into a single correspondence. In implementation, based on the type of terminal connected to the augmented reality device, the corresponding mode for that type of terminal is retrieved from the correspondence and set as the target response mode for responding to voice commands.

[0046] S104: If the target device is of the second type, or if no target device is detected, the target response mode for the voice command is determined to be the second response mode.

[0047] Since the second type of target device is a terminal that does not include the voice keyword detection technology service, the augmented reality device of this application utilizes the voice keyword detection technology service to enable normal voice interaction when connected to the second type of target device, so that voice interaction can be performed even when the augmented reality device is connected to the second type of target device. In addition, since the augmented reality device of this application is equipped with the voice keyword detection technology service, even when the augmented reality device is not connected to the target device, i.e., when the target device is not detected, the augmented reality device can use the corresponding target response mode to perform basic voice interactions such as adjusting brightness and volume.

[0048] In this application, by deploying a low-power voice keyword detection technology service in an augmented reality device, it is possible to identify voice commands with the help of the low-power voice keyword detection technology service without significantly increasing the power consumption or heat generation of the augmented reality device.

[0049] By implementing different target response modes for different types of devices, augmented reality devices can maintain their voice interaction capabilities regardless of the type of device they are connected to.

[0050] S105: The system responds to the voice command using the target response mode.

[0051] Based on the type of connected target device, the system responds to voice commands using different target response modes. Specifically, it analyzes the type of voice command using different target response modes and then processes that type of command. For example, if the voice command is analyzed as "increase volume," the output volume of the augmented reality device is increased. If the voice command is analyzed as "decrease screen brightness," the screen brightness of the augmented reality device is decreased. This enables the operation corresponding to the voice command.

[0052] In the configuration shown in S101-S105, the augmented reality device can communicate not only with terminals compatible with the augmented reality device and third-party terminals, but also perform voice interaction even when no terminal is connected. This demonstrates the flexibility and versatility of the augmented reality device.

[0053] Furthermore, in this application, a target response mode is determined based on the type of target device connected to the augmented reality device, and a response to voice commands is then implemented using this target response mode. Depending on the type of target device connected to the augmented reality device, different target response modes can be used to respond to voice commands. This ensures that the augmented reality device maintains its voice interaction functionality even when connected to different types of terminals. This provides technical support for the augmented reality device to perform voice interactions correctly even when connected to different target devices.

[0054] In one optional configuration, responding to the voice command using the target response mode includes the following:

[0055] If the target response mode is the first response mode, the voice command is sent to the target device to cause it to respond.

[0056] If the target response mode is the second response mode, the system determines whether the augmented reality device or the target device will respond based on the type of voice command.

[0057] Of these, the type of voice command indicates whether the voice command is a type of command to which the augmented reality device responds, or a type of command to which the target device responds.

[0058] As shown in Figure 2, in this application, there are mainly two types of responding entities: a target device and an augmented reality device. The augmented reality device mainly includes a voice keyword detection technology service and a voice command control application. When the target response mode is the first response mode, that is, when it is determined that the target response mode is the first response mode based on the type of target device connected to the augmented reality device being a first type, such as a company-developed terminal, a voice command is sent to the target device's operating system to perform full voice control. When the target response mode is the second response mode, that is, when it is determined that the target response mode is the second response mode based on the type of target device connected to the augmented reality device being a second type device, such as a third-party smartphone or computer, or when the target device has not been detected, a voice command is converted into an Event ID and notified to the voice command control application in the augmented reality device. The voice command control application determines whether the command is a control command for the augmented reality device itself. If the command is a control command for the augmented reality device itself, the voice command control application responds to the voice command, for example, by adjusting the volume or brightness. For example, if a voice command is a control command to a target device, the Event ID is converted to a Universal Keyboard (USB Keyboard, Universal SerialBus Keyboard) protocol ID, and the USB Keyboard ID is sent to the target device, which then responds to the voice command.

[0059] Specifically, as shown in Figure 3, when the voice keyword detection technology service of an augmented reality device is operational, the voice keyword detection technology service determines whether the target device connected to the augmented reality device is of type 1 or type 2. Depending on the type of connected target device, a different target response mode is used. When the target response mode is determined to be the first response mode, that is, based on the fact that the type of target device connected to the augmented reality device is type 1 (e.g., a proprietary terminal), the voice keyword detection technology service within the augmented reality device enters a dormant state, and voice commands are handed over to the target device's operating system, which then performs full voice control. Exemplarily, assuming that the target device connected to the augmented reality device is a proprietary terminal, since the proprietary terminal is equipped with the voice keyword detection technology service, there is usually a set of predefined voice commands on the proprietary terminal. For example, the volume up command adjusts the volume of multimedia data, the brightness up command adjusts the display brightness of the screen, and the mode switching command adjusts the display mode of multimedia data (e.g., from normal mode to 3D mode, or from 3D mode to normal mode).

[0060] When a user issues a corresponding voice command, the augmented reality device determines that the connected target device is a proprietary terminal and sends the collected voice command to the target device's operating system, enabling full voice control. A voice keyword detection technology service running on the proprietary terminal passes the voice command to the proprietary terminal. The proprietary terminal, specifically the application processing unit, responds to the voice command and performs control actions corresponding to the voice command, such as adjusting the volume, brightness, or switching modes.

[0061] If the voice command acquired by the augmented reality device is a voice adjustment command for adjusting the audio of the multimedia data output by the augmented reality device, and the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, the voice adjustment command will be passed on to the proprietary terminal, which will then adjust the audio of the multimedia data output by the augmented reality device.

[0062] If the voice command received by the augmented reality device is a command to adjust the display brightness of the multimedia data output by the augmented reality device, and the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, then the command to adjust the display brightness is passed to the proprietary terminal, which then adjusts the display brightness of the augmented reality device's screen.

[0063] If the voice command acquired by the augmented reality device is a command to adjust the display mode of the multimedia data output by the augmented reality device, for example, if the target device connected to the augmented reality device is a proprietary terminal compatible with the augmented reality device, the command to adjust the display brightness will be passed to the proprietary terminal, which will then adjust the display mode of the augmented reality device.

[0064] If the target response mode is determined to be the second response mode, that is, if the type of target device connected to the augmented reality device is a second type, such as a third-party mobile phone or computer, or if the target device has not been detected, the augmented reality device's voice keyword detection technology service will remain operational.

[0065] In one optional configuration, when the target response mode is the second response mode, determining which of the augmented reality device and the target device will respond based on the type of voice command includes the following:

[0066] If the voice command is of the first type, the augmented reality device responds to the voice command.

[0067] If the voice command is of a second type, the voice command is sent to the target device to elicit a response.

[0068] In this application, when the target response mode is the second response mode, the voice commands include two types. The first type of command is one that the augmented reality device can respond to, such as the volume adjustment command, the display brightness adjustment command, and the display mode adjustment command. The second type of command is one that the augmented reality device cannot respond to, but the target device can, such as the "back" command, the "confirm" command, and the "main menu" command.

[0069] In this application, when the target response mode is the second response mode, if the voice command is a command that the augmented reality device can respond to, such as a volume adjustment command, a display brightness adjustment command, and a display mode adjustment command, the augmented reality device will respond to the voice command. If the voice command is, for example, a "back" command, a "confirm" command, or a "main menu" command, the voice command is passed to the target device, and the target device will respond to the voice command.

[0070] Based on this, in this application, in the second response mode, different responding entities respond to voice commands depending on the type of voice command, thereby avoiding the problem of excessive load caused by only one of the responding entities responding to all voice commands. In this application, different types of voice commands are responded to by different responding entities, making it easy to implement in engineering and highly feasible.

[0071] In one optional configuration, if the voice command is a second type of command, sending the voice command to the target device to elicit a response includes:

[0072] If the voice command is a second type of command, the voice command is converted into a target type command, and the target type command is sent to the target device to elicit a response.

[0073] In this application, for example, if a voice command is one that cannot be responded to by an augmented reality device but can be responded to by a target device, it is necessary to convert the voice command into a type that the target device can recognize or respond to, thereby enabling a response to the voice command.

[0074] For example, assuming the target device connected to the augmented reality device is a third-party mobile phone or computer, when the voice keyword detection technology service in the augmented reality device identifies a voice command, it converts the voice command into an Event ID and notifies the voice command control application in the augmented reality device. Upon receiving the Event ID of the command, the voice command control application matches the Event ID of the command with a number corresponding to a predefined command function and provides corresponding feedback based on the predefined command function represented by the corresponding number. These predefined command functions are command functions represented by different numbers within a predefined set of commands; for example, the "increase volume" function is number 1, the "increase brightness" function is number 2, and the "switch mode" function is number 3.

[0075] If the command corresponding to the command event ID (Event Id) is of the first type, that is, if the command event ID (Event Id) corresponds to a number corresponding to a predefined command function on the augmented reality device, that is, if the command corresponding to the command event ID (Event Id) is a control-related command that the augmented reality device can respond to, such as a hardware-related command for the augmented reality device, such as volume adjustment, brightness adjustment, display switching, or mode switching, then the voice command control application executes the corresponding control operation directly via the System Application Programming Interface (API). If the command corresponding to the command event ID (Event Id) is of the second type, that is, if the command event ID (Event Id) corresponds to a number corresponding to a predefined command function on the target device, that is, if the command corresponding to the command event ID is a command that requires a response from the target device, then the command event ID (Event Id) is converted to a USB Keyboard protocol ID and sent to the connected target device. The target device responds to the USB Keyboard protocol ID and performs operations on the multimedia data output from the augmented reality device. For example, the "Back" command can be defined as the "F1" key on the target device's keyboard. When the user presses the "F1" key, the multimedia data output by the augmented reality device returns to the previous step. Alternatively, the "Main Menu" command can be defined as the "F2" key on the target device's keyboard. When the user presses the "F2" key, the augmented reality device pops up the main menu interface. Or, the "Confirm" command can be defined as the "Enter" key on the target device's keyboard. When the user presses the "Enter" key, the user can confirm their selection.

[0076] The execution process described above by the augmented reality device in this application can be realized by a high-performance dedicated chip mounted on the augmented reality device. The inclusion of a high-performance dedicated chip improves computational efficiency and enables rapid response to voice commands. Furthermore, this application utilizes a low-power voice keyword detection technology service, offering significant advantages such as low computational power requirements and low energy consumption. By incorporating AR functionality into the augmented reality device without significantly increasing power consumption, the control capabilities of the augmented reality device itself are greatly improved. Moreover, voice commands are customizable, highly scalable, and overcome the limitations of key input. By using the general-purpose USB keyboard protocol, the augmented reality device converts identified voice commands into target types that the target device can recognize or respond to. The target device's response to voice commands can be seen as the user responding to the voice command by operating keys (e.g., the "F1" key, the "F2" key, and the "Enter" key on the target device's keyboard). In this configuration, the augmented reality device is used as a control peripheral for the target device, providing voice interaction functionality to the target device at low cost. This will improve the competitiveness of augmented reality devices in terms of practical application.

[0077] In one optional configuration, the voice acquisition unit is used to acquire target voice data, which includes voice commands.

[0078] The above method further includes the following:

[0079] A voice command exists in the aforementioned target voice data. mosquito The voice command is obtained by determining whether or not it is true.

[0080] and / or transmit the target audio data to the target device.

[0081] In this configuration, the augmented reality device further comprises a keyword detection unit and a control application unit. The keyword detection unit includes a voice keyword detection technology service for determining the type of connected target device and identifying voice commands. The control application unit includes a voice command control application for determining the type of command and sending different types of commands to different responding entities. Target voice data includes voice commands and / or voice data generated when the augmented reality device outputs multimedia data.

[0082] It can be understood as follows: In this application, a voice command refers to a command input to an augmented reality device in the form of voice, and is a type of command data. The target voice data collected by the microphone array sensor may be command data. Furthermore, the voice data acquired by the microphone array sensor may be voice data generated when multimedia data is output. For example, when answering a call via an augmented reality device, the voice data collected by the microphone array sensor may be the content of the call. Alternatively, when projecting a movie via an augmented reality device, the voice data collected by the microphone array sensor may be the audio content in the movie.

[0083] The augmented reality device collects target audio data via a microphone array sensor and identifies whether the target audio data contains voice commands and / or audio data generated when the augmented reality device outputs multimedia data. If it identifies that the target audio data contains both voice commands and audio data generated when the augmented reality device outputs multimedia data, it can also separate these two types of audio data and perform different processing on them.

[0084] If the target audio data contains a voice command, the augmented reality device uses a keyword detection unit to determine the type of target device and use a different response mode accordingly. Specifically, if the target device is of the first type, the target device responds to the voice command. If the target device is of the second type, or if no target device is detected, the keyword detection unit identifies the voice command, and the control application unit determines the type of command, thereby selecting a different responding entity to respond to the voice command.

[0085] If the target audio data includes audio data generated when an augmented reality device outputs multimedia data, the audio data can be sent to the target device, which can then perform processing on the audio data, such as recording or word sorting.

[0086] In the above configuration, transmitting the collected target audio data to the augmented reality device and / or target device facilitates engineering implementation and ensures the proper use of the target audio data by the augmented reality device and / or target device.

[0087] In one optional configuration, the voice command is obtained by determining whether or not a voice command exists in the target voice data. And / or transmitting the target audio data to the target device includes the following:

[0088] The process involves applying noise reduction processing to the target audio data and obtaining the target audio data after noise reduction.

[0089] Voice commands are present in the target audio data after noise reduction. mosquito The voice command is obtained by determining whether or not it is true.

[0090] and / or transmit the noise-reduced target audio data to the target device.

[0091] In this configuration, the augmented reality device further includes a noise reduction unit, which performs noise reduction processing on the target audio data to obtain noise-reduced target audio data. The augmented reality device collects target audio data via a microphone array sensor, and the collected target audio data is subjected to noise reduction processing by the noise reduction unit to obtain noise-reduced target audio data. The noise-reduced target audio data is transmitted to the augmented reality device and / or the target device. If the target audio data contains a voice command, the augmented reality device uses a keyword detection unit to determine different target device types and thereby uses different response modes. If the target device type is the first type, the target device responds to the voice command. If the target device type is the second type, or if no target device is detected, the keyword detection unit identifies the voice command, and the control application unit determines different command types and thereby selects a different responding entity to respond to the voice command.

[0092] If the target audio data includes audio data generated when an augmented reality device outputs multimedia data, the noise-reduced audio data can be sent to the target device, which can then perform processing such as recording and word-awareness analysis on the noise-reduced audio data.

[0093] By applying noise reduction processing to the collected target audio data, the noise-reduced target audio data becomes noise-free or contains very little noise, thereby improving the quality of the audio data transmitted to the augmented reality device and / or target device, and enabling accurate responses to the target audio data.

[0094] In one optional configuration, the target audio data after noise reduction contains the voice command. mosquito By determining whether or not, the aforementioned voice command is obtained. And / or transmitting the noise-reduced target audio data to the target device includes the following:

[0095] Based on the noise-reduced target audio data, the first target audio data and the second target audio data are acquired. A voice command exists in the aforementioned first target voice data. mosquito By determining whether or not, the aforementioned voice command is obtained. and / or transmit the second target audio data to the target device.

[0096] In this configuration, the augmented reality device further includes a duplication and splitting unit, which duplicates (copies) the noise-reduced target audio data to obtain two noise-reduced target audio data, and then splits the two noise-reduced target audio data to obtain the first target audio data and the second target audio data. As shown in Figure 4, the augmented reality device transmits the acquired microphone audio data Audio In to a specific noise reduction unit to perform directional noise cancellation and obtain Clean Audio, which is noise-free or contains relatively little noise. The audio data, which is noise-free or contains relatively little noise, is then processed by the duplication and splitting unit to obtain the first target audio data and the second target audio data. This can be understood as follows: Both the first target audio data and the second target audio data are the same audio data as the Clean Audio, which is noise-free or contains relatively little noise after noise reduction. The second target audio data is transmitted directly to the Audio Out of the target device by a wired or wireless method via hardware, and the first target audio data is processed by the augmented reality device. Furthermore, if a voice command is present in the target voice data, the augmented reality device acquires the voice command, and the keyword detection unit determines the type of target device, thereby using a different response mode. Specifically, if the target device is of the first type, the target device responds to the voice command. If the target device is of the second type, or if no target device is detected, the keyword detection unit identifies the voice command, and the control application unit determines the type of command, thereby selecting a different responding entity to respond to the voice command.

[0097] By duplicating and splitting the noise-reduced target audio data, first and second target audio data are obtained. The first target audio data is processed by the augmented reality device and configured to perform voice command identification and respond to different responding subjects. The second target audio data is transmitted to the target device so that the collected target audio data can be used by other applications on the target device. For example, when a user makes a call, the collected target audio data is used as the call content by the calling application on the target device. Duplicating and splitting the noise-reduced target audio data ensures that the augmented reality device and the applications on the target device can obtain the necessary audio data without interfering with each other.

[0098] As shown in Figure 5, the voice interaction device provided by the embodiment of this application includes an acquisition unit 501, a first determination unit 502, a second determination unit 503, a third determination unit 504, and a response unit 505.

[0099] The acquisition unit 501 acquires the user's voice commands using the voice collection unit of the augmented reality device.

[0100] The first determination unit 502 determines the type of target device connected to the augmented reality device.

[0101] The second determination unit 503 determines that the target response mode to the voice command is the first response mode if the target device is of the first type.

[0102] The third determination unit 504 determines that the target response mode for the voice command is the second response mode if the target device is of the second type or if no target device is detected.

[0103] The response unit 505 responds to the voice command using the target response mode.

[0104] In one optional configuration, the response unit 505, when the target response mode is the first response mode, passes the voice command to the target device and causes the target device to respond. When the target response mode is the second response mode, it determines whether the augmented reality device or the target device will respond based on the type of the voice command.

[0105] In one optional configuration, the response unit 505 is configured to instruct the augmented reality device to respond to the voice command if the voice command is of a first type, and to send the voice command to the target device to cause it to respond if the voice command is of a second type.

[0106] In one optional configuration, the response unit 505 converts the voice command to a target type command if the voice command is a second type command, and sends the target type command to the target device to receive a response.

[0107] In one optional configuration, the voice acquisition unit is configured to acquire target voice data, which includes voice commands.

[0108] The apparatus further includes the following:

[0109] The voice data unit obtains a voice command by determining whether or not the target voice data contains a voice command, and / or transmits the target voice data to the target device. In one optional configuration, the apparatus further includes a noise reduction unit and an audio data unit. The noise reduction unit is used to perform noise reduction processing on the target audio data and to obtain the noise-reduced target audio data. Correspondingly, the audio data unit is used to obtain audio commands by determining whether or not the noise-reduced target audio data contains audio commands, and / or to transmit the noise-reduced target audio data to the target device.

[0110] In one optional configuration, the device further includes a duplication splitting unit. The duplication splitting unit obtains first target audio data and second target audio data based on noise-reduced target audio data. Correspondingly, the audio data unit further obtains an audio command by determining whether the first target audio data contains an audio command, and / or transmits the second target audio data to the target device.

[0111] It should be noted that the principle of the voice interaction device of the embodiment of this application is similar to the principle of the voice interaction method described above in relation to the problem that the device solves. Therefore, the implementation process, implementation principle, and beneficial effects of the device can all be described by referring to the explanation of the implementation process, implementation principle, and beneficial effects of the method described above, and any repetition will not be explained again.

[0112] Embodiments of this application provide an augmented reality device, the augmented reality device including at least the voice interaction device described in this application.

[0113] According to embodiments of this application, the application also provides electronic devices.

[0114] Figure 6 is a schematic block diagram showing an exemplary electronic device 600 used to carry out embodiments of the present application. An electronic device refers to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants (PDAs), servers, blade servers, mainframe computers, and other suitable computers. An electronic device may also represent various forms of mobile devices, such as personal digital assistants (PDAs), cell phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are illustrative and not intended to limit the implementation of the present application as described and / or claimed herein.

[0115] As shown in Figure 6, the electronic device 600 includes an arithmetic unit 601, which performs various appropriate operations and processes according to computer programs stored in read-only memory (ROM) 602 or computer programs loaded from storage unit 608 into random access memory (RAM) 603. The RAM 603 can also store various programs and data necessary for the operation of the electronic device 600. The arithmetic unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0116] Multiple components in the electronic device 600 are connected to the I / O interface 605, which includes input units 606 such as keyboards and mice, output units 607 such as various types of displays and speakers, storage units 608 such as disks and optical disks, and communication units 609 such as network cards, modems, and wireless communication transceivers. The communication units 609 enable the electronic device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0117] The arithmetic unit 601 can be a variety of general-purpose and / or dedicated process assemblies with processing and computing capabilities. Examples of the arithmetic unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various arithmetic units for executing machine learning model algorithms, a digital signal processing unit (DSP), and any suitable processor, controller, microcontroller, etc. The arithmetic unit 601 performs various methods and processes described in the preamble, such as a voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as a memory unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed into the electronic device 600 via ROM 602 and / or a communication unit 609. Once the computer program is loaded into RAM 603 and executed by the arithmetic unit 601, one or more steps of the voice interaction method described in the preamble can be performed. Alternatively, in other embodiments, the arithmetic unit 601 may be configured to perform any other method of voice interaction (e.g., via firmware).

[0118] Various embodiments of the systems and technologies described in the Preamble of this Specification can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), composite programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: one or more computer software programs which can run and / or interpret on a programmable system which includes at least one programmable processor, which may be a dedicated or general-purpose programmable processor, which can receive data and commands from a storage system, at least one input device, and at least one output device, and transmit data and commands to the storage system, the at least one input device, and the at least one output device.

[0119] The program code used to carry out the method of this application can be written in any combination of one or more programming languages. These program codes are provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, and when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are performed. The program code can run entirely on the machine, partially on the machine, run on the machine and partially on a remote machine as a standalone software package, or run entirely on a remote machine or server.

[0120] It should be understood as follows: Various forms of flows shown in the preamble can be used, and the order of steps can be changed, steps can be added, or steps can be deleted. For example, each step described in this application can be performed in parallel, sequentially, or in a different order. The text is not limited to achieving the desired results of the technical configuration disclosed in this application.

[0121] Furthermore, the terms “first” and “second” are used solely to describe the purpose and should not be understood as indicating or implying relative importance, or implicitly specifying the number of technical features shown. Therefore, features designated as “first” or “second” explicitly or implicitly include at least one such feature. In the specification of this application, “multiple” means two or more unless otherwise defined.

[0122] The above describes only specific embodiments of this application, and the scope of protection of this application is not limited thereto. Modifications or substitutions that are readily conceivable to a person skilled in the art within the technical scope disclosed herein are included in the scope of protection of this application. Therefore, the scope of protection of this application is determined by the scope of protection of the claims.

Claims

1. A voice interaction method, The augmented reality device's voice acquisition unit obtains the user's voice commands. To determine the type of target device connected to the augmented reality device, If the target device is of the first type, the target response mode to the voice command is determined to be the first response mode. If the target device is of the second type, or if no target device is detected, the target response mode for the voice command is determined to be the second response mode. Responding to the voice command using the target response mode, Includes, Responding to the voice command using the aforementioned target response mode is: If the target response mode is the first response mode, the voice command is sent to the target device to cause it to respond. If the target response mode is the second response mode, the system determines which of the augmented reality device and the target device will respond based on the type of voice command. The voice interaction method, characterized by including the following.

2. Determining which of the augmented reality device and the target device will respond based on the type of voice command is: If the voice command is of a first type, the augmented reality device responds to the voice command. If the voice command is of the second type, the voice command is sent to the target device to cause a response. The method according to claim 1, characterized by including

3. If the voice command is of a second type, sending the voice command to the target device to elicit a response is: The method according to claim 2, characterized in that, if the voice command is a command of a second type, the voice command is converted into a target type command, and the target type command is sent to the target device to cause a response.

4. The voice acquisition unit is used to acquire target voice data, and the target voice data includes voice commands. The above method further, The voice command is obtained by determining whether or not a voice command exists in the target voice data. The method according to claim 1, characterized by comprising and / or transmitting the target audio data to the target device.

5. The voice command is obtained by determining whether or not a voice command exists in the target voice data. and / or transmitting the target audio data to the target device is The process involves applying noise reduction processing to the target audio data and obtaining the target audio data after noise reduction. The voice command is obtained by determining whether or not a voice command exists in the target voice data after noise reduction. and / or transmit the noise-reduced target audio data to the target device. The method according to claim 4, characterized by including

6. The voice command is obtained by determining whether or not a voice command exists in the target voice data after noise reduction. And / or, transmitting the noise-reduced target audio data to the target device is: Based on the noise-reduced target audio data, the first target audio data and the second target audio data are acquired. The voice command is obtained by determining whether or not a voice command exists in the first target voice data. The method according to claim 5, characterized by comprising and / or transmitting the second target audio data to the target device.

7. A voice interaction device, Includes an acquisition unit, a first determination unit, a second determination unit, a third determination unit, and a response unit, The aforementioned acquisition unit acquires the user's voice commands using the voice collection unit of the augmented reality device. The first determination unit determines the type of target device connected to the augmented reality device, The second determination unit determines that the target response mode to the voice command is the first response mode when the target device is of the first type. The third determination unit determines that the target response mode to the voice command is the second response mode if the target device is of the second type or if no target device is detected. The voice interaction device is characterized in that the response unit responds to the voice command using the target response mode, and when the target response mode is a first response mode, it transmits the voice command to the target device to cause a response, and when the target response mode is a second response mode, it determines whether the augmented reality device or the target device will respond based on the type of the voice command.

8. An augmented reality device comprising at least the voice interaction device described in claim 7.

9. It is an electronic device, At least one processor, and Includes memory that is communicably connected to at least one processor, of which, The memory stores commands that can be executed by the at least one processor. The command is executed by the at least one processor, thereby enabling the at least one processor to perform the method according to any one of claims 1 to 6. The electronic device characterized by the above.

Citation Information

Patent Citations

  • Speech dialogue method, apparatus and system

    JP2019204074A

  • Electronic device, server and method of controlling the same

    JP2020028129A

  • Method, system and apparatus for providing a composite graphical assistant interface for controlling connected devices

    JP2021523456A

  • Systems and methods for augmented reality

    JP2022542363A

  • Operational command boundaries

    US20220358914A1