Voice interaction method, device and related equipment
By acquiring voice commands from augmented reality devices and adopting different response modes based on the target device type, the problem of voice interaction for augmented reality devices when not connected or connected to different target devices is solved, thus achieving the flexibility and versatility of the device.
Patent Information
- Application Number
- CN202310096785.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-01-12
AI Technical Summary
Existing augmented reality devices can only communicate with their compatible terminal devices, which limits their flexibility and versatility, and makes it impossible to perform normal voice interaction without connecting to the target device or with different target devices.
Augmented reality devices acquire voice commands through an audio acquisition unit, determine the target device type, and respond to the voice commands using different target response modes. This includes allowing the target device to respond when the target device is of type 1, or allowing either the augmented reality device or the target device to respond when the target device is of type 2 or when no target device is detected.
This enables augmented reality devices to perform normal voice interaction without being connected to or connected to different target devices, expanding their application scenarios and demonstrating the flexibility and versatility of the devices.
Smart Images

Figure CN116030810B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of voice interaction, and in particular to a voice interaction method and device and related equipment. BACKGROUND
[0002] Generally, an augmented reality device is used for displaying or collecting sensor data, and in most cases, the augmented reality device needs to be connected to a terminal device that is adapted to the augmented reality device to be used. SUMMARY
[0003] The present application provides a voice interaction method, device and related equipment to at least solve the above technical problems in the prior art.
[0004] According to a first aspect of the present application, a voice interaction method is provided, the method comprising:
[0005] obtaining a voice instruction of a user based on an audio collection unit of an augmented reality device;
[0006] determining a type of a target device connected to the augmented reality device;
[0007] when the target device is of a first type, determining that a target response mode for the voice instruction is a first response mode;
[0008] when the target device is of a second type or no target device is detected, determining that the target response mode for the voice instruction is a second response mode;
[0009] responding to the voice instruction in the target response mode.
[0010] In the above solution, the responding to the voice instruction in the target response mode comprises:
[0011] when the target response mode is the first response mode, handing over the voice instruction to the target device for response;
[0012] when the target response mode is the second response mode, determining that one of the augmented reality device and the target device responds based on a type of the voice instruction.
[0013] In the above solution, the determining that one of the augmented reality device and the target device responds based on the type of the voice instruction comprises:
[0014] when the voice instruction is of a first type, responding to the voice instruction by the augmented reality device;
[0015] when the voice instruction is of a second type, handing over the voice instruction to the target device for response.
[0016] In the above solution, when the voice instruction is a second type of instruction, the voice instruction is handed over to the target device for response.
[0017] When the voice instruction is a second type of instruction, the voice instruction is converted into a target type of instruction, and the target type of instruction is handed over to the target device for response.
[0018] In the above solution, the audio acquisition unit is configured to acquire target voice data, and the target voice data includes a voice instruction.
[0019] The method further includes:
[0020] The voice instruction is obtained by determining whether the voice instruction exists in the target voice data.
[0021] And / or, the target voice data is handed over to the target device.
[0022] In the above solution, the voice instruction is obtained by determining whether the voice instruction exists in the target voice data; and / or, the target voice data is handed over to the target device, including:
[0023] The target voice data is subjected to noise reduction processing to obtain noise-reduced target voice data.
[0024] The voice instruction is obtained by determining whether the voice instruction exists in the noise-reduced target voice data.
[0025] And / or, the noise-reduced target voice data is handed over to the target device.
[0026] In the above solution, the voice instruction is obtained by determining whether the voice instruction exists in the noise-reduced target voice data; and / or, the noise-reduced target voice data is handed over to the target device, including:
[0027] Based on the noise-reduced target voice data, first target voice data and second target voice data are obtained.
[0028] The voice instruction is obtained by determining whether the voice instruction exists in the first target voice data.
[0029] And / or, the second target voice data is handed over to the target device.
[0030] According to a second aspect of the present application, a voice interaction device is provided, and the device includes:
[0031] An acquisition unit is configured to acquire a voice instruction of a user based on an audio acquisition unit of an augmented reality device.
[0032] a first determining unit, configured to determine a type of a target device connected with the augmented reality device;
[0033] a second determining unit, configured to determine a target response mode of the voice instruction as a first response mode when the target device is of a first type;
[0034] a third determining unit, configured to determine the target response mode of the voice instruction as a second response mode when the target device is of a second type or no target device is detected;
[0035] a responding unit, configured to respond to the voice instruction in the target response mode.
[0036] According to a third aspect of the present application, an augmented reality device is provided, which at least comprises the voice interaction apparatus.
[0037] According to a fourth aspect of the present application, an electronic device is provided, which comprises:
[0038] at least one processor; and
[0039] a memory connected with the at least one processor in communication; wherein,
[0040] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method.
[0041] In the present application, the voice instruction of a user is acquired based on an audio acquisition unit of an augmented reality device, the type of a target device connected with the augmented reality device is determined, the target response mode of the voice instruction is determined as a first response mode when the target device is of a first type, the target response mode of the voice instruction is determined as a second response mode when the target device is of a second type or no target device is detected, and the voice instruction is responded to in the target response mode. Technical support is provided for the augmented reality device to normally perform voice interaction in the scenario of not connecting the target device or connecting different target devices.
[0042] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0043] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description read in conjunction with the accompanying drawings, in which:
[0044] In the drawings, identical or corresponding components are denoted by identical or corresponding reference numerals.
[0045] Figure 1 An implementation flowchart of a voice interaction method according to an embodiment of the present application is shown;
[0046] Figure 2 An implementation flowchart of different target response modes according to an embodiment of the present application is shown Figure 1 ;
[0047] Figure 3 An implementation flowchart of different target response modes according to an embodiment of the present application is shown Figure 2 ;
[0048] Figure 4 An implementation flowchart of an augmented reality device according to an embodiment of the present application is shown;
[0049] Figure 5 An implementation flowchart of a voice interaction device according to an embodiment of the present application is shown;
[0050] Figure 6 An implementation flowchart of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0051] In order to make the objectives, features and advantages of the present application more apparent and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0052] In the related art, an augmented reality device can only communicate with a terminal device adapted thereto, and cannot reflect the flexibility and multifunctionality of the augmented reality device.
[0053] It can be understood that the augmented reality (AR) device is a mainstream wearable device at present, and the smart voice interaction as a mainstream interaction mode of the augmented reality device can achieve hands-free and easy and fast input or control of the augmented reality device. Considering the influence of appearance, wearing, power consumption, heating and other factors, the augmented reality device is usually not used as a complex computing unit, but only as a display (projection) and sensor data acquisition function (such as image, audio, inertial measurement unit, etc.). Based on this, the augmented reality device can only be connected to a terminal device adapted thereto through a wired or wireless manner, and the collected sensor data is given to the terminal device for algorithm calculation. If the augmented reality device can also normally perform voice interaction when not connected to the terminal device or connected to other terminal devices, the functionality of the augmented reality device will be expanded. In this way, the widespread application of the augmented reality device in daily life can be laid a foundation.
[0054] The technical scheme of the embodiment of the present application relates to a voice interaction scheme. The augmented reality device can be connected with different types of target devices for voice interaction, and can also perform voice interaction without being connected to the target device, which embodies the multifunctionality and flexibility of the augmented reality device. Based on the obtained voice instruction and the type of the target device connected with the augmented reality device, different target response modes can be adopted to respond to the voice instruction. The technical support is provided for the augmented reality device to normally perform voice interaction in the scene of not being connected to the target device or being connected to different target devices, and the use scene of the augmented reality device is expanded.
[0055] The voice interaction method of the embodiment of the present application will be described in detail below.
[0056] The embodiment of the present application provides a voice interaction method, as shown in Figure 1 The method comprises the following steps.
[0057] S101: Obtain a voice instruction of a user based on an audio acquisition unit of an augmented reality device.
[0058] In this step, the augmented reality device is an electronic device that can perform AR interaction. The augmented reality device can be a smart wearable device, including but not limited to smart glasses, a smart watch. In the present application, the augmented reality device is taken as a split-type AR glasses as an example for description.
[0059] The augmented reality device comprises an audio acquisition unit, such as a microphone. In this step, the voice instruction of the user is obtained by the microphone in the manner of collecting the voice instruction of the user to the augmented reality device.
[0060] It can be understood that the microphone comprises a microphone array sensor (MicArray) for collecting the voice instruction of the user to the augmented reality device.
[0061] In practical applications, the user can issue a voice instruction to the augmented reality device when the augmented reality device is not in use. For example, the user can issue a voice instruction such as “please turn on the screen” or “please turn off the device” to the augmented reality device when the augmented reality device is in a black screen state.
[0062] The user can also issue a voice instruction to the augmented reality device when the augmented reality device is in use, such as when the augmented reality device is playing a movie or playing audio such as a song. That is, when the augmented reality device is outputting multimedia data, the voice instruction of the user is acquired based on the audio acquisition unit of the augmented reality device.
[0063] It can be understood that the augmented reality device usually outputs some multimedia data such as images and audio during AR interaction. When the augmented reality device is connected to a target device, the augmented reality device can project the video in the target device to the augmented reality device for output, such as projecting a movie in the target device to the augmented reality device for output. In this case, the multimedia data can refer to the video in the target device that can be projected by the AR device. In addition, the audio in the target device can be output through the augmented reality device, such as answering a call or a voice call through the augmented reality device. In this case, the multimedia data can refer to the audio in the target device that can be output by the AR device.
[0064] That is, when connected to a target device, the augmented reality device in the present application is a device that replaces the target device to output audio and video. This replacement scheme mainly considers that, in some application scenarios, it is much more convenient and has better output effect to use the augmented reality device as an audio and video output device than to use the target device as an audio and video output device. For example, in a projection application scenario, the wearer of the augmented reality device can experience an immersive effect by projecting the projectable video in the target device through the augmented reality device. For another example, in some crowded environments, such as rush hours in the morning and evening, it is not convenient to answer a call from the target device due to the crowdedness of the subway. Therefore, the wearer can answer the call through the split AR glasses worn by the wearer. The split AR glasses can be built in ordinary glasses worn by the wearer, and the wearer can answer the call through the split AR glasses, thereby avoiding the situation that the wearer cannot take the phone out of the pocket to answer the call in a crowded environment.
[0065] The user can issue a voice instruction to the augmented reality device when necessary. The augmented reality device, specifically the microphone, acquires the voice instruction.
[0066] Exemplarily, in the case that the augmented reality device is a split AR glasses, it is assumed that the current scene is that the user uses the split AR glasses to play audio and video, at this time, the multimedia data output by the split AR glasses is the audio and video information played by the user, when the user issues a voice instruction "turn up the volume", the voice instruction for audio and video playing can be acquired through the microphone collecting the voice instruction issued by the user, so as to increase the volume of the current audio and video playing through the response to the voice instruction.
[0067] S102: Determine the type of the target device connected with the augmented reality device.
[0068] In this step, the target device can be any device for voice interaction with the augmented reality device, such as a smart phone, a computer, a personal digital assistant, etc. The target device connected with the augmented reality device in the present application can be a terminal of different types. For example, the target device is a self-developed terminal, i.e. a terminal compatible with the augmented reality device. It can be understood that after the augmented reality device is produced by a manufacturer, there is usually a self-developed terminal produced by the manufacturer and compatible with the augmented reality device. The self-developed terminal can be understood as a normally sized terminal without a display screen or with a small display screen but with computing capability. The smart voice interaction function of the augmented reality device is realized by connecting the augmented reality device with the compatible self-developed terminal.
[0069] The target device can also be a terminal of other types, such as a third-party mobile phone or a third-party computer. In the present application, the smart voice interaction function of the augmented reality device is realized by connecting the augmented reality device with the terminal of other types. Exemplarily, when the augmented reality device is a split AR glasses, the split AR glasses can be connected with a terminal compatible therewith and also can be connected with a third-party terminal. In the present application, whether connected with a terminal compatible therewith or connected with a third-party terminal, the split AR glasses can perform smart voice interaction with the device connected with the split AR glasses.
[0070] In actual application, voice interaction is usually limited on the self-developed terminal as the mainstream interaction mode of the augmented reality device. That is, the user wants to use the augmented reality device, and has to purchase a self-developed terminal compatible therewith to normally perform voice interaction. In this case, on the one hand, the user will not additionally purchase a compatible terminal for cost and portability considerations in the case that the user already has a mobile phone. On the other hand, the augmented reality device is connected with a PC computer, and the PC as a computing unit will not be connected with a terminal. In the foregoing two scenarios, the voice interaction function of the augmented reality device will become unusable, which greatly limits the use scenarios of voice interaction under the AR glasses.
[0071] In the present application, the augmented reality device can access different types of terminals, and by determining the type of the target device connected to the augmented reality device, the voice interaction function of the augmented reality device is realized. The augmented reality device can also not access any type of terminal, and some basic voice interaction functions such as adjusting the volume and adjusting the brightness can be realized by the voice instruction control application itself.
[0072] In actual application, the augmented reality device can access different types of terminals, and can also not access any terminal, so that the user can choose to only purchase the augmented reality device, without having to purchase a self-developed terminal that is compatible with it, and directly connecting to the user's own mobile phone or computer can also realize the voice interaction function of the augmented reality device.
[0073] S103: When the target device is of the first type, determining that the target response mode to the voice instruction is a first response mode.
[0074] In the present application, the augmented reality device can access or connect to different types of terminals, and the target device type connected by the augmented reality device will be different, and the response mode to the voice instruction will also be different.
[0075] In the present application, the type of the target device includes a first type and a second type. The first type is a terminal that internally contains a voice keyword detection technology service, such as the self-developed terminal described above. The second type is a terminal that does not internally contain a voice keyword detection technology service, such as the third-party mobile terminal, computer terminal, and the like described above.
[0076] For example, assuming that the target device connected by the augmented reality device is the self-developed terminal described above, the corresponding target response mode to the voice instruction can be mode A (first response mode). Assuming that the target device connected by the augmented reality device is the other type of terminal (third-party mobile terminal, computer terminal, etc.) described above, the corresponding target response mode to the voice instruction can be mode B (second response mode).
[0077] In implementation, the augmented reality device identifies whether it is connected to a target device. If there is no connected device, it is determined that the target response mode to the voice instruction is a second response mode. If there is a connected device, the identity of the connected device is obtained, and based on the identity of the connected device, it is determined whether the connected device is of the first type or of the second type. If the identity of the connected device is identity A, and identity A is an identity representing a device of the first type, it is determined that the connected device is of the first type. If the identity of the connected device is identity B, and identity B is an identity representing a device of the second type, it is determined that the connected device is of the second type.
[0078] In the present application, based on whether the terminal connected with the augmented reality device is a terminal adapted to the augmented reality device or a third-party terminal not adapted to the augmented reality device, two response modes are pre-set. One of the modes is used when the terminal connected with the augmented reality device is a terminal adapted to the augmented reality device. The other mode is used when the terminal connected with the augmented reality device is a third-party terminal. The two different types of terminals and the mode to be used under each type of terminal are pre-set as a corresponding relationship. In implementation, based on the type of the terminal connected with the augmented reality device, the mode corresponding to the type of terminal in the corresponding relationship is found as the target response mode for responding to the voice instruction.
[0079] S104: When the target device is of the second type or no target device is detected, determining that the target response mode for the voice instruction is a second response mode.
[0080] Since the target device of the second type is a terminal not containing the voice keyword detection technology service, in order to ensure that the augmented reality device can still perform voice interaction when accessing the target device of the second type, the voice keyword detection technology service needs to be used in the augmented reality device to realize normal voice interaction of the augmented reality device when connecting the target device of the second type. At the same time, since the voice keyword detection technology service is included in the augmented reality device, when the augmented reality device is not connected with the target device, i.e., no target device is detected, the augmented reality device can also perform basic voice interaction, such as adjusting brightness and adjusting volume, by using the corresponding target response mode.
[0081] In the present application, a low-power voice keyword detection technology service is deployed in the augmented reality device. Without significantly increasing the power consumption or heat of the augmented reality device, the low-power voice keyword detection technology service is used to complete the recognition of the voice instruction.
[0082] By implementing different target response modes for different types of terminals, the ability of the augmented reality device to maintain voice interaction when accessing different types of terminals is satisfied.
[0083] S105: Responding to the voice instruction by using the target response mode.
[0084] By using different target response modes to respond to the voice instruction according to the type of the connected target device, i.e., by using different target response modes to analyze the voice instruction and process the instruction, it is exemplarily analyzed that the voice instruction is a voice instruction of "increasing volume" and the output volume of the augmented reality device is increased. It is analyzed that the voice instruction is an instruction of "decreasing screen brightness" and the screen brightness of the augmented reality device is decreased. The corresponding operation of the voice instruction is implemented.
[0085] In the scheme shown in S101-S105, the augmented reality device can not only communicate with the adapted terminal and the third-party terminal, but also perform voice interaction without accessing any terminal. The flexibility and multifunctionality of the augmented reality device are embodied.
[0086] In this application, the target response mode can be determined based on the type of the target device connected to the augmented reality device, and the response to the voice instruction is realized by using the target response mode. Based on the type of the target device connected to the augmented reality device, different target response modes can be used to respond to the voice instruction. The augmented reality device maintains the voice interaction capability in the case of accessing different types of terminals. Technical support is provided to realize the voice interaction of the augmented reality device in the scene of connecting different target devices.
[0087] In an optional scheme, the response to the voice instruction by using the target response mode comprises:
[0088] When the target response mode is the first response mode, the voice instruction is responded by the target device;
[0089] When the target response mode is the second response mode, it is determined that one of the augmented reality device and the target device responds based on the type of the voice instruction.
[0090] The type of the voice instruction represents the type of the instruction that is responded by the augmented reality device or the type of the instruction that is responded by the target device.
[0091] As Figure 2As shown, in this application, the responding subjects involved mainly have two types: target device and augmented reality device. The augmented reality device mainly includes voice keyword detection technology service and voice instruction control application. When the target response mode is the first response mode, i.e., the target device type connected based on the augmented reality device is the first type such as a self-developed terminal, and the target response mode is determined to be the first response mode, the voice instruction is handed over to the operating system of the target device for full voice control. When the target response mode is the second response mode, i.e., the target device type connected based on the augmented reality device is the second type such as a third-party mobile phone or computer, or no target device is detected, and the target response mode is determined to be the second response mode, the voice instruction is converted into an Event Id and notified to the voice instruction control application in the augmented reality device, and the voice instruction control application judges whether the instruction is an augmented reality device self-control instruction. If the instruction is an augmented reality device self-control instruction, the voice instruction control application responds to the voice instruction, such as volume adjustment, brightness adjustment, etc. If the voice instruction is a target device control instruction, the Event Id is converted into a Universal Serial Bus KeyBoard (USB KeyBoard) protocol Id, and the USB KeyBoard Id is handed over to the target device, and the target device responds to the voice instruction.
[0092] Specifically, as shown, Figure 3 The voice keyword detection technology service of the augmented reality device is in a running state, and the voice keyword detection technology service judges whether the target device connected with the augmented reality device is of the first type or of the second type. By judging the type of the target device connected, different target response modes are adopted. When the target response mode is the first response mode, i.e., the target device type connected based on the augmented reality device is the first type such as a self-developed terminal, and the target response mode is determined to be the first response mode, the voice keyword detection technology service in the augmented reality device enters a dormant state, and the voice instruction is handed over to the operating system of the target device for full voice control. Exemplarily, assuming that the target device connected with the augmented reality device is a self-developed terminal, since the self-developed terminal contains the voice keyword detection technology service, there are usually predefined voice instruction sets on the self-developed terminal, such as a volume-up instruction for adjusting the sound of multimedia data, an increase-brightness instruction for adjusting the display brightness of the screen, a mode-switching instruction for adjusting the display mode of multimedia data, such as adjusting from normal mode to 3D mode, or adjusting from 3D mode to normal mode, etc.
[0093] After the user issues a corresponding voice instruction, the augmented reality device determines that the accessed target device is a self-developed terminal, and can hand over the voice instruction collected by the augmented reality device to the operating system of the target device for full voice control. The voice keyword detection technology service running on the self-developed terminal hands over the voice instruction to the self-developed terminal. The self-developed terminal, specifically the application processing unit, responds to the voice instruction and performs a control action corresponding to the language instruction, such as sound adjustment, brightness adjustment, mode switching, etc.
[0094] If the voice instruction obtained by the augmented reality device is a (sound adjustment) instruction for adjusting the sound of the multimedia data output by the augmented reality device, and the target device connected with the augmented reality device is a self-developed terminal adapted to the augmented reality device, the sound adjustment instruction is handed over to the self-developed terminal for processing, so as to realize the adjustment of the sound of the multimedia data output by the augmented reality device through the self-developed terminal.
[0095] If the voice instruction obtained by the augmented reality device is an instruction for adjusting the display brightness of the multimedia data output by the augmented reality device, and the target device connected with the augmented reality device is a self-developed terminal adapted to the augmented reality device, the display brightness adjustment instruction is handed over to the self-developed terminal for processing, so as to realize the adjustment of the display brightness of the screen of the augmented reality device through the self-developed terminal.
[0096] If the voice instruction obtained by the augmented reality device is an instruction for adjusting the display mode of the multimedia data output by the augmented reality device, and the target device connected with the augmented reality device is a self-developed terminal adapted to the augmented reality device, the display mode adjustment instruction is handed over to the self-developed terminal for processing, so as to realize the adjustment of the display mode of the multimedia data output by the augmented reality device through the self-developed terminal.
[0097] When the target response mode is the second response mode, i.e., the target response mode is determined to be the second response mode based on the type of the target device connected with the augmented reality device being the second type such as a third-party mobile phone or computer, or no target device being detected, the voice keyword detection technology service of the augmented reality device remains running.
[0098] In an optional scheme, when the target response mode is the second response mode, based on the type of the voice instruction, it is determined that one of the augmented reality device and the target device responds, which includes:
[0099] When the voice instruction is the first type of instruction, the augmented reality device responds to the voice instruction;
[0100] When the voice instruction is the second type of instruction, the voice instruction is handed over to the target device for response.
[0101] In the present application, in the case that the target response mode is the second response mode, the voice instruction also includes two types. The first type instruction is an instruction that the augmented reality device can respond to, such as the aforementioned sound adjustment instruction, display brightness adjustment instruction, and display mode adjustment instruction. The second type instruction is an instruction that the augmented reality device cannot respond to, but the target device can respond to, such as the "return to the previous step" instruction, "determine" instruction, and "main menu" instruction.
[0102] In the present application, in the case that the target response mode is the second response mode, if the voice instruction is an instruction that the augmented reality device can respond to, such as the sound adjustment instruction, display brightness adjustment instruction, and display mode adjustment instruction, the augmented reality device responds to the voice instruction. If the voice instruction is an instruction such as the "return to the previous step" instruction, "determine" instruction, and "main menu" instruction, the voice instruction is handed over to the target device for response.
[0103] Based on this, in the second response mode in the present application, different response subjects respond to voice instructions according to the type of the voice instruction, which can avoid the problem of excessive burden caused by the fact that all voice instructions are responded to by one of the response subjects. In the present application, different types of voice instructions are responded to by different response subjects, which is easy to implement in engineering and has high feasibility.
[0104] In an optional scheme, when the voice instruction is a second type instruction, handing over the voice instruction to the target device for response includes:
[0105] When the voice instruction is a second type instruction, converting the voice instruction into a target type instruction, and handing over the target type instruction to the target device for response.
[0106] In the present application, if the voice instruction is an instruction that the augmented reality device cannot respond to, but the target device can respond to, the voice instruction needs to be converted into a type that the target device can recognize or respond to, and then the response to the voice instruction is realized.
[0107] Illustratively, assuming that the target device connected to the augmented reality device is a third-party mobile phone or computer terminal, when the voice keyword detection technology service in the augmented reality device identifies a voice instruction, the voice instruction is converted into an Event Id and notified to the voice instruction control application in the augmented reality device. After receiving the instruction Event Id, the voice instruction control application corresponds the instruction Event Id to a number corresponding to a predefined instruction function, and makes a corresponding feedback according to the predefined instruction function represented by the corresponding number. The predefined instruction function is an instruction function represented by a different number in a predefined instruction set, such as number 1 representing a sound increase function, number 2 representing a brightness increase function, number 3 representing a mode switching function, and the like.
[0108] When the instruction corresponding to the instruction Event Id is the first type of instruction, that is, when the instruction Event Id corresponds to the number corresponding to the instruction function predefined for the augmented reality device, that is, the instruction corresponding to the instruction Event Id is the self-control related instruction that the augmented reality device can respond to, such as sound adjustment, brightness adjustment, display switching, mode switching, and other instructions related to the hardware of the augmented reality device, the voice instruction control application directly completes the corresponding control operation through the system application programming interface (API, Application Programming Interface). When the instruction corresponding to the instruction Event Id is the second type of instruction, that is, when the instruction Event Id corresponds to the number corresponding to the instruction function predefined for the target device, that is, when the instruction corresponding to the instruction Event Id is an instruction that needs to be responded by the target device, the instruction Event Id is converted into a USB KeyBoard protocol Id and sent to the accessed target device. The target device responds to the USB KeyBoard protocol Id and performs an operation on the multimedia data output by the augmented reality device. Illustratively, the instruction "return to the previous step" is defined as the keyboard "F1" key of the target device, and when the user clicks the "F1" key, the multimedia data output by the augmented reality device will return to the previous step operation. Or "main menu" is defined as the keyboard "F2" key of the target device, and when the user clicks the "F2" key, the augmented reality device will pop up a main menu interface. Or it can also be defined as the keyboard "enter" key of the target device, and when the user clicks the "enter" key, the confirmation operation can be performed.
[0109] The aforementioned execution process of the augmented reality device in the application can be implemented by a high-performance special-purpose chip arranged in the augmented reality device, which can improve the computing efficiency and realize fast response to the voice instruction. In addition, in the application, the low-power voice keyword detection technology service is adopted, which has the advantages of low required computing power and low resource consumption. In the case of not significantly increasing the power consumption, the augmented reality device is endowed with AR capability, which greatly improves the self-control capability of the augmented reality device. In addition, the voice instruction can be customized and extended, which makes up for the shortcomings of control by keys. With the help of the general USB KeyBoard protocol, the augmented reality device converts the recognized voice instruction into a target type that can be recognized or responded by the target device, and regards the response of the target device to the voice instruction as the response of the voice instruction executed by the key (such as the aforementioned keyboard "F1" key of the target device, the keyboard "F2" key of the target device, and the keyboard "enter" key of the target device) of the user operation. This scheme is equivalent to regarding the augmented reality device as the control peripheral of the target device, and endows the target device with voice interaction capability in the case of low-cost access. In actual application, the product competitiveness of the augmented reality device is increased.
[0110] In an optional scheme, the audio acquisition unit is configured to acquire target voice data, and the target voice data includes the voice instruction.
[0111] The method further includes:
[0112] The voice instruction is acquired by determining whether the voice instruction exists in the target voice data.
[0113] The target voice data is delivered to the target device.
[0114] In the scheme, the augmented reality device further includes a keyword detection unit and a control application unit. The keyword detection unit includes a voice keyword detection technology service, which is configured to judge the type of the accessed target device and recognize the voice instruction. The control application unit includes a voice instruction control application, which is configured to judge the instruction type and deliver different types of instructions to different response subjects for response. The target voice data includes the voice instruction and / or voice data generated when the augmented reality device outputs multimedia data.
[0115] It can be understood that the voice instruction in the present application refers to an instruction input to the augmented reality device in the form of voice, which is a kind of instruction data. The target voice data collected by the microphone array sensor can be instruction data. In addition, the voice data collected by the microphone array sensor can be voice data generated when multimedia data is output, such as the voice data collected by the microphone array sensor when a call is answered through the augmented reality device, which can be the content of the call. Or when the projection of a movie is performed through the augmented reality device, the voice data collected by the microphone array sensor can be the audio content in the movie.
[0116] The augmented reality device collects target voice data through the microphone array sensor, identifies whether the target voice data contains voice instructions and / or whether the target voice data contains voice data generated when the augmented reality device outputs multimedia data. If it is identified that the target voice data contains both voice instructions and voice data generated when the augmented reality device outputs multimedia data, the two kinds of voice data can be separated to perform different processing on the two kinds of voice data.
[0117] When the target voice data contains voice instructions, the augmented reality device determines different target device types through the keyword detection unit, so as to adopt different response modes. Specifically, when the target device is of a first type, the target device responds to the voice instructions. When the target device is of a second type or no target device is detected, the keyword detection unit identifies the voice instructions, the control application unit determines different instruction types, and different response subjects are selected to respond to the voice instructions.
[0118] When the target voice data contains voice data generated when the augmented reality device outputs multimedia data, the voice data can be handed over to the target device, and the target device performs recording, semantic recognition and other processing on the voice data.
[0119] In the foregoing scheme, the collected target voice data is transmitted to the augmented reality device and / or the target device, which is easy to implement in engineering and ensures the normal use of the target voice data by the augmented reality device and / or the target device.
[0120] In an optional scheme, the voice instruction is obtained by determining whether the voice instruction exists in the target voice data; and / or, the target voice data is handed over to the target device, which includes:
[0121] The target voice data is subjected to noise reduction processing to obtain noise-reduced target voice data;
[0122] The voice instruction is obtained by determining whether the voice instruction exists in the noise-reduced target voice data;
[0123] and / or, the target device.
[0124] In this scheme, the augmented reality device further comprises a noise reduction unit for performing noise reduction processing on the target voice data to obtain noise-reduced target voice data. The augmented reality device collects target voice data through a microphone array sensor, and performs noise reduction processing on the collected target voice data through the noise reduction unit to obtain noise-reduced target voice data. The noise-reduced target voice data is transmitted to the augmented reality device end and / or the target device end. When the target voice data contains voice instructions, the augmented reality device end determines different target device types through the keyword detection unit, and thus adopts different response modes. When the target device type is the first type, the voice instructions are responded to by the target device. When the target device type is the second type or no target device is detected, the voice instructions are recognized through the keyword detection unit, and the control application unit determines different instruction types, and thus selects different response subjects to respond to the voice instructions.
[0125] When the target voice data contains voice data generated by the augmented reality device when outputting multimedia data, the noise-reduced voice data can be transmitted to the target device, and the target device can perform recording, semantic recognition, etc. on the noise-reduced voice data.
[0126] By performing noise reduction on the collected target voice data, the noise-reduced target voice data does not contain noise or contains less noise, which can improve the quality of the voice data transmitted to the augmented reality device end and / or the target device end, and thus accurate response to the target voice data can be realized.
[0127] In an optional scheme, the voice instructions are obtained by determining whether the noise-reduced target voice data contains voice instructions; and / or, the noise-reduced target voice data is transmitted to the target device, comprising:
[0128] Based on the noise-reduced target voice data, first target voice data and second target voice data are obtained;
[0129] The voice instructions are obtained by determining whether the first target voice data contains voice instructions;
[0130] and / or, the second target voice data is transmitted to the target device.
[0131] In this scheme, the augmented reality device further comprises a copy and shunt unit for copying the noise-reduced target voice data to obtain two copies of noise-reduced target voice data, and shunting the two copies of noise-reduced target voice data to obtain first target voice data and second target voice data. For example, Figure 4As shown, the augmented reality device sends the read microphone audio data Audio In to a specific noise reduction unit for directional noise reduction, obtaining enhanced audio data Clean Audio without noise or with less noise, and the audio data without noise or with less noise is obtained by the copy shunt unit to obtain the first target voice data and the second target voice data. It can be understood that the first target voice data and the second target voice data are the same voice data as the Clean Audio after noise reduction. The second target voice data is directly transmitted to the target device Audio Out by hardware wired or wireless mode, and the first target voice data is processed by the augmented reality device end. Further, when there is a voice instruction in the target voice data, the augmented reality device end acquires the voice instruction, judges different target device types through the keyword detection unit, and adopts different response modes. Specifically, when the target device is of the first type, the voice instruction is responded by the target device. When the target device is of the second type or no target device is detected, the voice instruction is recognized by the keyword detection unit, the control application unit judges different instruction types, and different response subjects are selected to respond to the voice instruction.
[0132] By copying and shunting the target voice data after noise reduction, the first target voice data and the second target voice data are obtained, the first target voice data is processed by the augmented reality device end, which is a scheme for voice instruction recognition and response by different response subjects. The second target voice data is transmitted to the target device end so that the collected target voice data can be used by other applications of the target device. Exemplarily, when the user makes a call, the collected target voice data is used as call content by the call application of the target device. Copying and shunting the target voice data after noise reduction can ensure that the applications of the augmented reality device end and the target device end can obtain the audio data they need without interfering with each other.
[0133] The embodiment of the present application provides a voice interaction device, such as Figure 5 As shown, the device comprises:
[0134] The acquisition unit 501 is configured to acquire a voice instruction of a user based on an audio acquisition unit of an augmented reality device.
[0135] The first determination unit 502 is configured to determine a type of a target device connected with the augmented reality device.
[0136] The second determination unit 503 is configured to determine that a target response mode of the voice instruction is a first response mode when the target device is of a first type.
[0137] The third determining unit 504 is configured to determine that the target response mode of the voice instruction is a second response mode when the target device is of a second type or no target device is detected.
[0138] The response unit 505 is configured to respond to the voice instruction in the target response mode.
[0139] In an optional implementation, the response unit 505 is configured to, when the target response mode is the first response mode, hand over the voice instruction to the target device for response; and when the target response mode is the second response mode, determine that one of the target device and the augmented reality device responds to the voice instruction based on a type of the voice instruction.
[0140] In an optional implementation, the response unit 505 is configured to, when the voice instruction is of a first type, respond to the voice instruction by the augmented reality device; and when the voice instruction is of a second type, hand over the voice instruction to the target device for response.
[0141] In an optional implementation, the response unit 505 is configured to, when the voice instruction is of the second type, convert the voice instruction into a target type instruction, and hand over the target type instruction to the target device for response.
[0142] In an optional implementation, the audio acquisition unit is configured to acquire target voice data, where the target voice data includes the voice instruction.
[0143] The apparatus further includes:
[0144] The voice data unit is configured to acquire the voice instruction by determining whether the voice instruction exists in the target voice data; and / or hand over the target voice data to the target device.
[0145] In an optional implementation, the voice data unit is configured to perform noise reduction processing on the target voice data to obtain noise-reduced target voice data; acquire the voice instruction by determining whether the voice instruction exists in the noise-reduced target voice data; and / or hand over the noise-reduced target voice data to the target device.
[0146] In an optional implementation, the voice data unit is configured to obtain first target voice data and second target voice data based on the noise-reduced target voice data; acquire the voice instruction by determining whether the voice instruction exists in the first target voice data; and / or hand over the second target voice data to the target device.
[0147] It should be noted that the voice interaction device of the embodiments of the present application has similar principles to the voice interaction method described above in solving problems, and therefore the implementation process, implementation principles and beneficial effects of the device can be seen from the description of the implementation process, implementation principles and beneficial effects of the method described above, and the repeated parts will not be described again.
[0148] The embodiments of the present application provide an augmented reality device, which at least comprises the voice interaction device described in the present application.
[0149] According to the embodiments of the present application, the present application further provides an electronic device.
[0150] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0151] As shown in Figure 6 The electronic device 600 includes a computing unit 601 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0152] Various components in the electronic device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc., an output unit 607, such as various types of displays, a speaker, etc., a storage unit 608, such as a magnetic disk, an optical disk, etc., and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0153] The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the voice interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the voice interaction method by any other suitable means, such as by means of firmware.
[0154] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0155] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, fully on a machine and partially on a remote machine or entirely on a remote machine or server.
[0156] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present application can be achieved, which is not limited herein.
[0157] In addition, the terms "first", "second", "third", etc. are used only for descriptive purposes and should not be construed as indicating or implying relative importance or an indicated number of technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0158] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A voice interaction method, characterized in that, The method comprises: obtaining a voice instruction of a user based on an audio acquisition unit of an augmented reality device; determining a type of a target device connected with the augmented reality device; when the target device is of a first type, determining that a target response mode for the voice instruction is a first response mode; when the target device is of a second type or no target device is detected, determining that the target response mode for the voice instruction is a second response mode; responding to the voice instruction in the target response mode; wherein the responding to the voice instruction in the target response mode comprises: when the target response mode is the first response mode, handing over the voice instruction to the target device for response; when the target response mode is the second response mode, determining that one of the augmented reality device and the target device responds based on a type of the voice instruction.
2. The method of claim 1, wherein, The determining that one of the augmented reality device and the target device responds based on the type of the voice instruction comprises: when the voice instruction is of a first type, responding to the voice instruction by the augmented reality device; when the voice instruction is of a second type, handing over the voice instruction to the target device for response.
3. The method of claim 2, wherein, The handing over the voice instruction to the target device for response when the voice instruction is of the second type comprises: when the voice instruction is of the second type, converting the voice instruction into a target type instruction, and handing over the target type instruction to the target device for response.
4. The method of claim 1, wherein, The audio acquisition unit is configured to acquire target voice data, and the target voice data comprises a voice instruction. The method further comprises: obtaining the voice instruction by determining whether the voice instruction exists in the target voice data; and / or, handing over the target voice data to the target device.
5. The method of claim 4, wherein, The obtaining the voice instruction by determining whether the voice instruction exists in the target voice data; and / or, handing over the target voice data to the target device comprises: performing noise reduction processing on the target voice data to obtain noise-reduced target voice data; obtaining the voice instruction by determining whether the voice instruction exists in the noise-reduced target voice data; and / or, handing over the noise-reduced target voice data to the target device.
6. The method of claim 5, wherein, The obtaining the voice instruction by determining whether the voice instruction exists in the noise-reduced target voice data; and / or, handing over the noise-reduced target voice data to the target device comprises: obtaining first target voice data and second target voice data based on the noise-reduced target voice data; obtaining the voice instruction by determining whether the voice instruction exists in the first target voice data; and / or, handing over the second target voice data to the target device.
7. A voice interaction device, characterized by The apparatus comprises: an obtaining unit configured to obtain a voice instruction of a user based on an audio acquisition unit of an augmented reality device; a first determining unit configured to determine a type of a target device connected with the augmented reality device; The second determining unit is configured to determine that the target response mode of the voice instruction is a first response mode when the target device is of the first type. The third determining unit is configured to determine that the target response mode of the voice instruction is a second response mode when the target device is of the second type or no target device is detected. The response unit is configured to respond to the voice instruction in the target response mode. The response unit is configured to, when the target response mode is the first response mode, hand over the voice instruction to the target device for response; and when the target response mode is the second response mode, determine that one of the target device and the augmented reality device responds to the voice instruction based on a type of the voice instruction.
8. An augmented reality device, characterized by The voice interaction apparatus at least comprises the voice interaction apparatus as claimed in claim 7.
9. An electronic device, comprising: The voice interaction apparatus comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as claimed in any one of claims 1-6.
Citation Information
Patent Citations
Equipment control method, equipment control device, equipment and medium
CN109658932A