Virtual reality voice interaction method
By obtaining user gaze information and interaction information to determine whether to start speech recognition, the problem of low start accuracy of voice assistant caused by inaccurate speech recognition in the prior art is solved, and a higher wake-up accuracy of voice recognition is achieved.
Patent Information
- Application Number
- CN202311864009.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the voice wake-up voice assistant is prone to inaccurate speech recognition, resulting in a low accuracy of voice assistant activation.
By acquiring the user gaze information and the first interactive information, it is determined whether the preset condition for turning on voice recognition is met based on the user gaze information and the first interactive information, and the voice recognition is activated when the condition is met. The first interactive information is image information or motion information of the user's body movement.
It improves the accuracy of voice recognition wake-up, avoids dependence on the accuracy of live recording voice recognition, and enhances the startup accuracy of voice assistants.
Smart Images

Figure CN120236575A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of virtual display, and in particular to a virtual reality voice interaction method. Background Art
[0002] With the development of computer technology and virtual display technologies such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and extended reality (XR), virtual display devices are increasingly widely used.
[0003] In order to facilitate the interaction and control of users with virtual display devices, voice assistants are usually installed in virtual display devices to achieve the interaction and control of virtual display devices through voice commands. Currently, the wake-up of voice assistants is generally performed by voice triggering, that is, the sound is recorded in real time through a microphone and voice detection is performed. When the set voice recognition wake-up voice is detected, voice recognition is started. This method of waking up a voice assistant by voice has a high requirement for the accuracy of voice recognition, and it is easy to have inaccurate voice recognition, resulting in a low accuracy rate of starting the voice assistant. Summary of the Invention
[0004] The embodiments of the present application provide a virtual reality voice interaction method, device, equipment, and storage medium to solve the technical problem that the method of waking up a voice assistant by voice in the related art is prone to inaccurate voice recognition, resulting in a low accuracy rate of starting the voice assistant, and effectively improve the accuracy rate of voice recognition wake-up.
[0005] In a first aspect, the embodiments of the present application provide a virtual reality voice interaction method applied to a virtual display device, including:
[0006] Obtain user gaze information and first interaction information;
[0007] If both the user gaze information and the first interaction information meet the preset conditions, then in response to the first interaction information, start voice recognition;
[0008] Wherein, the first interaction information is image information or motion information of the user's body movements.
[0009] This solution obtains user gaze information and first interaction information, determines whether the first preset condition for starting voice recognition is met according to the user gaze information and the first interaction information, and starts voice recognition when the first preset condition is met. The user can accurately start voice recognition through gaze and interaction actions, and the start of voice recognition does not need to rely on the accuracy of voice recognition of on-site recordings, effectively improving the accuracy rate of voice recognition wake-up.
[0010] Further, after obtaining the user's gaze information and the first interaction information, the following steps are also included:
[0011] Determine whether the user is gazing at the hand according to the user's gaze information and the first interaction information;
[0012] If it is determined that the user is gazing at the hand, then determine that both the user's gaze information and the first interaction information meet the first preset condition.
[0013] As described above, determining whether the user is gazing at the hand according to the user's gaze information and the first interaction information, and determining that both the user's gaze information and the first interaction information meet the first preset condition when it is determined that the user is gazing at the hand, accurately judges the timing to start voice recognition and improves the accuracy of virtual reality voice interaction.
[0014] Further, the determining whether the user is gazing at the hand according to the user's gaze information and the first interaction information includes:
[0015] Determine the gazing direction according to the user's gaze information;
[0016] Determine the hand position according to the first interaction information;
[0017] Determine whether the user is gazing at the hand according to the gazing direction and the hand position.
[0018] As described above, by determining the gazing direction according to the user's gaze information and the hand position according to the first interaction information, and accurately determining whether the user is gazing at the hand according to the gazing direction and the hand position, the accuracy of starting voice recognition is effectively improved.
[0019] Further, the starting of voice recognition in response to the first interaction information includes:
[0020] Determine the first gesture recognition information according to the first interaction information;
[0021] If the gesture corresponding to the first gesture recognition information is the first set gesture, then start voice recognition.
[0022] As described above, by determining the first gesture recognition information according to the first interaction information and starting voice recognition when the gesture corresponding to the first gesture recognition information is the first set gesture, the timing to start voice recognition is accurately judged and the accuracy of virtual reality voice interaction is improved.
[0023] Further, the starting of voice recognition if the gesture corresponding to the first gesture recognition information is the first set gesture includes:
[0024] If the gesture corresponding to the first gesture recognition information is a gesture from fingers together to fingers apart, or a gesture from a fist to an open palm, then voice recognition is activated.
[0025] As described above, by using the gesture of the hand from fingers together to fingers apart or the gesture of the hand from a fist to an open palm as the first set gesture to activate voice recognition, the accuracy of recognizing the first set gesture is relatively high, the user operation is simple, the user learning cost is low, and the efficiency and accuracy of the voice assistant are effectively improved.
[0026] Further, after obtaining the user's gaze information and the first interaction information, it further includes:
[0027] If both the user's gaze information and the first interaction information meet the first preset condition, then a virtual hand is displayed according to the first interaction information.
[0028] As described above, by displaying the user's virtual hand when both the user's gaze information and the first interaction information meet the first preset condition, it is convenient for the user to know that the virtual display device has recognized that the first preset condition is met, prompting the user to perform the first set gesture. At the same time, the user can understand the position of the hand and the change of the gesture through the displayed virtual hand, which is convenient for the user to operate the gesture and improves the startup efficiency of the voice assistant.
[0029] Further, after obtaining the user's gaze information and the first interaction information, it further includes:
[0030] If both the user's gaze information and the first interaction information meet the first preset condition, then a set light effect is rendered for the virtual hand displayed in the virtual display device.
[0031] As described above, by rendering a set light effect for the virtual hand when it is detected that the user is gazing at the hand, it is convenient for the user to know that the virtual display device has recognized that the user is gazing at the hand, prompting the user to perform the first set gesture and improving the startup efficiency of the voice assistant.
[0032] Further, after activating voice recognition in response to the first interaction information, it further includes:
[0033] A voice assistant interaction control is displayed. The voice assistant interaction control is used to display the interaction information of the voice assistant, and the display position of the voice assistant interaction control is determined according to the virtual hand displayed in the virtual display device.
[0034] As described above, through the voice assistant interaction control, the interaction situation with the voice assistant can be observed more intuitively, improving the usage experience of the voice assistant. And the voice assistant interaction control can move the position synchronously with the user's hand, making the interaction of the voice assistant more flexible and improving the user's usage experience.
[0035] Further, after starting voice recognition in response to the first interaction information, the following steps are also included:
[0036] Obtain second interaction information. If the second interaction information meets the second preset condition, then turn off the voice recognition.
[0037] As described above, by turning off the voice recognition when the second interaction information meets the second preset condition, the opening and closing of the voice assistant do not need to rely on the voice recognition accuracy of on-site recordings, effectively improving the accuracy of starting and closing the voice recognition.
[0038] Further, the step of turning off the voice recognition if the second interaction information meets the second preset condition includes:
[0039] Perform gesture recognition on the second interaction information to obtain second gesture recognition information;
[0040] If the gesture corresponding to the second gesture recognition information is the second set gesture, then turn off the voice recognition.
[0041] As described above, by turning off the voice recognition when the gesture corresponding to the second gesture recognition information is the second set gesture, the opening and closing of the voice assistant do not need to rely on the voice recognition accuracy of on-site recordings, effectively improving the accuracy of starting and closing the voice recognition.
[0042] Further, the step of turning off the voice recognition if the gesture corresponding to the second gesture recognition information is the second set gesture includes:
[0043] If the gesture corresponding to the second gesture recognition information is one or a combination of gestures from finger spreading to finger closing, palm spreading to palm clenching, and palm flipping, then turn off the voice recognition.
[0044] As described above, by using one or a combination of gestures from finger spreading to finger closing, palm spreading to palm clenching, and palm flipping of the hand as the second set gesture for turning off the voice recognition, the recognition accuracy of the second set gesture is relatively high, the user operation is simple, and the user learning cost is low, effectively improving the accuracy of turning off the voice recognition.
[0045] Further, the user gaze information includes eye movement recognition information and / or device orientation information. The eye movement recognition information is obtained by performing eye movement recognition through an eye movement recognition unit on the virtual display device. The first interaction information includes first external image information and / or first motion detection information. The first external image information is obtained by taking images through an image acquisition unit on the virtual display device. The first motion detection information is obtained by performing motion detection through a motion detection unit of an externally connected device.
[0046] As described above, the first interaction information of the user is reflected by the first external image information and / or the first motion detection information, accurately judging the user's limb movements, improving the accuracy of voice interaction, and accurately reflecting the user's gaze information through the eye movement recognition information and / or the device orientation information, improving the accuracy of voice interaction.
[0047] In a second aspect, an embodiment of the present application provides a virtual reality voice interaction device, which is applied to a virtual display device and includes an information acquisition module and a startup processing module, where:
[0048] The information acquisition module is configured to acquire user gaze information and first interaction information, where the first interaction information is image information or motion information of the user's limb movements;
[0049] The startup processing module is configured to, if the user gaze information and the first interaction information both meet the first preset condition, start voice recognition in response to the first interaction information.
[0050] In the embodiment of the present application, by acquiring the user gaze information and the first interaction information, it is determined whether the first preset condition for enabling voice recognition is met according to the user gaze information and the first interaction information, and voice recognition is started when the first preset condition is satisfied. The user can accurately start voice recognition through gaze and interaction actions. The startup of voice recognition does not need to rely on the voice recognition accuracy of on-site recording, effectively improving the voice recognition wake-up accuracy.
[0051] In a third aspect, an embodiment of the present application provides a virtual reality voice interaction device, including: a memory and one or more processors;
[0052] The memory is used to store one or more programs;
[0053] When the one or more programs are executed by the one or more processors, the one or more processors implement the virtual reality voice interaction method as described in the first aspect.
[0054] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer-executable instructions, and the computer-executable instructions are used to execute the virtual reality voice interaction method as described in the first aspect when executed by a computer processor. Description of the Drawings
[0055] Figure 1 is a flowchart of a virtual reality voice interaction method provided by an embodiment of the present application;
[0056] Figure 2 is a schematic block diagram of a virtual display device provided by an embodiment of the present application;
[0057] Figure 3It is a flowchart of another virtual reality voice interaction method provided by an embodiment of the present application;
[0058] Figure 4 It is a flowchart of another virtual reality voice interaction method provided by an embodiment of the present application
[0059] Figure 5 It is a schematic diagram of a set light effect display provided by an embodiment of the present application;
[0060] Figure 6 It is a schematic diagram of a first set gesture display provided by an embodiment of the present application;
[0061] Figure 7 It is a schematic diagram of the display of a voice assistant interaction control provided by an embodiment of the present application;
[0062] Figure 8 It is a schematic diagram of the response display of a voice assistant interaction control to an interrogative voice command provided by an embodiment of the present application;
[0063] Figure 9 It is a schematic diagram of the response display of a voice assistant interaction control to a command-type voice command provided by an embodiment of the present application;
[0064] Figure 10 It is a schematic diagram of a second set gesture display provided by an embodiment of the present application;
[0065] Figure 11 It is a schematic diagram of the structure of a virtual reality voice interaction device provided by an embodiment of the present application;
[0066] Figure 12 It is a schematic diagram of the structure of a virtual reality voice interaction device provided by an embodiment of the present application. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the specific embodiments of the present application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only for explaining the present application, rather than limiting the present application. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present application are shown in the drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the above-mentioned process can be terminated, but there can also be additional steps not included in the drawings. The above-mentioned process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0068] In the existing voice assistant interaction solutions for virtual display devices, the voice assistant is generally activated in response to a set wake word. When the microphone of the virtual display device receives the wake word, voice recognition is started. The user issues voice commands by speaking. After the voice assistant receives the voice command, it performs corresponding feedback actions. The activation of voice recognition requires the microphone to continuously record the on-site sound. In a scenario where multiple people are speaking simultaneously, it is easy to fail to recognize the wake word, resulting in the inability to correctly start voice recognition. At the same time, in the way of starting voice recognition by the wake word, it is easy to misjudge the user's intention due to voice recognition errors. For example, when the microphone receives a word that is homophonic or similar to the wake word, or when the wake word is obtained from a non-user, the voice assistant will be misactivated. The method of waking up the voice assistant by voice is prone to inaccurate voice recognition, and the startup accuracy rate of the voice assistant is relatively low.
[0069] Based on this, a virtual reality voice interaction method according to an embodiment of the present application is provided. The user can accurately start voice recognition by looking at the hand and initiating a first set gesture. The activation of voice recognition does not need to rely on the voice recognition accuracy of on-site recording, effectively solving the technical problem that the existing method of waking up the voice assistant by voice is prone to inaccurate voice recognition and the relatively low startup accuracy rate of the voice assistant, and effectively improving the wake-up accuracy rate of voice recognition.
[0070] Figure 1 A flowchart of a virtual reality voice interaction method provided by an embodiment of the present application is given. The virtual reality voice interaction method provided by the embodiment of the present application can be executed by a virtual reality voice interaction device. The virtual reality voice interaction device can be implemented in a hardware and / or software manner and integrated in a virtual reality voice interaction device (such as a virtual display device).
[0071] The following describes the virtual reality voice interaction method executed by the virtual reality voice interaction device as an example. The virtual reality voice interaction method can be applied to a virtual display device. Refer to Figure 1 , the virtual reality voice interaction method includes:
[0072] S110: Obtain user gaze information and first interaction information.
[0073] Exemplarily, obtain user gaze information and first interaction information in real time, and determine whether the user gaze information and the first interaction information meet a first preset condition. When both the user gaze information and the first interaction information meet the first preset condition, jump to step S120; otherwise, continuously obtain the user gaze information and the first interaction information until both the user gaze information and the first interaction information meet the first preset condition.
[0074] Among them, the first interaction information provided by this solution is the image information or motion information of the user's body movements. That is, the user's body movements can be determined based on the image information or motion information. When it is detected that the user's body movement is a set movement (such as a set gesture or action), it can be determined that the first interaction information meets the first preset condition.
[0075] In one embodiment, the first interaction information provided by this solution includes first external image information and / or first motion detection information. The first external image information is obtained by taking images through an image acquisition unit on the virtual display device, and the first motion detection information is obtained by performing motion detection on a motion detection unit of a connected external device (such as a finger ring, a watch, a remote control, etc.). This solution reflects the user's first interaction information through the first external image information and / or the first motion detection information, accurately judges the user's body movements, and improves the accuracy of voice interaction.
[0076] The user gaze information provided by this solution can reflect the orientation of the user or the virtual display device. When the orientation of the user or the virtual display device is a set orientation (such as the user or the virtual display device facing the hand or a set direction), it can be determined that the user gaze information meets the first preset condition. In a possible embodiment, the user gaze information provided by this solution includes eye movement recognition information and / or device orientation information. Among them, the eye movement recognition information can be obtained by performing eye movement recognition through an eye movement recognition unit on the virtual display device, and the device orientation information can be detected through a motion detection unit (such as an IMU module) configured on the virtual display device. This solution accurately reflects the user gaze information through the eye movement recognition information and / or the device orientation information, and improves the accuracy of voice interaction.
[0077] As Figure 2 As shown in the principle block diagram of a virtual display device provided, the virtual display device provided by this solution includes an eye movement recognition unit, a motion detection unit, an image acquisition unit, a voice acquisition unit, a processing unit, and a display unit. Among them, the eye movement recognition unit, the motion detection unit, the image acquisition unit, the voice acquisition unit, and the display unit are all connected to the processing unit.
[0078] In one embodiment, the processing unit can execute the virtual reality voice interaction method provided by the present solution, and the processing unit can control the picture displayed by the display unit. The display unit is installed inside the virtual display device. After the user wears the virtual display device, the display unit can observe the corresponding picture. The eye movement recognition unit can shoot the internal image in the virtual display device, perform eye movement recognition based on the internal image and output the corresponding eye movement recognition information to the processing unit, or send the internal image information to the processing unit, and the processing unit performs eye movement recognition based on the internal image information and determines the corresponding eye movement recognition information. The motion detection unit can detect the motion of the virtual display device, determine the device orientation information based on the motion detection result and output the corresponding device orientation information to the processing unit, or send the motion detection result to the processing unit, and the processing unit performs motion detection based on the internal image information and determines the corresponding device orientation information.
[0079] It needs to be explained that when a user wears a virtual display device, the internal image information captured records the user's eye image, and eye movement recognition of the internal image information can obtain the user's corresponding eye movement recognition information. The user's eye gaze direction and changes in the gaze direction can be identified based on the eye movement recognition information. The direction of the user or the virtual display device or the gaze position on the display screen can be determined based on the gaze direction. The image acquisition unit can be a camera installed on the virtual display device and facing the outside of the virtual display device. When the external image information captured by the image acquisition unit is displayed in the display unit, the displayed image is close to or consistent with the image observed by the user's eyes when the virtual display device is not worn.
[0080] Optionally, after acquiring the external image information, the image acquisition unit may perform gesture recognition on the external image information and output the corresponding hand recognition result (e.g., the hand position recognized in the external image) and the gesture recognition result to the processing unit. Alternatively, the external image information may be sent to the processing unit, which performs gesture recognition on the external image information and determines the corresponding hand recognition result (e.g., the hand position recognized in the external image) and the gesture recognition result. The virtual display device provided by this solution can effectively solve the problem that the voice assistant is easily triggered by the user himself or others by mistake by controlling the opening and closing of voice recognition based on user gaze information and interaction information, and by combining multiple recognition units, it provides users with a stable multimodal interaction experience.
[0081] In one embodiment, the voice acquisition unit can be a microphone, which can collect sound and perform voice recognition on the collected sound, and output the voice recognition result to the processing unit, or send the collected voice to the processing unit, which will perform voice recognition and determine the voice recognition result.
[0082] S120: If both the user gaze information and the first interaction information meet the first preset condition, then in response to the first interaction information, start voice recognition.
[0083] Exemplarily, when both the user gaze information and the first interaction information meet the first preset condition, in response to the first interaction information, start voice recognition. Optionally, after starting voice recognition, the user can be reminded that voice recognition has been started by means such as an animated diagram, a light effect prompt, or a voice prompt. The user can initiate a voice command by voice. After starting voice recognition, collect sound through a voice acquisition unit, perform voice recognition on the collected sound, and generate a corresponding voice command based on the voice recognition result. Input the voice command into the voice assistant, and the voice assistant can perform corresponding actions according to the voice command.
[0084] As described above, by obtaining the user gaze information and the first interaction information, determining whether the first preset condition for starting voice recognition is met according to the user gaze information and the first interaction information, and starting voice recognition when the first preset condition is satisfied, the user can accurately start voice recognition through gaze and interaction actions. The start of voice recognition does not need to rely on the voice recognition accuracy of on-site recording, effectively improving the voice recognition wake-up accuracy.
[0085] Based on the above embodiments, Figure 3 A flowchart of another virtual reality voice interaction method provided by an embodiment of the present application is given. This virtual reality voice interaction method is a concretization of the above virtual reality voice interaction method. Refer to Figure 3 and this virtual reality voice interaction method includes:
[0086] S210: Obtain the user gaze information and the first interaction information.
[0087] S220: Determine whether the user is gazing at the hand according to the user gaze information and the first interaction information. If it is determined that the user is gazing at the hand, then determine that both the user gaze information and the first interaction information meet the first preset condition.
[0088] Among them, the first interaction information is the image information or motion information of the user's body movements. The user gaze information provided by this solution includes eye movement recognition information and / or device orientation information, and the first interaction information includes first external image information and / or first motion detection information. Exemplarily, obtain the eye movement recognition information obtained by eye movement recognition through the eye movement recognition unit on the virtual display device and / or the device orientation information obtained by motion detection through the motion detection unit on the virtual display device, as well as the first external image information obtained by image capture through the image acquisition unit on the virtual display device and / or the first motion detection information obtained by motion detection through the motion detection unit of the connected external device, and determine whether the user is gazing at the hand according to the eye movement recognition information and / or device orientation information, and the first external image information and / or first motion detection information. For example, determine whether the position where the user is gazing as determined by the eye movement recognition information and / or device orientation information corresponds to the position of the user's hand reflected by the first external image information and / or first motion detection information (for example, the central position where the user is gazing as reflected by the eye movement recognition information is within the image range corresponding to the user's hand in the first external image information). If the position where the user is gazing corresponds to the position of the user's hand reflected by the first external image information and / or first motion detection information, or the duration during which the position where the user is gazing corresponds to the position of the user's hand reflected by the first external image information and / or first motion detection information reaches a set duration, it is considered that the user's gaze at the hand is detected. When it is determined that the user's gaze at the hand is detected, it can be determined that both the user gaze information and the first interaction information meet the first preset condition.
[0089] S230: If both the user gaze information and the first interaction information meet the first preset condition, then determine the first gesture recognition information according to the first interaction information. If the gesture corresponding to the first gesture recognition information is the first set gesture, then start voice recognition.
[0090] Exemplarily, when it is determined that both the user gaze information and the first interaction information meet the first preset condition, obtain the first gesture recognition information obtained by gesture recognition according to the first interaction information (first external image information and / or first motion detection information). Among them, when it is continuously detected that the user is gazing at the hand, continuously obtain the first gesture recognition information obtained by gesture recognition according to the first interaction information until the first set gesture is detected.
[0091] In one embodiment, when it is determined that the gesture corresponding to the first gesture recognition information is the first set gesture during the process of the user gazing at the hand, start voice recognition. This solution determines the first gesture recognition information according to the first interaction information and starts voice recognition when the gesture corresponding to the first gesture recognition information is the first set gesture, accurately judges the timing of starting voice recognition, and improves the accuracy of virtual reality voice interaction.
[0092] In one embodiment, when the user needs to wake up the voice assistant, the user can raise a hand (either the left or right hand), and look at the direction of the raised hand (the user can turn the eyesight of the glasses towards the direction of the raised hand, or turn the virtual display device towards the direction of the raised hand by turning the head). At this time, according to the user's gaze information and the first interaction information (such as eye movement recognition information and the first external image information), it can be determined whether the user is looking at the hand. At this time, the user can control the raised hand to execute a set gesture (such as a gesture of the hand opening from a closed state, waving the hand, swinging the hand, rotating the hand, etc.) to wake up the voice assistant.
[0093] In one embodiment, during the process of the user executing the set gesture, the first gesture recognition information obtained by gesture recognition based on the first external image information collected in real time can be acquired, and it can be determined whether the gesture corresponding to the first gesture recognition information is the first set gesture. If the gesture corresponding to the first gesture recognition information is not the first set gesture, then during the process of the user continuously looking at the hand, the first gesture recognition information obtained by gesture recognition based on the first external image information collected in real time is continuously acquired, and it is determined whether the gesture corresponding to the first gesture recognition information is the first set gesture until the first set gesture is detected or the user stops looking at the hand.
[0094] As described above, by acquiring the user's gaze information and the first interaction information, it is determined whether the first preset condition for enabling voice recognition is met according to the user's gaze information and the first interaction information, and voice recognition is started when the first preset condition is satisfied. The user can accurately start voice recognition through gaze and interaction actions. The start of voice recognition does not need to rely on the voice recognition accuracy of on-site recording, effectively improving the voice recognition wake-up accuracy. At the same time, it is determined whether the user is looking at the hand according to the user's gaze information and the first interaction information, and when it is determined that the user is looking at the hand, it is determined that both the user's gaze information and the first interaction information meet the first preset condition, accurately judging the timing of starting voice recognition and improving the virtual reality voice interaction accuracy.
[0095] Based on the above embodiment, Figure 4 a flowchart of another virtual reality voice interaction method provided by the embodiment of the present application is given. This virtual reality voice interaction method is a concretization of the above virtual reality voice interaction method. Refer to Figure 4 , this virtual reality voice interaction method includes:
[0096] S310: Acquire the user's gaze information and the first interaction information.
[0097] S320: Determine the gaze direction according to the user's gaze information, and determine the hand position according to the first interaction information.
[0098] S330: Determine whether the user is looking at the hand according to the gaze direction and the hand position.
[0099] Exemplarily, the user's gaze information and first interaction information are obtained in real time, the gaze direction is determined according to the user's gaze information, and the hand position is determined according to the first interaction information. For example, the eye movement recognition information obtained by the eye movement recognition unit for eye movement recognition and the first external image information obtained by the image acquisition unit for image shooting are obtained, and the user's gaze direction is determined according to the eye movement recognition information, or the device orientation information obtained by the motion detection unit for motion detection is obtained, and the user's gaze direction is determined according to the device orientation information. And the hand position of the hand detected in the first external image information is determined according to the hand recognition result of the first external image information, or the user's hand position is determined according to the first motion detection information.
[0100] Further, it is determined whether the user is gazing at the hand according to the determined gaze direction and hand position. Optionally, the gaze direction can be used to indicate the position in the picture displayed by the virtual display device that the user is gazing at, and the hand position can be used to indicate the range of the detected hand in the picture displayed by the virtual display device. When the position corresponding to the gaze direction is within the range corresponding to the hand position, or when the duration for which the position corresponding to the gaze direction is within the range corresponding to the hand position reaches a set duration, it can be considered that the user is gazing at the hand.
[0101] This solution accurately determines whether the user is gazing at the hand according to the gaze direction determined according to the user's gaze information and the hand position determined according to the first interaction information, effectively improving the accuracy of starting voice recognition.
[0102] In a possible embodiment, after obtaining the user's gaze information and first interaction information, the virtual reality voice interaction method provided by this solution may further include: if both the user's gaze information and the first interaction information meet the first preset condition, then display a virtual hand according to the first interaction information.
[0103] Exemplarily, when it is determined that both the user's gaze information and the first interaction information meet the first preset conditions, the hand position can be determined according to the first interaction information, and a virtual hand can be displayed in the virtual display device based on the hand position. For example, when it is determined that the user is gazing at the hand, hand recognition is performed on the first external image information obtained in real time to determine the hand recognition result in the first external image, and the user's virtual hand is rendered in the picture displayed by the virtual display device according to the hand recognition result. Among them, the shape and size of the virtual hand are close to or the same as those of the hand actually observed by the user with the naked eye. For example, when the virtual display device does not display the user's virtual hand (for example, when the user's virtual hand is hidden in the immersive mode), if it is detected that the user is gazing at the hand, the user's virtual hand is displayed. At this time, the user can know that the virtual display device has recognized that the user is gazing at the hand and can perform corresponding gesture actions. By displaying the user's virtual hand when both the user's gaze information and the first interaction information meet the first preset conditions, it is convenient for the user to know that the virtual display device has recognized that the conditions are met, prompting the user to perform the first set gesture. At the same time, the user can understand the position of the hand and the change of the gesture through the displayed virtual hand, which is convenient for the user to operate the gesture and improves the startup efficiency of the voice assistant.
[0104] In a possible embodiment, after obtaining the user's gaze information and the first interaction information, the virtual reality voice interaction method provided by this solution may further include: if both the user's gaze information and the first interaction information meet the first preset conditions, a set light effect is rendered on the virtual hand displayed in the virtual display device. Among them, the virtual hand is determined based on the first interaction information (for example, determined based on the user's hand recognized in the first external image information).
[0105] Exemplarily, when it is determined that both the user's gaze information and the first interaction information meet the first preset conditions, a set light effect is rendered according to the virtual hand displayed in the virtual display device. Among them, the virtual hand displayed in the virtual display device can be rendered based on the user's hand determined by performing hand recognition on the first external image information obtained in real time. The set light effect can be an animation effect (such as a flash effect, a text effect, etc.) played at the corresponding position of the virtual hand, or an effect on the virtual hand (such as highlighting, flashing the display of the virtual hand or the virtual hand contour, etc.).
[0106] Such as Figure 5A schematic diagram of a set light effect display is provided. When a user wears a virtual display device (the virtual display device is hidden in the figure), if it is detected that the user is looking at the hand, the set light effect is rendered at the palm position of the virtual hand displayed by the virtual display device (for example, the palm part of the virtual hand is highlighted). At this time, the user can observe the set light effect at the palm position of the virtual hand. The user can determine that the virtual display device has recognized the gaze on the hand, and can continue with the gesture action of the first set gesture to start voice recognition.
[0107] This solution sets the light effect for rendering the virtual hand when detecting that the user is looking at the hand, so that the user can understand that the virtual display device has recognized that the user is looking at the hand, prompting the user to perform the first set gesture, thereby improving the efficiency of starting the voice assistant.
[0108] S340: If it is determined that the user is gazing at the hand, determine that both the user gaze information and the first interaction information meet the first preset condition.
[0109] S350: If both the user gaze information and the first interaction information meet the first preset condition, voice recognition is initiated in response to the first interaction information.
[0110] In a possible embodiment, the first set gesture may be a gesture of the user's hand from closed to open. Based on this, when the virtual reality voice interaction method provided by the present solution starts voice recognition if the gesture corresponding to the first gesture recognition information is the first set gesture, it may be: if the gesture corresponding to the first gesture recognition information is a gesture of fingers closing together to fingers opening, or a gesture of hands clenching into a fist to an open palm, then voice recognition is started.
[0111] For example, the hand gesture from closed to open provided by this solution can be a gesture corresponding to the hand from fingers together to fingers open, or a gesture corresponding to the hand from fist to palm open. That is, after determining that the user is looking at the hand and obtaining the first gesture recognition information, when the gesture corresponding to the first gesture recognition information is a gesture from fingers together to fingers open, or a gesture from fist to palm open, voice recognition is started. Figure 6 A first set gesture display schematic diagram is provided, which shows the initial state, intermediate state and final state of two first set gestures, namely, the gesture corresponding to the hand from closing the fingers to opening the fingers, and the gesture corresponding to the hand from clenching the fist to opening the palm. Optionally, when the change of the hand from the initial state to the final state is detected, it can be considered that the first set gesture is seen.
[0112] In this solution, a gesture where the hand changes from fingers together to fingers apart, or a gesture where the hand changes from a fist to an open palm, is used as the first set gesture to initiate voice recognition. The accuracy of recognizing the first set gesture is relatively high, the user operation is simple, and the user learning cost is low, effectively improving the efficiency and accuracy of the voice assistant.
[0113] In a possible embodiment, after the virtual reality voice interaction method provided in this solution initiates voice recognition in response to the first interaction information, it may further include: displaying a voice assistant interaction control, which is used to display the interaction information of the voice assistant, and the display position of the voice assistant interaction control is determined according to the virtual hand displayed in the virtual display device.
[0114] Exemplarily, after determining that the first set gesture is detected and voice recognition is initiated, the display position of the voice assistant interaction control is determined according to the virtual hand displayed in the virtual display device, and the voice assistant interaction control is displayed at this display position. Among them, the display position of the voice assistant interaction control may be a position corresponding to the palm of the virtual hand. For example, the center position of the voice assistant interaction control coincides with the palm center position of the virtual hand, or the center position of the voice assistant interaction control is at a position above the palm center position of the virtual hand. At the same time, as the user's hand moves, the virtual hand and the voice assistant interaction control displayed in the virtual display device move synchronously with the user's hand.
[0115] Among them, the voice assistant interaction control provided in this solution is used to display the interaction information of the voice assistant, such as displaying the voice commands received by the voice assistant, displaying the relevant information of the actions executed based on the voice commands, displaying the reply information to the voice commands, etc. Optionally, the voice assistant interaction control provided in this solution may be displayed in the form of a speech bubble. At this time, when a voice command is issued to the voice assistant, the speech bubble can provide visual bubbles and voice feedback related information. As Figure 7 A schematic diagram of the display of a voice assistant interaction control is provided, where the voice assistant interaction control is displayed in the form of a speech bubble. During the movement of the user's hand, the virtual hand and the speech bubble displayed in the virtual display device move synchronously with the user's hand, achieving the display effect of the speech bubble following the user's hand synchronously.
[0116] After the voice assistant interaction control is displayed, different types of voice commands can be issued to the voice assistant through voice. The types of voice commands can be interrogative, command-type, etc. The voice assistant processes different types of voice commands differently and provides different feedback through the voice assistant interaction control. As Figure 8As shown in the response display schematic diagram of a voice assistant interaction control for an interrogative voice command, when an interrogative voice command "What's the weather like today?" is sent to the voice assistant by voice, after the voice assistant obtains relevant weather information, it broadcasts the weather information by voice and displays the corresponding weather information of "There will be showers today, and the temperature is 26°C" through the voice assistant interaction control. As Figure 9 As shown in the response display schematic diagram of a voice assistant interaction control for a command voice command, when a command voice command "Open the photo album" is sent to the voice assistant by voice, the voice assistant opens the photo album application in the virtual environment and displays the corresponding feedback information of "The photo album has been opened for you" through the voice assistant interaction control.
[0117] With this solution, the interaction with the voice assistant can be observed more intuitively through the voice assistant interaction control, improving the usage experience of the voice assistant. Moreover, the voice assistant interaction control can move synchronously with the user's hand, making the interaction with the voice assistant more flexible and enhancing the user's usage experience.
[0118] In a possible embodiment, after starting voice recognition in response to the first interaction information, the virtual reality voice interaction method provided by this solution may further include: obtaining second interaction information, and if the second interaction information meets the second preset condition, closing the voice recognition.
[0119] Exemplarily, after starting voice recognition, continuously obtain second interaction information (including second external image information and / or second motion detection information), and determine whether the second interaction information meets the second preset condition. For example, obtain the second external image information through the image acquisition unit, perform gesture recognition based on the second external image information to obtain the second gesture recognition information, and determine whether the gesture corresponding to the second gesture recognition information is the second set gesture, or obtain the second motion detection information through the motion detection unit, perform action recognition based on the second motion detection information to obtain the second motion recognition information, and determine whether the action corresponding to the second motion detection information is the second set action. When the gesture corresponding to the second gesture recognition information is the second set gesture and / or the action corresponding to the second motion detection information is the second set action, determine whether the second interaction information meets the second preset condition.
[0120] Among them, when the second interaction information does not meet the second preset condition, the voice assistant continues to run, and the user can continue to issue voice commands to the user assistant. When the second interaction information meets the second preset condition, the voice recognition is turned off. For example, when the user needs to turn off the voice recognition, the user can execute a second preset gesture. After detecting the second preset gesture executed by the user, the voice recognition is turned off, and the displayed voice assistant interaction control is closed. This solution turns off the voice recognition when the second interaction information meets the second preset condition. The opening and closing of the voice assistant do not depend on the voice recognition accuracy of the on-site recording, effectively improving the accuracy of turning on and off the voice recognition.
[0121] In one embodiment, when the virtual reality voice interaction method provided by this solution turns off the voice recognition if the gesture corresponding to the second gesture recognition information is the second preset gesture, it may be: if the gesture corresponding to the second gesture recognition information is one or a combination of gestures such as the gesture of spreading the fingers to closing the fingers, the gesture of spreading the palm to making a fist, and the gesture of turning the palm over, the voice recognition is turned off.
[0122] Exemplarily, the gesture for turning off the voice recognition provided by this solution may be the gesture corresponding to the gesture of spreading the fingers to closing the fingers of the hand, the gesture corresponding to the gesture of spreading the palm to making a fist of the hand, or the gesture corresponding to turning the palm over. That is, after starting the voice recognition, the second gesture recognition information is continuously obtained. When the gesture corresponding to the second gesture recognition information is one or a combination of gestures such as the gesture of spreading the fingers to closing the fingers, the gesture of spreading the palm to making a fist, and the gesture of turning the palm over, the voice recognition is turned off. As Figure 10 As shown in the schematic diagram of the display of a second preset gesture provided, during the operation of the voice assistant, when the second preset gesture of turning the palm over is detected, the voice recognition is turned off, and the displayed voice assistant interaction control is closed.
[0123] This solution uses one or a combination of gestures such as the gesture of spreading the fingers to closing the fingers of the hand, the gesture of spreading the palm to making a fist of the hand, and the gesture of turning the palm over as the second preset gesture for turning off the voice recognition. The accuracy of recognizing the second preset gesture is relatively high, the user operation is simple, and the user learning cost is low, effectively improving the accuracy of turning off the voice recognition.
[0124] As described above, by obtaining the user's gaze information and the first interaction information, determining whether the first preset condition for enabling speech recognition is met based on the user's gaze information and the first interaction information, and starting speech recognition when the first preset condition is satisfied, the user can accurately start speech recognition through gaze and interaction actions. The start of speech recognition does not depend on the speech recognition accuracy of on-site recording, effectively improving the speech recognition wake-up accuracy. At the same time, by obtaining the eye movement recognition information and the first external image information, determining whether the user is gazing at the hand based on the eye movement recognition information and the first external image information, when it is determined that the user is gazing at the hand, obtaining the first gesture recognition information obtained by gesture recognition based on the first external image information, and starting speech recognition when the gesture corresponding to the first gesture recognition information is the first set gesture. The user can accurately start speech recognition by gazing at the hand and initiating the first set gesture. The start of speech recognition does not depend on the speech recognition accuracy of on-site recording, effectively improving the speech recognition wake-up accuracy. At the same time, by determining the gaze direction according to the eye movement recognition and the hand position according to the first external image information, accurately determining whether the user is gazing at the hand based on the gaze direction and the hand position, effectively improving the accuracy of starting speech recognition. It also enables a more intuitive observation of the interaction with the voice assistant through the voice assistant interaction control, improving the usage experience of the voice assistant. Moreover, the voice assistant interaction control can move synchronously with the user's hand, making the interaction with the voice assistant more flexible and improving the user's usage experience.
[0125] Figure 11 FIG. shows a schematic structural diagram of a virtual reality voice interaction device provided by an embodiment of the present application. Refer to Figure 11 As shown in, the virtual reality voice interaction device includes an information acquisition module 31 and a start processing module 32.
[0126] Among them, the information acquisition module 31 is used to acquire the user's gaze information and the first interaction information, where the first interaction information is the image information or motion information of the user's body movements; the start processing module 32 is used to start speech recognition in response to the first interaction information if both the user's gaze information and the first interaction information meet the first preset condition.
[0127] As described above, by obtaining the user's gaze information and the first interaction information, determining whether the first preset condition for enabling speech recognition is met based on the user's gaze information and the first interaction information, and starting speech recognition when the first preset condition is satisfied, the user can accurately start speech recognition through gaze and interaction actions. The start of speech recognition does not depend on the speech recognition accuracy of on-site recording, effectively improving the speech recognition wake-up accuracy.
[0128] In a possible embodiment, the virtual reality voice interaction device further includes a first judgment module, and the first judgment module is used for:
[0129] Determine whether the user is gazing at the hand based on the user's gaze information and the first interaction information;
[0130] If it is determined that the user is gazing at the hand, it is determined that both the user's gaze information and the first interaction information meet the first preset condition.
[0131] In a possible embodiment, when the first determination module determines whether the user is gazing at the hand based on the user's gaze information and the first interaction information, it includes:
[0132] Determine the gazing direction based on the user's gaze information;
[0133] Determine the hand position based on the first interaction information;
[0134] Determine whether the user is gazing at the hand based on the gazing direction and the hand position.
[0135] In a possible embodiment, when the start processing module 32 starts voice recognition in response to the first interaction information, it includes:
[0136] Determine the first gesture recognition information based on the first interaction information;
[0137] If the gesture corresponding to the first gesture recognition information is the first set gesture, start voice recognition.
[0138] In a possible embodiment, when the start processing module 32 starts voice recognition if the gesture corresponding to the first gesture recognition information is the first set gesture, it includes:
[0139] If the gesture corresponding to the first gesture recognition information is a gesture from fingers together to fingers apart, or a gesture from a fist to an open palm, start voice recognition.
[0140] In a possible embodiment, the virtual reality voice interaction device further includes a virtual hand display module, and the virtual hand display module is used for:
[0141] If both the user's gaze information and the first interaction information meet the first preset condition, display a virtual hand according to the first interaction information.
[0142] In a possible embodiment, the virtual reality voice interaction device further includes a light effect rendering module, and the light effect rendering module is used for:
[0143] If both the user's gaze information and the first interaction information meet the first preset condition, render a set light effect on the virtual hand displayed in the virtual display device.
[0144] In a possible embodiment, the virtual reality voice interaction device further includes a control display module, and the control display module is used after the start processing module 32 starts voice recognition in response to the first interaction information:
[0145] A voice assistant interaction control is displayed. The voice assistant interaction control is used to display interaction information of the voice assistant, and the display position of the voice assistant interaction control is determined according to the virtual hand displayed in the virtual display device.
[0146] In a possible embodiment, the virtual reality voice interaction device further includes a shutdown processing module, and the shutdown processing module is configured to:
[0147] Obtain second interaction information, and if the second interaction information meets a second preset condition, turn off voice recognition.
[0148] In a possible embodiment, when the shutdown processing module turns off voice recognition if the second interaction information meets the second preset condition, it includes:
[0149] Perform gesture recognition on the second interaction information to obtain second gesture recognition information;
[0150] If the gesture corresponding to the second gesture recognition information is a second set gesture, turn off voice recognition.
[0151] In a possible embodiment, when the shutdown processing module turns off voice recognition if the gesture corresponding to the second gesture recognition information is the second set gesture, it includes:
[0152] If the gesture corresponding to the second gesture recognition information is one or a combination of gestures from finger spreading to finger closing, palm spreading to palm clenching, and palm flipping, turn off voice recognition.
[0153] In a possible embodiment, the user gaze information includes eye movement recognition information and / or device orientation information. The eye movement recognition information is obtained by performing eye movement recognition through an eye movement recognition unit on the virtual display device. The first interaction information includes first external image information and / or first motion detection information. The first external image information is obtained by taking an image through an image acquisition unit on the virtual display device, and the first motion detection information is obtained by performing motion detection on a motion detection unit of an externally connected device.
[0154] It should be noted that in the embodiments of the above virtual reality voice interaction device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present application.
[0155] The embodiments of the present application further provide a virtual reality voice interaction device, and this virtual reality voice interaction device can integrate the virtual reality voice interaction device provided by the embodiments of the present application. Figure 12It is a schematic structural diagram of a virtual reality voice interaction device provided by an embodiment of the present application. Refer to Figure 12 , the virtual reality voice interaction device includes: an input device 43, an output device 44, a memory 42, and one or more processors 41; the memory 42 is used to store one or more programs; when the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the virtual reality voice interaction method provided in the above embodiment. Among them, the input device 43, the output device 44, the memory 42, and the processor 41 can be connected through a bus or other means, Figure 12 Taking the connection through the bus as an example.
[0156] The memory 42, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the virtual reality voice interaction method provided in any embodiment of the present application (for example, the information acquisition module 31 and the startup processing module 32 in the virtual reality voice interaction device). The memory 42 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device. In addition, the memory 42 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 42 can further include a memory remotely set relative to the processor 41, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0157] The input device 43 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 44 can include display devices such as a display screen.
[0158] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 42, that is, implements the above virtual reality voice interaction method.
[0159] The above-provided virtual reality voice interaction device, equipment, and computer can be used to execute the virtual reality voice interaction method provided in any of the above embodiments, and have corresponding functions and beneficial effects.
[0160] An embodiment of the present application further provides a storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute the virtual reality voice interaction method provided in the above embodiment. The virtual reality voice interaction method includes: obtaining user gaze information and first interaction information; if both the user gaze information and the first interaction information meet a first preset condition, then in response to the first interaction information, start voice recognition; wherein, the first interaction information is image information or motion information of the user's body movements.
[0161] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media that may reside in different locations (such as in different computer systems connected via a network). The storage medium may store program instructions executable by one or more processors (such as specifically implemented as a computer program).
[0162] Of course, for the storage medium storing computer-executable instructions provided in an embodiment of the present application, the computer-executable instructions are not limited to the virtual reality voice interaction method provided above, and may also execute related operations in the virtual reality voice interaction method provided in any embodiment of the present application.
[0163] The virtual reality voice interaction device, equipment, and storage medium provided in the above embodiments can execute the virtual reality voice interaction method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, reference can be made to the virtual reality voice interaction method provided in any embodiment of the present application.
[0164] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments provided here. Various obvious changes, re-adjustments and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, it may also include more other equivalent embodiments, and the scope of the present application is determined by the scope of the claims.
Claims
1. A virtual reality voice interaction method, applied to a virtual display device, characterized in that include: Acquiring user gaze information and first interaction information; If both the user gaze information and the first interaction information meet the first preset condition, in response to the first interaction information, start voice recognition; The first interactive information is image information or motion information of the user's body movements.
2. The virtual reality voice interaction method according to claim 1, characterized in that After obtaining the user gaze information and the first interaction information, the method further includes: determining whether the user is gazing at the hand according to the user gaze information and the first interaction information; If it is determined that the user is gazing at the hand, it is determined that both the user gazing information and the first interaction information meet a first preset condition.
3. The virtual reality voice interaction method according to claim 1, wherein, The determining whether the user is gazing at the hand according to the user gaze information and the first interaction information includes: Determining a gaze direction according to the user gaze information; determining a hand position according to the first interaction information; It is determined whether the user is gazing at the hand according to the gaze direction and the hand position.
4. The virtual reality voice interaction method according to claim 1, characterized in that The step of starting speech recognition in response to the first interaction information includes: determining first gesture recognition information according to the first interaction information; If the gesture corresponding to the first gesture recognition information is a first set gesture, voice recognition is started.
5. The virtual reality voice interaction method according to claim 4, wherein If the gesture corresponding to the first gesture recognition information is a first set gesture, starting voice recognition includes: If the gesture corresponding to the first gesture recognition information is a gesture of fingers closing together to fingers opening apart, or a gesture of hand clenching into a fist to a palm opening, voice recognition is started.
6. The virtual reality voice interaction method according to claim 1, wherein After obtaining the user gaze information and the first interaction information, the method further includes: If the user gaze information and the first interaction information both meet the first preset condition, the virtual hand is displayed according to the first interaction information.
7. The virtual reality voice interaction method according to claim 1, characterized in that, After obtaining the user gaze information and the first interaction information, the method further includes: If both the user gaze information and the first interaction information meet the first preset condition, a light effect is set for rendering the virtual hand displayed in the virtual display device.
8. The virtual reality voice interaction method according to claim 1, wherein After starting the speech recognition in response to the first interaction information, the method further includes: Display a voice assistant interaction control, where the voice assistant interaction control is used to display interaction information of the voice assistant, and a display position of the voice assistant interaction control is determined according to a virtual hand displayed in the virtual display device.
9. The virtual reality voice interaction method according to claim 1, wherein After starting the speech recognition in response to the first interaction information, the method further includes: The second interaction information is obtained, and if the second interaction information meets the second preset condition, the voice recognition is turned off.
10. The virtual reality voice interaction method according to claim 9, wherein If the second interaction information meets the second preset condition, turning off the voice recognition includes: Performing gesture recognition according to the second interaction information to obtain second gesture recognition information; If the gesture corresponding to the second gesture recognition information is the second set gesture, voice recognition is turned off.
11. The virtual reality voice interaction method according to claim 10, wherein If the gesture corresponding to the second gesture recognition information is a second set gesture, turning off voice recognition includes: If the gesture corresponding to the second gesture recognition information is a combination of one or more of a gesture of spreading fingers to closing fingers, a gesture of opening the palm to making a fist, and a gesture of turning the palm over, voice recognition is turned off.
12. The virtual reality voice interaction method according to any one of claims 1-11, characterized in that, The user gaze information includes eye movement recognition information and / or device orientation information. The eye movement recognition information is obtained by performing eye movement recognition through an eye movement recognition unit on the virtual display device. The first interaction information includes first external image information and / or first motion detection information. The first external image information is obtained by performing image capture through an image acquisition unit on the virtual display device. The first motion detection information is obtained by performing motion detection on a motion detection unit of an externally connected device.