Virtual space interaction method and device, computer equipment, readable storage medium and program product
By combining hand trajectories and voice information, smart glasses can more accurately control target objects in the virtual space, solving the problem of inaccurate gesture control in existing technologies and achieving more efficient virtual space interaction.
Patent Information
- Application Number
- CN202510830390.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-19
AI Technical Summary
In the prior art, when smart glasses utilize gesture control, the recognition and control of complex gestures are not accurate enough.
By obtaining the hand trajectory in the virtual space, the gesture type and target object are determined, and the detailed control information is determined by combining the voice information. The target object is controlled by comprehensively utilizing the gesture and voice information.
It improves the control accuracy of virtual space interaction, reduces dependence on gesture operations, and improves the convenience and accuracy of operations.
Smart Images

Figure CN120669864A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a virtual space interaction method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art
[0002] Smart glasses, as a combination of AI (artificial intelligence) technology and wearable devices, are becoming a new focus in the technology industry.
[0003] When using smart glasses, there are cases in related technologies where hand gestures are used to achieve control, such as grabbing objects in virtual reality.
[0004] However, in related arts, gesture control is used, but when the gestures are complex, it is not possible to accurately recognize and control the gestures. Summary of the Invention
[0005] Based on this, it is necessary to provide a virtual space interaction method, device, computer equipment, computer-readable storage medium and computer program product that can improve accuracy in response to the above technical problems.
[0006] In a first aspect, the present application provides a virtual space interaction method, applied to smart glasses, comprising:
[0007] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0008] Acquire voice information and determine detailed control information based on the voice information;
[0009] The target object is controlled based on the coarse control information and the detailed control information.
[0010] In one embodiment, a hand trajectory in a virtual space is obtained, and a gesture type is determined based on the hand trajectory, including: obtaining multiple consecutive frames of images within a preset time period, identifying hand key points in the consecutive multiple frames of images respectively, and obtaining the position coordinates of the hand key points in the corresponding consecutive multiple frames of images; determining the hand trajectory using the position coordinates of the hand key points in the consecutive multiple frames of images; obtaining a preset trajectory-gesture type relationship, and determining the gesture type corresponding to the hand trajectory based on the trajectory-gesture type relationship.
[0011] In one embodiment, obtaining voice information includes: obtaining initial voice data, and performing endpoint detection on the initial voice data to obtain at least one intermediate voice data; and performing noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0012] In an optional embodiment, determining the detail control information based on the voice information includes: performing word segmentation processing on the voice information to obtain at least one word segmentation data; and querying a detail control data table using the at least one word segmentation data to obtain the detail control information.
[0013] In one embodiment, before controlling the target object according to the coarse control information and the detailed control information, the method further includes: obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, executing the step of controlling the target object according to the coarse control information and the detailed control information.
[0014] In one embodiment, before determining the target object and coarse control information according to the gesture type, the method further includes: obtaining an eyeball image and extracting the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, executing the step of controlling the target object according to the coarse control information and the detailed control information.
[0015] In a second aspect, the present application further provides a virtual space interaction device, comprising:
[0016] a first information determination module, configured to obtain a hand trajectory in a virtual space, determine a gesture type based on the hand trajectory, and determine a target object and rough control information based on the gesture type;
[0017] a second information determination module, configured to obtain voice information and determine detailed control information based on the voice information;
[0018] The control module is used to control the target object according to the coarse control information and the detailed control information.
[0019] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0020] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0021] Acquire voice information and determine detailed control information based on the voice information;
[0022] The target object is controlled based on the coarse control information and the detailed control information.
[0023] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0024] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0025] Acquire voice information and determine detailed control information based on the voice information;
[0026] The target object is controlled based on the coarse control information and the detailed control information.
[0027] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0028] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0029] Acquire voice information and determine detailed control information based on the voice information;
[0030] The target object is controlled based on the coarse control information and the detailed control information.
[0031] The above-mentioned virtual space interaction method, device, computer equipment, computer-readable storage medium and computer program product use hand trajectories to determine rough control information, use voice information to determine detailed control information, and implement node control information based on the rough control information and detailed control information, and use the detailed control information to supplement the rough control information, thereby lowering the operation threshold, avoiding excessive reliance on gestures, and improving control accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 1 is a flow chart of a virtual space interaction method according to an embodiment;
[0034] Figure 2 is a flowchart of a virtual space interaction method in another embodiment;
[0035] Figure 3 is a structural block diagram of a virtual space interaction device in one embodiment;
[0036] Figure 4 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0038] In one embodiment, Figure 1 As shown, a virtual space interaction method is provided. This embodiment uses the method applied to smart glasses as an example for illustration. It is understood that the method can also be applied to smart glasses, and can also be applied to a system including smart glasses and a server, and implemented through the interaction between the smart glasses and the server. In this embodiment, the method includes the following steps:
[0039] Step 102 : Acquire the hand trajectory in the virtual space, determine the gesture type and target object according to the hand trajectory, and determine rough control information according to the gesture type.
[0040] Among them, virtual space is used to represent the virtual environment corresponding to the real world, which is constructed by smart glasses through computer technology and network technology.
[0041] The hand trajectory is used to represent the trajectory of hand movement. Optionally, the smart glasses obtain hand movement data in the virtual space and determine the hand trajectory based on the hand movement data.
[0042] In one embodiment, multiple frames of images may be acquired for the smart glasses, and hand movement data may be determined based on the multiple frames of images, and a hand trajectory may be determined based on the determined hand movement data.
[0043] In one embodiment, the smart glasses include a camera device for realizing functions such as taking photos, recording videos, and image recognition.
[0044] In one embodiment, the smart glasses use a camera to obtain information about the surrounding environment and use the environmental information to generate a virtual space.
[0045] Optionally, the camera device in the smart glasses can be used to acquire multiple frames of images.
[0046] In an optional embodiment, a detection module can be set up to detect whether the hand moves in the virtual space. If so, the movement trajectory of the hand can be determined, and the hand trajectory can be determined based on the movement trajectory of the hand, and the gesture type can be determined based on the hand trajectory.
[0047] The gesture type is used to represent the types corresponding to different gestures.
[0048] The rough control information can be used to determine the preliminary control information corresponding to the target object. For example, the rough control information can be, for example, moving the target object or grasping the target object, and the preliminary control information can be, for example, moving the target object or grasping the target object.
[0049] In an optional embodiment, the coarse control information may be at least one. Optionally, when there are multiple coarse control information, multiple preliminary control information are generated to control the target object according to the multiple preliminary control information. For example, the coarse control information may be to move the target object horizontally to a preset distance and then vertically to a shelf. The multiple preliminary control information may be to move the target object horizontally to the preset distance and then vertically to the shelf. The smart glasses control the target object according to the multiple preliminary control information obtained.
[0050] Optionally, at least one preset coarse control information and at least one preset gesture type can be pre-set, and the preset coarse control information and the preset gesture type can be set correspondingly, and the association between the preset coarse control information and the preset gesture type can be stored so that when necessary, the association can be used to determine the coarse control information corresponding to the gesture type.
[0051] In one embodiment, after the gesture type is determined, the coarse control information corresponding to the gesture type is determined according to the correspondence between the coarse control information and the gesture type.
[0052] In one embodiment, a hand trajectory in a virtual space is obtained, and a gesture type is determined based on the hand trajectory, including: obtaining multiple consecutive frames of images within a preset time period, respectively identifying hand key points in the consecutive multiple frames of images, and obtaining position coordinates of the hand key points in the corresponding consecutive multiple frames of images; determining the hand trajectory using the position coordinates of the hand key points in the consecutive multiple frames of images; obtaining a preset trajectory-gesture type relationship, and determining the gesture type corresponding to the hand trajectory based on the trajectory-gesture type relationship.
[0053] Among them, the hand key points are used to determine the calibration points of the hand.
[0054] The trajectory gesture type relationship is used to characterize the association between the preset gesture type and the preset hand trajectory. Optionally, when the hand trajectory is obtained, the surgery type corresponding to the hand trajectory is determined based on the hand trajectory and the trajectory surgery type relationship.
[0055] In an optional embodiment, the key points of the hand may be finger joints. Optionally, there may be at least one key point of the hand.
[0056] The target object is used to represent the object that is desired to be controlled.
[0057] In one embodiment, the target object can be determined based on the hand trajectory. For example, the hand movement direction is determined based on the hand trajectory, and whether there is a preselected object within the hand movement direction range is identified based on the hand movement direction, and the target object is determined based on the preselected object.
[0058] In an exemplary embodiment, it is also possible to obtain an eyeball image and extract the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, perform steps of controlling the target object according to the coarse control information and the detailed control information.
[0059] In one embodiment, when there is only one pre-selected object, the pre-selected object may be identified as the target object.
[0060] In one embodiment, when there are multiple pre-selected objects, the coordinates of the pupils can be identified and matched with the center coordinates of the pre-selected objects. When a match is found, the pre-selected objects can be regarded as target objects.
[0061] Optionally, the smart glasses may also include an infrared camera to capture eye images.
[0062] In one embodiment, after obtaining the eyeball image, the smart glasses can process the eyeball image to determine the center coordinates of the eyeball, and match the obtained center coordinates of the eyeball with the center coordinates of the target object. If the center coordinates of the eyeball match the center coordinates of the preselected object, the preselected object can be considered as the target object, and the target object can be controlled based on the coarse control information and the detailed control information.
[0063] Step 104: Acquire voice information and determine detailed control information based on the voice information.
[0064] Step 106: Control the target object according to the coarse control information and the detailed control information.
[0065] The detailed control information is used to represent the subtle control of the target object, for example, the distance the target object is moved, the angle the target object is rotated, etc.
[0066] Optionally, when controlling a target object, the coarse control information and the detailed control information may be combined to more accurately control the target object. For example, if the coarse control information is to rotate the target object and the detailed control information is to rotate the target object by 30°, the control of the target object may be to rotate the target object by 30°.
[0067] In one embodiment, obtaining voice information includes: obtaining initial voice data, and performing endpoint detection on the initial voice data to obtain at least one intermediate voice data; and performing noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0068] Optionally, the initial voice data may be subjected to noise reduction processing and then endpoint detection may be performed to obtain at least one voice data, and voice information may be generated based on the voice data. Exemplarily, the obtained at least one voice data may be spliced to generate voice information.
[0069] Optionally, the noise reduction process may be spectral subtraction, adaptive filtering, Wiener filtering, etc.
[0070] In one embodiment, determining the detail control information based on the voice information includes: performing word segmentation processing on the voice information to obtain at least one word segmentation data; and querying the detail control data table using the at least one word segmentation data to obtain the detail control information.
[0071] The detail control data table is a preset table for storing the relationship between preset detail control data and preset voice information.
[0072] Optionally, after acquiring the voice information, the detail control data table may be queried to determine the detail control data corresponding to the voice information, and the target object may be controlled according to the detail control data.
[0073] In one embodiment, to ensure security, before controlling the target object based on the coarse control information and the detailed control information, the method further includes: obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, executing the step of controlling the target object based on the coarse control information and the detailed control information.
[0074] In one embodiment, a speech recognition model may be pre-trained, speech information may be input into the speech recognition model, and detailed control information corresponding to the speech information output by the speech recognition model may be obtained.
[0075] In one embodiment, after acquiring the voice information, the smart glasses use a voice recognition model to recognize the voice information so as to acquire detailed control information.
[0076] Optionally, when the smart glasses are interrupted by voice information for too long, the smart glasses can communicate with the cloud server and use the cloud server to recognize the voice information.
[0077] In an optional embodiment, after the coarse control information and the detailed control information are acquired, the coarse control information and the detailed control information are combined to realize control of the target object.
[0078] Optionally, a control instruction for the target object can be generated based on the coarse control information and the detailed control information, and the target object can be controlled according to the control instruction. For example, if the coarse control information is to grasp the target object, and the detailed control information is to reduce the target object to 50% of its original size, then the control instruction for the target object generated based on the coarse control information and the detailed control information can be to reduce the grasped target object to 50% of its original size, and the smart glasses can control the target object according to the control instruction.
[0079] In the above-mentioned virtual space interaction method, hand trajectories are used to determine rough control information, and voice information is used to determine detailed control information. Node control information is implemented based on the rough control information and detailed control information, and the detailed control information is used to supplement the rough control information, thereby lowering the operation threshold, avoiding excessive reliance on gestures, and improving control accuracy.
[0080] In one embodiment, Figure 2 As shown, a virtual space interaction method is provided. This embodiment uses the method applied to smart glasses as an example for illustration. It is understood that the method can also be applied to smart glasses, and can also be applied to a system including smart glasses and a server, and implemented through the interaction between the smart glasses and the server. In this embodiment, the method includes the following steps:
[0081] Step 202 : Acquire a plurality of consecutive image frames within a preset time period, identify the key points of the hand in the consecutive image frames respectively, and obtain the position coordinates of the key points of the hand in the corresponding consecutive image frames.
[0082] Step 204 : Determine the hand trajectory using the position coordinates of the hand key points in the continuous multi-frame image.
[0083] Step 206 : Acquire a preset trajectory gesture type relationship, determine the gesture type corresponding to the hand trajectory based on the trajectory gesture type relationship, and determine the target object and rough control information according to the gesture type.
[0084] Step 208: Acquire initial voice data and perform endpoint detection on the initial voice data to obtain at least one intermediate voice data.
[0085] Step 210: Perform noise reduction processing on at least one intermediate speech data to obtain speech information.
[0086] Step 212: perform word segmentation processing on the voice information to obtain at least one word segmentation data.
[0087] Step 214: Use at least one word segmentation data to query the detail control data table to obtain detail control information.
[0088] Step 216: Control the target object according to the coarse control information and the detailed control information.
[0089] In the above-mentioned virtual space interaction method, hand trajectories are used to determine rough control information, and voice information is used to determine detailed control information. Node control information is implemented based on the rough control information and detailed control information, and the detailed control information is used to supplement the rough control information, thereby lowering the operation threshold, avoiding excessive reliance on gestures, and improving control accuracy.
[0090] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0091] Based on the same inventive concept, embodiments of the present application also provide a virtual space interaction device for implementing the aforementioned virtual space interaction method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more virtual space interaction device embodiments provided below can be found in the aforementioned limitations of the virtual space interaction method and will not be further elaborated here.
[0092] In an exemplary embodiment, Figure 3 As shown, a virtual space interaction device 300 is provided, comprising: a first information determination module 302, a second information determination module 304 and a control module 306, wherein:
[0093] a first information determination module 302 for acquiring a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0094] The second information determination module 304 is used to obtain voice information and determine detailed control information based on the voice information;
[0095] The control module 306 is configured to control the target object according to the coarse control information and the detailed control information.
[0096] In one embodiment, the first information determination module is also used to obtain multiple consecutive frames of images within a preset time period, respectively identify the hand key points in the consecutive multiple frames of images, and obtain the position coordinates of the hand key points in the corresponding consecutive multiple frames of images; determine the hand trajectory using the position coordinates of the hand key points in the consecutive multiple frames of images; obtain a preset trajectory gesture type relationship, and determine the gesture type corresponding to the hand trajectory based on the trajectory gesture type relationship.
[0097] In one embodiment, the second information determination module is further used to obtain initial voice data, and perform endpoint detection on the initial voice data to obtain at least one intermediate voice data; and perform noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0098] In one embodiment, the second information determination module is further configured to perform word segmentation processing on the voice information to obtain at least one word segmentation data; and use the at least one word segmentation data to query the detail control data table to obtain the detail control information.
[0099] In one embodiment, the virtual space interaction device further includes a security verification module for obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, the step of controlling the target object according to the rough control information and the detailed control information is executed.
[0100] In one embodiment, the virtual space interaction device also includes a security verification module, which is also used to obtain an eyeball image and extract the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, the step of controlling the target object according to the coarse control information and the detailed control information is executed.
[0101] Each module in the aforementioned virtual space interaction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor within a computer device in the form of hardware, or can be stored in a computer device's memory in the form of software, allowing the processor to call and execute the corresponding operations of each module.
[0102] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 4As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a virtual space interaction method. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0103] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0104] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0105] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0106] Acquire voice information and determine detailed control information based on the voice information;
[0107] The target object is controlled based on the coarse control information and the detailed control information.
[0108] In one embodiment, when the processor executes the computer program, it also implements the following steps: obtaining the hand trajectory in the virtual space, and determining the gesture type based on the hand trajectory, including: obtaining multiple consecutive frames of images within a preset time period, respectively identifying the hand key points in the consecutive multiple frames of images, and obtaining the position coordinates of the hand key points in the corresponding consecutive multiple frames of images; using the position coordinates of the hand key points in the consecutive multiple frames of images to determine the hand trajectory; obtaining a preset trajectory-gesture type relationship, and determining the gesture type corresponding to the hand trajectory based on the trajectory-gesture type relationship.
[0109] In one embodiment, when the processor executes the computer program, it also implements the following steps: obtaining voice information, including: obtaining initial voice data and performing endpoint detection on the initial voice data to obtain at least one intermediate voice data; performing noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0110] In one embodiment, when the processor executes the computer program, the processor further implements the following steps: determining detailed control information based on the voice information, including: performing word segmentation processing on the voice information to obtain at least one word segmentation data; and using the at least one word segmentation data to query the detailed control data table to obtain the detailed control information. In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0111] In one embodiment, when the processor executes the computer program, the following steps are also implemented: before controlling the target object based on the coarse control information and the detailed control information, the method also includes: obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, executing the step of controlling the target object based on the coarse control information and the detailed control information.
[0112] In one embodiment, when the processor executes the computer program, it also implements the following steps: acquiring an eyeball image and extracting the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, executing the step of controlling the target object according to the coarse control information and the detailed control information.
[0113] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0114] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0115] Acquire voice information and determine detailed control information based on the voice information;
[0116] The target object is controlled based on the coarse control information and the detailed control information.
[0117] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: obtaining a hand trajectory in a virtual space, and determining a gesture type based on the hand trajectory, including: obtaining multiple consecutive frames of images within a preset time period, respectively identifying hand key points in the consecutive multiple frames of images, and obtaining the position coordinates of the hand key points in the corresponding consecutive multiple frames of images; determining the hand trajectory using the position coordinates of the hand key points in the consecutive multiple frames of images; obtaining a preset trajectory-gesture type relationship, and determining the gesture type corresponding to the hand trajectory based on the trajectory-gesture type relationship.
[0118] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: obtaining voice information, including: obtaining initial voice data and performing endpoint detection on the initial voice data to obtain at least one intermediate voice data; performing noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0119] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining detailed control information based on voice information, including: performing word segmentation processing on the voice information to obtain at least one word segmentation data; and querying a detailed control data table using the at least one word segmentation data to obtain the detailed control information. In one embodiment, when the computer program is executed by the processor, the following steps are also implemented:
[0120] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: before controlling the target object based on the coarse control information and the detailed control information, the method also includes: obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, executing the step of controlling the target object based on the coarse control information and the detailed control information.
[0121] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: acquiring an eyeball image and extracting the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, executing the step of controlling the target object according to the coarse control information and the detailed control information.
[0122] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:
[0123] Obtaining a hand trajectory in a virtual space, determining a gesture type and a target object based on the hand trajectory, and determining rough control information based on the gesture type;
[0124] Acquire voice information and determine detailed control information based on the voice information;
[0125] The target object is controlled based on the coarse control information and the detailed control information.
[0126] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: obtaining a hand trajectory in a virtual space, and determining a gesture type based on the hand trajectory, including: obtaining multiple consecutive frames of images within a preset time period, respectively identifying hand key points in the consecutive multiple frames of images, and obtaining the position coordinates of the hand key points in the corresponding consecutive multiple frames of images; determining the hand trajectory using the position coordinates of the hand key points in the consecutive multiple frames of images; obtaining a preset trajectory-gesture type relationship, and determining the gesture type corresponding to the hand trajectory based on the trajectory-gesture type relationship.
[0127] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: obtaining voice information, including: obtaining initial voice data and performing endpoint detection on the initial voice data to obtain at least one intermediate voice data; performing noise reduction processing on the at least one intermediate voice data to obtain voice information.
[0128] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining detailed control information based on voice information, including: performing word segmentation processing on the voice information to obtain at least one word segmentation data; and querying a detailed control data table using the at least one word segmentation data to obtain the detailed control information. In one embodiment, when the computer program is executed by the processor, the following steps are also implemented:
[0129] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: before controlling the target object based on the coarse control information and the detailed control information, the method also includes: obtaining voiceprint information in the voice information; if the voiceprint information matches the preset voiceprint information, executing the step of controlling the target object based on the coarse control information and the detailed control information.
[0130] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: acquiring an eyeball image and extracting the pupil center coordinates in the eyeball image; when the pupil center coordinates match the center coordinates of the target object, executing the step of controlling the target object according to the coarse control information and the detailed control information.
[0131] The following steps are implemented when the computer program is executed by the processor. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0132] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0133] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0134] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A virtual space interaction method, characterized in that: Applied to smart glasses, the method includes: Acquiring a hand trajectory in a virtual space, determining a gesture type and a target object according to the hand trajectory, and determining rough control information according to the gesture type; Acquiring voice information, and determining detailed control information based on the voice information; The target object is controlled according to the coarse control information and the detailed control information.
2. The method according to claim 1, characterized in that The obtaining of the hand trajectory in the virtual space and determining the gesture type according to the hand trajectory includes: Acquire a plurality of consecutive image frames within a preset time period, identify hand key points in the consecutive image frames respectively, and obtain position coordinates of the hand key points in the corresponding consecutive image frames; Determining a hand trajectory using the position coordinates of the hand key points in the continuous multiple frames of images; A preset trajectory gesture type relationship is acquired, and the gesture type corresponding to the hand trajectory is determined based on the trajectory gesture type relationship.
3. The method according to claim 1, characterized in that The acquiring of voice information includes: Acquire initial voice data, and perform endpoint detection on the initial voice data to obtain at least one intermediate voice data; Noise reduction processing is performed on the at least one intermediate speech data to obtain the speech information.
4. The method according to claim 1 or 3, characterized in that Determining detailed control information according to the voice information includes: Performing word segmentation processing on the voice information to obtain at least one word segmentation data; The detail control data table is searched using the at least one word segmentation data to obtain the detail control information.
5. The method according to claim 1, wherein Before controlling the target object according to the coarse control information and the detailed control information, the method further includes: Obtaining voiceprint information from the voice information; If the voiceprint information matches the preset voiceprint information, a step of controlling the target object according to the rough control information and the detailed control information is executed.
6. The method according to claim 1, characterized in that Before determining the target object and the rough control information according to the gesture type, the method further includes: Acquire an eyeball image, and extract pupil center coordinates from the eyeball image; When the pupil center coordinates match the center coordinates of the target object, the step of controlling the target object according to the coarse control information and the detailed control information is performed.
7. A virtual space interactive device, characterized in that: The device comprises: a first information determination module, configured to obtain a hand trajectory in a virtual space, determine a gesture type based on the hand trajectory, and determine a target object and rough control information based on the gesture type; a second information determination module, configured to obtain voice information and determine detail control information based on the voice information; A control module is configured to control the target object according to the rough control information and the detailed control information.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.