Virtual character full-body driving method and virtual reality device

CN118605716BActive Publication Date: 2026-08-11HISENSE VISUAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]目前,虚拟现实(Virtual Reality,VR)产品能够通过手柄和头戴式显示器(HeadMounted Displays,HMD),对手部和头部进行由内向外的追踪,实现手部和头部位姿的获取,但无法获取腿部、脚部或臀部等其他部位的位姿,这样,会导致虚拟人物(Avatar)缺失下半身,影响用户体验

Benefits of technology

[0037]本申请实施例提供的一种虚拟人物全身驱动方法及设备中,结合同一局域网中带有RGB相机的智能设备,实现目标用户全身动作的捕捉,由于智能设备成本低、普及范围广,因此可以满足大多数用户的使用需求。其中,游戏过程中,智能设备根据RGB相机采集的目标用户当前的人体RGB图像,获得人体各关键点的第一3D位姿,并传输给虚拟现实设备。虚拟现实设备结合手柄和头戴式显示器获得的手部和头部的核心关键点的第二3D位姿,确定全部关键节点精确的目标3D位姿,并根据全部关键点的目标3D位姿,对虚拟人物进行全身驱动,获得与当前游戏画面匹配的完整虚拟人物,从而解决由于手部和头部被遮挡导致上半身位姿缺失,以及手柄和头戴式显示器无法采集到下半身位姿的问题,获得完整的虚拟人物,同时,通过对两设备获得的第一3D位姿和第二3D位姿,获得全部关键点的目标3D位姿,可以提高关键点位姿的准确性,进而提高虚拟人物动作的准确性,有助于提升用户的游戏体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118605716B_ABST
    Figure CN118605716B_ABST
Patent Text Reader

Abstract

This application relates to the field of virtual reality technology, providing a method for driving the entire body of a virtual character and a virtual reality device. By utilizing smart devices within the same local area network, the method captures the full-body movements of a target user. In this method, the 3D poses of local key points acquired by the controllers and head-mounted display are fused with the 3D poses of the same key points acquired by the smart device, improving the accuracy of local pose calculation. Based on the accurate 3D poses of each key point, the 3D poses of all human body key points acquired by the smart device are updated, further improving the accuracy of global pose calculation. This allows for high-precision driving of the entire body of a virtual character without wearing a motion tracker. Furthermore, due to the low cost and wide availability of smart devices, this method can meet the needs of most users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual reality technology, and provides a method for driving the whole body of a virtual character and a virtual reality device. Background Technology

[0002] Currently, Virtual Reality (VR) products can track the hands and head from the inside out using controllers and head-mounted displays (HMDs) to capture the poses of the hands and head, but they cannot capture the poses of other body parts such as legs, feet, or hips. This results in the virtual avatar lacking a lower body, affecting the user experience.

[0003] In order to capture the full-body movements of virtual characters and obtain fully driven virtual characters, users usually wear motion rings, which can cause inconvenience to the user experience. Summary of the Invention

[0004] This application provides a method for driving the whole body of a virtual character and a virtual reality device for achieving full-body driving of a virtual character.

[0005] On one hand, embodiments of this application provide a method for driving the full body of a virtual character, applied to a virtual reality device, wherein the virtual reality device and a smart device containing an RGB camera are on the same local area network, and the method includes:

[0006] Establish a communication connection with the smart device and receive the first 3D pose of each key point in the human body key point set sent by the smart device. Each first 3D pose is obtained by the smart device based on the human body RGB image of the target user captured by the RGB camera when the target user performs actions according to the current game screen.

[0007] Based on the controller and head-mounted display included in the virtual display device, the second 3D pose of the core key points of the hand and head in the current action of the target user is obtained;

[0008] Based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point, the target 3D pose of all key points is obtained.

[0009] Based on the target 3D pose of all the key points, the virtual character corresponding to the target user is driven in its entirety to obtain a virtual character that matches the current game screen.

[0010] On the other hand, embodiments of this application provide a virtual reality device, including a controller and a head-mounted display, wherein the head-mounted display includes a processor, a memory, and a display screen, and the display screen, the memory, and the processor are connected via a bus;

[0011] The memory stores a computer program, and the processor performs the following operations according to the computer program:

[0012] Establish a communication connection with a smart device within the same local area network. The smart device includes an RGB camera. The smart device is used to obtain the first 3D pose of each key point in the human body key point set from the human body RGB image of the target user performing actions according to the current game screen captured by the RGB camera.

[0013] Receive the first 3D pose of each key point in the human body key point set sent by the smart device;

[0014] Based on the handle and the head-mounted display, obtain the second 3D pose of the core key points of the hand and head in the current action of the target user;

[0015] Based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point, the target 3D pose of all key points is obtained.

[0016] Based on the target 3D pose of all the key points, the virtual character corresponding to the target user is driven in its entirety to obtain a virtual character that matches the current game screen, and then displayed on the display screen.

[0017] Optionally, the processor obtains the target 3D pose of all key points based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point. Specifically, the operation is as follows:

[0018] Based on the key point labels, determine whether each core key point is concentrated in the human body key point cluster;

[0019] For any core key point that is not in the set of human body key points, the second 3D pose of the core key point is directly used as the target 3D pose of the core key point.

[0020] For any core key point in the human body key point set, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

[0021] Based on the target 3D pose of each core key point, the first 3D pose of each key point in the human body key point set is updated to obtain the target 3D pose of all key points.

[0022] Optionally, the processor fuses the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point. Specifically, the operation is as follows:

[0023] Calculate the pose variance between the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set.

[0024] Based on the pose variance, the fusion coefficient between the second 3D pose and the first 3D pose is obtained;

[0025] Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

[0026] Optionally, the formula for calculating the fusion coefficient is:

[0027]

[0028] in, The pose variance is represented by Var(), where Var() represents the variance function. k A represents the second 3D pose of the core key point at the current moment. k This represents the first 3D pose of the corresponding key points in the current human body keypoint set, where ω represents the fusion coefficient. The second variance represents the second 3D pose of the core key points. The first variance represents the first 3D pose of the corresponding key points in the human body key point set.

[0029] Optionally, the processor updates the first 3D pose of each key point in the human body key point set based on the target 3D pose of each core key point, thereby obtaining the target 3D pose of all key points. Specifically, the operation is as follows:

[0030] For each core critical point, perform the following operations:

[0031] Calculate the target 3D pose of the core key point and the pose difference between the target 3D pose of the corresponding key point in the human body key point set;

[0032] The target 3D pose of the partial key points is obtained by adding the pose difference to the first 3D pose of the partial key points, which are the human body key points and are associated with the core key points.

[0033] Optionally, the process of establishing a communication connection between the processor and the smart device is as follows:

[0034] Create a ServerSocket and listen for connection requests sent by the smart device after it creates a Socket;

[0035] Based on the connection request, a communication connection is established with the smart device, and a connection success message is sent to the smart device.

[0036] On the other hand, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for causing a computer device to perform the steps of the virtual character full-body driving method provided in embodiments of this application.

[0037] This application provides a method and device for full-body virtual character driving, which combines a smart device with an RGB camera on the same local area network to capture the full-body movements of the target user. Because smart devices are low-cost and widely available, they can meet the needs of most users. During gameplay, the smart device obtains the first 3D pose of each key point of the human body based on the RGB image of the target user captured by the RGB camera and transmits it to the virtual reality device. The virtual reality device, combined with the second 3D pose of the core key points of the hands and head obtained from the controller and head-mounted display, determines the precise target 3D pose of all key nodes. Based on the target 3D pose of all key points, the virtual character is driven in its entirety to obtain a complete virtual character matching the current game screen. This solves the problems of missing upper body pose due to occlusion of the hands and head, and the inability of the controller and head-mounted display to capture the lower body pose, resulting in a complete virtual character. Furthermore, by combining the first and second 3D poses obtained from the two devices to obtain the target 3D pose of all key points, the accuracy of the key point poses can be improved, thereby improving the accuracy of the virtual character's movements and enhancing the user's gaming experience.

[0038] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 A schematic diagram of a half-body virtual character provided in an embodiment of this application;

[0041] Figure 2 A schematic diagram illustrating the wearing of a sports ring on a user's leg, as provided in an embodiment of this application;

[0042] Figure 3 A schematic diagram illustrating the process of full-body motion capture using the related technologies provided in the embodiments of this application;

[0043] Figure 4 This is an overall architecture diagram of the virtual character full-body driving system provided in the embodiments of this application;

[0044] Figure 5 A flowchart illustrating the method for driving the full body of a virtual character as provided in this application embodiment;

[0045] Figure 6 This is a schematic diagram illustrating the process of establishing Socket communication as provided in an embodiment of this application.

[0046] Figure 7 A schematic diagram illustrating the process of extracting 3D poses of multiple key human body points for intelligent devices;

[0047] Figure 8 A schematic diagram of the human body key point set provided in the embodiments of this application;

[0048] Figure 9 A 3D pose flowchart for determining local core key points provided in the embodiments of this application;

[0049] Figure 10 This is a flowchart of the pose fusion method provided in the embodiments of this application;

[0050] Figure 11 A schematic diagram of pose variance before and after fusion provided for embodiments of this application;

[0051] Figure 12 A flowchart of the global pose update method provided in the embodiments of this application;

[0052] Figure 13 The full-body motion capture effect diagram provided in the embodiments of this application;

[0053] Figure 14A schematic diagram of a complete virtual character provided for an embodiment of this application;

[0054] Figure 15 This application provides an example of an interaction flowchart between a virtual reality device and a smart device.

[0055] Figure 16 This is a structural diagram of a virtual reality device provided in an embodiment of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0057] Currently, VR products can capture the poses of hands and head through controllers and HMDs, but cannot capture the poses of other body parts such as legs, feet, or buttocks. This results in virtual characters lacking the lower body, such as... Figure 1 As shown.

[0058] To capture the full-body movements of a virtual character and obtain a fully driven virtual character, a motion ring is usually worn on the user's lower body to collect the posture of the lower body.

[0059] For example, such as Figure 2 As shown, a motion ring is worn on the user's legs. The sensor built into the motion ring captures the user's lower body pose. Combined with the upper body pose captured by the controller and HMD, the motion of the whole body is captured in 3 degrees of freedom (DOF) to obtain a fully driven virtual character.

[0060] Considering the inconvenience caused to users wearing motion trackers, this technology employs deep learning algorithms to infer full-body pose data based on the controller and HMD, thereby obtaining a fully driven virtual character, such as... Figure 3 As shown, however, over time, misalignment from the previous moment may affect the accuracy of the pose in the next moment. There are also related technologies that use the Kinect camera to capture the pose of the user's lower body; however, the Kinect camera is expensive and has limited usage, as not every user owns one.

[0061] Therefore, this application provides a method for driving the entire body of a virtual character and a virtual reality device. The virtual reality device communicates with smart devices (such as mobile phones, televisions, etc.) on the same local area network, fusing the 3D poses of the core key points of the hands and head captured by its own controller and head-mounted display with the 3D poses of the corresponding key points captured by the RGB camera of the smart device. Based on the fusion result, the 3D poses of the remaining key points captured by the RGB camera of the smart device are updated to obtain accurate 3D poses of the entire body's key points, thereby achieving full-body driving of the virtual character. The entire process does not require the target user to wear a motion tracker, making it user-friendly. Furthermore, the low cost and wide availability of smart devices allow it to meet the needs of most users for a complete virtual character. By fusing and updating the 3D poses of key points, the accuracy of the 3D poses of key points is improved, thereby improving the accuracy of the virtual character's movements and enhancing the user's gaming experience.

[0062] See Figure 4 This is an overall architecture diagram of the virtual character full-body driving system provided in this application embodiment. First, a Socket communication connection is established between the smart device and the virtual reality device within the same local area network. Then, the RGB camera on the smart terminal captures the RGB image of the target user's human body during the game. The pre-trained keypoint extraction model deployed in the application launched on the smart device extracts the 3D pose of the human body's key points from the RGB image based on the model and transmits it to the virtual reality device via the network. Simultaneously, during the game, the controller and head-mounted display on the virtual display device also capture the 3D pose of the target user's hand and head key points. Combined with the 3D pose of the human body key points transmitted by the smart device, 3D pose fusion and updating operations are performed to obtain the accurate target 3D pose of the full-body key points. Based on the target 3D pose of the full-body key points, the virtual character's entire body is driven to move, resulting in a complete virtual character that matches the actions in the game screen.

[0063] based on Figure 4 The system architecture diagram shown is as follows: Figure 5 This embodiment of the present application provides a virtual character full-body driving implementation process, which is executed by a virtual reality device and mainly includes the following steps:

[0064] S501: Establishes a communication connection with smart devices.

[0065] In one implementation, the virtual reality device establishes a Socket communication connection with the smart device.

[0066] In specific implementation, such as Figure 6The diagram illustrates the Socket communication process between a virtual reality (VR) device and a smart device. The VR device acts as the server, creating a ServerSocket, while the smart device acts as the client, creating a Socket. The VR device listens for connection requests from the smart device using the Accept() function, establishes a Socket communication connection based on the request, and sends a connection success message to the smart device. After the communication connection is established, the VR device and the smart device transmit data using input streams (InputStream) and output streams (OutputStream). Once the transmission is complete, the Socket communication connection is closed, and communication resources are shut down.

[0067] S502: Receives the first 3D pose of each key point in the human body key point set sent by the smart device.

[0068] In the embodiments of this application, the smart device that communicates with the virtual reality device via Socket is equipped with a key point extraction model pre-trained based on a deep learning algorithm. Through this model, motion analysis can be performed on the RGB images of the human body captured by the RGB camera of the smart device to obtain the 3D pose of multiple human body key points.

[0069] In one example, the keypoint extraction model can be the MHPE (Modeling Keypoints and Poses) model. By annotating the acquired RGB images of the human body, a sample set is obtained to train the MHPE model. After training, the MHPE model can predict human pose information from monocular images. By deploying the trained MHPE model to an application on a smart device (e.g., an Android APK application), the smart terminal with the deployed MHPE model installed can communicate with the virtual reality device via a socket, transmitting the human pose information calculated by the MHPE model to the virtual reality device.

[0070] Taking a game scenario as an example, the target user wearing the virtual reality device stands within the field of view of the RGB camera of the smart device on the same local area network, so that the RGB camera can collect complete human body data of the target user. When the smart device receives a connection success message from the virtual reality device, it activates the RGB camera to collect the RGB image of the human body as the target user performs actions according to the current game screen. It then performs motion analysis on the current human body RGB image using a deployed model to obtain the first 3D pose of each key point in the human body keypoint set, and sends the first 3D pose of each keypoint to the virtual reality device.

[0071] See Figure 7This diagram illustrates the process of extracting 3D poses of multiple human body key points for intelligent devices. The Pose Encoder network extracts features from the RGB images of the human body captured by the RGB camera, and the Pose Decoder network decodes the extracted features to obtain the 2D coordinates of multiple human body key points. Through the conversion relationship between 2D and 3D, the 3D coordinates of multiple human body key points representing the human body pose are obtained.

[0072] like Figure 8 The diagram shown is a schematic diagram of the human body key point set provided in the embodiment of this application. The set contains 33 key points, each of which is uniquely represented by a label. The connection of the key points can represent the skeleton of the virtual character.

[0073] S503: Based on the controller and head-mounted display included in the virtual display device, obtain the second 3D pose of the core key points of the target user's hands and head in the current action.

[0074] Generally, virtual display devices have integrated pose sensors in their controllers and head-mounted displays. Based on these pose sensors, the poses of the controllers and head-mounted displays can be obtained. By subtracting fixed errors, the second 3D pose of the core key points of the target user's hands and head in the current action can be obtained.

[0075] The fixed errors for the hands and head can be set according to the actual measured data, and the fixed errors for the hands and head can be the same or different.

[0076] In one example, the core key points of the head are: Figure 8 The key point marked with a central symbol of 0, the core key point of the left hand is Figure 8 The key point numbered 15, the core key point on the right side is Figure 8 Key point number 16.

[0077] S504: Obtain the target 3D pose of all key points based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point.

[0078] In practical applications, the target user's environment is often complex, and they are constantly moving during gameplay. This may prevent the smart device's RGB camera from capturing the target user's RGB image. Therefore, based on the data captured by the RGB cameras, the 3D poses of key points from both devices can be fused and updated. For specific implementation details, please refer to [link to implementation details]. Figure 9 It mainly includes the following steps:

[0079] S5041: Based on the key point number, determine whether each core key point is concentrated in the human body key point cluster. If not, execute S5042; if yes, execute S5043.

[0080] according to Figure 8 It is known that each human body key point corresponds to a unique label. By using the label, it can be determined whether there are any identical key points obtained by the smart device and the key points obtained by the virtual reality device.

[0081] S5042: For any core key point that is not in the human body key point set, directly use the second 3D pose of the core key point as the target 3D pose of the core key point.

[0082] When the RGB camera is fully or partially occluded, the smart device cannot obtain the first 3D pose of the target user's head and / or hand key points; only the virtual reality device obtains the second 3D pose of the head and hand key points. That is, all or some of the core key points obtained by the virtual reality device are not in the set of human key points obtained by the smart device. For any core key point not in the set of human key points, the second 3D pose of that core key point obtained by the virtual reality device can be directly used as the target 3D pose of that core key point.

[0083] S5043: For any core key point in the human body key point set, fuse the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point.

[0084] When the RGB camera is not obstructed, both smart devices and virtual reality devices can obtain the 3D poses of key points of the target user's head and hands. Since the 3D poses of key points obtained by these two devices have certain errors, the optimal estimation method can be used to fuse the 3D poses of the core key points of the hands and head obtained simultaneously to improve the accuracy of the pose information.

[0085] For example, see a key point. Figure 10 This is a flowchart of the 3D pose fusion process for key points simultaneously acquired by virtual reality devices and smart devices. The process mainly includes the following steps:

[0086] S5043_1: Calculate the second 3D pose of the core key points and the pose variance between it and the first 3D pose of the corresponding key points in the human body key point set.

[0087] like Figure 11 As shown, assuming the second 3D pose of the core key points captured by the virtual reality device at time k is denoted as V. k The second variance of the second 3D pose of this key point is denoted as... The first 3D pose of the corresponding key points in the human body key point set obtained by the intelligent device is denoted as A. k The first variance of the first 3D pose of the corresponding key point is denoted as . Fusion pose variance for:

[0088]

[0089] Where Var() represents the variance function, ω represents the fusion coefficient, and ω∈[0,1]. When ω=0, it means that the second 3D pose acquired by the handheld device or head-mounted display is completely trusted; when ω=1, it means that the first 3D pose acquired by the smart device is completely trusted.

[0090] S5043_2: Obtain the fusion coefficient between the second 3D pose and the first 3D pose based on the pose variance.

[0091] When the above pose variance When the value is minimized, it represents the optimal pose estimation. At this point, the fusion coefficients can be solved.

[0092] S5043_3: Based on the fusion coefficient, the second 3D pose of the core key points and the first 3D pose of the corresponding key points in the human body key point set are fused to obtain the target 3D pose of the core key points.

[0093] Fusion target 3D pose P k The formula is expressed as follows:

[0094] P k =V k +ω(A k -V k )

[0095] That is when P obtained k Optimal.

[0096] S5044: Based on the target 3D pose of each core key point, update the first 3D pose of each key point in the human body key point set to obtain the target 3D pose of all key points.

[0097] The target 3D pose of each core key point after fusion is more accurate than the second 3D pose obtained by the virtual reality device and the first 3D pose obtained by the smart device. Since the human skeleton is represented by multiple key points and the relative poses between multiple key points, the target 3D pose of each core key point can be used to update the first 3D pose of the other key points to obtain the target 3D pose of the other key points, thereby improving the accuracy of the entire human motion capture.

[0098] In the embodiments of this application, during the process of updating the first 3D pose of each skeletal node, the relative distance between keypoints (i.e., bone length) remains unchanged. Taking a core keypoint as an example, the update process of the first 3D pose of the keypoint is as follows: Figure 12 As shown, it mainly includes the following steps:

[0099] S5044_1: Calculate the target 3D pose of the core key point and the pose difference between it and the first 3D pose of the corresponding key point in the human body key point set.

[0100] When the core keypoint is the head keypoint with label 0, calculate the pose difference between the target 3D pose P0 of the keypoint with label 0 and the first 3D pose A0; when the core keypoint is the left-hand keypoint with label 15, calculate the target 3D pose P of the keypoint with label 15. 15 With the first 3D pose A 15 The pose difference between the key points; when the core key point is the right-hand key point numbered 16, calculate the target 3D pose P of the key point numbered 16. 16 With the first 3D pose A 16 The pose difference between them.

[0101] S5044_2: The first 3D pose of the partial key points of the human body is obtained by concentrating the key points of the human body with the core key points and adding the pose difference.

[0102] by Figure 8 For example, when the core keypoint is the keypoint with head label number 0, its associated keypoints are keypoints numbered 0-10. In this case, the update formula for the first 3D pose of keypoints numbered 0-10 is:

[0103] P n '=P n +(P0-A0), n=0, 1, 2,...,10.

[0104] When the core keypoint is the left-hand keypoint numbered 15, its associated keypoints are keypoints numbered 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, and 31. In this case, the update formula for the first 3D pose of these keypoints is:

[0105] P n '=P n +(P 15 -A 15 ), n=11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31.

[0106] When the core keypoint is keypoint number 16 on the right-hand side, its associated keypoints are keypoints numbered 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, and 32. In this case, the update formula for the first 3D pose of these keypoints is:

[0107] P n '=P n +(P 16 -A 16 ), n=12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32.

[0108] After updating the first 3D pose of all key points numbered 0-33, a high-precision 3D pose of all key points of the whole body can be obtained.

[0109] S505: Based on the target 3D pose of all key points, perform full-body driving on the virtual character corresponding to the target user to obtain a virtual character that matches the current game screen.

[0110] In one example, the system iterates through 33 key points of the human body, calculates the target 3D pose of all key points, converts the 33 target poses into coordinates in a sphere coordinate system, and enables full-body driving of the virtual character corresponding to the target user, thereby obtaining a virtual character that matches the current game screen.

[0111] like Figure 13 The image shown is an effect diagram of full-body motion capture provided in the embodiment of this application. The upper left corner is the current game screen, the upper right corner is the upper body key points obtained based on the controller and head-mounted display of the virtual reality device, and the middle part is the complete human skeleton obtained after combining the facial key points collected by the RGB camera of the smart device.

[0112] It should be noted that, in Figure 13 The software can also display the spherical coordinates of each key point, as well as the controls for each key point, so that graphic designers can make later edits.

[0113] Using the aforementioned full-body driving method for virtual characters, a complete character model including the lower body can be obtained, such as... Figure 14 The image shown is an effect diagram of the full-body driving provided in the embodiment of this application. Full-body driving can improve the realism of the character model, thereby enhancing the user's immersive experience.

[0114] See Figure 15 This is a flowchart illustrating the interaction between a virtual reality device and a smart device provided in an embodiment of this application. The process mainly includes the following steps:

[0115] S1501: Start the VR game, the virtual reality device creates a ServeSocket.

[0116] S1502: Smart device creates Socket connection.

[0117] S1503: The smart device sends a connection request to the virtual reality device.

[0118] S1504: After receiving a connection request, the virtual reality device establishes a communication connection with the smart device and sends a connection success message to the smart device.

[0119] S1505: The smart device turns on the RGB camera and launches an application with a key point extraction model deployed on it.

[0120] S1506: The smart device extracts the first 3D pose of multiple human body key points from the RGB image of the human body captured by the RGB camera, and sends the first 3D pose of multiple human body key points to the virtual reality device. At the same time, the controller and head-mounted display of the virtual reality device capture the second 3D pose of the core key points of the hands and head.

[0121] S1507: The virtual reality device fuses the second 3D pose of each core key point with the first 3D pose of the same key point collected by the smart device to determine the target 3D pose of each core key point.

[0122] S 1 508: The virtual reality device updates the first 3D pose of multiple human body key points based on the target 3D pose of each core key point, and obtains the target 3D pose of all key points.

[0123] S1509: Perform coordinate transformation based on the target 3D pose of all key points.

[0124] S1510: Drives the virtual character based on the converted 3D coordinates.

[0125] S1511: After the game ends, the virtual reality device sends a disconnection message to the smart device.

[0126] S1512: After receiving the disconnection message, the smart device turns off the RGB camera and exits the application.

[0127] The virtual character full-body driving method provided in this application embodiment, combined with a smart device equipped with an RGB camera on the same local area network, achieves the capture of the target user's full-body movements. Because smart devices are low-cost and widely available, they can meet the needs of most users. During gameplay, the smart device obtains the first 3D pose of each key point of the human body based on the RGB image of the target user captured by the RGB camera, and transmits it to the virtual reality device. The virtual reality device combines the second 3D pose of the core key points of the hands and head obtained by the controller and head-mounted display with the first 3D coordinates of the same key points captured by the smart device, improving the accuracy of each core key point. Based on the fused target 3D pose, the first 3D pose of multiple human body key points captured by the smart device is updated, further improving the accuracy of all key points. Thus, based on the high-precision target 3D pose of multiple human body key points, the virtual character is driven in its entirety, obtaining a complete virtual character that matches the current game screen. This solves the problem that controllers and head-mounted displays cannot capture the lower body pose, resulting in a complete virtual character.

[0128] Based on the same technical concept, this application provides a virtual reality device that can implement the steps of the virtual character full-body driving method provided in the above embodiments.

[0129] See Figure 16 The virtual reality device includes a controller 1601 and a head-mounted display 1602. The head-mounted display 1602 includes a processor 1602_1, a memory 1602_2, and a display screen 1602_3. The display screen 1602_3, the memory 1602_2, and the processor 1602_1 are connected via a bus 1602_4.

[0130] The memory 1602_2 stores a computer program, and the processor 1602_1 performs the following operations according to the computer program:

[0131] Establish a communication connection with a smart device within the same local area network. The smart device includes an RGB camera. The smart device is used to obtain the first 3D pose of each key point in the human body key point set from the human body RGB image of the target user performing actions according to the current game screen captured by the RGB camera.

[0132] Receive the first 3D pose of each key point in the human body key point set sent by the smart device;

[0133] Based on the handle 1601 and the head-mounted display 1602, the second 3D pose of the core key points of the hand and head in the current action of the target user is obtained;

[0134] Based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point, the target 3D pose of all key points is obtained.

[0135] Based on the target 3D pose of all the key points, the virtual character corresponding to the target user is driven in its entirety to obtain a virtual character that matches the current game screen, and then displayed on the display screen 1602_3.

[0136] Optionally, the processor 1602_1 obtains the target 3D pose of all key points based on the first 3D pose of each key point in the human body key point set and the second 3D pose of each core key point. Specifically, the operation is as follows:

[0137] Based on the key point labels, determine whether each core key point is concentrated in the human body key point cluster;

[0138] For any core key point that is not in the set of human body key points, the second 3D pose of the core key point is directly used as the target 3D pose of the core key point.

[0139] For any core key point in the human body key point set, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

[0140] Based on the target 3D pose of each core key point, the first 3D pose of each key point in the human body key point set is updated to obtain the target 3D pose of all key points.

[0141] Optionally, the processor 1602_1 fuses the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point. Specifically, the operation is as follows:

[0142] Calculate the pose variance between the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set.

[0143] Based on the pose variance, the fusion coefficient between the second 3D pose and the first 3D pose is obtained;

[0144] Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

[0145] Optionally, the formula for calculating the fusion coefficient is:

[0146]

[0147] in, The pose variance is represented by Var(), where Var() represents the variance function. k A represents the second 3D pose of the core key point at the current moment. k This represents the first 3D pose of the corresponding key points in the current human body keypoint set, where ω represents the fusion coefficient. The second variance represents the second 3D pose of the core key points. The first variance represents the first 3D pose of the corresponding key points in the human body key point set.

[0148] Optionally, the processor 1602_1 updates the first 3D pose of each key point in the human body key point set according to the target 3D pose of each core key point, to obtain the target 3D pose of all key points. The specific operation is as follows:

[0149] For each core critical point, perform the following operations:

[0150] Calculate the target 3D pose of the core key point and the pose difference between the target 3D pose of the corresponding key point in the human body key point set;

[0151] The target 3D pose of the partial key points is obtained by adding the pose difference to the first 3D pose of the partial key points, which are the human body key points and are associated with the core key points.

[0152] Optionally, the process of establishing a communication connection between the processor 1602_1 and the smart device is as follows:

[0153] Create a ServerSocket and listen for connection requests sent by the smart device after it creates a Socket;

[0154] Based on the connection request, a communication connection is established with the smart device, and a connection success message is sent to the smart device.

[0155] It should be noted that, Figure 16 This is merely an example illustrating the hardware necessary for a virtual reality device to perform the full-body driving method steps for a virtual character provided in the embodiments of this application. Not shown, the virtual reality device may also include hardware from conventional VR products such as speakers, microphones, communication interfaces, power supplies, and left and right eyeglasses.

[0156] Examples of this application Figure 16The processor involved can be a central processing unit (CPU), a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0157] This application also provides a computer-readable storage medium for storing instructions that, when executed, can complete the virtual character full-body driving method described in the foregoing embodiments.

[0158] This application also provides a computer program product for storing a computer program used to execute the virtual character full-body driving method in the foregoing embodiments.

[0159] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0160] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0163] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for driving the entire body of a virtual character, characterized in that, Applied to virtual reality devices, wherein the virtual reality device and a smart device containing an RGB camera are on the same local area network, the method includes: Establish a communication connection with the smart device and receive the first 3D pose of each key point in the human body key point set sent by the smart device. Each first 3D pose is obtained by the smart device based on the human body RGB image of the target user captured by the RGB camera when the target user performs actions according to the current game screen. Based on the controllers and head-mounted display included in the virtual reality device, the second 3D pose of each core key point of the target user's hand and head in the current action is obtained; For each core key point that exists in the human body key point set, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point. For each core key point that does not exist in the human body key point set, the second 3D pose of the core key point is directly used as the target 3D pose of the core key point. With the bone length remaining constant, the pose difference between the target 3D pose of each core key point and the first 3D pose of the corresponding key point in the human body key point set is propagated to the key point associated with each core key point to obtain the target 3D pose of all key points. Based on the target 3D pose of all the key points, the virtual character corresponding to the target user is driven in its entirety to obtain a virtual character that matches the current game screen. The process of fusing the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point includes: Calculate the pose variance between the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set. Based on the pose variance, the fusion coefficient between the second 3D pose and the first 3D pose is obtained; Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

2. The method as described in claim 1, characterized in that, The process of fusing the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point includes: Calculate the second variance of the second 3D pose of each core key point, and calculate the first variance of the first 3D pose of the corresponding key point in the human body key point set. The pose variance is constructed by using fusion coefficients to jointly represent the first variance and the second variance, and the fusion coefficients are solved by minimizing the pose variance. Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

3. The method as described in claim 1, characterized in that, The formula for the pose variance is: when When the fusion coefficient is at its minimum, the formula for solving the problem is: in, This represents the pose variance. () represents the variance function. This represents the second 3D pose of the core key point at the current moment. This represents the first 3D pose of the corresponding key points in the current human body key point set. Represents the fusion coefficient. The second variance represents the second 3D pose of each of the core key points. The first variance represents the first 3D pose of the corresponding key point in the human body key point set.

4. The method as described in claim 1, characterized in that, The step of propagating the pose difference between the target 3D pose of each core key point and the first 3D pose of the corresponding key point in the human body key point set to the key points associated with each core key point to obtain the target 3D pose of all key points includes: For each core critical point, perform the following operations: Calculate the target 3D pose of the core key point and the pose difference between the target 3D pose of the corresponding key point in the human body key point set; The target 3D pose of the partial key points is obtained by adding the pose difference to the first 3D pose of the partial key points, which are the human body key points and are associated with the core key points.

5. The method according to any one of claims 1-4, characterized in that, The process of establishing a communication connection with the smart device includes: Create a ServerSocket and listen for connection requests sent by the smart device after it creates a Socket; Based on the connection request, a communication connection is established with the smart device, and a connection success message is sent to the smart device.

6. A virtual reality device, characterized in that, The device includes a handle and a head-mounted display, the head-mounted display including a processor, a memory, and a display screen, the display screen, the memory, and the processor being connected via a bus; The memory stores a computer program, and the processor performs the following operations according to the computer program: Establish a communication connection with a smart device within the same local area network. The smart device includes an RGB camera. The smart device is used to obtain the first 3D pose of each key point in the human body key point set from the human body RGB image of the target user performing actions according to the current game screen captured by the RGB camera. Receive the first 3D pose of each key point in the human body key point set sent by the smart device; Based on the handle and the head-mounted display, the second 3D pose of each core key point of the hand and head in the current action of the target user is obtained; For each core key point that exists in the human body key point set, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point. For each core key point that does not exist in the human body key point set, the second 3D pose of the core key point is directly used as the target 3D pose of the core key point. With the bone length remaining constant, the pose difference between the target 3D pose of each core key point and the first 3D pose of the corresponding key point in the human body key point set is propagated to the key point associated with each core key point to obtain the target 3D pose of all key points. Based on the target 3D pose of all the key points, the virtual character corresponding to the target user is driven in its entirety to obtain a virtual character that matches the current game screen, and then displayed on the display screen. The processor fuses the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point. The specific operation is as follows: Calculate the pose variance between the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set. Based on the pose variance, the fusion coefficient between the second 3D pose and the first 3D pose is obtained; Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

7. The virtual reality device as described in claim 6, characterized in that, The processor fuses the second 3D pose of the core key point with the first 3D pose of the corresponding key point in the human body key point set to obtain the target 3D pose of the core key point. The specific operation is as follows: Calculate the second variance of the second 3D pose of each core key point, and calculate the first variance of the first 3D pose of the corresponding key point in the human body key point set. The pose variance is constructed by using fusion coefficients to jointly represent the first variance and the second variance, and the fusion coefficients are solved by minimizing the pose variance. Based on the fusion coefficient, the second 3D pose of the core key point and the first 3D pose of the corresponding key point in the human body key point set are fused to obtain the target 3D pose of the core key point.

8. The virtual reality device as described in claim 6, characterized in that, The processor propagates the pose difference between the target 3D pose of each core key point and the first 3D pose of the corresponding key point in the human body key point set to the key points associated with each core key point, thereby obtaining the target 3D pose of all key points. The specific operation is as follows: For each core critical point, perform the following operations: Calculate the target 3D pose of the core key point and the pose difference between the target 3D pose of the corresponding key point in the human body key point set; The target 3D pose of the partial key points is obtained by adding the pose difference to the first 3D pose of the partial key points, which are the human body key points and are associated with the core key points.

Citation Information

Patent Citations

  • Room VR-oriented whole-body three-dimensional posture tracking method

    CN110570455A

  • Virtual human control and interaction method based on video stream

    CN110728739A

  • Fusion positioning system and method based on optical tracking and inertial tracking

    CN111947650A