Driving methods and electronic devices for remote 3D human body models
By matching and optimizing the driving data of the 3D human body model, the problem of low driving accuracy caused by rapid movement and data loss was solved, thus improving the user experience.
Patent Information
- Application Number
- CN202211135335.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-09-19
AI Technical Summary
In existing technologies, the driving accuracy of 3D human models is low and the user experience is poor. This is mainly due to motion blur of color images caused by rapid movement, self-occlusion problems under complex movements, and loss and delay of driving data in 3D remote communication systems.
By matching the driving data from the previous moment with the historical driving data, the target historical driving data is determined. Based on this data, the driving data for the current moment of the 3D human body model is predicted and optimized. The predicted driving data is then used to drive the 3D human body model.
It improved the accuracy of driving the 3D human body model, enhanced the user experience, and resolved the issue of driving anomalies.
Smart Images

Figure CN115439612B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional reconstruction technology, and in particular to a driving method and electronic device for a remote three-dimensional human body model. Background Technology
[0002] The 3D reconstruction of a human body model first requires acquiring information from various sensors as input, and then using a 3D reconstruction algorithm to reconstruct the human body model. In recent years, with the continuous development of imaging technology, visual 3D reconstruction technology based on RGB (RGB color mode) cameras has gradually become a research hotspot. Subsequently, the emergence of color depth cameras and the proposal of binocular stereo matching algorithms have further improved the quality and efficiency of human body 3D model reconstruction.
[0003] Building upon 3D human body reconstruction, the 3D human body model needs to be driven by the capture of human motion. Although existing methods have achieved relatively accurate human motion capture performance, they still rely on multi-camera acquisition systems, making them expensive and difficult to use. In recent years, with the significant improvement in the performance of mobile color and depth cameras and the continuous development of deep neural networks, single-view human motion capture methods based on machine learning have emerged. However, these methods are still limited by the limited single-view input information. For example, problems such as motion blur in color images due to rapid movement or self-occlusion of the human body under complex motion can occur, leading to inaccurate calculation of the driving data for the 3D human body model. Furthermore, in 3D remote communication systems, there is a high possibility of data loss and delay during transmission, resulting in driving anomalies during terminal applications. Therefore, existing technologies result in low driving accuracy for 3D human body models and a poor user experience. Summary of the Invention
[0004] This application provides a remote three-dimensional human body model driving method and electronic device, which is used to predict, complete and optimize the driving data of the three-dimensional human body model, solve the driving anomaly problem in the prior art, improve the accuracy of driving the three-dimensional human body model and improve the user experience.
[0005] In a first aspect, embodiments of this application provide a method for driving a remote three-dimensional human body model, comprising:
[0006] For any object undergoing remote 3D communication, if the object's 3D human model had driving data in the previous moment of the current moment, and the object's 3D human model did not have driving data in the current moment, then the driving data in the previous moment of the current moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data.
[0007] Based on the target historical driving data and the historical driving data, determine the predictive driving data of the object's three-dimensional human body model at the current moment;
[0008] The target driving data of the object's three-dimensional human body model at the current moment is obtained by using the prediction driving data at the current moment and the driving data at the previous moment.
[0009] The target-driven data at the current moment is used to drive the body and face of the object's three-dimensional human model.
[0010] A second aspect of this application provides an electronic device, including a processor and a memory, wherein the processor and the memory are connected via a bus;
[0011] The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program:
[0012] For any object undergoing remote 3D communication, if the object's 3D human model had driving data in the previous moment of the current moment, and the object's 3D human model did not have driving data in the current moment, then the driving data in the previous moment of the current moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data.
[0013] Based on the target historical driving data and the historical driving data, determine the predictive driving data of the object's three-dimensional human body model at the current moment;
[0014] The target driving data of the object's three-dimensional human body model at the current moment is obtained by using the prediction driving data at the current moment and the driving data at the previous moment.
[0015] The target-driven data at the current moment is used to drive the body and face of the object's three-dimensional human model.
[0016] According to a third aspect of the present invention, a computer storage medium is provided, the computer storage medium storing a computer program for performing the method as described in the first aspect.
[0017] In the above embodiments of this application, when it is determined that the 3D human body model of the object undergoing remote 3D communication had driving data in the previous moment of the current moment, and the 3D human body model of the object does not have driving data in the current moment, the driving data in the previous moment of the current moment is matched with the historical driving data of the object to determine the target historical driving data that matches the driving data in the previous moment. Then, based on the target historical driving data, the predicted driving data of the 3D human body model in the current moment is predicted. Finally, the predicted driving data is optimized using preset driving data and the driving data in the previous moment to obtain the target driving data. Thus, this embodiment realizes the prediction, completion, and optimization of the 3D human body model in the current moment, solves the problem of driving anomalies in the prior art, improves the accuracy of driving the 3D human body model, and improves the user experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 An exemplary illustration shows one of the application scenarios provided in the embodiments of this application;
[0020] Figure 2 The second example illustration shows a schematic diagram of an application scenario provided in an embodiment of this application;
[0021] Figure 3 An exemplary system architecture diagram of the driving method for a remote three-dimensional human body model provided in an embodiment of this application is shown;
[0022] Figure 4 One of the flowcharts illustrating the driving method for a remote three-dimensional human body model provided in an embodiment of this application is shown as an example;
[0023] Figure 5 An exemplary schematic diagram illustrates the process of determining target historical driving data provided in an embodiment of this application;
[0024] Figure 6 The second example is a schematic flowchart of the driving method for a remote three-dimensional human body model provided in an embodiment of this application;
[0025] Figure 7 An exemplary illustration shows one of the application scenarios of the remote three-dimensional human body model driving method provided in the embodiments of this application;
[0026] Figure 8 This paper exemplifies a second application scenario diagram of the remote three-dimensional human body model driving method provided in the embodiments of this application;
[0027] Figure 9 An exemplary schematic diagram of the structure of the remote three-dimensional human body model device provided in an embodiment of this application is shown;
[0028] Figure 10 An exemplary hardware structure diagram of the calibration device provided in an embodiment of this application is shown. Detailed Implementation
[0029] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0030] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation on its own.
[0031] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0032] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to be omnipresent but not exclusive; for example, a product or device comprising a series of components is not necessarily limited to those explicitly listed, but may include other components not explicitly listed or inherent to such product or device.
[0033] As used in this application, the term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.
[0034] The following is an overview of the ideas behind the embodiments of this application.
[0035] Current 3D human body model driving technologies suffer from issues such as motion blur in color images during rapid movement or self-occlusion of the human body during complex motion. These problems prevent accurate calculation of the driving data for the 3D human body model. Furthermore, in 3D remote communication systems, data loss and delays are likely to occur during transmission, leading to driving anomalies in terminal applications. Therefore, existing technologies result in low driving accuracy for 3D human body models and a poor user experience.
[0036] To address the issue of driving anomalies in existing technologies, this application provides a method for driving a remote 3D human body model. The method determines whether driving data existed for the 3D human body model of the object undergoing remote 3D communication in the previous moment, or whether driving data does not exist in the current moment. It then matches the driving data from the previous moment with the object's historical driving data to identify target historical driving data that matches the previous moment's data. Based on this target historical driving data, a predicted driving data for the 3D human body model at the current moment is then predicted. Finally, the predicted driving data is optimized using preset driving data and the previous moment's driving data to obtain the target driving data. Therefore, this embodiment achieves prediction, completion, and optimization of the 3D human body model at the current moment, solving the driving anomaly problem in existing technologies, improving the accuracy of 3D human body model driving, and enhancing the user experience.
[0037] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0038] Figure 1 An exemplary illustration shows an application scenario diagram of the remote three-dimensional human body model driving method provided in the embodiments of this application; such as Figure 1 As shown, this application scenario uses an electronic device as a VR device for illustration. This application scenario includes a camera 110, a server 120, a server 130, and a VR device 140. Servers 120 and 130 can be implemented using a single server or multiple servers. Servers 120 and 130 can be implemented using physical servers or virtual servers.
[0039] In one possible application scenario, for any object undergoing remote 3D communication, camera 120 acquires the object's geometric and texture data in real time. Server 120 then reconstructs a 3D human body model based on this data, obtaining a 3D human body model and driving data. Server 120 then sends the 3D human body model and driving data to server 130, which in turn sends them to VR device 140. If VR device 140 determines that the object's 3D human body model had driving data in the previous time step and does not have driving data in the current time step, it utilizes the driving data from the previous time step. The VR device 140 matches the motion data with the historical driving data of the object to determine the target historical driving data that matches the driving data of the previous moment. The driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, the VR device 140 determines the predicted driving data of the object's three-dimensional human model at the current moment. It then obtains the target driving data of the object's three-dimensional human model at the current moment by using the predicted driving data of the current moment and the driving data of the previous moment. Finally, the VR device 140 uses the target driving data of the current moment to drive the body and face of the object's three-dimensional human model and displays it.
[0040] like Figure 2 The diagram shown is a schematic diagram of another application scenario of the driving method for a remote three-dimensional human body model provided in the embodiments of this application. The application scenario includes a camera 110, a server 120, a server 130, a VR device 140, and a memory 150.
[0041] In one possible application scenario, for any object undergoing remote 3D communication, camera 120 acquires the object's geometric and texture data in real time and stores this data in memory 150. Then, server 120 reconstructs a 3D human body model based on the geometric and texture data retrieved from memory 150, obtaining a 3D human body model and driving data. Server 120 then sends the 3D human body model and driving data to server 130, which in turn sends them to VR device 140. VR device 140 determines that the object's 3D human body model had driving data in the previous time step, but does not have driving data in the current time step. The VR device 140 then matches the driving data from the previous moment with the historical driving data of the object to determine the target historical driving data that matches the driving data from the previous moment. This driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, the VR device 140 determines the predicted driving data of the object's 3D human model at the current moment. It then uses the predicted driving data from the current moment and the driving data from the previous moment to obtain the target driving data of the object's 3D human model at the current moment. Finally, the VR device 140 uses the target driving data from the current moment to drive and display the body and face of the object's 3D human model.
[0042] in, Figure 1 as well as Figure 2 The server 120 and camera 110, the server 120 and server 130, and the server 130 and VR device 140 can interact with each other through a communication network. The communication network can be either wireless or wired.
[0043] For example, server 120 can access the network via cellular mobile communication technology to communicate with camera 110 and server 130 respectively, and server 130 can access the network via cellular mobile communication technology to communicate with server 120 and VR device 140 respectively. The cellular mobile communication technology includes, for example, 5G technology.
[0044] Optionally, server 120 can access the network via short-range wireless communication to communicate with camera 110 and server 130 respectively, and server 130 can access the network via short-range wireless communication to communicate with server 120 and VR device 140 respectively. The short-range wireless communication method may include, for example, Wireless Fidelity (Wi-Fi) technology.
[0045] Furthermore, the description in this application only details a single camera 110, a single server 120, a single server 130, and a single VR device 140. However, those skilled in the art should understand that the illustrated camera 110, server 120, single server 130, and single VR device 140 are intended to represent the operation of the camera 110, server 120, server 130, and VR device 140 involved in the technical solution of this application, and are not intended to imply any limitation on the number, type, or location of the camera 110, server 120, server 130, and VR device 140. It should be noted that adding additional modules to or removing individual modules from the illustrated environment will not change the underlying concept of the exemplary embodiments of this application.
[0046] It should be noted that the driving method for remote 3D human body models proposed in this application is not only applicable to... Figure 1 and Figure 2 The application scenarios shown can also be applied to any driven device with a remote 3D human body model.
[0047] The following describes an exemplary embodiment of the driving method for a remote three-dimensional human body model, in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the methods and principles of this application, and the implementation of this application is not limited in any way in this respect.
[0048] Before providing a detailed description of the driving method for the remote 3D human body model in this application, the system architecture of this application will be described first, such as... Figure 3 The diagram shown is a system architecture diagram of the remote 3D human body model of this application. Figure 3 It includes the acquisition end, the cloud, and the rendering and display end.
[0049] The acquisition unit is responsible for capturing the 3D pose of the human body and reconstructing the 3D human body model. The acquisition unit primarily uses RGB or RGBD cameras to collect geometric and texture data for real-time reconstruction, and then uses a compatible host / workstation to perform related calculations on parameters such as driving data to obtain the 3D human body model and driving data.
[0050] The cloud-based system is primarily responsible for sending the 3D human body model and driving data determined by the acquisition end to the rendering and display end.
[0051] The rendering display end is mainly responsible for predicting, completing, and optimizing the driving data obtained from the cloud. Then, it renders the acquired 3D human body model and the processed driving data so that the driving data drives the 3D human body model and performs immersive 3D display in VR devices. It can also be connected to terminal devices such as TVs and mobile phones for 2D display.
[0052] like Figure 4 The diagram shown illustrates a flowchart of a remote 3D human body model driving method, which may include the following steps:
[0053] Step 401: For any object undergoing remote 3D communication, if the object's 3D human model had driving data in the previous moment and the object's 3D human model does not have driving data in the current moment, then the driving data in the previous moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data.
[0054] The historical driving data refers to the pre-stored driving data of each historical remote 3D communication of the object or the driving data of the object at each time point before the current time in the current remote 3D communication.
[0055] like Figure 5 The diagram shown illustrates the process for determining target historical driving data, including the following steps:
[0056] Step 501: Based on the target time driving data of the object and the driving data corresponding to each historical time in the historical driving data, obtain the similarity between the target time driving data and the driving data of each historical time, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time, and the target time and the historical time have the same length.
[0057] For example, if the target time is 10 seconds long and the historical driving data is 50 seconds long, then the historical times are 1-10 seconds, 2-11 seconds, 3-12 seconds, 4-13 seconds, and so on.
[0058] It should be noted that the length of the target time mentioned above is for illustrative purposes only and does not constitute a limitation on the length of the target time. The length of the target time can be set according to the actual situation.
[0059] The similarity between driving data can be determined by the distance between two driving data points; that is, the distance between two driving data points is used as the similarity between them. For example, Euclidean distance, shape distance, and pattern distance are all methods used in the prior art, and will not be elaborated upon in this embodiment.
[0060] In addition, the similarity between two driving data can be determined by a preset DTW (dynamic time warping) algorithm. The specific method for determining the similarity between driving data can be set according to the actual situation. This embodiment does not limit the specific method.
[0061] Step 502: The historical driving data with the highest similarity to the driving data at the target time is determined as the target historical driving data.
[0062] For example, if the historical driving data corresponding to the 20th to 30th second in the historical driving data has the highest similarity to the target time, then the historical driving data corresponding to the 20th to 30th second in the historical driving data is determined as the target historical driving data.
[0063] Step 402: Based on the target historical driving data and the historical driving data, determine the predicted driving data of the object's three-dimensional human body model at the current moment;
[0064] In one embodiment, step 402 may be implemented as: determining the driving data in the historical driving data that is located at the next historical moment of the target historical driving data as the predicted driving data for the current moment.
[0065] For example, if the length of a historical moment is 10 seconds, and the historical driving data from the 20th to the 30th second of the historical driving data is the target historical driving data, then the historical driving data from the 40th second of the historical driving data is determined as the predicted driving data for the current moment. If the length of a historical moment is 1 second, and the historical driving data from the 3rd second of the historical driving data is the target historical missing data, then the historical driving data from the 4th second of the historical driving data is determined as the predicted driving data for the current moment. Step 403: Using the predicted driving data for the current moment and the driving data from the previous moment, obtain the target driving data for the 3D human body model of the object at the current moment;
[0066] In one embodiment, step 403 may be specifically implemented as follows: using a preset interpolation algorithm to interpolate the predicted driving data at the current time and the driving data at the target time of the object to obtain the target driving data at the current time, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time.
[0067] For example, if the current time is the 11th second and the target time is from the 1st to the 10th second, and the historical driving data from the 20th to the 30th second is the predicted driving data for the 11th second, then the driving data from the 1st to the 10th second and the historical driving data from the 20th to the 30th second will be interpolated.
[0068] If the current time is the 11th second, the target time is the 10th second, and the historical driving data of the 30th second in the historical driving data is the predicted driving data of the 11th second, then the driving data of the 10th second and the historical driving data of the 30th second in the historical driving data are interpolated to obtain the target driving data of the 11th second of the current time.
[0069] The driving data includes facial driving data and body driving data. In this embodiment, the interpolation method used for facial driving data differs from that used for body driving data. Facial driving data is one-dimensional, and in this embodiment, a preset single-linear interpolation algorithm is used for interpolation. Body driving data is interpolated using a preset spherical linear interpolation algorithm.
[0070] It should be noted that the interpolation algorithm used in this embodiment is only for illustrative purposes and does not limit the specific interpolation method. The specific interpolation algorithm can be set according to the actual situation.
[0071] Step 404: Use the target-driven data at the current moment to drive the body and face of the object's three-dimensional human model.
[0072] The driving mechanism for the 3D human body model mainly consists of two parts: facial driving and body driving. Facial driving data from the target driving data is used to drive the 3D human body model's face, and body driving data from the target driving data is used to drive the 3D human body model's body.
[0073] Therefore, this application determines that when the 3D human model of the object undergoing remote 3D communication had driving data in the previous moment, and the 3D human model of the object does not have driving data in the current moment, it matches the driving data of the previous moment with the historical driving data of the object to determine the target historical driving data that matches the driving data of the previous moment. Then, based on the target historical driving data, it predicts the predicted driving data of the 3D human model in the current moment. Finally, it optimizes the predicted driving data using preset driving data and the driving data of the previous moment to obtain the target driving data. Thus, this embodiment realizes the prediction, completion, and optimization of the 3D human model in the current moment, solves the driving anomaly problem existing in the prior art, improves the accuracy of the 3D human model driving, and improves the user experience.
[0074] To ensure the integrity of the driving data, one embodiment may include the following three cases:
[0075] (1) If the three-dimensional human body model of the object has driving data in the previous time of the current time, and the three-dimensional human body model of the object has driving data in the current time, then the body and face of the three-dimensional human body model of the object are driven by the driving data in the current time.
[0076] The driving of the 3D human model mainly consists of two parts: facial driving and body driving. Facial driving data from the current driving data is used to drive the 3D human model's face, and body driving data from the current driving data is used to drive the 3D human model's body.
[0077] (2) If the three-dimensional human body model of the object does not have driving data in the previous time of the current time, and the three-dimensional human body model of the object does not have driving data in the current time, then the preset driving data is determined as the predicted driving data for the current time.
[0078] The preset driving data includes preset body driving data and preset facial driving data. This preset driving data is primarily used for default driving after the 3D human model is loaded. For example, the 3D human model's facial expressions will naturally include smiling and head movements, and the body will naturally move slightly from side to side.
[0079] It should be noted that the preset driver data can be set according to the actual situation, and this embodiment does not limit the specific value of the preset driver data.
[0080] (3) If the three-dimensional human body model of the object does not have driving data in the previous time of the current time, and the three-dimensional human body model of the object has driving data in the current time, then the body and face of the three-dimensional human body model of the object are driven by the driving data in the current time.
[0081] The driving of the 3D human model mainly consists of two parts: facial driving and body driving. Facial driving data from the current driving data is used to drive the 3D human model's face, and body driving data from the current driving data is used to drive the 3D human model's body.
[0082] To further connect the technical solutions in this application, the following is combined with... Figure 6 A detailed explanation may include the following steps:
[0083] Step 601: For any object undergoing remote 3D communication, determine whether the object's 3D human model had driving data in the previous time. If yes, proceed to step 602; otherwise, proceed to step 607.
[0084] Step 602: Determine whether the three-dimensional human body model of the object has driving data at the current moment. If yes, proceed to step 603; otherwise, proceed to step 604.
[0085] Step 603: Drive the body and face of the object's 3D human model using the driving data at the current moment;
[0086] Step 604: Based on the target time driving data of the object and the driving data corresponding to each historical time in the historical driving data, obtain the similarity between the target time driving data and the driving data of each historical time, wherein the target time is the previous time of the current time or a series of consecutive times adjacent to the current time, and the target time and the historical time have the same length, wherein the driving data includes facial driving data and body driving data;
[0087] Wherein, the historical driving data is the pre-stored driving data of each historical remote three-dimensional communication of the object or the driving data corresponding to each moment before the current moment in the current remote three-dimensional communication of the object;
[0088] Step 605: The historical driving data with the highest similarity to the driving data at the target time is determined as the target historical driving data;
[0089] Step 606: Determine the driving data in the historical driving data that is located at the next historical moment of the target historical driving data as the predicted driving data for the current moment;
[0090] Step 607: Determine whether the three-dimensional human body model of the object has driving data at the current moment. If yes, proceed to step 603; otherwise, proceed to step 608.
[0091] Step 608: Determine the preset driving data as the predicted driving data at the current moment;
[0092] Step 609: Use a preset interpolation algorithm to interpolate the predicted driving data at the current time and the driving data at the target time of the object to obtain the target driving data at the current time, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time.
[0093] Step 610: Use the target-driven data at the current moment to drive the body and face of the object's three-dimensional human model.
[0094] The driving method for the remote 3D human body model in this application will be described in detail below, taking into account specific application scenarios:
[0095] Scenario 1: Live Streaming
[0096] like Figure 7 As shown, the broadcaster user obtains their own geometric and texture data in real time via mobile phone 710. Server 720 then reconstructs a 3D human body model based on this data, obtaining a 3D human body model and driving data. Server 720 then sends the 3D human body model and driving data to server 730, which in turn sends them to VR device 740. If VR device 740 determines that the object's 3D human body model had driving data in the previous time step, and that the object's 3D human body model does not have driving data in the current time step, then it matches the driving data from the previous time step with the object's historical driving data. The VR device 740 determines target historical driving data that matches the driving data from the previous moment in the historical driving data, wherein the driving data includes facial driving data and body driving data; based on the target historical driving data and the historical driving data, the VR device 740 determines the predicted driving data of the object's 3D human model at the current moment, and obtains the target driving data of the object's 3D human model at the current moment through the predicted driving data at the current moment and the driving data from the previous moment; finally, the VR device 740 uses the target driving data at the current moment to drive the body and face of the object's 3D human model and displays it, and the viewer browses the broadcaster's live stream through the VR device 740.
[0097] Scenario 2: Remote Conferencing
[0098] like Figure 8 As shown, each user participating in the remote conference (User 1, User 2, and User 3) obtains their own geometric and texture data in real time via mobile phone 810. Based on this data, the corresponding server 820 reconstructs a 3D human body model, obtaining a 3D human body model and driving data. Server 820 then sends the 3D human body model and driving data to server 830, which in turn sends it to the VR devices 840 of other users. If VR device 840 determines that the object's 3D human body model had driving data in the previous moment and does not have driving data in the current moment, it uses the driving data from the previous moment and the object's historical driving data. The system performs matching to determine target historical driving data that matches the driving data from the previous moment. This driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, the VR device 840 determines the predicted driving data of the object's 3D human model at the current moment. Using the predicted driving data at the current moment and the driving data from the previous moment, the VR device 840 obtains the target driving data of the object's 3D human model at the current moment. Finally, the VR device 840 uses the target driving data at the current moment to drive and display the body and face of the object's 3D human model. Users attending the meeting can see the 3D human models of other users participating in the remote meeting through the VR device 840.
[0099] Scenario 3: Remote Video Scenario
[0100] like Figure 8As shown, each user participating in the remote video (User 1, User 2, and User 3) acquires their own geometric and texture data in real time via mobile phone 810. Based on this acquired geometric and texture data, the corresponding server 820 reconstructs a 3D human body model, obtaining a 3D human body model and driving data. Server 820 then sends the 3D human body model and driving data to server 830, which in turn sends it to the VR devices 840 of other video users. If VR device 840 determines that the object's 3D human body model had driving data in the previous moment and does not have driving data in the current moment, it uses the driving data from the previous moment and the object's historical driving data. The system performs matching to determine target historical driving data that matches the driving data from the previous moment. This driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, the VR device 840 determines the predicted driving data of the object's 3D human model at the current moment. Then, using the predicted driving data at the current moment and the driving data from the previous moment, the VR device 840 obtains the target driving data of the object's 3D human model at the current moment. Finally, the VR device 840 uses the target driving data at the current moment to drive and display the body and face of the object's 3D human model. Users performing remote video calls can see the 3D human models of other users performing video calls through the VR device 840.
[0101] It should be noted that in remote conferencing scenarios, each user needs to configure a capture terminal and a rendering terminal.
[0102] Based on the same inventive concept, the remote three-dimensional human body model driving method described above can also be implemented by a remote three-dimensional human body model driving device. The effect of this remote three-dimensional human body model driving device is similar to that of the aforementioned method, and will not be described again here.
[0103] Figure 9 This is a schematic diagram of the structure of a driving device for a remote three-dimensional human body model according to an embodiment of the present disclosure.
[0104] like Figure 9 As shown, the driving device 900 for the remote three-dimensional human body model disclosed herein may include a matching module 910, a first prediction driving data determination module 920, a target driving data determination module 930, and a first driving module 940.
[0105] The matching module 910 is used to match any object that performs remote 3D communication. If the object's 3D human body model had driving data in the previous moment of the current moment, and the object's 3D human body model does not have driving data in the current moment, then the driving data in the previous moment of the current moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data.
[0106] The first prediction driving data determination module 920 is used to determine the prediction driving data of the three-dimensional human body model of the object at the current moment based on the target historical driving data and the historical driving data.
[0107] The target-driven data determination module 930 is used to obtain the target-driven data of the three-dimensional human body model of the object at the current moment through the predicted driving data at the current moment and the driving data at the previous moment.
[0108] The first driving module 940 is used to drive the body and face of the three-dimensional human model of the object using the target driving data at the current moment.
[0109] In one embodiment, the apparatus further includes:
[0110] The second driving module 950 is used to drive the body and face of the three-dimensional human body model of the object using the driving data at the current moment if the three-dimensional human body model of the object has driving data at the previous moment and the three-dimensional human body model of the object has driving data at the current moment.
[0111] The second prediction driving data determination module 960 is used to determine the preset driving data as the prediction driving data at the current moment if the three-dimensional human body model of the object does not have driving data in the previous moment at the current moment, and the three-dimensional human body model of the object does not have driving data at the current moment.
[0112] The third driving module 970 is used to drive the body and face of the three-dimensional human body model of the object using the driving data at the current moment if there is no driving data for the three-dimensional human body model of the object in the previous moment and there is driving data for the three-dimensional human body model of the object in the current moment.
[0113] In one embodiment, the historical driving data is pre-stored driving data of each historical remote 3D communication of the object or driving data corresponding to each moment before the current moment in the current remote 3D communication of the object;
[0114] The matching module 910 is specifically used for:
[0115] Based on the target time driving data of the object and the driving data corresponding to each historical time in the historical driving data, the similarity between the target time driving data and the driving data of each historical time is obtained, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time, and the target time and the historical time have the same length.
[0116] The historical driving data with the highest similarity to the driving data at the target time is determined as the target historical driving data.
[0117] In one embodiment, the first prediction-driven data determination module 920 is specifically used for:
[0118] The driving data located at the next historical moment of the target historical driving data in the historical driving data is determined as the predicted driving data for the current moment.
[0119] In one embodiment, the target-driven data determination module 930 is specifically used for:
[0120] The predicted driving data at the current time and the driving data at the target time of the object are interpolated using a preset interpolation algorithm to obtain the target driving data at the current time. The target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time.
[0121] After introducing a driving method and apparatus for a remote three-dimensional human body model according to an exemplary embodiment of the present invention, the electronic device according to another exemplary embodiment of the present invention will be introduced next.
[0122] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as "circuit", "module", or "system".
[0123] In some possible implementations, the electronic device according to the present invention may include at least one processor and at least one computer storage medium. The computer storage medium stores program code that, when executed by the processor, causes the processor to perform the steps in the driving method for a remote three-dimensional human body model according to various exemplary embodiments of the present invention described above. For example, the processor may perform actions such as... Figure 4 Steps 401-404 are shown in the diagram.
[0124] The following reference Figure 10 To describe an electronic device 1000 according to this embodiment of the present invention. Figure 10 The electronic device 1000 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0125] like Figure 10 As shown, the electronic device 1000 is manifested in the form of a general electronic device. The components of the electronic device 1000 may include, but are not limited to: at least one processor 1001, at least one computer storage medium 1002, and a bus 1003 connecting different system components (including the computer storage medium 1002 and the processor 1001).
[0126] Bus 1003 represents one or more of several bus structures, including computer storage media bus or computer storage media controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0127] Computer storage medium 1002 may include readable media in the form of volatile computer storage media, such as random access computer storage medium (RAM) 1021 and / or cache storage medium 1022, and may further include read-only computer storage medium (ROM) 1023.
[0128] The computer storage medium 1002 may also include a program / utility 1025 having a set (at least one) of program modules 1024, such program modules 1024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0129] Electronic device 1000 can also communicate with one or more external devices 1004 (e.g., keyboard, pointing device, etc.), one or more devices that enable objects to interact with electronic device 1000, and / or any device that enables electronic device 1000 to communicate with one or more other electronic devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1005. Furthermore, electronic device 1000 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1006. As shown, network adapter 1006 communicates with other modules used in electronic device 1000 via bus 1003. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0130] In some possible implementations, various aspects of the remote three-dimensional human body model driving method provided by the present invention can also be implemented in the form of a program product, which includes program code that, when the program product is run on a computer device, causes the computer device to perform the steps in the remote three-dimensional human body model driving method according to various exemplary embodiments of the present invention described above.
[0131] The program product may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access computer storage media (RAM), read-only computer storage media (ROM), erasable programmable read-only computer storage media (EPROM or flash memory), optical fibers, portable compact disk read-only computer storage media (CD-ROM), optical computer storage media, magnetic computer storage media, or any suitable combination thereof.
[0132] The program product driving the remote three-dimensional human body model according to embodiments of the present invention can be a portable compact disc read-only computer storage medium (CD-ROM) and include program code, and can run on an electronic device. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0133] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0134] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0135] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the target electronic device, partially on the target device, as a standalone software package, partially on the target electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the target electronic device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external electronic device (e.g., via the Internet using an Internet service provider).
[0136] It should be noted that although several modules of the apparatus have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0137] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk computer storage media, CD-ROMs, optical computer storage media, etc.) containing computer-usable program code.
[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0140] These computer program instructions may also be stored in a computer-readable computer storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable computer storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0142] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for driving a remote three-dimensional human body model, characterized in that, The method includes: For any object undergoing remote 3D communication, if the object's 3D human model had driving data in the previous moment of the current moment, and the object's 3D human model did not have driving data in the current moment, then the driving data in the previous moment of the current moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, determine the predictive driving data of the object's three-dimensional human body model at the current moment; The target driving data of the object's three-dimensional human body model at the current moment is obtained by using the prediction driving data at the current moment and the driving data at the previous moment. The target-driven data at the current moment is used to drive the body and face of the object's three-dimensional human model.
2. The method according to claim 1, characterized in that, The method further includes: If the object's 3D human model had driving data in the previous time step and the object's 3D human model has driving data in the current time step, then the body and face of the object's 3D human model are driven using the driving data in the current time step; or, If the 3D human body model of the object did not have driving data in the previous time step, and the 3D human body model of the object does not have driving data in the current time step, then the preset driving data is determined as the predicted driving data for the current time step; or, If the three-dimensional human body model of the object did not have driving data in the previous time step, and the three-dimensional human body model of the object has driving data in the current time step, then the body and face of the three-dimensional human body model of the object are driven using the driving data in the current time step.
3. The method according to claim 1, characterized in that, The historical driving data is the pre-stored driving data of each historical remote 3D communication of the object or the driving data of the object at each time point before the current time in the current remote 3D communication; The step of matching the driving data from the previous moment with the historical driving data of the object to determine the target historical driving data that matches the driving data from the previous moment includes: Based on the target time driving data of the object and the driving data corresponding to each historical time in the historical driving data, the similarity between the target time driving data and the driving data of each historical time is obtained, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time, and the target time and the historical time have the same length. The historical driving data with the highest similarity to the driving data at the target time is determined as the target historical driving data.
4. The method according to claim 3, characterized in that, The step of determining the predictive driving data of the object's 3D human model at the current moment based on the target's historical driving data and the historical driving data includes: The driving data located at the next historical moment of the target historical driving data in the historical driving data is determined as the predicted driving data for the current moment.
5. The method according to claim 1, characterized in that, The step of obtaining the target driving data of the object's 3D human model at the current moment using the predicted driving data at the current moment and the driving data from the previous moment includes: The predicted driving data at the current time and the driving data at the target time of the object are interpolated using a preset interpolation algorithm to obtain the target driving data at the current time. The target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time.
6. An electronic device, characterized in that, It includes a processor and a memory, which are connected via a bus; The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program: For any object undergoing remote 3D communication, if the object's 3D human model had driving data in the previous moment of the current moment, and the object's 3D human model did not have driving data in the current moment, then the driving data in the previous moment of the current moment is matched with the object's historical driving data to determine the target historical driving data that matches the driving data in the previous moment. The driving data includes facial driving data and body driving data. Based on the target historical driving data and the historical driving data, determine the predictive driving data of the object's three-dimensional human body model at the current moment; The target driving data of the object's three-dimensional human body model at the current moment is obtained by using the prediction driving data at the current moment and the driving data at the previous moment. The target-driven data at the current moment is used to drive the body and face of the object's three-dimensional human model.
7. The electronic device according to claim 6, characterized in that, The processor is also configured to: If the three-dimensional human body model of the object has driving data in the previous time before the current time, and the three-dimensional human body model of the object has driving data in the current time, then the body and face of the three-dimensional human body model of the object are driven using the driving data in the current time. or, If the three-dimensional human body model of the object did not have driving data in the previous time of the current time, and the three-dimensional human body model of the object did not have driving data in the current time, then the preset driving data will be determined as the predicted driving data for the current time. or, If the three-dimensional human body model of the object did not have driving data in the previous time step, and the three-dimensional human body model of the object has driving data in the current time step, then the body and face of the three-dimensional human body model of the object are driven using the driving data in the current time step.
8. The electronic device according to claim 6, characterized in that, The historical driving data is the pre-stored driving data of each historical remote 3D communication of the object or the driving data of the object at each time point before the current time in the current remote 3D communication; The processor executes the process of matching the driving data from the previous moment with the historical driving data of the object to determine the target historical driving data that matches the driving data from the previous moment. Specifically, this is configured as follows: Based on the target time driving data of the object and the driving data corresponding to each historical time in the historical driving data, the similarity between the target time driving data and the driving data of each historical time is obtained, wherein the target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time, and the target time and the historical time have the same length. The historical driving data with the highest similarity to the driving data at the target time is determined as the target historical driving data.
9. The electronic device according to claim 8, characterized in that, The processor executes the process of determining the predictive driving data of the object's 3D human model at the current moment based on the target's historical driving data and the historical driving data, specifically configured as follows: The driving data located at the next historical moment of the target historical driving data in the historical driving data is determined as the predicted driving data for the current moment.
10. The electronic device according to claim 6, characterized in that, The processor executes the process of obtaining the target driving data of the object's 3D human model at the current moment using the predicted driving data at the current moment and the driving data from the previous moment, specifically configured as follows: The predicted driving data at the current time and the driving data at the target time of the object are interpolated using a preset interpolation algorithm to obtain the target driving data at the current time. The target time is the previous time of the current time or a series of consecutive times that are adjacent to the current time.
Citation Information
Patent Citations
Three-dimensional model generation method and device
CN113888696A
Human body reconstruction frame insertion method and related product
CN115035238A