Multi-camera-based digital human control method and device
By acquiring and fusing key rotation matrices through a multi-camera system, and combining joint visibility and inverse kinematics algorithms, the problem of inaccurate control of digital humans under single-camera control is solved, achieving higher joint rotation prediction accuracy and stability, and ensuring that the digital human stands stably.
Patent Information
- Application Number
- CN202511562356.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-27
AI Technical Summary
In existing technologies, digital human control methods based on a single camera suffer from inaccurate acquisition of key rotation matrices, leading to inaccurate control strategies and insufficient accuracy and stability in predicting joint rotations of the digital human. In particular, the stability and uniformity of footstep timing are poor in multi-camera environments, resulting in the digital human being unable to stand stably.
By employing a multi-camera system, the system acquires image sets from each camera, calculates the key rotation matrix corresponding to each camera, performs fusion prediction, and combines joint visibility and inverse kinematics algorithms to correct joint rotation amounts, thereby obtaining an accurate digital human joint rotation matrix and achieving stable control of the digital human.
It improves the accuracy of joint rotation prediction in digital humans, reduces the probability of foot floating, enhances the stability and consistency of digital human control, ensures that digital humans can stand stably, and improves the accuracy and stability of control.
Smart Images

Figure CN121582406A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a digital human control method and device based on multi-camera. BACKGROUND
[0002] With the development of the times, the era of big data has arrived, and the application of artificial intelligence has become the top priority of current companies. The three-dimensional human key point detection algorithm is a widely used algorithm, which has been applied in various fields, especially the human 3D key point detection in videos, which can be used for digital human driving. For example, a single camera can be used for limb driving control. SUMMARY
[0003] The present disclosure provides a digital human control method and device based on multi-camera, which can reduce the inaccuracy of key rotation matrix acquisition, improve the accuracy of joint rotation prediction of digital human, and improve the accuracy and stability of digital human control. The technical solution of the present disclosure is as follows: According to a first aspect of the embodiments of the present disclosure, a digital human control method based on multi-camera is provided, comprising: obtaining an image set collected by each camera in a camera set for a digital human, wherein the camera set comprises a plurality of cameras with different image collection angles; obtaining a first key rotation matrix corresponding to each camera according to the image set; fusing and predicting the first key rotation matrix corresponding to each camera to obtain a second key rotation matrix; obtaining a first key point corresponding to a target part of the digital human; obtaining a third key rotation matrix corresponding to the digital human according to the first key point and the second key rotation matrix, and controlling the digital human by using a control strategy corresponding to the third key rotation matrix.
[0004] According to some embodiments, the first key rotation matrix corresponding to each camera is obtained according to the image set, comprising: identifying the image set by using a prediction model to obtain first image information corresponding to each camera, wherein the image information comprises joint rotation information, second three-dimensional key points, and a rotation matrix of a root node; calculating the joint rotation information and the second three-dimensional key points by using an inverse kinematics algorithm to obtain a fourth joint rotation matrix; performing inverse rotation processing on the first joint rotation matrix and the rotation matrix of the root node to obtain the first joint rotation matrix corresponding to each camera.
[0005] According to some embodiments, the obtaining the first key point corresponding to the target part of the digital human comprises: The second three-dimensional key point is inversely rotated to obtain the first key point corresponding to the target part of the digital human.
[0006] According to some embodiments, wherein the image information further comprises joint visibility, the fusing and predicting the first key rotation matrix corresponding to each camera to obtain a second key rotation matrix comprises: According to the joint visibility value corresponding to each camera, a sum of joint visibility values is obtained. According to the first key rotation matrix corresponding to each camera, the joint visibility value corresponding to each camera, and the sum of joint visibility values, a weighted average processing is performed on the first key rotation matrix corresponding to each camera to obtain a second key rotation matrix.
[0007] According to some embodiments, wherein the image information further comprises joint visibility, the obtaining the third key rotation matrix corresponding to the digital human according to the first key point and the second key rotation matrix comprises: According to the first key point and the second key rotation matrix, a rotation amount of each joint in a joint set of the digital human is corrected to obtain a third key rotation matrix corresponding to the digital human, wherein the joint set comprises a foot, a first joint, and a second joint, the first joint is an upper level joint adjacent to the foot, and the second joint is an upper joint adjacent to the first joint.
[0008] According to some embodiments, the correcting the rotation amount of each joint in the joint set of the digital human according to the first key point and the second key rotation matrix to obtain the third key rotation matrix corresponding to the digital human comprises: A third key point corresponding to the foot of the digital human is obtained from the first key point. According to the confidence of the joint visibility, a weighted average processing is performed on the third key point corresponding to the foot to obtain three-dimensional key point space information corresponding to the foot. According to the second key rotation matrix, a first rotation amount corresponding to the first joint, and a second rotation amount corresponding to the second joint, a first increment of the first rotation amount and a second increment of the second rotation amount are obtained. According to the first increment and the second increment, an iterative processing is performed on the second key rotation matrix to obtain a third key rotation matrix.
[0009] According to some embodiments, the obtaining the first increment of the first rotation amount and the second increment of the second rotation amount according to the second key rotation matrix, the first rotation amount corresponding to the first joint, and the second rotation amount corresponding to the second joint comprises: applying the second key rotation matrix to the multi-person linear model using a forward kinematics algorithm to obtain a rotated multi-person linear model; extracting spatial coordinates corresponding to each joint in the joint set according to the rotated multi-person linear model; performing iterative processing on the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint according to the spatial coordinates corresponding to the foot in the spatial coordinates of each joint and the three-dimensional key point spatial information corresponding to the foot, to obtain the first increment of the first rotation amount and the second increment of the second rotation amount.
[0010] According to a second aspect of the embodiments of the present disclosure, a multi-camera-based digital human control device is provided, comprising: a set obtaining unit configured to obtain an image set collected by each camera in a camera set for a digital human, wherein the camera set comprises cameras with different image collection angles; a matrix obtaining unit configured to obtain a first key rotation matrix corresponding to each camera according to the image set; The matrix obtaining unit is further configured to perform fusion prediction on the first key rotation matrix corresponding to each camera to obtain a second key rotation matrix. a key point obtaining unit configured to obtain a first key point corresponding to a target part of the digital human; a digital human control unit configured to obtain a third key rotation matrix corresponding to the digital human according to the first key point and the second key rotation matrix, and control the digital human using a control strategy corresponding to the third key rotation matrix.
[0011] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the multi-camera-based digital human control method of any one of the preceding aspects.
[0012] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the multi-camera-based digital human control method of any one of the preceding aspects.
[0013] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of any one of the preceding aspects.
[0014] The technical solutions provided by the embodiments of the present disclosure at least have the following beneficial effects: In some or related embodiments, a set of images collected by each camera in a set of cameras for a digital human is obtained, wherein the set of cameras includes a plurality of cameras with different image collection angles; a first key rotation matrix corresponding to each camera is obtained according to the set of images; a second key rotation matrix is obtained by fusing and predicting the first key rotation matrix corresponding to each camera; a first key point corresponding to a target part of the digital human is obtained; a third key rotation matrix corresponding to the digital human is obtained according to the first key point and the second key rotation matrix, and the digital human is controlled by using a control strategy corresponding to the third key rotation matrix. Therefore, by correcting the key rotation matrix, the key rotation matrix can be obtained by decoupling the relationship between the pose and the collection angle of the camera, which can improve the accuracy of the key rotation matrix acquisition, reduce the situation that the control strategy is inaccurate due to inaccurate key rotation matrix acquisition, improve the accuracy of joint rotation prediction of the digital human, and reduce the situation that the timing stability and uniformity of the target part under multiple cameras are poor due to the pose estimation based on only the root node. The timing stability and uniformity of the target part under multiple cameras can be optimized, the probability of floating of the feet after driving the digital human can be reduced, the probability that the digital human cannot stand still can be reduced, the accuracy of digital human control can be improved, and the stability of digital human control can be improved.
[0015] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.
[0017] Figure 1 is a flowchart of a first digital human control method based on multiple cameras provided by the embodiments of the present disclosure; Figure 2 is a flowchart of a second digital human control method based on multiple cameras provided by the embodiments of the present disclosure; Figure 3 is a schematic diagram of the overall flow of a digital human control method based on multiple cameras provided by the embodiments of the present disclosure; Figure 4is a block diagram of a multi-camera-based digital human control device according to an exemplary embodiment; Figure 5 is an example schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0018] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.
[0019] The present disclosure embodiments propose a multi-camera-based digital human control method, device, electronic device and storage medium. In some embodiments, the multi-camera-based digital human control method and the information processing method, communication method, etc. can be replaced with each other, the multi-camera-based digital human control device and the information processing device, communication device, etc. can be replaced with each other, and the information processing system, communication system, etc. can be replaced with each other.
[0020] The present disclosure embodiments are not exhaustive, but only illustrate some embodiments, and are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional implementation manners in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, the steps of different embodiments or part or all of the steps of different embodiments can be combined arbitrarily, an embodiment can be combined with the optional implementation manners of other embodiments.
[0021] In the embodiments of the present disclosure, the terms and / or descriptions between the embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0022] The terms used in the embodiments of the present disclosure are only for the purpose of describing the specific embodiments, and not as a limitation on the present disclosure.
[0023] In the embodiments of the present disclosure, unless otherwise specified, the elements expressed in singular form, such as “one”, “a”, “the”, “above”, “said”, “preceding”, “this”, etc. can represent “one and only one”, or “one or more”, “at least one”, etc. For example, in the case of using articles such as “a”, “an”, “the” in English, the noun after the article can be understood as singular expression, or as plural expression.
[0024] In the embodiments of the present disclosure, "multiple" refers to two or more.
[0025] In some embodiments, the terms "at least one of," "one or more of," "a plurality of," "multiple," and the like can be replaced with each other.
[0026] The prefix words "first", "second", and the like in the embodiments of the present disclosure are merely used to distinguish different description objects, and do not constitute limitation on the position, order, priority, quantity, or content of the description objects. The description of the description objects should refer to the description in the context of the claims or embodiments, and should not constitute redundant limitation because of the use of the prefix words. For example, the description objects are "fields", and the ordinal words before "fields" in "first field" and "second field" do not limit the position or order between "fields", and "first" and "second" do not limit whether the "fields" modified thereby are in the same message or not, nor limit the order of "first field" and "second field". For another example, the description objects are "levels", and the ordinal words before "levels" in "first level" and "second level" do not limit the priority between "levels". For another example, the quantity of the description objects is not limited by the ordinal words, and can be one or more. For example, "first device", wherein the quantity of "devices" can be one or more. In addition, the objects modified by different prefix words can be the same or different, for example, the description objects are "devices", and "first device" and "second device" can be the same device or different devices, and the types thereof can be the same or different; for another example, the description objects are "information", and "first information" and "second information" can be the same information or different information, and the content thereof can be the same or different.
[0027] In some embodiments, a "terminal" or "terminal device" can be referred to as a "user equipment" (UE), a "user terminal," a "mobile station" (MS), a "mobile terminal" (MT), a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communication device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, etc.
[0028] In some embodiments, data, information, etc. can be acquired after consent from a user.
[0029] It should be noted that the terms "first", "second", and the like, herein do not necessarily have an ordinal, sequential, or chronologic implication, but are used to differentiate one element from another. It is to be understood that the terms so used in the description and claims are interchangeable under appropriate circumstances. The implementations described herein are not meant to be an exhaustive list of all implementations which can be made in accordance with the disclosure. Rather, the implementations are merely exemplary of the various ways in which the disclosure can be carried out. Other implementations can be devised without departing from the scope of the disclosure. The following examples are described in order to provide a more complete disclosure of the disclosure.
[0030] According to some embodiments, the limbs of the digital person can be driven, for example, by a single or multiple cameras. Among them, in the multi-camera-based digital person driving method, multiple pictures are directly used to estimate the human 3D key points, the joint rotation information based on the skinned multi-person linear model (SMPL), and the rotation driving is based on the orientation of a certain camera. The relationship between the camera and the human motion angle is not decoupled, which leads to the fact that while the model predicts the posture information, the output posture information is related to the camera angle and distance, and the joint rotation of the human body cannot be purely predicted, which reduces the prediction accuracy and stability. At the same time, only the root node (commonly used waist center point) is used to estimate the posture, and the timing stability and uniformity of the steps under multiple cameras are not further optimized, which further leads to the problem of floating steps of the digital person after driving, and the digital person cannot stand still.
[0031] Figure 1 is a flowchart of a first multi-camera-based digital person control method provided by the embodiments of the present disclosure, as shown in Figure 1 The multi-camera-based digital person control method can be used in a scene where the digital person is driven based on multiple cameras, and includes the following steps: In step S11, a set of images collected by each camera in a set of cameras for a digital person is obtained, wherein the set of cameras includes multiple cameras with different image collection angles. In some embodiments, the execution subject of the embodiments of the present disclosure can be an electronic device. The electronic device is not limited to a specific electronic device. For example, when the device identifier changes, the electronic device can also change accordingly. For example, when the structure of the electronic device changes, the electronic device can also change accordingly. Among them, the execution subject of the embodiments of the present disclosure can also be a server, which can be a single server or a server cluster, and the embodiments of the present disclosure do not limit it.
[0032] According to some embodiments, the camera can be a device for collecting images. The camera can collect pictures and videos, and the embodiments of the present disclosure do not limit it. The camera is not limited to a specific device. For example, when the device identifier corresponding to the camera changes, the camera can also change accordingly. For example, when the structure of the camera changes, the camera can also change accordingly. For example, different devices can collect different images. For example, images with different angles of view can be collected, and images with different resolutions can be collected. The embodiments of the present disclosure do not limit it.
[0033] In some embodiments, the camera device set may be, for example, a collection of at least one camera device. This camera device set does not specifically refer to a fixed set. For example, the camera device set may change when the number of camera devices included in the set changes. For example, the camera device set may change when a particular camera device in the set changes.
[0034] According to some embodiments, a digital human may also be referred to as a digital intelligent human. Here, a digital human can refer to a virtual human simulated through digital and artificial intelligence technologies, possessing intelligent and interactive characteristics, and capable of natural and realistic interaction with users. This digital human does not specifically refer to a particular fixed digital human. For example, when the digital human's identifier changes, the digital human may also change accordingly. For example, when the interactive capabilities of the digital human change, the digital human may also change accordingly.
[0035] In some embodiments, the image set may be, for example, a collection of at least one image captured by each camera in a set of camera devices. The at least one image included in this image set may be, for example, an image captured for a digital human. This image set is not specifically a fixed set. For example, the image set may change when the viewing angle corresponding to the set of camera devices changes. For example, the image set may also change when the acquisition time point corresponding to the image set changes.
[0036] According to some embodiments, the acquisition perspective may refer, for example, to the perspective from which each camera device captures data about the digital human. This acquisition perspective is not specifically defined by a fixed perspective. For example, when a camera device changes, its acquisition perspective may also change accordingly. For example, when a camera device receives an adjustment command for the acquisition perspective, the camera device may also change accordingly.
[0037] In some embodiments, an image set is acquired from each camera device in a camera device set for the digital human, wherein the camera device set includes multiple camera devices with different image acquisition angles.
[0038] In step S12, the first key rotation matrix corresponding to each camera device is obtained based on the image set; In some embodiments, the key rotation matrix can be, for example, a mathematical tool used in computer graphics and robotics to describe rotations in three-dimensional space. The key rotation matrix can be used to represent changes in orientation of an object or joint in three-dimensional space.
[0039] According to some embodiments, the first key rotation matrix may, for example, be a key rotation matrix obtained according to images collected by each camera. The first key rotation matrix is first used to distinguish from the remaining key rotation matrices. For example, when the image set changes, the first key rotation matrix may also change accordingly. For example, when the manner of obtaining the key rotation matrix changes, the first key rotation matrix may also change accordingly.
[0040] According to some embodiments, different cameras can correspond to different first key rotation matrices. When the image set is obtained, the first key rotation matrix corresponding to each camera can be obtained according to the image set.
[0041] In step S13, the first key rotation matrix corresponding to each camera is fused and predicted to obtain a second key rotation matrix. According to some embodiments, the second key rotation matrix may, for example, be a key rotation matrix obtained by fusing and predicting at least one key rotation matrix. The second key rotation matrix does not refer to a fixed key rotation matrix. For example, when the manner of obtaining the second key rotation matrix, i.e., the manner of fusing and predicting, changes, the second key rotation matrix may also change accordingly. For example, when the first key rotation matrix changes, the second key rotation matrix may also change accordingly.
[0042] In some embodiments, the first key rotation matrix corresponding to each camera can be fused and predicted to obtain a second key rotation matrix.
[0043] In step S14, the first key point corresponding to the target part of the digital person is obtained. According to some embodiments, the target part may, for example, be a basic part used to determine a control strategy. The target part may, for example, be a foot, and may also be determined according to a part determination instruction. The target part does not refer to a fixed part.
[0044] In some embodiments, the key point may, for example, be a point used to represent a unique feature in an image. The key point is not limited to an edge point, a corner point, a texture point, etc. Different images may correspond to different key points. Different key point obtaining manners may correspond to different key points.
[0045] In some embodiments, the first key point may, for example, be a key point corresponding to the target part of the digital person when the secondary control strategy is obtained. The first key point may, for example, be referred to as a first key point set, or at least one first key point. The first key point does not refer to a fixed key point. For example, when the time point or the obtaining manner of the first key point changes, the first key point may also change accordingly.
[0046] In some embodiments, the first key point corresponding to the target part of the digital human is obtained.
[0047] According to some embodiments, the execution order of steps S12, S13 and S14 is not limited. For example, step S12 and step S13 can be executed first, and then step S14 can be executed. For example, step S14 can be executed first, and then step S12 and step S13 can be executed. Step S14 can be executed simultaneously with step S12 and step S13.
[0048] In step S15, a third key rotation matrix corresponding to the digital human is obtained according to the first key point and the second key rotation matrix, and the digital human is controlled by using a control strategy corresponding to the third key rotation matrix.
[0049] According to some embodiments, the third key rotation matrix may, for example, be a key rotation matrix corresponding to the digital human, and the third key rotation matrix is not a fixed matrix. For example, when the first key point or the second key rotation matrix changes, the third key rotation matrix can also change accordingly.
[0050] In some embodiments, the control strategy may, for example, be used to indicate the strategy used for controlling the digital human. The control strategy is not a fixed strategy. For example, when the control strategy determination method changes, the control strategy can also change accordingly. For example, when the third key rotation matrix changes, the control strategy can also change accordingly. The control strategy can also be referred to as a driving strategy, and the present disclosure does not limit this.
[0051] According to some embodiments, the third key rotation matrix corresponding to the digital human can be obtained according to the first key point and the second key rotation matrix, and the digital human is controlled by using a control strategy corresponding to the third key rotation matrix. When the control strategy is obtained, the digital human can be controlled by using the control strategy.
[0052] In some or related embodiments, a set of images collected by each camera in a set of cameras for a digital human is obtained, wherein the set of cameras includes a plurality of cameras with different image collection perspectives; a first key rotation matrix corresponding to each camera is obtained according to the set of images; a second key rotation matrix is obtained by fusing and predicting the first key rotation matrix corresponding to each camera; a target part of the digital human corresponding to a first key point is obtained; a third key rotation matrix corresponding to the digital human is obtained according to the first key point and the second key rotation matrix, and a control strategy corresponding to the third key rotation matrix is used to control the digital human. Therefore, the key rotation matrix can be corrected, the key rotation matrix can be obtained by decoupling the relationship between the pose and the collection perspective of the camera, the accuracy of the key rotation matrix can be improved, the inaccuracy of the control strategy caused by the inaccuracy of the key rotation matrix can be reduced, the accuracy of the joint rotation prediction of the digital human can be improved, the situation that the timing stability and uniformity of the steps under the multi-camera are poor caused by the pose estimation based on only the root node can be reduced, the timing stability and uniformity of the target part under the plurality of cameras can be optimized, the probability of the floating of the feet after the driving of the digital human can be reduced, the probability that the digital human cannot stand still in place can be reduced, the accuracy of the control of the digital human can be improved, and the stability of the control of the digital human can be improved.
[0053] Figure 2 is a flowchart of a second multi-camera-based digital human control method provided by the embodiments of the present disclosure, as shown in Figure 2 , comprising the following steps: In step S21, a set of images collected by each camera in a set of cameras for a digital human is obtained, wherein the set of cameras includes a plurality of cameras with different image collection perspectives; The related processes can be as described above, and will not be described here again.
[0054] According to some embodiments, the camera of the embodiments of the present disclosure can be a camera, for example. Figure 3 is a schematic diagram of the overall process of a multi-camera-based digital human control method provided by the embodiments of the present disclosure, as shown in Figure 3 , a single image frame collected by each camera with different perspectives can be obtained, which are frame1, frame2, and framen respectively, and the images obtained by each camera are processed.
[0055] In step S22, a prediction model is used to identify the set of images to obtain first image information corresponding to each camera, wherein the image information includes joint rotation information, second three-dimensional key points, and a rotation matrix of a root node; The related processes can be as described above, and will not be described here again.
[0056] According to some embodiments, the prediction model may, for example, be a network constructed based on a multi-task prediction model, which may, without limitation, be a neural network, and may be a convolutional neural network. As shown in Figure 3 The image set may be input to the prediction model to obtain first image information. Specifically, n images in the image set may be input to the network constructed based on the multi-task prediction model, and the following outputs may be predicted: joint visibility Cn = [c1, c2, …, cn], joint rotation information w = [w1, w2, …, wn], second three-dimensional key points 3d key points P = [P1, P2, …, Pn], and a rotation matrix Rroot of a root node. The rotation matrix Rroot corresponds to each view (i.e., each camera).
[0057] In step S23, an inverse kinematics algorithm is used to calculate the joint rotation information and the second three-dimensional key points to obtain a fourth joint rotation matrix. The related processes may be as described above and will not be described again here.
[0058] The inverse kinematics (ik) algorithm is an important concept in robotics and computer graphics, which is used to calculate the joint angles of a robot or virtual character so that its end effector (such as the hand of a robot or the foot of a virtual character) reaches a specified target position and direction. In contrast to forward kinematics (Fk), which calculates the position and direction of the end effector based on known joint angles, the ik algorithm of the present disclosure may, for example, split the joint rotation calculation into swing and rotation calculations, and may input the joint axial rotation and spatial point position to the ik algorithm.
[0059] In some embodiments, the joint rotation information w and the second three-dimensional key points 3d key points P may be calculated by the ik algorithm to obtain a joint rotation matrix R, which is the fourth joint rotation matrix.
[0060] In step S24, the first joint rotation matrix and the rotation matrix of the root node are inversely rotated to obtain a first joint rotation matrix corresponding to each camera. The related processes may be as described above and will not be described again here.
[0061] According to some embodiments, the joint rotation matrix R, which is the fourth key rotation matrix, and the rotation matrix Rroot of the root node (label 0 of R) may be inversely rotated, and the specific formula is shown in formula (1): R[0] = R* (Rroot*R 反 ) T(1) wherein R 反 = [[1, 0, 0], [0, -1, 0], [0, 0, -1]] (R is obtained by the pixel point position prediction + ik algorithm, Rroot is the camera rotation about the root node directly predicted by the prediction model, and the offset of the two will obtain the root node R【0】 without camera rotation).
[0062] wherein R[0] represents the root node rotation calculated by the ik algorithm. After inverse rotation, the forward non-rotating limb motion information is obtained to eliminate the rotation direction caused by the camera angle, and a new key rotation matrix Rnew corresponding to each camera is obtained. Specifically, R【0】 in R is replaced by R【0】 calculated by S24, that is, Rnew can be obtained.
[0063] In step S25, the first key rotation matrix corresponding to each camera is fused and predicted to obtain a second key rotation matrix. wherein the related process can be as described above, which will not be repeated here.
[0064] According to some embodiments, wherein the image information further comprises joint visibility, the first key rotation matrix corresponding to each camera is fused and predicted to obtain a second key rotation matrix, comprising: According to the joint visibility value corresponding to each camera, the sum of the joint visibility values is obtained; According to the first key rotation matrix corresponding to each camera, the joint visibility value corresponding to each camera and the sum of the joint visibility values, the first key rotation matrix corresponding to each camera is weighted and averaged to obtain the second key rotation matrix. The visibility of each joint can be predicted for each camera to reduce the visual blind area in the human body driving process, and the weighted fusion of the joints is performed based on the joint visibility of multiple cameras to solve the driving fluctuation caused by occlusion and the like, thereby improving the stability and authenticity of the driving.
[0065] According to some embodiments, the Rnew and Pnew of each camera can be fused and predicted according to the joint visibility Cn, and the specific prediction scheme is as follows: the new key rotation matrix Rnew of each camera, that is, the first key rotation matrix, is multiplied by the Cn value of the corresponding camera and weighted and averaged to obtain the final joint rotation information: Rnewfinal, that is, the second key rotation matrix. The formula is shown in equation (3): Cn total = Cn1 + Cn2 + Cnn for multiple cameras; Rnewfinal = (Cn1 * Rnew1 + Cn2 * Rnew2 + Cnn * Rnewn) / Cn total (3) In step S26, the first key point corresponding to the target part of the digital human is obtained; The related process can be as described above, and details are not repeated here.
[0066] According to some embodiments, the first key point corresponding to each camera is obtained, including: The second three-dimensional key point is inversely rotated to obtain the first key point corresponding to each camera.
[0067] For example, the second three-dimensional key point 3d key point P can be inversely rotated, and formula (2) can be used for calculation to obtain Pnew.
[0068] Pnew=(Rroot *R 反 *P T ) T (2) In step S27, the third key rotation matrix corresponding to the digital human is obtained according to the first key point and the second key rotation matrix, and the control strategy corresponding to the third key rotation matrix is used to control the digital human.
[0069] The related process can be as described above, and details are not repeated here.
[0070] According to some embodiments, wherein the image information further includes joint visibility, the third key rotation matrix corresponding to the digital human is obtained according to the first key point and the second key rotation matrix, including: According to the first key point and the second key rotation matrix, the rotation amount of each joint in the joint set of the digital human is corrected to obtain the third key rotation matrix corresponding to the digital human, wherein the joint set includes the foot, the first joint and the second joint, the first joint is the upper joint adjacent to the foot, and the second joint is the upper joint adjacent to the first joint. Therefore, the rotation amount of the joint can be corrected, the accuracy of the key rotation matrix is improved, and the accuracy of the digital human control is improved.
[0071] According to some embodiments, according to the first key point and the second key rotation matrix, the rotation amount of each joint in the joint set of the digital human is corrected to obtain the third key rotation matrix corresponding to the digital human, including: The third key point corresponding to the foot of the digital human is obtained from the first key point; According to the confidence of the joint visibility, the third key point corresponding to the foot is weighted and averaged to obtain the three-dimensional key point space information corresponding to the foot; According to the second key rotation matrix, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint, the first increment of the first rotation amount and the second increment of the second rotation amount are obtained; According to the first increment and the second increment, the second key rotation matrix is iteratively processed to obtain a third key rotation matrix. Therefore, by decoupling the relationship between the 3d key points and the camera, the confidence of the point position is combined, the stable 3d key point space position coordinates are calculated, the footstep joint rotation is optimized and corrected, the driving result is coincided with the footstep space coordinates, the consistency of joint driving rotation and point position result is realized, the problems such as footstep floating in the driving process are solved, the stability of the footstep is improved, and the authenticity and accuracy of the driving are greatly improved.
[0072] According to some embodiments, according to the second key rotation matrix, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint, the first increment of the first rotation amount and the second increment of the second rotation amount are obtained, comprising: applying the second key rotation matrix to the multi-person linear model using the forward kinematics algorithm to obtain a rotated multi-person linear model; According to the rotated multi-person linear model, the space coordinates corresponding to each joint in the joint set are extracted; According to the space coordinates corresponding to the foot in the space coordinates of each joint and the three-dimensional key point space information corresponding to the foot, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint are iteratively processed to obtain the first increment of the first rotation amount and the second increment of the second rotation amount.
[0073] According to some embodiments, the first three-dimensional key point 3d key point about the footstep can be extracted from Pnew of each camera, Pnewfootstep (corresponding to the point position number of the foot, the space point position about the foot is taken out), and the confidence information of Cn is weighted and averaged to obtain the final 3d key point space information of the foot, that is, Pfoot, the formula is (4).
[0074] Pfoot= (Pnew1footstep* C1 + Pnewnfootstep* Cn) / n (4) Among them, the upper two joints of the footstep can be selected based on the joint chain, for example, the upper joint of the heel is the knee, and the upper joint of the knee is the thigh. The rotation amounts of the upper two joints are R1 and R2. For example, the first joint is the knee, and the second joint is the thigh.
[0075] According to some embodiments, based on Rnew final and R1, R2, an optimization iteration is performed, and the increments R1', R2' of R1 and R2 are superimposed on the original rotation. The specific process is as follows: Rnew final is applied to the smpl model using the fk algorithm, and the corresponding joint space coordinates are extracted according to the rotated smpl human model (extracted according to the number, and the point position coincides with the predicted point position), to obtain the space point position based on rotation, and the point position of the foot is set as Pfoot rot, the iteration target is set as Pfoot - Pfoot rot = 0, the iteration variable is R1+R1', R2+R2', and the iteration is ended after N iterations.
[0076] In some embodiments, R1' and R2' can be updated to Rnew final, that is, the corresponding joints R1, R2 in Rnew final are replaced by R1+R1', R2+R2', the optimization iteration of Rnew final is realized and output, and finally the final joint rotation information Rnew based on the optimization iteration is realized to drive the digital human.
[0077] In some embodiments, the above is a description for a single foot of a digital human, and the above method can be used for iteration for two feet.
[0078] In one or related embodiments, a prediction model can be used to identify the image set to obtain first image information corresponding to each camera, wherein the image information includes joint rotation information, second three-dimensional key points, and a rotation matrix of a root node; an iK algorithm is used to calculate the joint rotation information and the second three-dimensional key points to obtain a fourth joint rotation matrix; and inverse rotation processing is performed on the first joint rotation matrix and the rotation matrix of the root node to obtain a first joint rotation matrix corresponding to each camera. Therefore, the space point position can be predicted for each camera, and based on the ik algorithm, the rotation is solved by using the point position stability of the pixel, and the root node rotation of the model prediction is used, the two are offset, the relationship between the human joint posture and the camera angle is decoupled, the human driving is more real and stable, the influence of the camera angle is eliminated, the influence of the image collection angle of the camera on the digital human control can be reduced, and the accuracy of the digital human control can be improved.
[0079] A block diagram of a digital human control device based on multiple cameras is shown according to an example embodiment. Referring to Figure 4 The device 400 includes: The set acquisition unit 401 is configured to acquire an image set collected by each camera in a camera set for a digital human, wherein the camera set includes a plurality of cameras with different image collection angles. The matrix acquisition unit 402 is configured to acquire a first key rotation matrix corresponding to each camera according to the image set. The matrix obtaining unit 402 is further configured to perform fusion prediction on the first key rotation matrix corresponding to each camera device to obtain a second key rotation matrix. The key point obtaining unit 403 is configured to obtain a first key point corresponding to a target part of the digital human. The digital human control unit 404 is configured to obtain a third key rotation matrix corresponding to the digital human according to the first key point and the second key rotation matrix, and control the digital human by using a control strategy corresponding to the third key rotation matrix.
[0080] According to some embodiments, when the matrix obtaining unit 402 is used to obtain the first key rotation matrix corresponding to each camera device according to the image set, the matrix obtaining unit 402 is specifically configured to: identify the image set by using a prediction model to obtain first image information corresponding to each camera device, wherein the image information includes joint rotation information, second three-dimensional key points, and a rotation matrix of a root node; calculate the joint rotation information and the second three-dimensional key points by using an inverse kinematics algorithm to obtain a fourth joint rotation matrix; perform inverse rotation processing on the first joint rotation matrix and the rotation matrix of the root node to obtain the first joint rotation matrix corresponding to each camera device.
[0081] According to some embodiments, when the matrix obtaining unit 402 is used to obtain the first key point corresponding to the target part of the digital human, the matrix obtaining unit 402 is specifically configured to: perform inverse rotation processing on the second three-dimensional key points to obtain the first key point corresponding to the target part of the digital human.
[0082] According to some embodiments, wherein the image information further includes joint visibility, when the matrix obtaining unit 402 is used to obtain the second key rotation matrix by performing fusion prediction on the first key rotation matrix corresponding to each camera device, the matrix obtaining unit 402 is specifically configured to: obtain a sum of joint visibility values according to the joint visibility values corresponding to each camera device; perform weighted average processing on the first key rotation matrix corresponding to each camera device according to the first key rotation matrix corresponding to each camera device, the joint visibility values corresponding to each camera device, and the sum of the joint visibility values to obtain the second key rotation matrix.
[0083] According to some embodiments, wherein the image information further includes joint visibility, when the matrix obtaining unit 402 is used to obtain the third key rotation matrix corresponding to the digital human according to the first key point and the second key rotation matrix, the matrix obtaining unit 402 is specifically configured to: According to the first key point and the second key rotation matrix, a rotation amount of each joint in a joint set of the digital human is corrected to obtain a third key rotation matrix corresponding to the digital human, wherein the joint set includes a foot, a first joint and a second joint, the first joint is an upper joint adjacent to the foot, and the second joint is an upper joint adjacent to the first joint.
[0084] According to some embodiments, the matrix obtaining unit 402 is configured to, when correcting a rotation amount of each joint in a joint set of the digital human according to the first key point and the second key rotation matrix to obtain a third key rotation matrix corresponding to the digital human, specifically configured to: obtain a third key point corresponding to the foot of the digital human from the first key point; perform weighted average processing on the third key point corresponding to the foot according to the confidence of the joint visibility to obtain three-dimensional key point space information corresponding to the foot; obtain a first increment of the first rotation amount and a second increment of the second rotation amount according to the second key rotation matrix, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint; perform iterative processing on the second key rotation matrix according to the first increment and the second increment to obtain the third key rotation matrix.
[0085] According to some embodiments, the matrix obtaining unit 402 is configured to, when obtaining a first increment of the first rotation amount and a second increment of the second rotation amount according to the second key rotation matrix, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint, specifically configured to: apply the second key rotation matrix to the multi-person linear model using an fk algorithm to obtain a rotated multi-person linear model; extract spatial coordinates corresponding to each joint in the joint set according to the rotated multi-person linear model; perform iterative processing on the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint according to the spatial coordinates corresponding to the foot in the spatial coordinates of each joint and the three-dimensional key point space information corresponding to the foot to obtain the first increment of the first rotation amount and the second increment of the second rotation amount.
[0086] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and will not be described in details here.
[0087] In some or related embodiments, a set acquisition unit is used to acquire a set of images captured by each camera device in a set of camera devices for the digital human, wherein the set of camera devices includes multiple camera devices with different image acquisition angles; a matrix acquisition unit is used to acquire a first key rotation matrix corresponding to each camera device based on the image set; the matrix acquisition unit is also used to perform fusion prediction on the first key rotation matrix corresponding to each camera device to acquire a second key rotation matrix; a key point acquisition unit is used to acquire a first key point corresponding to the target part of the digital human; and a digital human control unit is used to acquire a third key rotation matrix corresponding to the digital human based on the first key point and the second key rotation matrix, and to control the digital human using a control strategy corresponding to the third key rotation matrix. Therefore, by correcting the key rotation matrix and decoupling the relationship between posture and the camera's acquisition perspective, the accuracy of key rotation matrix acquisition can be improved, reducing the possibility of inaccurate control strategies due to inaccurate key rotation matrix acquisition. This can improve the accuracy of joint rotation prediction in digital humans and reduce the problem of poor temporal stability and uniformity of footsteps under multiple cameras caused by posture estimation based solely on the root node. It can also optimize the temporal stability and uniformity of target parts under multiple camera devices, reduce the probability of foot floating after driving digital humans, and reduce the probability of digital humans being unable to stand stably in place, thereby improving the accuracy and stability of digital human control.
[0088] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0089] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0090] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0091] The computing unit 501 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above. For example, in some embodiments, the above-described methods can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the above-described methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the above-described methods by any other appropriate means, such as by means of firmware.
[0092] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0093] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0094] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0096] The systems and techniques described herein can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0097] The computer system can include clients and servers. The clients and the servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0098] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technical solutions disclosed in the present disclosure, and are not limited herein.
[0099] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A digital human control method based on multiple cameras, characterized in that, include: Obtain a set of images captured by each camera device in the camera device set for the digital human, wherein the camera device set includes multiple camera devices with different image acquisition angles; Based on the image set, obtain the first key rotation matrix corresponding to each camera device; The first key rotation matrix corresponding to each camera device is fused and predicted to obtain the second key rotation matrix; Obtain the first key point corresponding to the target part of the digital human; Based on the first key point and the second key rotation matrix, a third key rotation matrix corresponding to the digital human is obtained, and the digital human is controlled using the control strategy corresponding to the third key rotation matrix.
2. The method according to claim 1, characterized in that, The step of obtaining the first key rotation matrix corresponding to each camera device based on the image set includes: A prediction model is used to identify the image set and obtain the first image information corresponding to each camera device, wherein the image information includes joint rotation information, second three-dimensional key points and rotation matrix of the root node; The joint rotation information and the second three-dimensional key points are calculated using an inverse kinematics algorithm to obtain the fourth joint rotation matrix; The first joint rotation matrix and the root node rotation matrix are reverse rotated to obtain the first joint rotation matrix corresponding to each camera device.
3. The method according to claim 2, characterized in that, The step of obtaining the first key point corresponding to the target part of the digital human includes: The second three-dimensional key point is rotated in reverse to obtain the first key point corresponding to the target part of the digital human.
4. The method according to claim 2, characterized in that, in, The image information also includes joint visibility. The step of fusing and predicting the first key rotation matrix corresponding to each camera device to obtain the second key rotation matrix includes: Based on the joint visibility values corresponding to each camera device, the sum of the joint visibility values is obtained; Based on the first key rotation matrix corresponding to each camera device, the joint visibility value corresponding to each camera device, and the sum of the joint visibility values, a weighted average is performed on the first key rotation matrix corresponding to each camera device to obtain the second key rotation matrix.
5. The method according to claim 1, characterized in that, in, The image information also includes joint visibility. The step of obtaining the third key rotation matrix corresponding to the digital human based on the first key point and the second key rotation matrix includes: Based on the first key point and the second key rotation matrix, the rotation amount of each joint in the joint set of the digital human is corrected to obtain the third key rotation matrix corresponding to the digital human. The joint set includes the foot, the first joint, and the second joint. The first joint is the upper-level joint adjacent to the foot, and the second joint is the upper-level joint adjacent to the first joint.
6. The method according to claim 5, characterized in that, The step of correcting the rotation of each joint in the joint set of the digital human based on the first key point and the second key rotation matrix to obtain the third key rotation matrix corresponding to the digital human includes: Obtain the third key point corresponding to the feet of the digital human from the first key point; Based on the confidence level of the joint visibility, a weighted average is applied to the third key point corresponding to the foot to obtain the three-dimensional key point spatial information corresponding to the foot. Based on the second key rotation matrix, the first rotation amount corresponding to the first joint, and the second rotation amount corresponding to the second joint, obtain the first increment of the first rotation amount and the second increment of the second rotation amount; Based on the first increment and the second increment, the second key rotation matrix is iteratively processed to obtain the third key rotation matrix.
7. The method according to claim 6, characterized in that, The step of obtaining the first increment of the first rotation amount and the second increment of the second rotation amount based on the second key rotation matrix, the first rotation amount corresponding to the first joint, and the second rotation amount corresponding to the second joint includes: The second key rotation matrix is applied to the multi-person linear model using a forward kinematics algorithm to obtain the rotated multi-person linear model. Based on the rotated multi-person linear model, extract the spatial coordinates of each joint in the joint set; Based on the spatial coordinates of the foot and the spatial information of the three-dimensional key points of the foot in the spatial coordinates of each joint, the first rotation amount corresponding to the first joint and the second rotation amount corresponding to the second joint are iteratively processed to obtain the first increment of the first rotation amount and the second increment of the second rotation amount.
8. A digital human control device based on multiple cameras, characterized in that, include: A set acquisition unit is used to acquire a set of images captured by each camera device in the set of camera devices for the digital human, wherein the set of camera devices includes multiple camera devices with different image acquisition angles; The matrix acquisition unit is used to acquire the first key rotation matrix corresponding to each camera device based on the image set; The matrix acquisition unit is further configured to perform fusion prediction on the first key rotation matrix corresponding to each camera device to obtain the second key rotation matrix; A key point acquisition unit is used to acquire the first key point corresponding to the target part of the digital human; The digital human control unit is used to obtain the third key rotation matrix corresponding to the digital human based on the first key point and the second key rotation matrix, and to control the digital human using the control strategy corresponding to the third key rotation matrix.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multi-camera-based digital human control method as described in any one of claims 1 to 7.
10. A storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the multi-camera-based digital human control method as described in any one of claims 1 to 7.