Face and limb fusion capture method, virtual live broadcast system, equipment and medium

By using two physical cameras and a virtual camera to process images in the virtual live broadcast system, the problems of facial key point loss and bone-expression misalignment are solved, and high-precision virtual character synchronous driving is achieved.

CN120635830AActive Publication Date: 2025-09-12HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511107329.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-12
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In existing virtual live broadcast technology, facial key point loss and skeleton-expression misalignment occur frequently, resulting in key frame loss and reduced processing accuracy.

Method used

Two physical cameras are used to shoot the character's face and limbs respectively, and the images are processed by a virtual camera to perform color correction, resolution matching and distortion compensation. The facial key points and bone node coordinate data are extracted, integrated and spatially aligned to drive the live broadcast of the virtual character.

Benefits of technology

It effectively reduces keyframe loss, improves image processing accuracy, lowers facial recognition error rate, and achieves synchronous driving of full-body movements and micro-expressions, achieving the accuracy of a 100,000-level motion capture system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635830A_ABST
    Figure CN120635830A_ABST
Patent Text Reader

Abstract

The invention relates to a face and limb fusion capture method, a virtual live broadcast system, equipment and a medium, a rear end adopts a first virtual camera to collect and process a first image shot by the first camera to obtain a first video stream, and adopts a second virtual camera to collect and process a second image shot by the second camera to obtain a second video stream; the back end extracts face key point coordinate data from the first video stream, extracts skeleton node coordinate data from the second video stream, integrates the face key point coordinate data and the skeleton node coordinate data to obtain a single message body, and sends the single message body to the front end; the front end receives and analyzes the single message body to obtain a facial two-dimensional absolute pixel coordinate and a skeleton three-dimensional absolute pixel coordinate, and performs spatial registration on the facial two-dimensional absolute pixel coordinate and the skeleton three-dimensional absolute pixel coordinate to obtain a facial three-dimensional absolute pixel coordinate; the front end fuses the three-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton so as to drive the virtual character to live; and the phenomena of key frame loss and skeleton-expression dislocation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision and human-computer interaction technology, and in particular to a face and body fusion capture method, a virtual live broadcast system, a device and a medium. Background Art

[0002] With the development of the live broadcast industry, virtual live broadcast technology has emerged. The core principle of virtual live broadcast technology is to synchronize the movements of virtual images through sensor acquisition, data-driven modeling, and real-time rendering.

[0003] The relevant technology uses a monocular camera to simultaneously capture facial and body movements of a person. During the data-driven modeling stage, data conflicts are prone to occur when the AI ​​model recognizes images. For example, when a person turns his head, facial key points are easily lost, which in turn causes bone-expression misalignment (average error > 5mm) when the sensor data is integrated in Unity (3D development engine).

[0004] Currently, no effective solution has been proposed for the problems of keyframe loss and skeleton-expression misalignment in virtual live broadcasts. Summary of the Invention

[0005] Based on this, it is necessary to provide a facial and limb fusion capture method, virtual live broadcast system, equipment and medium that can reduce key frame loss and skeleton-expression misalignment to address the above technical problems.

[0006] In a first aspect, the present application provides a facial and limb fusion capture method, which is applied to a virtual live broadcast system, wherein the virtual live broadcast system includes a backend and a frontend, and the method includes:

[0007] The backend uses a first virtual camera to capture and process a first image captured by the first camera to obtain a first video stream, and uses a second virtual camera to capture and process a second image captured by the second camera to obtain a second video stream; wherein the first video stream includes a facial image of the user, and the second video stream includes an image of the user's body parts;

[0008] The backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end;

[0009] The front end receives and parses the single message body to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and performs spatial registration on the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face;

[0010] The front end fuses the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates to drive the live broadcast of the virtual character.

[0011] In some embodiments, the backend uses a first virtual camera to capture and process a first image captured by the first camera to obtain a first video stream, and uses a second virtual camera to capture and process a second image captured by the second camera to obtain a second video stream, including:

[0012] Cropping the first image according to the facial area of ​​the user, and performing color correction, resolution matching, and distortion compensation on the cropped first image to obtain the first video stream;

[0013] performing color correction, resolution matching, and distortion compensation processing on the second image to obtain the second video stream;

[0014] The first video stream and the second video stream are aligned according to timestamps.

[0015] In some embodiments, the backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end, including:

[0016] Extracting a plurality of facial key points including eyes, nose, and mouth from the first video stream, and obtaining coordinate data of the facial key points including two-dimensional coordinates and confidence levels;

[0017] Extracting multiple skeletal nodes including shoulders, hips, and knees from the second video stream to obtain skeletal node coordinate data including three-dimensional coordinates and visibility;

[0018] Defining a single message body containing multiple feature points, wherein each feature point is defined as an independent sub-message containing three-dimensional coordinates and confidence levels;

[0019] The facial key point coordinate data and the skeletal node coordinate data are integrated into the single message body through the repeated field, and the high-frequency field is set as a low-numbered field.

[0020] In some embodiments, the front end receives and parses the single message body, including:

[0021] The front end adds an incremental sequence number to each received single message body and detects whether a key frame is lost based on the sequence number; wherein the key frame includes a skeleton node data frame cut according to a preset time interval, or a facial reference point change frame;

[0022] When a key frame loss is detected, the front end triggers a retransmission request;

[0023] The backend will respond to the retransmission request, dynamically reduce the sampling accuracy of the skeleton nodes according to the bandwidth, and adjust the data sending interval within a preset time interval to meet the sending frequency of the key frames.

[0024] In some embodiments, the front end is deployed with a normalization-pixel coordinate bidirectional conversion engine; the front end receives and parses the single message body to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and spatially aligns the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face, including:

[0025] The normalization-pixel coordinate bidirectional conversion engine converts the facial key point coordinate data and the bone node coordinate data according to the image resolution of the front end to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton respectively;

[0026] The two-dimensional absolute pixel coordinates of the face are projected into a three-dimensional space based on calibration parameters of the first camera and the second camera to obtain the three-dimensional absolute pixel coordinates of the face; or the three-dimensional absolute pixel coordinates of the face are supplemented with Z-axis coordinate values ​​of the two-dimensional absolute pixel coordinates of the face based on depth estimation to obtain the three-dimensional absolute pixel coordinates of the face.

[0027] In some embodiments, the front end fuses the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates, including:

[0028] Assigning the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates to a target data object;

[0029] The target data object is serialized to generate a target sequence.

[0030] In a second aspect, the present application provides a virtual live broadcast system, comprising: an acquisition end, a back end, and a front end; wherein the acquisition end, the back end, and the front end are sequentially connected in communication;

[0031] The acquisition end includes a first camera and a second camera;

[0032] The backend creates a first virtual camera and a second virtual camera, wherein the first virtual camera is used to collect and process a first image captured by the first camera to obtain a first video stream, and the second virtual camera is used to collect and process a second image captured by the second camera to obtain a second video stream; wherein the first video stream contains a facial image of the user, and the second video stream contains an image of the user's limbs;

[0033] The backend is further configured to extract facial key point coordinate data from the first video stream, extract skeletal node coordinate data from the second video stream, integrate the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and send the single message body to the frontend;

[0034] The front end is used to receive and parse the single message body, obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and perform spatial registration on the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face;

[0035] The front end is also used to fuse the facial three-dimensional absolute pixel coordinates and the skeleton three-dimensional absolute pixel coordinates to drive the live broadcast of the virtual character.

[0036] In some embodiments, the first camera and the second camera are installed at a predetermined angle therebetween;

[0037] The focal length of the first camera is within a first preset focal length range, and the focal length of the second camera is within a second preset focal length range, wherein the first preset focal length range is smaller than the second preset focal length range;

[0038] The first camera is configured with a close-focus RGB camera with a frame rate not lower than a preset frame rate;

[0039] The second camera is equipped with a wide-angle RGB-D camera and supports SLAM spatial positioning.

[0040] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in the first aspect when executing the computer program.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0042] The aforementioned face and body fusion capture method, virtual live broadcast system, device, and medium utilize two physical cameras (a first camera and a second camera) to capture the face and body of a person, respectively, to generate two images. Two virtual cameras are then created to capture and process these two images, resulting in two videos. This mitigates conflicts between facial and body data, reduces keyframe loss, and thus reduces skeletal-expression misalignment. By connecting the virtual camera to the physical camera and backend, it acts as a frame buffer, further reducing frame drops caused by delays in feature point extraction. Furthermore, the virtual camera supports color correction, resolution matching, and distortion compensation, which can improve the accuracy of subsequent feature point extraction from the image. Therefore, this embodiment solves the keyframe loss and skeletal-expression misalignment issues encountered in related technologies, improving image processing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic diagram of the hardware structure of a virtual live broadcast system in one embodiment;

[0044] Figure 2 A schematic diagram of camera layout in one embodiment;

[0045] Figure 3 A schematic diagram of the hardware structure of a virtual live broadcast system in another embodiment;

[0046] Figure 4 1. A flowchart of a method for fusion capture of face and body parts according to an embodiment of the present invention;

[0047] Figure 5 is a flow chart of feature extraction and processing of dual-channel video streams in one embodiment;

[0048] Figure 6 This is a flowchart of a front-end receiving and parsing a single message body in one embodiment;

[0049] Figure 7 is a coordinate conversion flow chart of a normalization-pixel coordinate bidirectional conversion engine in one embodiment;

[0050] Figure 8 is a diagram of the internal structure of an electronic device in one embodiment;

[0051] Figure 9 FIG. 4 is a diagram showing the internal structure of an electronic device in another embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0053] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0054] In one embodiment, a virtual live broadcast system is provided. Figure 1 The hardware structure diagram of the virtual live broadcast system is as follows: Figure 1 As shown, the virtual live broadcast system includes: an acquisition end, a backend, and a frontend; wherein the acquisition end, the backend, and the frontend are sequentially connected in communication. The acquisition end, the backend, and the frontend are introduced below.

[0055] The acquisition end includes a first camera and a second camera. The first camera is used to capture the person's face, generating a first image. The second camera captures the person's body movements, generating a second image.

[0056] In some embodiments, Figure 2 A camera layout diagram is given. Figure 2As shown, the first camera and the second camera are installed at a preset angle (e.g., 60°). The focal length of the first camera is within a first preset focal length range (e.g., 0.5-1m), and the focal length of the second camera is within a second preset focal length range (e.g., 3-5m), wherein the first preset focal length range is smaller than the second preset focal length range. The first camera is positioned horizontally in front of the display screen and is equipped with a close-focus RGB camera with a frame rate of at least 120fps. Optionally, the first camera can also be equipped with a ring fill light. The second camera is installed at a preset top-down angle (e.g., 30°), is equipped with a wide-angle RGB-D camera, and supports SLAM spatial positioning. This arrangement makes it easy to cover the user's full range of motion, allowing for accurate capture of facial and body movements using ordinary cameras without the need for special equipment and at a low cost.

[0057] The back end can be an electronic device, in which a first virtual camera and a second virtual camera are created. The first virtual camera is used to collect and process the first image taken by the first camera to obtain a first video stream, and the second virtual camera is used to collect and process the second image taken by the second camera to obtain a second video stream; wherein the first video stream contains the user's facial image, and the second video stream contains the user's limbs image.

[0058] The backend is also used to extract facial key point coordinate data from the first video stream, extract skeletal node coordinate data from the second video stream, integrate the facial key point coordinate data and skeletal node coordinate data to obtain a single message body, and send the single message body to the front end.

[0059] The front end can be another electronic device, which is used to receive and parse a single message body, obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the bones, and spatially align the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the bones to obtain the three-dimensional absolute pixel coordinates of the face; the three-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the bones are merged to drive the live broadcast of the virtual character.

[0060] In this embodiment, the backend is deployed with OBS Studio video software, NVIDIA Broadcast software, and MediaPipe software. Two virtual cameras are created using OBS Studio to capture the video streams from two physical cameras. The video streams are then intercepted, processed, and transmitted to the frontend. NVIDIA Broadcast extracts 468 facial key points (with an error of <0.5 pixels), and MediaPipe extracts key skeletal nodes. The fusion ratio of facial and motion data is automatically adjusted based on the amplitude of movement (e.g., reducing facial weight during intense exercise). The frontend is equipped with an Nvidia 20 series or higher graphics card.

[0061] This embodiment uses two physical cameras (a first camera and a second camera) to capture the face and limbs of a person, respectively, to produce two images. Two virtual cameras are then created to capture and process these two images, resulting in two videos. This alleviates conflicts between facial and limb data, reduces keyframe loss, and thus reduces skeletal-expression misalignment. By connecting the virtual camera to the physical camera and the backend, it acts as a frame buffer, further reducing frame drops caused by delays in feature point extraction. Furthermore, the virtual camera supports color correction, resolution matching, and distortion compensation, which can improve the accuracy of subsequent feature point extraction from the image. Therefore, this embodiment solves the keyframe loss and skeletal-expression misalignment issues encountered in related technologies, improving image processing accuracy.

[0062] The virtual live broadcast system of this embodiment reduces the facial recognition error rate from 8.2% of the traditional solution to 2.1%, reduces energy consumption by 30% through heterogeneous computing (CPU + GPU collaboration), supports access to multi-platform SDKs such as ROS / Unity3D, eliminates data interference through physical isolation of dual cameras (motion + face), achieves the accuracy of a 100,000-level motion capture system (skeletal positioning error <2mm) with 2000-level equipment, develops a Unity Native plug-in to achieve synchronous driving of motion / facial data (latency <15ms), and solves the spatial coordinate system problem of ToF depth data and RGB skeleton data. The host can achieve synchronous driving of full-body movements and micro-expressions using an ordinary camera.

[0063] In one embodiment, Figure 3 Provides another hardware structure diagram of the virtual live broadcast system, such as Figure 3 As shown, the front end includes a host and a display screen, and the host is connected to the display screen. The camera layout is as follows:

[0064] The first camera and the second camera are installed at a preset angle (for example, 60 degrees) therebetween. The focal length of the first camera is within a first preset focal length range (for example, 0.5 to 1 meter), and the focal length of the second camera is within a second preset focal length range (for example, 3 to 5 meters). The first preset focal length range is smaller than the second preset focal length range.

[0065] The first camera, facing the display screen horizontally, features a close-focus RGB camera with a minimum frame rate. The second camera, mounted at a preset top-down angle (e.g., 30°), features a wide-angle RGB-D camera and supports SLAM spatial positioning. This setup allows for both high-precision macro capture and full-body coverage, ensuring seamless collaboration without blind spots. Accurately capturing facial and body movements can be achieved using standard cameras, eliminating the need for specialized equipment and maintaining a low cost.

[0066] In one embodiment, a face and body fusion capture method is provided, which is applied to Figures 1 to 3 The virtual live broadcast system of any embodiment. Figure 4 This is a flowchart of the face and body fusion capture method, as shown in Figure 4 As shown, the process includes the following steps:

[0067] In step S101, the backend uses a first virtual camera to collect and process a first image taken by a first camera to obtain a first video stream, and uses a second virtual camera to collect and process a second image taken by a second camera to obtain a second video stream; wherein, the first video stream contains a user's facial image, and the second video stream contains a user's limb image.

[0068] The backend can be an electronic device with OBS Studio video software deployed inside. Two virtual cameras can be created through OBS Studio video software to capture images from the first camera and the second camera respectively, and the captured images can be processed. After the images are adjusted to the best effect, feature extraction is performed to improve the accuracy of feature extraction.

[0069] In some embodiments, the first virtual camera crops the first image according to the user's facial area, and performs color correction, resolution matching, and distortion compensation on the cropped first image to obtain a first video stream; the second virtual camera performs color correction, resolution matching, and distortion compensation on the second image to obtain a second video stream; the first video stream and the second video stream are aligned according to the timestamps, and then subsequent feature extraction processing is continued to facilitate the alignment of the extracted skeletal nodes and facial joint points.

[0070] In step S102, the backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end.

[0071] NVIDIA Broadcast software and MediaPipe software are also deployed inside the backend to extract image features.

[0072] NVIDIA Broadcast is used to capture facial data and obtain facial key points. Specifically, NVIDIA Broadcast generates 20 facial key points (such as eyes, nose, and mouth) using the Vid2Vid Cameo model. The data structure can include coordinates (x, y) and confidence.

[0073] MediaPipe is used to capture motion data and obtain skeleton nodes. The MediaPipe coordinates are normalized to [0, 1] by default. Specifically, MediaPipe obtains 33 skeleton nodes (such as shoulder, hip, knee, and other joints) through the Holistic or Pose modules. Each skeleton node contains normalized coordinates (x, y, z) and visibility. For example, the left shoulder coordinates can be extracted using the following command:

[0074] landmarks[mp_pose.PoseLandmark.LEFT_SHOULDER.value].

[0075] By integrating facial key point coordinate data and skeleton node coordinate data into a single message body, storage space can be reduced and data transmission efficiency can be improved.

[0076] Step S103: The front end receives and parses the single message body to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and spatially aligns the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face.

[0077] The coordinates of facial key points are two-dimensional coordinates, and the coordinates of bone nodes are three-dimensional coordinates. Through coordinate system one and spatial calibration, the spatial alignment of the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the bones can be achieved, thereby unifying the two into a data structure containing three-dimensional coordinates and confidence.

[0078] In step S104, the front end fuses the 3D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton to drive the live broadcast of the virtual character.

[0079] The facial and skeleton 3D absolute pixel coordinates are assigned to a target data object. The target data object is serialized to generate a target sequence, which fuses the facial and skeleton 3D absolute pixel coordinates. The target sequence is then used to drive the virtual character live broadcast.

[0080] Specifically, when merging the 3D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton, the skeletal nodes (33) and facial key points (20) are populated into the PersonData message according to the protocol sequence, assigned to the PersonData data object, serialized into a protobuf, and ultimately transmitted to the application layer for driving. It should be noted that the functional architecture of the virtual live broadcast system can be divided into the data acquisition layer, data processing layer, transport layer, and application layer. The data acquisition layer is used to capture dual-channel images, and the data processing layer is used to obtain and process dual-channel images, sending the processing results to the transport layer, which then transmits them to the application layer for application.

[0081] The sample code for assigning a value to the PersonData data object is as follows:

[0082] person_data=PersonData()#Fill in skeleton data;

[0083] for landmark in MediaPipe_results.pose_landmarks.landmark:

[0084] person_data.pose_landmarks.add(x=landmark.x,y=landmark.y,

[0085] z=landmark.z,confidence=landmark.visibility)#Fill in facial data;

[0086] for landmark in nvidia_face_landmarks:

[0087] person_data.face_landmarks.add(x=landmark.x,y=landmark.y,

[0088] z=0.0,confidence=landmark.confidence)

[0089] person_data.timestamp_ms=current_timestamp.

[0090] The sample code for serialization to Protobuf is as follows:

[0091] serialized_data=person_data.SerializeToString().

[0092] In steps S101 to S104, two physical cameras (a first camera and a second camera) are set up to capture the face and limbs of the person, respectively, to obtain two images. Two virtual cameras are then created to capture and process these two images, resulting in two videos. This alleviates conflicts between facial and limb data, reduces keyframe loss, and thus reduces skeletal-expression misalignment. By connecting the virtual camera to the physical camera and the backend, it acts as a frame buffer, further reducing frame drops caused by delays in feature point extraction. Furthermore, the virtual camera supports color correction, resolution matching, and distortion compensation, which can improve the accuracy of subsequent feature point extraction from the image. Therefore, this embodiment solves the problems of keyframe loss and skeletal-expression misalignment in related technologies, improving image processing accuracy.

[0093] In one embodiment, Figure 5 A flowchart for feature extraction and processing of dual-channel video streams is provided, such as Figure 5 As shown, the backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end, including the following steps:

[0094] Step S201: extract multiple facial key points including eyes, nose, and mouth from the first video stream to obtain facial key point coordinate data including two-dimensional coordinates and confidence levels.

[0095] Specifically, NVIDIA Broadcast generates 20 facial key points (such as eyes, nose, mouth, etc.) through the Vid2Vid Cameo model. The data structure can contain coordinates (x, y) and confidence.

[0096] Step S202: extract multiple skeletal nodes including shoulders, hips, and knees from the second video stream to obtain skeletal node coordinate data including three-dimensional coordinates and visibility.

[0097] Specifically, MediaPipe obtains 33 skeletal nodes (such as shoulder, hip, knee, and other joints) through the Holistic or Pose modules. Each skeletal node contains normalized coordinates (x, y, z) and visibility. For example, the left shoulder coordinates can be extracted using the following command:

[0098] landmarks[mp_pose.PoseLandmark.LEFT_SHOULDER.value].

[0099] Step S203: defining a single message body including multiple feature points, wherein each feature point is defined as an independent sub-message including three-dimensional coordinates and confidence.

[0100] Based on the Potobuf protocol, 53 feature points (33 skeletal nodes + 20 facial key points) are compressed into a single message body, and each feature point is defined as an independent sub-message containing coordinates (such as x, y, z).

[0101] Step S204: Integrate the facial key point coordinate data and the skeleton node coordinate data into a single message body through the repeated field, and set the high-frequency field as a low-numbered field.

[0102] The repeated field consolidates the 53 feature points into a single message body, and frequently used fields (such as coordinate values) are assigned to low-numbered fields (for example, numbers 1 to 15). Considering that low-numbered fields require only one byte to encode (including the field number and field type), while high-numbered fields (for example, fields 16 to 2047) require two bytes, Protobuf uses varint encoding, and the tag calculation formula is (field_number << 3) | wire_type. For low-numbered fields, the entire tag can be represented in a single byte. Therefore, assigning frequently used fields to low-numbered fields and leveraging Protobuf's single-byte encoding feature reduces data transmission overhead and improves encoding efficiency. This improves transmission efficiency by over five times compared to JSON, supporting 8K@60FPS real-time streaming.

[0103] Specifically, the code example for defining a .proto file (containing composite messages for bones and faces) is as follows:

[0104] message PersonData {

[0105] message Landmark {

[0106] float x=1;

[0107] float y=2;

[0108] float z=3; / / If the facial data does not have a z-axis, it can be set to the default value;

[0109] float confidence=4;

[0110] }

[0111] repeated Landmark pose_landmarks=1; / / 33 bone nodes;

[0112] repeated Landmark face_landmarks = 2; / / 20 facial landmarks;

[0113] int64 timestamp_ms=3; / / unified timestamp;

[0114] }.

[0115] This embodiment integrates 33 skeleton node data (including 3D coordinates and visibility) from MediaPipe and 20 facial key point data (including 2D coordinates and confidence) from NVIDIA Broadcast, and defines a composite data structure through the Protobuf protocol to achieve unified encapsulation of multi-source spatiotemporal data.

[0116] In one embodiment, Figure 6 A flowchart of how the front end receives and parses a single message body is provided, such as Figure 6 As shown, the following steps are included:

[0117] Step S301: The front end adds an incremental sequence number to each received single message body and detects whether a key frame is missing based on the sequence number; wherein the key frame includes a skeleton node data frame cut at a preset time interval, or a facial reference point change frame;

[0118] Step S302: When a key frame loss is detected, the front end triggers a retransmission request;

[0119] In step S303, the backend will respond to the retransmission request, dynamically reduce the sampling accuracy of the skeleton nodes according to the bandwidth, and adjust the data sending interval within the preset time interval to meet the sending frequency of the key frames.

[0120] In this embodiment, an incremental sequence number is added to each data packet at the application layer, and packet loss is detected based on the sequence number. In this embodiment, frames corresponding to changes in the reference points of skeletal nodes and facial key points are designated as key frames. When key frame loss is detected, a retransmission request is immediately triggered, implementing key frame retransmission compensation.

[0121] On the other hand, the sampling accuracy of the skeleton nodes is dynamically reduced according to the bandwidth (for example, from extracting 30 skeleton nodes to extracting 20 skeleton nodes), and the data sending interval is adjusted between 20ms and 50ms, giving priority to ensuring the key frame frequency. When packet loss is detected, the rate is immediately reduced to 50% of the current value to achieve dynamic bit rate adjustment.

[0122] This embodiment reduces the key frame loss rate through key frame retransmission compensation and dynamic bit rate adjustment, ensuring a certain level of data integrity. In actual applications, this embodiment can maintain 95% data integrity under a 20% network packet loss rate.

[0123] In one embodiment, the front end is deployed with a normalization-pixel coordinate bidirectional conversion engine. Figure 7 The coordinate conversion flow chart of the normalized-pixel coordinate bidirectional conversion engine is provided, such as Figure 7 As shown, the front end receives and parses the single message body, obtains the 2D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton, and spatially aligns the 2D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton to obtain the 3D absolute pixel coordinates of the face, including the following steps:

[0124] Step S401: The normalization-pixel coordinate bidirectional conversion engine converts the facial key point coordinate data and the bone node coordinate data according to the front-end image resolution to obtain the facial two-dimensional absolute pixel coordinates and the bone three-dimensional absolute pixel coordinates respectively.

[0125] (1) Formula for converting normalized coordinates to absolute pixel coordinates:

[0126] pixel_coord=normalized_coord×image_resolution;

[0127] (2) Formula for converting absolute pixel coordinates to normalized coordinates:

[0128] normalized_coord=pixel_coord÷image_resolution;

[0129] Among them, pixel_coord represents the absolute pixel coordinate, normalized_coord represents the normalized coordinate, and image_resolution represents the image resolution.

[0130] In some application scenarios, a target sequence needs to synchronously drive virtual characters across multiple live streaming platforms. This means a single backend communicates with multiple front-end devices, which may have varying image resolutions, leading to device incompatibility. The normalized-to-pixel coordinate bidirectional conversion engine of this embodiment is compatible with hybrid data processing using MediaPipe's [0, 1] normalized coordinate system and NVIDIA's relative coordinate system, achieving cross-device consistency through dynamic image resolution adaptation.

[0131] Step S402: Project the 2D absolute pixel coordinates of the face into 3D space based on the calibration parameters of the first camera and the second camera to obtain the 3D absolute pixel coordinates of the face; or, supplement the Z-axis coordinate values ​​of the 2D absolute pixel coordinates of the face based on depth estimation to obtain the 3D absolute pixel coordinates of the face.

[0132] When fusing 3D data (such as the MediaPipe z-axis), you can project 2D facial coordinates into 3D space by calibrating camera parameters, or use MediaPipe's depth estimation to supplement the z value. For NVIDIA 2D facial data, use the MediaPipe depth estimation model to generate pseudo 3D coordinates, or perform stereoscopic reconstruction based on dual-camera calibration parameters.

[0133] Among them, supplementing the Z-axis coordinate value of the two-dimensional absolute pixel coordinate of the face based on depth estimation can be achieved through the following steps:

[0134] (1) 3DMM model matching: Match the 2D absolute pixel coordinates of the face with the preset 3DMM model, and solve the camera pose parameters through the PnP algorithm;

[0135] (2) Depth interpolation: Based on the matching results and camera parameters, the depth of the facial area is interpolated, and the corresponding Z-axis depth value is added to the two-dimensional absolute pixel coordinates of each face.

[0136] The sample code for 3DMM model matching and depth interpolation is as follows:

[0137] def estimate_face_depth_3dmm(face_2d_points,face_3d_model):

[0138] """

[0139] Estimate facial depth based on 3DMM model;

[0140] """

[0141] / / 1. Solve the camera posture through the PnP algorithm;

[0142] camera_matrix=get_camera_intrinsics()

[0143] success,rotation_vector,translation_vector=cv2.solvePnP(

[0144] face_3d_model.points,face_2d_points,camera_matrix,None);

[0145] / / 2. Project the 3D model to obtain the depth value;

[0146] projected_points=cv2.projectPoints(

[0147] face_3d_model.points,rotation_vector,translation_vector,camera_matrix,None);

[0148] / / 3. Interpolation calculation of the depth of each point on the face;

[0149] depth_map=interpolate_depth(face_2d_points, projected_points)

[0150] return depth_map.

[0151] In some embodiments, the configuration parameters of the virtual live broadcast system are as follows:

[0152] Camera configuration:

[0153] First camera (face camera): focal length 0.3-1.2 meters, frame rate ≥ 60fps, resolution ≥ 1080p;

[0154] Second camera (limb camera): focal length 2-6 meters, viewing angle ≥120°, supports depth detection;

[0155] Camera angle: 45°~90°, preferably 60°.

[0156] Data processing parameters:

[0157] Facial key points: 68-468 feature points, detection confidence ≥ 0.8;

[0158] Skeleton nodes: 25-33 joint points, visibility threshold ≥ 0.5;

[0159] Time synchronization accuracy: ±5 milliseconds;

[0160] Data transmission frequency: 30-120fps, adaptive adjustment.

[0161] Performance indicators:

[0162] End-to-end latency: ≤20 milliseconds;

[0163] Feature point positioning accuracy: ±2 mm;

[0164] Key frame loss rate: ≤1%.

[0165] In one embodiment, an electronic device is provided. The electronic device may be a backend server, and its internal structure diagram may be as follows: Figure 8As shown. The electronic device includes a processor, a memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store dual-channel video data. The network interface of the electronic device is used to communicate with an external terminal (such as a front-end live broadcast device) via a network connection. When the computer program is executed by the processor, a method for capturing facial and limb fusion is implemented.

[0166] In one embodiment, another electronic device is provided. The electronic device may be a front-end live broadcast device, and its internal structure diagram may be as shown in FIG. Figure 9 As shown. The electronic device includes a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal via wired or wireless communication. The wireless communication method can be achieved through Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for fusion capture of face and body parts. The display of the electronic device can be a liquid crystal display or an electronic ink display. The input device of the electronic device can be a touch layer covering the display, or keys, a trackball, or a touchpad provided on the electronic device housing, or an external keyboard, touchpad, or mouse.

[0167] Those skilled in the art will understand that Figure 8 and Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0169] The backend uses a first virtual camera to capture and process a first image captured by the first camera to obtain a first video stream, and uses a second virtual camera to capture and process a second image captured by the second camera to obtain a second video stream; wherein the first video stream includes a facial image of the user, and the second video stream includes an image of the user's body parts;

[0170] The backend extracts facial key point coordinate data from the first video stream and skeleton node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeleton node coordinate data to obtain a single message body, and sends the single message body to the frontend;

[0171] The front end receives and parses the single message body, obtains the 2D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton, and performs spatial registration on the 2D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton to obtain the 3D absolute pixel coordinates of the face;

[0172] The front end fuses the 3D absolute pixel coordinates of the face and the 3D absolute pixel coordinates of the skeleton to drive the live broadcast of the virtual character.

[0173] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0174] Cropping the first image according to the user's facial area, and performing color correction, resolution matching, and distortion compensation on the cropped first image to obtain a first video stream;

[0175] Performing color correction, resolution matching, and distortion compensation on the second image to obtain a second video stream;

[0176] The first video stream and the second video stream are aligned according to the timestamps.

[0177] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0178] Extracting multiple facial key points including eyes, nose, and mouth from the first video stream to obtain facial key point coordinate data including two-dimensional coordinates and confidence levels;

[0179] Extracting multiple skeletal nodes including shoulders, hips, and knees from the second video stream to obtain skeletal node coordinate data including three-dimensional coordinates and visibility;

[0180] Define a single message body containing multiple feature points, where each feature point is defined as an independent sub-message containing three-dimensional coordinates and confidence;

[0181] The facial key point coordinate data and the bone node coordinate data are integrated into a single message body through the repeated field, and the high-frequency fields are set as low-numbered fields.

[0182] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0183] The front end adds an incremental sequence number to each received single message body and detects whether there is a key frame loss based on the sequence number; wherein the key frame includes the skeleton node data frame cut at a preset time interval, or the facial reference point change frame;

[0184] When a key frame loss is detected, the front-end triggers a retransmission request;

[0185] Among them, the backend will respond to the retransmission request, dynamically reduce the sampling accuracy of the skeleton node according to the bandwidth, and adjust the data sending interval within the preset time interval to meet the sending frequency of the key frame.

[0186] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0187] The normalization-pixel coordinate bidirectional conversion engine converts the facial key point coordinate data and the bone node coordinate data according to the front-end image resolution, and obtains the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the bones respectively;

[0188] The 2D absolute pixel coordinates of the face are projected into 3D space based on calibration parameters of the first camera and the second camera to obtain the 3D absolute pixel coordinates of the face; alternatively, the Z-axis coordinate values ​​of the 2D absolute pixel coordinates of the face are supplemented based on depth estimation to obtain the 3D absolute pixel coordinates of the face.

[0189] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0190] Assigning the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates to the target data object;

[0191] Serialize the target data object to generate a target sequence.

[0192] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0193] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0194] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0195] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for capturing face and body fusion, characterized in that: Applied to a virtual live broadcast system, the virtual live broadcast system includes a backend and a frontend, and the method includes: The backend uses a first virtual camera to capture and process a first image captured by the first camera to obtain a first video stream, and uses a second virtual camera to capture and process a second image captured by the second camera to obtain a second video stream; wherein the first video stream includes a facial image of the user, and the second video stream includes an image of the user's body parts; The backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end; The front end receives and parses the single message body to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and performs spatial registration on the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face; The front end fuses the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates to drive the live broadcast of the virtual character.

2. The face and body fusion capture method according to claim 1, characterized in that: The backend uses a first virtual camera to collect and process a first image captured by the first camera to obtain a first video stream, and uses a second virtual camera to collect and process a second image captured by the second camera to obtain a second video stream, including: Cropping the first image according to the facial area of ​​the user, and performing color correction, resolution matching, and distortion compensation on the cropped first image to obtain the first video stream; performing color correction, resolution matching, and distortion compensation processing on the second image to obtain the second video stream; The first video stream and the second video stream are aligned according to timestamps.

3. The face and body fusion capture method according to claim 1, characterized in that: The backend extracts facial key point coordinate data from the first video stream, extracts skeletal node coordinate data from the second video stream, integrates the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and sends the single message body to the front end, including: Extracting a plurality of facial key points including eyes, nose, and mouth from the first video stream to obtain coordinate data of the facial key points including two-dimensional coordinates and confidence levels; Extracting multiple skeletal nodes including shoulders, hips, and knees from the second video stream to obtain skeletal node coordinate data including three-dimensional coordinates and visibility; Defining a single message body containing multiple feature points, wherein each feature point is defined as an independent sub-message containing three-dimensional coordinates and confidence levels; The facial key point coordinate data and the skeletal node coordinate data are integrated into the single message body through the repeated field, and the high-frequency field is set as a low-numbered field.

4. The face and body fusion capture method according to claim 3, characterized in that: The front end receives and parses the single message body, including: The front end adds an incremental sequence number to each received single message body and detects whether a key frame is lost based on the sequence number; wherein the key frame includes a skeleton node data frame cut according to a preset time interval, or a facial reference point change frame; When a key frame loss is detected, the front end triggers a retransmission request; The backend will respond to the retransmission request, dynamically reduce the sampling accuracy of the skeleton nodes according to the bandwidth, and adjust the data sending interval within a preset time interval to meet the sending frequency of the key frames.

5. The face and body fusion capture method according to claim 1, characterized in that: The front end is deployed with a normalization-pixel coordinate bidirectional conversion engine; the front end receives and parses the single message body, obtains the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and spatially aligns the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face, including: The normalization-pixel coordinate bidirectional conversion engine converts the facial key point coordinate data and the bone node coordinate data according to the image resolution of the front end to obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton respectively; The two-dimensional absolute pixel coordinates of the face are projected into a three-dimensional space based on calibration parameters of the first camera and the second camera to obtain the three-dimensional absolute pixel coordinates of the face; or the three-dimensional absolute pixel coordinates of the face are supplemented with Z-axis coordinate values ​​of the two-dimensional absolute pixel coordinates of the face based on depth estimation to obtain the three-dimensional absolute pixel coordinates of the face.

6. The face and body fusion capture method according to claim 1, characterized in that: The front end fuses the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates, including: Assigning the facial 3D absolute pixel coordinates and the skeleton 3D absolute pixel coordinates to a target data object; The target data object is serialized to generate a target sequence.

7. A virtual live broadcast system, characterized in that: include: An acquisition end, a back end, and a front end; wherein the acquisition end, the back end, and the front end are sequentially communicatively connected; The acquisition end includes a first camera and a second camera; The backend creates a first virtual camera and a second virtual camera, wherein the first virtual camera is used to collect and process a first image captured by the first camera to obtain a first video stream, and the second virtual camera is used to collect and process a second image captured by the second camera to obtain a second video stream; wherein the first video stream contains a facial image of the user, and the second video stream contains an image of the user's limbs; The backend is further configured to extract facial key point coordinate data from the first video stream, extract skeletal node coordinate data from the second video stream, integrate the facial key point coordinate data and the skeletal node coordinate data to obtain a single message body, and send the single message body to the frontend; The front end is used to receive and parse the single message body, obtain the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton, and perform spatial registration on the two-dimensional absolute pixel coordinates of the face and the three-dimensional absolute pixel coordinates of the skeleton to obtain the three-dimensional absolute pixel coordinates of the face; The front end is also used to fuse the facial three-dimensional absolute pixel coordinates and the skeleton three-dimensional absolute pixel coordinates to drive the live broadcast of the virtual character.

8. The virtual live broadcast system according to claim 7, characterized in that: The first camera and the second camera are installed at a preset angle; The focal length of the first camera is within a first preset focal length range, and the focal length of the second camera is within a second preset focal length range, wherein the first preset focal length range is smaller than the second preset focal length range; The first camera is configured with a close-focus RGB camera with a frame rate not lower than a preset frame rate; The second camera is equipped with a wide-angle RGB-D camera and supports SLAM spatial positioning.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Wear-free virtual live streaming method and device based on common camera

    CN112102451A

  • Virtual character animation generation method and device, storage medium and terminal

    CN114219878A

  • Method for controlling actions and facial expressions of digital human in meta universe and live broadcast

    CN115914660A

  • Human body action reconstruction system and method based on multi-modal input

    CN115937432A

  • Virtual character rendering method and device, equipment, storage medium and program product

    CN117097919A

Cited By

  • Interaction method and system based on Web front end and electronic equipment

    CN121349460A