Three-dimensional human body point cloud expansion and transmission method based on skeleton points
Through the three-dimensional human point cloud expansion and transmission method based on bone points, the problems of large bandwidth and low efficiency when transmitting three-dimensional human data in the prior art are solved, and more efficient data transmission and better reconstruction quality are achieved.
Patent Information
- Application Number
- CN202510109987.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the prior art, the required bandwidth is large due to the transmission of three-dimensional human body data, which affects the transmission speed and quality.
The three-dimensional human point cloud expansion and transmission method based on bone points is adopted. By calculating the coordinates of the three-dimensional bone point of the human body and sending it, a color point cloud of human body is generated and divided into multiple groups of point clouds to remove redundant information and form a mixed data frame for transmission.
Effectively remove background information and repeated information in multiple perspectives of the human body, reduce transmission bandwidth pressure, improve transmission efficiency and the quality of the receiving end to reconstruct the human body color point cloud.
Smart Images

Figure CN120147504A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of point cloud processing, and relates to a method for unfolding and transmitting three-dimensional human point clouds, specifically a method for unfolding and transmitting three-dimensional human point clouds based on skeletal points, which can be applied to the field of three-dimensional human virtual social interaction. Background Art
[0002] Three-dimensional human point clouds are a set of three-dimensional data obtained through three-dimensional scanning technology, RGB-D cameras, LiDAR, or other sensors that can represent the human body's external shape. Each point in the three-dimensional human point cloud has three-dimensional coordinates, usually (x, y, z), representing the spatial coordinates of different positions on the human body surface. This point cloud data can contain information such as the complete shape, posture, and movement of the human body and is widely used in multiple fields.
[0003] For three-dimensional human virtual social interaction, the basic idea of its implementation is as follows: At the sending end, color images and depth images of the human body from multiple perspectives are obtained through multiple RGB-D cameras, and these image data are encoded and compressed. Then, a remote connection is made to the receiving end and the data is transmitted. At the receiving end, the received data is decoded and the complete three-dimensional point cloud of the human body is reconstructed.
[0004] For example: The patent application with the application publication number CN115695441A and the name "Three-Dimensional Human Virtual Social Interaction System and Method Based on P2P Technology" discloses a method for transmitting human three-dimensional data. In this method, the collected color images are converted to YUV420P, the depth images are encoded for noise reduction, and the encoded color images and depth images are arranged to form a mixed data frame. This method can transmit all the data required for complete three-dimensional human body reconstruction. However, because the transmitted color images and depth images contain a lot of redundant information such as backgrounds and repeated perspectives of the human body, the required bandwidth for transmission is large, which in turn affects the speed and quality of transmitting three-dimensional data. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects existing in the above-mentioned prior art and propose a method for unfolding and transmitting three-dimensional human point clouds based on skeletal points to solve the technical problem of large required bandwidth caused by a large amount of redundant information in the transmitted data in the prior art.
[0006] To achieve the above purpose, the technical solution adopted by the present invention includes the following steps:
[0007] (1) Obtain color images and depth images of the human body from different perspectives;
[0008] Obtain color images and depth images of N perspectives through N RGB-D cameras, where N≥4;
[0009] (2) The sending end calculates the three-dimensional human body bone point coordinates and sends them.
[0010] The two-dimensional human body bone point coordinates of the j-th person in the color images captured by at least two RGB-D cameras Calculate its three-dimensional bone point coordinates, and send the human body bone point coordinate information to the receiving end, where n ∈ [2, N];
[0011] (3) The sending end generates a human body color point cloud.
[0012] Back-project each depth image, and color the textureless point cloud obtained by the back-projection corresponding to each color image to obtain the human body color point cloud corresponding to N depth images;
[0013] (4) The sending end divides the human body color point cloud.
[0014] Connect every two three-dimensional human body bone points into a bone segment, calculate the distance from each point in the human body color point cloud to each bone segment, and then divide the point to the bone segment with the closest distance to it, obtaining multiple groups of human body color point clouds with the same number as the bone segments;
[0015] (5) The sending end obtains a mixed data frame and sends it.
[0016] Map the position coordinate information and color information of each point in each group of human body color point clouds to the distance grid and color grid respectively, splice all the formed distance images and color images respectively, and then send the spliced mixed data frame to the receiving end;
[0017] (6) The receiving end obtains the three-dimensional human body point cloud transmission result.
[0018] The receiving end reconstructs the human body color point cloud through the human body bone point coordinate information sent in step (2) and the mixed data frame sent in step (5).
[0019] Compared with the prior art, the present invention has the following advantages:
[0020] The sending end of the present invention divides the human body color point cloud through the three-dimensional human body bone points, and forms a mixed data frame through the division result, which can remove the background information and the repeated information in multiple perspectives of the human body, avoid the defect of a large amount of redundant information in sending three-dimensional information in the prior art, reduce the bandwidth pressure of transmission, and thus improve the transmission efficiency and the quality of reconstructing the human body color point cloud at the receiving end. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart for implementing the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0022] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0023] Referring to Figure 1 , the present invention includes the following steps:
[0024] Step 1) Obtain the color image and depth image of the human body from different perspectives;
[0025] Obtain the color image and depth image from N perspectives through N RGB-D cameras, where N≥4;
[0026] Since the field of view of each RGB-D camera is limited, if you want to capture the complete human body data, multiple cameras are required. In this embodiment, N = 4, and four RGB-D cameras are respectively used to capture the color image and depth image of the front, left, right, and back of the human body.
[0027] Step 2) The sending end calculates the three-dimensional bone point coordinates of the human body and sends them;
[0028] The two-dimensional bone point coordinates of the jth human body in the color image captured by at least two RGB-D cameras Calculate its three-dimensional bone point coordinates and send the human bone point coordinate information to the receiving end, where n∈[2,N];
[0029] The calculation formula for the three-dimensional bone point coordinates of the human body is:
[0030]
[0031] Where, K n is the internal parameter matrix of the nth RGB-D camera, [R n1 |t n1 is the external parameter from the nth RGB-D camera to the first RGB-D camera, R n1 represents the rotation matrix, t n1 represents the translation vector, X j =(X w ,Y w ,Z w ) is the three-dimensional coordinate of the jth bone point.
[0032] In this embodiment, the two-dimensional human bone point coordinates on the color images of the front, left, and right are identified through the OpenPose human pose estimation library, and the three-dimensional bone point coordinates are calculated respectively through the bone point recognition results on the front and left, and the front and right color images. The bone point coordinate information is transmitted using the text channel of WebRTC.
[0033] Step 3) The sending end generates the human color point cloud;
[0034] Back-project each depth image, and color the textureless point cloud obtained by back-projecting each depth image with its corresponding color image to obtain the human body color point cloud corresponding to N depth images;
[0035] Calculate the spatial point P in the coordinate system of the first RGB-D camera corresponding to the pixel with coordinates (u, v) and value z in each depth image d =(P w ,P x ,P y ,P z ): And color the spatial point P w with the RGB value of the pixel at coordinates (u, v) in the color image corresponding to each depth image, where:
[0036]
[0037] Among them, represents the inverse result of K n .
[0038] Step 4) The sender divides the human body color point cloud;
[0039] Connect every two human body three-dimensional bone points into a bone segment, and after calculating the distance between each point in the human body color point cloud and each bone segment, divide the point into the bone segment closest to it to obtain multiple groups of human body color point clouds with the same number as the number of bone segments;
[0040] The calculation formula for the distance between each point in the human body color point cloud and each bone segment is:
[0041]
[0042] B - A=(x 2 -x 1 ,y 2 -y 1 ,z 2 -z 1 )
[0043] P - A=(x 0 -x 1 ,y 0 -y 1 ,z 0 -z 1 )
[0044] where t is the projection scale factor, P=(x 0 ,y 0 ,z 0 ) represents the coordinates of any point in the point cloud, A=(x 1 ,y 1 ,z1 ) and B=(x 2 , y 2 , z 2 ) are the coordinates of two 3D human skeleton points in the skeleton segment corresponding to P respectively. In this embodiment, 15 groups of segmented human color point clouds are obtained, representing different body parts of the human body.
[0045] Step 5) The sender obtains the mixed data frame and sends it;
[0046] Map the position coordinate information and color information of each point in each group of human color point clouds to the distance grid and color grid respectively, splice all the formed distance images and color images respectively, and then send the spliced mixed data frame to the receiver;
[0047] (5a) Take the center of the skeleton segment as the center of the sphere C=(C x , C y , C z ), and calculate the spherical coordinates S=(θ, φ, r) of the point with position coordinates a=(a x , a y , a z ) in the color point cloud belonging to this skeleton segment with respect to the center of the skeleton segment:
[0048]
[0049] where r is the Euclidean distance from the point to the center of the sphere, θ is the azimuth angle formed by the point with coordinates a and the positive X-axis direction on the horizontal plane, and φ is the elevation angle formed by the line connecting the point with coordinates a and the center of the sphere and the positive Z-axis direction;
[0050] (5b) Map the position coordinate information and color information of each point in each group of human color point clouds to a two-dimensional distance image and color image with a width and height of (360, 180) respectively. In this embodiment, the value ranges of the elevation angle and azimuth angle of the point are θ∈[0, 360°), φ∈[0, 180°), and their values correspond to the abscissa and ordinate in the image respectively, and the pixel value corresponds to the distance or color information from the point to the center of the sphere. Splice all the obtained distance images and color images to form a mixed data frame. In this embodiment, the width and height of the mixed data frame are (3240, 2000).
[0051] (5c) In this embodiment, the mixed data frame is sent to the receiver through the video channel of WebRTC.
[0052] Step 6) The receiver obtains the transmission result of the 3D human point cloud;
[0053] (6a) The receiving end splits the mixed data frame sent by the sending end through the video channel to obtain a depth image and a color image; at the same time, it connects the skeletal point coordinate information sent by the sending end through the text channel in the same connection manner as the sending end to obtain the same skeletal segments as the sending end;
[0054] (6c) Calculate the coordinates a=(a x ,C y ,C z ) of all three-dimensional points corresponding to the pixels in the depth image corresponding to the center C=(C x ,a y ,a z ) of the skeletal segment to obtain the human body color point cloud corresponding to the skeletal segment, where:
[0055]
[0056] where u and v are the abscissa and ordinate of the pixel points of the depth image respectively;
[0057] (6d) Reconstruct the human body color point cloud by coloring the corresponding three-dimensional points with the RGB values of the pixel points with the same pixel coordinates in the color image as those in the depth image.
Claims
1. A method for unfolding and transmitting a 3D human body point cloud based on skeleton points, characterized in that: The steps include: (1) Obtain color images and depth images of the human body from different perspectives; Obtain color images and depth images from N perspectives through N RGB-D cameras, where N ≥ 4; (2) The sending end calculates the coordinates of the three-dimensional skeleton points of the human body and sends them; The coordinates of the jth human 2D bone point in the color image captured by at least two RGB-D cameras Calculate the coordinates of its three-dimensional skeleton points and send the coordinate information of the human skeleton points to the receiving end, where n∈[2,N]; (3) The sending end generates a human body color point cloud; Back-project each depth image, and colorize the texture-free point cloud obtained by back-projection through each color image to obtain the human body color point cloud corresponding to N depth images; (4) The sending end divides the human body color point cloud; Connect every two human 3D bone points into bone segments, calculate the distance between each point in the human color point cloud and each bone segment, and then assign the point to the bone segment closest to it, to obtain multiple groups of human color point clouds with the same number of bone segments; (5) The sending end obtains the mixed data frame and sends it; The position coordinate information and color information of each point in each group of human body color point cloud are mapped to the distance grid and the color grid respectively, and all the distance images and color images formed by the mapping are spliced respectively, and then the spliced mixed data frame is sent to the receiving end; (6) The receiving end obtains the 3D human body point cloud transmission result; The receiving end reconstructs the human body color point cloud through the human body skeleton point coordinate information sent in step (2) and the mixed data frame sent in step (5).
2. The method according to claim 1, characterized in that The calculation formula for calculating the coordinates of the three-dimensional skeleton points of the human body described in step (2) is: Among them, K n is the internal parameter matrix of the nth RGB-D camera, [R n1 |t n1 ] The external reference from the nth RGB-D camera to the first RGB-D camera, R n1 represents the rotation matrix, t n1 represents the translation vector, X j =(X w ,Y w ,Z w ) is the three-dimensional coordinate of the jth bone point.
3. The method according to claim 1, characterized in that The steps of generating a human body color point cloud in step (3) are as follows: Through each depth image coordinates (u, v) value z d The pixel of the pixel calculates the spatial point P in the first RGB-D camera coordinate system corresponding to the pixel w =(P x ,P y ,P z ): The spatial point P is mapped to the RGB value of the pixel with coordinates (u, v) in the color image corresponding to each depth image. w Coloring is performed, where: in, K n The inverse result of .
4. The method according to claim 1, characterized in that: The distance between each point in the human body color point cloud and each bone segment is calculated using the following formula: BA=(x2-x1,y2-y1,z2-z1) PA=(x0-x1,y0-y1,z0-z1) Where t is the projection scale factor, P = (x0, y0, z0) represents the coordinates of any point in the point cloud, A = (x1, y1, z1) and B = (x2, y2, z2) are the coordinates of two three-dimensional human skeleton points in the skeleton segment corresponding to P.
5. The method according to claim 1, characterized in that The steps for obtaining the mixed data frame described in step (5) are as follows: (5a) The center of the bone segment is taken as the center of the sphere C = (C x ,C y ,C z ), calculate the position coordinates in the color point cloud belonging to the bone segment as a=(a x ,a y ,a z ) with respect to the spherical coordinates S = (θ, φ, r) of the center of the bone segment: Where r is the Euclidean distance from the point to the center of the sphere, θ is the azimuth angle formed by the point with coordinate a on the horizontal plane and the positive direction of the X axis, and φ is the elevation angle formed by the line connecting the point with coordinate a and the center of the sphere and the positive direction of the Z axis; (5b) The position coordinate information and color information of each point in each group of human body color point clouds are respectively mapped to a two-dimensional distance image and a color image with a width and height of (360, 180), and all the obtained distance images and color images are spliced to form a mixed data frame.
6. The method according to claim 1, characterized in that The receiving end in step (6) obtains the three-dimensional human body point cloud transmission result, and the implementation steps are as follows: (6a) The receiving end splits the mixed data frame sent by the sending end through the video channel to obtain a range image and a color image; at the same time, the receiving end connects the skeleton point coordinate information sent by the sending end through the text channel in the same connection mode as the sending end to obtain the same skeleton segment as the sending end; (6c) Calculate the center of the bone segment C = (C x ,C y ,C z ) corresponds to the coordinates of the three-dimensional points corresponding to all pixels in the distance image a=(a x ,a y ,a z ), and obtain the human body color point cloud corresponding to the bone segment, where: Among them, u and v are the horizontal and vertical coordinates of the pixel points of the depth image respectively; (6d) The color point cloud of the human body is reconstructed by coloring the corresponding three-dimensional points with the RGB values of the pixels in the color image whose pixel coordinates are the same as those in the depth image.
Citation Information
Patent Citations
Three-dimensional human body virtual social system and method based on P2P technology
CN115695441A
Method and system for dynamically reconstructing three-dimensional human body model in real time
CN108154551A
Kinectv2-based complete object real-time three-dimensional reconstruction method
CN110047144A
Action recognition method based on depth image and skeleton information
CN110263720A
3D human skeleton recognition and extraction method based on depth camera point cloud data
CN111681274A