A sparse-view multi-body joint reconstruction method based on dynamic graph convolutional network

By using RGB cameras to collect data from sparse perspectives, combining human body parameter model and dynamic graph convolution network, the limited number of cameras, occlusion and incomplete viewing angles in multi-person reconstruction are solved, and high-precision multi-person reconstruction and behavioral analysis are achieved to meet the real-time needs of home monitoring.

CN120259572BActive Publication Date: 2025-08-19HANGZHOU DIANZI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510748470.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-19
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Under sparse perspective conditions, the multi-person body reconstruction method has the problems of limited number of cameras and inability to obtain complete observation information, serious occlusion in multi-person interactive scenarios, lack of effective spatial correlation modeling, and existing 2D completion methods cannot handle 3D occlusion and missing view angles, and inaccurate initial pose estimation.

Method used

Data is collected using a small number of RGB cameras, combined with human body parameter model SMPL-X and 3D Gaussian splash reconstruction technology, data completion is performed through dynamic graph convolution network, initial posture is optimized, multi-person posture and motion recognition is realized, and Gaussian feature ball representation method is designed for adaptive densification, overcoming the problems of occlusion and incomplete viewing angle.

Benefits of technology

It significantly reduces the hardware requirements of the home monitoring system, improves the accuracy and stability of multi-person reconstruction, can handle multi-person interactions in complex scenarios, and achieves high-precision real-time reconstruction and behavioral analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259572B_ABST
    Figure CN120259572B_ABST
Patent Text Reader

Abstract

The present invention discloses a sparse-view multi-human joint reconstruction method based on a dynamic graph convolutional network. The method first uses an RGB camera to realize the collection of indoor RGB video data. Secondly, based on the RGB video data, a human body parameter model estimation algorithm is used to estimate the 3D posture of the human body, and a completion algorithm is used to generate and predict it to complete the reconstruction of the human body model. Finally, the real-time classification, positioning and tracking of multiple human body movements are realized through an action recognition algorithm. Based on the reconstructed human body model and human body 3D posture, the behavior patterns and interaction intentions of family members are analyzed. The present invention realizes a more flexible dynamic representation of the human body, overcomes the problem that the initial posture estimation may be inaccurate, and significantly improves the accuracy and stability of multi-human reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and three-dimensional reconstruction, and specifically relates to a method for collecting data based on a sparse-view camera and realizing multi-human joint reconstruction, posture estimation, motion recognition and tracking by utilizing a dynamic graph convolutional network. Background Art

[0002] In recent years, with the rapid development of deep learning and computer vision technologies, intelligent home monitoring systems have become increasingly popular. Human monitoring is a core component of these systems, including posture tracking, behavior understanding, and intention analysis. While significant breakthroughs have been made in areas such as human reconstruction and behavior recognition, supported by artificial intelligence, existing technologies still face numerous challenges.

[0003] Traditional human reconstruction methods primarily rely on multi-camera surround capture systems, such as marker-based motion capture systems (Motion Capture) or multi-camera array systems (Camera Rig). While these methods offer high reconstruction accuracy, they require stringent equipment quantity and deployment requirements, resulting in high costs and difficulty in practical deployment in typical home environments. In contrast, single-camera reconstruction methods, while simple to deploy and inexpensive, often experience significant degradation in reconstruction quality when faced with complex multi-person scenes, occlusions, and human interaction, making them difficult to meet practical application requirements. Therefore, as a compromise, utilizing a small number of cameras for smart home surveillance has become a practical and viable technical approach.

[0004] Significant progress has been made in human reconstruction technology using parametric models such as the Skinned Multi-Person Linear Model (SMPL) and its improved version, SMPL-X. These models represent human shape and pose using a predefined set of parameters, enabling the inference of a relatively accurate 3D human model from a small number of images. However, these methods still face challenges in handling complex scenes such as multi-person interactions and partial occlusions. This is particularly true in home surveillance environments with a limited number of cameras and sparse viewing angles, where performance and stability are difficult to guarantee.

[0005] 3D Gaussian Splatting has recently demonstrated great potential in scene reconstruction and rendering. Using three-dimensional Gaussian functions as the fundamental element of scene representation, this technique can efficiently represent and render complex scenes while providing real-time updates. Compared to traditional voxel or mesh representations, 3D Gaussian Splatting offers significant advantages in rendering speed and detail. However, its application to dynamic human reconstruction, particularly in scenarios involving multiple people with mutual occlusion, lacks a comprehensive and effective supervisory signal, and achieving complete reconstruction of multiple people in interactive scenarios remains a pressing technical challenge. Summary of the Invention

[0006] To address the above problems, the present invention provides a sparse-view multi-body joint reconstruction method for home intelligent monitoring. This method uses a small number of RGB cameras to collect indoor video data from different perspectives, and transmits the data to a central control computer via a wireless or wired network. The SMPL-X human parameter model estimation algorithm combined with 3D Gaussian splash reconstruction technology is used to realize multi-body posture, movement and behavior recognition. A completion algorithm based on a dynamic graph convolutional network is used to generate and supplement the missing data under sparse perspectives, thereby overcoming the problems of occlusion and incomplete perspective.

[0007] In the existing technology, the joint reconstruction of multiple human bodies under sparse viewing conditions mainly has the following problems:

[0008] 1. The limited number of cameras prevents the acquisition of complete multi-angle observation information, resulting in poor human body reconstruction quality;

[0009] 2. In multi-person interaction scenarios, there is a serious problem of people occluding each other, making it difficult to effectively distinguish and reconstruct each individual.

[0010] 3. Traditional methods lack effective modeling of human interactions and are unable to fully utilize the spatial correlation information between human bodies.

[0011] 4. Existing 2D completion methods cannot effectively handle occlusion and perspective loss in 3D space;

[0012] 5. The initial pose estimation results may be inaccurate and need to be optimized and adjusted during the reconstruction process.

[0013] In order to solve the above problems, the technical solution of the present invention is:

[0014] A sparse-view multi-human joint reconstruction method based on a dynamic graph convolutional network, the method comprising:

[0015] (1) Data collection step: using a small number of RGB cameras to collect indoor RGB video data;

[0016] (2) Multi-body joint reconstruction step: using the human body parameter model estimation algorithm to estimate the human body's 3D posture, using the 3D Gaussian eigensphere representation method, based on RGB video data, to perform 3D dynamic reconstruction and posture estimation of multiple interactive human bodies indoors, and synchronously optimize the initial posture data during the reconstruction process;

[0017] (3) In the missing data generation step, a completion algorithm based on a dynamic graph convolutional network is used to directly perform feature fusion on the connection relationship graph of the Gaussian feature ball, generate and predict missing data under sparse perspectives, and realize the construction of a complete human body model;

[0018] (4) Action recognition and tracking steps: using action recognition algorithms to achieve real-time classification, positioning, and tracking of multiple human actions;

[0019] (5) Behavior analysis and intention understanding step: Based on the reconstructed human body model and human 3D posture, the behavior patterns and interaction intentions of family members are analyzed.

[0020] Preferably, the multi-body joint reconstruction step specifically includes:

[0021] (2.1) Based on the RGB video data, we obtain an input image from each viewpoint and use a deep convolutional neural network to extract 2D keypoint information and segmentation information of the human body. Due to mutual occlusion in interactive scenes, the segmentation information may be incomplete. "1" indicates a human pixel, and "0" indicates a non-human pixel.

[0022] (2.2) Based on the detected 2D key point information, the SMPL-X parameter fitting algorithm is used to estimate the 3D posture parameters and shape parameters of the human body.

[0023] (2.3) The estimated SMPL-X shape parameters are converted into a human body patch mesh representation, where each vertex of the human body mesh representation is associated with a 3D Gaussian feature ball, thereby converting the human body mesh representation into a 3D Gaussian feature ball representation, where each Gaussian feature ball contains a learnable dynamic geometric feature, a learnable dynamic appearance feature and a learnable semantic feature, the geometric feature predicts its geometric properties (density, scale, angle) through a geometric property decoder, the appearance feature predicts its color properties through an appearance attribute decoder, and the semantic feature predicts its human ID properties through a semantic decoder.

[0024] (2.4) The aforementioned geometric and appearance attributes are rendered into a 2D image representation from the target perspective using a differentiable Gaussian sputtering rendering technique. The aforementioned human ID attributes are rendered into human segmentation information from the input image using a differentiable Gaussian sputtering rendering technique. The rendered 2D image representation and human segmentation information are compared with the input image and the estimated human segmentation information from the input image to calculate the rendering error.

[0025] (2.5) Based on the rendering error calculated above, a Gaussian eigensphere adaptive densification algorithm is proposed. For regions with large loss, the Gaussian eigenspheres in the corresponding regions are adaptively densified based on the loss magnitude to improve the reconstruction quality of the corresponding regions. The added Gaussian eigenspheres are initialized using the features of the closest Gaussian eigensphere and associated with the vertex in the closest human mesh representation.

[0026] Gaussian characteristic sphere adaptive densification algorithm, the implementation process is as follows:

[0027] First, initialize the Gaussian ball set, and set the error threshold ε and the split ratio upper limit r_max;

[0028] For each Gaussian ball g in the Gaussian ball set, calculate the rendering error e;

[0029] When the rendering error e is greater than the error threshold ε, the split ratio r = min(e / ε, r_max) is calculated;

[0030] Split the Gaussian ball g into r Gaussian balls and update the Gaussian ball set.

[0031] (2.6) Through the joint iterative optimization of steps 2.4 and 2.5 under multiple viewpoints, the geometric features, color features, and semantic features in the Gaussian sphere and their corresponding attribute decoders are learned, and the 3D pose parameters of the human body are optimized at the same time, the inaccuracy in the initial human pose estimation is corrected, and the human body segmentation capability and reconstruction quality in 3D space are enhanced.

[0032] Preferably, the missing data generating step specifically includes:

[0033] (3.1) Based on the human body mesh representation described in (2.3), the topological structure of the mesh is used to construct a connection relationship diagram between Gaussian feature balls.

[0034] (3.2) The mesh representation of each human body is used to obtain complete human body segmentation information through mesh rendering technology. This segmentation information is "ANDed" with the incomplete segmentation information in (2.1) (i.e., the human body segmentation information of the input image) pixel by pixel to obtain a mask of the visible area of the human body; the incomplete segmentation information in (2.1) is negated and "ANDed" with the complete human body segmentation information to obtain a mask of the invisible area of the human body.

[0035] (3.3) Projecting the visible area mask and the invisible area mask onto the 3D Gaussian feature sphere representation through inverse camera transformation projection, thereby obtaining a visible area Gaussian feature sphere and an invisible area Gaussian feature sphere.

[0036] (3.4) Construct a dynamic graph convolutional neural network (DGCN) and perform dynamic graph convolution calculations on the Gaussian feature ball connection relationship graph described in (3.1). Use the geometric features, color features, and semantic features of the Gaussian feature balls in the visible area to predict the corresponding features of the Gaussian feature balls in the invisible area.

[0037] (3.5) The geometric attribute decoder and color attribute decoder described in (2.3) decode the geometric and color features of the completed Gaussian feature sphere into complete geometric and color attributes, thereby achieving high-fidelity complete reconstruction under sparse viewing angles.

[0038] Preferably, the missing data generation step also includes a time domain consistency constraint, so that the geometric and texture attributes predicted by each decoder (MLP) maintain spatiotemporal consistency under multiple perspectives and multiple time points, increase the completion constraints, and ensure the rationality of the generated completion.

[0039] Preferably, the action recognition and tracking steps adopt an algorithm based on a combination of spatiotemporal interest points and dense trajectories to achieve accurate recognition and real-time tracking of multiple human actions in complex scenes.

[0040] Preferably, the behavior analysis and intention understanding step includes behavior pattern mining based on time series data and multi-person interaction relationship modeling, which can identify the interactive behaviors and intentions among family members.

[0041] The method, system, and device for sparse-view multi-body joint reconstruction based on a dynamic graph convolutional network provided by the present invention have the following beneficial effects:

[0042] The sparse viewpoint data collection method significantly reduces the hardware requirements and deployment costs of the home monitoring system, and effective coverage of the home environment can be achieved with 3-5 cameras.

[0043] An innovative Gaussian feature ball representation method is designed. Different from the original Gaussian ball representation, the Gaussian ball representation in the method is replaced by a Gaussian feature ball representation. The Gaussian feature ball contains geometric features, color features and semantic features, and uses the corresponding geometric attribute decoder, color attribute decoder and semantic decoder to predict geometric attributes, color attributes and human ID attributes, thereby achieving more flexible human body dynamic representation.

[0044] By synchronously optimizing the SMPL-X pose parameters during the reconstruction process, the problem of inaccurate initial pose estimation is overcome, and the accuracy and stability of multi-body reconstruction are significantly improved.

[0045] By utilizing the incomplete segmentation information of each human body as a supervisory signal, combined with Gaussian sputtering differentiable rendering technology, and the proposed reconstruction loss-guided adaptive Gaussian feature sphere densification technology, the accuracy of human body segmentation in 3D space is achieved, significantly improving the reconstruction quality of multi-person interaction scenes.

[0046] Based on the dynamic graph convolutional network, the convolution operation is performed directly on the connection relationship graph of the Gaussian feature balls. The features of the Gaussian feature balls in the visible area are used to predict the features of the Gaussian feature balls in the invisible area. No UV mapping intermediate step is required, which simplifies the processing flow and improves efficiency, effectively solving the problem of data missing under sparse perspective.

[0047] A behavior analysis system for multi-person interaction scenarios has been designed, which can understand the interaction intentions between family members and provide technical support for applications such as home safety monitoring and elderly health monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A flowchart for multi-body joint reconstruction and missing data generation of the present invention;

[0049] Figure 2 Schematic diagram of the combination of the human body parameter model SMPL-X and the 3D Gaussian characteristic sphere of the present invention;

[0050] Figure 3 This is a diagram of the data generation module architecture based on the dynamic graph convolutional network of the present invention;

[0051] Figure 4 This is a flowchart of the multi-person interaction behavior analysis of the present invention. DETAILED DESCRIPTION

[0052] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] like Figure 1 As shown in the figure, the present invention proposes a sparse view multi-body joint reconstruction method based on dynamic graph convolutional network, which is specifically implemented as follows:

[0054] Example 1: Data Collection and Transmission

[0055] The data acquisition system in this embodiment consists of 3-5 fixed-position RGB cameras. The cameras have a resolution of 1920×1080 pixels and a frame rate of 30 fps. The cameras are arranged according to the principle of "few but good," primarily covering key activity areas such as the living room and hallways. Key areas are also covered by at least two cameras with overlapping viewpoints to facilitate subsequent multi-view data fusion processing.

[0056] Data transmission supports both wireless (Wi-Fi 6, 802.11ax) and wired (Gigabit Ethernet) networks. Raw video captured by the camera is compressed using H.265 encoding and transmitted to a central control computer via the network. The transmission module uses the Real-Time Streaming Protocol (RTSP) to ensure low-latency transmission of video data, with an average transmission delay of less than 30ms.

[0057] Example 2: Joint reconstruction of multiple human bodies

[0058] like Figure 1 and Figure 2 As shown in the figure, joint reconstruction of multiple human bodies is one of the core innovations of the present invention. Through the innovative 3D Gaussian feature ball representation and reconstruction error-guided adaptive Gaussian feature ball densification technology, high-precision real-time reconstruction of multiple human bodies under sparse perspective is achieved.

[0059] The first step is to use a deep convolutional neural network to extract 2D human body key points and segmentation information for each human body from each perspective. This embodiment obtains human body segmentation information and locates 2D human body key points, identifying the positions of 23 key nodes for each human body.

[0060] In the second step, based on the 2D human key points, the SMPL-X human parameter fitting algorithm is used to estimate the 3D posture parameters and shape parameters of the human body. The SMPL-X model is defined as follows:

[0061]

[0062] in, is the shape parameter, is the attitude parameter, is the facial expression parameter, is the mixing weight, is the 3D key point, T is the human body mesh representation, is a linear blend skinning function. The parameters are estimated by optimizing the following objective function:

[0063] in, is the key point projection error, and are the prior constraints on shape parameters and pose parameters, 、 and is the loss weight coefficient. It should be noted that the initial SMPL-X pose parameters may be inaccurate, especially in scenes with occlusion and complex poses. Therefore, during the subsequent reconstruction process, we will simultaneously optimize these pose parameters to obtain more accurate reconstruction results.

[0064] The third step is to convert the estimated SMPL-X human body parameter model into a 3D Gaussian feature ball representation Specifically, a 3D Gaussian feature ball is initialized for each vertex in the human body mesh representation in the SMPL-X model corresponding to each human body. Different from the 3D Gaussian ball in the original Gaussian sputtering technology, the present invention abandons the direct representation of geometric attributes and color attributes, and adds the representation of semantic attributes. Specifically, each Gaussian feature ball uses geometric features , color characteristics and semantic features , and its corresponding geometric attribute decoder , color attribute decoder and semantic attribute decoder To represent geometric properties , color attributes and semantic attributes , is the position of the Gaussian characteristic sphere, is the opacity of the Gaussian characteristic sphere, is the variance of the Gaussian characteristic sphere.

[0065] Using dynamic and learnable feature representation instead of direct attribute representation improves the representation ability and flexibility of the model. By adding semantic attributes and using Gaussian sputtering differentiable rendering technology, it achieves effective fusion of 2D segmentation information to 3D segmentation information, solving the problem of accurate separation of multiple human bodies in interactive scenes.

[0066] The fourth step is to use Gaussian sputtering rendering technology to render the Gaussian feature sphere representation into a 2D image. And human body segmentation information , using the corresponding input image under the target perspective And pre-estimated human segmentation information , which can calculate the rendering error :

[0067]

[0068] In the fifth step, to achieve better segmentation accuracy at the edges of interacting human bodies, we proposed an error-guided Gaussian eigensphere adaptive densification algorithm. At the boundary of human body separation, the threshold for splitting the Gaussian eigenspheres is adaptively adjusted based on the size of the human body reconstruction error and the human body segmentation error. This achieves adaptive densification and utilizes a larger number of smaller Gaussian eigenspheres to improve reconstruction quality and separation accuracy. The pseudo code of this algorithm is as follows:

[0069] Input: Initial Gaussian sphere set , error threshold ε, split ratio upper limit r_max;

[0070] Output: optimized Gaussian ball set G'

[0071] 1. Initialize G'=

[0072] 2. For each Gaussian ball g_i in G':

[0073] 2.1 Calculating rendering error e_i

[0074] 2.2 If e_i >ε:

[0075] 2.2.1 Calculate the split ratio r = min(e_i / ε, r_max)

[0076] 2.2.2 Split g_i into r small Gaussian balls

[0077] 2.2.3 Update G'

[0078] 3. Return to G'

[0079] Through the above steps, the system can achieve simultaneous separation and reconstruction of up to 6 people. The reconstruction accuracy is 30% higher than that of traditional methods, and the frame rate reaches 100fps, meeting the needs of real-time applications.

[0080] Example 3: Missing Data Generation

[0081] like Figure 3 As shown, missing data generation is another key innovation of the present invention. A generation algorithm based on a dynamic graph convolutional network is used to intelligently infer and complete 3D human body areas that are invisible or occluded under sparse viewing angles.

[0082] The Gaussian feature representation of the human body output by joint reconstruction of multiple bodies suffers from insufficient sparse viewpoint observation and mutual occlusion in interactive scenes. This results in incomplete reconstruction of some areas due to a lack of effective supervision. We use the predefined connectivity graph of SMPL-X to construct a Gaussian feature ball connectivity graph for the 3D human body and complete it using a dynamic graph convolutional network.

[0083] In the first step, based on the human body mesh representation, which consists of a vertex set V and mutual connection relationships F, a connection relationship matrix A of all vertices can be constructed to obtain a connection relationship graph of Gaussian characteristic balls associated with all vertices.

[0084] In the second step, the human body mesh representation based on the SMPL-X parametric human body model can represent the complete human body geometry. Through mesh rendering technology, the 2D segmentation information of the complete human body can be obtained. , which is consistent with the incomplete human segmentation information estimated above Perform the "AND" operation to obtain the human body visible area mask , the invisible area mask can be calculated as .

[0085] The third step is to give the camera parameters C under the target perspective, through the inverse projection transformation , we can get the Gaussian characteristic sphere set of the visible area and the Gaussian sphere set of the invisible area .

[0086] The fourth step is to build a dynamic graph convolutional neural network (DGCN). On the Gaussian feature ball connection relationship graph, using the connection relationship matrix A, the Gaussian feature ball features of the visible area, including geometric features, appearance features, and semantic features, are used to predict the Gaussian ball features of the invisible area. The core operations of DGCN are:

[0087]

[0088] Indicates the Nodes in the layer The feature representation of is the neighborhood connection weight calculated by the connection relationship matrix A, N(i) represents the neighborhood connection weight with node The set of connected neighbor nodes, (·) is the activation function, and Respectively The convolution kernel weights and biases of the layer.

[0089] The innovation of this invention is to perform graph convolution directly on the connection relationship graph of Gaussian feature balls, without the intermediate step of UV mapping. By performing dynamic graph convolution calculation on Gaussian feature balls bound to the human body mesh, the Gaussian feature ball features of the visible area, including geometric features, color features and semantic features, are used to predict the Gaussian feature ball features of the invisible area. The eigenvector of To complete, the update formula is:

[0090]

[0091] in, Represents the feature transfer weight between the i-th node and the j-th node, satisfying the normalization constraint. F_j is the Gaussian feature ball of the closest area found by f_i in the visible area based on the connection relationship matrix A. .

[0092] In the fifth step, the Gaussian feature spheres of the visible area and the Gaussian feature spheres predicted in the invisible area are combined to obtain a complete human body reconstruction result through the aforementioned Gaussian sputtering rendering technology.

[0093] Through this process, the proposed solution fuses spatial and temporal information simultaneously, generating missing body data consistent with known parts and enabling complete reconstruction of the human body model. The generation process takes an average of 350 milliseconds, meeting near-real-time processing requirements.

[0094] Example 4: Action Recognition and Tracking

[0095] The action recognition and tracking of this embodiment are based on the spatiotemporal graph convolutional network (ST-GCN) and a two-stream network architecture to achieve action recognition and tracking of the reconstructed human body.

[0096] First, the human body key point sequence is represented as a spatiotemporal graph G=(V, E), where V is the node set (corresponding to the key points) and E is the edge set (corresponding to the skeletal connections). The spatiotemporal graph convolution operation is defined as:

[0097]

[0098] in, and are the input and output features, is the convolution kernel, is the adjacency matrix, and K is the set of convolution kernels.

[0099] Secondly, we introduce a feature extraction method based on spatiotemporal interest points and dense trajectories. By analyzing the motion trajectories of key points, we can capture the spatiotemporal characteristics of the action. Spatiotemporal interest points (STIPs) locate the action by detecting significant changes in the video area, while dense trajectories describe the detailed changes of the action by tracking the motion paths of multiple feature points.

[0100] Finally, a two-stream network architecture is used to simultaneously process spatial stream (single frame image) and temporal stream (optical flow) information. The prediction results of the two streams are weighted and fused to obtain the final action classification result:

[0101]

[0102] in, and are the prediction scores of spatial stream and temporal stream respectively, and is the weight coefficient.

[0103] This embodiment predefines 80 common household activity action categories, including daily movements such as walking, sitting, standing up, lying down, bending over, and falling, as well as household activities such as cooking, cleaning, and watching TV. The accuracy rate can reach over 92% in complex household scenarios, meeting real-time monitoring needs.

[0104] Example 5: Behavior Analysis and Intention Understanding

[0105] like Figure 4 As shown in the figure, behavior analysis and intention understanding are based on behavior pattern mining of time series data and multi-person interaction relationship modeling to analyze the behavior patterns and interaction intentions of family members.

[0106] First, we build a behavior sequence representation model to represent the identified action sequence as a time series feature vector:

[0107] ,in is the action characteristic at time t.

[0108] Secondly, design a multi-person interaction diagram ,in is a collection of human nodes, is the set of interaction edges. The relationship between nodes is modeled by the following attention mechanism:

[0109] ,

[0110] in, is the query vector, is the key vector, is a value vector, is the set of neighbors of node i.

[0111] Finally, through temporal pattern analysis and rule reasoning, the behavioral intentions of family members are identified, such as:

[0112] Identify abnormal behaviors (such as elderly people falling, children’s dangerous behaviors, etc.);

[0113] Understand daily activities (such as eating, resting, and entertainment);

[0114] Analyze the interactions among family members (e.g., conversations, assistance, etc.).

[0115] This embodiment can identify 10 basic intent categories and 30 complex interaction intentions, providing intelligent support for home security monitoring.

[0116] Example 6: System Integration and Deployment

[0117] The system of the present invention adopts a centralized processing architecture, with the central control computer (PC host equipped with a high-performance GPU) responsible for executing all computationally intensive tasks. The minimum configuration requirements for the central control computer are:

[0118] CPU: Intel Core i7-10700 or AMD Ryzen 7 3700X or above;

[0119] GPU: NVIDIA RTX 3060 12GB or above, supporting CUDA 11.0+;

[0120] Memory: 32GB DDR4;

[0121] Storage: 1TB SSD

[0122] Network: Supports Gigabit wired network and WiFi 6 wireless network.

[0123] The software architecture adopts a modular design, with each functional module communicating and exchanging data through standard interfaces. The system runs on Ubuntu 20.04 LTS or Windows 10 Professional, and the core algorithm library is implemented using PyTorch version 1.10 or later.

[0124] The system's end-to-end processing latency (from data acquisition to output) is controlled within 100ms, meeting the needs of real-time monitoring. The system supports 24 / 7 continuous operation and features fault recovery and data backup capabilities, ensuring the stability and reliability of the monitoring system.

[0125] Example effect verification:

[0126] The invention was tested and verified in a real-world home environment. The test environment was a typical three-bedroom residence with an area of approximately 120 square meters, where five cameras were deployed. The test scenarios included various complex scenarios such as daily life, multi-person interactions, and occlusion.

[0127] The test results are shown in Table 1:

[0128] The accuracy of multi-body reconstruction reaches 97.5% in the unobstructed case, 92.3% in the lightly obstructed case, and 85.6% in the heavy obstructed case.

[0129] Missing data generation quality: Through user research and quantitative evaluation, the structural similarity (SSIM) between generated data and real data reaches above 0.85, meeting the visual quality requirements.

[0130] Motion recognition accuracy: The accuracy of single-person motion recognition is 94.2%, and 91.5% in multi-person interaction scenarios.

[0131] System response time: The average end-to-end processing delay is 82ms, meeting real-time requirements.

[0132] System stability: Continuous operation for 72 hours without any problems, with an average CPU load of 42% and an average GPU load of 65%.

[0133] Compared with the prior art, the present invention has significant advantages in the following aspects:

[0134] The number of cameras required has been significantly reduced, from 10-20 in traditional systems to 3-5, significantly reducing deployment costs.

[0135] The multi-person processing capability is enhanced, and it can handle complex interaction scenarios of up to 6 people at the same time.

[0136] The reconstruction accuracy is improved, especially under occlusion and sparse viewing conditions, with the error reduced by more than 30% compared to traditional methods.

[0137] Real-time performance is optimized, with a frame rate of 100fps, meeting real-time monitoring needs.

[0138] With increased intelligence, it can understand the interaction intentions and behavior patterns of multiple people, providing more value for home security monitoring.

[0139] Table 1

[0140]

Claims

1. A sparse view multi-body joint reconstruction method based on dynamic graph convolutional network, characterized by: The following steps are involved: S1. Use RGB camera to collect indoor RGB video data; S2: Based on the RGB video data, the human body parameter model estimation algorithm is used to estimate the human body's 3D posture, and the completion algorithm is used to generate and predict it to complete the reconstruction of the human body model. The specific implementation process is as follows: Step 2.1: Use the human body parameter model estimation algorithm to estimate the 3D posture of the human body. Using the 3D Gaussian eigensphere representation method, based on RGB video data, perform 3D dynamic reconstruction and posture estimation of multiple interacting human bodies indoors. During the reconstruction process, the initial posture data is synchronously optimized. The specific implementation is as follows: Step 2.1.

1. Based on the RGB video data, obtain the input image of each viewpoint and use a deep convolutional neural network to extract the 2D key point information of the human body and the human body segmentation information of the input image; Step 2.1.2: Based on the detected 2D key point information, use the SMPL-X parameter fitting algorithm to estimate the 3D posture parameters and shape parameters of the human body; Step 2.1.3: Convert the estimated shape parameters into a human body mesh representation, wherein each vertex of the human body mesh representation is associated with a 3D Gaussian feature sphere, and convert the human body mesh representation into a 3D Gaussian feature sphere representation, wherein each Gaussian feature sphere contains a learnable dynamic geometric feature, a learnable dynamic appearance feature, and a learnable semantic feature. The geometric feature predicts its geometric attributes through a geometric attribute decoder, the appearance feature predicts its color attributes through an appearance attribute decoder, and the semantic feature predicts its human ID attributes through a semantic decoder. Step 2.1.4: Render the geometric attributes and appearance attributes into a 2D image representation at the target perspective using a differentiable Gaussian sputtering rendering technique; and render the human ID attributes into human segmentation information of the input image using a differentiable Gaussian sputtering rendering technique. The rendered 2D image representation and human segmentation information are compared with the input image and the estimated input image human segmentation information to calculate the rendering error respectively; Step 2.1.5: Based on the rendering error calculated above, a Gaussian eigensphere adaptive densification algorithm is proposed. The Gaussian eigenspheres of the corresponding area are adaptively densified based on the error. The added Gaussian eigenspheres are initialized with the features of the closest Gaussian eigensphere and associated with the vertex in the human mesh representation that is closest to them. Step 2.1.6: Through the joint iterative optimization of steps 2.1.4 and 2.1.5 under multiple viewpoints, learn the geometric features, color features, and semantic features in the Gaussian sphere and their corresponding attribute decoders, and simultaneously optimize the human 3D pose parameters to correct the inaccuracy in the initial human pose estimation; Step 2.2: Based on the results of 3D dynamic reconstruction and pose estimation, a completion algorithm based on a dynamic graph convolutional network is used to directly perform feature fusion on the connection relationship graph of the Gaussian feature spheres, generate and predict missing data under sparse viewpoints, and realize the construction of a complete human body model. S3. Based on the reconstructed human body model, the motion recognition algorithm is used to achieve real-time classification, positioning and tracking of multiple human body movements; and based on the reconstructed human body model and human 3D posture, the behavioral patterns and interaction intentions of family members are analyzed.

2. The sparse view multi-body joint reconstruction method based on dynamic graph convolutional network according to claim 1 is characterized in that: The Gaussian eigensphere adaptive densification algorithm adaptively adjusts the threshold of the Gaussian eigensphere splitting based on the human body reconstruction error and the human body segmentation error on the human body separation boundary to achieve adaptive densification. The implementation process is as follows: First, initialize the Gaussian ball set, and set the error threshold ε and the split ratio upper limit r_max; For each Gaussian ball g in the Gaussian ball set, calculate the rendering error e; When the rendering error e is greater than the error threshold ε, the split ratio r = min(e / ε, r_max) is calculated; Split the Gaussian ball g into r Gaussian balls and update the Gaussian ball set.

3. The sparse view multi-body joint reconstruction method based on dynamic graph convolutional network according to claim 2 is characterized in that: The specific implementation process of step 2.2 is as follows: Step 2.2.1: Based on the human body mesh representation, the topological structure of the mesh is used to construct a connection relationship diagram between Gaussian eigenspheres; Step 2.2.2: Use mesh rendering technology to obtain complete human segmentation information for each human body mesh representation. Perform a pixel-by-pixel AND operation on this segmentation information with the human body segmentation information of the input image in step 2.1.1 to obtain a human body visible area mask. Invert the human body segmentation information of the input image in step 2.1.1 and AND it with this complete human body segmentation information to obtain a human body invisible area mask. Step 2.2.3, project the visible area mask and the invisible area mask onto the 3D Gaussian feature sphere representation through the camera inverse transformation projection, and obtain the visible area Gaussian feature sphere and the invisible area Gaussian feature sphere respectively; Step 2.2.4: Construct a dynamic graph convolutional neural network (DGCN) and perform dynamic graph convolution on the Gaussian feature ball connection relationship graph. Use the geometric, color, and semantic features of the Gaussian feature balls in the visible area to predict the corresponding features of the Gaussian feature balls in the invisible area. Step 2.2.5: Decode the geometric and color features of the completed Gaussian feature sphere into complete geometric and color attributes through the geometric attribute decoder and color attribute decoder described in step 2.1.3, thereby achieving complete reconstruction under sparse viewing angles.

4. The sparse view multi-body joint reconstruction method based on dynamic graph convolutional network according to claim 3 is characterized in that: The step 2.2 also includes a time domain consistency constraint, so that the geometry and texture attributes predicted by the decoder maintain temporal and spatial consistency under multiple perspectives and multiple time points, thereby adding a completion constraint.

5. The sparse view multi-body joint reconstruction method based on dynamic graph convolutional network according to claim 1 is characterized in that: The action recognition algorithm adopts an algorithm based on the combination of spatiotemporal interest points and dense trajectories to achieve multi-human action recognition and real-time tracking.

Citation Information

Patent Citations

  • Multi-view human dynamic three-dimensional reconstruction method in multi-person closely interactive scene

    CN109242950A

  • Human body posture estimation method based on dynamic graph convolutional network

    CN119445672A

  • Dynamic human body modeling method based on three-dimensional Gaussian sputtering technology

    CN119722919A