Human body three-dimensional reconstruction method and system based on camera
By using a dual-fisheye camera robotic arm collaborative data acquisition system and a dynamic calibration module, the problems of poor data coordination and narrow field of view in human 3D reconstruction in surgical scenarios have been solved, achieving full field of view coverage and continuous point cloud generation, meeting the accuracy and real-time requirements of surgical scenarios.
Patent Information
- Application Number
- CN202511471217.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies for human 3D reconstruction and robotic arm collaboration in surgical scenarios suffer from problems such as poor data coordination, narrow field of view, reliance on manual perspective selection, low reconstruction accuracy, and lack of temporal correlation in point cloud fusion, making it difficult to meet clinical needs.
A dual-fisheye camera robotic arm collaborative data acquisition system is adopted, which constructs a dynamic calibration module, a panoramic image acquisition and stitching module, a viewpoint selection module guided by robotic arm pose, and a point cloud fusion module with spatiotemporal correlation and semantic weighting, to achieve data collaboration, dynamic calibration, intelligent viewpoint selection, and continuous point cloud generation.
It achieves full field-of-view coverage of the surgical scene, dynamically adapts to interference to ensure accurate calibration, intelligently selects the optimal viewpoint to improve reconstruction quality, and generates continuous and complete 3D human body point clouds, meeting the requirements of surgical scenes for accuracy, completeness and real-time performance.
Smart Images

Figure CN121392136A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional reconstruction, more particularly, it relates to a human three-dimensional reconstruction method and system based on a camera. BACKGROUND
[0002] In the human three-dimensional reconstruction and mechanical arm collaborative application in the surgical scene, the existing technology has not formed a systematic solution, and there are significant bottlenecks in each link, which is difficult to meet the needs of precision, field of view and real-time in the clinic. First, in the multi-source data acquisition stage, the camera and the surgical mechanical arm in the existing scheme are independent, and there is a lack of data collaboration mechanism. And the commonly used standard lens camera has obvious field of view limitation. The ordinary image can only obtain a small part of the information in the indoor scene due to the limitation of the angle of view. The standard lens can only obtain 15% of the field of view of the human visual system. The key area of the operation is prone to visual blind area. At the same time, the image data collected by the camera and the motion data output by the mechanical arm are not processed in time and space, which leads to the lack of unified data benchmark in the subsequent analysis, and the precision is limited by the data fragmentation. Secondly, in the camera parameter calibration link, the existing technology mainly depends on the static calibration board to complete the one-time parameter determination, which cannot deal with the dynamic interference such as slight displacement of the support and vibration of the mechanical arm in the operation process. The calibration parameters are prone to drift, which leads to the gaps in the subsequent image stitching and the deviation in the geometric positioning. Moreover, there is a lack of technology to generate multi-view reference images, and only a single panoramic image can be output, which cannot provide multi-dimensional basis for subsequent view selection. Thirdly, in the view selection and three-dimensional reconstruction link, the existing view selection relies on manual judgment and does not adjust dynamically combined with the motion state of the mechanical arm. Invalid view angles are often selected due to the obstruction of the instrument. The three-dimensional reconstruction is mainly based on single round and few view images, and the depth estimation error is large. Moreover, the point cloud fusion only uses simple geometric stitching and does not consider the difference in clinical importance of the operation site, which easily loses the details of the key organs. At the same time, the multi-round reconstruction results lack time sequence correlation, and cannot form a continuous and complete human three-dimensional model. Finally, in the cooperative control link of the mechanical arm and the vision system, the existing camera is fixedly installed and cannot dynamically adjust the view angle according to the obstruction, which easily leads to interruption of the reconstruction and is difficult to match the strict requirements of the operation scene. SUMMARY
[0003] In order to overcome the shortcomings of poor data collaboration, narrow field of view, and frequent blind area in key areas in the prior art; static calibration cannot resist dynamic interference; panoramic image lacks multi-view reference and view selection relies on manual selection; reconstruction precision is low and point cloud fusion lacks difference and time sequence correlation, the present application provides a human three-dimensional reconstruction method and system based on a camera.
[0004] The technical scheme of the present application is as follows:
[0005] A human three-dimensional reconstruction method based on a camera, comprising:
[0006] S1, construct a dual fisheye camera mechanical arm cooperative data acquisition system for acquiring multi-source data of a surgical scene and generating a spatiotemporal correlation dataset;
[0007] S2, construct a dual fisheye camera dynamic calibration module, process the spatiotemporal correlation dataset, and output internal and external parameters of the dual fisheye camera;
[0008] S3, construct a panoramic image acquisition and stitching module, process the spatiotemporal correlation dataset and the internal and external parameters of the dual fisheye camera, and output a stitched panoramic image and six virtual perspective plane images;
[0009] S4, construct a dynamic view selection module for mechanical arm pose guidance, process the six virtual perspective plane images, the spatiotemporal correlation dataset, and the internal and external parameters of the dual fisheye camera, and output an optimal view combination and an optimal view image corresponding to the optimal view combination;
[0010] S5, construct a human three-dimensional reconstruction module, process the optimal view image and the internal and external parameters of the dual fisheye camera, and output a preliminary human three-dimensional point cloud model;
[0011] S6, construct a spatiotemporal correlation and semantic weighted point cloud fusion module, process the preliminary human three-dimensional point cloud model, the spatiotemporal correlation dataset, and the semantic mask of the six virtual perspective plane images, and output a continuous and complete human three-dimensional point cloud model;
[0012] S7, construct a mechanical arm vision cooperative control module, process the optimal view combination, the spatiotemporal correlation dataset, and the continuous and complete human three-dimensional point cloud model, and output a cooperative control instruction and a continuously updated point cloud.
[0013] Further, in an embodiment, the step S1 includes the following steps:
[0014] S11, adopt two fisheye cameras with 180° view angles, fixedly installed through a symmetrical C-arm support of an operating table, to form the dual fisheye camera; a mechanical arm, carrying an end effector, with a built-in high-precision encoder; a hardware synchronization component, respectively connected to the dual fisheye camera and the mechanical arm, for realizing timing synchronization during data acquisition; a central controller, connected to the dual fisheye camera, the mechanical arm, and the hardware synchronization component, and denoted as a dual fisheye camera mechanical arm cooperative data acquisition unit;
[0015] S12, enable the dual fisheye camera mechanical arm cooperative data acquisition unit to synchronously capture dual fisheye original images of a surgical scene and end timing pose data of the mechanical arm, and the central controller applies a unified microsecond-level timestamp to the dual fisheye original images and the end timing pose data to generate the spatiotemporal correlation dataset.
[0016] Further, in an embodiment, the step S2 comprises the following steps:
[0017] S21, based on the spatio-temporal correlation data set, extracting a checkerboard calibration board image and a surgical instrument marker motion image, the checkerboard calibration board image being used for static calibration initialization, and the surgical instrument marker motion image containing the projection of the coded points of the surgical instrument marker in the dual fisheye original image, for dynamic optimization;
[0018] S22, based on the checkerboard calibration board image, calculating the intrinsic matrix K i and the extrinsic [R i |t i ] of the fisheye camera, the intrinsic matrix K i Specifically:
[0019]
[0020] Wherein, f x represents the focal length of the x-axis direction of the fisheye camera, f y represents the focal length of the y-axis direction of the fisheye camera, c x represents the coordinate of the principal point of the fisheye camera on the x-axis of the image coordinate system, and c y represents the coordinate of the principal point of the fisheye camera on the y-axis of the image coordinate system.
[0021] The extrinsic [R i |t i ] includes a rotation matrix R i and a translation vector t i , based on the intrinsic matrix K i and the extrinsic [R i |t i ], the projection matrix P i of the fisheye camera is calculated, the global coordinates P e of the end of the mechanical arm and the projection matrix P i Specifically as follows:
[0022] P e = [X e , Y e , Z e ] T
[0023] P i = K i [R i |t i ]
[0024] S23, setting a straight line trajectory, a circular trajectory and a simulated surgery trajectory based on the surgical instrument marker, the simulated surgery trajectory being used to simulate an operation path of the surgical instrument, synchronously collecting the encoded point projections, calculating theoretical pixel coordinates and actual pixel coordinates of each encoded point projection, and calculating pixel coordinate errors based on the theoretical pixel coordinates and the actual pixel coordinates, and iteratively optimizing an intrinsic matrix K based on the pixel coordinate errors i and an extrinsic matrix [R i |t i ] until the pixel coordinate errors are less than or equal to a preset value, and finally outputting the intrinsic matrix K i , the extrinsic matrix [R i |t i ] and a projection matrix P i .
[0025] Further, in an embodiment, the step S3 comprises the following steps:
[0026] S31, based on the intrinsic matrix K i and the extrinsic matrix [R i |t i ], performing distortion correction on the fisheye camera, and then performing denoising, brightness adjustment and color correction operations on the double fisheye original image to output two preprocessed 180° fisheye images;
[0027] S32, based on the improved ORB algorithm, extracting feature points and binary descriptors from each of the 180° fisheye images, screening feature point matching pairs based on Hamming distance, and removing false matching pairs through the RANSAC algorithm to output the corresponding relationship between the feature points of the two 180° fisheye images;
[0028] S33, obtaining a homography matrix H of the two 180° fisheye images through the corresponding relationship between the feature points, and performing perspective transformation alignment on one of the 180° fisheye images based on the homography matrix H, and then performing pixel-level fusion on the overlapping region with the other 180° fisheye image to generate the panoramic image;
[0029] S34, decoupling the panoramic image into the six virtual perspective plane images, the six virtual perspective plane images including a forward virtual perspective plane image, a backward virtual perspective plane image, a leftward virtual perspective plane image, a rightward virtual perspective plane image, an upward virtual perspective plane image and a downward virtual perspective plane image.
[0030] Further, in an embodiment, the step S4 comprises the following steps:
[0031] S41, based on a U-Net model, outputting semantic masks of the six virtual perspective plane images;
[0032] S42, extracting the global coordinate P of the end of the robot arm at the current moment based on the end timing pose data e , combined with the projection matrix P i , calculating the pixel coordinates u of the end of the robot arm on the six virtual perspective plane images i , which is specifically expressed as:
[0033]
[0034] u i = [u, v] T
[0035] When the pixel coordinates u i fall within the human body region in the semantic mask, the angle θ between the line connecting the end of the robot arm and the optical center of the corresponding fisheye camera and the surface normal vector of the human body is calculated, and the distance d from the end of the robot arm to the corresponding fisheye camera is calculated, and finally the occlusion probability P of the forward virtual perspective plane image, the backward virtual perspective plane image, the left virtual perspective plane image, the right virtual perspective plane image, the upward virtual perspective plane image and the downward virtual perspective plane image is calculated occl , the occlusion probability P occl is specifically as follows:
[0036]
[0037] where d max = 500mm, which is the maximum distance of the operating table workspace.
[0038] Further, in an embodiment, the step S4 further comprises the following steps:
[0039] S43, based on the end timing pose data, the type of surgical instrument in the surgical instrument marker and the semantic mask, calculating the occlusion probability heat map of the forward virtual perspective plane image, the backward virtual perspective plane image, the left virtual perspective plane image, the right virtual perspective plane image, the upward virtual perspective plane image and the downward virtual perspective plane image within 0.5-1 seconds in the future, wherein the greater the pixel value of the occlusion probability heat map represents the higher the occlusion probability of the corresponding region;
[0040] S44, performing static screening to screen the virtual perspective at the current moment whose occlusion probability P occl is less than the preset value, and then performing dynamic screening, which is specifically screening the virtual perspective within 0.5-1 seconds in the future from the static screening result whose occlusion probability P occlThe virtual view angles less than the preset value are continuously selected, and finally 2 or 3 virtual view angles are selected to form an optimal view angle combination, and a view angle image corresponding to the combination is extracted and defined as the optimal view angle image.
[0041] Further, in an embodiment, the step S5 comprises the following steps:
[0042] S51, using the improved ORB algorithm to extract feature points and binary descriptors of the optimal view angle image, associating the cross-view matching pairs of the optimal view angle combination based on the Hamming distance, solving the essential matrix E between the view angles in the optimal view angle combination by the 5-point method, combining the intrinsic matrix K i Solving the relative pose between the view angles, the relative pose including a relative rotation matrix R rel and a relative translation vector t rel , verifying that the re-projection error of the relative pose is less than or equal to 1.0 pixels;
[0043] S52, constructing a lightweight MVSNet variant model based on the MVSNet model, the lightweight MVSNet variant model being half of the MVSNet model and retaining 4 feature extraction layers;
[0044] S53, inputting the optimal view angle image and the relative pose into the lightweight MVSNet variant model to predict the depth value of each pixel in each of the optimal view angle images, and generating a dense depth map corresponding to each of the optimal view angle images based on the prediction result;
[0045] S54, performing a back projection operation on the dense depth map to obtain a sparse human body three-dimensional point cloud, specifically as follows:
[0046]
[0047] Wherein, Z represents the pixel depth value, (u, v) is the pixel coordinate, (c x , c y ) is the camera principal point coordinate;
[0048] Through the PMVS algorithm, the sparse human body three-dimensional point cloud is densified to form a preliminary human body three-dimensional point cloud model.
[0049] Further, in an embodiment, the step S6 comprises the following steps:
[0050] S61, based on the step S54, identifying a common view angle of multiple rounds of the preliminary human body three-dimensional point cloud model, the common view angle being a virtual view angle repeatedly appearing in the optimal view angle combinations corresponding to the preliminary human body three-dimensional point cloud models of different rounds, and based on the extrinsic parameters [R i |t iunifies all the preliminary point clouds to the same world coordinate system, performs fine registration on the overlapping areas of the multiple preliminary human three-dimensional point cloud models using an ICP algorithm, and outputs the spatially aligned multiple human three-dimensional point clouds;
[0051] S62, based on the MobileNetV2 improved lightweight CNN model, the human key parts in the multiple human three-dimensional point cloud models are extracted and weighted according to clinical importance, the point cloud is fused by combining the rigid body model of the mechanical arm movement and the elastic deformation model of the human tissue, the points with a weight greater than or equal to a preset value are preferentially retained in the high-weight part, and the point cloud in the low-weight part is smoothed by weighted average, and the point cloud after semantic weighted fusion is output.
[0052] Further, in an embodiment, the step S6 further comprises the following steps:
[0053] S63, based on the point cloud after semantic weighted fusion, a point cloud time sequence is established with a time stamp, the dynamic change of the human key part point cloud is predicted by Kalman filtering, a space-time constraint graph is constructed, the overall error is minimized by graph optimization, the human structure is completed by Poisson reconstruction in the non-overlapping area, the global optimization is triggered when the same optimal view combination appears again, the cumulative error is corrected, and finally the continuous and complete human three-dimensional point cloud model is generated.
[0054] A camera-based human three-dimensional reconstruction system applied to the camera-based human three-dimensional reconstruction method in the above embodiment, comprising:
[0055] Acquisition hardware for acquiring surgical scene images, outputting spatial pose data and adjusting the position of the object;
[0056] Mounting and supporting hardware for fixing the position of the acquisition hardware and carrying the object;
[0057] Calibration and auxiliary hardware for calibrating the basic parameters of the acquisition hardware and optimizing the accuracy of the parameters;
[0058] Control and synchronization hardware for controlling the acquisition hardware, the mounting and supporting hardware, and the calibration and auxiliary hardware.
[0059] The application according to the above scheme has the advantages of realizing cooperative acquisition and full field of view coverage of surgical scene data, dynamically adapting to interference to ensure accurate calibration, intelligently selecting optimal views to improve reconstruction quality, generating continuous point clouds through space-time semantic fusion, and realizing efficient cooperative control of the mechanical arm and vision to meet the requirements of precision, integrity and real-time of the surgical scene. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.
[0061] Figure 1 is a flowchart of the method for three-dimensional reconstruction based on multi-view point cloud in the embodiment.
[0062] Figure 2 is a flowchart of the system for three-dimensional reconstruction based on multi-view point cloud in the embodiment.
[0063] Figure 3 is a structural diagram of six virtual perspective planar images in the embodiment. DETAILED DESCRIPTION
[0064] The present application will be further described in combination with the drawings and embodiments:
[0065] Meanwhile, in order to better understand the specific implementation process, the names of the terms will be explained in more detail to help understand their meanings.
[0066] As shown in Figure 1 and Figure 3 , a method for three-dimensional reconstruction based on multi-view point cloud comprises the following steps:
[0067] S1, a dual fisheye camera mechanical arm cooperative data acquisition system is constructed, which is used for acquiring multi-source data of a surgical scene and generating a space-time correlated data set;
[0068] Step S1 comprises the following steps:
[0069] S11, two fisheye cameras with 180° view angle are fixedly installed on a symmetrical C-arm support of an operating table to form a dual fisheye camera; a mechanical arm is provided with an end effector and a built-in high-precision encoder; a hardware synchronization component is connected to the dual fisheye camera and the mechanical arm respectively, and is used for realizing time sequence synchronization during data acquisition; a central controller is connected to the dual fisheye camera, the mechanical arm and the hardware synchronization component, and is denoted as a dual fisheye camera mechanical arm cooperative data acquisition unit;
[0070] S12, the dual fisheye camera mechanical arm cooperative data acquisition unit is enabled to synchronously capture dual fisheye original images of a surgical scene and time sequence pose data of an end effector of a mechanical arm; the central controller marks the dual fisheye original images and the time sequence pose data with a unified microsecond-level time stamp to generate a space-time correlated data set;
[0071] In an embodiment, the dual fisheye camera unit: adopts two fisheye cameras with a 180° field of view, is fixedly installed on a surgical table symmetrical C-arm support, the support keeps the included angle of the optical axes of the two fisheye cameras at 175°±1°, ensures that the single-camera field of view has no overlapping redundancy and the scene edge is continuously covered, and the resolution of the camera is set to 4096×2160, and the frame rate is ≥30fps;
[0072] The surgical mechanical arm unit: carries an end effector (such as a needle holder or an electrotome), has a built-in high-precision encoder, and the positioning accuracy of the encoder is ≤0.05mm; the surgical mechanical arm unit can output the three-dimensional position of the end in real time, which is defined as the global coordinate of the end of the mechanical arm, and simultaneously output the quaternion orientation of the end, and the output frequency is ≥100Hz;
[0073] The hardware synchronization component: adopts a GPIO trigger signal to realize time sequence synchronization of data acquisition, and the synchronization error is ≤1ms; the hardware synchronization component is connected to the control systems of the dual fisheye camera unit and the surgical mechanical arm unit respectively;
[0074] Data acquisition and preprocessing: start the dual fisheye camera mechanical arm cooperative data acquisition system, and synchronously capture the dual fisheye original images of the surgical scene and the time sequence position and orientation data of the end of the surgical mechanical arm; wherein the storage format of the dual fisheye original images is PNG, the time sequence position and orientation data of the end of the surgical mechanical arm includes a timestamp, a global coordinate of the end of the mechanical arm and a quaternion orientation, and the storage format is TXT; the central control unit is used to mark a unified microsecond-level timestamp on the collected multi-source data; an image definition evaluation algorithm is used to screen the dual fisheye original images, and the images with a gradient threshold value ≤20 are determined as blurred images and are removed; based on the working space range of the surgical table, the time sequence position and orientation data of the end of the surgical mechanical arm are screened, and the positions and orientations beyond the working space are determined as abnormal data and are removed;
[0075] Output result: generate a space-time correlation data set, which includes the dual fisheye original images and the time sequence position and orientation data of the end of the surgical mechanical arm, the data set is stored in a compressed package format, and includes a data index table, and the index table marks the time stamp correlation of each data.
[0076] S2, a dual fisheye camera dynamic calibration module is constructed, the space-time correlation data set is processed, and the internal and external parameters of the dual fisheye camera are output;
[0077] Step S2 includes the following steps:
[0078] S21, based on the space-time correlation data set, a checkerboard calibration plate image and a surgical instrument marker movement image are extracted, the checkerboard calibration plate image is used for static calibration initialization, and the surgical instrument marker movement image includes the projection of the encoding point of the surgical instrument marker in the dual fisheye original image and is used for dynamic optimization;
[0079] S22, based on the checkerboard calibration plate image, the internal parameter matrix K of the fisheye camera is calculatedi and an extrinsic parameter [R i |t i ], the intrinsic parameter matrix K i Specifically:
[0080]
[0081] Wherein, f x represents the focal length of the x-axis direction of the fisheye camera, f y represents the focal length of the y-axis direction of the fisheye camera, c x represents the coordinate of the principal point of the fisheye camera on the x-axis of the image coordinate system, and c y represents the coordinate of the principal point of the fisheye camera on the y-axis of the image coordinate system.
[0082] The extrinsic parameter [R i |t i ] includes a rotation matrix R i and a translation vector t i , based on the intrinsic parameter matrix K i and the extrinsic parameter [R i |t i ], the projection matrix P i of the fisheye camera is calculated, the global coordinates P e of the end of the mechanical arm and the projection matrix P i are calculated.
[0083] P e =[X e ,Y e ,Z e ] T
[0084] P i =K i [R i |t i ]
[0085] S23, based on the surgical instrument marker, a straight line trajectory, a circular trajectory and a simulated surgery trajectory are set, the simulated surgery trajectory is used to simulate the operation path of the surgical instrument, the coded point projection is synchronously collected, the theoretical pixel coordinates and the actual pixel coordinates of each coded point projection are calculated, and based on the theoretical pixel coordinates and the actual pixel coordinates, the pixel coordinate error is calculated, and based on the pixel coordinate error, the intrinsic parameter matrix K i and the extrinsic parameter [R i |t i ] are iteratively optimized until the pixel coordinate error is less than or equal to a preset value, and finally the intrinsic parameter matrix K i , the extrinsic parameter [R i |t i ] and the projection matrix P i of the two fisheye cameras are output.
[0086] The double fisheye camera dynamic calibration module is composed of a surgical instrument marker, a static calibration initialization submodule, and a dynamic optimization submodule. The surgical instrument marker is a surgical forceps with 12 high-precision coded points, the coded point diameter is 2 mm, the coding accuracy is ≤0.01 mm, and the three-dimensional coordinates of the coded points are pre-calibrated and stored by a high-precision measuring instrument;
[0087] The input data is clear: from the spatiotemporal correlation data set output from S1, two types of calibration-related data are extracted, which are the checkerboard calibration plate image and the surgical instrument marker motion image. The checkerboard calibration plate image is used for static calibration initialization, and the checkerboard size is 10 mm×10 mm. The surgical instrument marker motion image contains the projection of the coded points in the double fisheye image, which is used for dynamic optimization.
[0088] The processing procedure is as follows: static calibration initialization: input the checkerboard calibration plate image into the static calibration initialization submodule, and preliminarily calculate the camera intrinsic matrix and camera extrinsic of each fisheye camera based on Zhang's calibration method. Finally, the initial projection matrix of the i-th fisheye camera is generated.
[0089] Dynamic optimization: control the surgical instrument marker to move along 3 preset trajectories in the operating table workspace, which are straight line trajectory, circular trajectory, and simulated surgical trajectory, respectively. The straight line trajectory ranges from 0-500 mm×0-300 mm×0-200 mm, the circular trajectory has a radius of 100 mm and a height of 150 mm, and the simulated surgical trajectory simulates the operation path of the surgical instrument. The coded point projections of the marker in the double fisheye image are synchronously collected. For each coded point, the theoretical pixel coordinates of the coded point are calculated. The actual pixel coordinates of the coded point are obtained, and the error Δu is calculated. The camera intrinsic matrix and camera extrinsic are iteratively optimized to minimize the average Δu of all coded points, until the average Δu is ≤0.5 mm, ensuring that the calibration accuracy reaches sub-millimeter level.
[0090] The output results are as follows: the intrinsic matrix, extrinsic, and projection matrix of the first fisheye camera are output, and the intrinsic matrix, extrinsic, and projection matrix of the second fisheye camera are output. All parameters are stored in XML format, and a calibration accuracy verification report is output, which contains the average Δu, maximum error, and minimum error data.
[0091] S3, a panoramic image acquisition and stitching module is constructed, which processes the spatiotemporal correlation data set and the intrinsic and extrinsic parameters of the double fisheye camera, and outputs the stitched panoramic image and six virtual perspective planar images.
[0092] Step S3 includes the following steps:
[0093] S31, based on the intrinsic matrix K i and the extrinsic [R i |t iThe distortion of the fisheye camera is corrected, and then the original double fisheye images are denoised, brightness adjusted and color corrected to output two pre-processed 180° fisheye images.
[0094] S32. Based on the improved ORB algorithm, feature points and binary descriptors are extracted for each 180° fisheye image. Feature point matching pairs are selected based on Hamming distance. Mismatched pairs are eliminated by RANSAC algorithm, and the feature point correspondence between two 180° fisheye images is output.
[0095] S33. Obtain the homography matrix H of two 180° fisheye images through the correspondence of feature points. Based on the homography matrix H, perform perspective transformation and alignment on one of the 180° fisheye images and then perform pixel-level fusion of the overlapping area with the other 180° fisheye image to generate a panoramic image.
[0096] S34. Decouple the panoramic image into six virtual view plane images, including a forward virtual view plane image, a backward virtual view plane image, a left virtual view plane image, a right virtual view plane image, an upward virtual view plane image, and a downward virtual view plane image.
[0097] The input consists of two types of data: the original dual fisheye images in the spatiotemporal correlation dataset output by S1, and the intrinsic and extrinsic parameters of the first fisheye camera and the second fisheye camera output by S2. The intrinsic and extrinsic parameters of the first and second fisheye cameras are used for fisheye distortion correction in the image preprocessing stage.
[0098] Module composition definition: Construct a 360° panoramic image acquisition and stitching module consisting of an image preprocessing submodule, a feature matching submodule, and a geometric transformation and fusion submodule;
[0099] Processing process: image preprocessing: input the dual fisheye original image into the image preprocessing submodule, perform fisheye distortion correction based on the camera intrinsic matrix and the external parameter output by S2; after correction, perform denoising, brightness adjustment and color correction operations in turn; the denoising operation adopts the combination of Gaussian filter and median filter, wherein the Gaussian filter kernel size is 3x3, the standard deviation σ is 1.0, and the median filter kernel size is 5x5; the brightness adjustment adopts histogram equalization, and the gray level is set to 256; the color correction adopts the gray world method in the white balance algorithm; after processing, output two 180° fisheye images after preprocessing; feature matching: input the two 180° fisheye images after preprocessing into the feature matching submodule, and extract image feature points and binary descriptors by using the improved ORB algorithm; the number of layers of the scale pyramid of the improved ORB algorithm is set to 8, the number of feature points is set to 2000, and the edge threshold is set to 31; filter the feature point matching pairs based on Hamming distance, and remove the false matching pairs by RANSAC algorithm, the number of iterations of RANSAC algorithm is set to 1000 times, and the inlier threshold is set to 2 pixels; keep the proportion of effective matching pairs ≥80%, and output the corresponding relationship of feature points of the two 180° fisheye images;
[0100] Geometric transformation and fusion: based on the corresponding relationship of feature points, calculate the homography matrix H of the two 180° fisheye images by 8-point method, and ensure that the reprojection error of the homography matrix is ≤1.5 pixels; perform perspective transformation alignment on one of the 180° fisheye images based on the homography matrix; perform pixel-level fusion on the overlapping area of the images by using the feathering algorithm, and the width of the overlapping area is ≤50 pixels; generate a 360° panoramic image after fusion, the image resolution is 8192x4096, the storage format is PNG, and the file size is 30-50MB; decouple the 360° panoramic image into six virtual perspective plane images in the forward, backward, left, right, upward and downward directions, and the resolution of the six images is uniformly set to 1920x1080, the storage format is PNG, and each image is labeled with the corresponding perspective direction and timestamp; output result: output the 360° panoramic image, and output six virtual perspective plane images, which are forward virtual perspective plane image, backward virtual perspective plane image, left virtual perspective plane image, right virtual perspective plane image, upward virtual perspective plane image and downward virtual perspective plane image.
[0101] Output result: output the 360° panoramic image, and output six virtual perspective plane images, which are forward virtual perspective plane image, backward virtual perspective plane image, left virtual perspective plane image, right virtual perspective plane image, upward virtual perspective plane image and downward virtual perspective plane image.
[0102] S4, construct a dynamic view selection module of the robot arm pose guidance, process six virtual view planar images, spatiotemporal correlation data sets and internal and external parameters of the fisheye camera, and output an optimal view combination and an optimal view image corresponding to the optimal view combination;
[0103] Step S4 includes the following steps:
[0104] S41, based on the U-Net model, output the semantic mask of the six virtual view planar images;
[0105] S42, based on the end time sequence pose data, extract the global coordinates P of the robot arm end at the current time e , combined with the projection matrix P i , calculate the pixel coordinates u of the robot arm end on the six virtual view planar images i , which is specifically represented as:
[0106]
[0107] u i =[u,v] T
[0108] When the pixel coordinates u i fall within the human body region in the semantic mask, calculate the included angle θ between the line connecting the robot arm end and the optical center of the corresponding fisheye camera and the surface normal vector of the human body, and calculate the distance d from the robot arm end to the corresponding fisheye camera, and finally calculate the occlusion probability P occl of the forward virtual view planar image, the backward virtual view planar image, the left virtual view planar image, the right virtual view planar image, the upward virtual view planar image and the downward virtual view planar image, the occlusion probability P occl is as follows:
[0109]
[0110] Where d max = 500mm, which is the maximum distance of the operating table workspace;
[0111] S43, based on the end time sequence pose data, the type of surgical instrument in the surgical instrument marker and the semantic mask, calculate the occlusion probability heat map of the forward virtual view planar image, the backward virtual view planar image, the left virtual view planar image, the right virtual view planar image, the upward virtual view planar image and the downward virtual view planar image within 0.5-1 seconds in the future, wherein the greater the pixel value of the occlusion probability heat map represents the higher the occlusion probability of the corresponding region;
[0112] S44, static screening, screening the occlusion probability P occlThe virtual view angle is less than the preset value, and dynamic screening is performed. Specifically, from the static screening result, the occlusion probability P of the future 0.5-1 seconds is screened occl The virtual view angle is less than the preset value, and finally 2 or 3 virtual view angles are selected to form an optimal view angle combination, and the view angle image corresponding to the combination is extracted, which is defined as the optimal view angle image.
[0113] The input data is clear: three types of data are input, which are six virtual view angle planar images (forward, backward, left, right, upward, downward) output by S3, the time sequence pose data of the surgical robot arm end (including the global coordinates of the robot arm end) in the space-time correlation data set output by S1, and the projection matrix of the first fisheye camera and the projection matrix of the second fisheye camera output by S2; wherein the projection matrix of the first and second fisheye cameras is used to calculate the projection position of the robot arm end in the virtual view angle planar image;
[0114] Module composition definition: a dynamic view angle selection module for robot arm pose guidance is constructed, which is composed of a semantic segmentation sub-module, an occlusion inference sub-module, an O-PredictNet occlusion prediction sub-module, and a view angle screening sub-module; wherein the O-PredictNet occlusion prediction sub-module is a lightweight CNN model, which inputs the past 500ms of time sequence pose data of the surgical robot arm end, the type of surgical instrument (obtained from the surgical robot arm control system), and the semantic mask of the current six virtual view angle planar images, and outputs the occlusion probability heat map of the six virtual view angle planar images in the future 0.5-1 seconds, the pixel value range of the heat map is 0-1, and the larger the pixel value represents the higher the occlusion probability of the corresponding region;
[0115] Processing process: semantic segmentation: input the six virtual view angle planar images into the semantic segmentation sub-module, which adopts a lightweight U-Net model, the model input size is 1920x1080, the number of convolution kernels is 64, 128, and 256 in turn, and the activation function adopts ReLU; six virtual view angle planar image semantic masks are output by the model, and the semantic mask pixel value is defined as 0=background area, 1=surgical instrument area, and 2=human body area;
[0116] Occlusion inference: for each virtual view angle planar image (corresponding to the ith virtual view angle), the global coordinates of the robot arm end at the current time Pe are extracted from the time sequence pose data of the surgical robot arm end; combined with the projection matrix of the corresponding ith fisheye camera, the pixel coordinates of the robot arm end on the virtual view angle planar image are calculated, if the pixel coordinates fall within the human body area (pixel value=2) in the semantic mask, the angle between the line of sight of the robot arm end and the normal vector of the human body surface is calculated, the distance from the robot arm end to the corresponding fisheye camera is calculated, and the occlusion probability of the virtual view angle is calculated;
[0117] O-PredictNet Occlusion Prediction: The past 500 ms of surgical robot end effector temporal pose data, surgical instrument type, and semantic mask of six virtual perspective planar images are input into the O-PredictNet Occlusion Prediction submodule. The model outputs the occlusion probability heat map of the six virtual perspective planar images in the future 0.5-1 seconds;
[0118] View angle screening: The view angle screening submodule performs a two-step screening operation. The first step is static screening, which screens virtual perspectives with an occlusion probability Poccl<0.3 at the current time. The second step is dynamic screening, which screens virtual perspectives with a continuously <0.3 occlusion probability in the future 0.5-1 seconds from the static screening results. Finally, 2-3 virtual perspectives are selected to form an optimal unoccluded view angle combination, and the corresponding view angle images are extracted, defined as the view angle images corresponding to the optimal unoccluded view angle combination;
[0119] Output results: The optimal unoccluded view angle combination is output, which includes the virtual perspective direction identifier in the combination, the current occlusion probability of each view angle, and the future predicted occlusion probability. At the same time, the view angle images corresponding to the optimal unoccluded view angle combination are output, and the image storage format is PNG, with the view angle direction and timestamp labeled.
[0120] S5, constructing a human three-dimensional reconstruction module, processing the optimal view angle image and the internal and external parameters of the dual fisheye camera, and outputting a preliminary human three-dimensional point cloud model;
[0121] Step S5 includes the following steps:
[0122] S51, using an improved ORB algorithm to extract feature points and binary descriptors of the optimal view angle image, associating the cross-view matching pairs of the optimal view angle combination based on Hamming distance, solving the essential matrix E between the view angles in the optimal view angle combination by 5-point method, and combining the intrinsic matrix K i to solve the relative pose between the view angles, which includes the relative rotation matrix R rel and the relative translation vector t rel , and verifying that the re-projection error of the relative pose is less than or equal to 1.0 pixels;
[0123] S52, constructing a lightweight MVSNet variant model based on the MVSNet model, which is half of the MVSNet model and retains 4 layers of feature extraction layers;
[0124] S53, inputting the optimal view angle image and the relative pose into the lightweight MVSNet variant model to predict the depth value of each pixel in each optimal view angle image, and generating a dense depth map corresponding to each optimal view angle image based on the prediction results;
[0125] S54, the back projection operation is performed on the dense depth map to obtain a sparse human three-dimensional point cloud, and the specific process is as follows:
[0126]
[0127] wherein Z represents a pixel depth value, (u, v) is a pixel coordinate, (c x , c y ) is a camera principal point coordinate;
[0128] The PMVS algorithm is used to perform point cloud densification operation on the sparse human three-dimensional point cloud to form a preliminary human three-dimensional point cloud model.
[0129] Input data: two types of data are input, which are the view images corresponding to the optimal unoccluded view combination output by S4 and the intrinsic matrix and extrinsic parameters of the first fisheye camera and the intrinsic matrix and extrinsic parameters of the second fisheye camera output by S2; wherein the intrinsic matrix and extrinsic parameters of the first and second fisheye cameras are used to calculate the relative pose between views and the back projection of the three-dimensional point cloud.
[0130] Module composition definition: a human three-dimensional reconstruction module based on optimal views is constructed, which is composed of a feature matching and pose estimation submodule, a dense depth map estimation submodule and a point cloud reconstruction submodule.
[0131] Processing process: feature matching and pose estimation: the view images corresponding to the optimal unoccluded view combination are input into the feature matching and pose estimation submodule, the improved ORB algorithm is used to extract the feature points and binary descriptors of each image, and the number of feature points of the improved ORB algorithm is set to 3000; the feature points between different view images are associated based on the Hamming distance to form a cross-view feature point matching pair; the essential matrix E between views is solved by the 5-point method, and the relative pose between views is solved by combining the camera intrinsic matrix, the relative pose includes a relative rotation matrix and a relative translation vector; the reprojection error of the relative pose is verified to be ≤1.0 pixels.
[0132] Dense depth map estimation: the view images corresponding to the optimal unoccluded view combination and the relative pose between views are input into the dense depth map estimation submodule, and the submodule uses a lightweight MVSNet variant model; the number of network channels of the model is 1 / 2 of the original MVSNet model, and 4 feature extraction layers are retained; the depth value of each pixel in each view image is predicted by the model, the depth value unit is mm, and the depth error is ≤0.3 mm; based on the prediction result, the dense depth map corresponding to each view image is generated.
[0133] Point cloud reconstruction: input the dense depth map, the intrinsic matrix and the extrinsic parameters of the first and second fisheye cameras into the point cloud reconstruction submodule. First, perform depth map back projection to obtain a sparse human three-dimensional point cloud. Then, perform point cloud densification using an improved PMVS algorithm with a neighborhood search radius of 0.02 m and a minimum matching point number of 15. Perform surface diffusion and optimization on the sparse human three-dimensional point cloud to supplement the three-dimensional points in the weak texture area and generate a dense human three-dimensional point cloud. Assign a weight to each point in the dense human three-dimensional point cloud based on the occlusion probability of each view output by S4, with the weight value = 1-occlusion probability of the corresponding view. Retain high-confidence three-dimensional points with a weight value ≥ 0.7 to form a preliminary human three-dimensional point cloud model.
[0134] Output result: output the preliminary human three-dimensional point cloud model, which is stored in PLY format and contains the three-dimensional coordinates (x, y, z) and weight values of each point. All point coordinates are unified in the world coordinate system.
[0135] S6, construct a spatiotemporal correlation and semantic weighting point cloud fusion module to process the multiple rounds of preliminary human three-dimensional point cloud models, the spatiotemporal correlation dataset, and the semantic masks of the six virtual perspective plane images, and output a continuous and complete human three-dimensional point cloud model.
[0136] Step S6 includes the following steps:
[0137] S61, based on step S54, identify the common view of the multiple rounds of preliminary human three-dimensional point cloud models, which is a virtual view that repeatedly appears in the optimal view combination corresponding to the preliminary human three-dimensional point cloud models of different rounds. Based on the extrinsic parameters [R i |t i ] corresponding to the common view, unify all preliminary point clouds in the same world coordinate system, and finely register the overlapping areas of the multiple rounds of preliminary human three-dimensional point cloud models using the ICP algorithm to output the spatially aligned multiple rounds of human three-dimensional point clouds.
[0138] S62, extract the human key parts in the multiple rounds of human three-dimensional point cloud models based on the improved lightweight CNN model of MobileNetV2 and assign weights according to their clinical importance. Combine the point cloud with the rigid body model of the mechanical arm motion and the elastic deformation model of the human tissue. Prioritize the points with high weight values greater than or equal to the preset value in the high-weight parts, and use weighted average smoothing for the point cloud in the low-weight parts. Output the semantically weighted fused point cloud.
[0139] S63, based on the semantic weighted fused point cloud, establish the point cloud time sequence with timestamp, predict the dynamic change of the key part point cloud of the human body with Kalman filter, construct the space-time constraint graph, minimize the overall error through graph optimization, complete the human body structure with Poisson reconstruction in non-overlapping areas, trigger global optimization when the same optimal view combination appears again, correct the cumulative error, and finally generate a continuous and complete human body three-dimensional point cloud model;
[0140] Input data is clear: input three types of data, which are the preliminary human body three-dimensional point cloud model output by multiple rounds of S5 (i.e. multiple preliminary human body three-dimensional point cloud models obtained after performing S5 at different times), the time sequence pose data of the surgical robot arm end in the space-time correlation data set output by S1 (including the global coordinates of the robot arm end at each time and the timestamp), and the semantic mask of the six virtual view images output by S4;
[0141] Module composition definition: construct a space-time correlation and semantic weighted point cloud fusion module composed of coordinate alignment sub-module, semantic weighted fusion sub-module and time correlation update sub-module;
[0142] Processing process: coordinate alignment (spatial correlation): input the multiple preliminary human body three-dimensional point cloud models into the coordinate alignment sub-module, first identify the common view angle between the multiple preliminary human body three-dimensional point cloud models, the common view angle is defined as the repeated virtual view angle in the optimal unobstructed view angle combination corresponding to the different rounds of preliminary point cloud; based on the fish-eye camera extrinsic parameter corresponding to the common view angle, unify all the preliminary human body three-dimensional point cloud models to the same world coordinate system; for the overlapping area of the multiple preliminary human body three-dimensional point cloud models, use ICP algorithm for fine registration, the iteration number of ICP algorithm is set to 50 times, and the convergence threshold is set to 0.01mm; the average distance error of the point cloud in the overlapping area after registration is ≤0.3mm, and the spatially aligned multiple human body three-dimensional point cloud models are output;
[0143] Semantic weighted fusion (semantic association): input the spatially aligned multi-round human body three-dimensional point cloud model and the semantic mask of six virtual perspective planar images into the semantic weighted fusion sub-module; extract the human body key parts in the point cloud through the lightweight CNN model improved based on MobileNetV2, and classify the human body key parts into heart region, liver region, limb region and trunk region; assign weights to each key part according to clinical importance, wherein the weight of the heart region is set to 1.0, the weight of the liver region is set to 0.9, the weight of the limb region is set to 0.6, and the weight of the trunk region is set to 0.7; perform multi-round point cloud fusion combining the rigid body model of the surgical robot arm motion and the elastic deformation model of the human body tissue, wherein the rigid body model of the surgical robot arm motion includes linear motion constraint and rotation motion constraint of surgical instruments, and the elastic deformation model of the human body tissue is based on the soft tissue deformation hypothesis of biomechanics; during the fusion process, high-confidence points (weight value ≥ 0.7) of high-weight parts are preferentially retained, and the point cloud of low-weight parts is smoothed by weighted average method, and the human body three-dimensional point cloud model after semantic weighted fusion is output.
[0144] Temporal association update: input the human body three-dimensional point cloud model after semantic weighted fusion and the time sequence data of the end of the surgical robot arm into the temporal association update sub-module; establish the time sequence relationship of multi-round point clouds based on time stamp; predict the dynamic changes of human body key part point clouds over time by using Kalman filtering algorithm, and the process noise covariance Q of the Kalman filtering algorithm is 1e-6I and the observation noise covariance R is 1e-4I; construct a space-time constraint graph, in which the nodes are point clouds at each time, and the edges are the constraint relationship between adjacent time point clouds; minimize the overall point cloud error by using a graph optimization algorithm; for the non-overlapping area of multi-round point clouds, use the Poisson reconstruction algorithm to complete the human body structure, and the reconstruction depth of the Poisson reconstruction algorithm is set to 10 levels; when the same optimal unoccluded view combination appears again, trigger the global optimization operation to correct the cumulative error in the multi-round point cloud fusion process; finally generate a continuous and complete human body three-dimensional point cloud model;
[0145] Output result: output the continuous and complete human body three-dimensional point cloud model, the storage format of which is PLY, the point cloud error is ≤0.8mm, the update delay is ≤50ms, and it contains the three-dimensional coordinates, weight values and time stamps of each point, and the coordinates are unified in the world coordinate system.
[0146] S7, construct a robot vision collaborative control module, process the optimal view combination, space-time association data set and continuous and complete human body three-dimensional point cloud model, and output collaborative control instructions and continuously updated point clouds.
[0147] Input data definition: input three types of data, respectively, the optimal unoccluded view combination (including the occlusion probability of each view) output by S4, the time sequence pose data of the surgical robot arm end in the spatiotemporal correlation data set output by S1 (real-time updated global coordinates and orientation of the robot arm end), and the continuous complete human body three-dimensional point cloud model output by S6;
[0148] Module composition definition: construct a robot arm-vision collaborative control module composed of a fixed camera control submodule, a vision-assisted robot arm control submodule, and a collaborative protocol submodule; wherein the collaborative protocol submodule is constructed based on the SurgArm-VisionProtocol protocol, and is used to realize the time sequence synchronization of the surgical robot arm motion instruction and the camera view control instruction;
[0149] Processing process: fixed camera control (basic mode): when the dual fisheye camera unit is fixedly installed (without vision-assisted robot arm), input the optimal unoccluded view combination and the time sequence pose data of the surgical robot arm end into the fixed camera control submodule; real-time monitor the position change of the surgical robot arm end, if the robot arm end enters a high-occlusion area of a virtual view (occlusion probability Poccl≥0.7), immediately generate a view switching instruction; send the view switching instruction to the central control unit, and switch to a backup optimal unoccluded view combination (the backup combination is the suboptimal combination selected in S4); the view switching delay is ≤30 ms, ensuring that the three-dimensional reconstruction process is not interrupted;
[0150] Vision-assisted robot arm control (enhanced mode, optional configuration): when the system is configured with a vision-assisted robot arm (used to adjust the position of the dual fisheye camera), input the optimal unoccluded view combination requirement into the vision-assisted robot arm control submodule; calculate the target pose of the vision-assisted robot arm according to the view requirement, and the target pose positioning accuracy is ≤0.1 mm; generate a robot arm pose adjustment instruction based on the target pose, and send the instruction to the vision-assisted robot arm control system; after receiving the instruction, the vision-assisted robot arm adjusts to the target pose within ≤100 ms, ensuring that the dual fisheye camera can collect view images corresponding to the optimal unoccluded view combination;
[0151] Collaborative synchronization: through the collaborative protocol submodule, the motion instruction (including the operation instruction of the end effector) of the surgical robot arm and the view control instruction (including the view switching instruction and the robot arm pose adjustment instruction) of the camera are stamped with a unified timestamp, ensuring the time sequence synchronization of the two types of instructions, and the synchronization delay is ≤1 ms; based on the continuous complete human body three-dimensional point cloud model output by S6, verify the matching of the surgical robot arm operation and the reconstruction result; if the projection position of the robot arm end in the point cloud and the actual image position error is ≥1 mm, generate an error correction instruction; send the error correction instruction to the corresponding execution unit (central control unit or vision-assisted robot arm control system) to adjust the camera view or the robot arm pose;
[0152] Output results: Output robotic arm-vision collaborative control commands, which include viewpoint switching commands, robotic arm pose adjustment commands, and error correction commands, stored in TXT format; Simultaneously output a continuously updated 3D human body point cloud model, stored in PLY format, updated in real time with surgical operations, with an update frequency ≥20Hz, and coordinates unified in the world coordinate system.
[0153] In the above embodiments, some algorithms, settings, and parameter operations are all conventional technical means in the field. They are necessary operations for carrying out activities and are all conventional adjustments in this embodiment. They will not be elaborated here.
[0154] like Figure 2 As shown, a camera-based human body 3D reconstruction system, applied to the multi-view point cloud-based 3D reconstruction method described in the above embodiments, includes:
[0155] Acquisition hardware is used to acquire surgical scene images, output spatial pose data, and adjust the position of objects;
[0156] Install and support hardware to fix the position of the data acquisition hardware and the object it supports;
[0157] Calibration and auxiliary hardware are used to calibrate the basic parameters of the acquisition hardware and optimize the accuracy of the parameters;
[0158] Control and synchronization hardware is used to control the acquisition hardware, installation and support hardware, and calibration and auxiliary hardware.
[0159] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
[0160] The present invention has been described above with reference to the accompanying drawings. Obviously, the implementation of the present invention is not limited to the above-described manner. Any improvements made using the inventive concept and technical solution of the present invention, or the direct application of the inventive concept and technical solution of the present invention to other situations without modification, are all within the protection scope of the present invention.
Claims
1. A camera-based human three-dimensional reconstruction method, characterized in that, The method comprises the following steps: S1, a dual fisheye camera mechanical arm cooperative data acquisition system is constructed, which is used for acquiring multi-source data of a surgical scene and generating a time-space correlation data set; S2, a dual fisheye camera dynamic calibration module is constructed, which is used for processing the time-space correlation data set and outputting internal and external parameters of the dual fisheye camera; S3, a panoramic image acquisition and stitching module is constructed, which is used for processing the time-space correlation data set and the internal and external parameters of the dual fisheye camera and outputting a stitched panoramic image and six virtual perspective plane images; S4, a dynamic view selection module of mechanical arm pose guidance is constructed, which is used for processing the six virtual perspective plane images, the time-space correlation data set and the internal and external parameters of the dual fisheye camera and outputting an optimal view combination and an optimal view image corresponding to the optimal view combination; S5, a human three-dimensional reconstruction module is constructed, which is used for processing the optimal view image and the internal and external parameters of the dual fisheye camera and outputting a preliminary human three-dimensional point cloud model; S6, a time-space correlation and semantic weighted point cloud fusion module is constructed, which is used for processing a plurality of preliminary human three-dimensional point cloud models, the time-space correlation data set and semantic masks of the six virtual perspective plane images and outputting a continuous and complete human three-dimensional point cloud model; S7, a mechanical arm vision cooperative control module is constructed, which is used for processing the optimal view combination, the time-space correlation data set and the continuous and complete human three-dimensional point cloud model and outputting cooperative control instructions and continuously updated point clouds.
2. The camera-based human three-dimensional reconstruction method of claim 1, wherein, The step S1 comprises the following steps: S11, two fisheye cameras with a 180° view angle are used, which are fixedly installed on a symmetrical C-arm support of an operating table to form the dual fisheye camera; a mechanical arm is used, which is provided with an end effector and a high-precision encoder; a hardware synchronization component is used, which is connected to the dual fisheye camera and the mechanical arm respectively and is used for realizing time sequence synchronization during data acquisition; a central controller is used, which is connected to the dual fisheye camera, the mechanical arm and the hardware synchronization component and is denoted as a dual fisheye camera mechanical arm cooperative data acquisition unit; S12, the dual fisheye camera mechanical arm cooperative data acquisition unit is started to synchronously capture dual fisheye original images of a surgical scene and time sequence pose data of the end effector of the mechanical arm; the central controller marks the dual fisheye original images and the time sequence pose data with a unified microsecond-level time stamp to generate the time-space correlation data set.
3. The camera-based human three-dimensional reconstruction method of claim 2, wherein, The step S2 comprises the following steps: S21, based on the time-space correlation data set, a checkerboard calibration board image and a surgical instrument marker motion image are extracted, the checkerboard calibration board image is used for static calibration initialization, and the surgical instrument marker motion image contains a coding point projection of a surgical instrument marker in the dual fisheye original image and is used for dynamic optimization; S22, based on the checkerboard calibration board image, calculating the intrinsic matrix K of the fisheye camera i and the extrinsic [R i |t i ], the intrinsic matrix K i Specifically: wherein f x represents the focal length of the fisheye camera in the x-axis direction, f y represents the focal length of the fisheye camera in the y-axis direction, c x represents the coordinate of the principal point of the fisheye camera in the x-axis of the image coordinate system, c y represents the coordinate of the principal point of the fisheye camera in the y-axis of the image coordinate system; extrinsic parameter [R i |t i ] includes a rotation matrix R i and a translation vector t i , based on the intrinsic parameter matrix K i and the extrinsic parameter [R i |t i ], the projection matrix P i of the fisheye camera, the global coordinates P e of the end of the mechanical arm and the projection matrix P i are calculated as follows: P e = [X e , Y e , Z e ] T P i = K i [R i |t i ] S23, based on the surgical instrument marker, setting a straight trajectory, a circular trajectory and a simulated surgery trajectory, the simulated surgery trajectory is used to simulate the operation path of the surgical instrument, synchronously collecting the encoded point projection, calculating the theoretical pixel coordinates and the actual pixel coordinates of each encoded point projection, and based on the theoretical pixel coordinates and the actual pixel coordinates, calculating the pixel coordinate error, and then based on the pixel coordinate difference, iteratively optimizing the intrinsic matrix K i and the extrinsic parameter [R i |t i ] until the pixel coordinate difference is less than or equal to a preset value, and finally outputting the intrinsic matrix K i , the extrinsic parameter [R i |t i ] and the projection matrix P i of the two fisheye cameras.
4. The camera-based human three-dimensional reconstruction method of claim 3, wherein, The step S3 comprises the following steps: S31, based on the internal parameter matrix K i and the external parameter [R i |t i ], the fisheye camera is rectified, and the double fisheye original image is denoised, brightness adjusted and color corrected, and two 180° fisheye images after preprocessing are output. S32, based on an improved ORB algorithm, feature points and binary descriptors of each of the 180° fisheye images are extracted, feature point matching pairs are screened based on a Hamming distance, and false matching pairs are removed through a RANSAC algorithm to output a corresponding relationship of feature points of two 180° fisheye images; S33, obtain a homography matrix H of the two 180-degree fisheye images through the feature point correspondence, perform perspective transformation alignment on one of the 180-degree fisheye images based on the homography matrix H, and then perform pixel-level fusion on the overlapping area of the 180-degree fisheye image and another 180-degree fisheye image to generate the panoramic image; S34, decouple the panoramic image into the six virtual perspective planar images, and the six virtual perspective planar images include a forward virtual perspective planar image, a backward virtual perspective planar image, a left virtual perspective planar image, a right virtual perspective planar image, an upward virtual perspective planar image, and a downward virtual perspective planar image.
5. The camera-based human three-dimensional reconstruction method of claim 4, wherein, The step S4 includes the following steps: S41, output semantic masks of the six virtual perspective planar images based on a U-Net model; S42, extracting the global coordinate P of the end of the mechanical arm at the current moment based on the end timing pose data e , combined with the projection matrix P i , calculating the pixel coordinate u of the end of the mechanical arm on the six virtual visual angle plane images i , which is specifically represented as: u i = [u, v] T When the pixel coordinate u i falls in the human body region in the semantic mask, the included angle θ between the line connecting the end of the mechanical arm and the optical center of the corresponding fisheye camera and the normal vector of the human body surface is calculated, the distance d from the end of the surgical mechanical arm to the corresponding fisheye camera is calculated, and finally the occlusion probability P of the front virtual perspective planar image, the rear virtual perspective planar image, the left virtual perspective planar image, the right virtual perspective planar image, the upper virtual perspective planar image and the lower virtual perspective planar image is calculated occl , the occlusion probability P occl is calculated as follows: where d max = 500 mm, the maximum distance of the operating table workspace.
6. The camera-based human three-dimensional reconstruction method of claim 5, wherein, The step S4 further includes the following steps: S43, calculate occlusion probability heat maps of the forward virtual perspective planar image, the backward virtual perspective planar image, the left virtual perspective planar image, the right virtual perspective planar image, the upward virtual perspective planar image, and the downward virtual perspective planar image within 0.5-1 seconds in the future based on the end time sequence pose data, the surgical instrument type in the surgical instrument marker, and the semantic masks, wherein a larger pixel value of the occlusion probability heat map represents a higher occlusion probability of the corresponding region; S44, performing static screening to screen the current time shielding probability P occl If the virtual view angle is less than the preset value, dynamic screening is further performed, specifically, from the static screening result, the shielding probability P in the future 0.5-1 seconds is screened occl If the virtual view angle is less than the preset value, dynamic screening is further performed, specifically, from the static screening result, the shielding probability P in the future 0.5-1 seconds is screened occl If the virtual view angle is less than the preset value, dynamic screening is further performed, specifically, from the static screening result, the shielding probability P in the future 0.5-1 seconds is screened 7. The camera-based human three-dimensional reconstruction method of claim 6, wherein, The step S5 includes the following steps: S51, the feature points and binary descriptors of the optimal view image are extracted by using the improved ORB algorithm, the cross-view matching pairs of the optimal view combination are associated based on the Hamming distance, the essential matrix E between the views in the optimal view combination is solved by using the 5-point method, and the intrinsic matrix K i The relative pose between the views is solved, and the relative pose includes a relative rotation matrix R rel and a relative translation vector t rel The reprojection error of the relative pose is verified to be less than or equal to 1.0 pixels. S52, construct a lightweight MVSNet variant model based on a MVSNet model, the lightweight MVSNet variant model is half of the MVSNet model, and four feature extraction layers are retained; S53, input the optimal perspective image and the relative pose into the lightweight MVSNet variant model to predict a depth value of each pixel in each optimal perspective image, and generate a dense depth map corresponding to each optimal perspective image based on a prediction result; S54, perform back projection operation on the dense depth map to obtain a sparse human body three-dimensional point cloud, and the operation is specifically as follows: wherein Z represents a pixel depth value, (u, v) is a pixel coordinate, (c x , c y ) is a camera principal point coordinate; Perform point cloud densification operation on the sparse human body three-dimensional point cloud through a PMVS algorithm to form a preliminary human body three-dimensional point cloud model.
8. The camera-based human three-dimensional reconstruction method of claim 7, wherein, The step S6 includes the following steps: S61、based on the step S54, identifying a common view angle of multiple rounds of the preliminary human body three-dimensional point cloud model, the common view angle being a virtual view angle repeatedly appearing in optimal view angle combinations corresponding to the preliminary human body three-dimensional point cloud model of different rounds, and based on the common view angle corresponding to the external parameter [R i |t i ], unifying all the preliminary point clouds to the same world coordinate system, performing fine registration on the overlapping areas of multiple rounds of the preliminary human body three-dimensional point cloud model using the ICP algorithm, and outputting the spatially aligned multiple round human body three-dimensional point cloud; S62, extract human body key parts in a multi-round human body three-dimensional point cloud model based on a MobileNetV2 improved lightweight CNN model, and weight the human body key parts according to clinical importance, combine a rigid body model of mechanical arm movement and an elastic deformation model of human body tissue to fuse point clouds, preferentially retain points with a weight greater than or equal to a preset value in a high-weight part, and use weighted average smoothing for point clouds in a low-weight part, and output point clouds after semantic weighted fusion.
9. The camera-based human three-dimensional reconstruction method of claim 8, wherein, The step S6 further includes the following steps: S63, based on the point clouds after semantic weighted fusion, establish a point cloud time sequence with a timestamp, predict dynamic changes of human body key part point clouds by Kalman filtering, construct a space-time constraint graph, minimize the overall error through graph optimization, use Poisson reconstruction to complete human body structure in a non-overlapping area, trigger global optimization when the same optimal perspective combination appears again, correct cumulative errors, and finally generate the continuous and complete human body three-dimensional point cloud model.
10. A camera-based human three-dimensional reconstruction system, characterized in that, The camera-based human three-dimensional reconstruction method according to any one of claims 1-9, comprising: acquisition hardware for acquiring surgical scene images, outputting spatial pose data, and adjusting the position of an object; mounting and supporting hardware for fixing the position of the acquisition hardware and carrying the object; calibration and auxiliary hardware for calibrating the basic parameters of the acquisition hardware and optimizing the accuracy of the parameters; control and synchronization hardware for controlling the acquisition hardware, the mounting and supporting hardware, and the calibration and auxiliary hardware.
Citation Information
Cited By
Depth-estimation-free pure rotation optimization unmanned aerial vehicle panoramic image splicing method
CN122089565A
A depth-estimation-free pure-rotation optimization unmanned aerial vehicle panoramic image stitching method
CN122089565B