Three-dimensional visual disjunction method and system for perception data
By combining time synchronization, spatial registration, and multi-scale feature fusion of multimodal data with Hessian matrix analysis and ray casting techniques, the problems of spatiotemporal misalignment and complex structure extraction of multimodal data were solved, achieving accurate separation and clear visualization of microscopic anisotropic structures.
Patent Information
- Application Number
- CN202511652560.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies suffer from spatiotemporal misalignment and registration errors when processing multimodal sensing data, making it difficult to accurately separate the microscopic anisotropic features of complex structures. This results in fractures or voids in the three-dimensional extraction results, particularly in geological core scanning and biological tissue microscopic imaging, where serious deviations occur.
By generating a spatiotemporally aligned multimodal data cube through time synchronization and spatial registration, multi-scale feature extraction and hierarchical fusion are employed, combined with Hessian matrix feature analysis and ray projection technology, and the data is dynamically projected onto a virtual observation space for interactive extraction and topological connectivity analysis, thereby generating an enhanced 3D visualization.
It achieves accurate identification and enhancement of anisotropic microstructures, improves the integrity and accuracy of complex structure extraction, overcomes the visual clutter problem of traditional methods, and supports clear structural observation from multiple perspectives and scales.
Smart Images

Figure CN121543064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent analysis technology for three-dimensional sensing data, specifically to a method and system for three-dimensional visualization extraction of sensing data. Background Technology
[0002] Accurate perception and understanding of the 3D environment is crucial. Currently, the industry widely uses heterogeneous sensors such as LiDAR, visual cameras, and millimeter-wave radar for environmental data acquisition. However, existing technologies face severe challenges in processing multimodal perception data: different sensors have inherent differences in sampling frequency, data format, and spatiotemporal reference, leading to spatiotemporal misalignment during data fusion; traditional point cloud registration methods lack sufficient registration accuracy for minute structural features in dynamic environments, especially in scenes with a large amount of repetitive textures or sparse features, where registration errors are further amplified. In addition, existing 3D visualization methods mostly rely on explicit geometric models, making it difficult to effectively represent and extract targets with complex internal structures or fuzzy boundaries, such as non-rigid structures like rock fissures and microvascular networks in biological tissues.
[0003] The existing technology has the following shortcomings:
[0004] Existing technologies have significant limitations in handling the 3D extraction of complex structures with microscopic anisotropy, particularly in fields such as geological core scanning and biological tissue microscopy. When faced with numerous randomly oriented, morphologically varied micron-sized tubular pores or fibrous structures, traditional volume rendering methods lack the ability to perceive feature orientation, making it impossible to accurately separate interwoven microstructures. Furthermore, segmentation methods based on global thresholds suffer from significant differences in local feature intensity, leading to numerous breaks or voids in the extraction results. More seriously, when these microstructures are anisotropically distributed in 3D space, existing methods struggle to maintain their topological connectivity, resulting in severe biases in subsequent quantitative analysis. This problem is particularly prominent in core pore structure analysis for energy exploration and microvascular network reconstruction in medical imaging, yet it remains unresolved. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for extracting three-dimensional visualization data to solve the problems mentioned above.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] A method for extracting three-dimensional visualization data includes the following steps:
[0008] S1: Acquire multimodal sensing data streams from heterogeneous sensors, and perform time synchronization and spatial registration on the multimodal sensing data streams to generate a spatiotemporally aligned multimodal data cube;
[0009] S2: Multi-scale feature extraction and hierarchical fusion of multimodal data cubes are performed. By decomposing and recombining the features of multimodal data cubes at different scales, a unified multi-scale feature field is generated.
[0010] S3: The unified multi-scale feature field is dynamically projected onto a virtual observation space. Through ray casting and volume rendering techniques, the geometric and semantic attributes in the feature field are directly synthesized into a two-dimensional image sequence with visual depth, thereby constructing a scene visualization representation that supports real-time visual exploration.
[0011] S4: Perform interactive extraction of the scene visualization representation based on the feature field gradient. By responding to the user's annotations on the two-dimensional image sequence, it is back-mapped to the multi-scale feature field, and the feature subset of the target structure is separated according to the intrinsic gradient information of the feature field.
[0012] S5: Based on the extracted feature subset, a dynamic parameter driving interface is associated in the real-time graphics pipeline. By adjusting the optical property mapping function, an enhanced 3D visualization of the target structure at any viewpoint and scale is synthesized and presented in real time.
[0013] As a further aspect of the present invention: the generation of the spatiotemporally aligned multimodal data cube specifically includes:
[0014] Cross-sensor clock drift compensation and data packet reordering are performed on the multimodal sensing data stream. A sequence alignment method based on dynamic time warping is adopted to unify sensor data with different sampling frequencies to the same time base and generate a time-synchronized continuous data frame sequence.
[0015] Construct a spatial transformation chain between sensors, optimize the rotation and translation parameters in the spatial transformation chain by solving the geometric feature correspondence in the observation data of adjacent sensors, and establish an accurate mapping relationship from each sensor coordinate system to a unified world coordinate system.
[0016] The time-synchronized continuous data frame sequence is projected onto a unified world coordinate system according to a precise mapping relationship. The discrete spatial sampling points are then converted into a spatiotemporally aligned multimodal data cube represented by a regular grid through voxelization.
[0017] As a further aspect of the present invention: the generation of a unified multi-scale feature field specifically includes:
[0018] An anisotropic scale space is constructed, and by setting different scale parameters along the three principal axes of the space, a three-dimensional convolution process is performed on the multimodal data cube to generate a multi-scale three-dimensional tensor sequence with direction specificity.
[0019] Directional enhancement is performed on each scale feature in the multi-scale three-dimensional tensor sequence. By calculating the eigenvalue combination of the three-dimensional Hessian matrix in the neighborhood of each voxel, tubular and sheet-like structural features are extracted and enhanced to generate directionally enhanced feature tensors.
[0020] Nonlinear superposition of directional enhancement feature tensors at different scales allows for the fusion of coarse-scale semantic information with fine-scale geometric details while maintaining the continuity of spatial structure, thereby generating a unified multi-scale feature field.
[0021] As a further aspect of the present invention: the dynamic projection of the unified multi-scale feature field onto a virtual observation space specifically includes:
[0022] A virtual observation sphere centered on the observer's position is established, and the three-dimensional spatial coordinates in the multi-scale feature field are mapped to the two-dimensional spherical coordinates of the observation sphere through spherical projection transformation.
[0023] A multi-resolution latitude and longitude grid is constructed on the surface of the observation sphere, and the grid density is adaptively adjusted according to the observation distance. A high-density grid is used in areas close to the observer, and a low-density grid is used in areas far away.
[0024] By dynamically updating the spherical projection transformation parameters according to the change of observation perspective, a continuous and smooth mapping from multi-scale feature fields to virtual observation space is achieved, establishing a two-way correspondence between feature field data and observation space.
[0025] As a further aspect of the present invention: the construction of a scene visualization representation supporting real-time visual exploration specifically includes:
[0026] Sampling rays are emitted from each pixel in the virtual observation space, and three-dimensional sampling with adaptive step size is performed along the ray path in the multi-scale feature field.
[0027] Optical property transformation is performed on the multi-scale features of each sampling point, converting geometric features into optical absorption coefficients and semantic features into scattering intensity;
[0028] The optical effects are accumulated along each light path, and the final pixel color and depth value are obtained through depth synthesis calculation;
[0029] Based on the continuous changes in the observation perspective, a sequence of two-dimensional images with correct occlusion relationships and depth information is generated in real time to construct a scene visualization representation that supports multi-view exploration.
[0030] As a further aspect of the present invention: the interactive extraction of the scene visualization representation based on the feature field gradient specifically includes:
[0031] A bidirectional mapping relationship between the feature field space and image pixels is established, and the user's annotation operation on the two-dimensional image sequence is converted into the initial boundary constraint in the three-dimensional space. The initial spatial position set in the multi-scale feature field is obtained by back projection calculation.
[0032] Based on the initial set of spatial locations, gradient flow field-based region growing is performed in a multi-scale feature field. By tracking the changing direction of the feature field gradient, the growing region is expanded along the path with the highest feature similarity to form a candidate feature subset.
[0033] Topological connectivity analysis and boundary optimization are performed on candidate feature subsets. By calculating the rate of change vector of the feature field in three-dimensional space, connected regions with similar gradient features are identified. Based on the preset structural integrity threshold, feature subsets of the target structure are accurately separated from the multi-scale feature field.
[0034] As a further aspect of the present invention: the execution of region growth based on gradient flow field specifically includes:
[0035] Calculate the three-dimensional gradient vector at each voxel position in the multi-scale feature field to construct a gradient flow field describing the direction and intensity of feature changes, wherein the gradient vector is obtained by calculating the spatial difference of the eigenvalues between adjacent voxels;
[0036] Starting from each position in the initial spatial location set, path tracing is performed simultaneously along both the positive and negative directions of the gradient vector in the gradient flow field. During the tracing process, the feature similarity between adjacent voxels is calculated in real time, and the growth path is extended only when the similarity is higher than a set threshold.
[0037] Intersecting growth paths are merged by analyzing the consistency of characteristic distribution in the intersection area of the paths, and adjacent paths with similar gradient change patterns are merged into the same growth region.
[0038] Based on the spatial distribution of all growth paths, a connectivity graph structure is constructed. By removing isolated branches and filling internal holes, a spatially continuous and boundary-complete subset of candidate features is formed.
[0039] As a further aspect of the present invention: the formation of a spatially continuous and boundary-complete subset of candidate features further includes:
[0040] Based on the spatial distribution of candidate feature subsets, the feature field curvature distribution of each voxel position is calculated. By analyzing the consistency of curvature changes between adjacent voxels, connected regions with similar curvature features are divided into the same topological unit.
[0041] Boundary voxel optimization is performed on the divided topological units. By calculating the difference in characteristic gradients between the boundary voxel and its adjacent inner and outer voxels, the boundary position is adjusted to the position with the largest gradient change, thus forming a boundary transition.
[0042] Based on the optimized boundary, the structural integrity of each topological unit is verified. By detecting voids and fractures inside the unit, interpolation repair is performed using the continuity of the feature field.
[0043] The verified topological units are merged according to their spatial adjacency, and the feature subset of the target structure is accurately separated from the multi-scale feature field based on the preset structural integrity threshold.
[0044] As a further aspect of the present invention: the real-time synthesis and presentation of the enhanced 3D visualization of the target structure at any viewpoint and scale specifically includes:
[0045] Construct multiple optical transfer function groups, design independent optical transmission paths for different semantic features in the feature subset, and realize beam splitting rendering of different structures by establishing the mapping relationship between feature values and optical parameters;
[0046] Based on the spatial relationship between the observation viewpoint and the target structure, the sampling density of each pixel is dynamically calculated. Adaptive supersampling is used in the structural boundary region, while the basic sampling rate is maintained in the uniform region, so as to achieve a balance between detail preservation and rendering efficiency.
[0047] During the rendering process, depth-aware color enhancement is implemented, dynamically adjusting color saturation and brightness based on pixel depth values to make the colors of near structures vivid and the colors of distant structures soft, thereby enhancing the sense of depth and layering in the scene.
[0048] By dynamically adjusting the depth of field effect through real-time refocusing technology, the blur radius is automatically calculated based on the user's interaction focus, so that the focus area is clearly presented and the non-focus area is naturally blurred.
[0049] A three-dimensional visualization and extraction system for perceptual data, comprising:
[0050] The data synchronization and registration module is used to acquire multimodal sensing data streams from heterogeneous sensors, and to perform time synchronization and spatial registration on the multimodal sensing data streams to generate a spatiotemporally aligned multimodal data cube.
[0051] The feature fusion module is used to extract and hierarchically fuse multi-scale features of multimodal data cubes. It generates a unified multi-scale feature field by decomposing and recombining the features of multimodal data cubes at different scales.
[0052] The dynamic visualization module is used to dynamically project a unified multi-scale feature field onto a virtual observation space. Through ray casting and volume rendering techniques, the geometric and semantic attributes in the feature field are directly synthesized into a two-dimensional image sequence with visual depth, thereby constructing a scene visualization representation that supports real-time visual exploration.
[0053] The interactive extraction module is used to perform interactive extraction of scene visualization representation based on feature field gradient. By responding to user annotations on the two-dimensional image sequence, it back-maps to a multi-scale feature field and separates the feature subset of the target structure based on the intrinsic gradient information of the feature field.
[0054] The enhanced rendering module, based on the extracted feature subset, associates a dynamic parameter driving interface in the real-time graphics pipeline. By adjusting the optical property mapping function, it synthesizes and renders an enhanced 3D visualization of the target structure at any viewpoint and scale in real time.
[0055] The beneficial effects of this invention are:
[0056] (1) By introducing directional enhancement processing based on Hessian matrix feature analysis, this invention can accurately identify and enhance the local directional features of anisotropic microstructures, effectively solving the problems of fracture and topological structure destruction that occur when extracting anisotropic structures such as tubular and fibrous structures using traditional methods, and significantly improving the integrity and accuracy of complex structure extraction.
[0057] (2) By constructing a set of multiple optical transfer functions and combining them with a depth-aware color enhancement mechanism, this invention achieves differentiated rendering of target structures with different depths and semantic features, overcoming the visual confusion problem of traditional volume rendering when expressing complex internal structures, and enabling operators to more clearly distinguish and observe structural details at different scales. Attached Figure Description
[0058] The invention will now be further described with reference to the accompanying drawings.
[0059] Figure 1 This is a flowchart of the method of the present invention;
[0060] Figure 2 This is a flowchart of the system in this invention. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Please see Figure 1 As shown, the present invention is a method for extracting three-dimensional visualization data, comprising the following steps:
[0063] S1: Acquire multimodal sensing data streams from heterogeneous sensors, and perform time synchronization and spatial registration on the multimodal sensing data streams to generate a spatiotemporally aligned multimodal data cube;
[0064] S2: Multi-scale feature extraction and hierarchical fusion of multimodal data cubes are performed. By decomposing and recombining the features of multimodal data cubes at different scales, a unified multi-scale feature field is generated.
[0065] S3: The unified multi-scale feature field is dynamically projected onto a virtual observation space. Through ray casting and volume rendering techniques, the geometric and semantic attributes in the feature field are directly synthesized into a two-dimensional image sequence with visual depth, thereby constructing a scene visualization representation that supports real-time visual exploration.
[0066] S4: Perform interactive extraction of the scene visualization representation based on the feature field gradient. By responding to the user's annotations on the two-dimensional image sequence, it is back-mapped to the multi-scale feature field, and the feature subset of the target structure is separated according to the intrinsic gradient information of the feature field.
[0067] S5: Based on the extracted feature subset, a dynamic parameter driving interface is associated in the real-time graphics pipeline. By adjusting the optical property mapping function, an enhanced 3D visualization of the target structure at any viewpoint and scale is synthesized and presented in real time.
[0068] In S1, multimodal sensing data streams from heterogeneous sensors are acquired, and the multimodal sensing data streams are time-synchronized and spatially registered to generate a spatiotemporally aligned multimodal data cube, specifically including:
[0069] During data acquisition, at least two different types of sensors are used simultaneously to collect environmental data. LiDAR acquires 3D point cloud data, a vision camera acquires 2D image data, and millimeter-wave radar acquires target point data. All sensors provide a time reference through a unified clock source, recording precise timestamp information in each data packet.
[0070] During time synchronization, clock drift of each sensor is first compensated. The drift trend of each sensor's clock is detected by calculating the timestamp difference between adjacent data packets. A sequence alignment method based on dynamic time warping is employed to unify sensor data from different sampling frequencies to the same time base. The specific calculation process includes: first, extracting feature sequences from each sensor's data stream; then, calculating the cumulative distance matrix between feature sequences; and finally, finding the minimum cost path through backtracking to determine the correspondence between sensor data frames. This process resamples data from different sampling frequencies to the same time point, generating a continuous sequence of time-synchronized data frames.
[0071] During spatial registration, a spatial transformation chain is constructed between sensors. Geometric features are extracted from the observation data of adjacent sensors to establish feature correspondences. Specifically, feature points with significant geometric characteristics are extracted from the lidar point cloud and visual images, and the 3D coordinates of these feature points in their respective sensor coordinate systems are calculated. Then, the optimal solution for the rotation and translation parameters is calculated using singular value decomposition (SVD). The SVD calculation process includes: first, calculating the centroid coordinates of the two point sets; then, calculating the covariance matrix; and finally, performing SVD on the covariance matrix to obtain the rotation matrix and translation vector. This process establishes a precise mapping relationship from each sensor coordinate system to a unified world coordinate system.
[0072] After spatiotemporal alignment, the time-synchronized continuous data frame sequence is projected onto a unified world coordinate system according to the established precise mapping relationship. Discrete spatial sampling points are converted into a regular grid representation through voxelization. The voxel size is set according to application requirements, typically a 0.1m × 0.1m × 0.1m cubic grid. For each voxel, feature values are calculated based on the sensor data points falling within that voxel. If the same voxel contains data from multiple sensors, it is weighted and fused according to preset weighting coefficients. This process ultimately generates a spatiotemporally aligned multimodal data cube, where each voxel contains fused observations from different sensors and has a unified timestamp and spatial coordinates.
[0073] In S2, multi-scale feature extraction and hierarchical fusion are performed on the multimodal data cube. By decomposing and recombining the features of the multimodal data cube at different scales, a unified multi-scale feature field is generated, specifically including:
[0074] When constructing an anisotropic scale space, different scale parameters are set along the three principal axes of the space. The scale parameter σx is set for the X-axis, with a value ranging from 0.5 to 2.0; the scale parameter σy is set for the Y-axis, with a value ranging from 0.5 to 2.0; and the scale parameter σz is set for the Z-axis, with a value ranging from 0.3 to 1.5. The multimodal data cube is then convolved using a 3D Gaussian convolution kernel. The expression for the 3D Gaussian function is: G(x,y,z)=exp(-(x...)...) 2 / 2σx 2 +y 2 / 2σy 2 +z 2 / 2σz 2In this convolution, exp is an exponential function with base e, G(x,y,z) represents a three-dimensional Gaussian function, and the kernel size is automatically determined based on the scale parameter, typically a multiple of 6σ+1. The convolution operation is repeated for each scale combination to generate a direction-specific multi-scale three-dimensional tensor sequence. This sequence contains feature representations at multiple scales, where large-scale features capture overall structural information, while small-scale features preserve detailed information.
[0075] During directional enhancement, for each scale feature in the multi-scale 3D tensor sequence, the 3D Hessian matrix within the neighborhood of each voxel is calculated. The Hessian matrix is a 3×3 symmetric matrix composed of the second-order partial derivatives of the eigenfunctions along the three coordinate axes. The specific calculation process is as follows: first, the gradient of each voxel in the X, Y, and Z directions is calculated; then, the gradient of the gradient is calculated to obtain the second-order partial derivative. The eigenvalues of the Hessian matrix are obtained by solving the characteristic equation, which has the form det(H-λI)=0, where H is the Hessian matrix, λ is the eigenvalue, I is the identity matrix, and det represents the determinant of the matrix. The eigenvalues are sorted by their absolute values and denoted as λ1, λ2, and λ3. The local geometric structure is determined based on the sign and relative magnitude of the eigenvalues: when the absolute value of λ1 is much larger than the other two eigenvalues and is negative, the region belongs to a tubular structure; when the absolute values of λ1 and λ2 are much larger than λ3 and are both negative, the region belongs to a sheet-like structure. By adjusting the magnitude of the eigenvalues, the saliency of the corresponding structure is enhanced, generating a directionally enhanced feature tensor.
[0076] In multi-scale fusion, directional enhancement feature tensors at different scales are nonlinearly superimposed. The nonlinear superposition employs a voxel-by-voxel fusion strategy. For each spatial location, the fusion weights for each scale are determined based on its feature saliency. Feature saliency is calculated based on the feature response intensity at different scales; higher response intensity is assigned a greater weight. Coarse-scale feature tensors primarily provide semantic information, while fine-scale feature tensors primarily provide geometric details. During the fusion process, the feature tensors at each scale are first normalized to the same numerical range, then linearly combined according to the weight coefficients, and finally processed by a nonlinear activation function. The nonlinear activation function uses a modified sigmoid function, with the form f(x) = 1 / (1 + exp(-α(x-β)), where α controls the steepness of the function, ranging from 2.0 to 5.0, and β controls the center position of the function, ranging from 0.3 to 0.7. This process, while maintaining the continuity of the spatial structure, achieves effective fusion of coarse-scale semantic information and fine-scale geometric details, generating a unified multi-scale feature field.
[0077] In S3, a unified multi-scale feature field is dynamically projected onto a virtual observation space. Through ray casting and volume rendering techniques, the geometric and semantic attributes of the feature field are directly synthesized into a two-dimensional image sequence with visual depth, thereby constructing a scene visualization representation that supports real-time visual exploration. Specifically, this includes:
[0078] When establishing a virtual observation space, an observation sphere centered on the observer's viewpoint is first defined. The radius of this sphere is determined based on the spatial extent of the scene to be visualized, typically set to 1.2 to 1.5 times the radius of the smallest circumscribed sphere capable of enclosing the entire multi-scale feature field. The three-dimensional Cartesian coordinates in the feature field are converted to spherical coordinates through a spherical projection transformation, where the horizontal azimuth θ ranges from 0 to 360 degrees, and the vertical elevation φ ranges from -90 to 90 degrees. The projection transformation is calculated as follows: first, the direction vector of the feature point relative to the viewpoint is calculated; then, the three-dimensional coordinates of this vector are converted to azimuth and elevation values in the spherical coordinate system.
[0079] When constructing a multi-resolution latitude and longitude grid on the surface of an observation sphere, the grid density is dynamically adjusted based on the distance between each region on the sphere and the observer. Regions within 2 meters of the observer use a high-density grid with a 1-degree interval between meridians and parallels; regions between 2 and 5 meters use a medium-density grid with a 2-degree interval between meridians and parallels; and regions beyond 5 meters use a low-density grid with a 5-degree interval between meridians and parallels. This multi-resolution grid structure reduces computational complexity while maintaining visual quality.
[0080] During ray projection sampling, a sampling ray is emitted from each pixel in the virtual observation space. The sampling ray originates from the viewpoint, passes through the spherical grid cell corresponding to the current pixel, and enters the multi-scale feature field. Adaptive step-size sampling is performed along the ray path, with an initial sampling step size set to 0.1 meters. When a large gradient change in feature values is detected, the step size is automatically reduced to 0.02 meters for finer sampling. At each sampling point, the feature value at that location is obtained from the multi-scale feature field using trilinear interpolation.
[0081] During the optical property conversion process, the geometric feature values of the sampling points are converted into optical absorption coefficients. The conversion formula for the optical absorption coefficient is: Optical Absorption Coefficient = Basic Absorption Coefficient + Geometric Feature Value × Absorption Scaling Factor, where the basic absorption coefficient is set to 0.1 and the absorption scaling factor is set to 0.8. Semantic feature values are converted into scattering intensity using the formula: Scattering Intensity = Semantic Feature Value × Scattering Intensity Coefficient, where the scattering intensity coefficient is set to 1.2. These parameters can be adjusted according to specific application scenarios.
[0082] In the light accumulation and synthesis stage, optical effects are accumulated sequentially along each light path from near to far. A backward-to-forward synthesis order is used, and the color contribution of the current sampling point is calculated using the formula: Color Contribution = Current Point Color Value × Current Point Opacity × Path Cumulative Transmittance. Path cumulative transmittance represents the proportion of light that is not absorbed during its propagation from the viewpoint to the current sampling point; its value decreases as the absorption coefficient accumulates at each point along the path. The final pixel color value is the sum of the color contributions of all sampling points along the path, and the pixel depth value is the distance between the first sampling point that reaches the preset opacity threshold and the viewpoint.
[0083] Based on continuous changes in the observation perspective, the above-mentioned ray projection and volume rendering processes are repeated in real time to generate a sequence of two-dimensional images with correct spatial relationships and depth information. When the viewpoint position changes, the spherical projection transformation parameters are recalculated, and the mapping relationship between the multi-scale feature field and the virtual observation space is updated to ensure that the scene visualization representation maintains spatial consistency under different observation angles.
[0084] In S4, the scene visualization representation is interactively extracted based on the feature field gradient. By responding to user annotations on the two-dimensional image sequence, it is back-mapped to a multi-scale feature field. Based on the intrinsic gradient information of the feature field, a feature subset of the target structure is separated, specifically including:
[0085] When establishing a bidirectional mapping relationship, the three-dimensional sampling rays corresponding to each pixel in the two-dimensional image sequence are first recorded. When a user performs annotation operations on the image, the corresponding sampling ray is found based on the pixel coordinates of the annotation point, and all voxel positions that the ray passes through in the feature field are recorded as the initial candidate set. The initial spatial position set is determined by back projection calculation. Specifically, from the ray paths corresponding to the annotation points, the 3 to 5 sampling points with the highest feature response values are selected as growth starting points. The three-dimensional coordinates of these starting points are calculated through the correspondence between image pixel coordinates and observation spherical coordinates, and then mapped to the specific voxel positions in the multi-scale feature field according to the inverse transformation from spherical coordinates to world coordinates.
[0086] When calculating the three-dimensional gradient vector, for each voxel in the multi-scale feature field, the rate of change of its eigenvalues in the X, Y, and Z directions is calculated separately. The gradient in the X direction is obtained by calculating the difference in eigenvalues between the current voxel and its right-hand neighbor; the gradient in the Y direction is obtained by calculating the difference in eigenvalues between the current voxel and its upper-hand neighbor; and the gradient in the Z direction is obtained by calculating the difference in eigenvalues between the current voxel and its upper-layer neighbor. The gradient values in the three directions are combined to form the three-dimensional gradient vector of that voxel. The gradient vectors of all voxels together form the gradient flow field describing the variation law of the feature field.
[0087] During bidirectional path tracing, growth proceeds simultaneously from each initial spatial location in both the positive and negative directions of the gradient vector. In the positive direction, the feature similarity between the current voxel and the next candidate voxel is calculated using the formula 1 minus the absolute value of the difference between the two voxel feature values. When the similarity is higher than 0.85, the candidate voxel is included in the growth region, and growth continues along the gradient direction. The negative direction uses the same logic but proceeds in the opposite direction of the gradient vector. Throughout the growth process, the positions of all included voxels and their connections are recorded in real time.
[0088] When handling path intersections, the spatial convergence of different growth paths is detected. When the spatial distance between the voxels of two paths is less than 2 voxel units, the consistency of the feature distribution of the two paths in the intersection region is calculated. The consistency calculation method is to compare the feature value distribution curves of the two paths in a 3×3×3 neighborhood around the intersection point. If the correlation coefficient of the two curves is greater than 0.9, they are considered to belong to the same feature structure, and path fusion processing is performed.
[0089] When constructing the connectivity graph structure, all voxels in the growth path are treated as nodes of the graph, and the connections between adjacent voxels are treated as edges. A depth-first search algorithm is used to identify connected components in the graph, removing isolated connected components with fewer than 10 nodes. For each connected component, it is checked whether there are any unincluded holes. If holes are found, the nearest neighbor interpolation method is used to fill the hole regions, forming a spatially continuous subset of candidate features.
[0090] When calculating the curvature distribution of the characteristic field, the curvature of the characteristic field at each voxel location is calculated based on the gradient flow field. The curvature calculation employs the central difference method, estimating the curvature value by the rate of change of the gradient between the current voxel and its six neighboring voxels. Specifically, the process involves first calculating the change of the gradient vector in each direction, then constructing a curvature tensor matrix, and finally obtaining the principal curvature value through eigenvalue analysis of the matrix. Adjacent voxels with curvature values differing by less than 0.1 are grouped into the same topological unit.
[0091] When optimizing boundary voxels, each topological unit's boundary voxel is analyzed. The characteristic gradient differences between the boundary voxel and its internal neighboring voxels are calculated, as well as the characteristic gradient differences with its external neighboring voxels. The boundary position is adjusted to the location where the sum of the internal and external gradient differences is greatest, i.e., the region with the most significant feature change. During the adjustment process, the overall connectivity of the topological unit is maintained to avoid creating new holes or breaks.
[0092] When verifying structural integrity, a 3D connectivity check is performed on each topological unit. This involves traversing all voxels within the unit to check for any unvisited isolated regions. If a break is found, interpolation repair is performed along the feature gradient direction at the break point, with the interpolation weights determined based on the feature similarity of adjacent voxels. The continuity and consistency of the original features are maintained during the repair process.
[0093] When finally separating the feature subsets, the validated topological units are merged according to their spatial adjacency. The merging criteria are that the feature similarity between adjacent units is greater than 0.8 and the boundary curvature changes continuously. The merged feature subsets need to meet the structural integrity threshold requirement, that is, the connectivity of all voxels within the subset reaches 100%, and the average gradient intensity of the boundary voxels is greater than 0.5. Feature subsets that meet these conditions are separated from the multi-scale feature field and used as the final target structural feature subsets.
[0094] In S5, based on the extracted feature subset, a dynamic parameter-driven interface is associated in the real-time graphics pipeline. By adjusting the optical property mapping function, enhanced 3D visualizations of the target structure at arbitrary viewpoints and scales are synthesized and presented in real time. Specifically, this includes:
[0095] When constructing a multi-transfer function group, corresponding optical parameters are set for different semantic feature value ranges within the feature subset. Three to five independent transfer functions are set, each responsible for processing feature values within a specific semantic range. The transfer functions are implemented using lookup tables, mapping the input feature values to color and opacity values. For each transfer function, its feature value range is defined; for example, feature values of 0.0 to 0.3 are mapped to the blue family, 0.3 to 0.6 to the green family, and 0.6 to 1.0 to the red family. Opacity mapping uses a linear relationship; higher feature values correspond to higher opacity. The base opacity coefficient is set to 0.1, and the maximum opacity coefficient is set to 0.9. During rendering, the corresponding transfer function is selected for optical property calculation based on the semantic feature value of each sampling point.
[0096] When dynamically calculating the sampling density, the number of samples per pixel is determined based on the spatial relationship between the observation viewpoint and the target structure. The base sampling rate is set to emit one sampling ray per pixel. In structural boundary regions, the boundary position is identified by analyzing the depth difference between adjacent pixels. When the depth difference between adjacent pixels is greater than twice the depth threshold, the region is determined to be a boundary region. The depth threshold is dynamically calculated based on the scene scale and is set to 0.5% of the scene's diagonal length. Adaptive supersampling is used in boundary regions, increasing the sampling rate to four sampling rays per pixel, with the sampling rays evenly distributed within the pixel region. In uniform regions where feature values change smoothly, the base sampling rate is maintained. This differentiated sampling strategy preserves boundary details while maintaining rendering efficiency.
[0097] When implementing depth-aware color enhancement, the color attributes of each pixel are dynamically adjusted based on its depth value. The depth value is calculated by sampling the distance between the intersection point of the ray and the target structure. The color saturation adjustment formula is: Adjusted saturation = Base saturation × (1 - Depth attenuation factor × Normalized depth value), where the base saturation is 1.0 and the depth attenuation factor is 0.6. The brightness adjustment formula is: Adjusted brightness = Base brightness × (1 - Depth attenuation factor × Normalized depth value), where the base brightness is 1.0. The normalized depth value is obtained by dividing the actual depth value by the maximum visible distance of the scene, and its value ranges from 0 to 1. Through this adjustment, the color saturation and brightness of near structures are higher, and the colors gradually become softer as the distance increases.
[0098] To achieve real-time refocusing, the blur radius of each pixel is calculated based on the user-specified focal distance. The focal distance is obtained through user interaction and ranges from 0.5 meters to the maximum visible distance of the scene. The blur radius calculation formula is: Blur radius = Base blur coefficient × |Current pixel depth - Focal distance| / Focal distance, where the base blur coefficient is set to 3.0. For pixels whose depth value differs from the focal distance within the depth of field, the blur radius is set to 0, maintaining complete sharpness. The depth of field range is set according to optical parameters, typically ranging from 5% to 10% of the focal distance. During rendering, each pixel is sampled from its neighborhood based on its blur radius, and a natural blur effect is achieved through Gaussian weighted averaging. Focal areas remain sharp, while non-focal areas exhibit a corresponding blur effect based on their degree of defocus.
[0099] Please see Figure 2 As shown, a three-dimensional visualization and extraction system for perceptual data includes:
[0100] The data synchronization and registration module is used to acquire multimodal sensing data streams from heterogeneous sensors, and to perform time synchronization and spatial registration on the multimodal sensing data streams to generate a spatiotemporally aligned multimodal data cube.
[0101] The feature fusion module is used to extract and hierarchically fuse multi-scale features of multimodal data cubes. It generates a unified multi-scale feature field by decomposing and recombining the features of multimodal data cubes at different scales.
[0102] The dynamic visualization module is used to dynamically project a unified multi-scale feature field onto a virtual observation space. Through ray casting and volume rendering techniques, the geometric and semantic attributes in the feature field are directly synthesized into a two-dimensional image sequence with visual depth, thereby constructing a scene visualization representation that supports real-time visual exploration.
[0103] The interactive extraction module is used to perform interactive extraction of scene visualization representation based on feature field gradient. By responding to user annotations on the two-dimensional image sequence, it back-maps to a multi-scale feature field and separates the feature subset of the target structure based on the intrinsic gradient information of the feature field.
[0104] The enhanced rendering module, based on the extracted feature subset, associates a dynamic parameter driving interface in the real-time graphics pipeline. By adjusting the optical property mapping function, it synthesizes and renders an enhanced 3D visualization of the target structure at any viewpoint and scale in real time.
[0105] The working principle of this invention is as follows: First, it generates a spatiotemporally aligned multimodal data cube by processing perceptual data from various heterogeneous sensors such as LiDAR, visual cameras, and millimeter-wave radar through time synchronization and spatial registration. Next, it constructs a multi-scale feature space using anisotropic 3D Gaussian convolution, enhances directional structural features through Hessian matrix feature analysis, and generates a unified multi-scale feature field through nonlinear fusion. Then, it establishes a virtual observation sphere for dynamic projection, and uses ray casting and volume rendering techniques to transform the feature field into a two-dimensional image sequence with depth information. Based on this, it establishes a bidirectional mapping relationship between the image and the feature field, and achieves precise separation of feature subsets of the target structure from the feature field based on gradient flow field-guided region growing and topological connectivity analysis. Finally, it constructs multiple optical transfer function sets, combines adaptive sampling, depth-sensing color enhancement, and real-time refocusing techniques, and generates enhanced 3D visualizations in the real-time graphics pipeline, supporting multi-view, multi-scale interactive exploration. This method realizes a complete technical chain from raw perceptual data to 3D visualization of the target structure, solving key technical problems such as multimodal data fusion, feature field construction, interactive extraction, and real-time enhanced rendering.
[0106] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A method for three-dimensional visualization extraction of perception data, characterized in that, The method comprises the following steps: S1: acquiring a multi-modal perception data stream from a heterogeneous sensor, and performing time synchronization and spatial registration on the multi-modal perception data stream to generate a spatio-temporally aligned multi-modal data cube; S2: performing multi-scale feature extraction and hierarchical fusion on the multi-modal data cube, generating a unified multi-scale feature field by decomposing and reorganizing the features of the multi-modal data cube in different scale spaces; S3: dynamically projecting the unified multi-scale feature field into a virtual observation space, directly combining the geometric and semantic attributes in the feature field into a two-dimensional image sequence with visual depth through ray casting and volume rendering techniques, thereby constructing a scene visualization representation supporting real-time visual exploration; S4: performing interactive extraction of the scene visualization representation based on the feature field gradient, mapping the target structure feature subset to the multi-scale feature field in reverse through the user's annotation on the two-dimensional image sequence, and separating out the target structure feature subset according to the internal gradient information of the feature field; S5: based on the extracted feature subset, associating a dynamic parameter driving interface in a real-time graphics pipeline, and adjusting the optical property mapping function to synthesize and present an enhanced three-dimensional visualization picture of the target structure under any viewing angle and scale in real time.
2. The method of claim 1, wherein, The generation of the spatio-temporally aligned multi-modal data cube specifically comprises: Compensating for cross-sensor clock drift and reordering data packets of the multi-modal perception data stream, using a dynamic time warping-based sequence alignment method to unify sensor data of different sampling frequencies to the same time reference, generating a time-synchronized continuous data frame sequence; Constructing a spatial transformation chain between sensors, optimizing the rotation and translation parameters in the spatial transformation chain by solving the correspondence relationship of geometric features in the observation data of adjacent sensors, and establishing an accurate mapping relationship from each sensor coordinate system to a unified world coordinate system; Projecting the time-synchronized continuous data frame sequence into the unified world coordinate system according to the accurate mapping relationship, and converting the discrete spatial sampling points into a regular grid representation of the spatio-temporally aligned multi-modal data cube through voxelization processing.
3. The method of claim 1, wherein the method further comprises: The generation of the unified multi-scale feature field specifically comprises: Constructing an anisotropic scale space, generating a multi-scale three-dimensional tensor sequence with direction specificity by setting different scale parameters along the three principal axes of space for three-dimensional convolution processing of the multi-modal data cube; Enhancing the direction of each scale feature in the multi-scale three-dimensional tensor sequence, extracting and strengthening tubular and sheet structure features by calculating the eigenvalue combination of the three-dimensional Hessian matrix in each voxel neighborhood, and generating a direction-enhanced feature tensor; Nonlinearly superimposing direction-enhanced feature tensors of different scales to fuse coarse-scale semantic information and fine-scale geometric details while maintaining the continuity of the spatial structure, and generating a unified multi-scale feature field.
4. The method of claim 1, wherein the method further comprises: The dynamic projection of the unified multi-scale feature field into a virtual observation space specifically comprises: Establishing a virtual observation sphere centered on the observer's position, and mapping the three-dimensional spatial coordinates in the multi-scale feature field to the two-dimensional spherical coordinates of the observation sphere through spherical projection transformation; A multi-resolution latitude-longitude grid is constructed on the surface of the observation sphere, and the grid density is adaptively adjusted according to the observation distance, that is, a high-density grid is used in the area close to the observer, and a low-density grid is used in the area far from the observer; The spherical projection transformation parameters are dynamically updated according to the change of the observation angle, the continuous and smooth mapping of the multi-scale feature field to the virtual observation space is realized, and the bidirectional correspondence between the feature field data and the observation space is established.
5. The method of claim 1, wherein the method further comprises: The scene visualization representation supporting real-time visual exploration is constructed, specifically including: A sampling light ray is emitted from each pixel point of the virtual observation space, and three-dimensional sampling with adaptive step is performed in the multi-scale feature field along the light ray path; The multi-scale features of each sampling point are converted into optical properties, that is, the geometric features are converted into optical absorption coefficients, and the semantic features are converted into scattering intensities; The optical effects are accumulated along each light ray path, and the final pixel color and depth value are calculated through depth compositing; A two-dimensional image sequence with correct occlusion relationship and depth information is generated in real time according to the continuous change of the observation angle, and a scene visualization representation supporting multi-view exploration is constructed.
6. The method of claim 1, wherein, The scene visualization representation is interactively extracted based on the feature field gradient, specifically including: A bidirectional mapping relationship between the feature field space and the image pixels is established, the labeling operation of the user on the two-dimensional image sequence is converted into the initial boundary constraint in the three-dimensional space, and the initial spatial position set in the multi-scale feature field is calculated through reverse projection; Based on the initial spatial position set, region growing based on the gradient flow field is performed in the multi-scale feature field, the change direction of the feature field gradient is tracked, the region is grown along the path with the maximum feature similarity, and a candidate feature subset is formed; The topological connectivity analysis and boundary optimization are performed on the candidate feature subset, the connectivity region with similar gradient features is identified by calculating the change rate vector of the feature field in the three-dimensional space, and the feature subset of the target structure is accurately separated from the multi-scale feature field according to the preset structural integrity threshold.
7. A method for three-dimensional visualization of perceptual data according to claim 6, characterized in that, The region growing based on the gradient flow field is performed, specifically including: A three-dimensional gradient vector of each voxel position in the multi-scale feature field is calculated, and a gradient flow field describing the feature change direction and intensity is constructed, wherein the gradient vector is obtained by calculating the spatial difference of the feature values between adjacent voxels; Starting from each position in the initial spatial position set, path tracking is simultaneously performed in the positive and negative directions of the gradient vector in the gradient flow field, and the feature similarity between adjacent voxels is calculated in real time during the tracking process, and the path is continued to grow only when the similarity is higher than the set threshold; The intersecting growth paths are fused, and adjacent paths with similar gradient change patterns are merged into the same growth region by analyzing the feature distribution consistency of the path intersection region; Based on the spatial distribution of all growth paths, a connectivity graph structure is constructed, and a spatially continuous and boundary complete candidate feature subset is formed by removing isolated branches and filling internal cavities.
8. The method of claim 6, wherein the method further comprises: The spatially continuous and boundary complete candidate feature subset is further included in Based on the spatial distribution of the candidate feature subset, a feature field curvature distribution of each voxel position is calculated, and by analyzing the consistency of the curvature change between adjacent voxels, a connected region with similar curvature features is divided into the same topological unit; The boundary voxels of the divided topological unit are optimized, the feature gradient difference between the boundary voxels and their adjacent internal and external voxels is calculated, the boundary position is adjusted to the position with the largest gradient change, and a boundary transition is formed; Based on the optimized boundary, the structural integrity of each topological unit is verified, the internal cavities and fractures of the unit are detected, and the feature field continuity is used for interpolation repair; The topological units that pass the verification are merged according to the spatial adjacency relationship, and the feature subset of the target structure is accurately separated from the multi-scale feature field according to the preset structural integrity threshold.
9. The method of claim 1, wherein, The real-time synthesis and presentation of the enhanced three-dimensional visualization picture of the target structure under any viewing angle and scale, specifically includes: A multiple optical transfer function group is constructed, an independent light transmission path is designed for different semantic features in the feature subset, a mapping relationship between feature values and optical parameters is established, and the light rendering of different structures is realized; According to the spatial relationship between the observation angle and the target structure, the sampling density of each pixel is dynamically calculated, the adaptive oversampling is used in the structure boundary area, the basic sampling rate is maintained in the uniform area, the balance between detail preservation and rendering efficiency is realized, and the depth perception color enhancement is implemented in the rendering process. The depth perception color enhancement is implemented in the rendering process, the color saturation and brightness are dynamically adjusted according to the pixel depth value, the color of the near structure is bright, the color of the far structure is soft, and the depth level of the scene is enhanced. Through the real-time refocusing technology, the depth of field effect is dynamically adjusted, the blur radius is automatically calculated according to the user interaction focus, the focus area is clearly presented, and the non-focus area is naturally blurred.
10. A three-dimensional visual analytics system for perception data, characterized by A three-dimensional visualization extraction method for sensing data is used to execute any one of claims 1-9, comprising: A data synchronization registration module is used to acquire multi-modal sensing data streams from heterogeneous sensors, and to perform time synchronization and spatial registration on the multi-modal sensing data streams to generate a spatio-temporal aligned multi-modal data cube; A feature fusion module is used to extract and hierarchically fuse multi-scale features from the multi-modal data cube, to generate a unified multi-scale feature field by decomposing and recombining the features of the multi-modal data cube in different scale spaces; A dynamic visualization module is used to dynamically project the unified multi-scale feature field to a virtual observation space, to directly combine the geometric and semantic attributes in the feature field into a two-dimensional image sequence with visual depth through ray casting and volume rendering techniques, and to construct a scene visualization representation supporting real-time visual exploration; An interactive extraction module is used to perform interactive extraction based on feature field gradients on the scene visualization representation, to respond to user annotations on the two-dimensional image sequence, to reversely map to the multi-scale feature field, and to separate the feature subset of the target structure according to the internal gradient information of the feature field. The enhanced presentation module is based on the extracted feature subset, associates a dynamic parameter driven interface in a real-time graphics pipeline, and synthesizes and presents an enhanced three-dimensional visualization picture of the target structure at any viewing angle and scale in real time by adjusting the optical property mapping function.