A method for efficient decoupling and object removal based on neural radiance field scene

By analyzing the projection cross-relationships and pixel structure of multi-view image datasets, an overlap interference factor is constructed, high-interference blocks are separated, and object attribution calibration is performed. This solves the artifact problem in object removal in complex indoor scenes and achieves efficient and accurate object culling.

CN121937615BActive Publication Date: 2026-07-21NINGBO MEIXIANG INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610385577.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-07-21
Estimated Expiration
2046-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively separate multiple independent objects in complex indoor scenes, leading to artifacts or incorrect reconstruction of the background structure when objects are removed. This is particularly true in furniture display scenes, where overlapping foreground and background voxels cause incorrect segmentation.

Method used

By analyzing the projection cross relationship, projection density, and pixel structure dispersion in a multi-view image dataset, an overlap interference factor is constructed to identify and separate high-interference blocks. Based on the color response path and depth drift path, object attribution is performed to remove objects to be removed.

Benefits of technology

It enables the automatic extraction of highly interfering regions without relying on object semantic labels or image masks, avoiding human segmentation errors, ensuring the accuracy and data-driven nature of object culling operations, and realizing the digital representation of multi-dimensional object behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937615B_ABST
    Figure CN121937615B_ABST
Patent Text Reader

Abstract

The application provides a method for efficiently decoupling and removing objects based on a neural radiation field scene, and relates to the technical field of data processing.The method comprises the following steps: acquiring a multi-view image dataset to identify overlapping views, analyzing the projection intersection relationship of different spatial voxels in different views to obtain an overlapping interference factor, extracting an image region with an overlapping interference factor higher than a preset interference threshold as a high-interference block, clustering and segmenting the spatial voxels in the high-interference block according to depth continuity and color gradient trends to obtain a candidate object set, tracking the color response path and depth drift path of the candidate object in each view image, matching the color derivative change rate and spatial coordinate difference between the two paths to obtain a reflection offset value, identifying a candidate object with a reflection offset value higher than the average value as a to-be-removed object, and constructing a voxel shielding range area of the to-be-removed object, and then removing the to-be-removed object from the multi-view image dataset.The application can accurately identify the voxel region of the to-be-removed object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for efficient decoupling and object removal based on neural radiation field scenes. Background Technology

[0002] Existing techniques typically employ a full-scene joint optimization approach to train neural radiation field models, enabling them to reconstruct the volumetric density and color distribution of a scene given a set of multi-view images. However, in complex indoor scenes containing multiple independent objects, the model may fail to semantically separate and model different objects. Due to limitations in object-level operations, existing techniques introduce explicit or implicit object masks to perform regional optimization of targets, thereby enabling scene editing operations.

[0003] These methods typically rely on accurate semantic segmentation results and perform local fine-tuning of the neural radiation field based on the segmented regions to achieve explicit control over the target. However, in furniture display scenarios, when removing sofas or chairs, the segmentation regions may have incorrect boundaries due to the overlap in depth and color between the foreground and background voxels, leading to obvious artifacts or incorrect reconstruction of the background structure after target removal. Summary of the Invention

[0004] The purpose of this invention is to provide an efficient decoupling and object removal method based on neural radiation field scenes, aiming to solve the problems mentioned in the background art.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] A method for efficient decoupling and object removal in a neural radiation field scene, the method comprising:

[0007] S100: Acquire multi-view image dataset and perform view overlap recognition based on it. Analyze the projection intersection relationship of different spatial voxels in different views. Then, combine the projection distribution density and pixel structure dispersion of each spatial voxel to obtain the overlap interference factor.

[0008] S200: Based on the overlap interference factor, the interference region is separated, and the image region with the overlap interference factor higher than the preset interference threshold is extracted as the high interference block. The coordinate range of the block in the three-dimensional voxel space is then marked to obtain the localization region data.

[0009] S300 performs spatial structure clustering based on the positioning area data, and clusters and segments the spatial voxels in the high interference block according to the depth continuity and color gradient trend to obtain a candidate object set.

[0010] S400: Fit the reflection trajectory based on the candidate object set, track the color response path and depth drift path of the candidate object in the image from each viewpoint, match the rate of change of the color derivative and the spatial coordinate difference between the two paths, and obtain the reflection offset value.

[0011] S500 performs object attribution calibration based on reflection offset values, identifies candidate objects with reflection offset values ​​higher than the average value as objects to be removed, constructs their voxel occlusion range regions, and then removes them from the multi-view image dataset.

[0012] Further, in S100, a multi-view image dataset is acquired and view overlap recognition is performed based on it. The projection intersection relationship of different spatial voxels in different views is analyzed. Then, combined with the projection distribution density and pixel structure dispersion of each spatial voxel, the overlap interference factor is obtained, including:

[0013] Based on the multi-view image dataset, voxel matching is performed to extract the voxel mapping relationship of the same spatial position in different image frames, and a voxel projection trajectory set is obtained.

[0014] Based on the voxel projection trajectory set, projection cross analysis is performed to calculate the number of projection coverages and positional offsets of each voxel in multiple viewpoints, and preliminary cross relationship data is obtained.

[0015] Based on the preliminary cross-relationship data, the projection distribution density is calculated, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data.

[0016] Pixel structure discreteness is evaluated based on voxel density data. The consistency of pixel distribution direction and shape diffusion range of each voxel in multiple frames of images are analyzed to obtain structural discrete data.

[0017] The overlap interference factor is obtained by jointly constructing data based on voxel density data and structural discrete data, and by weighting and fusing the density and discrete values ​​of each voxel.

[0018] Furthermore, based on the preliminary cross-relationship data, the projection distribution density is calculated, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data, including:

[0019] Based on the preliminary cross-relationship data, spatial voxel aggregation and labeling are performed. The overlapping pixel regions of each voxel in different viewpoints are labeled and encoded to obtain the aggregation labeling matrix.

[0020] Inter-frame coverage intensity is measured based on the aggregated marker matrix, and the frequency of each marker voxel in multiple frames is counted to obtain inter-frame aggregated frequency data.

[0021] Voxel region compression is performed based on inter-frame aggregation frequency data. The concentration of each voxel projection range in the spatial coordinate system is evaluated to obtain compression ratio data.

[0022] The voxel density data is obtained by merging the compression ratio data and inter-frame aggregation frequency data, and fusing the repetition rate and concentration of each voxel.

[0023] Furthermore, pixel structure discreteness is evaluated based on voxel density data. The consistency of pixel distribution direction and shape diffusion range of each voxel in multiple frames of images are analyzed to obtain structural discrete data, including:

[0024] Multi-frame pixel vector extraction is performed based on voxel density data to extract the pixel coordinates of each voxel in the multi-view image, thus obtaining a multi-frame pixel location set.

[0025] Pixel orientation offset is calculated based on the pixel position set of multiple frames. The angle difference between pixels of each voxel under different viewpoints is measured to obtain the orientation consistency matrix.

[0026] Based on the orientation consistency matrix, the spatial diffusion radius is measured, and the farthest offset distance of the projected position of each voxel in different viewpoints is counted to obtain the shape diffusion range data.

[0027] A joint evaluation is performed based on the directional consistency matrix and shape diffusion range data, and the directional offset consistency and shape diffusion range are integrated to obtain discrete structural data.

[0028] Further, in step S200, interference regions are separated based on the overlap interference factor. Image regions with overlap interference factors higher than a preset interference threshold are extracted as high-interference blocks, and their coordinate range in three-dimensional voxel space is determined to obtain localization region data, including:

[0029] Interference value threshold mapping is performed based on the overlap interference factor, and voxels in all image regions are mapped according to the overlap interference factor value to form an interference level layer.

[0030] Block extraction is performed based on the interference level layer. Voxels with overlapping interference factors higher than the preset interference threshold are extracted to obtain a preliminary set of high-interference image regions.

[0031] Based on the preliminary set of high-interference image regions, viewpoint consistency correction is performed, and the high-interference image regions in each image viewpoint are mapped to the three-dimensional voxel coordinate system to generate a set of spatial interference coordinate points.

[0032] Based on the spatial interference coordinate point set, boundary wrapping and coordinate encapsulation are performed to construct the minimum closed boundary volume of each high interference region, generating positioning region data.

[0033] Furthermore, S300 performs spatial structure clustering based on the positioning area data, segmenting the spatial voxels within high-interference blocks according to depth continuity and color gradient trends to obtain a candidate object set, including:

[0034] Based on the location area data, voxel depth layers are divided. All voxels in the high-interference block are sorted according to their depth values ​​in the Z-axis direction and divided into multiple depth partitions according to continuous intervals.

[0035] Color gradient calculation is performed based on depth partitioning, calculating the rate of change of color value of each voxel and its neighboring voxels in RGB space to obtain color gradient data;

[0036] Directional aggregation evaluation is performed based on color gradient data, and the consistency of color gradient directions of adjacent voxels is analyzed to obtain color direction fusion data.

[0037] Based on the data fused from depth partitioning and color orientation, voxel clustering is performed, and voxel points are grouped according to depth continuity and color orientation consistency to obtain a candidate object set.

[0038] Furthermore, S400 performs reflection trajectory fitting based on the candidate object set, tracks the color response path and depth drift path of the candidate object in images from various viewpoints, matches the rate of change of the color derivative and the spatial coordinate difference between the two paths, and obtains the reflection offset value, including:

[0039] Multi-view pixel trajectory extraction is performed based on the candidate object set. The position change sequence of each candidate object in each view image is extracted to obtain the color response trajectory set and the depth drift trajectory set.

[0040] Color derivative analysis is performed based on the color response trajectory set. The rate of change of the first derivative of pixel values ​​between consecutive frames in each color trajectory is calculated to obtain color derivative change data.

[0041] Spatial offset is calculated based on the depth drift trajectory set, and the spatial coordinate difference of each trajectory in the three-dimensional voxel coordinate system is measured to obtain spatial coordinate offset data.

[0042] The color derivative change data and spatial coordinate offset data are matched and fused together to obtain the reflection offset value by weighting and fusing the color derivative change rate of the corresponding trajectory with its spatial coordinate difference.

[0043] Furthermore, multi-view pixel trajectory extraction is performed based on the candidate object set, extracting the position change sequence of each candidate object in the images from each viewpoint, resulting in a color response trajectory set and a depth drift trajectory set, including:

[0044] Based on the candidate object set, perform intra-frame pixel region localization of the image, mark the boundary range of each object in each viewpoint image, and obtain viewpoint image position mapping data;

[0045] Based on the viewpoint image position mapping data, cross-frame pixel centroid calculation is performed, the pixel set of the object region in each frame is extracted and its pixel centroid coordinates are calculated to obtain the frame sequence coordinate set.

[0046] Color value tracking is performed based on the frame sequence coordinate set, and the color vector of the pixel centroid position in each frame is extracted to obtain the color response sequence.

[0047] Based on the color response sequence and the frame sequence coordinate set, construct the color response trajectory set and the depth drift trajectory set respectively.

[0048] Furthermore, based on the color response sequence and frame sequence coordinate set, a color response trajectory set and a depth drift trajectory set are constructed, including:

[0049] Based on the frame sequence coordinate set, the depth direction coordinates are arranged, and the pixel centroid coordinates of each frame are combined in time order to form a spatial drift path, thus obtaining the initial displacement trajectory data.

[0050] The displacement vector is solved based on the initial displacement trajectory data, the coordinate change vector between adjacent frames is calculated, the depth displacement path is constructed, and the depth drift trajectory set is obtained.

[0051] Based on the color response sequence, inter-frame color change sampling is performed to extract continuous change data of color values ​​over time, thus obtaining the color response sample sequence;

[0052] The color response sample sequence is assembled into a temporal structure, and the color change data is organized into response paths corresponding to the frame sequence positions to generate a color response trajectory set.

[0053] Furthermore, in S500, object attribution is performed based on the reflection offset value. Candidate objects with reflection offset values ​​higher than the mean are identified as objects to be removed, and their voxel occlusion range regions are constructed. These are then removed from the multi-view image dataset, including:

[0054] Based on the reflection offset values, statistical analysis of the offset values ​​is performed, and the mean and standard deviation of the reflection offset values ​​of all candidate objects are calculated to obtain the attribution determination threshold.

[0055] High-offset objects are filtered based on the attribution determination threshold. Candidate objects whose reflection offset values ​​exceed the mean are extracted and marked as objects to be removed, thus obtaining a list of objects to be removed.

[0056] Based on the list of objects to be removed, the occlusion range is extracted, and the boundary envelope of each object to be removed in the voxel coordinate system is extracted to construct voxel occlusion range data.

[0057] Data culling is performed based on voxel occlusion range data. Voxel pixel information that intersects with the occlusion range is removed from the multi-view image dataset to complete the object culling process.

[0058] The above-described solution of the present invention has at least the following beneficial effects:

[0059] This invention first analyzes the projection relationships of voxels in the same physical space from different image perspectives, establishes a projection cross-relationship model, statistically analyzes the overlap of each voxel's position in multiple perspectives, and combines its projection density on the image with the consistency of pixel structure direction to form an overlap interference factor. This factor reflects the degree of cross-occurrence of each spatial voxel in different image perspectives and the complexity of its manifestation. It can not only capture the stacking relationship of projection areas in the image, but also introduce local texture discreteness into it, so that the overlap interference factor has the ability to express spatial repeatability and image structure complexity at the same time, providing data support for subsequent high-interference block separation.

[0060] This invention constructs a spatial interference level layer by analyzing overlapping interference factors, and identifies regions in this layer where the interference factor is higher than a preset threshold. These high-interference regions are projected from the two-dimensional image into a three-dimensional spatial voxel coordinate system, generating a spatial interference point set that corresponds consistently across multiple frames. This allows for the automatic extraction of spatial blocks directly from the voxel overlap relationship of the image based on the interference intensity, without relying on object semantic labels or image masks. It has a completely data-driven characteristic, avoiding the interference of human segmentation errors and semantic model failures on subsequent processing, and providing a raw data foundation for structural recognition and object inference in high-interference regions.

[0061] This invention calculates the centroid position of the pixel region of a candidate object in a multi-view image, extracts its displacement trajectory between frames, and records the pixel color vector of the centroid position in each frame based on the displacement path, forming a color response trajectory and a depth drift trajectory. These reflect the surface color response behavior and spatial position transformation pattern of the candidate object in the image sequence, respectively, and are used for difference matching and dynamic determination. This forms a complete object dynamic modeling structure, making time-axis processing of static image data possible and realizing the digital expression of multi-dimensional object behavior.

[0062] This invention performs differential processing on the color response trajectory, calculates the first derivative of the color change between frames, quantifies the changing trend of the color response speed of the object surface under different viewpoints, and uses the position transformation of each frame in the spatial drift trajectory to calculate the spatial coordinate difference. It records the position jump amplitude of the voxel in the three-dimensional coordinate system, and after standardizing the two data dimensions respectively, it performs weighted fusion to obtain the reflection offset value, which reflects the joint manifestation of the color response anomaly and spatial variation amplitude of a single candidate object in the temporal structure. Candidate objects that exhibit drastic color response changes or unstable spatial displacement are screened out and removed as potential abnormal targets.

[0063] This invention performs voxel analysis on high-interference blocks, focusing on the voxel depth continuity along the Z-axis and the consistency of color gradient directions between voxels. It sorts the voxels in each high-interference region by depth, dividing them into multiple continuous depth levels. Then, it calculates the gradient change rate of adjacent voxels in the RGB color space in each level, evaluates the similarity of color gradient directions between each group of voxels, and finally classifies voxels with continuous depth and consistent color directions into the same category, thus more realistically simulating the spatial morphology and appearance change trends of objects in the three-dimensional world. Attached Figure Description

[0064] Figure 1 This is a flowchart of an efficient decoupling and object removal method based on a neural radiation field scene provided by an embodiment of the present invention. Detailed Implementation

[0065] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0066] like Figure 1 As shown, embodiments of the present invention propose an efficient decoupling and object removal method based on a neural radiation field scene, the method comprising:

[0067] S100: Acquire multi-view image dataset and perform view overlap recognition based on it. Analyze the projection intersection relationship of different spatial voxels in different views. Then, combine the projection distribution density and pixel structure dispersion of each spatial voxel to obtain the overlap interference factor.

[0068] S200: Based on the overlap interference factor, the interference region is separated, and the image region with the overlap interference factor higher than the preset interference threshold is extracted as the high interference block. The coordinate range of the block in the three-dimensional voxel space is then marked to obtain the localization region data.

[0069] S300 performs spatial structure clustering based on the positioning area data, and clusters and segments the spatial voxels in the high interference block according to the depth continuity and color gradient trend to obtain a candidate object set.

[0070] S400: Fit the reflection trajectory based on the candidate object set, track the color response path and depth drift path of the candidate object in the image from each viewpoint, match the rate of change of the color derivative and the spatial coordinate difference between the two paths, and obtain the reflection offset value.

[0071] S500 performs object attribution calibration based on reflection offset values, identifies candidate objects with reflection offset values ​​higher than the average value as objects to be removed, constructs their voxel occlusion range regions, and then removes them from the multi-view image dataset.

[0072] In this embodiment of the invention, S100, a multi-view image dataset is acquired and view overlap recognition is performed based on it. The projection cross relationship of different spatial voxels in different views is analyzed. Then, combined with the projection distribution density and pixel structure dispersion of each spatial voxel, an overlap interference factor is obtained. The multi-view projection behavior of voxels in three-dimensional space is quantified into an overlap interference factor, which characterizes the foreground and background overlapping areas caused by view changes between images, effectively avoiding the problem of inaccurate boundary judgment by the semantic model. S200, interference region separation is performed based on the overlap interference factor. Image regions with overlap interference factors higher than a preset interference threshold are extracted as high interference blocks, and their coordinate range in three-dimensional voxel space is marked to obtain the positioning region data. This realizes the accurate projection of interference regions from image space to voxel space, providing spatial constraints for subsequent clustering and modeling. S300, spatial structure clustering is performed based on the positioning region data. Spatial voxels in the high interference blocks are clustered and segmented according to depth continuity and color gradient trend to obtain a candidate object set. This realizes the structural division of high interference voxels in space, and can build clear boundaries between voxels with similar colors or overlapping voxels, preventing the phenomenon of multiple objects being confused or mis-clustered.

[0073] S400: Based on the candidate object set, reflection trajectory fitting is performed to track the color response path and depth drift path of the candidate object in the images from various viewpoints. The color derivative change rate and spatial coordinate difference between the two paths are matched to obtain the reflection offset value. This unifies and quantifies the color discontinuity and positional jump of the object, avoiding the problem of judging whether the target is an independent object based solely on visual similarity, and providing structured indicator support for subsequent object removal. S500: Based on the reflection offset value, object attribution is performed. Candidate objects with reflection offset values ​​higher than the average value are identified as objects to be removed, and their voxel occlusion range area is constructed. Then, they are removed from the multi-view image dataset. The object screening and removal are completed through quantitative indicators, ensuring that the object removal operation is based on behavioral trajectory and image change data, avoiding artifact and background misjudgment phenomena.

[0074] Obtaining a multi-view image dataset specifically includes:

[0075] Obtaining a multi-view image dataset is a crucial prerequisite for constructing a neural radiation field model, as there is a direct data foundation and processing dependency between them. As a neural network model based on viewpoint synthesis and volumetric rendering, the neural radiation field relies on image information from multiple different viewpoints to model the color and density of any point in 3D space. Therefore, only after obtaining an image dataset containing sufficient viewpoint differences and complete spatial coverage can accurate voxel reconstruction and volumetric modeling of the scene be performed, enabling the neural radiation field to represent the complete scene.

[0076] Addressing the challenge of object-level semantic decoupling in neural radiation fields, this paper proposes a method to analyze interference regions in a scene before the reconstruction stage by processing spatial intersections, pixel distribution density, and color structure discreteness in multi-view images. All these analyses are based on multi-view image data. Multi-view images not only provide the image frame sequences required for training the neural radiation field but also offer essential data for subsequent construction of spatial voxel mapping, identification of interference regions, and extraction of reflection trajectories. Voxel information in the image data is reconstructed into three-dimensional spatial points through pose estimation and projection calculation, thus participating in the volumetric modeling of the neural field and constituting the input to the entire radiation field network.

[0077] In a preferred embodiment of the present invention, S100, a multi-view image dataset is acquired and view overlap recognition is performed based on it. The projection intersection relationship of different spatial voxels in different views is analyzed, and then, combined with the projection distribution density and pixel structure dispersion of each spatial voxel, an overlap interference factor is obtained, including:

[0078] Based on the multi-view image dataset, voxel matching is performed to extract the voxel mapping relationship of the same spatial position in different image frames, and a voxel projection trajectory set is obtained.

[0079] Based on the voxel projection trajectory set, projection cross analysis is performed to calculate the number of projection coverages and positional offsets of each voxel in multiple viewpoints, and preliminary cross relationship data is obtained.

[0080] Based on the preliminary cross-relationship data, the projection distribution density is calculated, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data.

[0081] Pixel structure discreteness is evaluated based on voxel density data. The consistency of pixel distribution direction and shape diffusion range of each voxel in multiple frames of images are analyzed to obtain structural discrete data.

[0082] The overlap interference factor is obtained by jointly constructing data based on voxel density data and structural discrete data, and by weighting and fusing the density and discrete values ​​of each voxel.

[0083] In this embodiment of the invention, voxel matching is performed based on a multi-view image dataset to extract voxel mapping relationships at the same spatial location in different image frames, resulting in a voxel projection trajectory set. This establishes a geometric correspondence between three-dimensional spatial voxels and two-dimensional images, constructing voxel mapping trajectories in multiple image frames. This provides a unified analytical foundation for subsequent cross-relationship identification, density assessment, and structural separation. Projection cross-analysis is performed based on the voxel projection trajectory set to calculate the number of projection coverage times and positional offset amplitudes of each voxel in multiple views, obtaining preliminary cross-relationship data. The projection consistency and view coverage features exhibited by voxels in multiple views are extracted, providing a basis for determining whether they are in a spatially stable structural region. Projection distribution density is calculated based on the preliminary cross-relationship data, statistically analyzing the aggregation intensity of each voxel in the spatial projection region to obtain voxel density data. This method describes the structural stability of voxel regions within an image sequence, providing an input dimension for subsequent interference factor assessment. It evaluates pixel structural dispersion based on voxel density data, analyzing the consistency of pixel distribution direction and shape diffusion range of each voxel across multiple frames to obtain structural discrete data. This data characterizes the degree of change in the projection path direction and spatial divergence of a voxel in the image sequence, effectively identifying boundary voxels, structural transition voxels, or non-rigid object voxels, providing a basis for subsequent high-interference block localization. Furthermore, it jointly constructs a weighted fusion of density and discrete values ​​for each voxel based on voxel density and structural discrete data, obtaining an overlap interference factor. This achieves a fusion expression of density and dispersion, screening out regions with potential structural aliasing, viewpoint repetition, or occlusion overlap from large-scale voxel data, providing a foundation for subsequent region identification with quantitative judgment criteria.

[0084] Specifically, voxel matching is performed based on a multi-view image dataset to extract voxel mapping relationships for the same spatial location in different image frames, resulting in a voxel projection trajectory set, which includes:

[0085] First, camera pose parameters are analyzed on the input multi-view image dataset. Based on the extrinsic and intrinsic parameter matrices corresponding to each frame, a unified camera projection model is constructed. Then, a 3D voxel mesh structure is built for the target scene. Perspective projection is performed on the center point of each spatial voxel, mapping it to the 2D coordinate system of all view images to obtain its projected pixel position in each frame. By setting image boundary judgment rules and depth occlusion thresholds, it is determined whether the current voxel is visible in the view image. If visible, its corresponding pixel coordinates are recorded; otherwise, it is marked as invisible. After traversing all image frames, the pixel projection path and visibility identifier of each voxel in different view images are obtained. Each voxel and its mapping results in all view images are recorded one-to-one, forming a voxel projection trajectory set.

[0086] In a preferred embodiment of the present invention, the projection distribution density is calculated based on preliminary cross-relationship data, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data, including:

[0087] Based on the preliminary cross-relationship data, spatial voxel aggregation and labeling are performed. The overlapping pixel regions of each voxel in different viewpoints are labeled and encoded to obtain the aggregation labeling matrix.

[0088] Inter-frame coverage intensity is measured based on the aggregated marker matrix, and the frequency of each marker voxel in multiple frames is counted to obtain inter-frame aggregated frequency data.

[0089] Voxel region compression is performed based on inter-frame aggregation frequency data. The concentration of each voxel projection range in the spatial coordinate system is evaluated to obtain compression ratio data.

[0090] The voxel density data is obtained by merging the compression ratio data and inter-frame aggregation frequency data, and fusing the repetition rate and concentration of each voxel.

[0091] In this embodiment of the invention, spatial voxel aggregation and labeling are performed based on preliminary cross-relationship data. Overlapping pixel regions of each voxel in different viewpoints are labeled and encoded to obtain an aggregation label matrix. This matrix allows for querying the projection hit status of any spatial voxel across multiple frames, providing an efficient indexing mechanism for subsequent statistical calculations. Inter-frame coverage intensity is measured based on the aggregation label matrix, and the frequency of recurrence of each labeled voxel in multiple frames is statistically analyzed to obtain inter-frame aggregation frequency data, reflecting the degree of repetition of spatial voxels in multi-view images. Voxel region compression is performed based on the inter-frame aggregation frequency data to evaluate the concentration of each voxel's projection range in the spatial coordinate system, obtaining compression ratio data to quantitatively distinguish the spatial distribution consistency between different voxels. Finally, the compression ratio data and inter-frame aggregation frequency data are combined and calculated to fuse the repetition rate and concentration of each voxel, obtaining voxel density data. This data reflects the structural importance and aggregation characteristics of voxels in the entire image set, providing a quantitative basis for subsequent interference intensity determination and high-interference region extraction.

[0092] Specifically, voxel region compression is performed based on inter-aggregate frequency data to assess the concentration of each voxel's projection range in the spatial coordinate system, yielding compression ratio data, including:

[0093] First, for each frame of the multi-view image dataset, the projection coordinates of all spatial voxels in the image are extracted, and a mapping table between voxel indices and image projection positions is established. Then, based on the inter-frame projection cross-relationships of voxels recorded in the preliminary cross-relationship data, the set of all projection points for each voxel in multiple frames is statistically analyzed. For each voxel, its set of two-dimensional coordinate points projected in all image frames is uniformly projected back into a three-dimensional spatial coordinate system, and based on the distribution characteristics of these coordinate points in three-dimensional space, a minimum spatial bounding box containing all projection points of that voxel is constructed. This bounding box can be obtained by calculating the difference between the maximum and minimum coordinates of all projection points, thus forming the three-dimensional projection volume boundary of the voxel. Next, the number of projection points of the voxel in all image frames is recorded, and the ratio of the three-dimensional volume of the bounding box to the number of projection points is taken as the spatial compression ratio of the voxel. To unify the measurement standard, the compression ratios of all voxels need to be normalized so that their value range is controlled within [0,1], forming a compression ratio dataset.

[0094] Specifically, based on the compression ratio data and inter-frame aggregation frequency data, the repetition rate and concentration of each voxel are merged to obtain voxel density data, which includes:

[0095] Obtain the inter-frame aggregated frequency value for each voxel. Space compression ratio value ,in The indices are voxels, representing the degree of repetition of the voxel in the temporal dimension and the degree of projective aggregation in the spatial dimension, respectively. To construct a voxel density index that simultaneously considers both temporal and spatial aggregation characteristics, [the following is used:] ... and Normalization was performed to ensure that the two metrics were comparable at the same scale. Considering that high inter-frame frequency represents high cohesion, while high compression ratio represents low cohesion, voxel density values ​​were constructed... When doing so, the compression ratio needs to be inverted, i.e., calculated... To form a spatial concentration index Next, set the weighted fusion parameters. According to the formula Perform linear weighted fusion to obtain the voxel density value for each voxel. .

[0096] In a preferred embodiment of the present invention, pixel structure discreteness is evaluated based on voxel density data, and the consistency of pixel distribution direction and shape diffusion range of each voxel in multiple frames of images are analyzed to obtain structural discrete data, including:

[0097] Multi-frame pixel vector extraction is performed based on voxel density data to extract the pixel coordinates of each voxel in the multi-view image, thus obtaining a multi-frame pixel location set.

[0098] Pixel orientation offset is calculated based on the pixel position set of multiple frames. The angle difference between pixels of each voxel under different viewpoints is measured to obtain the orientation consistency matrix.

[0099] Based on the orientation consistency matrix, the spatial diffusion radius is measured, and the farthest offset distance of the projected position of each voxel in different viewpoints is counted to obtain the shape diffusion range data.

[0100] A joint evaluation is performed based on the directional consistency matrix and shape diffusion range data, and the directional offset consistency and shape diffusion range are integrated to obtain discrete structural data.

[0101] In this embodiment of the invention, multi-frame pixel vector extraction is performed based on voxel density data to extract the pixel coordinates of each voxel in multi-view images, resulting in a multi-frame pixel position set. The voxel space and image space are then correlated at the positional level to construct pixel coordinate trajectories, providing a data foundation for subsequent orientation analysis. Pixel orientation offset is calculated based on the multi-frame pixel position set, measuring the angle difference between pixels of each voxel under different viewpoints to obtain an orientation consistency matrix, effectively identifying whether its projection changes have stable directionality. Spatial diffusion radius is measured based on the orientation consistency matrix, statistically analyzing the farthest offset distance of each voxel's projection position in different viewpoints to obtain shape diffusion range data, determining the degree of projection dispersion of voxels under viewpoint changes, and providing a quantitative basis for subsequent analysis. Joint evaluation is performed based on the orientation consistency matrix and shape diffusion range data, fusing orientation offset consistency and shape diffusion range to obtain structural discrete data. By combining the two feature dimensions of orientation consistency and diffusion range, it is determined whether the voxel belongs to a locally consistent object or background discrete noise.

[0102] Specifically, spatial diffusion radius is measured based on the orientation consistency matrix, and the farthest offset distance of the projected position of each voxel in different viewpoints is calculated to obtain shape diffusion range data, which includes:

[0103] First, after extracting the pixel vectors of voxels from the multi-view images and establishing their orientation consistency matrix, the radius of spatial diffusion is measured based on the sequence of projection position coordinates of each voxel in each frame. Specifically, the set of projection pixel positions of a voxel in multiple frames is defined as follows: ,in Indicates the first The projected coordinates of the voxel in the frame image. Using each coordinate point in this set as a reference point, calculate its Euclidean distance from all other points. , satisfy The combination of these operations will be used to perform the following distance calculation operation: Iterate through all combinations of projection positions, selecting the maximum value as the maximum projection offset of the voxel, which is defined as its diffusion radius in the image coordinate domain. Record the maximum distance value corresponding to each voxel as its diffusion index. The data of this type for all voxels are then aggregated to form a set of shape diffusion range data for all voxels.

[0104] Specifically, a joint evaluation is performed based on the directional consistency matrix and shape diffusion range data, fusing directional offset consistency and shape diffusion range to obtain discrete structural data, including:

[0105] In completing the direction consistency matrix Construction and spatial diffusion radius After obtaining the data, the stability of the fused voxel in terms of orientation change trend and its distribution range in terms of viewpoint projection position change are analyzed. First, the mean value of the orientation consistency of the voxel is obtained by averaging the cosine values ​​of all inter-frame angles in the orientation consistency matrix. and with As a score for directional dispersion, it reflects the instability of directional changes. Next, the diffusion radius... Normalization is performed using the global maximum diffusion radius value. Standardize it to obtain a spatial dispersion index with uniform scale. Then, the two indicators are weighted and fused proportionally to construct the structural dispersion value of the voxel. The calculation formula is as follows: ,in and These are the fusion weights for the two indicators, and the specific weights can be set empirically or adaptively adjusted according to the data characteristics in different scenarios. Under this weighted scoring mechanism, voxels with weak directional consistency and strong spatial diffusion will receive higher structural dispersion scores, reflecting their inconsistencies and behavioral instabilities in the image sequence, making them suitable for identification as interfering voxels or boundary region voxels; conversely, they can be considered as target component voxels with good structural consistency. After completing the above calculations, the structural dispersion scores of all voxels will form structural dispersion data.

[0106] In a preferred embodiment of the present invention, S200, interference regions are separated based on the overlap interference factor, image regions with overlap interference factors higher than a preset interference threshold are extracted as high-interference blocks, and their coordinate range in three-dimensional voxel space is calibrated to obtain positioning region data, including:

[0107] Interference value threshold mapping is performed based on the overlap interference factor, and voxels in all image regions are mapped according to the overlap interference factor value to form an interference level layer.

[0108] Block extraction is performed based on the interference level layer. Voxels with overlapping interference factors higher than the preset interference threshold are extracted to obtain a preliminary set of high-interference image regions.

[0109] Based on the preliminary set of high-interference image regions, viewpoint consistency correction is performed, and the high-interference image regions in each image viewpoint are mapped to the three-dimensional voxel coordinate system to generate a set of spatial interference coordinate points.

[0110] Based on the spatial interference coordinate point set, boundary wrapping and coordinate encapsulation are performed to construct the minimum closed boundary volume of each high interference region, generating positioning region data.

[0111] In this embodiment of the invention, interference value threshold mapping is performed based on the overlap interference factor. Voxels in all image regions are mapped according to the overlap interference factor value to form an interference level layer. The original voxel-level overlap interference factor values ​​are transformed into a continuous interference layer in the image space, realizing an intuitive expression of the interference strength of each region in the image. Block extraction is performed based on the interference level layer to extract voxels with overlap interference factors higher than a preset interference threshold, resulting in a preliminary set of high-interference image regions. This achieves numerical separation of noise structures in the image and avoids missegmentation problems caused by semantic ambiguity or background aliasing. Viewpoint consistency correction is performed based on the preliminary set of high-interference image regions. The high-interference image regions in each image viewpoint are mapped to a three-dimensional voxel coordinate system to generate a spatial interference coordinate point set. This realizes spatial fusion processing from the image domain to the voxel domain, providing a unified spatial expression method for multi-view image data and eliminating occlusion and projection differences under different viewpoints. Boundary encapsulation and coordinate encapsulation are performed based on the spatial interference coordinate point set to construct the minimum closed boundary volume of each high-interference region, generating positioning region data. The high-interference point set is encapsulated into regular or irregular geometric bodies to achieve a spatially closed expression of the data region.

[0112] Specifically, based on the overlap interference factor, an interference value threshold mapping is performed, mapping voxels in all image regions according to the overlap interference factor value to form an interference level layer, which includes:

[0113] First, a projection mapping relationship between image space and voxel space is established. For each frame, using known camera intrinsic parameters (such as focal length and principal point position) and extrinsic parameters (such as rotation and translation matrices), the 3D coordinates in voxel space are projected back onto the 2D plane of the image, locating their pixel coordinates in the image. Then, a 2D matrix with the same resolution as the image is constructed in each frame, initialized to zero, and all voxels mapped to the image in the current frame are iterated through, with their overlap interference factor values ​​assigned to the corresponding pixel positions. Finally, an interference level layer is formed on each frame, with pixels as units and the overlap interference factor as the grayscale value.

[0114] Specifically, based on the spatial interference coordinate point set, boundary wrapping and coordinate encapsulation are performed to construct the minimum closed boundary volume for each high-interference region, generating positioning region data, which includes:

[0115] After completing the unified viewpoint mapping of high-interference regions, a spatial interference coordinate point set is obtained. This point set consists of voxel coordinates obtained by back-projecting all image regions with overlapping interference factors exceeding a preset threshold into a three-dimensional voxel space. This point set represents a set of three-dimensional discrete coordinate data, distributed at different spatial locations, corresponding to multiple high-interference structural regions.

[0116] First, the coordinates of all voxels in the point set are aggregated according to spatial connectivity. A 3D eight-neighborhood search strategy can be used, whereby the region connectivity of the point set is marked by the 26 neighboring voxels around a voxel, identifying multiple independent interference clusters with spatial continuity. Each interference cluster is considered an independent candidate interference region. A boundary volume is encapsulated for each region individually. The specific steps are as follows: read the coordinates of all voxels in the cluster, construct a 3D axis-aligned bounding box, traverse the coordinate values ​​of all points on the X, Y, and Z axes respectively, record the minimum and maximum values, and define the encapsulation range of the interference cluster in 3D space. This minimum closed bounding box can be represented as: , , This refers to the coordinate range of the encapsulated body. Finally, after all high-interference clusters have completed the encapsulation operation, the coordinate range of their encapsulated bodies is recorded as the positioning region data.

[0117] In a preferred embodiment of the present invention, in step S300, spatial structure clustering is performed based on the positioning area data. Spatial voxels within high-interference blocks are clustered and segmented according to depth continuity and color gradient trends to obtain a candidate object set, including:

[0118] Based on the location area data, voxel depth layers are divided. All voxels in the high-interference block are sorted according to their depth values ​​in the Z-axis direction and divided into multiple depth partitions according to continuous intervals.

[0119] Color gradient calculation is performed based on depth partitioning, calculating the rate of change of color value of each voxel and its neighboring voxels in RGB space to obtain color gradient data;

[0120] Directional aggregation evaluation is performed based on color gradient data, and the consistency of color gradient directions of adjacent voxels is analyzed to obtain color direction fusion data.

[0121] Based on the data fused from depth partitioning and color orientation, voxel clustering is performed, and voxel points are grouped according to depth continuity and color orientation consistency to obtain a candidate object set.

[0122] In this embodiment of the invention, voxel depth layers are divided based on the positioning area data. All voxels in the high-interference block are sorted according to their depth values ​​along the Z-axis and divided into multiple depth partitions according to continuous intervals. Voxels with continuous shapes in three-dimensional space but potentially inconsistent colors are aggregated into an analysis unit, providing a clear local spatial range for subsequent color gradient analysis. Color gradients are calculated based on the depth partitions, calculating the rate of change of color values ​​of each voxel and its neighboring voxels in RGB space to obtain color gradient data. This quantifies the degree of change of each voxel and its neighboring voxels in color space, laying the foundation for evaluating color direction consistency. Directional aggregation evaluation is performed based on the color gradient data, analyzing the color gradient direction consistency of adjacent voxels to obtain color direction fusion data. This achieves directional modeling of color change trends between voxels, avoiding classification based solely on color value differences. Voxel clustering is performed based on the depth partitions and color direction fusion data. Voxel points are grouped according to depth continuity and color direction consistency to obtain a candidate object set. This achieves multi-dimensional voxel clustering that integrates depth topology information and color response trends, ensuring the clarity of boundaries for subsequent behavior trajectory tracking of candidate objects.

[0123] Specifically, based on the positioning area data, voxel depth layers are divided. All voxels within high-interference blocks are sorted according to their depth values ​​along the Z-axis, and then divided into multiple depth partitions based on continuous intervals. This includes:

[0124] First, the depth coordinates of all voxels in the positioning area data are extracted along the Z-axis. The Z-axis is typically the projected depth direction relative to the camera's viewpoint. All voxel Z-values ​​are sorted in ascending order to form a continuous depth sequence. After sorting, difference analysis is performed on this depth sequence to identify intervals with significant depth jumps. Using a preset depth interval threshold as a criterion, voxels with continuous depth changes within the threshold range are grouped into the same depth level, forming multiple depth partitions. To prevent excessive voxel fragmentation, a minimum voxel count threshold can be set to constrain the depth layers. If the number of voxels in a layer is lower than the set value, it is merged with adjacent layers. After the depth partitioning operation is completed, the system labels the depth layer of each voxel with a number, which is used as the grouping basis for spatial clustering in subsequent color analysis.

[0125] Specifically, directional aggregation evaluation is performed based on color gradient data, analyzing the consistency of color gradient directions among adjacent voxels to obtain color direction fusion data, which includes:

[0126] After completing the depth layer partitioning, for each voxel set within the depth partition, the pixel value in the RGB color space is extracted for each voxel, and the color gradient vector between the voxel and its directly adjacent voxels is calculated. The color gradient vector is calculated in the following form: for voxel A and its adjacent voxel B, the color gradient vector... It consists of the differences between the RGB components, denoted as . After obtaining the color gradient vectors between all voxels, the cosine of the angle between adjacent color gradient vectors is further calculated to evaluate their directional consistency. Let the color gradient vectors of voxel A and its adjacent voxels B, C, D, etc., be respectively... Then calculate separately The cosine values ​​of the same direction are calculated, and their average value is used as the directional consistency index for voxel A. The closer this index value is to 1, the more consistent the direction of its color gradient change is within the neighborhood. The directional consistency values ​​of all voxels are mapped to a direction blending layer to reflect whether the direction of color change has a concentrated trend.

[0127] Specifically, voxel clustering is performed based on the fusion of depth partitioning and color orientation data. Voxel points are grouped according to depth continuity and color orientation consistency to obtain a candidate object set, which includes:

[0128] Within each depth partition, voxel points with directional consistency indices exceeding a set fusion threshold are selected as seed points, and density-based clustering algorithms are used for aggregation and expansion. During clustering, in addition to considering color direction consistency, a Euclidean spatial distance constraint is introduced between voxels; only voxels whose spatial distance does not exceed a set radius and whose directional consistency is greater than the threshold are allowed to be grouped into the same class. In voxel clustering, if multiple seed points exist in a region, their color direction consistency center vector is used as the guiding direction, and voxel points are weighted and filtered according to vector angle differences to construct cluster boundaries with consistent color directionality. This process is repeated until voxels within all depth partitions are classified, ultimately forming multiple non-overlapping voxel cluster sets. These sets are then encapsulated according to spatial boundaries to output candidate object sets.

[0129] In a preferred embodiment of the present invention, S400, reflection trajectory fitting is performed based on the candidate object set, tracing the color response path and depth drift path of the candidate object in images from various viewpoints, matching the rate of change of the color derivative and the spatial coordinate difference between the two paths, and obtaining the reflection offset value, including:

[0130] Multi-view pixel trajectory extraction is performed based on the candidate object set. The position change sequence of each candidate object in each view image is extracted to obtain the color response trajectory set and the depth drift trajectory set.

[0131] Color derivative analysis is performed based on the color response trajectory set. The rate of change of the first derivative of pixel values ​​between consecutive frames in each color trajectory is calculated to obtain color derivative change data.

[0132] Spatial offset is calculated based on the depth drift trajectory set, and the spatial coordinate difference of each trajectory in the three-dimensional voxel coordinate system is measured to obtain spatial coordinate offset data.

[0133] The color derivative change data and spatial coordinate offset data are matched and fused together to obtain the reflection offset value by weighting and fusing the color derivative change rate of the corresponding trajectory with its spatial coordinate difference.

[0134] In this embodiment of the invention, multi-view pixel trajectory extraction is performed based on the candidate object set. The position change sequence of each candidate object in the image from each viewpoint is extracted to obtain a color response trajectory set and a depth drift trajectory set. This ensures the physical consistency and positional accuracy of the object trajectory extraction and provides a unified data interface for subsequent color response and depth drift trajectory analysis. Color derivative analysis is performed based on the color response trajectory set to calculate the rate of change of the first derivative of pixel values ​​between consecutive frames in each color trajectory, obtaining color derivative change data and capturing the trend and frequency of color changes of the object under multiple views. Spatial offset calculation is performed based on the depth drift trajectory set to measure the spatial coordinate difference of each trajectory in the three-dimensional voxel coordinate system, obtaining spatial coordinate offset data, reflecting the relative displacement trend of the object under different views. Matching and fusion calculation is performed based on the color derivative change data and the spatial coordinate offset data. The color derivative change rate of the corresponding trajectory and its spatial coordinate difference are weighted and fused to obtain the reflection offset value. This maps the object behavior data from a time series to a numerical index, facilitating rapid sorting, classification, and filtering in a large-scale candidate object set.

[0135] Specifically, the color derivative change data and spatial coordinate offset data are matched and fused together. The color derivative change rate of the corresponding trajectory is weighted and fused with its spatial coordinate difference to obtain the reflection offset value, including:

[0136] For each color trajectory in the color response trajectory set, a first-order forward differencing method is used to calculate the color vectors between adjacent frames, obtaining the color derivative rate of change sequence corresponding to that color trajectory. The inter-frame derivative value can be represented as the root mean square difference of three channels: R, G, and B, forming a single-channel color rate of change sequence. Simultaneously, using the depth drift trajectory set, the spatial coordinate changes of the voxel center points of the same object in consecutive image frames are obtained, and the spatial coordinate offset sequence of the object is constructed by calculating the three-dimensional Euclidean distance between frames. To perform fusion operations at a uniform scale, the color derivative change sequence and the spatial coordinate offset sequence are normalized separately. Max-min normalization can be used to scale the two sequences to a uniform range, such as 0-1, to avoid imbalances in the fusion results due to different numerical scales. After normalization, two weighting coefficients are defined to represent the proportions of the color derivative and spatial offset in the fusion calculation. The choice of weighting coefficients can be set according to scene complexity or object characteristics; for example, a larger color derivative coefficient can be set in scenes with severe color interference, and a larger spatial offset coefficient can be set in backgrounds with multiple targets occluding. The reflection offset value between each frame is obtained by weighted summation of the color derivative and the space.

[0137] In a preferred embodiment of the present invention, multi-view pixel trajectory extraction is performed based on the candidate object set, extracting the position change sequence of each candidate object in the image from each viewpoint, to obtain a color response trajectory set and a depth drift trajectory set, including:

[0138] Based on the candidate object set, perform intra-frame pixel region localization of the image, mark the boundary range of each object in each viewpoint image, and obtain viewpoint image position mapping data;

[0139] Based on the viewpoint image position mapping data, cross-frame pixel centroid calculation is performed, the pixel set of the object region in each frame is extracted and its pixel centroid coordinates are calculated to obtain the frame sequence coordinate set.

[0140] Color value tracking is performed based on the frame sequence coordinate set, and the color vector of the pixel centroid position in each frame is extracted to obtain the color response sequence.

[0141] Based on the color response sequence and the frame sequence coordinate set, construct the color response trajectory set and the depth drift trajectory set respectively.

[0142] In this embodiment of the invention, intra-frame pixel region localization is performed based on the candidate object set, and the boundary range of each object in each viewpoint image is marked to obtain viewpoint image position mapping data. This ensures the accurate positioning of candidate objects in multi-view images and converts two-dimensional image information into three-dimensional spatial data, providing accurate initial data for subsequent pixel trajectory calculation and color response path tracking. Cross-frame pixel centroid calculation is performed based on the viewpoint image position mapping data. The pixel set of the object region in each frame is extracted, and its pixel centroid coordinates are calculated to obtain a frame sequence coordinate set. This provides accurate positional information for the subsequent construction of color response trajectories and depth drift trajectories, ensuring matching and tracking accuracy between multiple frames. Color value tracking is performed based on the frame sequence coordinate set. The color vector at the location of the pixel centroid in each frame is extracted to obtain a color response sequence, reflecting the color characteristics of candidate objects changing over time. Based on the color response sequence and frame sequence coordinate set, a color response trajectory set and a depth drift trajectory set are constructed respectively, comprehensively capturing the changing behavior of candidate objects from both color and spatial dimensions, providing data support for subsequent object removal and attribution calibration.

[0143] In a preferred embodiment of the present invention, a color response trajectory set and a depth drift trajectory set are constructed based on the color response sequence and the frame sequence coordinate set, respectively, including:

[0144] Based on the frame sequence coordinate set, the depth direction coordinates are arranged, and the pixel centroid coordinates of each frame are combined in time order to form a spatial drift path, thus obtaining the initial displacement trajectory data.

[0145] The displacement vector is solved based on the initial displacement trajectory data, the coordinate change vector between adjacent frames is calculated, the depth displacement path is constructed, and the depth drift trajectory set is obtained.

[0146] Based on the color response sequence, inter-frame color change sampling is performed to extract continuous change data of color values ​​over time, thus obtaining the color response sample sequence;

[0147] The color response sample sequence is assembled into a temporal structure, and the color change data is organized into response paths corresponding to the frame sequence positions to generate a color response trajectory set.

[0148] In this embodiment of the invention, a color response trajectory set and a depth drift trajectory set are constructed based on the color response sequence and the frame sequence coordinate set, respectively, including:

[0149] Based on the frame sequence coordinate set, depth direction coordinates are arranged, and the pixel centroid coordinates of each frame are combined in temporal order to form a spatial drift path, obtaining initial displacement trajectory data, which accurately captures the positional changes of the object in the image sequence. Displacement vectors are solved based on the initial displacement trajectory data, and coordinate change vectors between adjacent frames are calculated to construct depth displacement paths, resulting in a depth drift trajectory set, achieving accurate modeling of the object's motion trajectory in space. Inter-frame color change sampling is performed based on the color response sequence, extracting continuous color value changes over time to obtain a color response sample sequence, providing continuous data on the object's color changes in multi-view images. Temporal structure assembly is performed based on the color response sample sequence, organizing the color change data into response paths corresponding to the frame sequence positions, generating a color response trajectory set, effectively capturing the details of the object's color changes.

[0150] Specifically, the displacement vector is solved based on the initial displacement trajectory data, the coordinate change vector between adjacent frames is calculated, a depth displacement path is constructed, and a depth drift trajectory set is obtained, which includes:

[0151] First, the centroid coordinates of the candidate objects are extracted from each frame of the image, and the coordinate differences between adjacent frames are calculated. These coordinate differences reflect the object's motion in space. The centroid coordinates of each frame can be obtained using a weighted average method, where the weights are typically determined by the pixel's brightness or color intensity. Next, the displacement vector between each frame is obtained by calculating the difference in the centroid coordinates of each pair of adjacent image frames. This displacement vector represents the spatial position change of the candidate object from one frame to another in three-dimensional space. After the displacement vectors are calculated, they are arranged in chronological order to form a depth displacement path. Each displacement vector represents the spatial change of the candidate object between different points in time. Connecting these displacement vectors forms a complete depth displacement path, describing the continuous spatial motion of the object in multi-view images. Finally, the displacement vectors calculated between all frames are integrated into a depth drift trajectory set.

[0152] Specifically, based on the color response sample sequence, a temporal structure assembly is performed, organizing the color change data into response paths corresponding to the frame sequence positions to generate a color response trajectory set, which includes:

[0153] First, color information of candidate object regions is extracted from each frame of the image. This color information typically refers to the color value of each pixel in the image, with common color spaces including RGB and HSV. By extracting this color information, the color response data of the object in each frame can be obtained. Next, the color information of each frame is arranged in chronological order and matched with the frame's timestamp or frame number to form a complete color response sample sequence. This sequence reflects the color changes of the object under multiple viewpoints, demonstrating the continuous change of color over time. After extracting the color response sample sequence, temporal structure assembly is required. The process of temporal structure assembly involves associating the color response sample sequence with the temporal information of the image frames to ensure the temporal order and consistency of the color data. This step guarantees the temporal order of the color data, ensuring that the color change of each frame corresponds precisely to its corresponding time point. Finally, all the color response data is organized into a color response trajectory set.

[0154] In a preferred embodiment of the present invention, S500, object attribution is performed based on the reflection offset value, candidate objects with reflection offset values ​​higher than the mean are identified as objects to be removed, and their voxel occlusion range regions are constructed, and then they are removed from the multi-view image dataset, including:

[0155] Based on the reflection offset values, statistical analysis of the offset values ​​is performed, and the mean and standard deviation of the reflection offset values ​​of all candidate objects are calculated to obtain the attribution determination threshold.

[0156] High-offset objects are filtered based on the attribution determination threshold. Candidate objects whose reflection offset values ​​exceed the mean are extracted and marked as objects to be removed, thus obtaining a list of objects to be removed.

[0157] Based on the list of objects to be removed, the occlusion range is extracted, and the boundary envelope of each object to be removed in the voxel coordinate system is extracted to construct voxel occlusion range data.

[0158] Data culling is performed based on voxel occlusion range data. Voxel pixel information that intersects with the occlusion range is removed from the multi-view image dataset to complete the object culling process.

[0159] In this embodiment of the invention, a statistical analysis of the reflection offset values ​​is performed to calculate the mean and standard deviation of the reflection offset values ​​of all candidate objects, thereby obtaining an attribution threshold. This ensures the statistical analysis of the reflection offset values ​​of candidate objects, quantifies the behavioral differences between different candidate objects, and provides a quantitative basis for subsequent screening of high-offset objects. High-offset objects are then screened based on the attribution threshold, extracting candidate objects whose reflection offset values ​​exceed the mean and marking them as objects to be removed, thus obtaining a list of objects to be removed and achieving precise screening of high-offset objects. The occlusion range is extracted based on the list of objects to be removed, extracting the boundary envelope of each object to be removed in the voxel coordinate system to construct voxel occlusion range data. This ensures that the spatial range of the objects to be removed is accurately defined, providing a clear spatial data basis for subsequent removal operations and avoiding erroneous or incomplete removal. Finally, data removal is performed based on the voxel occlusion range data, removing voxel pixel information that intersects with the occlusion range from the multi-view image dataset, completing the object removal process. This ensures accurate removal of objects to be removed from the multi-view image data and avoids erroneous image reconstruction and artifact generation.

[0160] Specifically, the occlusion range is extracted based on the list of objects to be removed, and the boundary envelope of each object to be removed in the voxel coordinate system is extracted to construct voxel occlusion range data, including:

[0161] First, based on the list of objects to be removed, the position and voxel data of each object in 3D space are extracted. This requires mapping the pixel region of each object in the image to a voxel coordinate system in 3D space. Each object typically consists of multiple voxel points with a specific distribution in space. After obtaining the spatial information of the objects to be removed, the system calculates the spatial boundaries of the object voxels to generate the minimum bounding volume, i.e., the minimum closed boundary volume, for each object. Through this process, the system can determine the 3D spatial extent of the objects to be removed, ensuring accurate identification of the boundary region of each object. These boundary volumes are constructed by connecting the farthest points of the voxel coordinates, ensuring coverage of all voxel points to be removed. Voxel occlusion range data is then generated based on the constructed boundaries.

[0162] Specifically, data culling is performed based on voxel occlusion range data. This involves removing voxel pixel information that intersects with the occlusion range from the multi-view image dataset to complete the object culling process. This includes:

[0163] First, by comparing the extracted voxel occlusion range with the voxel coordinates of each frame in the image, the system determines which image regions intersect with the voxel occlusion range of the object to be removed. Each voxel pixel in the image contains position and color information. The system calculates the position of these pixels in 3D space and performs overlap detection with the occlusion area of ​​the object to be removed, accurately identifying pixel regions within the occlusion range. Next, for these voxel pixels that intersect with the occlusion range, the system removes them by replacing their color values ​​with the background color or transparent pixels, ensuring that the content of the object to be removed is no longer displayed in the image. Alternatively, the system can directly delete these voxel data to completely remove the relevant content from the image. After these culling operations, the system performs subsequent image restoration using image processing techniques such as image smoothing and filling to ensure that the gaps left after removal are properly filled, maintaining image quality and integrity. Finally, the object to be removed is successfully culled from the multi-view image, leaving a clear and structurally complete image.

[0164] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for efficient decoupling and object removal based on neural radiation field scenes, characterized in that, The method includes: Acquire multi-view image datasets and perform view overlap recognition based on them. Analyze the projection intersection relationship of different spatial voxels in different viewpoints. Then, combine the projection distribution density and pixel structure dispersion of each spatial voxel to obtain the overlap interference factor. Interference regions are separated based on the overlap interference factor. Image regions with overlap interference factors higher than the preset interference threshold are extracted as high interference blocks, and their coordinate range in the three-dimensional voxel space is marked to obtain the localization region data. Based on the location area data, spatial structure clustering is performed. Spatial voxels in high-interference blocks are clustered and segmented according to depth continuity and color gradient trend to obtain a candidate object set. Based on the candidate object set, the reflection trajectory is fitted, the color response path and depth drift path of the candidate object in the image of each viewpoint are tracked, and the color derivative change rate and spatial coordinate difference between the two paths are matched to obtain the reflection offset value. Object attribution is performed based on reflection offset values. Candidate objects with reflection offset values ​​higher than the average value are identified as objects to be removed, and their voxel occlusion range regions are constructed. These objects are then removed from the multi-view image dataset. The overlap interference factor is obtained by combining the projection distribution density of each spatial voxel and the pixel structure dispersion, including: Multi-frame pixel vector extraction is performed based on voxel density data to extract the pixel coordinates of each voxel in the multi-view image, thus obtaining a multi-frame pixel location set. Pixel orientation offset is calculated based on the pixel position set of multiple frames. The angle difference between pixels of each voxel under different viewpoints is measured to obtain the orientation consistency matrix. Based on the orientation consistency matrix, the spatial diffusion radius is measured, and the farthest offset distance of the projected position of each voxel in different viewpoints is counted to obtain the shape diffusion range data. A joint evaluation is performed based on the directional consistency matrix and shape diffusion range data, and the directional offset consistency and shape diffusion range are integrated to obtain discrete structural data. The overlap interference factor is obtained by jointly constructing data from voxel density data and discrete structural data. The tracking of the color response path and depth drift path of the candidate object in images from various viewpoints includes: Based on the candidate object set, perform intra-frame pixel region localization of the image, mark the boundary range of each object in each viewpoint image, and obtain viewpoint image position mapping data; Based on the viewpoint image position mapping data, cross-frame pixel centroid calculation is performed, the pixel set of the object region in each frame is extracted and its pixel centroid coordinates are calculated to obtain the frame sequence coordinate set. Color value tracking is performed based on the frame sequence coordinate set, and the color vector of the pixel centroid position in each frame is extracted to obtain the color response sequence. Based on the frame sequence coordinate set, the depth direction coordinates are arranged, and the pixel centroid coordinates of each frame are combined in time order to form a spatial drift path, thus obtaining the initial displacement trajectory data. The displacement vector is solved based on the initial displacement trajectory data, the coordinate change vector between adjacent frames is calculated, the depth displacement path is constructed, and the depth drift trajectory set is obtained. Based on the color response sequence, inter-frame color change sampling is performed to extract continuous change data of color values ​​over time, thus obtaining the color response sample sequence; The color response sample sequence is assembled into a temporal structure, and the color change data is organized into response paths corresponding to the frame sequence positions to generate a color response trajectory set.

2. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 1, characterized in that, Acquire a multi-view image dataset and perform view overlap recognition based on it. Analyze the projection intersection relationship of different spatial voxels in different views, and then combine the projection distribution density and pixel structure dispersion of each spatial voxel to obtain the overlap interference factor, including: Based on the multi-view image dataset, voxel matching is performed to extract the voxel mapping relationship of the same spatial position in different image frames, and a voxel projection trajectory set is obtained. Based on the voxel projection trajectory set, projection cross analysis is performed to calculate the number of projection coverages and positional offsets of each voxel in multiple viewpoints, and preliminary cross relationship data is obtained. Based on the preliminary cross-relationship data, the projection distribution density is calculated, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data. Pixel structure discreteness is evaluated based on voxel density data. The consistency of pixel distribution direction and shape diffusion range of each voxel in multiple frames of images are analyzed to obtain structural discrete data. The overlap interference factor is obtained by jointly constructing data based on voxel density data and structural discrete data, and by weighting and fusing the density and discrete values ​​of each voxel.

3. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 2, characterized in that, Based on the preliminary cross-relationship data, the projection distribution density is calculated, and the aggregation intensity of each voxel in the spatial projection region is statistically analyzed to obtain voxel density data, including: Based on the preliminary cross-relationship data, spatial voxel aggregation and labeling are performed. The overlapping pixel regions of each voxel in different viewpoints are labeled and encoded to obtain the aggregation labeling matrix. Inter-frame coverage intensity is measured based on the aggregated marker matrix, and the frequency of each marker voxel in multiple frames is counted to obtain inter-frame aggregated frequency data. Voxel region compression is performed based on inter-frame aggregation frequency data. The concentration of each voxel projection range in the spatial coordinate system is evaluated to obtain compression ratio data. The voxel density data is obtained by merging the compression ratio data and inter-frame aggregation frequency data, and fusing the repetition rate and concentration of each voxel.

4. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 3, characterized in that, Interference regions are separated based on the overlap interference factor. Image regions with overlap interference factors higher than a preset interference threshold are extracted as high-interference blocks, and their coordinate range in three-dimensional voxel space is determined to obtain the localization region data, including: Interference value threshold mapping is performed based on the overlap interference factor, and voxels in all image regions are mapped according to the overlap interference factor value to form an interference level layer. Block extraction is performed based on the interference level layer. Voxels with overlapping interference factors higher than the preset interference threshold are extracted to obtain a preliminary set of high-interference image regions. Based on the preliminary set of high-interference image regions, viewpoint consistency correction is performed, and the high-interference image regions in each image viewpoint are mapped to the three-dimensional voxel coordinate system to generate a set of spatial interference coordinate points. Based on the spatial interference coordinate point set, boundary wrapping and coordinate encapsulation are performed to construct the minimum closed boundary volume of each high interference region, generating positioning region data.

5. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 4, characterized in that, Spatial structure clustering is performed based on the location area data. Spatial voxels within high-interference blocks are clustered and segmented according to depth continuity and color gradient trends to obtain a candidate object set, including: Based on the location area data, voxel depth layers are divided. All voxels in the high-interference block are sorted according to their depth values ​​in the Z-axis direction and divided into multiple depth partitions according to continuous intervals. Color gradient calculation is performed based on depth partitioning, calculating the rate of change of color value of each voxel and its neighboring voxels in RGB space to obtain color gradient data; Directional aggregation evaluation is performed based on color gradient data, and the consistency of color gradient directions of adjacent voxels is analyzed to obtain color direction fusion data. Based on the data fused from depth partitioning and color orientation, voxel clustering is performed, and voxel points are grouped according to depth continuity and color orientation consistency to obtain a candidate object set.

6. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 5, characterized in that, Reflection trajectories are fitted based on the candidate object set, tracing the color response path and depth drift path of the candidate objects in images from various viewpoints. The rate of change of the color derivative and the spatial coordinate difference between the two paths are matched to obtain the reflection offset value, including: Multi-view pixel trajectory extraction is performed based on the candidate object set. The position change sequence of each candidate object in each view image is extracted to obtain the color response trajectory set and the depth drift trajectory set. Color derivative analysis is performed based on the color response trajectory set. The rate of change of the first derivative of pixel values ​​between consecutive frames in each color trajectory is calculated to obtain color derivative change data. Spatial offset is calculated based on the depth drift trajectory set, and the spatial coordinate difference of each trajectory in the three-dimensional voxel coordinate system is measured to obtain spatial coordinate offset data. The color derivative change data and spatial coordinate offset data are matched and fused together to obtain the reflection offset value by weighting and fusing the color derivative change rate of the corresponding trajectory with its spatial coordinate difference.

7. The method for efficient decoupling and object removal based on neural radiation field scene according to claim 6, characterized in that, Object attribution is performed based on reflection offset values. Candidate objects with reflection offset values ​​higher than the mean are identified as objects to be removed, and their voxel occlusion range regions are constructed. These objects are then removed from the multi-view image dataset, including: Based on the reflection offset values, statistical analysis of the offset values ​​is performed, and the mean and standard deviation of the reflection offset values ​​of all candidate objects are calculated to obtain the attribution determination threshold. High-offset objects are filtered based on the attribution determination threshold. Candidate objects whose reflection offset values ​​exceed the mean are extracted and marked as objects to be removed, thus obtaining a list of objects to be removed. Based on the list of objects to be removed, the occlusion range is extracted, and the boundary envelope of each object to be removed in the voxel coordinate system is extracted to construct voxel occlusion range data. Data culling is performed based on voxel occlusion range data. Voxel pixel information that intersects with the occlusion range is removed from the multi-view image dataset to complete the object culling process.

Citation Information

Patent Citations

  • Shielding removing method based on nerve radiation field

    CN116977360A

  • Label display method and system in three-dimensional scene and virtual engine

    CN118051159A