Anti-collision monitoring system based on depth estimation and instance segmentation fused three-dimensional model
Through the multi-branch network structure and cross-frame identity association mechanism, combined with multi-scale spatiotemporal occlusion mode analysis, the target segmentation and depth data are optimized, and the segmentation boundary overlap and identity jump problems of target recognition in complex environments are solved, achieving high-precision three-dimensional modeling and collision detection.
Patent Information
- Application Number
- CN202510846726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing instance segmentation algorithms are prone to problems such as overlapping segmentation boundaries and jumping target identity in situations where the target is severely blocked or the target appearance is highly similar, resulting in a decrease in target recognition capabilities, affecting the accuracy and real-time nature of three-dimensional modeling and collision prediction.
The multi-branch network structure is used to extract global semantics and local texture features, combined with cross-frame identity association and conflict detection mechanisms, and through multi-scale spatiotemporal occlusion mode analysis and boundary motion consistency verification, quadratic difference correction is performed, the spatial positioning information of the target is optimized, and it is integrated into the global coordinate system.
It improves the accuracy and comprehensiveness of target segmentation, maintains the consistency of target identity, significantly optimizes spatial positioning information, builds a high-precision and coherent three-dimensional semantic scenario, and improves the real-time and reliability of the anti-collision monitoring system.
Smart Images

Figure CN120356173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving security, and more specifically, to an anti-collision monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation. Background Art
[0002] In complex environments, dense target scenarios widely exist in fields such as autonomous driving, robotic swarm operations, and intelligent security. These scenarios are usually accompanied by high occlusion of targets, irregular distribution, and complexity of appearance features. In this environment, three-dimensional modeling technologies based on depth estimation and instance segmentation are widely used to achieve accurate identification and real-time modeling of static and dynamic objects in the scene. This technology provides important support for dynamic environment perception and collision risk prediction, but its effectiveness depends on accurate target segmentation, stable identity calibration, and high-quality depth estimation results.
[0003] Existing instance segmentation algorithms are prone to problems such as overlapping segmentation boundaries and jumping of target identities in cases of severe occlusion or highly similar target appearances. Since instance segmentation mainly relies on feature extraction and discrimination of single-frame images, when local areas are occluded or targets are densely distributed, the algorithm's ability to identify target boundaries significantly decreases, resulting in identity mismatches or segmentation errors between targets. These mismatched information is amplified in subsequent three-dimensional modeling and collision prediction, which may cause distortion of model reconstruction or deviation of motion trajectory prediction, ultimately affecting the anti-collision warning accuracy and real-time performance of the system. This problem is particularly prominent in multi-target dynamic scenarios, and it is urgent to improve the technical path to achieve more robust instance segmentation and target recognition. To solve the above problems, a technical solution is provided. Summary of the Invention
[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an anti-collision monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation, which realizes efficient feature extraction of global semantics and local textures through a multi-branch network structure, ensuring the accuracy and comprehensiveness of target segmentation; combines cross-frame identity association and conflict detection mechanisms to effectively maintain the consistency of target identities and reduce mismatch problems caused by occlusion or similar appearances; adopts multi-scale spatio-temporal occlusion pattern analysis and boundary motion consistency verification, uses a fusion algorithm to comprehensively evaluate occlusion complexity and segmentation robustness, identifies high-risk occlusion areas and performs second-order differential correction, significantly optimizing the spatial positioning information of targets; finally, integrates the optimized segmentation and depth data into the global coordinate system to construct a high-precision and coherent three-dimensional semantic scene, providing stable and high-quality input data for collision detection and trajectory prediction, so as to solve the problems proposed in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A collision prevention monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation, comprising: a feature extraction module, an identity matching module, an occlusion correction module, and a three-dimensional integration module;
[0007] Feature extraction module: Based on a multi-branch network structure, it extracts features of global semantics and local texture from the input image, generates a target segmentation mask and class information, and stores the results in a data structure , and transfers the data structure to the identity matching module;
[0008] Identity matching module: Using the segmentation mask and class information in the data structure , combined with the target tracking and multi-frame recognition strategies of adjacent frames, it performs cross-frame identity matching and conflict detection, and uses the re-identification mechanism to correct boundary jumps or identity mismatches caused by occlusion. The updated segmentation result and identity identifier are stored in the data structure , and transfers the data structure to the occlusion correction module;
[0009] Occlusion correction module: Fusing the corrected segmentation result and depth estimation data in the data structure , it uses multi-scale spatio-temporal occlusion pattern analysis and boundary motion consistency verification, and comprehensively evaluates the occlusion complexity and segmentation robustness through a non-linear fusion algorithm. It performs secondary differential correction on high-risk occlusion areas and outputs optimized spatial position information, which is stored in the data structure , and transfers the data structure to the three-dimensional integration module;
[0010] Three-dimensional integration module: Integrates the three-dimensional coordinates and segmentation annotations of all targets into the global coordinate system, constructs a complete three-dimensional semantic scene, and stores the integrated three-dimensional model in the data structure , providing stable and high-quality input data for collision detection and trajectory prediction.
[0011] In a preferred embodiment, the specific processing logic of the feature extraction module is as follows:
[0012] During specific processing, first, for the input image a global backbone network and a local branch network are established, and let represent the global semantic feature map, and let represent the local texture feature map; then, without changing the original resolutions of the global semantic feature map and the local texture feature map, through the position homomorphism transformation operator perform multi-scale local calibration in the feature coordinate system, and calculate the following fused feature map: ; where Represents the characteristic coordinates Perform isometric projection and differential translation in the local texture feature map according to the learnable pose offset matrix Represents the fused multi-dimensional features, i.e., the fused feature map; then use a multi-level classification decoder for the fused feature map to obtain the target segmentation mask And class information , where the multi-level classification decoder filters the global features in the semantic dimension, discriminates the edge differences in the texture dimension, and finally outputs the candidate segmentation regions; finally, the target segmentation mask, class information, and their coordinate mapping relationships are stored in the data structure In the form of In
[0013] In a preferred embodiment, the specific processing logic of the identity matching module is as follows:
[0014] For each target instance stored in the data structure , extract its spatio-temporal feature vector , where Represents the unique identifier of the target; feature extraction is implemented through a multi-modal feature encoder;
[0015] Use the dynamic time warping algorithm to compare the target feature vector in the previous frame with the target feature vector in the current frame, and calculate the feature sequence similarity score ;
[0016] For target pairs with a feature sequence similarity score lower than the matching threshold, use the re-identification module to re-determine using the high-dimensional fusion vector of the global semantic feature and the local texture feature; the re-identification process optimizes the discriminative ability of the target feature through a generative adversarial network; the re-identification recalculates the feature sequence similarity score according to the optimized high-dimensional fusion vector and re-matches the target identity according to the new score.
[0017] In a preferred embodiment, the specific processing logic of the identity matching module further includes the following:
[0018] For the case where there are still multiple pairs with a matching score higher than the matching threshold after passing through the re-identification mechanism, use a multi-object matching algorithm based on graph optimization to construct a target matching graph, and extract the unique identity correspondence through the maximum matching subgraph; the specific steps include:
[0019] Construct graph nodes: each node represents a target instance;
[0020] Construct graph edges: connect the nodes according to the feature sequence similarity score And ;
[0021] Apply the maximum weight matching algorithm to extract the optimal target identity correspondence relationship. The formula is as follows: ; where is the final matching set; update the target information after identity matching and conflict detection correction to the data structure , where: is the corrected segmentation mask; is the updated category information; is the new coordinate mapping relationship; is the target identity identifier.
[0022] In a preferred embodiment, the specific processing logic of the occlusion correction module is as follows:
[0023] S3.1, First, perform multi-scale spatio-temporal analysis on the segmentation mask and the corresponding depth data ; capture the occlusion dynamics of the target object at different time frames and spatial scales through spatio-temporal occlusion density evaluation; specifically, for each target , within the time interval , evaluate its occlusion degree at each frame . The calculation formula is as follows: ; where is the spatial occlusion density weight function, reflecting the influence of different spatial positions on the occlusion degree; obtain the occlusion pattern sequence of the target at each time frame and spatial scale through multi-scale analysis;
[0024] S3.2, Based on the occlusion pattern sequence, use the holographic cointegration transformation operator to transform the spatio-temporal occlusion dynamics into a high-dimensional vector representation; perform convolution and pooling operations on the sequence through a multi-scale time convolutional network to capture the occlusion consistency and complexity at different time scales, and generate the occlusion holographic cointegration vector ; the specific calculation process is: ; where and respectively represent multi-scale convolution and multi-scale pooling operations.
[0025] In a preferred embodiment, the specific processing logic of the occlusion correction module further includes the following content:
[0026] S3.3, To evaluate the robustness of the segmentation boundary in a dynamic environment, define the dynamic boundary continuous tensor , and perform multi-level spatial feature extraction on the curvature change and motion continuity of the segmentation boundary ; specifically, for each target , at each frame , calculate the boundary curvature change rate and the boundary motion consistency score , the formula is as follows: ; where represents the boundary curvature of the target in frame , is the boundary motion consistency score, is a small constant to prevent division by zero; integrate the curvature change rate and the motion consistency score through a multi-level fusion operator to generate a dynamic boundary continuous tensor : .
[0027] In a preferred embodiment, the specific processing logic of the occlusion correction module further includes the following:
[0028] S3.4, comprehensively calculate the occluded holographic cointegration vector and the dynamic boundary continuous tensor to generate a cross-domain harmonic kernel moment ; during the fusion process, an adaptive threshold mechanism and a non-linear mapping function are used to dynamically adjust the weight distribution in scenarios with different occlusion complexities and boundary robustness, and the formula is as follows: ; where and are learnable adjustment parameters, is an offset constant; based on the classification evaluation of the cross-domain harmonic kernel moment and the risk threshold, classify the target as a high-risk occlusion area: if the cross-domain harmonic kernel moment is greater than the risk threshold, the target in the occlusion area is evaluated as high-risk;
[0029] S3.5, for the target in the occlusion area evaluated as high-risk, perform quadratic difference correction to optimize its spatial position information; micro-refine the segmentation mask and its mapped position in the depth map, and the formula is as follows: ; where is the refinement coefficient, represents the gradient operation on the cross-domain harmonic kernel moment, which is used to guide the adjustment direction and amplitude of the segmentation mask; through quadratic difference correction, generate an optimized segmentation mask and the corresponding spatial position information , and store them in the data structure .
[0030] In a preferred embodiment, the specific processing logic of the 3D integration module is as follows:
[0031] S4.1, first, for each target in the data structure Spatial location information Global coordinate system conversion: set the global coordinate system and use the known sensor position parameters to convert the spatial position of each target into global coordinates : ;in, Indicates the target The three-dimensional position vector in the global coordinate system;
[0032] S4.2, using depth data and optimized segmentation mask , generating each target 3D point cloud ; The specific operations are as follows: ;in, is the image pixel coordinate, is the corresponding depth value; then, the 3D point clouds of all targets are aligned through global coordinate transformation and fused into the global point cloud middle: ;in, Represents a spatial transformation operation on a point cloud.
[0033] The specific processing logic of the 3D integration module also includes the following:
[0034] S4.3, the data structure The category information in the global point cloud is integrated with the semantic labels of the points; the specific operation is to , according to their target The category information is annotated to generate a 3D point cloud with semantic labels : ;in, represents the semantic category label; then, a spatial consistency optimization algorithm is applied Optimize the 3D point cloud with semantic labels, eliminate noise points, fill data holes, and ensure the consistency and coherence of semantic labels in 3D space: ; Generate 3D semantic point cloud through spatial consistency optimization algorithm ;
[0035] S4.4, the optimized 3D semantic point cloud and the corresponding semantic labels and spatial coordinate information are stored in a structured manner in the global data structure ,in Represents the set of semantic labels for all target instances.
[0036] The technical effects and advantages of the anti-collision monitoring system based on depth estimation and instance segmentation fusion three-dimensional model of the present invention are as follows:
[0037] The present invention realizes efficient feature extraction of global semantics and local textures through a multi-branch network structure, ensuring the accuracy and comprehensiveness of target segmentation; combines cross-frame identity association and conflict detection mechanisms to effectively maintain the consistency of target identities and reduce mismatch problems caused by occlusion or appearance similarity; adopts multi-scale spatio-temporal occlusion pattern analysis and boundary motion consistency verification, and uses a non-linear fusion algorithm to comprehensively evaluate occlusion complexity and segmentation robustness, identify high-risk occlusion areas and perform second-order difference correction, significantly optimizing the spatial positioning information of the target; finally, integrates the optimized segmentation and depth data into the global coordinate system to construct a high-precision and coherent three-dimensional semantic scene, providing stable and high-quality input data for collision detection and trajectory prediction. It improves the spatial perception ability and warning accuracy of the system in complex dynamic environments, enhances the real-time performance and reliability of the anti-collision monitoring system, is widely applicable to fields such as autonomous driving, robot navigation, and intelligent security, and significantly improves the adaptability and practicality of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 FIG. is a schematic structural diagram of an anti-collision monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1: Figure 1 An anti-collision monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation according to the present invention is provided, including: a feature extraction module, an identity matching module, an occlusion correction module, and a three-dimensional integration module;
[0041] Feature extraction module: Based on a multi-branch network structure, perform feature extraction of global semantics and local textures on the input image, generate a target segmentation mask and category information, and store the results in a data structure and transfer the data structure to the identity matching module;
[0042] Identity matching module: Utilize the segmentation mask and category information in the data structure , combine the target tracking and multi-frame recognition strategies of adjacent frames, perform cross-frame identity matching and conflict detection, and use the re-identification mechanism to correct boundary jumps or identity mismatches caused by occlusion. The updated segmentation result and identity identifier are stored in the data structure and transfer the data structure Transferred to the occlusion correction module;
[0043] Occlusion correction module: Fuse the corrected segmentation result and depth estimation data in the data structure Use multi-scale spatio-temporal occlusion pattern analysis and boundary motion consistency verification to comprehensively evaluate the occlusion complexity and segmentation robustness through a non-linear fusion algorithm. Perform quadratic difference correction on high-risk occlusion areas and output optimized spatial position information, which is stored in the data structure , and transfer the data structure To the 3D integration module;
[0044] 3D integration module: Integrate the 3D coordinates and segmentation annotations of all targets into the global coordinate system, construct a complete 3D semantic scene, and store the integrated 3D model in the data structure , providing stable and high-quality input data for collision detection and trajectory prediction.
[0045] In a dynamic environment of a dense scene, the distribution of target objects usually has high complexity and diversity, including problems such as occlusion, overlap, and blurred feature boundaries of targets. In this case, to provide a reliable segmentation basis for subsequent cross-frame recognition and correction, it is necessary to fully exploit the global semantic and local texture features of the image and fuse multi-scale information to ensure the accuracy and integrity of the initial target mask and class annotation. This step is the initial step of the segmentation and recognition task in the entire system, and its result will directly determine the stability and accuracy of the subsequent processing modules.
[0046] The specific processing logic of the feature extraction module is as follows:
[0047] During specific processing, first establish a global backbone network and a local branch network for the input image , let represent the global semantic feature map, and let represent the local texture feature map; then, without changing the original resolution of the global semantic feature map and the local texture feature map, perform multi-scale local calibration in the feature coordinate system through the position homomorphic transformation operator to calculate the following fused feature map: ; where represents the feature coordinate, perform isometric projection and differential translation in the local texture feature map according to the learnable pose offset matrix, represent the fused multi-dimensional feature, that is, the fused feature map; then use a multi-level classification decoder on the fused feature map to obtain the target segmentation mask and class information , where the multi-level classification decoder filters the global features in the semantic dimension, discriminates edge differences in the texture dimension, and finally outputs candidate segmentation regions; finally, the target segmentation mask, category information, and their coordinate mapping relationships are stored in the data structure in the form of for the cross-frame target tracking and conflict detection unit in the identity matching module to call; after completion, the preliminary object segmentation that may exist in a dense environment has obtained the identification mask and category information, providing an identification reference for subsequent occlusion correction and 3D modeling, where the preliminary target mask and category prediction respectively represent the preliminary target distribution in the scene in the spatial dimension and the semantic dimension.
[0048] The position homomorphic transformation operator is an advanced mathematical tool used to achieve spatial alignment of global semantic features and local texture features in a multi-branch network. Its core technical features include multi-scale geometric correction, pose adjustment, and isometric projection. This operator first receives the local texture feature map from the local branch network, geometrically corrects and offsets the position of the local features through a learnable rotation matrix and translation vector to ensure its alignment with the global semantic features in the same coordinate system. Then, the position homomorphic transformation operator applies convolutional kernels at multiple scales to refine the adjusted local features, enhancing the robustness and expressiveness of the features and reducing the feature misalignment problem caused by scale differences. Finally, the local feature map after geometric correction and pose adjustment is stitched and fused with the global semantic feature map in the feature coordinate system to generate a unified fused feature map, thus achieving the consistency and complementarity of global and local features in the spatial dimension.
[0049] Through the joint extraction and correction of global semantics and local texture features, a target segmentation mask and category information with complete preliminary segmentation information are generated, covering the preliminary characterization of the target distribution in the spatial dimension and the semantic dimension. These data are stored in a structured form and become an important input for subsequent cross-frame target tracking and occlusion correction. The results of this module ensure the accuracy and robustness of the preliminary segmentation of objects in a dense scene under complex background conditions.
[0050] In an environment with multiple dense and dynamically changing targets, the identity consistency of target objects and the accuracy of segmentation boundaries are the keys to ensuring the accuracy of subsequent 3D modeling and collision prediction. Existing instance segmentation algorithms are difficult to effectively maintain the continuous identity of targets during single-frame processing, especially prone to mismatching and boundary jumps in the case of target occlusion and high appearance similarity. Therefore, it is necessary to ensure the continuity of target identity and the stability of segmentation boundaries through cross-frame identity association and conflict detection, combined with a multi-frame recognition strategy. This process not only requires accurate target tracking but also uses a re-identification mechanism to dynamically correct identity mismatches caused by occlusion to ensure the accuracy of subsequent 3D modeling and the reliability of collision prediction.
[0051] The specific processing logic of the identity matching module is as follows:
[0052] Deeply analyze the preliminary target segmentation mask and class information generated by the feature extraction module, and extract the spatio-temporal features of each target instance. These features include the shape, color, texture of the target, and its position in space, which are comprehensively encoded by a multi-modal feature encoder. The encoded feature vectors can fully reflect the multi-dimensional attributes of the target, ensuring high distinctiveness and uniqueness of different targets in the feature space.
[0053] For each target instance stored in the data structure extract its spatio-temporal feature vector , where represents the unique identifier of the target. Feature extraction is achieved through a multi-modal feature encoder as follows: ; where is the segmentation mask of target , is its class information, is its coordinate mapping relationship. The spatio-temporal feature vector comprehensively considers multi-dimensional information such as the shape, color, texture, and spatial position of the target, ensuring high distinctiveness and uniqueness of the features.
[0054] Use the dynamic time warping algorithm to compare the target feature vector in the previous frame with the target feature vector in the current frame to calculate the similarity score of the feature sequence. The dynamic time warping algorithm can effectively handle the motion trajectories and feature changes of the target in different time frames. By setting a similarity threshold, pairs of targets with high similarity are screened out to complete cross-frame identity matching. This process ensures the identity consistency of the target in consecutive frames and reduces mis-matching caused by motion or occlusion.
[0055] For target pairs with a feature sequence similarity score lower than the matching threshold, use the re-identification module and use the high-dimensional fusion vector of the global semantic feature and the local texture feature for re-determination. The re-identification process optimizes the discriminative ability of the target features through a generative adversarial network as follows: ; The re-identification module recalculates the similarity score of the feature sequence based on the optimized high-dimensional fusion vector and re-matches the target identity according to the new score.
[0056] For the case where there are still multiple pairs of matching scores higher than the matching threshold after the re-identification mechanism, a multi-object matching algorithm based on graph optimization is adopted to construct a target matching graph, and the unique identity correspondence is extracted through the maximum matching subgraph. The specific steps are as follows:
[0057] Construct graph nodes: Each node represents a target instance.
[0058] Construct graph edges: According to the feature sequence similarity score and connect the nodes.
[0059] Among them, represents the similarity score optimized by the re-identification mechanism. This score is recalculated based on the high-dimensional fusion feature vector optimized by the generative adversarial network, aiming to improve the accuracy of target matching and correct the mis-matching problems caused by occlusion or high appearance similarity in the initial matching.
[0060] Apply the maximum weight matching algorithm to extract the optimal target identity correspondence. The formula is as follows: ; where is the final matching set.
[0061] Update the target information after identity matching and conflict detection correction to the data structure , where:
[0062] is the corrected segmentation mask.
[0063] is the updated category information.
[0064] is the new coordinate mapping relationship.
[0065] is the target identity identifier.
[0066] This data structure will be used as the input of the subsequent occlusion correction module to ensure the consistency of target identity and the accuracy of segmentation boundaries, providing a reliable data basis for high-precision 3D modeling and collision prediction.
[0067] Through cross-frame identity association and conflict detection, the consistency maintenance of target identity and the dynamic correction of segmentation boundaries are realized. The dynamic time warping algorithm is used for preliminary matching, and the re-identification module and graph optimization algorithm are used to solve the problems of low similarity and multiple matching conflicts, ensuring the accurate correspondence of target identity and the stability of segmentation boundaries. Finally, the updated data structure It contains target information that has been identity-calibrated and segmentation-corrected, providing a solid foundation for subsequent depth estimation and spatial position information optimization, and significantly improving the recognition accuracy and robustness of the anti-collision monitoring system in complex dynamic scenarios.
[0068] In the identity matching module of the present invention, only the target identity identifier is introduced in the identity matching module, and this identifier is not included in the data structures of other modules. Specifically, as the core link for cross-frame identity association and conflict detection, the identity matching module is responsible for ensuring the consistency of target identities and the accuracy of segmentation boundaries. Therefore, it is necessary to introduce the target identity identifier to track and correct the identities of each target in consecutive frames. However, in the subsequent occlusion correction module and 3D integration module, although the target identity identifier is not explicitly mentioned in the data structure, it is actually maintained indirectly through the segmentation mask , class information and spatial position information Specifically, the optimized segmentation mask and spatial position information have been adjusted and corrected based on the identity matching results of the identity matching module, enabling the implicit transmission and application of identity information during the 3D integration process. Therefore, this design ensures the continuity and consistency of the identity identifier throughout the overall system process without the need to repeatedly store the identifier in the data structure of each step, thereby optimizing the simplicity and processing efficiency of the data structure.
[0069] After completing the cross-frame identity association and conflict detection of the identity matching module, the data structure already contains the corrected target segmentation mask, class information, coordinate mapping relationship, and target identity identifier. However, in dense and dynamically changing scenarios, complex occlusions between target objects and minor changes in segmentation boundaries may still lead to inaccurate spatial position information. To further improve the accuracy of 3D model construction, it is necessary to combine depth estimation data to perform more detailed analysis and correction on the segmentation results. Through multi-scale spatio-temporal occlusion pattern analysis and boundary motion consistency verification, using a non-linear fusion algorithm, comprehensively evaluate the occlusion complexity and segmentation robustness, and perform second-order difference correction on high-risk occlusion regions to output more accurate spatial position information, ensuring the high reliability of 3D model construction and collision prediction.
[0070] The specific processing logic of the occlusion correction module is as follows:
[0071] S3.1, First, perform multi-scale spatio-temporal analysis on the segmentation mask output by the identity matching module and the corresponding depth data. By introducing spatio-temporal occlusion density evaluation, capture the occlusion dynamics of target objects at different time frames and spatial scales. Specifically, for each target , in the time interval within each frame degree of occlusion is evaluated, and the calculation formula is as follows: ; where is the spatial occlusion density weight function, reflecting the influence of different spatial positions on the degree of occlusion. Through multi-scale analysis, an occlusion pattern sequence of the target at each time frame and spatial scale is obtained , providing a basis for subsequent occlusion dynamic capture and quantification.
[0072] Among them, depth data refers to the depth information of each pixel point or spatial position in the scene obtained through depth estimation technology. Specifically, depth data can be obtained using a variety of sensor technologies, such as depth cameras, lidar, or stereo vision systems. These sensors can capture the position distance of objects in the scene relative to the camera or sensor, thereby generating a two-dimensional image or three-dimensional point cloud data containing depth values. In the specific implementation process, depth data is usually stored in the form of a two-dimensional depth map or three-dimensional point cloud, and each data point contains the depth value of the corresponding pixel or spatial position. The depth value corresponding to each pixel point in the depth map represents the distance from that point to the sensor. In the three-dimensional point cloud, each point represents its specific coordinate position in three-dimensional space. Through the accurate acquisition and processing of depth data, the system can achieve high-precision spatial perception and modeling of complex dynamic scenes, providing reliable spatial information support for the collision avoidance monitoring system.
[0073] S3.2, based on the occlusion pattern sequence, use the holographic cointegration transformation operator to transform the spatio-temporal occlusion dynamics into a high-dimensional vector representation. This transformation performs convolution and pooling operations on the sequence through a multi-scale time convolutional network to capture the occlusion consistency and complexity at different time scales, generating an occlusion holographic cointegration vector . The specific calculation process is as follows: ; where and represent multi-scale convolution and multi-scale pooling operations respectively, ensuring that the features of occlusion dynamics at different time scales are effectively captured and integrated.
[0074] S3.3, to evaluate the robustness of the segmentation boundary in a dynamic environment, define a dynamic boundary continuous tensor , and perform multi-level spatial feature extraction on the curvature change and motion continuity of the segmentation boundary . Specifically, for each target , within each frame , calculate its boundary curvature change rate and the boundary motion consistency score , and the formulas are as follows: ; where represents the target At the frame the boundary curvature, is the boundary motion consistency score, is a small constant to prevent division by zero.
[0075] Among them, the segmentation boundary is a set of pixel points in the target segmentation mask. These pixel points are located at the boundary between the target area and the background or other target areas, defining the spatial range of the target instance. The segmentation boundary is generated by extracting the pixel transition points between the value 1 (representing the target area) and 0 (representing the background or other areas) in the segmentation mask. These boundaries not only describe the contour of the target but also provide the local geometric characteristics of the target in the scene. In a dynamic scene, the change of the segmentation boundary between consecutive frames is a key factor in measuring the target motion consistency and boundary robustness. For example, through the curvature analysis of the segmentation boundary, abnormal changes caused by occlusion or motion can be identified, providing a basis for further segmentation optimization.
[0076] Integrate the curvature change rate and the motion consistency score through a multi-level fusion operator to generate a dynamic boundary continuous tensor : ; This tensor comprehensively reflects the continuity and stability of the target segmentation boundary in a multi-frame dynamic environment.
[0077] Among them, the multi-level fusion operator is a mathematical tool for comprehensive analysis of multi-dimensional features through hierarchical decomposition, cross-layer correlation, and dynamic weighted integration. The specific operations include dividing the input features (such as the boundary curvature change rate and the motion consistency score) into multiple levels (such as low-level local features and high-level global features) according to attributes. After extracting the feature representations for each level of features, methods such as inner product, convolution, or adaptive weighting are used to analyze the correlation between different levels, and a non-linear mapping function is combined to dynamically weight and fuse the low-level and high-level features, finally generating a unified representation with both local fine information and global trends, providing basic support for the robustness evaluation of the segmentation boundary or occlusion dynamic capture in complex scenes; The formulaic representation is , where and represent the low-level and high-level features respectively, is the cross-layer correlation analysis function, is the non-linear activation function, is the fusion result.
[0078] S3.4. Comprehensively calculate the occlusion holographic cointegration vector and the dynamic boundary continuous tensor to generate a cross-domain harmonic kernel moment . In this fusion process, an adaptive threshold mechanism and a non-linear mapping function are used to dynamically adjust the weight distribution in scenarios with different occlusion complexities and boundary robustness. The formula is as follows: ; Among them, and are learnable adjustment parameters, is an offset constant, ensuring the adaptability and stability of the fused cross-domain harmonic kernel moments in different scenarios. Based on the classification evaluation of the cross-domain harmonic kernel moments and the risk threshold, the target is classified as a high-risk or low-risk occlusion area: if the cross-domain harmonic kernel moment is greater than the risk threshold, the target in the occlusion area is evaluated as high-risk.
[0079] S3.5. For the target in the occlusion area evaluated as high-risk , perform second-order difference correction to optimize its spatial position information. Micro-refine the segmentation mask and its mapped position in the depth map, and the formula is as follows: ; Among them, is the refinement coefficient, represents the gradient operation on the cross-domain harmonic kernel moment, which is used to guide the adjustment direction and amplitude of the segmentation mask. Through second-order difference correction, an optimized segmentation mask and the corresponding spatial position information are generated and stored in the data structure , providing high-precision input data for the 3D modeling and collision prediction of the subsequent 3D integration module.
[0080] Through multi-scale spatio-temporal occlusion pattern quantization and boundary motion consistency verification, the occlusion correction module effectively combines the corrected instance segmentation results and depth estimation data, and uses holographic cointegration vectors and boundary continuous tensors for complexity and robustness evaluation. Generate cross-domain harmonic kernel moments through a non-linear fusion algorithm, dynamically adjust the weight distribution in the occlusion complexity and segmentation robustness scenarios, and achieve accurate discrimination and second-order difference correction of high-risk occlusion areas. Finally, the optimized spatial position information is systematically stored, providing a highly accurate and reliable data foundation for 3D model construction and collision prediction in the 3D integration module. Significantly improves the spatial perception ability and warning accuracy of the anti-collision monitoring system in complex dynamic scenarios, and ensures the efficiency and accuracy of 3D model construction and collision prediction.
[0081] After completing the second-order difference correction of the high-risk occlusion area in the occlusion correction module, the data structure already contains the optimized segmentation mask , class information and accurate spatial position information However, only these optimized segmentation and position information are not sufficient to construct a complete and consistent 3D semantic scene. To achieve high-precision 3D model construction and ensure the reliability of collision detection and trajectory prediction, it is necessary to integrate this optimized object information into the global coordinate system to form a unified and coherent 3D semantic scene. Through precise coordinate transformation, data alignment, and fusion processing, ensure the accurate positioning and semantic consistency of each object in 3D space, thereby providing stable and high-quality input data for subsequent collision detection and trajectory prediction.
[0082] The specific processing logic of the 3D integration module is as follows:
[0083] S4.1, First, perform a global coordinate system transformation on the spatial position information of each object in the data structure . Set the global coordinate system and use the known sensor position parameters (including the translation vector and the rotation matrix ) to convert the spatial position of each object into global coordinates : ; where represents the 3D position vector of object in the global coordinate system. This transformation ensures that all object instances are located within a unified spatial reference frame, laying the foundation for subsequent 3D modeling.
[0084] S4.2, Use the depth data and the optimized segmentation mask to generate the 3D point cloud of each object . The specific operation is as follows: ; where is the image pixel coordinate and is the corresponding depth value.
[0085] Subsequently, align the 3D point clouds of all objects through global coordinate transformation and fuse them into the global point cloud : ; where represents the spatial transformation operation of the point cloud. The fused global point cloud covers the accurate positions and shapes of all object instances in the global coordinate system, providing a detailed data basis for the construction of the 3D semantic scene.
[0086] S4.3, Integrate the category information in the data structure with the points in the global point cloud for semantic labeling. The specific operation is to label each point according to the category information of the object to which it belongs, generating a 3D point cloud with semantic labels: ; among which, represents the semantic category label.
[0087] Subsequently, apply the spatial consistency optimization algorithm to further optimize the 3D point cloud with semantic labels, eliminate noise points, fill data holes, and ensure the consistency and coherence of semantic labels in 3D space: ; through the spatial consistency optimization algorithm, generate a high-quality 3D semantic point cloud , providing accurate and stable spatial semantic information for collision detection and trajectory prediction.
[0088] The spatial consistency optimization algorithm is an advanced processing technology used to enhance the coherence and accuracy of 3D semantic point cloud data in the spatial dimension. The algorithm ensures the consistency and stability of the positions and semantic labels of each target instance in the global coordinate system through a multi-step process. First, the algorithm applies graph-based filtering techniques to identify and remove noise points in the point cloud, ensuring the cleanliness of the data. Subsequently, the spatial neighborhood clustering method is used to aggregate adjacent and semantically identical point clouds into continuous entities to eliminate segmentation breaks caused by sensor errors or dynamic environmental changes. In addition, the algorithm introduces geometric constraint conditions, and through curvature analysis and normal vector consistency checks, further corrects the subtle deviations of the target boundaries to ensure the smoothness and coherence of the boundaries. To handle moving targets in dynamic scenes, the algorithm combines time series information and applies motion compensation techniques to correct the spatial displacements caused by target movement and maintain the dynamic consistency of the 3D model. Finally, the spatial consistency optimization algorithm integrates semantic labels and spatial position information through multi-level feature fusion and optimization strategies to generate a high-quality, coherent, and semantically consistent 3D semantic point cloud.
[0089] S4.4, Structurally store the optimized 3D semantic point cloud, as well as the corresponding semantic labels and spatial coordinate information, into the global data structure , where represents the set of semantic labels of all target instances. This data structure ensures the ordered storage of all 3D spatial information and semantic labels, facilitating the efficient invocation and processing by subsequent collision detection and trajectory prediction modules.
[0090] Through precise coordinate system transformation, generation and fusion of multi-target three-dimensional point clouds, integration of semantic labels, and spatial consistency optimization, the 3D integration module successfully integrates the optimized object segmentation and depth estimation data into the global coordinate system, constructing a complete and highly accurate 3D semantic scene. The three-dimensional semantic point cloud stored in a structured manner and its set of semantic labels provide stable and detailed input data for subsequent collision detection and trajectory prediction. Through strict data processing and fusion logic, this module ensures the accuracy and consistency of 3D model construction, significantly improving the spatial perception ability and warning accuracy of the anti-collision monitoring system in complex dynamic scenarios, and ensuring the efficiency and reliability of the system in practical applications.
[0091] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0092] Only some exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.
[0093] It should be noted that in this article, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0094] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A collision prevention monitoring system based on a three-dimensional model integrating depth estimation and instance segmentation, characterized in that, Including: A feature extraction module, an identity matching module, an occlusion correction module, and a 3D integration module; Feature extraction module: Based on a multi-branch network structure, it extracts features of global semantics and local texture from the input image, generates a target segmentation mask and class information, and stores the results in a data structure , and transfers the data structure to the identity matching module; Identity matching module: Utilize the segmentation mask and class information in the data structure , combine the object tracking and multi-frame recognition strategies of adjacent frames, perform cross-frame identity matching and conflict detection, and use the re-identification mechanism to correct the boundary jumps or identity mismatches caused by occlusion. The updated segmentation results and identity identifiers are stored in the data structure , and transfer the data structure to the occlusion correction module; Occlusion correction module: Fusion data structure The corrected segmentation result and depth estimation data in are used to analyze the multi-scale spatio-temporal occlusion pattern and verify the boundary motion consistency. The occlusion complexity and segmentation robustness are comprehensively evaluated through a non-linear fusion algorithm. The quadratic difference correction is performed on the high-risk occlusion area, and the optimized spatial position information is output and stored in the data structure , and the data structure is passed to the 3D integration module; 3D Integration Module: Integrate the 3D coordinates and segmentation annotations of all targets into the global coordinate system to construct a complete 3D semantic scene. The integrated 3D model is stored in the data structure , providing stable and high-quality input data for collision detection and trajectory prediction.
2. The anti-collision monitoring system based on the three-dimensional model fused with depth estimation and instance segmentation according to claim 1, characterized in that, The specific processing logic of the feature extraction module is as follows: In specific processing, first, for the input image a global backbone network is established and a local branch network , such that represents the global semantic feature map, and let represent the local texture feature map; Then, without changing the original resolution of the global semantic feature map and the local texture feature map, perform multi-scale local calibration in the feature coordinate system through the position homomorphic transformation operator to calculate the following fused feature map: ; where represents the feature coordinates, perform isometric projection and differential translation in the local texture feature map according to the learnable pose offset matrix, represents the fused multi-dimensional feature, that is, the fused feature map; then use a multi-level classification decoder for the fused feature map to obtain the target segmentation mask and class information , where the multi-level classification decoder filters the global features in the semantic dimension and discriminates the edge differences in the texture dimension, and finally outputs the candidate segmentation regions; finally, store the target segmentation mask, class information, and their coordinate mapping relationships in the form of to the data structure .
3. The anti-collision monitoring system based on the three-dimensional model integrating depth estimation and instance segmentation according to claim 2, characterized in that, The specific processing logic of the identity matching module is as follows: For each target instance stored in the data structure extract its spatio-temporal feature vector , where represents the unique identifier of the target; feature extraction is implemented by a multi-modal feature encoder; Using the dynamic time warping algorithm, compare the target feature vector in the previous frame with the target feature vector in the current frame, and calculate the similarity score of the feature sequence ; For target pairs with a feature sequence similarity score lower than the matching threshold, a re-identification module is used to re-determine using the high-dimensional fusion vector of the global semantic feature and the local texture feature; The re-identification process optimizes the discriminative ability of the target feature through a generative adversarial network; the re-identification recalculates the feature sequence similarity score based on the optimized high-dimensional fusion vector and re-matches the target identity according to the new score.
4. The anti-collision monitoring system based on the three-dimensional model integrating depth estimation and instance segmentation according to claim 3, characterized in that, The specific processing logic of the identity matching module also includes the following: For cases where there are still multiple pairs with a matching score higher than the matching threshold after the re-identification mechanism, a multi-object matching algorithm based on graph optimization is adopted to construct a target matching graph, and the unique identity correspondence is extracted through the maximum matching subgraph; The specific steps include: Constructing graph nodes: Each node represents a target instance; Construct graph edges: Based on the feature sequence similarity score and connect nodes; Apply the maximum weight matching algorithm to extract the optimal target identity correspondence. The formula is as follows: ; where is the final matching set; update the target information after identity matching and conflict detection correction to the data structure , where is the corrected segmentation mask; is the updated category information; is the new coordinate mapping relationship; is the target identity identifier.
5. The anti-collision monitoring system based on the three-dimensional model fused with depth estimation and instance segmentation according to claim 4, characterized in that, The specific processing logic of the occlusion correction module is as follows: S3.
1. First, perform multi-scale spatio-temporal analysis on the segmentation mask and the corresponding depth data ; By evaluating the spatio-temporal occlusion density, capture the occlusion dynamics of the target object at different time frames and spatial scales; Specifically, for each target , within the time interval , evaluate its occlusion degree in each frame using the following formula: ; where is the spatial occlusion density weight function, reflecting the influence of different spatial positions on the occlusion degree; Through multi-scale analysis, obtain the occlusion pattern sequence of the target at each time frame and spatial scale; S3.2, based on the occlusion mode sequence, adopt the holographic cointegration transformation operator Dynamically transform the spatio-temporal occlusion into a high-dimensional vector representation; perform convolution and pooling operations on the sequence through a multi-scale temporal convolutional network to capture the occlusion consistency and complexity at different time scales, and generate an occlusion holographic cointegration vector ; The specific calculation process is as follows: Among them, and respectively represent multi-scale convolution and multi-scale pooling operations.
6. The anti-collision monitoring system based on the three-dimensional model fused by depth estimation and instance segmentation according to claim 5, characterized in that, The specific processing logic of the occlusion correction module also includes the following: S3.
3. To evaluate the robustness of the segmentation boundary in a dynamic environment, a dynamic boundary continuous tensor is defined , and multi-level spatial feature extraction is performed on the curvature change and motion continuity of the segmentation boundary . Specifically, for each target , at each frame , calculate its boundary curvature change rate and the boundary motion consistency score , and the formulas are as follows ; where represents the boundary curvature of the target at frame , is the boundary motion consistency score is a small constant to prevent division by zero; integrate the curvature change rate and the motion consistency score through a multi-level fusion operator to generate the dynamic boundary continuous tensor : .
7. The anti-collision monitoring system based on the three-dimensional model fused by depth estimation and instance segmentation according to claim 6, characterized in that, The specific processing logic of the occlusion correction module also includes the following: S3.
4. Synthesize and calculate the occluded holographic co-integration vector and the dynamic boundary continuous tensor to generate the cross-domain harmonic kernel moment. During the fusion process, an adaptive threshold mechanism and a non-linear mapping function are used to dynamically adjust the weight distribution in scenarios with different occlusion complexities and boundary robustness. The formula is as follows: where and are learnable adjustment parameters, is an offset constant; based on the classification evaluation of the cross-domain harmonic kernel moment and the risk threshold, the target is classified as a high-risk occlusion area: if the cross-domain harmonic kernel moment is greater than the risk threshold, the target in the occlusion area is evaluated as high-risk. S3.
5. For the target in the occluded area evaluated as high risk , perform second-order difference correction to optimize its spatial position information; micro-refine the segmentation mask and its mapped position in the depth map, and the formula is as follows: ; where is the refinement coefficient, represents the gradient operation on the cross-domain harmonic kernel moment, which is used to guide the adjustment direction and amplitude of the segmentation mask; through second-order difference correction, an optimized segmentation mask and the corresponding spatial position information are generated and stored in the data structure .
8. The anti-collision monitoring system based on the three-dimensional model fused with depth estimation and instance segmentation according to claim 7, characterized in that, The specific processing logic of the 3D integration module is as follows: S4.
1. First, perform a global coordinate system transformation on the spatial position information of each target in the data structure . Set up a global coordinate system, and use the known sensor position parameters to convert the spatial position of each target into global coordinates : ; Among them, represents the target in the three-dimensional position vector in the global coordinate system; S4.2, using the depth data and the optimized segmentation mask , generate the three-dimensional point cloud of each target ; The specific operations are as follows: ; Among them, is the image pixel coordinate, is the corresponding depth value; Subsequently, align the three-dimensional point clouds of all targets through global coordinate transformation and fuse them into the global point cloud : ; Among them, represents the spatial transformation operation of the point cloud .
9. The anti-collision monitoring system based on the three-dimensional model fused by depth estimation and instance segmentation according to claim 8, characterized in that, The specific processing logic of the 3D integration module also includes the following: S4.3, integrate the category information in the data structure with the points in the global point cloud for semantic labeling; specifically, for each point , label it according to the category information of the target it belongs to, and generate a three-dimensional point cloud with semantic labels : ; where represents the semantic category label; subsequently, apply the spatial consistency optimization algorithm to optimize the three-dimensional point cloud with semantic labels, eliminate noise points, fill data holes, and ensure the consistency and coherence of the semantic labels in three-dimensional space: ; through the spatial consistency optimization algorithm, generate a three-dimensional semantic point cloud ; S4.4, Structurally store the optimized three-dimensional semantic point cloud, as well as the corresponding semantic labels and spatial coordinate information, into the global data structure , where represents the set of semantic labels of all target instances.
Citation Information
Patent Citations
Irregular object pose estimation method and device based on depth camera
CN113450408A
Two-stage multi-modal three-dimensional instance segmentation method
CN114494276A
Three-dimensional instance segmentation method and system based on dense and sparse convolution fusion
CN116452940A
Real-time video analytics for traffic conflict detection and quantification
US20180253973A1
Cited By
Intelligent video monitoring management method and system based on big data
CN120976831A
Intelligent video monitoring management method and system based on big data
CN120976831B
Helmet collision detection accuracy improving method and system based on multi-factor cooperation
CN121502306A
Helmet collision detection accuracy improvement method and system based on multi-factor coordination
CN121502306B
Traffic accident scene detection method and device
CN121811342A