A multi-modal image semantic segmentation method and system for video analysis

CN122821518APending Publication Date: 2026-09-25北京星期七科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010420.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]为了解决现有技术的不足,本发明公开了一种面向视频分析的多模态图像语义分割方法及系统,旨在解决夜间低照度环境下,因RGB特征被噪声污染导致跨模态注意力权重分布平坦化,进而引发语义分割掩码边界模糊及类别置信度下降的技术问题

Benefits of technology

[0016]综合而言,本发明提供的一种面向视频分析的多模态图像语义分割方法及系统,其中,方法通过引入局部可用性识别与受限注意力匹配机制,有效解决了夜间低照度环境下RGB特征噪声污染导致的注意力权重平坦化问题。通过动态模态控制系数,系统能够根据局部场景状态自适应调整各模态贡献,在RGB失真区域自动增强激光雷达与热成像的约束作用。此外,通过多模态边界一致性仲裁与各向异性扩散滤波,系统能够有效消除物理边界冲突,实现目标边缘的锐利化修正。该方案显著提升了复杂光照条件下语义分割掩码的边界精度与类别置信度,确保了自动驾驶感知系统的鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821518A_ABST
    Figure CN122821518A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of automatic driving and computer vision, and discloses a multi-modal image semantic segmentation method and system for video analysis, which comprises the following steps: performing space-time alignment and feature extraction on visible light images, infrared thermal images and laser radar point clouds; for non-empty query points, generating a local availability label by combining visible light local state information and infrared and laser radar response information; generating a local modal control coefficient based on the label, and performing restricted attention matching on a query vector and a key vector to obtain a restricted attention weight distribution; generating an initial fusion feature map by using a cross-modal attention fusion algorithm; obtaining multi-modal boundary consistency information, correcting the boundary of the initial fusion feature map to obtain a corrected fusion feature map; and finally generating a semantic segmentation mask graph. Through dynamic modal control and boundary consistency correction, the precision and robustness of semantic segmentation in a low-illumination environment are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving and computer vision technology, and in particular to a multimodal image semantic segmentation method and system for video analysis. Background Technology

[0002] In the field of autonomous driving, existing multimodal collaborative semantic segmentation frameworks typically utilize cross-modal attention mechanisms to fuse RGB, thermal imaging, and LiDAR information. Under sufficient lighting conditions, this mechanism effectively combines RGB textures with LiDAR geometric contours. However, in low-light environments at night, the signal-to-noise ratio of RGB images drops sharply, causing target texture features to be overwhelmed by noise. In this situation, during cross-modal attention computation, the LiDAR query point is forced to search for a match in a noisy RGB feature map. Because the target region and background noise have minimal differences in the similarity metric space, the attention weight distribution tends to flatten, preventing the query point from accurately focusing. A large amount of weight is allocated to the background noise, leading to blurred and irregular boundaries of the target semantic mask in the fused feature map, and even the degradation of the target's internal semantic category to the background. Existing technologies struggle to effectively prevent the contamination of high-confidence geometric queries by noisy features when a particular modal feature is severely contaminated by noise and loses its metric significance.

[0003] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention discloses a multimodal image semantic segmentation method and system for video analysis. It aims to solve the technical problem that in low-light environments at night, the RGB features are contaminated by noise, leading to the flattening of cross-modal attention weight distribution, which in turn causes blurred semantic segmentation mask boundaries and decreased category confidence.

[0005] The technical solution of the present invention is as follows: In a first aspect, this invention discloses a multimodal image semantic segmentation method for video analysis, used to perform semantic segmentation on detected targets in low-light driving environments, including: Spatiotemporal alignment and feature extraction were performed on the synchronously acquired visible light images, infrared thermal images, and lidar point clouds of the target to obtain visible light feature maps that characterize visible light texture and semantics, infrared thermal feature maps that characterize thermal radiation contours and saliency, and lidar point cloud feature maps that characterize lidar geometric contours and depth information. For each non-empty query point in the lidar point cloud feature map, obtain the state information of the aligned visible light image in the corresponding local area, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the state information and response information. Based on local availability labels, local-level modal control coefficients are generated. Then, based on the modal control coefficients, constrained attention matching is performed on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map to obtain the constrained attention weight distribution. Using a cross-modal attention fusion algorithm, based on a constrained attention weight distribution, visible light feature maps, infrared thermal feature maps, and lidar point cloud feature maps are weighted and fused to generate an initial fused feature map; The multimodal boundary consistency information of the initial fused feature map near the target edge is obtained to correct the boundary of the initial fused feature map, resulting in a corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the semantic boundary of the fused feature, the physical depth boundary of the lidar, and the infrared thermal boundary. A semantic segmentation mask map is generated based on the modified fusion feature map.

[0006] This technical solution effectively blocks the contamination of geometric queries by noisy RGB features under low illumination through local availability identification and restricted attention matching. Combined with boundary consistency correction, it ensures the boundary sharpness and category accuracy of semantic segmentation masks in complex environments.

[0007] Furthermore, for each non-empty query point in the lidar point cloud feature map, the steps of obtaining the state information of the aligned visible light image in the corresponding local region, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same region, and generating a local availability label based on the state information and response information include: Traverse each non-empty query point in the lidar point cloud feature map and obtain the projection coordinates of the non-empty query point on the aligned visible light image; with the projection coordinates as the center and according to the preset pixel window size or the search radius determined based on the depth and curvature of the non-empty query point, delineate the local neighborhood corresponding to the query point as the visible light local region. Within a local visible light region, local signal-to-noise ratio and texture gradient information are obtained from the aligned visible light image as state information; At the same local neighborhood location, temperature gradient and thermal radiation contrast information are extracted from infrared thermal feature map as first response information, and depth variance and local curvature information are extracted from lidar point cloud feature map as second response information. The state information, first response information, and second response information are input into the state-aware network or a preset logical judgment rule to generate a local availability label for a local area. The local availability label is marked as a valid area label, a low-confidence area label, or a distorted area label.

[0008] Furthermore, the steps for generating local-level modal control coefficients based on local availability labels include: For each pixel location on the visible light feature map, read the local availability label corresponding to that pixel location; When the read local availability tag is a valid area tag, the visible light mode control coefficient of the pixel location is set as the first weight, and the infrared thermal mode control coefficient and lidar mode control coefficient of the pixel location are set as auxiliary weights; the first weight is greater than the auxiliary weight. When the read local availability tag is a low-confidence area tag, the visible light mode control coefficient of that pixel location is reduced to the second weight, and the infrared thermal mode control coefficient and lidar mode control coefficient of that pixel location are increased to the equivalent weight; the second weight is less than the first weight, and the equivalent weight is greater than the auxiliary weight; When the read local availability label is a distorted area label, the visible light mode control coefficient of that pixel location is attenuated to a third weight value approaching zero, and the infrared thermal mode control coefficient and lidar mode control coefficient of that pixel location are increased to the dominant weight value; the third weight value is less than the second weight value, and the dominant weight value is greater than the equivalent weight value; Output the modal control coefficients for this pixel location.

[0009] Furthermore, based on the modal control coefficients, the steps of performing constrained attention matching between the query vector corresponding to the lidar point cloud feature map and the key vectors corresponding to the visible light feature map and the infrared thermal feature map to obtain the constrained attention weight distribution include: Obtain the depth value and local curvature information of the query vector; Based on the depth value and local curvature information of the query vector, the two-dimensional search radius is calculated, and the initial candidate matching range of the query vector is determined accordingly. Within the initial candidate matching range, local availability labels are read pixel by pixel. Pixels marked as distorted regions are included in the noise rejection range, while pixel positions marked as valid regions and low-confidence regions are retained as the corrected candidate matching range. Based on the corrected candidate matching range and modality control coefficients, constrained attention matching is performed on the query vector and key vector to obtain a constrained attention weight distribution.

[0010] Furthermore, based on the corrected candidate matching range and modality control coefficients, the steps for performing constrained attention matching on the query vector and key vector to obtain the constrained attention weight distribution include: Calculate the initial similarity between the query vector and each key vector within the corrected candidate matching range; Based on the Euclidean distance between the position of each key vector within the corrected candidate matching range and the projection center of the query vector, a spatial distance penalty term is applied to the initial similarity. Based on the modal control coefficients at the locations of each key vector, the similarity after spatial distance penalty is weighted and modulated. Among them, the similarity of key vectors located within the noise exclusion range is forcibly assigned a preset minimum negative value. The modulated similarity is normalized to generate a constrained attention weight distribution.

[0011] Furthermore, the steps for obtaining multimodal boundary consistency information of the initial fused feature map near the target edge include: Obtain the depth boundary map corresponding to the aligned lidar point cloud, the thermal radiation boundary map corresponding to the aligned infrared thermal image, and the semantic boundary response map of the initial fused feature map; Morphological dilation operations are performed on the depth boundary map and the thermal radiation boundary map respectively to obtain the dilated depth boundary map and the dilated thermal radiation boundary map. Calculate the spatial intersection of the expansion depth boundary map and the expansion thermal radiation boundary map to obtain the boundary intersection region; Based on the spatial distance between depth boundary pixels and thermal radiation boundary pixels within the boundary intersection region, identify physical boundary conflict zones within the boundary intersection region; For the physical boundary conflict zone, the local connectivity length and depth gradient magnitude of the depth boundary map are evaluated to obtain the depth state; at the same time, the temperature gradient direction symmetry and attenuation width of the thermal radiation boundary map are evaluated to obtain the thermal state. Based on the depth and thermal states, within the physical boundary conflict zone, one of the following operations may be selectively performed: preserve the blocking properties of the depth boundary map, perform morphological shrinkage on the thermal radiation boundary map or generate a smooth blocking line based on the peak line of the semantic boundary response map, the shrunken thermal radiation boundary and the discontinuous depth boundary. Spatially stitching the boundary outside the physical boundary conflict zone with the boundary generated inside the physical boundary conflict zone yields a composite conduction blocking mask, which is the multimodal boundary consistency information.

[0012] Furthermore, based on multimodal boundary consistency information, the boundaries of the initial fused feature map are corrected to obtain the corrected fused feature map. The steps include: Obtain motion information of the target region corresponding to the composite conduction blocking mask; Obtain local deformation information of the initial fused feature map; Based on motion information and local deformation information, the composite conduction blocking mask is adjusted to obtain the adjusted composite conduction blocking mask. Anisotropic diffusion filtering is used, with the adjusted composite conduction blocking mask as the conduction coefficient control term, to perform iterative diffusion processing on the initial fused feature map to obtain the corrected fused feature map.

[0013] Furthermore, the steps for obtaining motion information of the target region corresponding to the composite conduction blocking mask include: Identify multiple moving targets within a composite conduction blocking mask area; For each moving target individual, calculate the motion vector of that moving target individual; Determine whether there is a spatial occlusion relationship between moving target individuals; When there is a spatial occlusion relationship between moving target individuals, the background motion information or the motion information of the occluded target within the occluded area is suppressed, and the motion vector of the occluded target individual is inferred. By integrating the motion vectors of individual moving targets, motion information of the target region corresponding to the composite conduction blocking mask is generated.

[0014] Furthermore, the steps for generating a semantic segmentation mask map based on the corrected fusion feature map include: Obtain the feature gradient distribution and local connectivity of the local edge regions of the corrected fused feature map; The fine contour information of the corresponding local edge region is obtained from the depth boundary map of the aligned lidar point cloud and the thermal radiation boundary map of the aligned infrared thermal image. Based on feature gradient distribution, local connectivity, depth boundary map and refined contour information, determine whether there is residual blur or discontinuity in the local edge region of the corrected fused feature map. If residual blur or discontinuity exists, guided filtering or morphological reconstruction is performed on the local edge regions of the corrected fused feature map based on the depth boundary map and refined contour information. The processed and fused feature map is input into the semantic segmentation head to generate a semantic segmentation mask map.

[0015] Secondly, the present invention also discloses a multimodal image semantic segmentation system for video analysis, comprising: The feature extraction module is used to perform spatiotemporal alignment and feature extraction on the synchronously acquired visible light image, infrared thermal image and lidar point cloud of the target, to obtain visible light feature map representing visible light texture and semantics, infrared thermal feature map representing thermal radiation contour and saliency, and lidar point cloud feature map representing lidar geometric contour and depth information. The status response acquisition module is used to acquire the status information of the aligned visible light image in the corresponding local area for each non-empty query point in the lidar point cloud feature map, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the status information and response information. The constrained attention matching module is used to generate local-level modal control coefficients based on local availability labels, and to perform constrained attention matching on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map according to the modal control coefficients, so as to obtain the constrained attention weight distribution. The fusion feature map generation module is used to generate an initial fusion feature map by weighted fusion of visible light feature map, infrared thermal feature map and lidar point cloud feature map based on a constrained attention weight distribution using a cross-modal attention fusion algorithm. The boundary correction module is used to obtain multimodal boundary consistency information of the initial fused feature map near the target edge, and to correct the boundary of the initial fused feature map to obtain a corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the fused feature semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary. The mask image generation module is used to generate a semantic segmentation mask image based on the corrected fusion feature map.

[0016] In summary, this invention provides a multimodal image semantic segmentation method and system for video analysis. The method effectively addresses the attention weight flattening problem caused by RGB feature noise contamination in low-light nighttime environments by introducing local availability recognition and constrained attention matching mechanisms. Through dynamic modal control coefficients, the system can adaptively adjust the contributions of each modality according to the local scene state, automatically enhancing the constraint effects of LiDAR and thermal imaging in RGB distortion areas. Furthermore, through multimodal boundary consistency arbitration and anisotropic diffusion filtering, the system can effectively eliminate physical boundary conflicts and achieve sharpening correction of target edges. This scheme significantly improves the boundary accuracy and category confidence of semantic segmentation masks under complex lighting conditions, ensuring the robustness of autonomous driving perception systems. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a multimodal image semantic segmentation method for video analysis provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the structure of a multimodal image semantic segmentation system for video analysis provided in an embodiment of the present invention.

[0019] Labeling Explanation: 210, Feature Extraction Module; 220, State Response Acquisition Module; 230, Restricted Attention Matching Module; 240, Fusion Feature Map Generation Module; 250, Boundary Correction Module; 260, Mask Map Generation Module. Detailed Implementation

[0020] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of this invention. The components of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0022] In low-light nighttime environments for autonomous driving or smart city roadside applications, vehicles are often equipped with visible light cameras, infrared thermal imagers, and LiDAR. However, the signal-to-noise ratio of visible light images drops sharply at night, and the textures of dark targets (such as pedestrians and vehicles without lights) are severely obscured by noise. Traditional multimodal processing uses LiDAR point clouds as the query benchmark to match and fuse textures in the visible light feature map. However, in low-light conditions, the similarity between the target and the background is extremely small, leading to a flattening of the cross-modal attention weight distribution. Query points mistakenly assign weights to background noise regions, blurring the semantic mask boundaries of the target and causing internal categories to degenerate into the background. To address the aforementioned problems of attention weight failure and segmentation boundary distortion caused by low-light noise, this invention proposes a novel multimodal image semantic segmentation processing architecture.

[0023] Firstly, please see Figure 1 This invention discloses a multimodal image semantic segmentation method for video analysis, used for semantic segmentation of detected targets in low-light driving environments, comprising: S1. Spatiotemporal alignment and feature extraction are performed on the synchronously acquired visible light image, infrared thermal image and lidar point cloud of the target to obtain visible light feature map representing visible light texture and semantics, infrared thermal feature map representing thermal radiation contour and saliency, and lidar point cloud feature map representing lidar geometric contour and depth information. S2. For each non-empty query point in the lidar point cloud feature map, obtain the state information of the aligned visible light image in the corresponding local area, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the state information and response information. S3. Based on the local availability label, generate local-level modal control coefficients, and according to the modal control coefficients, perform constrained attention matching on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map to obtain the constrained attention weight distribution. S4. Using a cross-modal attention fusion algorithm, based on a constrained attention weight distribution, the visible light feature map, infrared thermal feature map, and lidar point cloud feature map are weighted and fused to generate an initial fused feature map. S5. Obtain the multimodal boundary consistency information of the initial fused feature map near the target edge, and use it to correct the boundary of the initial fused feature map to obtain the corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the fused feature semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary. S6. Generate a semantic segmentation mask map based on the corrected fusion feature map.

[0024] Specifically, a visible light feature map refers to a high-dimensional feature matrix containing two-dimensional visual details such as color and texture, extracted from a visible light image through a feature extraction network. It is not essentially a simple copy of the original pixel values, but rather a semantic representation formed after processing such as convolution, downsampling, and nonlinear activation. An infrared thermal feature map refers to a feature matrix reflecting the distribution of thermal radiation intensity and temperature profile of an object's surface, extracted from infrared thermal imaging data. This feature map focuses more on the thermal difference between the target and the background, rather than color differences. A lidar point cloud feature map refers to a feature matrix containing geometric spatial information such as depth and curvature, formed by mapping three-dimensional lidar point cloud data onto a two-dimensional plane after voxelization or projection processing. A non-empty query point refers to a point where a valid point cloud echo exists at the corresponding location after projection, and which is not judged as an invalid point or isolated noise point after preprocessing. For example, the shoulders of pedestrians in front of vehicles, the edges of vehicle bodies, and roadside guardrails often form relatively obvious geometric structure responses in the point cloud.

[0025] In this scenario, the query vector (Query) consists of geometric feature points provided by the LiDAR, while the key vector consists of pixel features on the visible light and infrared feature maps. When the query vector and a certain key vector are more semantically and spatially consistent, that location receives higher weight in subsequent fusion.

[0026] Local availability labels are status markers generated for the visible light reliability of a local area corresponding to a query point. They distinguish whether the area can be directly trusted, requires careful reference, or should be heavily suppressed during attention matching. This label is generated by combining visible light status, infrared response, and lidar response.

[0027] Modal control coefficients refer to the proportion of contributions of visible light, infrared thermal, and lidar modes at a certain local location. Their function is not simply to turn a mode on or off, but to smoothly adjust the participation level of each mode under different local conditions.

[0028] Multimodal boundary consistency information refers to a type of boundary constraint information used to describe whether the semantic boundaries of fused features, the physical depth boundaries of LiDAR, and the thermal boundaries of infrared thermography are aligned, how they are offset, and how much conflict they have. This information is used to further refine the boundaries after fusion, making the final mask closer to the real physical contour.

[0029] In practice, the closest image frame in time is obtained through hardware timestamp matching, and motion compensation is performed on the LiDAR point cloud using the vehicle's pose estimation system, such as inertial measurement unit data. Then, a pre-calibrated sensor intrinsic and extrinsic parameter matrix is ​​invoked to uniformly project and align the 3D point cloud and infrared thermal map onto the 2D pixel coordinate system of the visible light image. Specifically, the visible light image is continuously acquired by an RGB camera mounted on a pole in front of the vehicle or on the roadside, the infrared thermal image is synchronously acquired by an infrared thermal imaging camera whose field of view overlaps with that of the RGB camera, and the LiDAR point cloud is output by a LiDAR mounted on the roof, front of the vehicle, or on the roadside. All three types of sensors have their acquisition timestamps written through a unified clock source or time synchronization module. For pre-calibrated sensor intrinsic and extrinsic parameter matrices, offline calibration can be used. For example, with the vehicle stationary, a calibration board with corner points or high reflectivity markers is placed at multiple different distances and orientations, and RGB images, infrared images, and point cloud data are collected respectively. Corner points of the calibration board are extracted from the RGB images, planes or edges of the calibration board are extracted from the point cloud, and the center of the thermally marked region is extracted from the infrared image. This establishes the correspondence between the three sensors. After iterative optimization using multiple sets of samples, stable intrinsic and extrinsic parameters are obtained. After calibration, the calibration results are stored as a configuration file and directly called during online inference. For feature extraction, deep convolutional networks such as ResNet-50 can be used to extract visible light and infrared feature maps, and 3D networks such as VoxelNet can be used to extract point cloud features and map them into 2D feature maps. To facilitate implementation, the visible light and infrared branches can share the same level of backbone structure but not the same weights, in order to avoid the two imaging mechanisms being too different and causing mutual interference in feature representation; the point cloud branch first performs voxelization and voxel feature encoding, and then projects the encoding results onto a two-dimensional plane aligned with the image to form the lidar point cloud feature map required for subsequent attention matching.

[0030] In the process of generating local availability labels, the core objective is to pre-determine the reliability of each region of the visible light image before fusion, thereby avoiding invalid visible light noise features in low-light environments from participating in attention matching and addressing the inherent defect of attention weight flattening. This method uses effective non-empty query points from the LiDAR as a spatial benchmark to delineate matching local neighborhoods. It collects two imaging state indicators from the visible light image: signal-to-noise ratio and texture gradient. Simultaneously, it collects infrared modal temperature gradient and thermal radiation contrast, as well as two geometric response indicators from the LiDAR: modal depth variance and local curvature. This achieves multi-dimensional joint reliability assessment, avoiding the problem of inaccurate judgments based on single modal indicators. Combining a lightweight state-aware network or fixed logic judgment rules, all local regions are divided into three categories: effective regions, low-confidence regions, and distorted regions. Effective regions correspond to target areas with relatively good lighting and complete, clear visible light textures, exhibiting high reliability of visible light features. Low-confidence regions correspond to transitional areas with weak lighting and slightly damaged textures, where visible light features are subject to noise interference. Distorted regions correspond to areas with severe darkness at night, where visible light textures are completely submerged by noise, rendering visible light features completely ineffective and without reference value. Based on this refined labeling result, the weights of the three modalities can be dynamically adjusted pixel by pixel to accurately shield invalid visible light features in noisy areas, ensuring that subsequent attention matching is calculated based solely on valid features.

[0031] In the process of generating modal control coefficients and performing constrained attention matching, the aim is to use the aforementioned labels to intervene in attention calculation, breaking the inherent pattern of indiscriminate cross-modal attention matching, and achieving global constraints and local noise reduction in the attention calculation process. This invention dynamically allocates contribution weights for three modalities pixel-by-pixel based on the availability labels of different regions: In visible light reliable regions, visible light texture features are prioritized to supplement details, while infrared and LiDAR auxiliary weights are weakened; in visible light low-reliability regions, the proportions of the three modal features are balanced, taking into account visual texture, thermal radiation information, and geometric depth information; in visible light completely distorted regions, visible light feature weights are directly suppressed, and feature matching is completed entirely using infrared thermal contours and LiDAR depth geometric features. Simultaneously, the attention search range is adaptively constrained by combining the depth and curvature information of the LiDAR point cloud itself, directly classifying pixels marked as distorted regions into noise rejection intervals, and additionally adding a spatial distance penalty term to suppress interference from irrelevant features at long distances. This completely avoids invalid noise features preempting attention weights, solving the problems of attention weight flattening and indistinguishability between target and background features in low-light scenes, allowing attention weights to accurately focus on the effective feature regions of the real target.

[0032] The main purpose of acquiring multimodal boundary consistency information and performing boundary correction is to address potential edge roughness issues after fusion. After obtaining the initial fused feature map through cross-modal attention fusion, the visible light semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary naturally exhibit pixel-level spatial offsets and local boundary conflicts due to differences in imaging principles and detection dimensions. Direct segmentation can result in defects such as jagged target edges, contour adhesion, and edge breakpoints. This invention constructs multimodal boundary consistency information based on the degree of spatial offset and matching consistency of the three boundaries, and performs targeted fine-grained edge correction. In a specific implementation, the standard Canny edge detection operator can be used to extract the semantic edges, lidar depth edges, and infrared thermal edges of the initial fused feature map, and calculate the coordinate differences of the three types of edges on the two-dimensional pixel plane. If the pixel offset difference exceeds a preset boundary offset tolerance threshold, a boundary conflict is determined in the region, requiring boundary correction processing. The preset pixel distance is the boundary offset tolerance range, which can be statistically calibrated based on the dataset validation set or adaptively set according to the feature map resolution to adapt to input images of different scales. This scheme abandons the crude correction method of global translation and optimizes only the local neighborhood of the target edge. Using a bilinear interpolation algorithm, the offset semantic edge feature vector is locally aligned and resampled to the high-precision, high-reliability LiDAR depth boundary coordinates, preserving the complete features of the target's interior to the greatest extent possible and avoiding damage to the target's semantic information during correction. After boundary alignment, the boundary mask is fine-tuned by combining target motion information and local deformation information. Anisotropic diffusion filtering is used to smooth edge spikes, ultimately eliminating edge distortion caused by multimodal fusion. Finally, guided filtering and morphological reconstruction are applied to the residual blurred and discontinuous regions at the edges of the corrected fused feature map. The optimized corrected fused feature map is then input into the semantic segmentation head of a fully convolutional network (FCN), which outputs the predicted probabilities of each category pixel by pixel, ultimately generating a semantic segmentation mask map with continuous edges, accurate contours, and a close fit to the real physical boundaries. To more clearly illustrate the implementation of this step, the following processing details can be added: The initial fused feature map is usually first subjected to channel compression or edge response mapping to obtain a single-channel response map suitable for edge extraction; the depth edge of the LiDAR can be accurately extracted from the depth change position of adjacent pixels in the depth map, and the infrared thermal edge is extracted based on the temperature gradient change position, ensuring that the three types of boundary extraction benchmarks are consistent and further improving the boundary correction accuracy.

[0033] Through the aforementioned overall technical solution, this invention alters the logic of indiscriminate global matching in traditional cross-modal attention mechanisms. Its working principle involves introducing a local availability assessment mechanism before performing attention similarity calculations to identify areas severely contaminated by noise. A restricted attention matching mechanism then forcibly severs the matching path between high-confidence geometric query points from the LiDAR and high-intensity background noise features. Subsequently, after feature fusion, multimodal boundary consistency information is introduced again as a post-processing constraint, forcing the semantic boundary to converge towards the physical boundary. This dual-guarantee logic of pre-matching constraints and post-fusion correction enables the system to effectively prevent erroneous aggregation of noisy background information in low-light nighttime scenarios, ensuring that the target area is primarily filled with reliable geometric and thermal radiation information. This achieves the technical effect of maintaining clear target mask boundaries and accurate semantic category determination even under adverse lighting conditions.

[0034] In real-world nighttime urban road scenarios, lighting conditions are often extremely complex and unevenly distributed. For example, a section of the road may be partially illuminated by bright streetlights, while an area just a few meters away may be completely dark. If we rely solely on a global brightness threshold to determine the usability of an area, it is highly likely that dark targets at the edge of streetlights will be misjudged as unusable, or noisy areas affected by glare from oncoming headlights will be misjudged as usable.

[0035] To achieve more refined and accurate local state perception, in some preferred embodiments, for each non-empty query point in the lidar point cloud feature map, the steps of acquiring the state information of the aligned visible light image in the corresponding local region, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same region, and generating a local availability label based on the state information and response information include: Traverse each non-empty query point in the lidar point cloud feature map and obtain the projection coordinates of the non-empty query point on the aligned visible light image; with the projection coordinates as the center and according to the preset pixel window size or the search radius determined based on the depth and curvature of the non-empty query point, delineate the local neighborhood corresponding to the query point as the visible light local region. Within a local visible light region, local signal-to-noise ratio and texture gradient information are obtained from the aligned visible light image as state information; At the same local neighborhood location, temperature gradient and thermal radiation contrast information are extracted from infrared thermal feature map as first response information, and depth variance and local curvature information are extracted from lidar point cloud feature map as second response information. The state information, first response information, and second response information are input into the state-aware network or a preset logical judgment rule to generate a local availability label for a local area. The local availability label is marked as a valid area label, a low-confidence area label, or a distorted area label.

[0036] In this specific implementation scheme, a multi-dimensional local feature evaluation system is constructed. Specifically, firstly, high-frequency components are extracted from the visible light image using the Laplacian operator, and the true local signal-to-noise ratio and texture gradient are calculated by combining the local brightness variance. Simultaneously, the difference between the highest and lowest local temperatures (thermal radiation contrast) is calculated on the infrared thermogram, and the dispersion of local depth values ​​(depth variance) is statistically analyzed on the point cloud feature map. To enable direct implementation of this step, the following sequence can be followed: First, based on the projected coordinates of each non-empty query point, extract the same local neighborhood from the three aligned datasets. The calculation logic for determining the search radius based on the depth and curvature of the non-empty query points can be referenced from the method for determining the two-dimensional search radius in the constrained attention matching section later. In the visible light local region, perform weak denoising first, then statistically analyze the mean brightness, brightness fluctuation, and high-frequency response strength to distinguish between two cases: overall darkness but the structure remains, and overall darkness where the structure is submerged by noise. In the infrared local region, in addition to calculating the thermal radiation contrast, check whether the thermal gradient is continuously distributed along the target edge to avoid mistaking isolated thermal noise points for real targets. In the point cloud local region, in addition to depth variance, check whether the number of valid points meets the minimum support condition to avoid misjudgment caused by sparse echoes. Subsequently, concatenate these three sets of data into a state feature vector and input it into a pre-trained lightweight multilayer perceptron (MLP, i.e., state-aware network). This network will output three probability values, corresponding to the valid, low-confidence, and distorted states, respectively. For example, when the visible light signal-to-noise ratio is extremely low, but the thermal radiation contrast is extremely high and the depth variance is small (indicating that there is a real physical entity here, but it is not visible to the naked eye), the network will identify it as a low-confidence area tag instead of simply discarding it.

[0037] A pre-trained lightweight multilayer perceptron (MLP) can be trained as follows: Multiple sets of synchronized visible light images, infrared thermal images, and LiDAR point clouds are collected from nighttime road scenes. Local regions are first labeled as valid, low-confidence, or distorted using a combination of manual annotation and semi-automatic rule-based methods. During manual annotation, the visible light region, its corresponding infrared region, and the corresponding point cloud projection region are simultaneously displayed on the annotation interface, allowing the annotator to determine the presence of a real target structure based on whether all three indicators support the existence of such a region. The semi-automatic rule-based method first filters out obviously distorted regions, such as those with high noise in visible light, lack of continuous thermal response in infrared, or no stable geometric support in the point cloud, which are then manually verified. After annotation, features such as local signal-to-noise ratio, texture gradient, thermal radiation contrast, temperature gradient, depth variance, and local curvature are used as input samples to train the MLP classifier. When the recognition results for the three states on the validation set stabilize and the recall capability for low-confidence regions meets expectations, the network can be solidified as a state-aware network for online inference. If a network approach is not used, a preset logical judgment rule can be adopted. For example, first determine whether the visible light state is obviously distorted, and then use the infrared and point cloud responses to decide whether to downgrade to low confidence or directly judge it as distorted.

[0038] After obtaining fine-grained local availability labels, simply using a binary mask (either 0 or 1) to control feature participation will result in severe feature fragmentation at the boundary between effective and distorted regions. This hard truncation disrupts the semantic continuity within the target, making it difficult for downstream segmentation networks to extract coherent contextual features.

[0039] To achieve a smooth transition and dynamic transfer of modal contributions, in some preferred embodiments, the step of generating local-level modal control coefficients based on local availability labels includes: For each pixel location on the visible light feature map, read the local availability label corresponding to that pixel location; When the read local availability tag is a valid area tag, the visible light mode control coefficient of the pixel location is set as the first weight, and the infrared thermal mode control coefficient and lidar mode control coefficient of the pixel location are set as auxiliary weights; the first weight is greater than the auxiliary weight. When the read local availability tag is a low-confidence area tag, the visible light mode control coefficient of that pixel location is reduced to the second weight, and the infrared thermal mode control coefficient and lidar mode control coefficient of that pixel location are increased to the equivalent weight; the second weight is less than the first weight, and the equivalent weight is greater than the auxiliary weight; When the read local availability label is a distorted area label, the visible light mode control coefficient of that pixel location is attenuated to a third weight value approaching zero, and the infrared thermal mode control coefficient and lidar mode control coefficient of that pixel location are increased to the dominant weight value; the third weight value is less than the second weight value, and the dominant weight value is greater than the equivalent weight value; Output the modal control coefficients for this pixel location.

[0040] In this specific implementation scheme, the system establishes a state-based dynamic weight allocation mechanism. In practice, a set of consecutive floating-point numbers can be pre-set as weight parameters. For example, the first weight is set to 0.8, with an auxiliary weight set to 0.1, ensuring that visible light texture dominates under good lighting conditions; the second weight is set to 0.4, with an equivalent weight set to 0.3, allowing the three modes to work together and complement each other in low-confidence regions; the third weight is set to 0.05 (approaching zero but not zero to maintain weak connectivity for gradient backpropagation), with a dominant weight set to 0.475, allowing thermal and geometric information to take over completely in distorted regions. The system reads the label pixel by pixel and assigns corresponding control coefficients by looking up a table, ultimately outputting three modality control coefficient maps with the same resolution as the feature map. To more clearly illustrate the source of this pre-set parameter, it can be obtained through offline validation set search: First, prepare a dataset containing various scenarios such as low-light conditions at night, local glare, backlight, and heat source interference. Try multiple weight combinations that satisfy the condition that the visible light weight is highest in the effective area and lowest in the distorted area, while infrared and LiDAR dominate in the distorted area. Then, compare the performance of each combination in terms of boundary sharpness, target recall rate, and false detection suppression, and select the combination with the best overall performance as the deployment parameters. If different vehicle models, different camera gain strategies, or different infrared imaging devices lead to significant differences in modal quality, the first weight, second weight, third weight, auxiliary weight, equivalent weight, and dominant weight can also be used as configurable parameters and re-determined through small-scale scene calibration before deployment. The working principle of this specific implementation scheme is to transform discrete availability labels into continuous modal control coefficients, which directly affect the weighted calculation of subsequent feature fusion.

[0041] In traditional attention mechanisms, the query vector typically needs to be matched against all key vectors globally across the entire image. In nighttime scenes, even with modality control coefficients, global search still carries significant risks: distant background noise pixels, after undergoing multiple nonlinear transformations in the network, may have feature vectors that coincidentally exhibit a high degree of mathematical agreement with the query vector, resulting in artificially inflated matching responses.

[0042] To completely block the interference of long-distance noise at the physical space level, in some preferred embodiments, the step of performing constrained attention matching between the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map, based on the modal control coefficients, to obtain the constrained attention weight distribution includes: Obtain the depth value and local curvature information of the query vector; Based on the depth value and local curvature information of the query vector, the two-dimensional search radius is calculated, and the initial candidate matching range of the query vector is determined accordingly. Within the initial candidate matching range, local availability labels are read pixel by pixel. Pixels marked as distorted regions are included in the noise rejection range, while pixel positions marked as valid regions and low-confidence regions are retained as the corrected candidate matching range. Based on the corrected candidate matching range and modality control coefficients, constrained attention matching is performed on the query vector and key vector to obtain a constrained attention weight distribution.

[0043] In this specific embodiment, the system introduces a spatial constraint mechanism based on three-dimensional geometric priors. During specific implementation, the system extracts the physical depth and local curvature inherent to LiDAR query points. According to the principle of perspective projection, closer targets occupy larger pixel areas on an image, so the two-dimensional search radius has an inverse relationship with the depth value; meanwhile, feature changes are more dramatic in regions with larger curvature (such as vehicle edges), and the two-dimensional search radius needs to be appropriately expanded to accommodate deformation. The two-dimensional search radius can be determined by means of table lookup or piecewise rules: for example, depth is first divided into near-distance, mid-distance and far-distance intervals, then each interval is divided into flat regions and high-change regions according to the curvature state, which correspond to small, moderate, and large search radius levels respectively. This mapping relationship can be obtained through offline statistics, and the specific method is: in labeled training samples, count the actual offset distribution of the same physical point in the aligned image under different depth and curvature conditions, then select a local neighborhood scale that can cover most real matching positions, and form a preset mapping relationship of depth interval - curvature state - search radius level. The foregoing establishes the mapping relationship of depth interval - curvature state - search radius level through offline statistics, and this level is only a gear division, not a directly usable pixel radius; a fixed reference pixel radius value is pre-configured for each search radius level, so the actual pixel size can be converted from the level. The specific implementation mode is as follows: pre-calibrate the reference pixel radius corresponding to different levels offline: for example, the small radius level configuration corresponds to pixel radius r1, the medium radius level configuration corresponds to pixel radius r2, and the large radius level configuration corresponds to pixel radius r3, which satisfies r1 < r2 < r3. After the system queries and obtains the search radius level, it directly matches the reference pixel radius preset for this level, takes the projection coordinates of the query vector on the image as the center and the matched pixel radius as the side length to construct a square neighborhood or a circular neighborhood, and this neighborhood is the initial candidate matching range corresponding to the query vector. Subsequently, the system performs an intersection operation between this range and the local availability label map, removes pixels marked as distorted regions within the range, and forms a corrected candidate matching range with irregular shape but high reliability. To further improve stability, after removing distorted regions, it is also possible to check whether the number of remaining candidate points is lower than the minimum matching support condition; if it is lower than this condition, candidate positions that are closer to the query point in spatial distance are allowed to be preferentially supplemented from low-confidence regions, instead of restarting global search.

[0044] After the corrected candidate matching range is determined, the matching value of each key vector in the range for the query vector is not completely equal. In the physical world, features belonging to the same target are usually closely adjacent in space. If only the inner product similarity of feature vectors is calculated, matching weights may be evenly scattered within the candidate range, and a focused attention center cannot be formed.

[0045] To further enhance the spatial locality of matching and completely suppress noise, the steps for obtaining the constrained attention weight distribution by performing restricted attention matching on the query vector and key vector based on the corrected candidate matching range and modality control coefficients include: Calculate the initial similarity between the query vector and each key vector within the corrected candidate matching range; Based on the Euclidean distance between the position of each key vector within the corrected candidate matching range and the projection center of the query vector, a spatial distance penalty term is applied to the initial similarity. Based on the modal control coefficients at the locations of each key vector, the similarity after spatial distance penalty is weighted and modulated. Among them, the similarity of key vectors located within the noise exclusion range is forcibly assigned a preset minimum negative value. The modulated similarity is normalized to generate a constrained attention weight distribution.

[0046] In this specific implementation, the system introduces distance decay and forced penalty mechanisms into the similarity metric space. Specifically, the initial similarity matrix is ​​first obtained by calculating the dot product of the query vector and the key vector. Next, the system calculates the two-dimensional Euclidean distance from each pixel within the candidate range to the center of the circle, and generates a spatial distance penalty term based on the distance; the farther the pixel is from the center, the greater the reduction in its similarity. The penalty method can employ a monotonically decreasing lookup table function, a piecewise decay rule, or a preset distance level weight, without being limited to a specific function form. Subsequently, the similarity is modulated using the aforementioned generated modality control coefficients. Most importantly, for pixels falling within the noise exclusion range, the system does not simply multiply by zero, but instead forces their similarity to a minimal negative value that is completely negligible during normalization. This minimal negative value is not limited to a single, unique value, as long as it does not make an observable contribution to the final attention weight under the chosen data type and normalization implementation. To facilitate engineering implementation, a masking constant can be set in the program, and numerical stability tests should be conducted before deployment to confirm that this constant will not cause overflow or abnormal gradients. Finally, the entire modulated similarity matrix is ​​input into the Softmax function for normalization.

[0047] After feature fusion, although noise inside the target is effectively suppressed, severe boundary conflicts often arise at the target edges due to differences in the physical imaging characteristics of different sensors. For example, heated targets (such as pedestrians) are prone to thermal halo effects in infrared thermal maps, causing the extracted thermal boundary to expand significantly outward compared to the actual physical contour. Meanwhile, when scanning target edges, lidar, limited by the beam divergence angle, is prone to pixel mixing, resulting in discontinuous or inward contraction of the depth boundary. Simply aligning or merging the two as in conventional methods would create chaotic control signals at the edges, making subsequent boundary corrections ineffective.

[0048] To address the misleading issue caused by conflicts in multimodal physical boundaries, in practical applications, the steps for obtaining multimodal boundary consistency information of the initial fused feature map near the target edge include: Obtain the depth boundary map corresponding to the aligned lidar point cloud, the thermal radiation boundary map corresponding to the aligned infrared thermal image, and the semantic boundary response map of the initial fused feature map; Morphological dilation operations are performed on the depth boundary map and the thermal radiation boundary map respectively to obtain the dilated depth boundary map and the dilated thermal radiation boundary map. Calculate the spatial intersection of the expansion depth boundary map and the expansion thermal radiation boundary map to obtain the boundary intersection region; Based on the spatial distance between depth boundary pixels and thermal radiation boundary pixels within the boundary intersection region, identify physical boundary conflict zones within the boundary intersection region; For the physical boundary conflict zone, the local connectivity length and depth gradient magnitude of the depth boundary map are evaluated to obtain the depth state; at the same time, the temperature gradient direction symmetry and attenuation width of the thermal radiation boundary map are evaluated to obtain the thermal state. Based on the depth and thermal states, within the physical boundary conflict zone, one of the following operations may be selectively performed: preserve the blocking properties of the depth boundary map, perform morphological shrinkage on the thermal radiation boundary map or generate a smooth blocking line based on the peak line of the semantic boundary response map, the shrunken thermal radiation boundary and the discontinuous depth boundary. Spatially stitching the boundary outside the physical boundary conflict zone with the boundary generated inside the physical boundary conflict zone yields a composite conduction blocking mask, which is the multimodal boundary consistency information.

[0049] In this specific implementation scheme, the system constructs a rigorous physical boundary conflict arbitration mechanism. Specifically, the system first utilizes morphological dilation to expand the response range of the depth and thermal boundaries, and then quickly identifies conflict zones where the two boundaries do not overlap or have significant offsets by finding their intersection. The morphological dilation structuring element is not limited to a unique size; smaller local structuring elements can be selected based on the boundary map resolution and target scale to ensure that slight offsets are covered without incorrectly connecting adjacent independent targets. Upon entering the conflict zone, the system calculates the number of connected pixels (local connectivity length) of the depth boundary and the gradient decay width of the thermal boundary, respectively. If the depth connectivity is good and the gradient is large, but the thermal boundary decays widely (indicating severe thermal halo), the system determines it to be a strong depth, weak thermal state. In this case, the thermal boundary is directly erased, and only the depth boundary is retained. If the depth boundary is severely discontinuous, but the thermal boundary outline is clear, the system determines it to be a weak depth, strong thermal state. In this case, based on the degree of outward diffusion of the thermal boundary, a morphological contraction operation is performed on the thermal boundary towards the target center. If neither is reliable, the semantic boundary peak line of the initial fusion feature is extracted, and combined with the coordinates of the contracted thermal boundary and the discontinuous depth boundary, a virtual smooth blocking line is generated. To enable the smooth blocking line to be implemented directly, the following method can be used: First, extract the center line with the strongest semantic boundary response within the conflict zone, and then compare this center line with the local orientation of the thermal boundary and the depth boundary. If the three orientations are basically consistent, smoothing is performed mainly based on the semantic boundary center line. If the semantic boundary is locally missing, it is first supplemented along the continuous direction of the depth boundary, and its outward expansion range is limited by the contracted thermal boundary. Finally, the arbitrated boundary is spliced ​​with the boundary of the non-conflict zone to generate a globally unique composite conduction blocking mask.

[0050] Specifically, the depth state can be determined by the local connectivity length, boundary continuity, and whether depth abrupt changes are stable; the thermal state can be determined by whether the temperature gradient direction is symmetrically distributed around the target edge, whether the outward decay of the thermal boundary is too wide, and whether there are isolated hotspots. These states can be determined either through preset logical rules or by a small, offline-trained classifier. If preset logical rules are used, a state determination table can be established. For example, if the depth is continuous and abrupt changes are stable—the depth boundary is preferentially preserved; if the thermal boundary is wide and diffuse—thermal boundary contraction is performed; if both are weak—semantic boundary peak line compensation is invoked.

[0051] After obtaining an accurate composite conduction blocking mask, if it is directly applied to the static feature map correction, when facing high-speed moving targets (such as pedestrians crossing the road or fast-moving vehicles), due to the sensor sampling time difference and the motion blur of the target itself, the position of the static mask may have a slight spatial misalignment with the actual semantic position in the feature map. This misalignment will cause the correction process to block features at the wrong location, thus fragmenting the originally complete semantic features.

[0052] To adapt to boundary correction in dynamic scenarios, in some preferred embodiments, the boundaries of the initial fused feature map are corrected based on multimodal boundary consistency information to obtain the corrected fused feature map. The steps include: Obtain motion information of the target region corresponding to the composite conduction blocking mask; Obtain local deformation information of the initial fused feature map; Based on motion information and local deformation information, the composite conduction blocking mask is adjusted to obtain the adjusted composite conduction blocking mask. Anisotropic diffusion filtering is used, with the adjusted composite conduction blocking mask as the conduction coefficient control term, to perform iterative diffusion processing on the initial fused feature map to obtain the corrected fused feature map.

[0053] In this specific implementation scheme, the system introduces motion compensation and partial differential equation filtering techniques. Specifically, the system first extracts the target's motion vector and the local optical flow deformation field of the feature map. Motion information can be obtained from consecutive video frames: for example, by estimating dense optical flow from visible light images of the previous and current moments, estimating geometric displacement from point clouds of two consecutive frames, and then combining this with target tracking results to obtain the main motion direction and displacement trend of the target region. Local deformation information can be obtained by comparing the response changes of the fused features near the edges of the previous and subsequent frames, used to distinguish between rigid translation and non-rigid deformation. Using this dynamic information, a small affine transformation or elastic deformation adjustment is performed on the composite conduction blocking mask to make it conform to the instantaneous state of the current feature map. "Small" means that the adjustment range is limited to reasonable motion displacement between adjacent frames and will not cross into other independent target regions. Subsequently, the system constructs an anisotropic diffusion filtering process. In this process, the adjusted composite conduction blocking mask is directly used as the conduction coefficient control term: at the location of the mask lines, i.e., the target boundary location, conduction is significantly suppressed; in the areas inside and outside the mask, conduction remains allowed. The system uses the initial fused feature map as the initial state and performs multiple iterations until the internal features of the target become smooth and the boundary leakage phenomenon is suppressed. For ease of implementation, the number of iterations can be determined as a deployment parameter through a validation set, or it can be set to terminate when the change in the boundary region after two consecutive iterations is lower than a preset stopping condition.

[0054] When extracting target motion information to adjust the mask, if there are multiple targets intersecting in the scene (e.g., a pedestrian walking behind a parked vehicle, or two vehicles driving side by side), simple global optical flow calculation will confuse the motion vectors of the foreground target and the background occluded target, resulting in serious deviations in the calculated motion information, which in turn leads to incorrect mask adjustment direction.

[0055] To obtain accurate individual motion information in complex occlusion scenarios, in some implementations, the steps for obtaining motion information of the target region corresponding to the composite conduction blocking mask include: Identify multiple moving targets within a composite conduction blocking mask area; For each moving target individual, calculate the motion vector of that moving target individual; Determine whether there is a spatial occlusion relationship between moving target individuals; When there is a spatial occlusion relationship between moving target individuals, the background motion information or the motion information of the occluded target within the occluded area is suppressed, and the motion vector of the occluded target individual is inferred. By integrating the motion vectors of individual moving targets, motion information of the target region corresponding to the composite conduction blocking mask is generated.

[0056] In this specific implementation scheme, the system constructs motion analysis logic based on individual tracking and occlusion inference. Specifically, firstly, instance segmentation or clustering algorithms are used to distinguish independent moving individuals within the masked region. This individual identification can be jointly completed based on semantic regions in the modified fusion feature map, point cloud clustering results, and positional continuity in consecutive frames. By comparing the feature positions of consecutive frames, the initial motion vector of each individual is calculated. Next, the system uses depth information provided by LiDAR to determine the spatial hierarchy between individuals. When a shallow foreground target is found to occlude a deeper background target, the system actively masks pixels in the overlapping area when calculating the motion vector of the background target to prevent interference from foreground motion. For background targets that are completely or partially occluded, the system introduces a Kalman filter, using its historical motion trajectory and the local motion trend of the currently unoccluded portion to infer its true motion vector within the occluded region. To more clearly illustrate how this filter is obtained and used, the following approach can be adopted: When the target is first detected, an independent motion state record is established for it, including the direction of position changes, displacement trends, and changes in the size of the visible area in the previous few frames; when the target enters an occlusion state, the prediction is no longer entirely dependent on the current frame observation, but rather combines historical states for short-term prediction; when the target reappears, the predicted trajectory is corrected using new observation results. Finally, all the corrected and inferred individual vectors are integrated into a global motion information field.

[0057] Although anisotropic diffusion filtering can create sharp boundaries on a macroscopic level, at the micro-pixel level of the feature map (especially for extremely complex non-rigid target edges, such as a pedestrian's fingers or a fluttering hem of clothing), due to discretization errors in the diffusion equation, small residual blurs or isolated breaks may still exist in local edge pixels. If a feature map with these minor flaws is directly fed into the segmentation head, the final output mask edges may appear jagged or have individual pixels stuck together.

[0058] To provide pixel-level edge polishing in the final stage, in some preferred embodiments, the step of generating a semantic segmentation mask map based on the modified fused feature map includes: Obtain the feature gradient distribution and local connectivity of the local edge regions of the corrected fused feature map; The fine contour information of the corresponding local edge region is obtained from the depth boundary map of the aligned lidar point cloud and the thermal radiation boundary map of the aligned infrared thermal image. Based on feature gradient distribution, local connectivity, depth boundary map and refined contour information, determine whether there is residual blur or discontinuity in the local edge region of the corrected fused feature map. If residual blur or discontinuity exists, guided filtering or morphological reconstruction is performed on the local edge regions of the corrected fused feature map based on the depth boundary map and refined contour information. The processed and fused feature map is input into the semantic segmentation head to generate a semantic segmentation mask map.

[0059] In this specific implementation, the system adds a micro-edge quality inspection and reconstruction process before entering the segmentation head. Specifically, the system calculates the gradient magnitude distribution of feature vectors in the edge regions of the corrected fusion feature map and checks the eight-neighbor connectivity of high-gradient pixels. Simultaneously, it retrieves the original high-resolution depth boundary and thermal radiation fine contour as reference benchmarks. If a feature gradient fails to reach a local peak or connectivity is interrupted, the system determines that residual fuzziness exists. In this case, the system uses the high-resolution fine contour information as a guide map to perform guided filtering on the feature map of that local region; or it uses morphological closing operations to reconstruct and connect discontinuous feature responses. The fine contour information can be obtained as follows: after extracting the thermal boundary from the infrared thermogram, the boundary is refined to retain the continuous main contour; after extracting the depth abrupt change line from the depth boundary map, isolated short edges and obvious noise edges are removed, and then both are locally aligned with the semantic boundary in the corrected fusion feature map to obtain a reference contour for micro-repair. The choice between guided filtering and morphological reconstruction depends on the type of residual problem: if the edges are slightly blurred but the contours are continuous, guided filtering is preferred; if there are breakpoints, holes, or local adhesions at the edges, morphological reconstruction is preferred. After this micro-processing, the feature map is fed into a semantic segmentation head consisting of multiple deconvolution layers for final category mapping. Alternatively, an upsampling and convolution segmentation head can achieve the same function, as long as it outputs a pixel-level category probability map corresponding to the spatial resolution of the input image.

[0060] Secondly, see Figure 2The present invention also discloses a multimodal image semantic segmentation system for video analysis, used to perform the above-described method, the system comprising: The feature extraction module 210 is used to perform spatiotemporal alignment and feature extraction on the visible light image, infrared thermal image and lidar point cloud of the target acquired simultaneously, to obtain a visible light feature map that represents the visible light texture and semantics, an infrared thermal feature map that represents the thermal radiation contour and saliency, and a lidar point cloud feature map that represents the lidar geometric contour and depth information. The state response acquisition module 220 is used to acquire the state information of the aligned visible light image in the corresponding local area for each non-empty query point in the lidar point cloud feature map, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the state information and response information. The constrained attention matching module 230 is used to generate local-level modal control coefficients based on local availability labels, and to perform constrained attention matching on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map according to the modal control coefficients, so as to obtain the constrained attention weight distribution. The fusion feature map generation module 240 is used to generate an initial fusion feature map by weighted fusion of visible light feature map, infrared thermal feature map and lidar point cloud feature map based on a constrained attention weight distribution using a cross-modal attention fusion algorithm. The boundary correction module 250 is used to obtain multimodal boundary consistency information of the initial fused feature map near the target edge, and to correct the boundary of the initial fused feature map to obtain a corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the fused feature semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary. The mask image generation module 260 is used to generate a semantic segmentation mask image based on the modified fusion feature map.

[0061] To facilitate system-level implementation, each module can be deployed on a unified computing platform either as a software thread or an inference subgraph, or it can be broken down into preprocessing units, fusion units, and post-processing units according to the sensor processing chain. The feature extraction module 210 and the restricted attention matching module 220 can be preferentially deployed on hardware accelerators with parallel computing capabilities, while the state response acquisition module 230 and the boundary correction module 250 can operate collaboratively with the main control processor. During system operation, the three types of sensor data first enter a unified cache queue, are synchronized in time, and then sent to the feature extraction module 210. The data then flows sequentially through the state response acquisition module 220, the restricted attention matching module 230, the fusion feature map generation module 240, the boundary correction module 250, and the mask map generation module 260, ultimately outputting a semantic segmentation mask map, which can be further provided to target detection, trajectory prediction, path planning, or roadside event analysis modules.

[0062] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal image semantic segmentation method for video analysis, used for semantic segmentation of detected targets in low-light driving environments, characterized in that, include: Spatiotemporal alignment and feature extraction are performed on the visible light image, infrared thermal image and lidar point cloud of the target acquired simultaneously to obtain visible light feature map representing visible light texture and semantics, infrared thermal feature map representing thermal radiation contour and saliency, and lidar point cloud feature map representing lidar geometric contour and depth information. For each non-empty query point in the lidar point cloud feature map, obtain the state information of the aligned visible light image in the corresponding local area, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the state information and the response information. Based on the local availability label, local-level modal control coefficients are generated. Based on the modal control coefficients, constrained attention matching is performed on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map to obtain a constrained attention weight distribution. Using a cross-modal attention fusion algorithm, based on the constrained attention weight distribution, the visible light feature map, the infrared thermal feature map, and the lidar point cloud feature map are weighted and fused to generate an initial fused feature map; The multimodal boundary consistency information of the initial fused feature map near the target edge is obtained to correct the boundary of the initial fused feature map and obtain a corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the fused feature semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary. Based on the modified fusion feature map, a semantic segmentation mask map is generated.

2. The multimodal image semantic segmentation method for video analysis according to claim 1, characterized in that, The step of obtaining the state information of the aligned visible light image in the corresponding local region for each non-empty query point in the lidar point cloud feature map, and the response information of the infrared thermal feature map and the lidar point cloud feature map in the same region, and generating a local availability label based on the state information and the response information includes: Traverse each non-empty query point in the lidar point cloud feature map and obtain the projection coordinates of the non-empty query point on the aligned visible light image; with the projection coordinates as the center and according to the preset pixel window size or the search radius determined based on the depth and curvature of the non-empty query point, delineate the local neighborhood corresponding to the query point as the visible light local region. Within the visible light local region, local signal-to-noise ratio and texture gradient information are obtained from the aligned visible light image as state information; At the same local neighborhood location, temperature gradient and thermal radiation contrast information are extracted from the infrared thermal feature map as first response information, and depth variance and local curvature information are extracted from the lidar point cloud feature map as second response information. The state information, the first response information, and the second response information are input into a state-aware network or a preset logical judgment rule to generate a local availability label for the local area. The local availability label is marked as a valid area label, a low-confidence area label, or a distorted area label.

3. The multimodal image semantic segmentation method for video analysis according to claim 2, characterized in that, The step of generating local-level modal control coefficients based on the local availability label includes: For each pixel location on the visible light feature map, read the local availability label corresponding to that pixel location; When the read local availability tag is a valid area tag, the visible light mode control coefficient of the pixel location is set as the first weight, and the infrared thermal mode control coefficient and the lidar mode control coefficient of the pixel location are set as auxiliary weights; the first weight is greater than the auxiliary weight. When the read local availability tag is a low-confidence area tag, the visible light mode control coefficient of the pixel location is reduced to the second weight, and the infrared thermal mode control coefficient and lidar mode control coefficient of the pixel location are increased to the equivalent weight; the second weight is less than the first weight, and the equivalent weight is greater than the auxiliary weight; When the read local availability tag is a distorted area tag, the visible light mode control coefficient of the pixel position is attenuated to a third weight value close to zero, and the infrared thermal mode control coefficient and lidar mode control coefficient of the pixel position are increased to the dominant weight value; the third weight value is less than the second weight value, and the dominant weight value is greater than the equivalent weight value; Output the modal control coefficient at that pixel location.

4. The multimodal image semantic segmentation method for video analysis according to claim 2, characterized in that, The step of performing constrained attention matching on the query vector corresponding to the lidar point cloud feature map and the key vector corresponding to the visible light feature map and the infrared thermal feature map based on the modal control coefficients to obtain the constrained attention weight distribution includes: Obtain the depth value and local curvature information of the query vector; Based on the depth value and local curvature information of the query vector, the two-dimensional search radius is calculated, and the initial candidate matching range of the query vector is determined accordingly. Within the initial candidate matching range, the local availability label is read pixel by pixel, wherein the pixel position marked as a distortion region is included in the noise rejection range, and the pixel position marked as a valid region and the pixel position marked as a low confidence region are retained as the corrected candidate matching range. Based on the corrected candidate matching range and the modality control coefficient, constrained attention matching is performed on the query vector and the key vector to obtain the constrained attention weight distribution.

5. A multimodal image semantic segmentation method for video analysis according to claim 4, characterized in that, The step of performing constrained attention matching on the query vector and the key vector based on the corrected candidate matching range and the modality control coefficient to obtain the constrained attention weight distribution includes: Calculate the initial similarity between the query vector and each key vector within the corrected candidate matching range; Based on the Euclidean distance between the position of each key vector within the corrected candidate matching range and the projection center of the query vector, a spatial distance penalty term is applied to the initial similarity; Based on the modal control coefficients at the locations of each key vector, the similarity after spatial distance penalty is weighted and modulated, wherein the similarity of key vectors located within the noise rejection range is forcibly assigned a preset minimum negative value. The modulated similarity is normalized to generate the constrained attention weight distribution.

6. The multimodal image semantic segmentation method for video analysis according to claim 1, characterized in that, The step of obtaining the multimodal boundary consistency information of the initial fused feature map near the target edge includes: Obtain the depth boundary map corresponding to the aligned lidar point cloud, the thermal radiation boundary map corresponding to the aligned infrared thermal image, and the semantic boundary response map of the initial fused feature map; Morphological dilation operations are performed on the depth boundary map and the thermal radiation boundary map respectively to obtain the dilated depth boundary map and the dilated thermal radiation boundary map. Calculate the spatial intersection of the expansion depth boundary map and the expansion thermal radiation boundary map to obtain the boundary intersection region; Based on the spatial distance between depth boundary pixels and thermal radiation boundary pixels within the boundary intersection region, identify the physical boundary conflict zone within the boundary intersection region; For the physical boundary conflict zone, the local connectivity length and depth gradient magnitude of the depth boundary map are evaluated to obtain the depth state; at the same time, the temperature gradient direction symmetry and attenuation width of the thermal radiation boundary map are evaluated to obtain the thermal state. Based on the depth state and the thermal state, within the physical boundary conflict zone, one of the following operations may be selectively performed: retaining the blocking properties of the depth boundary map, performing morphological shrinkage on the thermal radiation boundary map or generating a smooth blocking line based on the peak line of the semantic boundary response map, the shrunken thermal radiation boundary and the discontinuous depth boundary. Spatially concatenate the boundary outside the physical boundary conflict zone with the boundary generated inside the physical boundary conflict zone to obtain a composite conduction blocking mask, which is the multimodal boundary consistency information.

7. A multimodal image semantic segmentation method for video analysis according to claim 6, characterized in that, The step of correcting the boundaries of the initial fused feature map based on the multimodal boundary consistency information to obtain the corrected fused feature map includes: Obtain motion information of the target region corresponding to the composite conduction blocking mask; Obtain local deformation information of the initial fused feature map; Based on the motion information and the local deformation information, the composite conduction blocking mask is adjusted to obtain the adjusted composite conduction blocking mask. Using anisotropic diffusion filtering, with the adjusted composite conduction blocking mask as the conduction coefficient control term, the initial fused feature map is iteratively diffused to obtain the corrected fused feature map.

8. A multimodal image semantic segmentation method for video analysis according to claim 7, characterized in that, The step of obtaining the motion information of the target region corresponding to the composite conduction blocking mask includes: Identify multiple moving target individuals within the composite conduction blocking mask area; For each of the moving target individuals, calculate the motion vector of the moving target individual; Determine whether there is a spatial occlusion relationship between the moving target individuals; When there is a spatial occlusion relationship between the moving target individuals, the background motion information or the motion information of the occluded target within the occlusion area is suppressed, and the motion vector of the occluded target individual is inferred. By integrating the motion vectors of the individual moving targets, motion information of the target region corresponding to the composite conduction blocking mask is generated.

9. A multimodal image semantic segmentation method for video analysis according to claim 7, characterized in that, The step of generating a semantic segmentation mask map based on the modified fusion feature map includes: Obtain the feature gradient distribution and local connectivity of the local edge regions of the modified fused feature map; The fine contour information of the corresponding local edge region is obtained from the depth boundary map of the aligned lidar point cloud and the thermal radiation boundary map of the aligned infrared thermal image. Based on the feature gradient distribution, the local connectivity, the depth boundary map, and the refined contour information, it is determined whether there is residual blur or discontinuity in the local edge regions of the corrected fused feature map; If residual blur or discontinuity exists, guided filtering or morphological reconstruction is performed on the local edge regions of the corrected fused feature map based on the depth boundary map and the refined contour information. The processed modified fusion feature map is input into the semantic segmentation head to generate a semantic segmentation mask map.

10. A multimodal image semantic segmentation system for video analysis, used to perform the method according to any one of claims 1 to 9, characterized in that, The system includes: The feature extraction module is used to perform spatiotemporal alignment and feature extraction on the visible light image, infrared thermal image and lidar point cloud of the target acquired simultaneously, to obtain a visible light feature map that represents the visible light texture and semantics, an infrared thermal feature map that represents the thermal radiation contour and saliency, and a lidar point cloud feature map that represents the lidar geometric contour and depth information. The status response acquisition module is used to acquire the status information of the aligned visible light image in the corresponding local area for each non-empty query point in the lidar point cloud feature map, as well as the response information of the infrared thermal feature map and the lidar point cloud feature map in the same area, and generate a local availability label based on the status information and the response information. The constrained attention matching module is used to generate local-level modal control coefficients based on the local availability label, and to perform constrained attention matching on the query vector corresponding to the lidar point cloud feature map and the key vectors corresponding to the visible light feature map and the infrared thermal feature map according to the modal control coefficients, so as to obtain a constrained attention weight distribution. The fusion feature map generation module is used to utilize a cross-modal attention fusion algorithm to perform weighted fusion of the visible light feature map, the infrared thermal feature map, and the lidar point cloud feature map based on the constrained attention weight distribution, and generate an initial fusion feature map. The boundary correction module is used to obtain multimodal boundary consistency information of the initial fused feature map near the target edge, and to correct the boundary of the initial fused feature map to obtain a corrected fused feature map. The multimodal boundary consistency information is used to characterize the spatial offset and matching consistency among the fused feature semantic boundary, the lidar physical depth boundary, and the infrared thermal boundary. The mask image generation module is used to generate a semantic segmentation mask image based on the modified fusion feature map.