Self-adaptive semantic segmentation method in unsupervised field

Through cross-modal feature extraction, feature alignment and fusion, and unsupervised adaptive training, the semantic segmentation problem of cross-modal data is solved, high-precision semantic segmentation is achieved, and the adaptability and stability of the security monitoring system is improved.

CN120259644APending Publication Date: 2025-07-04SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510190021.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing unsupervised field adaptive semantic segmentation methods are difficult to effectively process cross-modal data, especially visible light and infrared images, resulting in target recognition errors or omissions, which cannot meet the actual needs of security monitoring.

Method used

The method of cross-modal feature extraction, feature alignment based on generative adversarial networks, multimodal feature fusion and unsupervised field adaptive training is adopted, and the precise extraction, alignment and fusion of different modal features is achieved by introducing technical means such as SENet and Spat i a l TransformerNetworks, dual-path generative adversarial networks, dynamic weighted fusion and U-Net architecture improvement.

Benefits of technology

The semantic segmentation accuracy and adaptability of cross-modal data is significantly improved, the accuracy of extracting color features of people in visible light images is improved to 85%, the completeness of extracting thermal radiation features of vehicles in infrared images is increased to 80%, and the accuracy of segmentation in park scenes is increased from 65% to 85%, reducing the cost of data labeling and ensuring the stable operation of the security monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259644A_ABST
    Figure CN120259644A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised field adaptive semantic segmentation method, and relates to the technical field of semantic segmentation, and the method comprises the following steps: cross-modal feature extraction, generative adversarial network-based feature alignment, multi-modal feature fusion, and unsupervised field adaptive training: carrying out unsupervised field adaptive training by using the difference between a source domain and a target domain; according to the method, innovation is made in the aspects of feature extraction, alignment, fusion and unsupervised training for cross-modal data of visible light and infrared images. The feature extraction accuracy of the visible light image is improved from 70% to 85%, and the feature extraction accuracy of the infrared light image is improved from 60% to 80%; the modal feature cosine similarity is increased from 0.4 to 0.75, and the object recognition accuracy is improved from 40% to 70%; the boundary segmentation accuracy is improved from 65% to 85%; the average intersection-to-parallel ratio of a target domain is increased from 0.5 to 0.65, the labeling cost is reduced, the model adaptability is enhanced, and stable operation of a security monitoring system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic segmentation, and specifically provides an unsupervised domain adaptation semantic segmentation method. Background Art

[0002] In the technical field of semantic segmentation, unsupervised domain adaptation semantic segmentation methods are of great significance for improving segmentation accuracy and adaptability. In many practical application scenarios, such as high-end intelligent security monitoring, it is often necessary to process data of different modalities. Typically, visible light images and infrared images are simultaneously used for scene segmentation.

[0003] Currently, most unsupervised domain adaptation semantic segmentation methods are designed for single-modal data. When applied to cross-modal data, due to different feature representations and distributions of different-modal data, many challenges will be faced. For example, visible light images mainly focus on visual features such as color, texture, and shape, and the rich color information does not exist in infrared images; while infrared images focus on the intensity and distribution of thermal radiation, and its thermal radiation information cannot be reflected in visible light images.

[0004] Taking the security monitoring scenario as an example, to ensure the safety and order of places such as parks, it is necessary to accurately perform semantic segmentation on various targets such as people, vehicles, and buildings, which requires the simultaneous use of information from visible light and infrared images. However, existing unsupervised domain adaptation methods are difficult to directly process this cross-modal data and difficult to fully consider the differences between modalities, which leads to frequent target recognition errors or omissions in the monitoring screen and cannot meet the actual needs of security monitoring.

[0005] Therefore, how to effectively fuse and adapt the features of different-modal data in an unsupervised manner to achieve accurate semantic segmentation has become an urgent and unconventional problem to be solved. Solving this problem is of great necessity for improving the performance of semantic segmentation in cross-modal data processing and meeting the requirements of actual application scenarios such as security monitoring.

[0006] In view of this, an unsupervised domain adaptation semantic segmentation method is provided to overcome the above problems. Summary of the Invention

[0007] The purpose of the present invention is to provide an unsupervised domain adaptation semantic segmentation method to solve the problems raised in the above background art.

[0008] To solve the above technical problems, an unsupervised domain adaptation semantic segmentation method provided by the present invention includes the following steps:

[0009] Cross-modal Feature Extraction:

[0010] For visible light images, based on the ResNet network pre-trained on a large-scale natural image dataset, an attention mechanism module combining the channel attention mechanism of SENet and the spatial attention mechanism of Spatial Transformer Networks is introduced. The channel feature description vector is obtained through global average pooling:

[0011]

[0012] where H and W are the height and width of the feature map respectively, c is the number of channels, and then the channel attention weights are obtained through two fully connected layers, and the feature map F res obtained through the ResNet network is weighted to obtain F se , in the Spatial Transformer Networks spatial attention mechanism, F se is transformed through the affine transformation parameters to obtain the final feature;

[0013] For infrared images, a lightweight convolutional neural network based on residual connection and multi-scale convolutional kernel fusion is constructed. 3x3, 5x5, and 7x7 multi-scale convolutional kernels are used to capture different-scale thermal radiation features. The different-scale feature maps F 3x3 、F 5x5 、F 7x7 obtained through the multi-scale convolutional kernels are concatenated by channels to obtain F concat , in the residual connection, assuming the input of a certain layer of the network is x, and the output y is obtained through operations such as convolution, then the final output:

[0014] F residual = x + y;

[0015] Feature alignment based on generative adversarial network: Construct a dual-path generative adversarial network, which includes two independent generators and a shared discriminator. One generator converts the visible light image features into a representation close to the infrared image feature distribution, and the other generator performs the reverse conversion. During training, an auxiliary loss function combining cosine similarity and Euclidean distance is introduced:

[0016]

[0017] where λ1 and λ2 are weight coefficients, and an alternating training strategy is adopted;

[0018] Multi-modal feature fusion: Adopt a dynamic weighted fusion method, and dynamically adjust the weights according to the feature channel variance and semantic relevance. Assume that the visible light image features and infrared image features after feature alignment are Calculate the feature channel variance var vis (c), var ir(c), calculate the semantic correlation score through a pre-trained semantic embedding model, and the fusion weight:

[0019]

[0020] w ir w(c) = 1 - w vis (c),

[0021] The fused feature:

[0022]

[0023] Input the fused feature into a semantic segmentation network improved based on the U-Net architecture, and introduce a multi-scale feature pyramid module and an adaptive context aggregation module;

[0024] Unsupervised domain adaptation training: Utilize the differences between the source domain and the target domain for unsupervised domain adaptation training. Adopt a method combining contrast learning and clustering to construct contrast sample pairs. For the source domain feature sample s i , divide the target domain feature samples into different clusters C1, C2, …, C through the DBSCAN clustering algorithm n , calculate the distance d(s i , center(C j )) between the source domain sample and the center of each cluster, and find the closest cluster C k , find the sample in C k with the closest Euclidean distance to s i as the positive sample pair, and randomly select samples from other clusters as the negative sample pair. The contrast loss function:

[0025]

[0026] where τ is the temperature parameter.

[0027] Furthermore, in the cross-modal feature extraction step, when training the attention mechanism module for visible light images, optimize the calculation of channel attention weights by adjusting the number of neurons and activation functions in the fully connected layer.

[0028] Furthermore, when constructing a lightweight convolutional neural network for infrared images, adjust the arrangement order and number of convolutional layers of the multi-scale convolutional kernels to optimize the extraction efficiency and accuracy of different-scale thermal radiation features of infrared images.

[0029] Furthermore, in the feature alignment step based on the generative adversarial network, an improved generator and discriminator network structure is adopted, including but not limited to increasing the number of hidden layers of the generator and improving the discrimination algorithm of the discriminator, which is used to enhance the alignment ability of the dual-path generative adversarial network for different modality features.

[0030] Furthermore, in the multi-modal feature fusion step, transfer learning is used to fine-tune the pre-trained semantic embedding model for adapting to the semantic relevance calculation in the security monitoring scenario.

[0031] Furthermore, in the semantic segmentation network improved based on the U-Net architecture, the fusion method of feature maps with different scales in the multi-scale feature pyramid module is adjusted, including but not limited to weighted fusion or cascaded fusion, which is used to capture the semantic information of targets with different sizes.

[0032] Furthermore, when the adaptive context aggregation module aggregates context information, it adopts an attention mechanism based on position encoding to consider the spatial position information of features.

[0033] Furthermore, in the unsupervised domain adaptation training step, a strategy of dynamically adjusting the temperature parameter is adopted. A larger value is set in the early stage of training to accelerate the convergence speed, and the value is decreased in the later stage of training to improve the discrimination ability of the model.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] Improve the cross-modal feature extraction ability: For visible light images, an attention mechanism module combining the channel attention mechanism of SENet and the spatial attention mechanism of Spatial Transformer Networks is introduced. Through global average pooling, a fully connected layer is used to calculate the channel attention weights and combined with affine transformation, which can more accurately extract effective features such as color, texture, and shape, avoiding information loss caused by modality differences. When identifying the personnel in the park, the accuracy of extracting the clothing color features of the personnel is improved from 70% when ResNet is used alone to 85%. For infrared images, a lightweight convolutional neural network based on residual connection and multi-scale convolutional kernel fusion is constructed. Different scale convolutional kernels are used to capture thermal radiation features at different scales, and the residual connection is combined to solve the problem of gradient disappearance, effectively extracting features such as thermal radiation distribution, edges, and temperature changes in infrared images, improving the infrared image feature extraction ability, and the integrity of extracting vehicle thermal radiation features in infrared images is improved from 60% to 80%, providing richer and more accurate feature information for subsequent feature fusion and semantic segmentation.

[0036] Effectively align the feature distributions of different modalities: A dual-path Generative Adversarial Network (Dual-path GAN) is constructed to achieve bidirectional feature transformation through two independent generators and a shared discriminator. An auxiliary loss function combining cosine similarity and Euclidean distance and an alternating training strategy are introduced to effectively solve the problem of large differences in the feature distributions of different modality data. After 100 iterations of training, the cosine similarity between the two modality features is increased from the initial 0.4 to 0.75, enabling the subsequent semantic segmentation model to better process cross-modal data and significantly improving the accuracy and reliability of semantic segmentation. In the recognition of objects in complex scenarios in the park, the recognition accuracy of objects that were difficult to recognize due to feature differences in the original infrared images is increased from 40% to 70% after feature alignment.

[0037] Achieve efficient fusion of multi-modal features: The dynamic weighted fusion method is adopted to dynamically adjust the weights according to the feature channel variance and semantic relevance. The fused features are input into a semantic segmentation network improved based on the U-Net architecture, and a multi-scale feature pyramid module and an adaptive context aggregation module are introduced to fully exploit the advantages of different modality data and effectively integrate the complementary information in visible light images and infrared images. In the task of segmenting the boundaries between buildings and the surrounding environment in the park, the segmentation accuracy is increased from 65% of the original single-modal data to 85%, enabling more accurate recognition and segmentation of different semantic categories, reducing the occurrence of misjudgment and missed judgment, and significantly improving the semantic segmentation accuracy and adaptability to complex scenarios.

[0038] Improve the adaptability and generalization ability of the model in the target domain: Unsupervised domain adaptation training is carried out using the differences between the source domain and the target domain. A method combining contrast learning and clustering is adopted to construct contrast sample pairs, effectively avoiding the difficulty of annotating a large number of samples on cross-modal data and reducing the data annotation cost and time. Through this training method, the mean Intersection over Union (mIoU) of the model in the target domain is increased from 0.5 of traditional contrast learning to 0.65, enabling the model to better adapt to the specificity of the target domain and improving the adaptability and generalization ability of the model in the target domain. Even if there are large differences between the source domain data and the target domain data, such as different park layouts and building styles, the model can accurately perform semantic segmentation on the scenes in the target domain, ensuring the stable operation of the security monitoring system. Description of the Drawings

[0039] Figure 1 This is the schematic diagram of an unsupervised domain adaptation semantic segmentation method of the present invention. Detailed Embodiments

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Please refer to Figure 1 , the present invention provides a technical solution:

[0042] Refer to Figure 1 As shown, an embodiment of an unsupervised domain adaptation semantic segmentation method:

[0043] I. Specific scenario

[0044] This embodiment is applied to the field of high-end intelligent security monitoring. Taking a large comprehensive park with a complex environment as the background, the park contains high-density building clusters, open squares, underground parking lots, and green landscape areas. In such a scenario, the security monitoring system needs to perform accurate real-time semantic segmentation on various targets such as people, vehicles, buildings, roads, and vegetation to ensure the safety and order of the park.

[0045] The park is equipped with advanced visible light cameras and infrared thermal imaging cameras. When the daylight is sufficient during the day, the visible light camera can clearly present scene details such as the clothing characteristics of people and the exterior color of vehicles with high resolution and rich color information; while in special environments such as at night, low light, or smoke-filled, the infrared camera can clearly distinguish the positions of vehicles and people in scenes such as underground parking lots by capturing the thermal radiation information of objects.

[0046] However, visible light images and infrared images belong to different modal data. Visible light images are mainly based on visual features such as color, texture, and shape, while infrared images focus on the intensity and distribution of thermal radiation. Traditional unsupervised domain adaptation semantic segmentation methods are difficult to fully consider the differences between modalities, resulting in incorrect or missed target recognition in the monitoring screen and unable to meet the security requirements of the park.

[0047] II. Method steps, inference basis, and beneficial effects

[0048] (1) Cross-modal feature extraction

[0049] For visible light images: Based on the ResNet network pre-trained on a large-scale natural image dataset, an attention mechanism module combining the channel attention mechanism of SENet and the spatial attention mechanism of Spatial Transformer Networks is introduced.

[0050] The ResNet network performs excellently in natural image feature extraction. However, when dealing with cross-modal data, it is difficult to focus on key information according to the characteristics of visible light images. The SENet channel attention mechanism can weight the features of different channels by learning the dependence relationships between channels, enhancing the information of key channels;

[0051] The Spatial Transformer Networks spatial attention mechanism can perform adaptive transformation on different spatial positions of the image, focusing on the target area. The combination of the two enables the network to pay more attention to the regional and channel information related to the target semantics.

[0052] Formula: Assume the input visible light image is I vis , and the feature map F is obtained through the ResNet network res . In the SENet channel attention mechanism, the channel feature description vector Z is obtained through global average pooling c ,

[0053]

[0054] where H and W are the height and width of the feature map respectively, and c is the number of channels. Then, the channel attention weight w is obtained through two fully connected layers c , and F res is weighted to obtain F se . In

[0055] the Spatial Transformer Networks spatial attention mechanism, F se is transformed through the affine transformation parameter θ to obtain the final feature F vis . In practical applications, these parameters are continuously optimized through training, so that when identifying the personnel in the park, the human body contour and the color features of key parts can be extracted more accurately. For example, in a large number of experiments, the accuracy of extracting the color features of the personnel's clothing is increased from 70% when ResNet is used alone to 85%.

[0056] Significantly improve the accuracy and effectiveness of visible light image feature extraction, avoid information loss caused by modal differences, and provide richer and more accurate color, texture, shape and other feature information for subsequent processing.

[0057] For infrared images: Construct a lightweight convolutional neural network based on residual connection and multi-scale convolution kernel fusion. Use 3x3, 5x5 and 7x7 multi-scale convolution kernels to capture different scale thermal radiation features, and solve the problem of gradient disappearance caused by the increase of network depth through residual connection.

[0058] Convolution kernels of different scales have different receptive fields. Small-scale convolution kernels (such as 3x3) focus on detailed information, while large-scale convolution kernels (such as 7x7) capture the overall thermal radiation distribution. Fusing multi-scale convolution kernels can comprehensively extract the features of infrared images. Residual connections enable the network to effectively transmit gradients even when deepening, ensuring the training effect.

[0059] Formula: Let the input infrared image be I ir , and different-scale feature maps F 3x3 , F 5x5 , F 7x7 are obtained through multi-scale convolution kernels. They are concatenated by channels to obtain F concat . In the residual connection, assuming the input of a certain layer of the network is x, and the output y is obtained after operations such as convolution, the final output is:

[0060] F residual = x + y.

[0061] In actual training, compared with the network without residual connections, the convergence speed of the network with residual connections is increased by 30%, and the integrity of the extraction of vehicle thermal radiation features in infrared images is increased from 60% to 80%.

[0062] Effectively extract features such as thermal radiation distribution, edges, and temperature changes in infrared images, improve the ability to extract infrared image features, and provide high-quality infrared features for subsequent processing.

[0063] (2) Feature alignment based on generative adversarial network (GAN)

[0064] Construct a dual-path generative adversarial network (Dual-path GAN), which includes two independent generators and a shared discriminator. One generator converts the features of visible light images into a representation close to the distribution of infrared image features, and the other generator performs the reverse conversion. During training, an auxiliary loss function combining cosine similarity and Euclidean distance is introduced, and an alternating training strategy is adopted.

[0065] Traditional single-path generative adversarial networks can only perform one-way feature conversion, while dual-path generative adversarial networks can perform two-way conversion, more comprehensively aligning the features of the two modalities. Introducing an auxiliary loss function based on feature similarity measurement can guide the generator to more accurately achieve feature alignment from the perspective of feature similarity. The alternating training strategy enables the two generators and the discriminator to better play against each other, ensuring that the features of the two modalities are closer in distribution.

[0066] Formula: Let the features of visible light images be F vis , and the features of infrared images be F ir . Generator G1 converts F vis into Generator G2 converts F ir into The discriminator D determines whether the input features are real features or generated features. The traditional adversarial loss function is:

[0067]

[0068] Auxiliary loss function:

[0069]

[0070] where λ1 and λ2 are weight coefficients, which are adjusted to 0.5 and 0.3 through multiple experiments. In the alternating training, when training a generator each time, other network parameters are fixed. After 100 iterations of training, the cosine similarity of the two modal features is improved from the initial 0.4 to 0.75.

[0071] It effectively solves the problem of large differences in the feature distributions of different modal data, enabling the subsequent semantic segmentation model to better process cross-modal data and greatly improving the accuracy and reliability of semantic segmentation. For example, in the target recognition of complex scenarios in the park, the recognition accuracy of objects that were difficult to identify due to feature differences in the original infrared images is increased from 40% to 70% after feature alignment.

[0072] (III) Multi-modal feature fusion

[0073] The dynamic weighted fusion method is adopted, and the weights are dynamically adjusted according to the feature channel variance and semantic relevance. The fused features are input into a semantic segmentation network improved based on the U-Net architecture, introducing a multi-scale feature pyramid module and an adaptive context aggregation module.

[0074] Measuring the importance of features only based on variance is not comprehensive enough. Combining semantic relevance can more accurately judge the importance of features for different semantic categories. The multi-scale feature pyramid module can fuse feature information of different scales and capture the semantics of targets of different sizes; the adaptive context aggregation module adaptively aggregates the context according to the feature semantics and spatial information to better process semantic segmentation in complex scenarios.

[0075] Formula: Let the visible light image feature be F vis and the infrared image feature be F ir After feature alignment, it is Calculate the feature channel variance var vis (c), var ir (c), and calculate the semantic relevance score sim vis-ir (c) through a pre-trained semantic embedding model. Then the fusion weight:

[0076]

[0077] w ir (c) = 1 - w vis (c),

[0078] Fused features:

[0079]

[0080] In the improved U-Net network, the multi-scale feature pyramid module fuses features through convolution and pooling operations at different scales, and the adaptive context aggregation module adaptively aggregates context using the attention mechanism. In the task of segmenting the boundaries between campus buildings and the surrounding environment, the segmentation accuracy has been improved from 65% of the original single-modal data to 85%.

[0081] Fully exploit the advantages of different modal data, effectively integrate complementary information, more accurately identify and segment different semantic categories, reduce misjudgment and missed judgment, and significantly improve the semantic segmentation accuracy and adaptability to complex scenarios.

[0082] (4) Unsupervised domain adaptation training

[0083] Use the differences between the source domain (visible light and infrared image data collected from other similar large campuses) and the target domain (current campus security monitoring data) for unsupervised domain adaptation training, and adopt a method combining contrast learning and clustering to construct contrast sample pairs.

[0084] Traditional contrast learning methods randomly select contrast sample pairs without considering the potential relationships between samples. Combining the clustering method, first cluster the target domain feature samples, and then select positive sample pairs in similar clusters, which can better mine the common features between the source domain and the target domain and the specific features of the target domain, and improve the adaptability and generalization ability of the model.

[0085] Formula: For the source domain feature sample s i , divide the target domain feature samples into different clusters C1, C2, …, C n . Calculate the distances between the source domain samples and the centers of each cluster

[0086] d(s i , center(C j )) to find the closest cluster C k , and find the sample k in C i with the closest Euclidean distance to s as the positive sample pair, and randomly select samples from other clusters as the negative sample pair. Contrast loss function:

[0087]

[0088] where τ is the temperature parameter, which is adjusted to 0.1 through experiments. After training, the mean intersection over union (mIoU) of the model in the target domain has been improved from 0.5 of traditional contrast learning to 0.65.

[0089] It effectively avoids the difficulty of cross-modal data annotation, reduces costs and time, enables the model to better adapt to the target domain specificity, improves the adaptability and generalization ability in the target domain, and ensures the stable operation of the security monitoring system.

[0090] 3. Summary:

[0091] Cross-modal feature extraction: For visible light images, we introduce an attention mechanism module that combines SENet and Spatial TransformerNetworks based on the pre-trained ResNet network, and use formula calculations to focus on key areas and channel information to improve feature extraction accuracy. For example, the accuracy of color feature extraction of people's clothing has increased from 70% to 85%. For infrared images, we build a lightweight network based on residual connections and multi-scale convolution kernel fusion, using the advantages of convolution kernels of different scales and residual connections to ensure gradient transfer, and improve feature extraction completeness. For example, the completeness of vehicle thermal radiation feature extraction has increased from 60% to 80%.

[0092] Feature alignment based on generative adversarial network (GAN): A dual-path generative adversarial network (Dua l-pathGAN) is constructed to achieve bidirectional feature conversion through two independent generators and a shared discriminator, and an auxiliary loss function and alternating training strategy combining cosine similarity and Euclidean distance are introduced. According to the formula calculation, after 100 iterations of training, the cosine similarity of the two modal features increased from 0.4 to 0.75, solving the problem of large differences in feature distribution of different modal data, so that subsequent semantic segmentation models can better handle cross-modal data, such as the accuracy of infrared image object recognition in complex scenes in the park increased from 40% to 70%.

[0093] Multimodal feature fusion: Dynamic weighted fusion method is used to dynamically adjust the weights according to feature channel variance and semantic relevance, and the fused features are input into the improved U-Net network, and a multi-scale feature pyramid module and an adaptive context aggregation module are introduced. More accurate feature fusion is achieved through formula calculation. In the task of segmenting the boundary between campus buildings and the surrounding environment, the segmentation accuracy is increased from 65% of a single modality to 85%, fully tapping the advantages of different modal data and integrating complementary information.

[0094] Unsupervised domain adaptive training: Unsupervised training is performed using the difference between the source domain and the target domain, and a method based on contrastive learning and clustering is used to construct contrast sample pairs. Through formula calculation, the mean intersection over union (mIoU) of the model in the target domain is increased from 0.5 of traditional contrastive learning to 0.65, avoiding the difficulty of cross-modal data labeling and improving the adaptability and generalization ability of the model in the target domain.

[0095] In summary, through the design of four key steps: cross-modal feature extraction, feature alignment, multi-modal feature fusion, and unsupervised domain adaptation training, this method effectively solves the domain adaptation problem faced by cross-modal data in semantic segmentation, achieves more accurate semantic segmentation in high-end intelligent security monitoring scenarios, provides reliable and efficient support for the security system, and has important practical application value and technological leadership.

Claims

1. An unsupervised domain adaptation semantic segmentation method, characterized in that, Including the following steps: Cross-modal feature extraction: For visible light images, based on the ResNet network pre-trained on a large-scale natural image dataset, an attention mechanism module combining the channel attention mechanism of SENet and the spatial attention mechanism of SpatialTransformerNetworks is introduced, and the channel feature description vector is obtained through global average pooling: Where H and W are the height and width of the feature map respectively, c is the number of channels, and then through two fully connected layers to obtain the channel attention weights, and the feature map F obtained through the ResNet network is weighted res to obtain F se , in the Spatial Transformer Networks spatial attention mechanism, F is transformed through the affine transformation parameters se to obtain the final feature; For infrared images, a lightweight convolutional neural network based on residual connection and multi-scale convolutional kernel fusion is constructed. 3x3, 5x5, and 7x7 multi-scale convolutional kernels are used to capture thermal radiation features at different scales, and different-scale feature maps F 3x3 , F 5x5 , F 7x7 are concatenated by channel to obtain F concat . In the residual connection, assuming the input of a certain layer of the network is x and the output y is obtained through operations such as convolution, the final output is: F residual = x + y; Feature alignment based on generative adversarial network: Construct a two-path generative adversarial network, including two independent generators and a shared discriminator. One generator converts the visible light image features into a representation close to the infrared image feature distribution, and the other generator performs the reverse conversion. During training, an auxiliary loss function combining cosine similarity and Euclidean distance is introduced: Where λ1 and λ2 are weight coefficients, and an alternating training strategy is adopted; Multi-modal feature fusion: The dynamic weighted fusion method is adopted to dynamically adjust the weights according to the variance of feature channels and semantic relevance. Suppose the visible light image features and infrared image features are Calculate the variance var of the feature channels vis (c), var ir (c). Calculate the semantic relevance score through a pre-trained semantic embedding model. Then the fusion weight is: w ir (c) = 1 - w vis (c), Fused features: The fused features are input into a semantic segmentation network improved based on the U-Net architecture, and a multi-scale feature pyramid module and an adaptive context aggregation module are introduced; Unsupervised domain adaptation training: Unsupervised domain adaptation training is performed using the difference between the source domain and the target domain. A method based on contrastive learning and clustering is used to construct contrast sample pairs. For the source domain feature sample s i , the target domain feature samples are divided into different clusters C1, C2, ..., C n , calculate the distance d(s) between the source domain sample and the center of each cluster i ,center(C j )), find the closest cluster C k , in C k Find the i Sample with the closest Euclidean distance As positive sample pairs, randomly select samples from other clusters As a negative sample pair, contrast loss function: Where τ is the temperature parameter.

2. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: In the cross-modal feature extraction step, when training the attention mechanism module for visible light images, the calculation of the channel attention weights is optimized by adjusting the number of neurons and activation functions in the fully connected layer.

3. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: When constructing a lightweight convolutional neural network for infrared images, the arrangement order of multi-scale convolutional kernels and the number of convolutional layers are adjusted to optimize the extraction efficiency and accuracy of different-scale thermal radiation features of infrared images.

4. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: In the feature alignment step based on the generative adversarial network, an improved generator and discriminator network structure are adopted, including but not limited to increasing the number of hidden layers of the generator and improving the discrimination algorithm of the discriminator, to enhance the alignment ability of the two-path generative adversarial network for different-modal features.

5. The unsupervised domain adaptation semantic segmentation method according to claim 1, wherein: In the multi-modal feature fusion step, transfer learning is used to fine-tune the pre-trained semantic embedding model to adapt to the semantic relevance calculation in the security monitoring scenario.

6. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: In the semantic segmentation network improved based on the U-Net architecture, the fusion method of different-scale feature maps in the multi-scale feature pyramid module is adjusted, including but not limited to weighted fusion or cascade fusion, to capture the semantic information of targets of different sizes.

7. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: When the adaptive context aggregation module aggregates context information, an attention mechanism based on position encoding is adopted to consider the spatial position information of the features.

8. The unsupervised domain adaptation semantic segmentation method according to claim 1, characterized in that: In the unsupervised domain adaptation training step, a strategy of dynamically adjusting the temperature parameter is adopted. A larger value is set in the early stage of training to accelerate the convergence speed, and the value is reduced in the later stage of training to improve the discrimination ability of the model.