A target segmentation method based on multi-source cross attention fusion

CN119131396BActive Publication Date: 2026-08-07NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2024-09-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明的目的是在于针对现有技术受单一数据源的信息局限等因素导致目标分割精度低、泛化能力的不足,提供一种基于多源交叉注意力融合的目标分割方法,用于在不同领域下结合多源数据源的目标分割,提高分割精度,以解决背景技术中提出的问题

Benefits of technology

[0026]本发明解决了单一数据源的信息局限的问题,利用多源数据信息融合可以自动、准确地将目标进行精确分类,并且将该区域准确地分割出来,用于在不同领域下结合多源数据源的目标分割,提高分割精度,可供遥感领域、医学领域等各个领域参考与服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131396B_ABST
    Figure CN119131396B_ABST
Patent Text Reader

Abstract

The application discloses a target segmentation method based on multi-source cross attention fusion, and particularly relates to the technical field of image processing and target segmentation, and comprises the following steps: acquiring multi-source image data, and performing corresponding preprocessing operations on the data set; constructing a segmentation network based on multi-source image data fusion; introducing a double-branch encoder, a self-built CAFFM (cross attention feature fusion module) and a feature decoder CU2C module to construct a segmentation network based on multi-source cross attention fusion, which is used for training and prediction of multi-source data target segmentation in different fields; feeding the data and corresponding labels in the training set into the segmentation network, so as to train the segmentation network for the multi-source data target segmentation task; iteratively updating parameters by calculating the loss function of the segmentation network; and stopping the training when the accuracy of the network test set no longer improves after a certain number of training rounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image segmentation and target segmentation technology, and specifically to a target segmentation method based on multi-source cross-attention fusion. Background Technology

[0002] Image segmentation is a crucial area of ​​artificial intelligence. Analyzing large amounts of image data enables the construction of more suitable segmentation network models, thereby automatically identifying and classifying specific targets in different domains, improving the accuracy and efficiency of target segmentation. In image segmentation tasks across various domains, a single information source often cannot provide sufficient detail, especially in complex backgrounds or with blurred boundaries. For example, in the retinal arteriovenous vessel segmentation task, which involves the accurate differentiation and segmentation of arteries and veins in retinal images, there is a certain challenge because arteries and veins often exhibit similar structural features in images. Furthermore, due to complex backgrounds, intersecting vessels, and limitations in image resolution, single-modality image data often struggles to yield accurate segmentation results. Similarly, in remote sensing image land cover classification tasks, which involve the accurate differentiation and segmentation of different land covers in remote sensing images, there is also a challenge. Due to the complexity and diversity of a single SAR image, traditional land cover classification methods based on manual rules and mathematical models are often limited by the limitations of feature selection and classification algorithms, making it difficult to handle complex land cover conditions and extract detailed features. Although SAR images have the ability to overcome complex environments, they contain limited information, making it difficult to obtain accurate land cover classification results. Therefore, designing a general and effective segmentation method is crucial for multiple domains.

[0003] Due to the complexity and diversity of image segmentation tasks across different domains, single-source segmentation methods lack consideration for multifaceted image features, thus affecting segmentation accuracy. For example, in the field of medical image segmentation, in 2022, Hu et al. proposed a multi-scale interactive network with an A / V discriminator for retinal artery and vein identification. This model achieved high accuracy on three publicly available fundus image datasets (AV-DRIVE, HRF, and LES-AV). However, some small blood vessels remained difficult to segment. In the field of remote sensing images, in 2021, Huang et al. proposed a novel unsupervised learning method to learn discriminative physical perception features of synthetic aperture radar (SAR) images using deep neural networks. They designed a convolutional neural network to extract high-level features from detected SAR images. However, while SAR images have the ability to overcome complex environments, they contain limited information, while optical images have more detail but are less feasible for complex environments. Therefore, with the development of deep learning and multimodal data fusion technologies, multi-source cross-attention fusion methods have gradually become a research hotspot in the field of retinal vessel arteriovenous segmentation. By combining multi-source data from different imaging modalities, these methods can capture richer vascular feature information, thus significantly improving segmentation accuracy. For example, color fundus images in retinal imaging and optical coherence tomography (OCTA) angiography each have unique advantages. Color fundus images can provide color and morphological information of blood vessels, while OCTA can display the three-dimensional structure and depth information of blood vessels. However, information from these modalities alone may not be sufficient for the accurate differentiation and segmentation of arteries and veins. In the field of remote sensing image segmentation, the fusion of multi-source data has also shown great potential. Traditional remote sensing image segmentation methods are usually based on single optical images or SAR data, making it difficult to simultaneously consider image details and environmental adaptability. To overcome these challenges, segmentation methods based on multi-source fusion of optical and SAR images can be adopted. Therefore, designing a segmentation network based on multi-source cross-attention fusion, combining information sources from different modalities in different domains, can obtain richer contextual information and improve the accuracy of target segmentation.

[0004] In addition to the above, target segmentation methods based on multi-source cross-attention fusion face three major challenges: First, accuracy is crucial, requiring the method to accurately identify different types of targets across various domains, overcoming data source differences and fusion difficulties. Second, real-time performance is critical, demanding rapid processing and analysis of large amounts of image data to provide timely segmentation results. Finally, versatility is paramount, requiring the system to possess strong generalization capabilities, adapting to image sources across different domains and handling various data formats and resolutions. Therefore, designing a target segmentation method based on multi-source cross-attention fusion is of significant practical importance in overcoming these challenges. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies, such as low target segmentation accuracy and insufficient generalization ability due to limitations in information from a single data source. This invention provides a target segmentation method based on multi-source cross-attention fusion, which combines multiple data sources in different domains to improve segmentation accuracy and solve the problems mentioned in the background.

[0006] In various fields, the multi-source cross-attention fusion target segmentation method can improve the target segmentation effect by integrating multi-source data in the field, making full use of the advantages of different data sources, and extracting richer feature information. It has wide application value in various fields such as smart healthcare, smart remote sensing, and land planning.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a target segmentation method based on multi-source cross-attention fusion, comprising:

[0008] Acquire multi-source image data and perform corresponding preprocessing operations on the dataset;

[0009] A segmentation network based on multi-source image data fusion is constructed by introducing a dual-branch encoder, a self-built CAFFM (Cross-Attention Feature Fusion Module), and a feature decoder CU2C module. This network is used for training and prediction of multi-source data target segmentation in different domains.

[0010] The data and corresponding labels in the training set are fed into the segmentation network to train the segmentation network for multi-source data target segmentation tasks. The parameters are updated iteratively by calculating the loss function of the segmentation network. After a certain number of training rounds, the accuracy of the network on the test set no longer improves, and training is stopped.

[0011] Select the optimal model weights and test the training effect of the model in the multi-source test set;

[0012] The test model was evaluated on its performance in retinal arteriovenous segmentation in the medical field and on land cover classification in the remote sensing field.

[0013] Preferred, such as Figure 1As shown, the preprocessing operations are as follows: Before training the segmentation model for a general segmentation task, image preprocessing is first performed on the multi-source data. The image preprocessing differs for different multi-source datasets in different fields. For example, in the multi-source land cover classification task in the field of remote sensing, the image preprocessing first reads the single-channel SAR image data. Considering that single-channel SAR images only have phase and amplitude features, but in reality, the phase of a single-channel SAR system does not provide useful information, while amplitude (or intensity) is the only useful information, the amplitude feature of the SAR image is extracted using the absolute value function. The optical image data is then read and converted into a grayscale image. An edge detection factor is used to extract the outline of land cover in the optical image from the grayscale image, and the edge information is enhanced onto the original optical image, thereby improving the distinguishability between different land cover and facilitating feature learning. After image preprocessing, the dataset is divided into training set, validation set, and test set in a ratio of 8:1:1. Finally, random horizontal rotation and random vertical flipping are used for data augmentation.

[0014] Preferred, such as Figure 1 As shown, in the retinal arteriovenous segmentation task in the medical field, image preprocessing takes into account that the high contrast of blood vessels in optical fundus images helps to identify the coarse shape of blood vessels, while the detailed information in OCTA images helps to accurately distinguish arteries and veins. Therefore, the optical fundus images are first grayscaled and CLAHE contrast enhanced, and the OCTA images are denoised and data augmented. The amount of data in the dataset is increased by randomly cropping the optical and corresponding OCTA fundus images to improve the generalization ability of the model. After image preprocessing, the dataset is divided into training set, validation set and test set in a ratio of 8:1:1. Finally, random horizontal rotation and random vertical flipping are used for data augmentation.

[0015] Preferably, the segmentation network based on multi-source image data fusion includes a dual-branch encoder module (Encoder), a self-built CAFFM module (Cross-Attention Feature Fusion Module), and a CU2C feature decoder module (Decoder). Firstly, considering that single-source target segmentation methods lack consideration of many detailed image features, thus affecting classification accuracy, a dual-branch encoder is used to extract features from multi-source image data of the same target, such as... Figure 2 As shown in the Encoder;

[0016] The dual-branch encoder module uses ResNet-50 as the backbone network, and selects shallower and deeper feature layers from it respectively. For low-level features, the output of the corresponding shallow Stage 2 layer of the network is selected, and the low-level features of data source D1 and data source D2 are labeled as follows. and For high-level features, the output of the deepest layer, Stage 4, of the corresponding network was selected and labeled as the high-level features of data source D1 and data source D2, respectively. and

[0017] To address the complex background issues of optical images, pre-trained ResNet50 weights, trained on the ImageNet benchmark dataset, are loaded into the optical feature extraction branch of the encoder. This leverages previously learned feature representations, accelerating the model's learning of feature information in optical images and the feature convergence process. Since non-optical images lack optical features, the pre-trained weights are not processed. To fully utilize the salient and complementary features of images D1 and D2 for more accurate land cover classification, a self-constructed CAFFM module is adopted, as shown in the following diagram. Figure 2 As shown in CAFFM, the CAFFM module converts the low-level feature maps of data source D1 and data source D2 output from stage 2. and High-level feature maps of data sources D1 and D2 output by stage4 and CAFFM feature fusion is performed separately. The CAFFM module is an improvement on the attention fusion model.

[0018] The specific improvements are as follows:

[0019] First, perform a similarity operation on the Q features of data source D1 and the K features of data source D2, that is, compare the similarity between each element in the features of D1 and the corresponding element in the features of D2, and determine the output of the V features of D2 accordingly.

[0020] Simultaneously, similarity operations are performed on the Q features of D2 and the K features of D1, that is, comparing the similarity between each element in the D2 features and the corresponding element in the D1 features, and determining the output of the V features of D1 accordingly; the two features are then fused using a Hadamard product to obtain the final output. and In short, the cross-modal attention fusion module quantifies the similarity of features between different modalities and then adjusts the weights of features based on these similarities to achieve cross-modal information interaction and fusion.

[0021] at last, By comparing with low-level feature maps D1 and D2 and The low-level fused feature maps D1 and D2 are obtained by concatenation and convolution. By comparing with D1 and D2 high-level feature maps and The high-level fusion feature maps D1 and D2 are obtained by splicing, ASPP, and upsampling. Finally, to ensure that the final segmentation output better displays the learned features, a CU2C decoder module was adopted. This module consists of a concatenation layer, an upsampling layer, and two convolutional layers. First, the low-level feature maps D1 and D2 obtained from the Encoder are fused. High-level fusion feature maps of D1 and D2 Channel concatenation is performed, and then the feature map size is restored to the original image size through an upsampling layer. Finally, the feature map is convolved twice to obtain the output image of target segmentation.

[0022] Preferably, the segmentation network achieves optimality through a combined loss function of CrossEntropyLoss and Di ceLoss, which is as follows:

[0023] Total Loss=α×CrossEntropyLoss+β×DiceLoss;

[0024] In the formula, CrossEntropyLoss is the cross-entropy loss, DiceLoss is the Dice loss, and α and β are weight parameters used to balance the relative importance of CrossEntropyLoss and DiceLoss in the total loss, that is, to balance classification accuracy and overall segmentation consistency. Furthermore, before model training, the initial weight of each target category is set according to the proportion of the target category in the data, and the category with the lower proportion is given higher attention.

[0025] The present invention has the following advantages:

[0026] This invention solves the problem of information limitations from a single data source. By using multi-source data fusion, targets can be automatically and accurately classified and the region can be accurately segmented. This invention can be used for target segmentation in different fields by combining multiple data sources, thereby improving segmentation accuracy. It can be used as a reference and service in various fields such as remote sensing and medicine. Attached Figure Description

[0027] Figure 1 A diagram illustrating the data preprocessing process provided for this invention;

[0028] Figure 2 A diagram illustrating the segmentation network framework based on multi-source cross-attention fusion provided by this invention;

[0029] Figure 3 This is a diagram of the segmentation network framework based on optical and SAR cross-attention fusion provided in this embodiment;

[0030] Figure 4The diagram shows the optical and SAR data preprocessing process provided for this implementation case.

[0031] Figure 5 This is a segmentation result diagram provided in this embodiment;

[0032] Figure 6 This is a diagram of the segmentation network framework based on light and OCTA cross-attention fusion provided in this embodiment. Detailed Implementation

[0033] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Implementation Case 1: The target segmentation method based on multi-source cross-attention fusion provided by this invention is used in optical and SAR ground cover classification tasks in the field of remote sensing. First, the input optical and SAR image data are preprocessed. Then, the data and labels from the training set are input into the segmentation network for training. The network weights for classifying ground covers in the images are trained. After a certain number of iterations, the average accuracy of the network reaches a fitted level, and the network weights are saved. Optical and SAR image data from the test set are then input to obtain the corresponding ground cover classification result images, completing the ground cover classification of multi-source images. (Refer to...) Figure 3 As shown, the land cover classification method based on optical and SAR cross-attention fusion mainly includes the following steps:

[0035] Step 1: The WHU-OPT-SAR dataset, comprising 100 optical images (OPT) of 5556×3704 pixels each, covering an area of ​​50,000 square kilometers with diverse terrain and vegetation in a province, along with synthetic aperture radar (SAR) images of the same area and corresponding pixel-level annotations for background, farmland, city, village, water body, forest, road, and eight other categories, is used as the dataset for this study. To prevent overfitting and improve generalization performance, the dataset is preprocessed. The preprocessing process is as follows: Figure 4 As shown.

[0036] First, single-channel SAR image data is read. Considering that single-channel SAR images only have phase and amplitude features, but in reality, the phase of a single-channel SAR system does not provide useful information, while amplitude (or intensity) is the only useful information, the amplitude feature of the SAR image is extracted using an absolute value function. Then, optical image data is read and converted into grayscale images. Edge detection factors are used to extract the contours of ground features in the optical images from the grayscale images, and the edge information is enhanced onto the original optical images, thereby improving the distinguishability between different ground features and facilitating feature learning. The dataset is divided into training, validation, and test sets in an 8:1:1 ratio. Finally, random horizontal rotation and random vertical flipping are used for data augmentation.

[0037] Step Two: Model Building. The model of this invention comprises three parts: a dual-branch encoder, a self-built CAFFM (Cross-Attention Feature Fusion Module), and a CU2C (Decoder) feature decoder module. Firstly, considering that single-source land cover classification methods lack consideration of many detailed image features, thus affecting classification accuracy, a dual-branch encoder is used to extract features from optical and SAR images, such as… Figure 3 As shown in the Encoder.

[0038] The Encoder module uses ResNet-50 as the backbone network, selecting both shallow and deep feature layers from it. For low-level features, the output of the corresponding shallow Stage 2 layer of the network is selected, and the low-level features of optical and SAR are labeled as follows. and For high-level features, the output of the deepest Stage 4 layer of the corresponding network was selected and labeled as high-level features for optical and SAR respectively. and To address the complex background issues of optical images, pre-trained ResNet50 weights, trained on the ImageNet benchmark dataset, are loaded into the optical feature extraction branch of the encoder. This leverages previously learned feature representations, accelerating the model's learning of ground feature information in optical images and the convergence process of these features. Since SAR images lack optical features, the pre-trained weights are not processed. To fully utilize the salient and complementary features of both optical and SAR images for more accurate ground feature classification, a self-constructed CAFFM module is adopted, as shown in the following diagram. Figure 3 As shown in CAFFM, this module outputs optical and SAR low-level feature maps from stage 2. and Stage 4 outputs high-level optical and SAR feature maps. and CAFFM feature fusion is performed separately. The CAFFM module is an improved cross-attention feature fusion module based on the attention fusion model. First, a similarity operation is performed on the optical Q features and the SAR K features, that is, the similarity between each element in the optical features and the corresponding element in the SAR features is compared, and the output of the SAR V features is determined accordingly. Simultaneously, a similarity operation is performed on the SAR Q features and the optical K features, that is, the similarity between each element in the SAR features and the corresponding element in the optical features is compared, and the output of the optical V features is determined accordingly. The two features are then fused using a Hadamard product to obtain the final feature. and In short, the cross-modal attention fusion module quantifies the similarity of features between different modalities and then adjusts the feature weights based on these similarities to achieve cross-modal information interaction and fusion. Finally, By combining with optical and SAR low-level feature maps and By stitching and convolution, a low-level fused feature map of optical and SAR elements is obtained. By combining with optical and SAR advanced feature maps and By stitching, ASPP, and upsampling, optical and SAR advanced fusion feature maps are obtained. Finally, to ensure that the final classification result map better displays the learned features, a CU2C decoder module was adopted. This module consists of a concatenation layer, an upsampling layer, and two convolutional layers. First, the low-level fused feature map of optical SAR obtained from the Encoder is processed... Advanced fusion feature map of light SAR Channel stitching is performed, and then the feature map size is restored to the original image size through an upsampling layer. Finally, the feature map is convolved twice to obtain the output image of land cover classification.

[0039] Step 3: Model Training. This invention was performed on a workstation equipped with an NVIDIA IA-RTX4070 GPU, using Python as the programming language and the PyTorch deep learning framework. Adam was used as the optimizer for model training, and a combined loss function of CrossEntropyLoss and DiceLoss was used as the loss function for model training. α and β were set to 1.2 and 0.8, respectively, with a batch size of 16 and a learning rate of 0.001. Before model training, initial weights were set for eight categories—background, farmland, city, village, water body, forest, road, and others—based on their proportion in the data: [0.0, 0.0162, 0.1189, 0.0959, 0.0392, 0.01487, 0.5864, 0.3186]. Categories with lower proportions received higher attention. The optical and SAR images of the training data are input into a segmentation network based on optical and SAR cross-attention fusion for training. After 40 training rounds, the model gradually fits the model. During each iteration of training, the model parameters with the best training effect are saved.

[0040] Step 4: Model Evaluation. To verify the model's performance, it is tested on a test set. The optical and SAR images in the test set have also undergone preprocessing and are fed into the optimal weight network to output the predicted segmentation results. The segmentation results are shown in the image below. Figure 5 As shown.

[0041] Implementation Case 2: The segmentation method based on multi-source cross-attention fusion provided by this invention is applied to retinal arteriovenous segmentation in the medical field. First, the input optical and OCTA image data are preprocessed. Then, the data and labels from the training set are input into the segmentation network for training. The network weights for land cover classification in the images are trained. After a certain number of iterations, the average accuracy of the network reaches a fitted level, and the network weights are saved. Optical and OCTA images from the test set are then input to obtain the corresponding retinal arteriovenous segmentation results, completing the multi-source image arteriovenous segmentation task. (Refer to...) Figure 3 As shown, the retinal arteriovenous vessel segmentation method based on optical and OCTA cross-attention fusion mainly includes the following steps:

[0042] Step 1: 40 optical images of the retina and fundus, along with their corresponding OCTA images, were used as the dataset. To prevent overfitting and improve generalization performance, the dataset was preprocessed. Considering the high contrast of blood vessels in optical fundus images, which helps in recognizing their coarse morphology, and the detailed information in OCTA images, which helps in accurately distinguishing arteries and veins, the optical fundus images were first grayscaled and CLAHE contrast enhanced. The OCTA images underwent denoising and data augmentation. The dataset size was increased by randomly cropping the optical and corresponding OCTA fundus images to improve the model's generalization ability. The dataset was divided into training, validation, and test sets in an 8:1:1 ratio. Finally, random horizontal rotation and random vertical flipping were used for data augmentation.

[0043] Step Two: Model Building. The model of this invention comprises three parts: a dual-branch encoder, a self-built CAFFM (Cross-Attention Feature Fusion Module), and a CU2C (Decoder) feature decoder module. Firstly, considering that single-source target segmentation methods lack consideration of many detailed image features, thus affecting classification accuracy, a dual-branch encoder is used to extract features from optical images and OCTA images, such as… Figure 6 The encoder module uses ResNet-50 as its backbone network, selecting both shallow and deep feature layers. For low-level features, the output of the corresponding shallow Stage 2 layer is selected, and the low-level features of optical and OCTA are labeled as follows. and For high-level features, the output of the deepest layer, Stage 4, of the corresponding network was selected and labeled as the high-level features of optical and OCTA, respectively. and To address the complex background issues in fundus retinal images, which differ in features from ordinary optical images, the pre-trained weights are not processed. To fully utilize the salient and complementary features of optical and OCTA images for more accurate arteriovenous segmentation, a self-constructed CAFFM module is adopted, as shown in the following diagram. Figure 3 As shown in CAFFM, this module outputs the optical and OCTA low-level feature maps from stage 2. and Stage 4 outputs optical and OCTA high-level feature maps. and CAFFM feature fusion is performed separately. The CAFFM module is an improved cross-attention feature fusion module based on the attention fusion model. First, a similarity operation is performed on the Q features of optics and the K features of OCTA, that is, comparing the similarity between each element in the optical features and the corresponding element in the OCTA features, and determining the output of the V features of OCTA accordingly. Simultaneously, a similarity operation is performed on the Q features of OCTA and the K features of optics, that is, comparing the similarity between each element in the OCTA features and the corresponding element in the optical features, and determining the output of the V features of optics accordingly. The two features are then fused using a Hadamard product to obtain the final result. and In short, the cross-modal attention fusion module quantifies the similarity of features between different modalities and then adjusts the feature weights based on these similarities to achieve cross-modal information interaction and fusion. Finally, By combining with optics, OCTA low-level feature maps and By stitching and convolution, a low-level fused feature map of optical and SAR elements is obtained. Through integration with optics, OCTA advanced feature maps and The optical and OCTA advanced fusion feature maps are obtained by stitching, ASPP, and upsampling. Finally, to ensure that the final classification result map better displays the learned features, a CU2C decoder module was adopted. This module consists of a concatenation layer, an upsampling layer, and two convolutional layers. First, the low-level fused feature map of optical OCTA obtained from the Encoder is processed. Advanced Fusion Feature Map of OCTA Channel stitching is performed, and then the feature map size is restored to the original image size through an upsampling layer. Finally, the feature map is convolved twice to obtain the output image of retinal arteriovenous segmentation.

[0044] Step 3: Model Training. This invention was performed on a workstation equipped with an NVIDIA IA-RTX4070 GPU, using Python as the programming language and the PyTorch deep learning framework. Adam was used as the optimizer for model training, and a combined loss function of CrossEntropyLoss and DiceLoss was used as the loss function for model training. α and β were set to 0.8 and 1.2, respectively, with a batch size of 16 and a learning rate of 0.001. Before model training, initial weights were set for the background, arteries, and veins based on their proportion in the data [0.0, 0.6674, 0.5372], with higher attention given to categories with lower proportions. Optical and OCTA images of the training data were input into a segmentation network based on optical and OCTA cross-attention fusion for training. After 75 training iterations, the model gradually fit the desired performance. During each iteration, the parameters of the model with the best training effect were saved.

[0045] Step 4: Model Evaluation. To verify the model's performance, it was tested on a test set. The optical and OCTA images in the test set were also preprocessed and fed into the optimal weight network to output the predicted segmentation results. Its multi-source cross-attention fusion retinal arteriovenous segmentation algorithm outperforms segmentation algorithms based on single optical images in terms of performance and accuracy.

[0046] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A target segmentation method based on multi-source cross-attention fusion, which acquires multi-source image data and performs corresponding preprocessing operations on the dataset, characterized in that: include: A segmentation network based on multi-source image data fusion is constructed. By introducing a dual-branch encoder, a self-built CAFFM module, and a feature decoder CU2C module, a segmentation network based on multi-source cross-attention fusion is built for training and prediction of multi-source data target segmentation in different domains. The data and corresponding labels in the training set are fed into the segmentation network to train the segmentation network for multi-source data target segmentation tasks. The parameters are updated iteratively by calculating the loss function of the segmentation network. After a certain number of training rounds, the accuracy of the network on the test set no longer improves, and training is stopped. The segmentation network based on multi-source image data fusion includes a dual-branch encoder module, a self-built CAFFM module, and a feature decoder CU2C module; the dual-branch encoder is used to extract features from multi-source image data of the same target. The dual-branch encoder module uses ResNet-50 as the backbone network, and selects shallower and deeper feature layers from it respectively. For low-level features, the output of the corresponding shallow Stage 2 layer of the network is selected, and the low-level features of data source D1 and data source D2 are labeled as follows. and For high-level features, the output of the deepest layer, Stage 4, of the corresponding network was selected and labeled as the high-level features of data source D1 and data source D2, respectively. and ; The CAFFM module converts the low-level feature maps from data source D1 and data source D2 output by stage 2. and High-level feature maps of data sources D1 and D2 output by stage4 and CAFFM feature fusion is performed separately. The CAFFM module is an improvement on the attention fusion model. The specific improvements are as follows: First, perform a similarity operation on the Q features of data source D1 and the K features of data source D2, that is, compare the similarity between each element in the features of D1 and the corresponding element in the features of D2, and determine the output of the V features of D2 accordingly. Simultaneously, similarity operations are performed on the Q features of D2 and the K features of D1, that is, comparing the similarity between each element in the D2 features and the corresponding element in the D1 features, and determining the output of the V features of D1 accordingly; the two features are then fused using a Hadamard product to obtain the final output. and ; at last, By comparing with low-level feature maps D1 and D2 and The low-level fused feature maps D1 and D2 are obtained by concatenation and convolution. , By comparing with D1 and D2 high-level feature maps and The high-level fusion feature maps D1 and D2 are obtained by splicing, ASPP, and upsampling. ; The CU2C decoder module is used, which consists of a concatenation layer, an upsampling layer, and two convolutional layers. First, the low-level fused feature maps D1 and D2 obtained from the encoder are processed. High-level fusion feature maps of D1 and D2 Channel concatenation is performed, and then the feature map size is restored to the original image size through an upsampling layer. Finally, the feature map is convolved twice to obtain the output image of target segmentation.

2. The target segmentation method based on multi-source cross-attention fusion according to claim 1, characterized in that: The preprocessing operations are as follows: Before training the segmentation model for the general segmentation task, image preprocessing is first performed on the multi-source data. The amplitude features of the SAR image are extracted using the absolute value function. The optical image data is read and converted into grayscale images. The edge detection factor is used to extract the contours of ground objects in the optical image from the grayscale image, and the edge information is enhanced onto the original optical image. After image preprocessing, the dataset is divided into training set, validation set, and test set in a ratio of 8:1:

1. Finally, random horizontal rotation and random vertical flipping are used for data augmentation.

3. The target segmentation method based on multi-source cross-attention fusion according to claim 1, characterized in that: In the medical field, the retinal arteriovenous segmentation task is first performed by converting the optical fundus images to grayscale and enhancing their contrast with CLAHE. The OCTA images are then denoised and augmented. The dataset size is increased by randomly cropping the optical and corresponding OCTA fundus images. After image preprocessing, the dataset is divided into training, validation, and test sets in an 8:1:1 ratio. Finally, random horizontal rotation and random vertical flipping are used for data augmentation.

4. The target segmentation method based on multi-source cross-attention fusion according to claim 1, characterized in that: The segmentation network achieves its optimality through a combined loss function of CrossEntropyLoss and DiceLoss, which is as follows: ; In the formula, It is cross-entropy loss. yes loss, and These are all weight parameters.

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • Medical image segmentation method based on cross-enhanced self-attention

    CN117495872A