Coronary artery angiography image segmentation method based on double-flow collaborative network

By building a dual-flow collaborative network, combining the main segmentation network and edge segmentation network, the problems of noise interference and scale differences in coronary angiography image segmentation are solved, and high-precision and robust vascular segmentation are achieved, which improves the accuracy and efficiency of coronary heart disease diagnosis.

CN120298419APending Publication Date: 2025-07-11HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276960.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing coronary angiography image segmentation method is difficult to achieve accurate segmentation when facing complex scenes, noise interference and vascular structures of different scales, resulting in the impact of diagnostic accuracy, especially in primary medical institutions with prominent diagnostic pressure.

Method used

The method based on dual-stream collaborative network is adopted to build the main segmentation network and edge segmentation network, and the global semantic information is captured through the MSFE module, the NIEE module suppresses noise, and the CAKC module extracts fine-grained edge features, and dynamically fusion features at different learning stages through the AEGF module to achieve high-precision and robust vascular segmentation.

Benefits of technology

It improves the segmentation accuracy and robustness of coronary angiography images, ensures edge clarity and the integrity of vascular topology, and improves the accuracy and efficiency of coronary heart disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298419A_ABST
    Figure CN120298419A_ABST
Patent Text Reader

Abstract

The invention discloses a coronary artery angiography image segmentation method based on a double-flow collaborative network. According to the method, the double-flow collaborative network is constructed based on the idea of keeping different resolutions and introducing auxiliary information. An MSFE module is provided in a main segmentation network branch to capture global semantic information of different levels, and the ability to understand a coronary artery complex structure is enhanced. An NIEE module is proposed in an edge segmentation network branch to suppress noise interference and enhance an edge contour, and then a CAKC module is utilized to extract fine-grained edge information, so that edge detail expression is enhanced. In different learning stages, an AEGF module is used to realize dynamic fusion of global semantic features of the main segmentation network and branch local detail features of the edge segmentation network so as to obtain an optimal learning effect. And finally, performing resolution alignment on all the information streams to obtain a final output result. The method can effectively suppress noise interference, enhance edge contour features, and realize high precision and robustness of blood vessel segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical image segmentation method, and particularly to a coronary angiography image segmentation method based on a dual-stream collaborative network. Background Art

[0002] With factors such as the continuous development of social economy and the rapid transformation of the national lifestyle, the prevalence of cardiovascular diseases in China has been showing a continuous upward trend and ranks first in the death composition ratio of urban and rural residents. Clinically, coronary angiography technology is usually used as the "gold standard" for the diagnosis of coronary heart disease. However, the existing diagnosis mode mainly relies on the subjective visual estimation of doctors for angiography images, which is easily interfered by doctors' personal experience and subjective factors, affecting the diagnostic accuracy. Especially in primary medical institutions, due to the limited number of experienced doctors, the diagnostic pressure is more prominent. Therefore, developing an artificial intelligence-based coronary angiography image analysis technology to assist diagnosis and improve medical efficiency is a technical problem that urgently needs to be solved at present. Among them, the image segmentation technology is the core component, and the quality of its segmentation result will directly determine the correctness of the subsequent diagnosis result.

[0003] With the progress of society, the image segmentation technology has made great progress. Nowadays, more and more researchers have focused on deep learning methods. However, although the traditional encoding-decoding structure is still a popular technology today, it still has certain limitations in dealing with complex scenes, extracting detailed features, and suppressing noise interference. Especially in the field of medical image segmentation, traditional methods often have difficulty achieving accurate segmentation when facing complex blood vessel morphologies and high-noise environments. In addition, most of the existing technologies only focus on blood vessel segmentation under a certain specific dataset, for example, they overly emphasize the feature extraction of small blood vessels or global blood vessels, and fail to take into account blood vessel structures of different scales and actual diagnostic needs. This has prompted researchers to continuously explore more advanced network structures and optimization strategies to further improve the accuracy and robustness of segmentation.

[0004] Due to the complexity of blood vessel structures (such as morphological diversity) and the inevitable loss of key information features during the upsampling and downsampling processes of traditional encoding-decoding networks, both result in missing details and structural breaks in the segmentation results, causing obstacles to subsequent diagnoses. In actual scenarios, images are often accompanied by artifacts and noise interference, which significantly reduce the quality of edge feature extraction. Since the most important feature representation of blood vessels is through edge depiction, this problem greatly affects the practicality of the segmentation results. Existing deep learning methods use fixed receptive field convolutions and the same feature extraction strategies, resulting in limited modeling capabilities for the feature expressions of blood vessels with different scale differences, which often causes imbalances in performance between different scale segmentation tasks. That is, when segmenting small blood vessels, the model tends to ignore their details; when segmenting major blood vessels, it may lead to insufficient global coherence. The occurrence of this problem greatly reduces the diagnostic value of the segmentation results. Summary of the Invention

[0005] Aiming at the problems of large background noise interference, insufficient contrast, and local blurring of blood vessel boundaries in coronary angiography images, as well as the problems of insufficient adaptability and balance of existing segmentation methods under different task conditions, the present invention provides a coronary angiography image segmentation method based on a dual-stream collaborative network. This method can effectively suppress noise interference, enhance edge contour features, and at the same time take into account the extraction and fusion of global semantic information and local details, thereby achieving high-precision and robustness in blood vessel segmentation and solving the problem of segmentation imbalance in complex imaging environments.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] A coronary angiography image segmentation method based on a dual-stream collaborative network, comprising the following steps:

[0008] Step 1. Construct a coronary angiography image dataset, and the specific steps are as follows:

[0009] Step 1-1. Obtain a coronary angiography image dataset;

[0010] Step 1-2. Divide the dataset into a training set, a validation set, and a test set;

[0011] Step 1-3. During the training stage, preprocess the original images used. The preprocessing operations include data augmentation and color adjustment. Data augmentation includes randomly rotating the coronary angiography images input for network training within the range of -10° to 10°, mirror transformation, Gaussian blur, randomly scaling within the range of 0.5 to 2.0, randomly cropping the images, and changing the random aspect ratio within the range of 0.7 to 1.3. Color adjustment includes changing the hue of the coronary angiography images input for network training within the range of -0.1 to 0.1, adjusting the saturation within the range of 0.7 to 1.7, and adjusting the brightness within the range of 0.7 to 1.3;

[0012] Step 2: Construct a dual-stream collaborative network for coronary angiography image segmentation, which specifically includes the following steps:

[0013] Step 2-1: Construct a dual-stream collaborative network including a main segmentation network and an edge segmentation network:

[0014] The input coronary angiography image data is fed into the main segmentation network and the edge segmentation network in parallel. Among them, the edge segmentation network consists of a high-level encoding module M bio and a dynamic receptive field edge module M drf and is composed of. The feature F bio obtained via M bio is expressed as follows:

[0015] F bio = M bio (I)

[0016] The feature F drf obtained via M drf is expressed as follows:

[0017] F drf = M drf (J)

[0018] Among them, F bio is the feature map obtained through M bio , F drf is the feature map obtained through M drf , and J is the feature map fed into M drf ;

[0019] At the same time, the input feature I enters the main segmentation network. The main segmentation network is composed of multiple multi-scale information extraction modules M global,i (i = 1, 2,..., n), where n is the number of multi-scale information extraction modules. The feature output of the main segmentation network is shown as follows:

[0020]

[0021] Finally, the feature maps extracted by the main segmentation network and the edge segmentation network are uniformly processed and fused through a soft self-alignment module based on sub-pixel convolution to obtain the final output result;

[0022] Step 2-2: In the edge segmentation network, design an efficient encoding module (Neuro-Inspired Efficient Encoding Module, NIEE) to suppress image noise and enhance vascular edge features. The specific steps are as follows:

[0023] Step 221: DWT decomposes the image in the horizontal and vertical directions through the low-pass filter L and the high-pass filter H respectively, as shown in the following formula:

[0024]

[0025] Among them, X[i,k] is an element of the original image, and its position is at (i,k), X L [i,j] is the low-frequency part of the signal, X H [i,j] is the high-frequency part of the signal, and H[k-j] and L[k-j] are the weights of the high-pass and low-pass filters respectively;

[0026] Step 222: Use LH to convolve X L and X H again in the row direction and column direction respectively to generate four sub-bands: low-frequency LL, high-frequency horizontal LH, high-frequency vertical HL, and high-frequency diagonal HH, as shown in the following formula:

[0027]

[0028] Among them, L[k-j] and H[k-j] are the low-frequency part and high-frequency part of the signal obtained through wavelet transform respectively; LL[i,j], LH[i,j], HL[i,j], and HH[i,j] are the obtained low-frequency sub-band, high-frequency horizontal sub-band, high-frequency vertical sub-band, and high-frequency diagonal sub-band respectively;

[0029] Step 223: For the decomposed frequency components, design the AEGF module to further simulate the noise selective suppression and key feature highlighting mechanism in the LM area. The AEGF module calculates the importance weight W A through adaptive global average pooling and per-channel convolution, and dynamically amplifies or suppresses each frequency component, and the formula is as follows:

[0030]

[0031] Among them, W1 and W2 are the per-channel convolution weights respectively, H and W are the height and width of the input feature map respectively, X(i,j) represents the pixel value of the feature map at the i-th row and j-th column, and σ is the Sigmoid activation function; the enhanced high-frequency component is as follows:

[0032] X' = W A ⊙X

[0033] Among them, W A is the importance weight, X is the input feature map, ⊙ is the element-wise multiplication, and X' is the resulting feature map;

[0034] Step 224: Apply independent convolutional branches to each high-frequency component LH, HL, and HH to extract local edge features, and sum the LH and HL components element-wise and then further fuse them through convolution to accurately describe the directional morphological features of blood vessels; independently process the HH to capture the blood vessel edge information distributed along the diagonal direction. Finally, combine the fused high-frequency components with the low-frequency component to generate an enhanced feature map. The formula is as follows:

[0035]

[0036] where F low is the low-frequency feature, is the enhanced feature of each high-frequency component;

[0037] Step 23: Design a Cascaded Adaptive Kernel Convolution Module (CAKC) in the edge segmentation network to enhance the recognition ability of the coronary artery edge structure through the dynamic receptive field of cascaded adaptive kernels. The specific steps are as follows:

[0038] Use the sobel operator to extract features in the x and y directions, and then introduce the deformable convolution technology AKConv to design the CAKC module. The cascade structure is defined as:

[0039] Input = x1

[0040] x n = F n-1 (x n-1 ) + x n-1

[0041] Output = x n

[0042] where n represents the input feature map of the nth layer, and F n-1 (x) represents the convolutional layer operation of the (n - 1)th layer;

[0043] AKConv expands the regular sampling grid through a generation algorithm, and the sampling coordinate P n is defined as:

[0044] P n = Generate(R regular ) + Offset

[0045] where Generate(x) is the algorithm for generating the initial grid, and R regularis the initial rule sampling network, and Offset is the cheap amount dynamically adjusted; the sampling grid takes the upper left (0, 0) as the sampling origin, and finally generates a sampling grid of any size and shape. Based on this, the convolution operation corresponding to the position P0 is defined as:

[0046] Conv(P0) = Σω·(P0 + P n )

[0047] where ω represents the convolution kernel parameter, and P n is the coordinate of the sampling point generated in the previous step, and the updated sampling coordinate is:

[0048] P new = P0 + P n

[0049] Step 24. Design a multi-scale information extraction module (Multi-Scale Feature Extraction Module, MSFE) in the main segmentation network to enhance the recognition and segmentation capabilities of coronary artery structures at different scales. At the same time, in order to achieve adaptive spatial feature enhancement, a parallel spatial modulation module (Parallel Spatial Encoding Module, PSE) is designed within the MSFE module to further extract and enhance spatial features. Through the multi-scale feature extraction mechanism, the recognition and segmentation capabilities of different-scale structures of the coronary artery are enhanced. The specific steps are as follows:

[0050] Step 24-1. The MSFE module consists of three convolution channels with parallel inputs, namely the Conv 1×1 channel, the Conv 3×3 channel, and the Conv 5×5 channel. Among them, PSE modules are added to both the Conv 3×3 channel and the Conv 5×5 channel for feature extraction and enhancement. Specifically, first, the input feature map x is generated into an initial feature map through three aggregation methods: the average feature map D, the maximum feature map U, and an additional average feature map F. They are used to retain overall information, highlight significant features, and enrich spatial details respectively, forming diverse feature representations to facilitate the capture of subtle spatial information in the input. Next, D, U, and F are concatenated in the channel dimension to obtain a combined feature, and then mapped to a single feature space K through a convolution operation to complete spatial feature integration to meet the requirements of different receptive fields. After convolution processing, an adaptive spatial attention matrix A is generated through the Sigmoid activation function to assign weights to each position, reflecting the spatial relationship between pixels;

[0051] Step 242: The EMA (Efficient Multi-Scale Attention Module) module is introduced into the MFSE module. The EMA module consists of three core mechanisms: feature grouping, parallel sub-structures, and cross-channel learning mechanisms. The feature grouping mechanism first divides the given feature map along the channel dimension into G sub-feature maps, denoted as X = [X0, X i ,..., X G-1 . These sub-maps are then used by the EMA module to capture multi-scale spatial information through its different receptive field paths and simultaneously perform cross-channel information interaction modeling, where: in the shared 1×1 branch, a pair of direction-aware feature maps are first obtained. The global information embedding of the c-th channel is shown as follows:

[0052]

[0053] where, z h and z w respectively represent the average pooling results of the c-th channel in the height and width directions. X' c,i,j represents the feature value at the position (i, j) in the feature map of the c-th channel. H and W respectively represent the sizes in the height and width directions of the feature map; for the vector obtained after parallel transformation encoding through different angles, continue to perform C1 transformation with a shared 1×1 convolution, then split it in the height dimension h, and finally use the non-linear sigmoid function to fit the two-dimensional binomial distribution on the linear convolution. The formula is as follows:

[0054] w h , w w = σ(Split(C1([z h , z w )), h)

[0055] where, [,] is the concatenation operation along the spatial dimension, C1 is the Concat+Conv(1×1) operation, w h and w w are the convolution weights in the height and width directions respectively, σ is the Sigmoid function, and Split is the splitting operation in the input feature space dimension;

[0056] Step 25: Design an Adaptive Feature Fusion Module (AEGF) to automatically control the weights of the edge information flow and the global information flow. The specific steps are as follows:

[0057] Step 251: Given that the feature representations of the global feature output and the edge detection output are respectively and Divide E and G into N sub-channels respectively, and control the contribution degrees of the global and edge features of each sub-channel through a gating mechanism; use the global feature x c,i,j as the main to obtain the dynamic weight as shown in the following formula:

[0058]

[0059] where x c,i,j is the value of the c-th channel in the input feature map at the position (i, j), and W3 and W4 are linear transformation matrices respectively;

[0060] Step 252: Perform a spatial mapping operation on the bilateral feature maps in different channels. The edge pixel information in the sub-channels is adaptively adjusted according to the importance of the features, so as to achieve selective enhancement or suppression of different spatial positions. The final fused feature map formula is as follows:

[0061]

[0062] where χ fuse is the finally fused feature map, i is the number of sub-channels, and are the edge feature and the global feature respectively;

[0063] Step 3: Input the data-augmented training set and validation set into the main segmentation network and the edge segmentation network to obtain a pre-trained coronary artery segmentation network model;

[0064] Step 4: Input the test set into the pre-trained coronary artery segmentation network model to evaluate the performance of the network.

[0065] Compared with the prior art, the present invention has the following advantages:

[0066] The present invention proposes a dual-stream collaborative network, which coordinates the processing of global semantic information and edge detail information through a mechanism of gradual refinement and independent encoding, thereby achieving high-precision segmentation of coronary artery vessels. The main segmentation network adopts the MSFE module, which captures multi-level semantic information through different-resolution networks and convolutional operations, greatly enhancing the context semantic expression ability and ensuring the network's accurate perception of the global structure. At the same time, the edge segmentation network branch focuses on strengthening the extraction of edge features: firstly, the NIEE module effectively suppresses image noise and enhances the clarity of the edge contour; subsequently, the CAKC module further extracts fine-grained edge features and refines the edge expression. In the entire segmentation process, the edge segmentation network branch and the main segmentation network branch perform feature interaction and integration through the AEGF module at different stages, combining the local details of edge features with global semantic information to ensure that the segmentation result has edge clarity and the integrity of the vascular topological structure. Finally, through the resolution alignment and optimized integration of the feature stream, the high precision and high integrity of the output result are ensured, thereby improving the diagnostic value of vascular segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is a flowchart of an automatic segmentation method for coronary angiography images based on a dual-stream collaborative network;

[0068] Figure 2 is the overall structure diagram of the coronary artery segmentation network proposed by the present invention;

[0069] Figure 3 is the structure diagram of the NIEE module;

[0070] Figure 4 is the structure diagram of the CAKC module;

[0071] Figure 5 is the structure diagram of the MSFE module;

[0072] Figure 6 is the pseudo-code diagram of the AEGF module;

[0073] Figure 7 is the overall implementation flowchart of the automatic segmentation method for coronary angiography images based on a dual-stream collaborative network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The technical solutions of the present invention will be further described below with reference to the accompanying drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.

[0075] The present invention provides a coronary angiography image segmentation method based on a dual-stream collaborative network. First, the dataset images are processed. The private dataset is established by selecting key frames at an appropriate viewpoint in the end-diastolic phase of the heart, while the public dataset is directly used. Secondly, based on the idea of maintaining different resolutions and introducing auxiliary information, the present invention constructs a dual-stream collaborative network. In the main segmentation network branch, the MSFE module is proposed to capture global semantic information at different levels and enhance the understanding ability of the complex structure of the coronary arteries. In the edge segmentation network branch, the NIEE module is proposed to suppress noise interference and enhance the edge contour. Subsequently, the proposed CAKC module is used to extract fine-grained edge information, thereby strengthening the expression of edge details. At different learning stages, the proposed AEGF module is used to achieve the dynamic fusion of the global semantic features of the main segmentation network and the local detail features of the edge segmentation network branch to obtain the best learning effect. Finally, all information flows are aligned in resolution to obtain the final output result. The segmentation result obtained in this way not only has high pixel-level accuracy but also can maintain clear edges and structural coherence in a complex imaging environment, providing high-quality clinical auxiliary support for the diagnosis of coronary heart disease. As Figure 1 shown, the specific steps are as follows:

[0076] A1. Construct a coronary angiography image dataset and send this dataset into the subsequent network for training. The specific steps are as follows:

[0077] B1. Obtain a coronary angiography image dataset.

[0078] In the present invention, two different coronary angiography image datasets are used for verification. One original coronary angiography image dataset is provided by the Department of Cardiology, the Second Affiliated Hospital of Harbin Medical University, and all data are real data. Another coronary angiography dataset comes from the Department of Cardiology, Kazakh National Medical University, and is used as a public dataset for all researchers.

[0079] B2. Label the private dataset and divide it into a training set, a validation set, and a test set, and directly divide the public dataset into a training set, a validation set, and a test set.

[0080] In this step, professional doctors conduct systematic ground-truth labeling on the dataset provided by the Department of Cardiology, the Second Affiliated Hospital of Harbin Medical University. And the two datasets are respectively divided into a training set, a validation set, and a test set.

[0081] B3. Perform preprocessing operations on the original data.

[0082] To enhance the generalization ability of the model, preprocessing operations are first performed on the network input data. The preprocessing operations include data augmentation and color adjustment. The data augmentation methods include: random rotation within the range of -10° to 10°, mirror transformation, Gaussian blur, random scaling with a scaling factor in the range of 0.5 to 2.0, random cropping of the image, and random aspect ratio change within the range of 0.7 to 1.3 for filling; in addition, color adjustment is also performed, including hue change within the range of -0.1 to 0.1, saturation adjustment within the range of 0.7 to 1.7, and brightness adjustment within the range of 0.7 to 1.3.

[0083] A2. Construct a two-stream collaborative network for coronary angiography image segmentation, which includes a main segmentation network and an edge segmentation network. Its structure is as Figure 2 shown. The MSFE module is proposed in the mainstream segmentation network, the NIEE module and the CAKC module are proposed in the edge network, and the AEGF module is proposed to perform the feature fusion task. The sub-pixel convolution module is introduced in the upsampling part, and the specific steps are as follows:

[0084] B4. Construct a two-stream collaborative network including a main segmentation network and an edge segmentation network.

[0085] First, the input coronary angiography image data is fed into the edge segmentation network and the main segmentation network in parallel. The edge segmentation network consists of two modules: the high-level encoding module M bio and the dynamic receptive field edge module M drf . M bio simulates the mechanism of the mammalian visual cortex to capture complex edge information in the input image; M drf enhances the adaptability to irregular edges by dynamically adjusting the size and shape of the receptive field. The feature F bio obtained through M bio can be expressed as follows:

[0086] F bio = M bio (I) (1)

[0087] The feature F drf obtained through M drf can be expressed as follows:

[0088] F drf = M drf (J) (2)

[0089] where F bio is the feature map obtained through M bio , F drf is the feature map obtained through M drf , and J is the input to M drfThe feature map. Meanwhile, the input feature I enters the main segmentation network, which consists of multiple multi-scale information extraction modules M global,i (i = 1, 2, ..., n), where n is the number of multi-scale information extraction modules. By designing convolutional structures of different scales, the accurate extraction of global structural information is ensured, thereby ensuring the integrity of the topological structure of the segmentation result. The feature output of the main segmentation network is shown as follows:

[0090]

[0091] During the process of the input information flowing through each branch of the network respectively, the feature information between different branches will be interacted and fused multiple times to fully explore and utilize the advantages of their respective features. Finally, the feature maps extracted by each branch are uniformly processed and fused through a soft self-alignment module based on sub-pixel convolution, so as to obtain a more consistent final output result.

[0092] B5. Design the NIEE module in combination with the development of frontier neurology to suppress image noise and enhance vascular edge features. The structure of the NIEE module is as Figure 3 shown.

[0093] Neuroscience research shows that contour recognition in the mammalian visual system is closely related to object segmentation, especially in the secondary visual area (LM area) of the mouse visual cortex. This area can identify high-order statistical features in textures through efficient encoding and reduce noise interference irrelevant to the task, enabling the visual system to focus on key edge and shape information. Inspired by this, the present invention designs a biological information-based edge feature enhancement method, called the NIEE module, to effectively distinguish noise from real edge features, thereby significantly improving the accuracy of vascular segmentation. This method helps to accurately extract vascular edges in complex medical images and improve the robustness and accuracy of the segmentation results. The present invention mainly uses wavelet transform technology for multi-frequency decomposition and carefully designs various convolutional layers or weight attention to enhance or weaken features at different frequencies to achieve the ultimate goal of simulating the LM area.

[0094] The DWT decomposes the image in the horizontal and vertical directions through a low-pass filter L and a high-pass filter H respectively, as shown in the following formula:

[0095]

[0096] where X[i, k] is an element of the original image at position (i, k). X L [i, j] is the low-frequency part of the signal, and X H [i, j] is the high-frequency part of the signal. H[k - j] and L[k - j] are the weights of the high-pass and low-pass filters respectively.

[0097] After that, for X L and X H convolution is performed again using LH in the row direction and column direction respectively to generate four sub-bands: low-frequency (LL), high-frequency horizontal (LH), high-frequency vertical (HL), and high-frequency diagonal (HH), as shown in the following formula:

[0098]

[0099] where L[k - j] and H[k - j] are the low-frequency part and high-frequency part of the signal obtained through wavelet transform respectively; LL[i, j], LH[i, j], HL[i, j], and HH[i, j] are the obtained low-frequency sub-band, high-frequency horizontal sub-band, high-frequency vertical sub-band, and high-frequency diagonal sub-band respectively.

[0100] These sub-bands respectively correspond to the frequency components of different scales and directions of the image. Among them, the LL sub-band simulates the ability of the LM region to integrate the low-frequency information of the overall image structure and retains the global background and rough contours; while the LH, HL, and HH sub-bands simulate the multi-directional selective sensitivity of the LM region to texture features and capture the high-frequency texture details in the horizontal, vertical, and diagonal directions. This decomposition process essentially realizes the ability of the LM region to capture complex high-order statistical features through multi-directional and multi-scale analysis in neural coding, while reducing the redundant interference between different frequency features.

[0101] Next, for the decomposed frequency components, the present invention designs an AEGF module to further simulate the noise selective suppression and key feature highlighting mechanisms of the LM region. The AEGF module calculates the importance weight W through adaptive global average pooling and per-channel convolution A , and dynamically amplifies or suppresses each frequency component, and the formula is as follows:

[0102]

[0103] where W1 and W2 are the per-channel convolution weights respectively, H and W are the height and width of the input feature map respectively, X(i, j) represents the pixel value of the feature map at the i-th row and j-th column, and σ is the Sigmoid activation function. By multiplying the weight W A element-wise with the original feature, the key details related to the vascular segmentation task are strengthened, and at the same time, the generation of noise interference is effectively suppressed. The enhanced high-frequency component is as follows:

[0104] X' = W A ⊙X (11)

[0105] where W A is the importance weight, X is the input feature map, ⊙ is element-wise multiplication, and X' is the resulting feature map.

[0106] Subsequently, independent convolutional branches are applied to each high-frequency component (LH, HL, HH) to extract local edge features. The horizontal (LH) and vertical (HL) components are summed element-wise and further fused through convolution to accurately describe the directional morphological features of blood vessels. The diagonal component (HH) is processed independently to capture the blood vessel edge information distributed along the diagonal direction. This design directly corresponds to the LM region's ability to compactly encode high-order statistical features and local texture patterns. Finally, the fused high-frequency components are combined with the low-frequency component (LL) to generate an enhanced feature map. The formula is as follows:

[0107]

[0108] where F low is the low-frequency feature, is the enhanced feature of each high-frequency component. The fused feature map is further processed through subsequent convolutional layers and pooling to further reduce the feature space and provide support for the downstream blood vessel segmentation task.

[0109] B6. Design the CAKC module to enhance the recognition ability of the coronary artery edge structure by cascading the dynamic receptive fields of variable kernel convolutions. The structure of the CAKC module is as Figure 4 shown.

[0110] In this method, the present invention first uses the sobel operator to extract features in the x and y directions, and then introduces a new deformable convolution technique AKConv and designs the CAKC module. The principle is as follows:

[0111] The cascade structure is defined as:

[0112] Input = x1 (13)

[0113] x n = F n-1 (x n-1 ) + x n-1 (14)

[0114] Output = x n (15)

[0115] where n is the input feature map of the nth layer, and F n-1 (x) represents the operation of the (n - 1)th convolutional layer. In the present invention, n = 2 is taken.

[0116] AKConv provides a flexible convolution mechanism that allows the convolutional kernel to have adjustable parameter numbers and sampling shapes, thus breaking through the limitations of the fixed local window and fixed sampling shape in traditional convolutions. Traditional convolution operations are based on a regular sampling grid. For example, the sampling network R of a 3×3 convolution can be defined as:

[0117] R = {(-1, -1), (-1, 0),... (0, 1), (1, 1)} (16)

[0118] However, this regular sampling grid cannot meet the requirements of irregular convolution kernels. AKConv extends this regular sampling grid through a generation algorithm, and the sampling coordinates P n are defined as

[0119] P n = Generate(R regular ) + Offset (17)

[0120] where Generate(x) is the algorithm for generating the initial grid, R regular is the initial regular sampling network, and Offset is the cheap amount for dynamic adjustment. The sampling grid takes the upper left (0, 0) as the sampling origin and finally generates a sampling grid of any size and shape. Based on this, the convolution operation corresponding to the position P0 is defined as:

[0121] Conv(P0) = ∑ω · (P0 + P n ) (18)

[0122] where ω represents the convolution kernel parameters, and P n is the sampling point coordinate generated in the previous step. The updated sampling coordinates are:

[0123] P new = P0 + P n (19)

[0124] AKConv overcomes the regularity limitations of traditional convolutions and the application scope limitations of existing deformable convolutions, supports convolution kernel sampling of any size and shape, can efficiently extract features of irregular sampling shapes, and demonstrates great flexibility and adaptability. Moreover, AKConv also supports linear increase and decrease of the number of convolution parameters, helps optimize hardware performance, and is especially suitable for lightweight model applications, thus reducing the computational overhead of the model.

[0125] B7. Design the MSFE module. Through the multi-scale feature extraction mechanism, enhance the ability to identify and segment different-scale structures of the coronary artery. The structure of the MSFE module is as Figure 5 shown.

[0126] In the segmentation of coronary angiography images, there are often significant differences in the segmentation effects of the model on blood vessels of different scales (such as small blood vessels and major blood vessels). To improve the adaptability and generalization ability of the model on blood vessels of different scales, the present invention proposes the MSFE module. The MSFE module simulates different resolution feature extraction strategies by introducing three different scales of convolution, enabling the extraction of blood vessel information of different scales on feature maps of different resolutions, thereby comprehensively capturing the global semantics and detailed features of blood vessels. As Figure 5 shown, the MSFE module mainly consists of three convolution channels with parallel inputs, namely the Conv 1×1 channel, the Conv 3×3 channel, and the Conv 5×5 channel. Except for the Conv 1×1 convolution channel, the remaining convolution channels will undergo feature extraction and enhancement by two PSE modules. Next, the principles of these two modules will be introduced in detail.

[0127] As Figure 5 shown, the proposed PSE module aims to achieve the function of spatial feature enhancement, that is, to extract key spatial information from the initial feature map, thereby constructing a spatial attention matrix. Specifically, first, the input feature map x is generated into an initial feature map through three aggregation methods: the average feature map D, the maximum feature map U, and an additional average feature map F, which are used to retain overall information, highlight significant features, and enrich spatial details respectively, forming a diverse feature representation to facilitate the capture of subtle spatial information in the input. Next, D, U, and F are concatenated in the channel dimension to obtain a combined feature, and are mapped to a single feature space K through a convolution operation to complete spatial feature integration to adapt to different receptive field requirements. This convolution operation not only compresses the feature dimension but also strengthens the model's attention to key regions by selectively aggregating spatial information. After convolution processing, an adaptive spatial attention matrix A is generated through the Sigmoid activation function, which assigns weights to each position, reflecting the spatial relationship between pixels. This attention matrix provides a global context perspective, enabling the model to dynamically adjust its attention to spatial features, thereby focusing on key regions in complex backgrounds, effectively enhancing the ability to capture spatial details, and optimizing the overall feature expression.

[0128] EMA module (Efficient Multi-ScaleAttentionModule): The present invention introduces the EMA mechanism into the MFSE module, which is a computational unit that enhances the model's cross-dimensional feature learning ability. EMA captures pixel-level pairwise relationships by retaining multi-channel information and performing cross-dimensional interaction modeling while maintaining the computational overhead. The blood vessel information in coronary angiography images is sparse, and the segmentation results are prone to discontinuity or blurred details. This poses higher requirements for the model's cross-dimensional feature learning ability. The application of EMA not only effectively improves the segmentation accuracy of blood vessel regions but also enhances the generalization ability of the model.

[0129] The EMA module consists of three core mechanisms: feature grouping, parallel sub-structures, and cross-channel learning mechanism. The feature grouping mechanism first divides the given feature map along the channel dimension into G sub-feature maps, denoted as X = [X0, X i ,..., X G-1 . These sub-maps are then used by the EMA module to capture multi-scale spatial information through its different receptive field paths, while performing cross-channel information interaction modeling, specifically by using shared 1×1 branch paths and dedicated 3×3 kernel paths. In the shared 1×1 branch, first a pair of direction-aware feature maps are obtained, and the global information embedding of the c-th channel is shown as follows:

[0130]

[0131] where z h and z w represent the average pooling results of the c-th channel in the height and width directions respectively, X' c,i,j represents the feature value at the position (i, j) in the feature map of the c-th channel, and H and W represent the sizes in the height and width directions of the feature map. This pair of vectors obtained by parallel transformation encoding through different angles are then further transformed by a shared 1×1 convolution for C1 transformation. Then it is segmented in the height dimension h, and finally the non-linear sigmoid function is used to fit the two-dimensional binomial distribution on the linear convolution, and the formula is as follows:

[0132] w h , w w = σ(Split(C1([z h , z w )), h) (22)

[0133] where [,] is the concatenation operation along the spatial dimension, C1 is the Concat+Conv(1×1) operation, w h and w w are the convolution weights in the height and width directions respectively, σ is the Sigmoid function, and Split is the segmentation operation in a specific dimension (the height dimension in the present invention).

[0134] B8. The present invention designs the AEGF module. By automatically controlling the weights of the edge information flow and the global information flow, in order to obtain the best learning effect, the pseudo-code diagram of the AEGF module is as Figure 6 shown.

[0135] By introducing an adaptive scheduling mechanism, it can help the network to preferentially learn the topological structure outline of the image at the initial stage of training, and as the training progresses, gradually focus on more detailed information, thereby optimizing the final segmentation result. This asymptotic refinement method based on topological structure has advantages in image segmentation tasks with obvious topological features. Therefore, the present invention proposes an AEGF module. Based on the idea of adaptive topological optimization, through the dynamic fusion weights of dual-end information flow features, it better fuses edge and global features, so as to achieve the best learning effect at different training stages.

[0136] Given that the feature representations of the global feature output and the edge detection output are respectively and To achieve refined weight allocation, first, E and G are respectively divided into N sub-channels (N = 4), and the contribution degrees of the global and edge features of each sub-channel are controlled through a gating mechanism. Taking the global feature x c,i,j (x c,i,j is the value of the c-th channel at the position (i, j) in the input feature map) as the main one, the obtained dynamic weight is shown as follows:

[0137]

[0138] where W3 and W4 are linear transformation matrices respectively.

[0139] Since the pixels at different positions in the feature map contribute differently to the dynamic weight, the present invention further performs a spatial mapping operation on the bilateral feature maps in different channels. Through this spatial mapping, the edge pixel information in the sub-channels can be adaptively adjusted according to the importance of the features, so as to achieve selective enhancement or suppression of different spatial positions. Specifically, in each sub-channel, the module weights the pixel features according to the dynamic weight and concatenates all sub-channels along the channel dimension. The formula for the final fused feature map is shown as follows:

[0140]

[0141] where χ fuse is the finally fused feature map, i is the number of sub-channels, is the dynamic weight obtained above, and are the edge feature and the global feature respectively.

[0142] A3. Input the test set into the coronary artery segmentation network, and use the obtained pre-trained weights to obtain the segmentation result.

[0143] Integrating the foregoing image segmentation methods, as Figure 7 shown, the overall implementation process can be specifically divided into the following steps:

[0144] First, input the augmented training set of data;

[0145] Secondly, the data information will be input into the MSFE module and the NIEE module in parallel;

[0146] Secondly, the information processed by the NIEE module is input into the CAKC module;

[0147] Secondly, the information from different feature streams undergoes feature fusion through the AEGF module;

[0148] Secondly, obtain the pre-trained model;

[0149] Finally, use the test set to evaluate the network.

[0150] Dataset:

[0151] In the present invention, the dataset for the global vascular segmentation task comes from the Department of Cardiology, the Second Affiliated Hospital of Harbin Medical University. A total of 58 coronary angiography images of patients were collected, and the data strictly complied with the approved protocol. The images are in DICOM format, including angiography images and videos, with an imaging rate of 15 frames per second, a bit depth of 24 bits, and a resolution of 512×512 pixels. The coronary angiography images include three major branches (LAD, LCX, RCA) and multiple derivative branches. Due to the limitations of two-dimensional imaging, vascular overlap and eccentric lesions may occur, and multi-position and multi-angle imaging are required. In this study, left and right coronary angiography images were collected at multiple projection angles.

[0152] The dataset for the main vascular segmentation task comes from the Department of Cardiology, the National Medical University of Kazakhstan, and includes 1500 X-ray films of patients suspected of having coronary heart disease. The images were collected by Philips Azurion 3 and Siemens Artis Zee angiography machines, with a size of 512×512 pixels. All data has been approved by the ethics committee. The data is stored in DICOM format, including metadata such as patient information and screening date, and sensitive information has been deleted under privacy protection.

[0153] Evaluation metrics:

[0154] To evaluate the effectiveness of network segmentation, the present invention uses 95% Hausdorff distance (HD), F1_Score (F1), Accuracy (ACC), Precision, and IoU. In the pixel-level segmentation task, the confusion matrix contains four basic values: true positive (TP), true negative (TN), false positive (FP), and false negative (FN). Except for 95% HD, the other four evaluation metrics can be calculated by the following formulas:

[0155]

[0156] Among them, TP refers to the number of coronary artery pixels in the predicted image, TN refers to the number of background pixels in the predicted image, FP is the false positive, that is, the number of background pixels misidentified as vascular pixels, and FN is the false negative, that is, the number of vascular pixels misidentified as background pixels. Dice is used to evaluate the similarity between the vascular segmentation result and the ground truth, IOU is used to evaluate the overlapping degree between the segmentation result and the ground truth, Recall represents the number of pixels correctly predicted as coronary arteries among all coronary artery pixels, and Accuracy is the proportion of pixels predicted as vessels among all correctly predicted samples.

[0157] To evaluate the quality of the segmentation boundary, the present invention excludes the most extreme 5% distance values and uses the 95% HD metric to represent the maximum spatial difference between the predicted mask and the ground truth. In the HD calculation, the segmentation result and the ground truth are regarded as two subsets in the metric space, and the calculation expression is as follows:

[0158] HD = max{sup p∈P inf g∈G (p, g), sup g∈G inf p∈P (p, g)} (29)

[0159] Experimental Results and Analysis:

[0160] On the dataset of main coronary artery vessel segmentation, the objective performance indicators of each method are shown in Table 1. As can be seen from the table, the proposed model has achieved the best performance in four indicators: IOU (85.78%), ACC (98.93%), F1 Score (83.66%) and 95% HD (34.49), significantly outperforming other comparison methods. Especially in the F1 Score, because the proposed model is highly sensitive to the class imbalance problem, it can provide a key supplementary perspective on the vessel recognition performance. This is manifested in that the proposed model can more accurately reflect the prediction ability of the model for the minority class (i.e., vessels). Compared with the second-ranked MCDAU-Net method, the F1 Score of the proposed model has increased by 0.44%, while compared with the worst-performing PIDNet method, it has been significantly improved by 10.07%. In terms of the edge-related indicator (95% HD), the proposed model also performs excellently, reducing by 0.61 compared with the second-ranked MCDAU-Net method and significantly decreasing by 42.52% compared with the worst-performing FRUnet method, further demonstrating its superiority. The excellent performance of these objective indicators provides important evidence support for the effectiveness of the proposed method.

[0161] Table 1 Comparison of Results of Different Segmentation Networks

[0162]

[0163] Table 2 shows the segmentation performance metrics of the model of the present invention and other models on the private dataset, which is a dataset dedicated to evaluating the segmentation task covering all coronary artery vessels, including the main trunk and small vessels. Within a limited number of training cycles, the model of the present invention achieved the best performance in all four metrics. Compared with the segmentation tasks only for the main vessels, the pixel-level accuracy of most models in the comprehensive vessel segmentation task has been improved because the comprehensive segmentation task requires capturing more small vessel details, and the tendency of the model to over-segment instead becomes an advantage in this scenario - the more detailed the segmentation, the more beneficial it is to performance improvement. The only exception is the ACC metric, whose performance in the comprehensive segmentation task has decreased. This is because this metric reflects the overall classification performance of the model for vascular and non-vascular pixels, and the number of non-vascular pixels in the comprehensive segmentation task has relatively decreased, resulting in a relatively low overall weight of ACC. Specifically, compared with the IMFF-Net method with the sub-optimal F1 Score, the F1 Score of the model of the present invention has increased by 0.72%, and compared with the worst-performing model SkelCon, the improvement amplitude is as high as 12.44%. In terms of the 95% HD metric, the model of the present invention also shows outstanding performance, reducing by 2.46% compared with IMFF-Net and 26.21% compared with SkelCon. These results indicate that the model of the present invention has demonstrated excellent segmentation ability in the comprehensive segmentation task, especially in maintaining the fitting degree of the segmentation edge with the ground truth and dealing with the imbalance between positive and negative samples.

[0164] Table 2 Comparison of Results of Different Segmentation Networks

[0165]

[0166] Ablation Experiments:

[0167] To further explore the influence of each part module in the dual-stream collaborative network on the coronary artery segmentation performance, the present invention conducted a large number of ablation experiment studies on the NIEE module, CAKC module, AEGF module, MSFE module, and SubPixel module based on the constructed private coronary angiography image dataset. The experimental results are shown in Table 3.

[0168] Table 3 Results of Ablation Experiments

[0169]

[0170]

[0171] The ablation study results in Table 3 show that the individual NIEE module, CAKC module, AEGF module, MSFE module, and SubPixel module are effective in improving the segmentation performance, and the network shows the best segmentation performance when all modules act simultaneously. Especially for the F1 Score, which is more concerned about the positive class task, the recall rate increases by 4.58% when all modules act simultaneously compared with the original backbone network. Since coronary artery segmentation with a higher F1 Score usually has greater clinical value in diagnosing coronary artery stenosis and other coronary artery diseases, among all the network structures compared in the ablation study, the dual-stream collaborative segmentation network is the best choice for coronary artery segmentation.

[0172] The automatic segmentation method of coronary angiography images based on the dual-stream collaborative network disclosed in the present invention aims at the problems of large background noise, low contrast, local blur of blood vessel boundaries in coronary angiography images, and the resulting high false positive rate of existing methods. The MSFE module is proposed in the main network branch to capture global semantic information at different levels and enhance the understanding ability of the complex structure of the coronary artery. The NIEE module is proposed in the edge branch to suppress noise interference and enhance the edge contour, and then the proposed CAKC module is used to extract fine-grained edge information, thereby strengthening the expression of edge details. At different learning stages, the proposed AEGF module is used to realize the dynamic fusion of the global semantic features of the main network and the local detail features of the edge branch, effectively combining global and local information. Finally, all information flows are aligned in resolution to obtain the output. The segmentation result obtained in this way not only has high pixel-level accuracy, but also can maintain clear edges and structural coherence in a complex imaging environment, providing high-quality clinical auxiliary support for the diagnosis of coronary heart disease. Experimental results show that the method can effectively extract, fuse, and correct mixed features of coronary angiography images, and is significantly better than other coronary angiography image segmentation comparison models in reducing the false positive rate of blood vessel segmentation.

Claims

1. A method for segmenting coronary angiography images based on a dual-stream collaborative network, characterized in that The method includes the following steps: Step 1: Construct a coronary angiography image dataset. The specific steps are as follows: Step 1.1: Obtain a coronary angiography image dataset; Step 1.2: Divide the dataset into a training set, a validation set, and a test set; Step 1.3: In the training stage, preprocess the original images used; Step 2: Construct a two-stream collaborative network for coronary angiography image segmentation, which specifically includes the following steps: Step 2.1: Construct a two-stream collaborative network including a main segmentation network and an edge segmentation network: The input coronary angiography image data is sent into the main segmentation network and the edge segmentation network in parallel, where the edge segmentation network consists of the high-level encoding module M bio and the dynamic receptive field edge module M drf and the feature F bio obtained via M bio is expressed as follows: F bio = M bio (I) Via M drf The obtained feature F drf Is represented as follows: F drf = M drf (J) Among them, F bio is the feature map obtained through M bio , and F drf is the feature map obtained through M drf , and J is the feature map fed into M drf ; Meanwhile, the input feature I enters the main segmentation network, which consists of multiple multi-scale information extraction modules M global,i (i = 1, 2,..., n), where n is the number of multi-scale information extraction modules. The feature output of the main segmentation network is shown as follows: Finally, the feature maps extracted by the main segmentation network and the edge segmentation network are uniformly processed and fused through a soft self-alignment module based on sub-pixel convolution, so as to obtain the final output result; Step 2.2: In the edge segmentation network, design an efficient encoding module NIEE to suppress image noise and enhance vascular edge features; Step 2.3: Design a cascaded adaptive kernel convolution module CAKC in the edge segmentation network to enhance the recognition ability of the coronary artery edge structure through the dynamic receptive field of cascaded adaptive kernels; Step 2.4: Design a multi-scale information extraction module MSFE in the main segmentation network to enhance the recognition and segmentation ability of coronary artery structures at different scales. At the same time, in order to achieve adaptive spatial feature enhancement, a parallel spatial modulation module PSE is designed within the MSFE module to further extract and enhance spatial features, and the recognition and segmentation ability of coronary artery structures at different scales is enhanced through a multi-scale feature extraction mechanism; Step 2.5: Design an adaptive feature fusion module AEGF by automatically controlling the weights of the edge information flow and the global information flow; Step 3: Input the data-augmented training set and validation set into the main segmentation network and the edge segmentation network to obtain a pre-trained coronary artery segmentation network model; Step 4: Input the test set into the pre-trained coronary artery segmentation network model to evaluate the effect of the network.

2. The coronary angiography image segmentation method based on a dual-stream collaborative network according to claim 1, wherein In the above Step 1.3, the preprocessing operations include data augmentation and color adjustment.

3. The method for segmenting coronary angiography images based on a dual-stream collaborative network according to claim 2, wherein The data augmentation includes randomly rotating the coronary angiography images input for network training within the range of -10° to 10°, mirror transformation, Gaussian blur, randomly scaling within the range of 0.5 to 2.0, randomly cropping the images, and randomly changing the aspect ratio within the range of 0.7 to 1.

3.

4. The coronary angiography image segmentation method based on a dual-stream collaborative network according to claim 2, wherein The color adjustment includes performing a hue change within the range of -0.1 to 0.1, a saturation adjustment within the range of 0.7 to 1.7, and a brightness adjustment within the range of 0.7 to 1.3 on the coronary angiography images input for network training.

5. The coronary angiography image segmentation method based on a dual-stream collaborative network according to claim 1, wherein The specific steps of the above Step 2.2 are as follows: Step 2.2.1: DWT first decomposes the image in the horizontal and vertical directions through a low-pass filter L and a high-pass filter H, as shown in the following formula: where X[i,k] is an element of the original image at position (i,k), and X L [i,j] is the low-frequency part of the signal, and X H [i,j] is the high-frequency part of the signal, and H[k - j] and L[k - j] are the weights of the high-pass and low-pass filters, respectively; Step Two Two Two, perform operations on X L and X H Use LH to perform convolution again in the row direction and column direction respectively to generate four sub-bands: low-frequency LL, high-frequency horizontal LH, high-frequency vertical HL, and high-frequency diagonal HH, as shown in the following formula: where L[k - j] and H[k - j] are respectively the low-frequency part and the high-frequency part of the signal obtained through wavelet transform; LL[i, j], LH[i, j], HL[i, j], and HH[i, j] are respectively the obtained low-frequency sub-band, high-frequency horizontal sub-band, high-frequency vertical sub-band, and high-frequency diagonal sub-band; Step Two Two Three: For the decomposed frequency components, design an AEGF module to further simulate the noise selective suppression and key feature highlighting mechanisms in the LM region. The AEGF module calculates the importance weight W through adaptive global average pooling and per-channel convolution A , and dynamically amplifies or suppresses each frequency component selectively. The formula is as follows: Among them, W1 and W2 are the per-channel convolution weights respectively, H and W are the height and width of the input feature map respectively, X(i,j) represents the pixel value of the feature map at the i-th row and j-th column, and σ is the Sigmoid activation function; the enhanced high-frequency components are as follows: X' = W A ⊙X Among them, W A is the importance weight, X is the input feature map, ⊙ is the element-wise multiplication, and X' is the resulting feature map; Step 224: Apply independent convolution branches to each of the high-frequency components LH, HL, and HH to extract local edge features, and sum the LH and HL components element-wise and then further fuse them through convolution to accurately describe the directional morphological features of blood vessels; process HH independently to capture the blood vessel edge information distributed along the diagonal direction. Finally, combine the fused high-frequency components with the low-frequency components to generate the enhanced feature map, and the formula is as follows: Among them, F low is a low-frequency feature, and is the enhanced feature of each high-frequency component.

6. The coronary angiography image segmentation method based on a dual-stream collaborative network according to claim 1, characterized in that The specific steps of the above Step 23 are as follows: Use the sobel operator to extract features in the x and y directions, and then introduce the deformable convolution technology AKConv, design the CAKC module, and the cascade structure is defined as: Input=x1 x n = F n-1 (x n-1 ) + x n-1 Output = x n where n represents the input feature map of the nth layer, and F n-1 (x) represents the convolutional layer operation of the (n - 1)th layer; AKConv extends the regular sampling grid by a generation algorithm, and the sampling coordinates P n are defined as: P n = Generate(R regular ) + Offset where Generate(x) is the algorithm for generating the initial grid, and R regular is the initial regular sampling network, and Offset is the dynamically adjusted offset; the sampling grid takes the upper left (0, 0) as the sampling origin, and finally generates a sampling grid of any size and shape. Based on this, the convolution operation corresponding to the position P0 is defined as: Conv(P0) = Σω·(P0 + P n ) where ω represents the convolution kernel parameter, and P n is the coordinate of the sampling point generated in the previous step. The updated sampling coordinate is: P new = P0 + P n .

7. The method for segmenting coronary angiography images based on a two-stream collaborative network according to claim 1, characterized in that The specific steps of the above Step 24 are as follows: Step 241: The MSFE module consists of three convolution channels with parallel inputs, namely the Conv 1×1 channel, the Conv3×3 channel, and the Conv 5×5 channel. Among them, in the Conv 3×3 channel and the Conv 5×5 channel, the PSE module is added for feature extraction and enhancement. Specifically, first, the input feature map x generates the initial feature maps through three aggregation methods: the average feature map D, the maximum feature map U, and the additional average feature map F. They are used to retain the overall information, highlight the significant features, and enrich the spatial details respectively, forming diverse feature representations to facilitate capturing the subtle spatial information in the input; next, concatenate D, U, and F in the channel dimension to obtain the combined feature, and map it to a single feature space K through convolution operations to complete the spatial feature integration to adapt to different receptive field requirements; after the convolution process, generate the adaptive spatial attention matrix A through the Sigmoid activation function to assign weights to each position and reflect the spatial relationship between pixels; Step 242: An EMA module is introduced into the MFSE module. The EMA module consists of three core mechanisms: feature grouping, parallel sub-structures, and cross-channel learning mechanism. The feature grouping mechanism first divides the given feature map along the channel dimension into G sub-feature maps, denoted as X = [X0, X i ,..., X G-1 . These sub-maps are then used by the EMA module to capture multi-scale spatial information using its different receptive field paths and simultaneously perform cross-channel information interaction modeling, where: in the shared 1×1 branch, a pair of direction-aware feature maps is first obtained, and the global information embedding of the c-th channel is shown as follows: where z h and z w represent the average pooling results of channel c in height and width respectively, and X′ c,i,j represents the feature value at position (i, j) in the feature map of the c-th channel; for the vectors obtained by parallel transformation encoding through different angles, continue to perform C1 transformation using a shared 1×1 convolution, then segment in the height dimension h, and finally use a non-linear sigmoid function to fit the two-dimensional binomial distribution on the linear convolution. The formula is as follows: w h ,w w = σ(Split(C1([z h ,z w )),h) Among them, [,] is a concatenation operation along the spatial dimension, C1 is a Concat+Conv(1×1) operation, w h and w w are the convolution weights in the height and width directions respectively, σ is the Sigmoid function, and Split is a splitting operation on the spatial dimension of the input feature.

8. The method for segmenting coronary angiography images based on a dual-stream collaborative network according to claim 1, wherein The specific steps of the above Step 25 are as follows: Step 251. Given that the feature representations of the global feature output and the edge detection output are respectively and Divide E and G into N sub-channels respectively, and control the contribution degrees of the global and edge features of each sub-channel through a gating mechanism; With the global feature x c,i,j as the main, the obtained dynamic weight is as shown in the following formula: where x c,i,j is the value of the c-th channel at position (i, j) in the input feature map, and W3 and W4 are linear transformation matrices respectively; Step 252: Perform spatial mapping operations on the bilateral feature maps in different channels, and the edge pixel information in the sub-channels is adaptively adjusted according to the importance of the features, so as to realize the selective enhancement or suppression of different spatial positions. Finally, the fused feature map formula is shown as follows: Among them, χ fuse is the finally fused feature map, i is the number of sub-channels, and are the edge feature and the global feature respectively. The present invention discloses a coronary angiography image segmentation method based on a dual-stream collaborative network. The method constructs a dual-stream collaborative network based on the idea of maintaining different resolutions and introducing auxiliary information. In the main segmentation network branch, an MSFE module is proposed to capture global semantic information at different levels and enhance the ability to understand the complex structure of coronary arteries. In the edge segmentation network branch, an NIEE module is proposed to suppress noise interference and enhance the edge contour, and then a CAKC module is used to extract fine-grained edge information, thereby strengthening the expression of edge details. In different learning stages, an AEGF module is used to achieve the dynamic fusion of the global semantic features of the main segmentation network and the local detail features of the edge segmentation network branch to obtain the best learning effect. Finally, all information flows are aligned in resolution to obtain the final output result. This method can effectively suppress noise interference, enhance edge contour features, and achieve high-precision and robustness in vascular segmentation.