Training method for blood vessel recognition model, blood vessel recognition method, device, electronic device, storage medium and computer program product
Through the multi-frame fusion algorithm L-MSFUnet, a vascular recognition model containing LSTM and MSFUnet networks is trained using multiple DSA images, solving the problem that it is difficult to obtain a complete vascular tree in a single image frame, and achieving more accurate vascular recognition and more comprehensive vascular tree generation.
Patent Information
- Application Number
- CN202510402251.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In interventional surgery, it is difficult to obtain a complete abdominal aortic vascular tree in a single DSA image frame, resulting in inaccurate vascular identification and may cause serious medical accidents.
Using the multi-frame fusion algorithm L-MSFUnet, a vascular recognition model containing the LSTM network and multiple MSFUnet networks is trained using multiple angiographic images, and a more comprehensive coverage of the abdominal aortic vascular tree is generated using multi-frame data.
It effectively improves the accuracy of vascular identification, reduces the risk of medical accidents, and generates a more comprehensive vascular tree structure through multi-frame data processing.
Smart Images

Figure CN119919773B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for training a blood vessel recognition model, a blood vessel recognition method, an apparatus, an electronic device, a storage medium, and a computer program product. Background Art
[0002] In interventional surgery, it is crucial to accurately identify the aorta and its branch vessels from Digital Subtraction Angiography (DSA) images. In endovascular aneurysm repair of the abdominal aorta, in order to make the blood vessels visible, doctors often need to inject contrast agents into the blood vessels multiple times to display different blood vessel sites, so as to guide the selection of the stent type and the placement site. Moreover, since the blood is constantly flowing, it is necessary to perform long-term X-ray irradiation on the abdomen of the organism, and this will obtain a frame sequence containing multiple DSA images.
[0003] In the related art, when using a DSA image sequence for blood vessel recognition, mainly a most representative image frame is selected from the DSA image sequence to determine the shape and position of the blood vessels. However, due to the large abdominal space, it is difficult to obtain a complete abdominal aortic blood vessel tree in a single image frame, which may lead to inaccurate blood vessel recognition and thus cause serious medical accidents. Summary of the Invention
[0004] The present disclosure provides a method for training a blood vessel recognition model, a blood vessel recognition method, an apparatus, an electronic device, a storage medium, and a computer program product, so as to at least solve the problem in the related art that it is difficult to obtain a complete abdominal aortic blood vessel tree in a single image frame, which may lead to inaccurate blood vessel recognition and thus cause serious medical accidents. The technical solutions of the present disclosure are as follows:
[0005] According to the first aspect of the embodiments of the present disclosure, a method for training a blood vessel recognition model is provided. The blood vessel recognition model includes an LSTM network and multiple MSFUnet networks. The training method includes: obtaining multiple blood vessel images, where the multiple blood vessel images are images obtained by injecting a contrast agent into the blood vessels of an organism; dividing the multiple blood vessel images into multiple image groups, where the multiple image groups correspond one-to-one to the multiple MSFUnet networks, and each image group includes at least two of the blood vessel images; inputting each image group into the MSFUnet network corresponding to the image group to obtain a first blood vessel recognition result output by the MSFUnet network; inputting multiple first blood vessel recognition results corresponding one-to-one to the multiple MSFUnet networks into the LSTM network to obtain a final blood vessel recognition result; calculating a first loss based on the first blood vessel recognition result output by each MSFUnet network and the first image label bound to the image group corresponding to the MSFUnet network, where the first image label is an image label including the blood vessel morphology annotated based on at least two of the blood vessel images included in the image group; calculating a second loss based on the final blood vessel recognition result output by the LSTM network and the second image label bound to the multiple blood vessel images, where the second image label is an image label including the blood vessel morphology annotated based on the multiple blood vessel images; training the blood vessel recognition model by adjusting the parameters of the MSFUnet network based on the first loss and adjusting the parameters of the LSTM network based on the second loss.
[0006] Optionally, the MSFUnet network includes a dual attention module, a channel attention module, a plurality of encoder modules, a plurality of residual convolution modules, and a plurality of optical flow modules; inputting each image group into the MSFUnet network corresponding to the image group to obtain a first blood vessel recognition result output by the MSFUnet network includes: encoding the image group through the plurality of encoder modules to obtain an encoding result; inputting the encoding result into the dual attention module to obtain a first attention result; performing optical flow processing on the first attention result through the plurality of optical flow modules to obtain an optical flow result; for the first residual convolution module among the plurality of residual convolution modules, obtaining a residual convolution result of the residual convolution module based on the first attention result, the output of the encoder module at the same layer as the residual convolution module, and the output of the optical flow module at the same layer as the residual convolution module; for each of the second to nth residual convolution modules among the plurality of residual convolution modules, obtaining a residual convolution result of the residual convolution module based on the output of the previous-layer residual convolution module of the residual convolution module, the output of the encoder module at the same layer as the residual convolution module, and the output of the optical flow module at the same layer as the residual convolution module, where n is an integer greater than or equal to 2; obtaining a second attention result through the channel attention module based on the first attention result and the residual convolution results of each residual convolution module in the plurality of residual convolution modules; and obtaining the first blood vessel recognition result of the MSFUnet network based on the residual convolution result output by the residual convolution module located at the last layer in the plurality of residual convolution modules and the second attention result.
[0007] Optionally, the training method further includes: for each MSFUnet network, calculating a third loss based on the optical flow result output by the optical flow module located at the last layer in the plurality of optical flow modules included in the MSFUnet network and the first image label; and adjusting the parameters of the MSFUnet network by adjusting the parameters of the MSFUnet network based on both the first loss and the third loss.
[0008] Optionally, each encoder module includes a downsampling module; the encoding of the image group by the multiple encoder modules to obtain an encoding result includes: for the first encoder module, obtaining the downsampled features of the average pooling branch and the downsampled features of the max pooling branch of the image group through the downsampling module included in the encoder module, and splicing the downsampled features of the average pooling branch and the downsampled features of the max pooling branch to obtain an output result; for each encoder module other than the first encoder module among the multiple encoder modules, obtaining the downsampled features of the average pooling branch and the downsampled features of the max pooling branch of the output result of the previous encoder module through the downsampling module included in the encoder module, and splicing the downsampled features of the average pooling branch and the downsampled features of the max pooling branch to obtain an output result; taking the output result of the last encoder module among the multiple encoder modules as the encoding result; wherein, the downsampled features of the average pooling branch are obtained by the following method: performing an average pooling operation on the image group or the output result of the previous encoder module to obtain an average pooling result; sequentially performing a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the average pooling result to obtain a first convolution kernel; convolving the first convolution kernel with the average pooling result to obtain the downsampled features of the average pooling branch; the downsampled features of the max pooling branch are obtained by the following method: performing a max pooling operation on the image group or the output result of the previous encoder module to obtain a max pooling result; sequentially performing a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the max pooling result to obtain a second convolution kernel; convolving the second convolution kernel with the max pooling result to obtain the downsampled features of the max pooling branch.
[0009] According to a second aspect of the embodiments of the present disclosure, there is provided a blood vessel recognition method, which is implemented based on a blood vessel recognition model trained by the training method according to the present disclosure. The blood vessel recognition method includes: obtaining multiple target blood vessel images to be recognized, wherein the multiple target blood vessel images to be recognized are images obtained by injecting a contrast agent into the blood vessels of an organism; dividing the multiple target blood vessel images to be recognized into multiple image groups to be recognized, wherein the multiple image groups to be recognized correspond one-to-one to the multiple MSFUnet networks, and each image group to be recognized includes at least two of the target blood vessel images to be recognized; inputting each image group to be recognized into the MSFUnet network corresponding to the image group to be recognized to obtain a first blood vessel recognition target result output by the MSFUnet network; inputting the multiple first blood vessel recognition target results corresponding one-to-one to the multiple MSFUnet networks into the LSTM network to obtain a final blood vessel recognition target result.
[0010] According to a third aspect of the embodiments of the present disclosure, there is provided a training device for a blood vessel recognition model. The blood vessel recognition model includes an LSTM network and a plurality of MSFUnet networks. The training device includes: a blood vessel image acquisition module configured to acquire a plurality of blood vessel images, where the plurality of blood vessel images are images obtained by injecting a contrast agent into the blood vessels of an organism; an image group division module configured to divide the plurality of blood vessel images into a plurality of image groups, where the plurality of image groups correspond one-to-one to the plurality of MSFUnet networks, and each image group includes at least two of the blood vessel images; a first blood vessel recognition result acquisition module configured to input each image group into the MSFUnet network corresponding to the image group to obtain a first blood vessel recognition result output by the MSFUnet network; a final blood vessel recognition result acquisition module configured to input a plurality of first blood vessel recognition results corresponding one-to-one to the plurality of MSFUnet networks into the LSTM network to obtain a final blood vessel recognition result; a first loss calculation module configured to calculate a first loss based on the first blood vessel recognition result output by each MSFUnet network and a first image label bound to the image group corresponding to the MSFUnet network, where the first image label is an image label including the blood vessel morphology annotated based on at least two of the blood vessel images included in the image group; a second loss calculation module configured to calculate a second loss based on the final blood vessel recognition result output by the LSTM network and a second image label bound to the plurality of blood vessel images, where the second image label is an image label including the blood vessel morphology annotated based on the plurality of blood vessel images; a training module configured to train the blood vessel recognition model by adjusting the parameters of the MSFUnet network based on the first loss and adjusting the parameters of the LSTM network based on the second loss.
[0011] Optionally, the MSFUnet network includes a dual attention module, a channel attention module, a plurality of encoder modules, a plurality of residual convolution modules, and a plurality of optical flow modules; the first vascular recognition result acquisition module is configured to: encode the image group through the plurality of encoder modules to obtain an encoding result; input the encoding result into the dual attention module to obtain a first attention result; perform optical flow processing on the first attention result through the plurality of optical flow modules to obtain an optical flow result; for the first residual convolution module in the plurality of residual convolution modules, based on the first attention result, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module, obtain the residual convolution result of this residual convolution module; for each of the second to nth residual convolution modules in the plurality of residual convolution modules, based on the output of the previous layer's residual convolution module of this residual convolution module, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module, obtain the residual convolution result of this residual convolution module, where n is an integer greater than or equal to 2; obtain a second attention result through the channel attention module based on the first attention result and the residual convolution result of each residual convolution module in the plurality of residual convolution modules; obtain the first vascular recognition result of the MSFUnet network based on the residual convolution result output by the residual convolution module located at the last layer in the plurality of residual convolution modules and the second attention result.
[0012] Optionally, the training device further includes: a third loss calculation module configured to calculate a third loss for each MSFUnet network based on the optical flow result output by the optical flow module located at the last layer in the plurality of optical flow modules included in this MSFUnet network and the first image label; the training module is configured to: adjust the parameters of the MSFUnet network based on both the first loss and the third loss.
[0013] Optionally, each encoder module includes a downsampling module; the first vascular recognition result acquisition module is configured to: for the first encoder module, obtain the downsampled features of the average pooling branch and the downsampled features of the max pooling branch of the image group through the downsampling module included in the encoder module, and splice the downsampled features of the average pooling branch and the downsampled features of the max pooling branch to obtain an output result; for each encoder module other than the first encoder module among the multiple encoder modules, obtain the downsampled features of the average pooling branch and the downsampled features of the max pooling branch of the output result of the previous encoder module through the downsampling module included in the encoder module, and splice the downsampled features of the average pooling branch and the downsampled features of the max pooling branch to obtain an output result; use the output result of the last encoder module among the multiple encoder modules as the encoding result; wherein, the downsampled features of the average pooling branch are obtained by the following method: perform an average pooling operation on the image group or the output result of the previous encoder module to obtain an average pooling result; sequentially perform a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the average pooling result to obtain a first convolution kernel; perform a convolution of the first convolution kernel and the average pooling result to obtain the downsampled features of the average pooling branch; the downsampled features of the max pooling branch are obtained by the following method: perform a max pooling operation on the image group or the output result of the previous encoder module to obtain a max pooling result; sequentially perform a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the max pooling result to obtain a second convolution kernel; perform a convolution of the second convolution kernel and the max pooling result to obtain the downsampled features of the max pooling branch.
[0014] According to a fourth aspect of the embodiments of the present disclosure, a blood vessel recognition device is provided. The blood vessel recognition device is implemented based on a blood vessel recognition model trained by the training method according to the present disclosure. The blood vessel recognition device includes: a target blood vessel image acquisition module configured to acquire multiple target blood vessel images to be recognized, where the multiple target blood vessel images to be recognized are images obtained by injecting a contrast agent into the blood vessels of an organism; a to-be-recognized image group division module configured to divide the multiple target blood vessel images to be recognized into multiple to-be-recognized image groups, where the multiple to-be-recognized image groups correspond one-to-one to the multiple MSFUnet networks, and each to-be-recognized image group includes at least two of the target blood vessel images to be recognized; a first blood vessel recognition target result acquisition module configured to input each to-be-recognized image group into the MSFUnet network corresponding to the to-be-recognized image group to obtain a first blood vessel recognition target result output by the MSFUnet network; and a blood vessel recognition final target result acquisition module configured to input the multiple first blood vessel recognition target results corresponding one-to-one to the multiple MSFUnet networks into the LSTM network to obtain a blood vessel recognition final target result.
[0015] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory for storing executable instructions of the processor; where the processor is configured to execute the instructions to implement the training method or the blood vessel recognition method of the blood vessel recognition model according to the present disclosure.
[0016] According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method or the blood vessel recognition method of the blood vessel recognition model according to the present disclosure.
[0017] According to a seventh aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, where the computer program, when executed by a processor, implements the training method or the blood vessel recognition method of the blood vessel recognition model according to the present disclosure.
[0018] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0019] In the present disclosure, by using the multi-frame fusion algorithm L-MSFUnet, that is, by using multiple angiography images for blood vessel recognition, more data volume and richer data content are considered, and thus a more comprehensive abdominal aortic vascular tree can be generated, that is, the accuracy of blood vessel recognition can be effectively improved, and thus the occurrence of medical accidents can be effectively avoided.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0021] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0022] Figure 1 is a flowchart showing a method for training a blood vessel recognition model according to an exemplary embodiment of the present disclosure;
[0023] Figure 2 is a schematic structural diagram showing a blood vessel recognition model according to an exemplary embodiment of the present disclosure;
[0024] Figure 3 is a schematic diagram showing intra-frame position encoding and temporal frame encoding according to an exemplary embodiment of the present disclosure;
[0025] Figure 4 is a schematic structural diagram showing the structure of SFB according to an exemplary embodiment of the present disclosure;
[0026] Figure 5 is a schematic structural diagram showing the structure of an optical flow module (Flow) according to an exemplary embodiment of the present disclosure;
[0027] Figure 6 is a schematic structural diagram showing the structure of a residual convolution module ResDe according to an exemplary embodiment of the present disclosure;
[0028] Figure 7 is a schematic structural diagram showing the structure of a CA module according to an exemplary embodiment of the present disclosure;
[0029] Figure 8 is a schematic structural diagram showing the structure of a downsampling module (FreqPass) according to an exemplary embodiment of the present disclosure;
[0030] Figure 9 is a schematic comparison diagram showing the comparison between the single-frame prediction segmentation results of different blood vessel regions and the blood vessel fusion segmentation results obtained from the complete sequence according to an exemplary embodiment of the present disclosure;
[0031] Figure 10 is a flowchart showing a blood vessel recognition method according to an exemplary embodiment of the present disclosure;
[0032] Figure 11 is a block diagram showing a training device for a blood vessel recognition model according to an exemplary embodiment of the present disclosure;
[0033] Figure 12 is a block diagram showing a blood vessel recognition device according to an exemplary embodiment of the present disclosure;
[0034] Figure 13 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0035] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0037] Abdominal aortic aneurysm is a local permanent dilation of the abdominal aorta. Endovascular aneurysm repair (EVAR) has rapidly developed into the preferred method for treating abdominal aortic aneurysm with relatively low morbidity and mortality. In interventional surgery, it is crucial to accurately identify the aortic blood vessels and their branch vessels from DSA images. At this time, doctors often need to inject contrast agent into the user's blood vessels multiple times to display different blood vessel sites, and then determine the type selection and placement site of the implanted stent. In a complete sequence of blood vessel image frames, there are often many noises, including the intestines, kidneys, complex fine blood vessels, relatively blurred blood vessel edges with mural thrombus, and so on. In addition, due to the large abdominal space, it is difficult to obtain a complete angiogram of the abdominal aortic vascular tree in a single image. Inaccurate blood vessel identification may lead to incorrect placement of endovascular grafts, and thus cause surgical failure. In addition, traditional methods usually only focus on the segmentation of the abdominal aorta or local blood vessels, ignoring the key point of constructing a complete vascular tree including branch vessels.
[0038] To solve the above problems in the related art, the present disclosure provides a method for training a blood vessel recognition model, a blood vessel recognition method, a device, an electronic device, a storage medium, and a computer program product. By using the multi-frame fusion algorithm L-MSFUnet, that is, by using multiple angiography images for blood vessel recognition, more data volume and richer data content are considered, and thus it can be ensured to generate a more comprehensive abdominal aortic vascular tree, that is, the accuracy of blood vessel recognition can be effectively improved, and thus the occurrence of medical accidents can be effectively avoided.
[0039] Figure 1FIG. 0 is a flowchart showing a method for training a blood vessel recognition model according to an exemplary embodiment of the present disclosure. The blood vessel recognition model may include a Long Short-Term Memory (LSTM) network and multiple Multi-frame Sequence Fusion U-net (MSFUnet) networks that combine multiple frames of sequences.
[0040] Referring to Figure 1 , in step 101, multiple blood vessel images may be obtained. Among them, the multiple blood vessel images may be images obtained by injecting a contrast agent into the blood vessels of an organism.
[0041] In step 102, the multiple blood vessel images may be divided into multiple image groups. Among them, the multiple image groups may correspond one-to-one to multiple MSFUnet networks, and each image group may include at least two blood vessel images. Figure 2 FIG. 10 is a schematic structural diagram of a blood vessel recognition model according to an exemplary embodiment of the present disclosure. Referring to Figure 2 , the overall structure of the blood vessel recognition model, that is, the L-MSFUne model, may be composed of an encoder-decoder segmentation network MSFUnet network and an LSTM network module. Exemplarily, Figure 2 FIG. 14 shows 3 MSFUnet networks. Each MSFUnet network may correspond to an image group, and each image group may include 3 blood vessel images arranged in sequence.
[0042] It should be noted that the input images may also be divided into patches of size 5, and the division step size may be 4. Through this overlapping division operation, the edge effect of the image can be reduced. In addition, when encoding, in addition to performing position encoding within the frame, the time information of the images may also be encoded according to the input order of the blood vessel images. Exemplarily, Figure 2 the time encodings of the 3 frames of blood vessel images included in each image group shown in FIG. 19 may be: 0, 1, 2, respectively.
[0043] Figure 3 FIG. 23 is a schematic diagram showing intra-frame position encoding and time frame encoding according to an exemplary embodiment of the present disclosure. Referring to Figure 3 , a single image group may include 3 frames of blood vessel images. The "0, 1, 2, 3" under the first frame of blood vessel image may represent the position encodings of 4 image regions included in this frame of blood vessel image, and the arrangement positions of the 4 image regions within the first frame of blood vessel image may be, but are not limited to: "from left to right, from top to bottom". And the "0, 0, 0, 0" under the first frame of blood vessel image may represent that the time encoding of this frame of blood vessel image is 0.
[0044] "0, 1, 2, 3" under the blood vessel image of the second frame can represent the position encoding of 4 image regions included in this frame of blood vessel image, and the arrangement positions of these 4 image regions within the second frame of blood vessel image can be, but are not limited to: "from left to right, from top to bottom". Also, "1, 1, 1, 1" under the second frame of blood vessel image can represent that the time encoding of this frame of blood vessel image is 1.
[0045] "0, 1, 2, 3" under the blood vessel image of the third frame can represent the position encoding of 4 image regions included in this frame of blood vessel image, and the arrangement positions of these 4 image regions within the third frame of blood vessel image can be, but are not limited to: "from left to right, from top to bottom". Also, "2, 2, 2, 2" under the third frame of blood vessel image can represent that the time encoding of this frame of blood vessel image is 2. In this way, by introducing time frame encoding, the feature continuity across multiple frames can be enhanced.
[0046] In step 103, each image group can be input into the MSFUnet network corresponding to this image group to obtain the first blood vessel recognition result output by this MSFUnet network. Exemplarily, as described above, each image group can contain 3 consecutive frames of DSA images, and these DSA images can be sampled to 512 512 in size.
[0047] According to an exemplary embodiment of the present disclosure, the MSFUnet network can include a double-attention module, a Channel Attention (CA) module, multiple encoder modules (SpatialFrequency Block, SFB), multiple residual convolution modules (resdecoder block, ResDe), and multiple optical flow modules (flow attention).
[0048] Referring to Figure 2 , the image group can be encoded through multiple encoder modules to obtain an encoding result. Then, the encoding result can be input into the double-attention module to obtain a first attention result. Next, the first attention result can be subjected to optical flow processing through multiple optical flow modules to obtain an optical flow result.
[0049] Then, for the first residual convolution module among the multiple residual convolution modules, the residual convolution result of this residual convolution module can be obtained based on the first attention result, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module; for each of the second to nth residual convolution modules among the multiple residual convolution modules, the residual convolution result of this residual convolution module can be obtained based on the output of the previous-layer residual convolution module of this residual convolution module, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module, where n can be an integer greater than or equal to 2.
[0050] Next, the second attention result can be obtained through the channel attention module based on the first attention result and the residual convolution result of each residual convolution module among the multiple residual convolution modules. Then, based on the residual convolution result output by the residual convolution module at the last layer among the multiple residual convolution modules and the second attention result, the first blood vessel recognition result of the MSFUnet network can be obtained. Specifically, the Kronecker product operation can be performed on the residual convolution result output by the residual convolution module at the last layer and the second attention result, and thus the first blood vessel recognition result of the MSFUnet network can be obtained.
[0051] Figure 4 is a schematic structural diagram of an SFB showing an exemplary embodiment according to the present disclosure. Referring to Figure 4 , the SFB module can be composed of 2 self-attention modules and a downsampling module (FreqPass). When calculating the self-attention feature, the feature can be divided into 8 8-patch calculation windows, and the self-attention feature (WMSA) can be calculated within the calculation window. In addition, in order to improve the context connection of the entire image, a window shifting operation (SWMSA) can be adopted in the second self-attention module. The self-attention module can adopt a multi-head self-attention mechanism, that is, different numbers can be adopted between different layers. Exemplarily, from the first layer to the fifth layer, they can be set to: 4, 8, 16, 32, 64 respectively.
[0052] It should be noted that referring to Figure 2, the decoder included in the MSFUnet network can be composed of three parts, namely: the ResDe module, the Flow module, and the CA module. After the feature extraction of five layers of SFB, the extracted features can be input into the double-attention layer, and then the output results of the double-attention layer can be used as the input of the optical flow attention module (Flow) and the input of the residual convolution module (ResDe) respectively. Exemplarily, the "double-attention layer" can be an SFB without FreqPass.
[0053] Figure 5 FIG. is a schematic structural diagram showing an optical flow module (Flow) according to an exemplary embodiment of the present disclosure. Referring to Figure 5 , the optical flow features can be calculated through three adjacent frames of images. Among them, the first frame and the second frame can be used as the first set of optical flow features; the second frame and the third frame can be used as another set of features of the optical flow. The optical flow features can be extended in the forward direction and can be used as the input of the cross attention (Cross Attention) respectively. The cross attention module can adopt a window-based multihead self-attention mechanism (Window-based Multihead Cross Attention, WMCA) for feature extraction.
[0054] Specifically, first, a self-attention calculation can be performed on the forward features (for example, such as the WMSA module in Figure 5 ), and the result can be used as the query (Q) in the cross attention module. Then, the backward features can also be used to calculate the key (K) and value (V) of the cross attention. Through the above steps, the feature result of the cross attention can be finally calculated.
[0055] Next, the feature result of this cross attention can be input into the GlobalCorrelation Module for the prediction of the optical flow. The global correlation can be calculated by dot product similarity and normalized using the Softmax function. In addition, in order to ensure the consistency of the optical flow feature information between different scales (feature layers) and promote the propagation accuracy from low-resolution optical flow features to high-resolution optical flow features, the deep optical flow features and the optical flow features of the current layer can be further calculated by cross attention (for example, the Global Attention module in Figure 5 ). In this way, by setting the decoding structure integrating the optical flow characteristics, the spatio-temporal continuity of the blood vessels in the image sequence can be ensured.
[0056] As described above, in the decoder, the input of each layer of ResDe can be composed of the output of the previous layer of ResDe and the skip connection between the same layers. When applying the skip connection, the output of the encoder module SFB of the same layer and the output of the Flow module of the same layer can be fed into ResDe. Figure 6 is a schematic structural diagram showing the residual convolutional module ResDe according to an exemplary embodiment of the present disclosure. Refer to Figure 6 , ResDe can include an upsampling layer and two residual blocks with the same operations. Among them, the residual block can include convolution (conv), batch normalization (BN), and rectified linear unit (ReLU). Exemplarily, the upsampling layer can perform upsampling using content-aware feature recombination (Carafe). Compared with the traditional bilinear upsampling method, Carafe upsampling has a larger receptive field and can better retain the semantic information of the feature map.
[0057] It should be noted that after passing through the double-attention module and multiple residual convolutional modules ResDe, five output features with different resolutions can be obtained. On the premise of keeping the number of channels unchanged, the resolutions can be extended to the same scale (h / 4×w / 4). These five output features with different resolutions can be input into the channel attention module (CA) to obtain attention features with channel weights. Figure 7 is a schematic structural diagram showing the CA module according to an exemplary embodiment of the present disclosure. Refer to Figure 7 , the CA module can extract information of different channels through global average pooling (GAP) operation and global max pooling (GMP) operation. This mechanism allows the network to learn global information, can selectively strengthen important features, and can also suppress less useful information. In this way, the expression ability of the feature channels can be enhanced through the CA module, thereby improving the representation ability of the network.
[0058] It should be noted that, referring to Figure 2 , the prediction results of multiple MSFUnet networks can be input into the LSTM network module in the order of frames. Further, the prediction results of the MSFUnet network and the optical flow features output by the last optical flow module among the multiple optical flow modules included in the MSFUnet network can be fed into the LSTM network module at the same time. And its size can be: [n, 1, 4, 512, 512], where n can be the number of frames of the sequence. The LSTM network module can be composed of a convolutional layer, 4 LSTM layers, and a linear layer.
[0059] Specifically, the convolutional layer can be a point convolution operation, with the number of input channels being 4, the number of output channels being 2, the convolutional kernel being 1, and the stride being 1. Then, the result output by the convolutional layer can be size-transformed. Exemplarily, it can be deformed from [n, 1, 2, 512, 512] to [1, n, 512 512 2], and it can be sent to the LSTM layer. Additionally, the number of hidden layers of the LSTM layer can be set to 4, and the hidden layer size can be 1024. Next, the output result of the LSTM layer can pass through a linear layer and can be re-deformed to [1, 2, 512, 512]. In this way, the segmentation results from different frameworks can be further fused through the LSTM network module, and then a comprehensive abdominal aortic vascular tree can be generated.
[0060] It should be noted that due to the influence of the contrast agent perfusion rate, blood vessel shape, and imaging principle on the imaging effect, DSA images often have uneven blood vessel imaging and blurred blood vessel tissue edges. To solve this problem, a FreqPass module with Hamming filtering can be used to improve the uniformity within the blood vessel class and the difference at the edges between classes such as blood vessel tissue.
[0061] According to an exemplary embodiment of the present disclosure, each encoder module (SFB) can include a downsampling module (FreqPass). Figure 8 FIG. is a schematic structural diagram of a downsampling module (FreqPass) according to an exemplary embodiment of the present disclosure.
[0062] For the first encoder module, the downsampling features of the average pooling branch (avg pool) and the max pooling branch (max pool) of the image group can be obtained through the downsampling module included in the encoder module, and the downsampling features of the average pooling branch and the max pooling branch can be concatenated to obtain an output result. Further, the concatenation result of the downsampling features can be operated on using point convolution to obtain an encoded result, that is, the final downsampling feature Xdown (with a size of [b, 2c, h / 2, w / 2]) can be obtained.
[0063] For each encoder module among multiple encoder modules except the first encoder module, the average pooling branch's downsampled features and the max pooling branch's downsampled features of the output result of the previous encoder module can be obtained through the downsampling module included in this encoder module, and the average pooling branch's downsampled features and the max pooling branch's downsampled features can be concatenated to obtain the output result. In this way, the output result of the last encoder module among multiple encoder modules can be used as the encoding result for inputting into the double-attention module.
[0064] It should be noted that the aforementioned "average pooling branch's downsampled features" can be obtained in the following way:
[0065] For this image group or the output result of the previous encoder module, for example, for the input feature X (with size [b, c, h, w]), an average pooling operation can be performed to obtain the average pooling result. Then, a convolution operation, a Pixel Shuffle operation, and a Softmax operation can be sequentially performed on this average pooling result, and thus features with size [b, hw / 4, 3, 3] can be obtained. Figure X Next, a Hamming window dot multiplication operation can be used to obtain the first convolution kernel Xw (with size [b, 9, h / 2, w / 2]). Then, the first convolution kernel Xw can be convolved with the average pooling result to obtain the average pooling branch's downsampled features.
[0066] The aforementioned "max pooling branch's downsampled features" can be obtained in the following way:
[0067] For this image group or the output result of the previous encoder module, for example, for the input feature X (with size [b, c, h, w]), a max pool operation can be performed to obtain the max pooling result. Then, a convolution operation, a Pixel Shuffle operation, and a Softmax operation can be sequentially performed on this max pooling result to obtain features with size [b, hw / 4, 3, 3]. Figure X Next, a Hamming window dot multiplication operation can be used to obtain the second convolution kernel Xw (with size [b, 9, h / 2, w / 2]). Then, the second convolution kernel Xw can be convolved with the max pooling result to obtain the max pooling branch's downsampled features. In this way, by using the feature extraction layer with the FreqPass module, the uniformity of blood vessels in intra-class images and the inter-class differences can be effectively improved.
[0068] In step 104, multiple first blood vessel recognition results corresponding to multiple MSFUnet networks can be input into the LSTM network to obtain the final blood vessel recognition result.
[0069] In step 105, a first loss can be calculated based on the first blood vessel recognition result output by each MSFUnet network and the first image label bound to the image group corresponding to the MSFUnet network. The first image label can be an image label including the blood vessel morphology annotated based on at least two blood vessel images included in the image group, that is, the first image label can be an image label including the blood vessel morphology annotated by a doctor based on at least two blood vessel images included in the image group.
[0070] Exemplarily, referring to Figure 2 , a total of 3 MSFUnet networks are shown. Taking the first MSFUnet network as an example. The image group corresponding to the first MSFUnet network includes a total of 3 blood vessel images, and the doctor can annotate the blood vessel morphology based on these 3 blood vessel images according to his own experience as the first image label bound to the image group corresponding to the first MSFUnet network.
[0071] In step 106, a second loss can be calculated based on the final blood vessel recognition result output by the LSTM network and the second image label bound to multiple blood vessel images, where the second image label can be an image label including the blood vessel morphology annotated based on multiple blood vessel images.
[0072] Specifically, the second loss can be a sequence-level loss (pre_sequence, label_sequence) calculated by the difference between the result after fusing the single-frame prediction image output by the MSFUnet network through the LSTM network module and the sequence label, and this loss can be calculated using Focal Loss. The aforementioned second image label can be an image label including the blood vessel morphology annotated by a doctor based on multiple blood vessel images, that is, the doctor can annotate the blood vessel morphology based on the multiple blood vessel images initially input into the blood vessel recognition model as the second image label.
[0073] In step 107, the blood vessel recognition model can be trained by adjusting the parameters of the MSFUnet network based on the first loss and adjusting the parameters of the LSTM network based on the second loss.
[0074] According to an exemplary embodiment of the present disclosure, for each MSFUnet network, a third loss can also be calculated based on the optical flow result output by the optical flow module (flow) of the last layer among the multiple optical flow modules (flow) included in the MSFUnet network and the first image label. Exemplarily, referring toFigure 2 , the third loss can be calculated based on the optical flow result output by the last optical flow module (flow) among the five optical flow modules (flow) included in the MSFUnet network and the first image label. Next, the parameters of the MSFUnet network can be adjusted based on both the aforementioned first loss and the third loss.
[0075] It should be noted that the sum of the first loss and the third loss can also be referred to as the frame-level loss, that is, the frame-level loss can be obtained by calculating the first image label and the prediction results (pre_MSF, label_frame) of the MSFUnet model and the prediction results (pre_flow, label_frame) of the optical flow module and adding them together. And both (pre_MSF, label_frame) and (pre_flow, label_frame) can be calculated using Focal Loss and weighted Dice Loss.
[0076] Specifically, Focal Loss can be calculated by the following formula:
[0077]
[0078]
[0079] where M is the number of classes, is the value after one-hot encoding of the sample class, is the predicted probability belonging to class c, is 0.25, and γ is 2.
[0080] Weighted Dice Loss, that is, can be calculated by the following formula:
[0081]
[0082]
[0083]
[0084] where M is the number of classes, n is the number of regions within a single frame, B represents the boundary length of the target, S represents the size of the target, X is the set of predicted values, and Y is the set of true values. In this way, the use of the hybrid loss function can significantly improve the segmentation accuracy of blood vessels.
[0085] It should be noted that the dataset adopted in this disclosure can include abdominal DSA data of 93 patients. The collected cases can be randomly divided into 3 groups: data of 60 patients can be assigned to the training set; data of 15 patients can be assigned to the validation set; data of 18 patients can be assigned to the test set. Moreover, all patient data can be exported as images with a resolution of 1024 pixels. Among them, the training set can include 2588 images, the validation set can include 576 images, and the test set can include 475 images. The network proposed in this disclosure can be but is not limited to: Python3.11.8. And all experiments can be carried out under the PyTorch framework, and an NVIDIA V100 Tensor Core GPU can also be used.
[0086] In addition, in this disclosure, the segmentation results of blood vessels can be evaluated based on the average Dice coefficient and the average HD95. Among them, the average Dice coefficient can be expressed as follows:
[0087]
[0088]
[0089] where X is the set of predicted values, Y is the set of true values, and n is the number of samples.
[0090] HD can be defined as follows:
[0091]
[0092] where 95%HD (HD95) is similar to the maximum HD, but it is based on the 95th percentile of the distances between the boundary points of X and Y. The average HD95 can be defined as follows:
[0093]
[0094] where n is the number of samples.
[0095] It should be noted that in this disclosure, the proposed network can be trained using the Adam optimizer, the initial learning rate can be 0.001, and the aforementioned hybrid loss function can be used as the training objective. To evaluate the performance of the model, Mean Dice and Mean HD95 can be used as metrics. The experimental results can be shown in Table 1:
[0096]
[0097] Table 1 Experimental Results
[0098] The test results show that the method provided by the present disclosure has achieved superior performance in the segmentation results of DSA images (see Table 1 for details); moreover, the experimental results show that the method provided by the present disclosure performs excellently in both single-frame and sequence metrics. Specifically, the method provided by the present disclosure has obtained the highest scores in the single-frame average Dice and average HD95 evaluations. In addition, the method provided by the present disclosure also provides the best performance in the sequence mean Dice evaluation and has obtained significantly favorable results in the sequence mean HD95 metric.
[0099] Figure 9 is a schematic comparison diagram showing the single-frame prediction segmentation results of different vascular regions according to an exemplary embodiment of the present disclosure and the vascular fusion segmentation results obtained from the complete sequence. Referring to Figure 9 , the first row of this figure shows four frames uniformly extracted from the DSA image sequence of a patient, specifically the 11th frame, the 21st frame, the 31st frame, and the 41st frame. The last image shows that the 24th frame is selected as a representative intermediate sequence frame for direct comparison and visualization of the results.
[0100] From the single-frame segmentation results (Frame 1 and Frame 2) and the sequence fusion results (Sequence), it can be observed that the method provided by the present disclosure has achieved excellent performance in segmenting branched blood vessels, especially in the celiac trunk branches, superior mesenteric artery, and renal artery. In the segmentation results of Frame 1 and Frame 2, there is an overlap between the superior mesenteric artery and the abdominal aorta, and only ScaleFormer and the method provided by the present disclosure have successfully distinguished these two structures, thus providing good segmentation results. For the segmentation of aneurysms and common iliac arteries, most methods have shown high accuracy. However, in the results of Frame 4, Swin-Unet and DAEFormer show obvious problems of insufficient segmentation. Regarding the segmentation of the right iliac artery, Swin-Unet shows obvious missegmentation in both Frame 3 and Frame 4. For the left iliac artery, only ScaleFormer and the method provided by the present disclosure have achieved satisfactory segmentation results in Frame 2, Frame 3, Frame 4, and the sequence results (Sequence). Further visualization of the results superimposed on the DSA images shows that the method provided by the present disclosure has optimal segmentation performance, which is particularly obvious in the case of complex branches and blurred vascular boundaries, that is, the method provided by the present disclosure effectively improves the integrity and accuracy of vascular segmentation.
[0101] The present disclosure provides a multi-frame fusion algorithm L-MSFUnet for abdominal aortic vessel segmentation based on DSA images. Compared with existing methods, while segmenting single-frame vessels, the present disclosure can effectively extract the complete vascular tree structure in the sequence images, thereby effectively improving the integrity and accuracy of vessel segmentation.
[0102] Specifically, by introducing temporal frame encoding, the feature continuity across multiple frames can be enhanced; by adopting a feature extraction layer with a FreqPass module, the uniformity of vessels within the class and the inter-class difference in the images can be effectively improved; by integrating a decoding structure with optical flow characteristics, the spatio-temporal continuity of vessels in the image sequence can be ensured; by further fusing the segmentation results from different frames through an LSTM module, a comprehensive abdominal aortic vascular tree can be generated; the use of a hybrid channel attention module and a hybrid loss function significantly improves the segmentation accuracy. It can be seen that the vascular recognition model provided by the present disclosure can improve the accuracy of vessel extraction and provide more reliable decision-making support for clinical treatment.
[0103] Figure 10 FIG. is a flowchart showing a vascular recognition method according to an exemplary embodiment of the present disclosure, and the vascular recognition method can be implemented based on a vascular recognition model trained by the training method according to the present disclosure.
[0104] Referring to Figure 10 , in step 1001, multiple target vascular images to be recognized can be obtained, where the multiple target vascular images to be recognized can be images obtained by injecting a contrast agent into the blood vessels of an organism.
[0105] In step 1002, the multiple target vascular images to be recognized can be divided into multiple groups of images to be recognized, where the multiple groups of images to be recognized can correspond to multiple MSFUnet networks one by one, and each group of images to be recognized can include at least two target vascular images to be recognized.
[0106] It should be noted that the input images can also be divided into patches of size 5, and the division step size can be 4. Through this overlapping division operation, the edge effect of the image can be reduced. In addition, when encoding, in addition to performing positional encoding within the frame, the temporal information of the images can also be encoded according to the input order of the vascular images.
[0107] In step 1003, each group of images to be recognized can be input into the MSFUnet network corresponding to the group of images to be recognized, and a first vascular recognition target result output by the MSFUnet network can be obtained.
[0108] In step 1004, multiple first blood vessel recognition target results corresponding to multiple MSFUnet networks can be input into the LSTM network to obtain the final blood vessel recognition target result.
[0109] Figure 11 FIG. is a block diagram showing a training device 1100 of a blood vessel recognition model according to an exemplary embodiment of the present disclosure. The blood vessel recognition model may include an LSTM network and multiple MSFUnet networks.
[0110] Referring to Figure 11 , the training device 1100 of the blood vessel recognition model may include a blood vessel image acquisition module 1101, an image group division module 1102, a first blood vessel recognition result acquisition module 1103, a final blood vessel recognition result acquisition module 1104, a first loss calculation module 1105, a second loss calculation module 1106, and a training module 1107.
[0111] The blood vessel image acquisition module 1101 may acquire multiple blood vessel images, where the multiple blood vessel images may be images obtained by injecting a contrast agent into the blood vessels of an organism.
[0112] The image group division module 1102 may divide the multiple blood vessel images into multiple image groups, where the multiple image groups may correspond one-to-one to multiple MSFUnet networks, and each image group may include at least two blood vessel images.
[0113] The first blood vessel recognition result acquisition module 1103 may input each image group into the MSFUnet network corresponding to the image group to obtain the first blood vessel recognition result output by the MSFUnet network.
[0114] According to an exemplary embodiment of the present disclosure, the MSFUnet network may include a double-attention module, a channel attention module (CA), multiple encoder modules (SpatialFrequency Block, SFB), multiple residual convolution modules (resdecoder block, ResDe), and multiple optical flow modules (flow attention).
[0115] The first blood vessel recognition result acquisition module 1103 may encode the image group through multiple encoder modules to obtain an encoding result. Then, the first blood vessel recognition result acquisition module 1103 may input the encoding result into the double-attention module to obtain a first attention result. Next, the first blood vessel recognition result acquisition module 1103 may perform optical flow processing on the first attention result through multiple optical flow modules to obtain an optical flow result.
[0116] Then, for the first residual convolution module among the multiple residual convolution modules, the first vascular recognition result acquisition module 1103 can obtain the residual convolution result of this residual convolution module based on the first attention result, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module; for each of the second to nth residual convolution modules among the multiple residual convolution modules, the first vascular recognition result acquisition module 1103 can obtain the residual convolution result of this residual convolution module based on the output of the previous-layer residual convolution module of this residual convolution module, the output of the encoder module at the same layer as this residual convolution module, and the output of the optical flow module at the same layer as this residual convolution module, where n can be an integer greater than or equal to 2.
[0117] Next, the first vascular recognition result acquisition module 1103 can obtain a second attention result through the channel attention module based on the first attention result and the residual convolution result of each residual convolution module in the multiple residual convolution modules. Then, the first vascular recognition result acquisition module 1103 can obtain the first vascular recognition result of the MSFUnet network based on the residual convolution result output by the residual convolution module located at the last layer in the multiple residual convolution modules and the second attention result. Specifically, the first vascular recognition result acquisition module 1103 can perform a Kronecker product operation on the residual convolution result output by the residual convolution module at the last layer and the second attention result, and thus can obtain the first vascular recognition result of the MSFUnet network.
[0118] According to an exemplary embodiment of the present disclosure, each encoder module (SFB) can include a downsampling module (FreqPass).
[0119] For the first encoder module, the first vascular recognition result acquisition module 1103 can obtain the downsampling features of the average pooling branch (avg pool) and the downsampling features of the max pooling branch (max pool) of this image group through the downsampling module included in this encoder module, and can concatenate the downsampling features of the average pooling branch and the downsampling features of the max pooling branch to obtain an output result. Further, an operation can also be performed on the concatenation result of the downsampling features using pointwise convolution to obtain an encoding result, that is, to obtain the final downsampling feature Xdown (with a size of [b, 2c, h / 2, w / 2]).
[0120] For each encoder module among multiple encoder modules except the first encoder module, the first vascular recognition result acquisition module 1103 can obtain the downsampled features of the average pooling branch and the downsampled features of the max pooling branch of the output result of the previous encoder module through the downsampling module included in this encoder module, and can splice the downsampled features of the average pooling branch and the downsampled features of the max pooling branch to obtain the output result. In this way, the output result of the last encoder module among multiple encoder modules can be used as the encoding result for inputting into the double-attention module.
[0121] It should be noted that the aforementioned "downsampled features of the average pooling branch" can be obtained in the following manner:
[0122] Perform average pooling operation on the image group or the output result of the previous encoder module. For example, perform average pooling operation on the input feature X (with size [b, c, h, w]) to obtain the average pooling result. Then, perform convolution operation, pixel rearrangement operation, and Softmax operation on this average pooling result in sequence, and then a feature with size [b, hw / 4, 3, 3] can be obtained. Figure X s. Next, perform Hamming window dot multiplication operation to obtain the first convolution kernel Xw (with size [b, 9, h / 2, w / 2]). Then, perform convolution on the first convolution kernel Xw and the average pooling result to obtain the downsampled features of the average pooling branch.
[0123] The aforementioned "downsampled features of the max pooling branch" can be obtained in the following manner:
[0124] Perform max pooling operation on the image group or the output result of the previous encoder module. For example, perform max pooling operation on the input feature X (with size [b, c, h, w]) to obtain the max pooling result. Then, perform convolution operation, pixel rearrangement operation, and Softmax operation on this max pooling result in sequence to obtain a feature with size [b, hw / 4, 3, 3]. Figure X s. Next, perform Hamming window dot multiplication operation to obtain the second convolution kernel Xw (with size [b, 9, h / 2, w / 2]). Then, perform convolution on the second convolution kernel Xw and the max pooling result to obtain the downsampled features of the max pooling branch. In this way, by using the feature extraction layer with the FreqPass module, the uniformity of blood vessels in intra-class images and the inter-class differences can be effectively improved.
[0125] The vascular recognition final result acquisition module 1104 can input multiple first vascular recognition results corresponding to multiple MSFUnet networks into the LSTM network to obtain the vascular recognition final result.
[0126] The first loss calculation module 1105 can calculate a first loss based on the first blood vessel recognition result output by each MSFUnet network and the first image label bound to the image group corresponding to the MSFUnet network. The first image label can be an image label including blood vessel morphology annotated based on at least two blood vessel images included in the image group, that is, the first image label can be an image label including blood vessel morphology annotated by a doctor based on at least two blood vessel images included in the image group.
[0127] The second loss calculation module 1106 can calculate a second loss based on the final blood vessel recognition result output by the LSTM network and the second image label bound to multiple blood vessel images, where the second image label can be an image label including blood vessel morphology annotated based on multiple blood vessel images.
[0128] Specifically, the second loss can be a sequence-level loss (pre_sequence, label_sequence) calculated by the difference between the result after fusing the single-frame prediction image output by the MSFUnet network through the LSTM network module and the sequence label, and this loss can be calculated using Focal Loss. The aforementioned second image label can be an image label including blood vessel morphology annotated by a doctor based on multiple blood vessel images, that is, the doctor can annotate the blood vessel morphology based on his own experience on multiple blood vessel images initially input to the blood vessel recognition model as the second image label.
[0129] The training module 1107 can train the blood vessel recognition model by adjusting the parameters of the MSFUnet network based on the first loss and adjusting the parameters of the LSTM network based on the second loss.
[0130] According to an exemplary embodiment of the present disclosure, the training device 1100 of the above blood vessel recognition model may further include a third loss calculation module. For each MSFUnet network, the third loss calculation module can also calculate a third loss based on the optical flow result output by the optical flow module at the last layer among the multiple optical flow modules (flow) included in the MSFUnet network and the first image label. Next, the training module 1107 can adjust the parameters of the MSFUnet network based on both the aforementioned first loss and the third loss.
[0131] Figure 12 FIG. is a block diagram showing a blood vessel recognition device 1200 according to an exemplary embodiment of the present disclosure. The blood vessel recognition device 1200 can be implemented based on the blood vessel recognition model trained according to the training method of the present disclosure.
[0132] Refer to Figure 12, the vascular recognition device 1200 may include a target vascular image acquisition module 1201, a to-be-recognized image group division module 1202, a first vascular recognition target result acquisition module 1203, and a final vascular recognition target result acquisition module 1204.
[0133] The target vascular image acquisition module 1201 can acquire multiple target vascular images to be recognized. Among them, the multiple target vascular images to be recognized can be images obtained by injecting a contrast agent into the blood vessels of an organism.
[0134] The to-be-recognized image group division module 1202 can divide the multiple target vascular images to be recognized into multiple to-be-recognized image groups. Among them, the multiple to-be-recognized image groups can correspond to multiple MSFUnet networks one by one, and each to-be-recognized image group can contain at least two target vascular images to be recognized.
[0135] It should be noted that the input image can also be divided into patches of size 5, and the division step size can be 4. Through this overlapping division operation, the edge effect of the image can be reduced. In addition, when encoding, in addition to performing position encoding within the frame, the time information of the image can also be encoded according to the input order of the vascular images.
[0136] The first vascular recognition target result acquisition module 1203 can input each to-be-recognized image group into the MSFUnet network corresponding to the to-be-recognized image group, and obtain the first vascular recognition target result output by the MSFUnet network.
[0137] The final vascular recognition target result acquisition module 1204 can input the multiple first vascular recognition target results corresponding to the multiple MSFUnet networks one by one into the LSTM network, and obtain the final vascular recognition target result.
[0138] Figure 13 is a block diagram showing an electronic device 1300 according to an exemplary embodiment of the present disclosure.
[0139] Refer to Figure 13 , the electronic device 1300 includes at least one memory 1301 and at least one processor 1302. Instructions are stored in the at least one memory 1301. When the instructions are executed by the at least one processor 1302, a training method or a recognition method of a vascular recognition model according to an exemplary embodiment of the present disclosure is executed.
[0140] As an example, the electronic device 1300 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 1300 does not have to be a single electronic device, but can also be a collection of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 1300 can also be a part of an integrated control system or system manager, or can be configured as a portable electronic device that interfaces with a local or remote device (e.g., via wireless transmission).
[0141] In the electronic device 1300, the processor 1302 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0142] The processor 1302 can run instructions or code stored in the memory 1301, where the memory 1301 can also store data. The instructions and data can also be sent and received over a network via a network interface device, where the network interface device can use any known transmission protocol.
[0143] The memory 1301 can be integrated with the processor 1302, for example, by arranging RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 1301 can include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 1301 and the processor 1302 can be operatively coupled or can communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 1302 can read files stored in the memory.
[0144] In addition, the electronic device 1300 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 1300 can be connected to each other via a bus and / or a network.
[0145] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the above-mentioned training method or blood vessel recognition method of the blood vessel recognition model. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0146] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program that, when executed by a processor, implements the training method or blood vessel recognition method of the blood vessel recognition model according to the present disclosure.
[0147] According to the training method, blood vessel recognition method, device, electronic device, storage medium, and computer program product of the blood vessel recognition model according to the present disclosure, by using the multi-frame fusion algorithm L-MSFUnet, that is, by using multiple angiography images for blood vessel recognition, more data volume and richer data content are considered, and thus a more comprehensive abdominal aortic vascular tree can be generated, that is, the accuracy of blood vessel recognition can be effectively improved, and thus the occurrence of medical accidents can be effectively avoided.
[0148] According to an exemplary embodiment of the present disclosure, by introducing temporal frame encoding, the feature continuity across multiple frames can be enhanced.
[0149] According to an exemplary embodiment of the present disclosure, by setting a decoding structure integrating optical flow characteristics, the spatio-temporal continuity of blood vessels in an image sequence can be ensured.
[0150] According to an exemplary embodiment of the present disclosure, the expression ability of feature channels can be enhanced through a CA module, thereby improving the representation ability of the network.
[0151] According to an exemplary embodiment of the present disclosure, the segmentation results from different frameworks can be further fused through an LSTM network module, and then a comprehensive abdominal aortic vascular tree can be generated.
[0152] According to an exemplary embodiment of the present disclosure, by adopting a feature extraction layer with a FreqPass module, the uniformity of blood vessels in intra-class images and the inter-class differences can be effectively improved.
[0153] According to an exemplary embodiment of the present disclosure, by using a hybrid loss function, the segmentation accuracy of blood vessels can be significantly improved.
[0154] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0155] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for training a blood vessel recognition model, characterized in that: The blood vessel recognition model includes an LSTM network and multiple MSFUnet networks, and the training method includes: Acquiring a plurality of vascular images, wherein the plurality of vascular images are images obtained by injecting a contrast agent into a blood vessel of a living body; Dividing the plurality of vascular images into a plurality of image groups, wherein the plurality of image groups correspond one-to-one to the plurality of MSFUnet networks, and each image group contains at least two of the vascular images; Input each image group into the MSFUnet network corresponding to the image group to obtain a first blood vessel recognition result output by the MSFUnet network; Inputting a plurality of first blood vessel recognition results corresponding to the plurality of MSFUnet networks one by one into the LSTM network to obtain a final blood vessel recognition result; The MSFUnet network includes a dual attention module, a channel attention module, multiple encoder modules, multiple residual convolution modules and multiple optical flow modules; The step of inputting each image group into the MSFUnet network corresponding to the image group to obtain a first blood vessel recognition result output by the MSFUnet network includes: Encoding the image group through the multiple encoder modules to obtain an encoding result; Inputting the encoding result into the dual attention module to obtain a first attention result; Performing optical flow processing on the first attention result through the multiple optical flow modules to obtain an optical flow result; For a first residual convolution module among the multiple residual convolution modules, based on the first attention result, an output of an encoder module at the same layer as the residual convolution module, and an output of an optical flow module at the same layer as the residual convolution module, obtaining a residual convolution result of the residual convolution module; For each residual convolution module from the second to the nth of the multiple residual convolution modules, based on the output of the residual convolution module of the previous layer of the residual convolution module, the output of the encoder module of the same layer as the residual convolution module, and the output of the optical flow module of the same layer as the residual convolution module, obtain the residual convolution result of the residual convolution module, where n is an integer greater than or equal to 2; Obtaining, by the channel attention module, a second attention result based on the first attention result and the residual convolution result of each residual convolution module in the plurality of residual convolution modules; Based on the residual convolution result output by the residual convolution module located at the last layer among the multiple residual convolution modules and the second attention result, a first blood vessel recognition result of the MSFUnet network is obtained.
2. The training method according to claim 1, characterized in that: The training method further comprises: Calculate a first loss based on a first blood vessel recognition result output by each MSFUnet network and a first image label bound to an image group corresponding to the MSFUnet network, wherein the first image label is an image label containing blood vessel morphology annotated based on at least two blood vessel images included in the image group; Calculating a second loss based on the final result of blood vessel recognition output by the LSTM network and a second image label bound to the plurality of blood vessel images, wherein the second image label is an image label containing blood vessel morphology annotated based on the plurality of blood vessel images; The blood vessel recognition model is trained by adjusting the parameters of the MSFUnet network based on the first loss and adjusting the parameters of the LSTM network based on the second loss.
3. The training method according to claim 2, characterized in that: The training method further comprises: For each MSFUnet network, based on the optical flow result output by the optical flow module located at the last layer among the multiple optical flow modules included in the MSFUnet network and the first image label, calculate the third loss; The adjusting the parameters of the MSFUnet network based on the first loss comprises: By adjusting the parameters of the MSFUnet network based on both the first loss and the third loss.
4. The training method according to claim 1, characterized in that: Each encoder module contains a downsampling module; The step of encoding the image group by using the multiple encoder modules to obtain an encoding result includes: For the first encoder module, obtain the down-sampling features of the average pooling branch and the down-sampling features of the maximum pooling branch of the image group through the down-sampling module included in the encoder module, and splice the down-sampling features of the average pooling branch and the down-sampling features of the maximum pooling branch to obtain an output result; For each encoder module except the first encoder module among the multiple encoder modules, obtain the down-sampling features of the average pooling branch and the down-sampling features of the maximum pooling branch of the output result of the previous encoder module through the down-sampling module included in the encoder module, and splice the down-sampling features of the average pooling branch and the down-sampling features of the maximum pooling branch to obtain an output result; Using the output result of the last encoder module among the multiple encoder modules as the encoding result; The down-sampling features of the average pooling branch are obtained in the following way: Perform an average pooling operation on the image group or the output result of the previous encoder module to obtain an average pooling result; perform a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the average pooling result in sequence to obtain a first convolution kernel; convolve the first convolution kernel with the average pooling result to obtain a down-sampling feature of the average pooling branch; The down-sampling features of the maximum pooling branch are obtained in the following way: Perform a maximum pooling operation on the image group or the output result of the previous encoder module to obtain a maximum pooling result; perform a convolution operation, a pixel rearrangement operation, a Softmax operation, and a Hamming window dot multiplication operation on the maximum pooling result in sequence to obtain a second convolution kernel; convolve the second convolution kernel with the maximum pooling result to obtain a downsampling feature of the maximum pooling branch.
5. A blood vessel identification method, characterized in that: The blood vessel recognition method is implemented based on a blood vessel recognition model trained by a training method according to any one of claims 1 to 4, and the blood vessel recognition method includes: Acquire a plurality of target blood vessel images to be identified, wherein the plurality of target blood vessel images to be identified are images obtained by injecting contrast agents into blood vessels of a living body; Dividing the plurality of target blood vessel images to be identified into a plurality of image groups to be identified, wherein the plurality of image groups to be identified correspond one-to-one to the plurality of MSFUnet networks, and each image group to be identified includes at least two target blood vessel images to be identified; Input each image group to be identified into the MSFUnet network corresponding to the image group to be identified, and obtain the first blood vessel identification target result output by the MSFUnet network; A plurality of first blood vessel recognition target results corresponding one-to-one to the plurality of MSFUnet networks are input into the LSTM network to obtain a final blood vessel recognition target result.
6. A training device for a blood vessel recognition model, characterized in that: The blood vessel recognition model includes an LSTM network and multiple MSFUnet networks, and the training device includes: A blood vessel image acquisition module is configured to acquire a plurality of blood vessel images, wherein the plurality of blood vessel images are images obtained by injecting a contrast agent into a blood vessel of a living body; An image group division module is configured to divide the plurality of blood vessel images into a plurality of image groups, wherein the plurality of image groups correspond one-to-one to the plurality of MSFUnet networks, and each image group contains at least two of the blood vessel images; A first blood vessel recognition result acquisition module is configured to input each image group into the MSFUnet network corresponding to the image group, and obtain a first blood vessel recognition result output by the MSFUnet network; a blood vessel identification final result acquisition module, configured to input a plurality of first blood vessel identification results corresponding to the plurality of MSFUnet networks one by one into the LSTM network to obtain a blood vessel identification final result; The MSFUnet network includes a dual attention module, a channel attention module, multiple encoder modules, multiple residual convolution modules and multiple optical flow modules; The first blood vessel recognition result acquisition module is specifically configured as follows: Encoding the image group through the multiple encoder modules to obtain an encoding result; Inputting the encoding result into the dual attention module to obtain a first attention result; Performing optical flow processing on the first attention result through the multiple optical flow modules to obtain an optical flow result; For a first residual convolution module among the multiple residual convolution modules, based on the first attention result, an output of an encoder module at the same layer as the residual convolution module, and an output of an optical flow module at the same layer as the residual convolution module, obtaining a residual convolution result of the residual convolution module; For each residual convolution module from the second to the nth of the multiple residual convolution modules, based on the output of the residual convolution module of the previous layer of the residual convolution module, the output of the encoder module of the same layer as the residual convolution module, and the output of the optical flow module of the same layer as the residual convolution module, obtain the residual convolution result of the residual convolution module, where n is an integer greater than or equal to 2; Obtaining, by the channel attention module, a second attention result based on the first attention result and the residual convolution result of each residual convolution module in the plurality of residual convolution modules; Based on the residual convolution result output by the residual convolution module located at the last layer among the multiple residual convolution modules and the second attention result, a first blood vessel recognition result of the MSFUnet network is obtained.
7. A blood vessel identification device, characterized in that: The blood vessel recognition device is implemented based on a blood vessel recognition model trained by the training method according to any one of claims 1 to 4, and the blood vessel recognition device includes: a target blood vessel image acquisition module, configured to acquire a plurality of target blood vessel images to be identified, wherein the plurality of target blood vessel images to be identified are images obtained by injecting a contrast agent into a blood vessel of a living body; The to-be-recognized image group division module is configured to divide the plurality of to-be-recognized target blood vessel images into a plurality of to-be-recognized image groups, wherein the plurality of to-be-recognized image groups correspond one-to-one to the plurality of MSFUnet networks, and each to-be-recognized image group includes at least two of the to-be-recognized target blood vessel images; A first blood vessel recognition target result acquisition module is configured to input each to-be-recognized image group into the MSFUnet network corresponding to the to-be-recognized image group, and obtain a first blood vessel recognition target result output by the MSFUnet network; The blood vessel identification final target result acquisition module is configured to input a plurality of first blood vessel identification target results corresponding to the plurality of MSFUnet networks one by one into the LSTM network to obtain a blood vessel identification final target result.
8. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the blood vessel recognition model training method according to any one of claims 1 to 4, or to implement the blood vessel recognition method according to claim 5.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the blood vessel recognition model training method as described in any one of claims 1 to 4, or execute the blood vessel recognition method as described in claim 5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for training a blood vessel recognition model according to any one of claims 1 to 4, or implements the method for blood vessel recognition according to claim 5.
Citation Information
Patent Citations
Liver vessel ultrasound image target identification and tracking method based on improved U-net network and LSTM network
CN113763309A
Intracranial great vessel image processing method and system, electronic equipment and medium
CN117809122A