Lung infection region segmentation method and device, electronic equipment and storage medium
By employing multi-scale fusion and feature interaction methods, the problem of insufficient segmentation accuracy in lung infection regions was solved, achieving more accurate segmentation of lung infection regions and improving the accuracy and efficiency of diagnosis.
Patent Information
- Application Number
- CN202210789640.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-07-06
AI Technical Summary
The current technology lacks sufficient precision in segmenting the lung infection area, leading to false positives and false negatives, which affects diagnostic efficiency and accuracy.
By employing a multi-scale fusion and feature interaction approach, and combining feature extraction, multi-scale fusion, global information extraction, and feature complementarity modules, the segmentation accuracy of lung infection regions is improved.
It enables more precise segmentation of lung infection areas, improving diagnostic accuracy and efficiency, and reducing false positives and false negatives.
Smart Images

Figure CN115345890B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of medical image processing, and particularly relates to a lung infection area segmentation method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the increasing number of confirmed COVID-19 patients, the number of patient lung CT data is also gradually increasing. Artificially labeling the infection area is a tedious and repetitive task that requires a high level of professional knowledge and clinical experience. Therefore, designing an automatic segmentation model for COVID-19 infection area can improve the efficiency of labeling and lay a good foundation for further analysis and diagnosis. Computer-aided systems (CAD) can assist doctors in lung diagnosis through anomaly detection algorithms, thereby greatly improving the efficiency of the examination. However, lung infection segmentation is a difficult task, and the infection area and other areas are not obvious in color and texture, and the infection boundary is blurred, which can lead to low segmentation accuracy and even false detection of the lesion area, missing the best treatment time. SUMMARY
[0003] The present disclosure provides a lung infection area segmentation method and device, electronic equipment and storage medium, which can improve the segmentation accuracy of the lung infection area.
[0004] According to an aspect of the present disclosure, a lung infection area segmentation method is provided, which includes:
[0005] performing feature extraction processing on a lung image to obtain a plurality of high-level features;
[0006] performing multi-scale fusion processing on the high-level features to obtain high-level enhanced features;
[0007] performing global information extraction on the high-level enhanced features to obtain global guidance features;
[0008] performing multi-layer feature interaction processing on the global guidance features and the high-level features to obtain lung features, the lung features being used to represent a lung infection area of the lung image.
[0009] In some possible implementations, the performing multi-scale fusion processing on the high-level features to obtain high-level enhanced features includes:
[0010] performing multi-scale fusion processing on each of the high-level features to obtain multi-scale fusion features;
[0011] performing channel splicing processing on the multi-scale fusion features to obtain the high-level enhanced features;
[0012] The performing multi-scale fusion processing on each of the high-level features to obtain multi-scale fusion features includes:
[0013] respectively performing feature processing of the high-level feature by multiple branches to obtain branch features;
[0014] performing fusion processing on the branch features to obtain multi-scale fusion features corresponding to the high-level feature.
[0015] In some possible implementation manners, the multiple branches include three branch groups, and the respectively performing feature processing of the high-level feature by multiple branches to obtain branch features includes:
[0016] performing first convolution processing on the high-level feature on all the branches to obtain corresponding first convolution features;
[0017] performing second convolution processing on each of the first convolution features on the second branch group to obtain second convolution features;
[0018] performing channel fusion on the first convolution features of the first branch group and each of the second convolution features of the second branch group to obtain third convolution features;
[0019] performing addition fusion on the third convolution features and the first convolution features of the third branch group to obtain the branch features.
[0020] In some possible implementation manners, the performing global information extraction on the high-level enhanced feature to obtain a global guidance feature includes:
[0021] performing feature remodeling on the high-level enhanced feature to obtain a remodeled feature;
[0022] adding the remodeled feature and the position encoding feature to obtain an encoding fusion feature;
[0023] performing global information extraction on the encoding fusion feature to obtain the global guidance feature.
[0024] In some possible implementation manners, the performing multi-layer feature interaction processing on the global guidance feature and the high-level feature to obtain a lung feature includes:
[0025] inputting the global guidance feature and the high-level feature into a feature complement module group, and performing feature interaction on the global guidance feature and the high-level feature by using feature complement modules in the feature complement module group in sequence to obtain the lung feature;
[0026] The feature complement module group includes multiple feature complement modules, input features of the feature complement modules include a global guidance feature, a corresponding high-level feature, and a refined feature output by a previous feature complement module, and the refined feature input by a first feature complement module is the global guidance feature.
[0027] In some possible implementation manners, the performing feature interaction on the global guidance feature and the high-level feature by using the feature complementary module in the feature complementary module set comprises:
[0028] obtaining a forward weight and a reverse weight by using the global guidance feature and the high-level feature;
[0029] performing multiplication processing on the refined feature output by the previous feature complementary module and the forward weight and the reverse weight respectively to obtain a forward correction feature and a reverse correction feature;
[0030] performing reinforcement processing on the refined feature by using the forward correction feature and the reverse correction feature respectively to obtain a forward reinforcement feature and a reverse reinforcement feature;
[0031] performing the reinforcement processing on the forward reinforcement feature and the reverse reinforcement feature to obtain the refined feature output by the current feature complementary module.
[0032] In some possible implementation manners, the obtaining a forward weight and a reverse weight by using the global guidance feature and the high-level feature comprises:
[0033] performing addition and activation processing on the global guidance feature and the high-level feature to obtain the forward weight;
[0034] performing negation on the forward weight to obtain the reverse weight;
[0035] and / or
[0036] the input feature of the reinforcement processing comprises a first input feature and a second input feature, and the performing reinforcement processing comprises:
[0037] performing multiplication processing on the first input feature and the second input feature to obtain a product feature;
[0038] performing addition fusion processing on the product feature and the first input feature and the second input feature respectively to obtain a first fusion feature and a second fusion feature respectively;
[0039] performing convolution processing on the first fusion feature and the second fusion feature respectively, and performing feature splicing in a channel dimension to obtain a third fusion feature;
[0040] performing local enhancement and global enhancement processing on the third fusion feature to obtain an output feature as a reinforcement feature or a refined feature.
[0041] According to a second aspect of the present disclosure, a lung infection segmentation device is provided, which comprises:
[0042] The feature extraction module is configured to perform feature extraction processing on the lung image to obtain a plurality of high-level features.
[0043] The fusion module is configured to perform multi-scale fusion on the high-level features to obtain high-level enhanced features.
[0044] The global information extraction module is configured to perform global information extraction on the high-level enhanced features to obtain global guidance features.
[0045] The feature complementation module is configured to perform multi-level feature interaction processing on the high-level features by using the global guidance features to obtain lung features, which are used to represent lung infection regions of the lung image.
[0046] According to a third aspect of the present disclosure, an electronic device is provided, which includes:
[0047] a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method according to any one of the first aspect.
[0048] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method according to any one of the first aspect.
[0049] In the embodiments of the present disclosure, the complementation of multi-content information of lung image information can be completed, the shallow features are discarded during encoding, the high-level features containing rich semantic information are aggregated, and the integrated semantic information is used to extract rich global information. In the decoding process, the global information is used as guidance, and multi-level feature fusion and interaction are performed, the connection between features is effectively established, and thus the lung infection region is more accurately segmented.
[0050] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure.
[0051] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0052] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the technical solutions of the present disclosure together with the specification.
[0053] Figure 1 a flowchart of a lung infection region segmentation method according to an embodiment of the present disclosure is shown;
[0054] Figure 2A structural schematic diagram of a segmentation network of a lung infection area according to an embodiment of the present disclosure is shown.
[0055] Figure 3 A method flowchart for obtaining advanced enhanced features according to an embodiment of the present disclosure is shown.
[0056] Figure 4 A structural schematic diagram of a fusion module RFB performing multi-scale fusion according to an embodiment of the present disclosure is shown.
[0057] Figure 5 A flowchart for performing interaction processing by a feature complementary module according to an embodiment of the present disclosure is shown.
[0058] Figure 6 A structural schematic diagram of a feature complementary module according to an embodiment of the present disclosure is shown.
[0059] Figure 7 A structural schematic diagram of an interaction enhancement module according to an embodiment of the present disclosure is shown.
[0060] Figure 8 A structural schematic diagram of a local-global attention module according to an embodiment of the present disclosure is shown.
[0061] Figure 9 A segmentation result for a COVID-19 dataset according to an embodiment of the present disclosure is shown.
[0062] Figure 10 A block diagram of a lung infection area segmentation device according to an embodiment of the present disclosure is shown.
[0063] Figure 11 A block diagram of an electronic device 800 according to an embodiment of the present disclosure is shown.
[0064] Figure 12 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION:
[0065] Various exemplary embodiments, features, and aspects of the present disclosure will be explained in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote like elements or components. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.
[0066] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0067] The term "and / or", as used herein, merely describes association between associated objects, and can indicate that three cases can exist, for example, A and / or B can indicate that the three cases of A alone, A and B together, and B alone can exist. In addition, the term "at least one of", as used herein, indicates any one of a plurality or any combination of at least two of a plurality, for example, includes at least one of A, B, and C can indicate any one or more elements selected from the set consisting of A, B, and C.
[0068] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main ideas of the present disclosure.
[0069] The execution subject of the lung infection area segmentation method provided by the present disclosure can be an image processing device, for example, the method can be executed by a terminal device or a server or other processing device, wherein the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the lung infection area segmentation method can be realized by a processor calling computer readable instructions stored in a memory.
[0070] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length, the present disclosure will not be described again.
[0071] Figure 1 A flowchart of a lung infection area segmentation method according to an embodiment of the present disclosure is shown as follows, Figure 1 As shown, the lung infection area segmentation method comprises:
[0072] S10: performing feature extraction processing on the lung image to obtain a plurality of high-level features;
[0073] In some possible implementation manners, the lung image can be a CT (Computed Tomography, Computed Tomography) image, and in other implementation manners, it can also be a medical image obtained by other imaging types. In addition, the lung image of the embodiment of the present disclosure is a two-dimensional image, but the present disclosure does not make specific limitation on this.
[0074] In some possible implementation manners, the feature extraction processing can be performed by using a feature extraction network, and the feature extraction network can include a pyramid network, a residual network, or the like. The feature extraction network can obtain low-level to high-level multi-level features respectively. The high-level features are selected for subsequent processing in the embodiments of the present disclosure to enrich the global information of the features.
[0075] S20: performing multi-scale fusion on the high-level features to obtain high-level enhanced features;
[0076] In some possible implementation manners, the multi-scale fusion processing can be performed on each high-level feature respectively to obtain high-level feature information of different scales, and then the multi-scale information is fused to obtain the high-level enhanced features.
[0077] S30: performing global information extraction on the high-level enhanced features to obtain global guidance features;
[0078] In some possible implementation manners, the global information extracted by the present disclosure can be used as prior guidance in the decoding process. The features can provide rich global context information, which can be used for preliminary positioning of the lung infection area. The subsequent decoding processing is performed by using the global guidance features, and the segmentation accuracy of the infection area can be improved.
[0079] S40: performing feature interaction processing on the global guidance features and the high-level features to obtain lung features, where the lung features are used to represent the lung infection area of the lung image.
[0080] In some possible implementation manners, the obtained global guidance features and high-level features are interactively fused, which can establish the connection between the global features and the high-level features, and the lung infection area can be segmented more accurately. In the embodiments of the present disclosure, the infection area to be segmented can be any focus area of pneumonia, for example, the present disclosure can be used to segment the infection area of COVID-19.
[0081] In the embodiments of the present disclosure, in order to overcome the problem of insufficient segmentation accuracy of the lung infection area, the present disclosure proposes a lung infection area segmentation method based on multi-content complementation. In the encoding process, the shallow features of the lung image are discarded, the high-level features containing rich semantic information are integrated, and the global guidance features of the infection area are established by using the integrated high-level features. In the decoding process, the connection between the global guidance features and the high-level features is established by using multi-level feature interaction, the features are effectively fused, and the lung infection area is segmented more accurately.
[0082] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The embodiments of the present disclosure can first obtain a lung image, where the lung image can be a two-dimensional lung CT image, and the manner of obtaining the lung image can include at least one of the following manners:
[0083] A) Directly use the CT device to collect the image of the lung; in the embodiment of the present disclosure, the CT device can be any manufacturer's device, and the present disclosure does not make specific limitations hereon.
[0084] B) Transmit and receive the lung image through the electronic device; the embodiment of the present disclosure can receive the lung image transmitted by other electronic devices through communication, which can include wired communication and / or wireless communication, and the present disclosure does not make specific limitations hereon.
[0085] C) Read the lung image stored in the database; the embodiment of the present disclosure can read the locally stored lung image or the lung image stored in the server according to the received data reading instruction, and the present disclosure does not make specific limitations hereon.
[0086] After obtaining the lung image, the segmentation processing of the infected area can be performed on the lung image. Figure 2 The structure diagram of the segmentation network of the infected area of the lung according to the embodiment of the present disclosure is shown. Specifically, the embodiment of the present disclosure can first perform feature extraction processing on the lung image to obtain multi-level features from low level to high level. Specifically, the embodiment of the present disclosure adopts a residual network (Res2Net deep convolutional neural network) as an encoder to extract image features, and when encoding, the last three layers of the Res2Net network can be extracted as high-level features R i , i = 3, 4, 5}. After obtaining the high-level features, multi-scale fusion processing can be performed on each high-level feature to obtain high-level enhanced features. i
[0087] Figure 3 The method flowchart for obtaining the high-level enhanced features according to the embodiment of the present disclosure is shown. The multi-scale fusion processing is performed on the high-level features to obtain the high-level enhanced features, which includes:
[0088] S21: Multi-scale fusion processing is performed on each high-level feature to obtain multi-scale fusion features;
[0089] S21: Channel splicing processing is performed on the multi-scale fusion features to obtain the high-level enhanced features.
[0090] In the embodiment of the present disclosure, feature fusion can be performed on high-level features including different information to obtain multi-scale fusion features. Further fusion processing of the multi-scale fusion features can obtain high-level enhanced features.
[0091] The multi-scale fusion processing is performed on each high-level feature to obtain multi-scale fusion features, which includes performing feature processing of multiple branches on each high-level feature to obtain branch features, and performing fusion processing on the branch features to obtain the multi-scale fusion features corresponding to the high-level features.
[0092] Figure 4 A structure diagram of a fusion module RFB performing multi-scale fusion according to an embodiment of the present disclosure is shown. Among them, for each high-level feature, a multi-scale feature fusion processing can be performed by adopting the fusion module RFB respectively, to obtain the corresponding multi-scale fusion feature. Specifically, a plurality of branch groups of feature processing can be performed on the high-level feature, wherein each branch group includes at least one feature processing branch, and the feature processing modes of different branch groups are different. The embodiment of the present disclosure can include three branch groups, wherein, Figure 4 The b1 branch is the first branch group, b2-b4 is the second branch group, and b5 is the third branch group. The above grouping manner is only exemplary, and other grouping manners can be adopted in other embodiments, and the number of branches in the branch group can be set according to the requirements.
[0093] Based on Figure 4 The embodiment of the present disclosure, the respective high-level features are processed by a plurality of branches to obtain branch features, including: performing first convolution processing on the high-level features on all branches to obtain corresponding first convolution features; performing second convolution processing on each of the first convolution features on the second branch group to obtain second convolution features; performing channel fusion on the first convolution features of the first branch group and each of the second convolution features of the second branch group to obtain third convolution features; and performing addition fusion on the third convolution features and the first convolution features of the third branch group to obtain the branch features.
[0094] In one example, each feature processing branch includes a convolution block performing first convolution processing, such as 1*1 convolution in the channel direction in the embodiment of the present disclosure, through which the first convolution feature with reduced channel number can be obtained, and the channel number C of each first convolution feature is the same (such as 64). The first convolution features of each branch in the first branch group and the second branch group are not subjected to subsequent feature processing, and are used for feature fusion in subsequent processes. For the feature processing branches in the second branch group, different dilation rates can be used to perform dilated convolution on the first convolution features, for example, the embodiment of the present disclosure can use dilation rates of 3, 5 and 7 to perform dilated convolution on the first convolution features, to obtain feature information with different receptive fields. The embodiment of the present disclosure performs three-layer convolution on the three branches b k (k=2, 3, 4) of the second branch group, which includes a convolution kernel of 1*(2k-1), a convolution kernel of (2k-1)*1, and a dilated convolution kernel of 3*3 with a dilation rate of 2k-1, to obtain the second convolution features of the three branches b k (k=2, 3, 4) of the second branch group. Then, the first convolution features of the first branch group and each of the second convolution features of the second branch group can be subjected to channel fusion to obtain third convolution features, specifically, the first convolution features of the first branch group and each of the second convolution features of the second branch group are concatenated to obtain the third convolution features.k The features obtained by the k = 1, 2, 3, 4 branches are spliced along the channel dimension and reduced to 64 by 1 × 1 convolution of the convolution kernel to obtain the third convolutional feature b c Then, the third convolutional feature and the first convolutional feature of the third branch group are added to obtain the branch feature. Specifically, b c The branch is added to the b5 branch, and the addition result is input to the ReLU activation function as a whole to obtain the multi-scale fusion feature corresponding to the high-level feature, i.e., the multi-scale feature {f3, f4, f5} of {R3, R4, R5}. Finally, the features {f3, f4, f5} are spliced in the channel to obtain the high-level enhanced feature F g Wherein, the dimension of the high-level enhanced feature is represented as C × H × W, C represents the number of channels, H represents the length, and W represents the width, wherein C = 192.
[0095] In the multi-scale feature fusion process in the embodiments of the present disclosure, multi-branch feature processing is adopted, and different scale feature information is extracted by using different expansion rates, thereby providing different receptive fields. The RFB simulates the receptive field of human vision to enhance the multi-scale feature extraction capability of the network, thereby providing high discriminative features and robust features, which is beneficial to improve the segmentation accuracy.
[0096] In the case of obtaining the high-level enhanced feature, the global guidance feature can be obtained by using the high-level enhanced feature to preliminarily obtain the feature information expressing the infection area. The embodiments of the present disclosure perform global information extraction on the high-level enhanced feature to obtain the global guidance feature, including: performing feature remodeling on the high-level enhanced feature to obtain a remodeled feature; adding the remodeled feature and a position encoding feature to obtain an encoding fusion feature; performing global information extraction on the encoding fusion feature to obtain the global guidance feature.
[0097] In one example, first, the high-level enhanced feature F g is linearly remodeled to obtain C one-dimensional features (C × HW). For example, the torch.view() function can be used to convert the high-level enhanced feature into a one-dimensional feature, and then a fully connected layer is used to project the number of C from 192 to 384 to obtain the remodeled feature after linear processing. Then, the remodeled feature is directly added to the position encoding PE to obtain the encoding fusion feature F' gwherein the position encoding can improve the perception of the model to the position information, the position encoding is a parameter newly learned by the network during the training process, and after the training is completed, the position encoding is an invariant tensor, and the number thereof is also C, and the size is 1xHW, which constitutes the dimension of CxHW. The encoded fusion feature is obtained by using the sum of the position encoding and the reshaped feature, and then the obtained encoded fusion feature is transmitted to the Transformer network for global information extraction. Finally, the global guidance feature R is obtained after outputting after multiple layers of the Transformer g . The above process can be represented by the following formula:
[0098]
[0099] R g =Transformer(F' g )
[0100] wherein FC(·) represents a full connection layer, R(·) represents a reshaping operation, PE represents a position encoding, represents addition, and Transformer(·) represents a Transformer layer. One layer of the Transformer includes a multi-head self-attention (MSA) and a multi-layer perceptron (MLP) sublayer, and layer normalization (LN) is inserted before the two sublayers, and residual connection is performed after the two sublayers. It is worth noting that the number of Transformer layers in the present disclosure is set to 4, and the number of multi-head self-attention is set to 6, but this is not a specific limitation of the present disclosure.
[0101] After obtaining the global guidance feature, the global guidance feature can be used to perform multi-layer feature interaction processing on the advanced feature, and then a lung feature is obtained, which is used to represent the lung infection area of the lung image. Specifically, the global guidance feature and the advanced feature can be input into a feature complementary module in the present disclosure, and the feature complementary module in the feature complementary module is used to sequentially perform feature interaction on the global guidance feature and the advanced feature, respectively, to obtain the lung image; wherein the feature complementary module includes a plurality of feature complementary modules, the input features of the feature complementary module include the global guidance feature, the corresponding advanced feature, and the refined feature output by the previous feature complementary module, and the refined feature input by the first feature complementary module is the global guidance feature.
[0102] Specifically, the number of feature complementary modules in the feature complementary module group can be the same as the number of high-level features obtained by the feature extraction process. For example, the embodiment of the present disclosure can include three feature complementary modules FCM, and each feature complementary module FCM is used to perform feature interaction between one high-level feature and the global guidance feature. By performing feature interaction with high-level features of different scales, the obtained features can be continuously refined, so that the features have more detailed feature information. The input of the first layer FCM of the present disclosure is the global guidance feature R g The input of the second layer FCM is the global guidance feature R g The input of the third layer FCM is the global guidance feature R g The input of the third layer FCM is the global guidance feature R
[0103] Figure 5 The flowchart for performing interaction processing by the feature complementary module in the embodiment of the present disclosure is shown, wherein the feature interaction between the global guidance feature and the high-level feature is performed by using the feature complementary module in the feature complementary module group, which includes:
[0104] S41: obtaining forward weight and reverse weight by using the global guidance feature and the high-level feature;
[0105] S42: performing multiplication processing on the refined feature output by the previous feature complementary module and the forward weight and reverse weight respectively to obtain forward correction feature and reverse correction feature;
[0106] S43: performing reinforcement processing on the refined feature by using the forward correction feature and the reverse correction feature respectively to obtain forward reinforcement feature and reverse reinforcement feature;
[0107] S44: performing the reinforcement processing on the forward reinforcement feature and the reverse reinforcement feature to obtain the refined feature output by the feature complementary module.
[0108] In some possible implementations, the global guidance feature and the high-level feature can be added, and the added result can be activated (such as sigmoid processing) to obtain the forward weight; and the forward weight is negated to obtain the reverse weight. The negation method is to subtract the forward weight by 1 to obtain the reverse weight. Specifically, Figure 6 The structural diagram of the feature complementary module according to the embodiment of the present disclosure is shown, and the global guidance feature R g The high-level feature Ri The addition is performed, and then a Sigmoid function is passed to obtain the forward weight The forward feature weight is then negated to obtain the reverse weight The above process can be described by the following formula:
[0109]
[0110]
[0111] wherein Sig(·) represents the Sigmoid function, and Θ(·) represents the reverse operation, that is, subtracting all inputs from all matrices, wherein all elements in the matrix are 1.
[0112] After obtaining the forward weight and the reverse weight, the refined features input to the feature complementary module are respectively subjected to reinforcement processing by using the forward weight and the reverse weight, wherein the refined features input to the first feature complementary module are global guidance features, and the refined features input to the remaining feature complementary modules are output features of the previous feature complementary module.
[0113] The process of performing feature reinforcement processing by the embodiments of the present disclosure can include: multiplying the refined features input to the feature complementary module with the forward weight and the reverse weight respectively to obtain forward correction features and reverse correction features The obtained correction features contain unique properties of the forward features and the reverse features; then, the refined features are respectively subjected to reinforcement processing by using the forward correction features and the reverse correction features to obtain forward reinforcement features and reverse reinforcement features The above-mentioned interaction processing can enable the forward features and the reverse features to be continuously interactively enhanced, thereby ensuring consistent learning of the forward reinforcement features and the reverse reinforcement features The above-mentioned entire process can be described by the following formula:
[0114]
[0115]
[0116]
[0117]
[0118]
[0119] wherein represents a multiplication operation, IEM(·) represents an interaction enhancement module, used to perform the enhancement processing.
[0120] Figure 7 A structural schematic diagram of an interaction enhancement module according to an embodiment of the present disclosure is shown. As shown in FIG. 7, the input of the interaction enhancement module IEM includes two features, and the present disclosure defines the input features as a first input feature and a second input feature. When different enhancement processing procedures are performed, the first input feature and the second input feature represent different feature information. For example, when a forward correction feature performs enhancement processing on the refined feature, the first input feature and the second input feature can be the forward correction feature and the refined feature input by the current feature complementary module, respectively. Correspondingly, the output feature is a forward enhanced feature. When a reverse correction feature performs enhancement processing on the refined feature, the first input feature and the second input feature can be the reverse correction feature and the refined feature input by the current feature complementary module, respectively. Correspondingly, the output feature is a reverse enhanced feature. In addition, when the forward enhanced feature and the reverse enhanced feature perform the enhancement processing, the first input feature and the second input feature are the forward enhanced feature and the reverse enhanced feature, respectively. At this time, the output feature is the refined feature output by the current feature complementary module.
[0121] The following will be described in combination with Figure 7The input features of the reinforcement processing include a first input feature and a second input feature, and the execution of the reinforcement processing includes: performing multiplication processing on the first input feature and the second input feature to obtain a product feature; performing addition fusion processing on the product feature and the first input feature and the second input feature respectively to obtain a first fusion feature and a second fusion feature respectively; performing convolution processing, such as 3*3 convolution, on the first fusion feature and the second fusion feature respectively, and then performing feature splicing in the channel dimension on the first fusion feature H1 and the second fusion feature H2 after the convolution processing to obtain a third fusion feature; performing local enhancement and global enhancement processing on the third fusion feature to aggregate local information and global information to enhance the representation ability of the feature, and then obtaining an output feature as a reinforcement feature or a refined feature. The local enhancement and global enhancement processing can be realized by a local-global attention module LGCA module. Before performing the enhancement processing, 3*3 convolution processing can be first performed on the third fusion feature, the enhancement processing is performed on the LGCA module after the convolution processing, and 3*3 convolution is performed on the enhanced feature, and finally the output feature as the forward / reverse reinforcement feature or the refined feature is obtained. It is worth noting that three IEMs are used in each FCM module in the embodiment of the present disclosure to process the features. For the sake of convenience, the following formula only expresses the processing process of the forward feature reinforcement, and the processing processes of the remaining two IEMs are consistent with the forward feature reinforcement process, except that the inputs are different. Specifically, the input of the IEM is and F i+1 The above process can be described by the following formula:
[0122]
[0123]
[0124]
[0125] wherein, represents an addition operation, Cat(·) represents a channel splicing operation, LGCA(·) represents a local-global attention module, Conv 3×3 represents 3*3 convolution processing.
[0126] The processing process of the local-global attention module is described below, Figure 8A structural diagram of a local-global attention module according to an embodiment of the present disclosure is shown. The local enhancement and global enhancement processing are performed on the third fusion feature to obtain an output feature as a strengthened feature or a refined feature, including: obtaining a local affinity attention matrix of the third fusion feature by using a local attention branch, and obtaining a local attention fusion feature of the third fusion feature based on the local affinity attention matrix; obtaining a global affinity attention matrix of the third fusion feature by using a global attention branch, and obtaining a global attention fusion feature of the third fusion feature based on the global affinity attention matrix; fusing the local affinity attention matrix and the global affinity attention matrix to obtain a fusion attention matrix, and obtaining a local-global attention fusion feature of the third fusion feature by using the fusion attention matrix; multiplying the local-global attention fusion feature with the local attention fusion feature and the global attention fusion feature respectively to obtain a first attention feature and a second attention feature, and obtaining the output feature of the LGCA module by using a fusion feature of the first attention feature and the second attention feature.
[0127] In the local-global attention module, the present disclosure aims to enhance the representation ability of the feature from the perspective of local self-attention and global self-attention, and the specific structure of the LGCA is shown in 8. It is assumed that the feature input into the LGCA (such as the third fusion feature) is denoted as X ∈ Z B×C×H×W , where B is the batch size, for the local self-attention branch A, first input into the graph convolutional network GCN after 1 × 1 convolution to perform node feature aggregation, so that the features between adjacent nodes can be effectively fused, and then the feature is reshaped to generate two features with sizes Z B×HW×C and Z B×C×HW , and then the two features are multiplied to generate a local affinity attention matrix with a size of Z B×HW×HW . Subsequently, the present disclosure performs a Softmax operation and a dimension summation Sum operation on the local affinity attention matrix in sequence, and then reshapes the feature to Z B×1×H×W , and then multiplies it with the input third fusion feature X to obtain a local attention fusion feature, which is endowed with local feature properties on the basis of the original feature X.
[0128] For the global self-attention branch C, the present disclosure also performs 1 × 1 convolution and graph convolution on the third fusion feature X, and then reshapes the feature to generate two features with sizes Z BHW×C and Z C×BHW , and then performs matrix multiplication to generate a global affinity attention matrix with a size of Z BHW×BHW . Similarly, the present disclosure performs a Softmax operation and a dimension summation Sum operation on the global affinity attention matrix in sequence, and then reshapes the feature to ZB×1×H×W Then, the global attention fusion feature is obtained by multiplying the input third fusion feature X, and the original feature X is endowed with the global feature attribute.
[0129] In order to effectively learn the local feature and the global feature, the local-global interaction branch B is used to interact the local affinity attention matrix and the global affinity attention matrix. Specifically, the local affinity attention matrix is first reshaped, and the size is converted to Z BHW×HW , and then the dimension sum Sum is performed to generate a new feature, and the size is Z BHW . Similarly, the global affinity attention matrix is dimensionally summed Sum to generate a new feature, that is, the size is Z BHW , and then the generated new features are added to obtain a fusion attention matrix, and then the fusion attention matrix is reshaped to generate a new feature with a size of Z B×1×H×W , and then multiplied by the input original feature X, so that a new feature is obtained, that is, the size is Z B×C×H×W . Next, the new feature size is reshaped to Z B×C×HW , and then graph convolution is performed, and then maximum pooling MaxPool, average pooling AvgPool and Softmax operations are performed in turn, and finally the local-global attention fusion feature is obtained, and the size of the feature is Z B×1×H×W . Then, respectively multiplied by the local self-attention fusion feature and the global self-attention fusion feature, the first attention feature and the second attention feature are obtained. Through the above configuration, the representation ability of the original feature X on the local feature and the global representation can be enhanced. Finally, the first attention feature and the second attention feature are respectively executed 1x1 convolution processing, and then channel connection is performed to further fuse the local self-attention branch and the global self-attention branch by performing channel splicing, and then the spliced feature is subjected to 3x3 convolution to reduce the number of channels by half, and the output feature of the LGCA module is obtained.
[0130] Through the above embodiments, the fusion and enhancement of local information and global information can be realized, and through the interactive fusion and reinforcement of forward and reverse features, the interaction and detail complementation of features are realized, the expression ability of the features is improved, and the segmentation accuracy is improved.
[0131] In addition, the refined feature output by the third feature complementary module can be used as a lung feature to determine a lung infection region, wherein the pixel points with pixel values greater than a threshold value in the lung feature can be used to determine the lung infection region, such as a threshold value of 0.5. The above is only an exemplary description, and is not specifically limited herein.
[0132] In some embodiments of the present disclosure, the obtained lung feature can be further processed by a chunking process, for example, the lung feature can be split according to the length and width dimensions, such as m equal parts, where m is the number of chunks in the dimension. In some embodiments of the present disclosure, the H and W dimensions of the lung feature can be split into 3 equal parts, respectively, to obtain 9 sub-chunks. In some embodiments of the present disclosure, the number of chunks in different dimensions can be the same or different, which is not limited in the present disclosure. After obtaining the chunks, the lung infection area can be identified for each sub-chunk fusion feature to obtain the confidence of the lung infection area in each sub-chunk. The probability of including the lung infection area in the sub-chunk can be identified by using an activation function (sigmoid), and the probability value is determined as the confidence. Then, the sub-chunk with the highest confidence is screened out by using the maximum value operation (torch.max), and the sub-chunk fusion lung feature with the highest confidence is used to determine a new lung feature map. The confidence value is between 0 and 1.
[0133] The feature map of the sub-chunk with the highest confidence and the lung feature map can be used to obtain a new lung feature map. Specifically, the lung feature can be processed by dimension reduction to obtain a feature with the same size as the sub-chunk fusion feature, such as performing 3*3 convolution to realize the dimension reduction. Then, the dimension-reduced feature and the feature of the sub-chunk with the highest confidence are connected in the channel (torch.cat) to obtain a connection feature. The connection feature is then processed by convolution (such as 3*3 convolution) to change the channel to 1 to obtain a new lung feature map. Through the above process, the detailed features of the lung feature can be further extracted to improve the segmentation accuracy.
[0134] The detailed enhancement process of some embodiments of the present disclosure can be represented by the following formula:
[0135] l1, l2,...,l9 = Chunk3(Chunk3(F,2),3)
[0136] l max =arr(sigmoid(l1),sigmoid(l2),...,sigmoid(l9))
[0137] R1 = Conv3(Cat(F, l max ))
[0138] where Chunk3(F,dim) represents a dimension chunking operation that divides the feature F into 3 equal parts in the target dimension dim; l1, l2,...,l9 represent the sub-chunks after channel chunking, and l maxrepresents that the sub-block with the highest target confidence is contained, sigmoid(·) represents a binary classification operation, arr(·) represents a maximum value operation, Cat(·) represents a channel splicing operation. Conv3(·) represents a two-dimensional convolution operation with a convolution kernel of 3. R1 represents a new lung feature, and F represents a lung feature.
[0139] In the case of obtaining the new lung feature, segmentation of the lung infection area can be performed using the new lung feature. The new lung feature of the embodiment of the present disclosure can more accurately express the feature information of the lung infection area by fusing local features, thereby improving the segmentation accuracy.
[0140] In combination with the processing procedures of the above-mentioned modules, the present disclosure can complement the multi-content information of the lung image information. In the encoding, the shallow features are discarded, the high-level features containing rich semantic information are aggregated, and the integrated semantic information is used to extract rich global information. In the decoding process, the global information is used as a guide, and the multi-level feature fusion interaction is performed, thereby effectively establishing the connection between the features, so that the lung infection area is more accurately segmented.
[0141] Next, the training process of the lung image segmentation network is described by the embodiment of the present disclosure. Figure 2 As shown in the structural schematic diagram, the training data used by the embodiment of the present disclosure includes a public lung segmentation dataset COVID-19CT, or can also include clinically collected lung images. The above-mentioned lung images include lung infections of different degrees, such as COVID-19 infection areas, or can be other types of pneumonia lesions. The present disclosure does not make specific limitations on this. The specific training process is as follows:
[0142] First, the training set and the test set are divided according to the proportion of 8:2 from the training data set. The test set uses the weight in the test training process to select the best weight. Considering the utilization rate of the display memory, the embodiment of the present disclosure uniformly adjusts the input picture to 352x352. The gradient descent algorithm in the training process selects the Adam algorithm. The advantages are high computational efficiency, less memory required, can solve the problem of containing high noise or sparse gradient, and the hyperparameters can be intuitively explained and only need a small amount of parameter adjustment. The selected loss function combines the weighted IoU (Intersection Over Union, IoU) loss and the binary cross entropy (Binary Cross Entropy, BCE) loss, which is expressed as:
[0143]
[0144] wherein represents the weighted IoU loss, BCE loss representing global constraint and local constraint (pixel level). During the training process, the three local features F5, F4, F3 respectively output by the three feature complementary modules FCM can be up-sampled to the same size as the real label Mask, and compared with the mask to calculate the overall loss. Specifically, the mean value of the loss of each feature F5, F4, F3 can be used as the network loss, and the network loss can be used for back propagation to update the network parameters until the training condition is met. The training condition can include that the training round number reaches a preset value (such as a value greater than 5000), or the loss is less than a loss threshold value, such as a value less than 0.01, or the accuracy of the verification process does not increase continuously after a preset number of rounds, and the preset number of rounds is an integer greater than 50. The present disclosure does not make specific limitations on this.
[0145] The effects of the present disclosure can be further illustrated by the following experiments. All the frameworks of the present disclosure are implemented by using the PyTorch framework, and the training process is accelerated by an RTX3090. The initial learning rate of the Adam optimization algorithm selected by the present disclosure is 1e-4, the batch size is set to 5, and the image size of all training and testing processes is 352x352.
[0146] Figure 9 For the segmentation results of the COVID-19 data set of the embodiments of the present disclosure, compared with UNet, UNet++ and Inf-Net, the present disclosure can accurately locate the lung infection area, thereby realizing accurate segmentation.
[0147] Those skilled in the art can understand that in the above method of the specific implementation, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0148] In addition, the present disclosure also provides a lung infection area segmentation device, an electronic device, a computer readable storage medium, and a program, all of which can be used to implement any one of the lung infection area segmentation methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated here.
[0149] Figure 10 A block diagram of a lung infection area segmentation device according to an embodiment of the present disclosure is shown, as shown in Figure 10 The lung infection area segmentation device comprises:
[0150] The feature extraction module 10 is configured to perform feature extraction processing on the lung image to obtain a plurality of high-level features.
[0151] The fusion module 20 is configured to perform multi-scale fusion on the high-level features to obtain high-level enhanced features.
[0152] The global information extraction module 30 is configured to perform global information extraction on the high-level enhancement feature to obtain a global guidance feature.
[0153] The feature complementation module 40 is configured to perform multi-level feature interaction processing on the global guidance feature and the high-level feature to obtain a lung feature, where the lung feature is used to represent a lung infection area of the lung image.
[0154] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules for performing the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For briefness, details are not described here.
[0155] The embodiments of the present disclosure also provide a computer-readable storage medium having computer program instructions stored therein, and the computer program instructions are executed by a processor to implement the above method. The computer-readable storage medium can be a non-volatile computer-readable storage medium.
[0156] The embodiments of the present disclosure also provide an electronic device, which includes a processor, and a memory for storing processor-executable instructions, and the processor is configured to implement the above method.
[0157] The electronic device can be provided as a terminal, a server, or other forms of devices.
[0158] Figure 11 A block diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. The electronic device 800 can be, for example, a terminal such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0159] Reference Figure 11 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0160] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the methods described above. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0161] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or nonvolatile memory, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0162] The power component 806 supplies power for various components of the electronic device 800. The power component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0163] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. When the electronic device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the back camera can receive external multimedia data. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0164] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.
[0165] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keyboard, a click wheel, a button, etc. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0166] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components of the electronic device 800, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature changes of the electronic device 800. The sensor component 814 can include an optical sensor for detecting ambient light, a proximity sensor configured to detect proximity of an object, a motion sensor, a temperature sensor, a magnetic sensor, an acceleration sensor, a gyroscope sensor, or a pressure sensor.
[0167] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and another device. The electronic device 800 can access a wireless network based on a corresponding communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0168] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, to perform the above-described methods.
[0169] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to complete the above-described methods.
[0170] Figure 12 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server. Referring to FIG. 19, the electronic device 1900 includes one or more processors 1910, memory 1920 and at least one interface 1940. In an example embodiment, the electronic device 1900 can further include a bus 1950. In an example embodiment, the bus 1950 can be connected to the one or more processors 1910, the memory 1920, and the at least one interface 1940. In an example embodiment, the bus 1950 can include a path that enables OBI (On Board Interconnect) or CI (Chip Interconnect) communication. Figure 12The electronic device 1900 includes a processing component 1922, which is further composed of one or more processors, and memory resources represented by the memory 1932 for storing instructions, such as application programs, executable by the processing component 1922. The application programs stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.
[0171] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0172] In exemplary embodiments, a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions, is also provided, which can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0173] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0174] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a
[0175] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0176] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0177] The computer readable program instructions can also be loaded onto a computing / processing device, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computing / processing device, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computing / processing device, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0178] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, to cause a series of operational steps to be performed on the computer to produce a computer-implemented process. The instructions can also cause one or more processors of a computer or other programmable data processing apparatus to
[0179] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0180] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0181] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive of the disclosure. Many modifications and variations of the described embodiments are possible in light of this disclosure. It is intended that the scope of the disclosure be limited not by this detailed description, but rather by the claims appended hereto. The choice of words in the description is intended to be a choice of emphasis on features deemed to be key in expressing the principles of the present disclosure. Other words can be substituted therefor without departing from the essence of the disclosure.
Claims
1. A lung infection segmentation method, characterized in that, The method comprises the following steps: performing feature extraction processing on a lung image to obtain a plurality of high-level features; performing multi-scale fusion processing on the high-level features to obtain high-level enhanced features; performing global information extraction on the high-level enhanced features to obtain global guidance features; performing multi-layer feature interaction processing on the global guidance features and the high-level features to obtain lung features, wherein the lung features are used to represent lung infection areas of the lung image; the step of performing multi-scale fusion processing on the high-level features to obtain high-level enhanced features comprises: performing multi-scale fusion processing on each of the high-level features to obtain multi-scale fusion features; performing channel splicing processing on the multi-scale fusion features to obtain the high-level enhanced features; wherein the step of performing multi-scale fusion processing on each of the high-level features to obtain multi-scale fusion features comprises: performing feature processing of a plurality of branches on each of the high-level features to obtain branch features; performing fusion processing on the branch features to obtain multi-scale fusion features corresponding to the high-level features; the plurality of branches comprises three branch groups, and the step of performing feature processing of a plurality of branches on each of the high-level features to obtain branch features comprises: performing first convolution processing on the high-level features on all branches to obtain corresponding first convolution features; performing second convolution processing on each of the first convolution features on a second branch group to obtain second convolution features, wherein the second convolution processing comprises performing dilated convolution on the first convolution features using different dilation rates; performing channel fusion on the first convolution features of the first branch group and each of the second convolution features of the second branch group to obtain third convolution features; performing addition fusion on the third convolution features and the first convolution features of the third branch group to obtain the branch features.
2. The method of claim 1, wherein, the step of performing global information extraction on the high-level enhanced features to obtain global guidance features comprises: performing feature remodeling on the high-level enhanced features to obtain remodeled features; adding the remodeled features and position encoding features to obtain encoding fusion features; performing global information extraction on the encoding fusion features to obtain the global guidance features.
3. The method of claim 1, wherein, the step of performing multi-layer feature interaction processing on the global guidance features and the high-level features to obtain lung features comprises: inputting the global guidance features and the high-level features into a feature complementarity module, and sequentially performing feature interaction on the global guidance features and the high-level features by using feature complementarity modules in the feature complementarity module to obtain the lung features; wherein the feature complementarity module comprises a plurality of feature complementarity modules, and the input features of the feature complementarity modules comprise global guidance features, corresponding high-level features, and refined features output by a previous feature complementarity module, and the refined features input by the first feature complementarity module are the global guidance features.
4. The method of claim 3, wherein, the step of performing feature interaction on the global guidance features and the high-level features by using the feature complementarity modules in the feature complementarity module comprises: obtaining forward weights and reverse weights by using the global guidance features and the high-level features; The refined features output by the previous feature complementary module are respectively multiplied with the forward weight and the backward weight to obtain forward correction features and backward correction features; The refined features are respectively reinforced by using the forward correction features and the backward correction features to obtain forward reinforced features and backward reinforced features; The forward reinforced features and the backward reinforced features are subjected to the reinforcement processing to obtain the refined features output by the current feature complementary module.
5. The method of claim 4, wherein, The forward weight and the backward weight are obtained by using the global guidance feature and the high-level feature, including: The global guidance feature and the high-level feature are subjected to addition and activation processing to obtain the forward weight; The forward weight is negated to obtain the backward weight; And / or The input features of the reinforcement processing include first input features and second input features, and the reinforcement processing includes: The first input features and the second input features are subjected to multiplication processing to obtain product features; The product features are respectively added to the first input features and the second input features to obtain first fusion features and second fusion features; The first fusion features and the second fusion features are respectively subjected to convolution processing and channel dimension feature splicing to obtain third fusion features; The third fusion features are subjected to local enhancement and global enhancement processing to obtain output features as reinforced features or refined features.
6. A lung infected region segmentation apparatus characterized by comprising: It includes: A feature extraction module is configured to perform feature extraction processing on a lung image to obtain a plurality of high-level features; A fusion module is configured to perform multi-scale fusion on the high-level features to obtain high-level enhanced features; A global information extraction module is configured to perform global information extraction on the high-level enhanced features to obtain a global guidance feature; A feature complementary module is configured to perform multi-level feature interaction processing on the global guidance feature and the high-level features to obtain a lung feature, which is used to represent a lung infection area of the lung image; The multi-scale fusion processing on the high-level features to obtain high-level enhanced features includes: Each of the high-level features is subjected to multi-scale fusion processing to obtain multi-scale fusion features; The multi-scale fusion features are subjected to channel splicing processing to obtain the high-level enhanced features; The multi-scale fusion processing on each of the high-level features to obtain multi-scale fusion features includes: Each of the high-level features is subjected to feature processing of a plurality of branches to obtain branch features; The branch features are subjected to fusion processing to obtain multi-scale fusion features corresponding to the high-level features; The plurality of branches includes three branch groups, and the feature processing of a plurality of branches on each of the high-level features to obtain branch features includes: Each of the high-level features is subjected to first convolution processing on all branches to obtain corresponding first convolution features; Each of the first convolution features is subjected to second convolution processing on the second branch group to obtain second convolution features; wherein the second convolution processing includes performing dilated convolution on the first convolution features by using different dilation rates; The channel fusion is performed on the first convolutional feature of the first branch group and each second convolutional feature of the second branch group to obtain a third convolutional feature; The addition fusion is performed on the third convolutional feature and the first convolutional feature of the third branch group to obtain the branch feature.
7. An electronic device, comprising: Comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method of any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Endoscopic image-based large intestine polyp segmentation method and system and related components
CN113838047A