Chest radiography lesion detection method based on enhanced feature extraction and fusion
The enhanced feature extraction and fusion method for YOLO series algorithms addresses the information loss in traditional FPNs by aligning and fusing multi-scale features, improving small target detection in chest X-ray images with increased precision and efficiency.
Patent Information
- Application Number
- CN202510486007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing YOLO series algorithms have the loss of information during feature fusion in chest flat film lesions detection, resulting in low detection accuracy of small targets, and high-resolution feature map processing and calculation complexity.
The N-1 level feature map is introduced to the neck network, and the feature map alignment and fusion are realized through reference feature map size transformation and channel dimension splicing, combined with the global coordinate attention mechanism, the feature extraction and fusion efficiency is improved and the calculation complexity is reduced.
It improves the accuracy and generalization ability of chest radiograph lesions detection, reduces missed diagnosis and misdiagnosis, and maintains the efficiency of the model.
Smart Images

Figure CN120318203A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing, and particularly to a method for detecting chest lesions based on enhanced feature extraction and fusion. Background Art
[0002] CXR (Chest X-Ray) refers to a chest X-ray film, chest radiograph, or plain chest film, which is an image formed by X-rays penetrating chest tissues and can help doctors evaluate structures such as the heart, lungs, and bones. Figure 1 A chest radiograph example is shown. It can be seen that a chest radiograph is a high-resolution overlapping image, characterized by mutual occlusion between tissue organs and a relatively high proportion of small target lesions.
[0003] YOLO series algorithms are commonly used in lesion detection of medical images such as CT and have high detection efficiency. YOLO series algorithms usually include a backbone network Backbone, a neck network Neck, and a head output layer Head. The backbone network Backbone is used to extract multi-scale features of medical images. The neck network Neck is used to fuse multi-scale features and enhance the detection ability for targets of different sizes. The head output layer Head generates the final detection results based on the features output by the neck network Neck, including target category, location (bounding box), and confidence.
[0004] In YOLO series algorithms, the neck structure generally uses a Feature Pyramid Network (FPN) and its improved versions to achieve the fusion of multi-level features. However, these information fusion methods have an obvious defect: when cross-layer information fusion is required, the traditional FPN-like structure cannot transmit information losslessly, which hinders the YOLO network from better performing information fusion. Specifically, as Figure 2 shown, in the feature fusion structure of the traditional FPN, the P3, P4, and P5 layers are arranged in sequence from top to bottom. When the P3 layer wants to use the information of the P4 layer, it can directly access and fuse the information of the P4 layer. Figure 2The Fuse module in the Chinese text represents fusion. When the P3 layer wants to utilize the information of the P5 layer, it must fuse the information of the P4 layer and the P5 layer and then fuse it with the information of the P3 layer to indirectly obtain the information of the P5 layer. This transmission method may result in a large amount of information loss during the calculation process. The information interaction between layers can only exchange the information selected in the intermediate process, and the unselected information will be discarded during the transmission process, resulting in the information of a certain layer being able to only fully utilize the adjacent layer, while weakening the help provided by other layers. Therefore, when the YOLO series of algorithms are used for lesion recognition and localization in chest radiographs, the above defects will result in the loss of a large number of subtle features in chest radiographs, leading to low detection accuracy for small targets. Although in related technologies, to improve the detection performance of small targets, a detection head is directly added using a high-resolution feature map, the high-resolution feature map requires a large number of pixels to process each target, thus greatly increasing the computational complexity of the model. Summary of the Invention
[0005] This application aims to at least solve the technical problems existing in the prior art and provides a method for detecting chest radiograph lesions based on enhanced feature extraction and fusion.
[0006] In a first aspect, this application provides a method for detecting chest radiograph lesions based on enhanced feature extraction and fusion, including: obtaining a chest radiograph; inputting the chest radiograph into a trained object detection model to obtain a lesion detection result; marking the lesion detection result on the chest radiograph and outputting the marked chest radiograph; the object detection model includes: a backbone network: including a feature extraction network with N levels connected in sequence, and the feature extraction networks at N levels respectively extract feature maps of different scales of the chest radiograph; inputting the N - 1 feature maps obtained by the feature extraction networks from level 2 to level N into the neck network; where the level index of the feature extraction network is from 1 to N, the N - 1 feature maps include 1 reference feature map and N - 2 non-reference feature maps, the level index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3; a neck network: taking the size of the reference feature map as a reference, performing size transformation on the non-reference feature maps; splicing the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a spliced feature map; performing fusion processing on the spliced feature map to obtain a fused feature map; performing distribution processing on the fused feature map to obtain level distribution feature maps at one or more preset distribution levels; after performing inverse size transformation on the level distribution feature maps at one or more preset distribution levels, respectively fusing them with the feature maps corresponding to the preset distribution levels to obtain fused features at one or more preset distribution levels, and the level index of the preset distribution level is greater than 2 and less than N; fusing the fused features at one or more preset distribution levels and the feature map of level N to obtain multiple detection features; a head output layer: obtaining a lesion detection result based on the multiple detection features.
[0007] Second aspect, the present application provides a chest radiograph lesion detection device based on enhanced feature extraction and fusion, which is used to implement the chest radiograph lesion detection method based on enhanced feature extraction and fusion provided in the first aspect of the present invention, including: a chest radiograph acquisition module for acquiring a chest radiograph; a detection module for inputting the chest radiograph into a trained target detection model to obtain a lesion detection result; a marking output module for marking the lesion detection result on the chest radiograph and outputting the marked chest radiograph; the target detection model includes: a backbone network: including N hierarchical feature extraction networks connected in sequence, and the N hierarchical feature extraction networks respectively extract feature maps of different scales of the chest radiograph; inputting the N-1 feature maps obtained by the feature extraction networks from layer 2 to layer N into the neck network; wherein, the layer index of the feature extraction network is from 1 to N, the N-1 feature maps include 1 reference feature map and N-2 non-reference feature maps, the layer index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3; a neck network: taking the size of the reference feature map as a reference, performing size transformation on the non-reference feature maps; splicing the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a spliced feature map; performing fusion processing on the spliced feature map to obtain a fused feature map; performing distribution processing on the fused feature map to obtain hierarchical distribution feature maps of more than one preset distribution layer; performing inverse size transformation on the hierarchical distribution feature maps of more than one preset distribution layer and then fusing them with the feature maps corresponding to the preset distribution layers respectively to obtain fused features of more than one preset distribution layer, and the layer index of the preset distribution layer is greater than 2 and less than N; fusing the fused features of more than one preset distribution layer and the feature map of layer N to obtain multiple detection features; a head output layer: obtaining a lesion detection result based on the multiple detection features.
[0008] Third aspect, the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the chest radiograph lesion detection method based on enhanced feature extraction and fusion provided in the first aspect of the present invention.
[0009] The beneficial technical effects of the present invention:
[0010] (1) In the present invention, N-1 hierarchical feature maps are introduced from the backbone network to the neck network. In the neck network, taking the feature map of a certain layer between layer 2 and layer N as the reference feature map, and taking the size of the reference feature map as a reference, performing size transformation on the non-reference feature maps so that the sizes of the non-reference feature maps are the same as the size of the reference feature map, thereby completing the alignment of the N-1 feature maps. This alignment method can reduce information loss during the alignment process of the feature maps.
[0011] (2) Concatenate the reference feature map and the non-reference feature map after size transformation in the channel dimension to obtain a concatenated feature map, perform fusion processing on the concatenated feature map to obtain a fused feature map, perform distribution processing on the fused feature map to obtain hierarchical distribution feature maps of one or more preset distribution levels, fuse the fused features of one or more preset distribution levels and the feature map of level N to obtain multiple detection features, and the head output layer obtains a lesion detection result based on the multiple detection features. It can be seen that the present application introduces the feature map of level 2 with high resolution into the neck network for feature fusion, so that there are more chest radiograph details in the hierarchical distribution feature map and richer semantic information is obtained;
[0012] (3) The feature map of level N has a lower resolution but a larger receptive field, which helps the object detection model to more effectively capture the overall structural information of the chest radiograph. Fusing it with the feature maps of other levels in the neck network helps the object detection model to better adapt to objects of different sizes and improve the detection performance and generalization ability;
[0013] (4) Introduce the feature map of level N into the head output layer: Since its resolution is low, the number of pixels required to process each object is small, so that the inherent efficiency enjoyed by the YOLO series models is maintained and the object recognition accuracy is improved without significantly increasing the computational complexity;
[0014] Therefore, through the combination of the above technical means, the object detection model provided by the present application improves the object detection accuracy, especially the detection accuracy of small objects, without significantly increasing the computational complexity. The lesion detection method provided by the present application helps to reduce missed diagnoses and misdiagnoses, and provides more comprehensive and reliable diagnostic basis for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic flowchart of a chest radiograph lesion detection method based on enhanced feature extraction and fusion in a preferred embodiment of the present invention;
[0016] Figure 2 is a schematic network structure diagram of an object detection model in a preferred embodiment of the present invention;
[0017] Figure 3 is a schematic network structure diagram of the G-CA module in a preferred embodiment of the present invention;
[0018] Figure 4 is a schematic structure diagram of an electronic device in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0020] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0021] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0022] The execution subject of the chest X-ray lesion detection method based on enhanced feature extraction and fusion provided by the present invention includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the chest X-ray lesion detection method based on enhanced feature extraction and fusion can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0023] The present invention provides a chest X-ray lesion detection method based on enhanced feature extraction and fusion. In a preferred embodiment, please refer to Figure 1 as shown, the chest X-ray lesion detection method includes:
[0024] Step S1, obtaining a chest X-ray.
[0025] Exemplarily, a chest radiograph to be detected can be obtained from a chest radiograph imaging device through a network interface or a data cable, such as a digital radiography (DR) system, a computed radiography (CR) system, a mobile X-ray machine (bedside DR), etc. Alternatively, the chest radiograph to be detected can be read from a memory or database storing the chest radiograph.
[0026] Step S2: Input the chest radiograph into the trained object detection model to obtain lesion detection results.
[0027] Exemplarily, the lesion detection results include the target bounding box coordinates, confidence scores, and class labels of multiple lesion targets. The lesion targets are not limited to cysts, nodules, effusions, etc.
[0028] Step S3: Mark the lesion detection results on the chest radiograph and output the marked chest radiograph.
[0029] Exemplarily, as Figure 2 the output image in, mark the bounding boxes in the chest radiograph according to the target bounding box coordinates in the lesion detection results, and display the confidence scores and class labels in the bounding boxes.
[0030] In this embodiment, preferably, see Figure 2 , the object detection model includes:
[0031] Backbone network: It includes a feature extraction network with N layers connected in sequence. The feature extraction networks of the N layers respectively extract feature maps of different scales of the chest radiograph; input the N - 1 feature maps obtained by the feature extraction networks from layer 2 to layer N into the neck network; wherein, the layer index of the feature extraction network is from 1 to N, the N - 1 feature maps include 1 reference feature map and N - 2 non-reference feature maps, the layer index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3;
[0032] Neck network: Taking the size of the reference feature map as a reference, perform size transformation on the non-reference feature maps; splice the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a spliced feature map; perform fusion processing on the spliced feature map to obtain a fused feature map; perform distribution processing on the fused feature map to obtain layer-distributed feature maps of one or more preset distribution layers; after performing inverse size transformation on the layer-distributed feature maps of one or more preset distribution layers, fuse them with the feature maps corresponding to the preset distribution layers respectively to obtain fused features of one or more preset distribution layers, and the layer index of the preset distribution layer is greater than 2 and less than N; fuse the fused features of one or more preset distribution layers and the feature map of layer N to obtain multiple detection features.
[0033] Head output layer: Obtain lesion detection results based on multiple detection features.
[0034] In this embodiment, the backbone network is used to extract the lesion target features in the chest X-ray, and feature maps of different scales are obtained respectively. The neck network is used to fuse the feature maps of different scales, so as to obtain richer semantic information and more accurate target position information. The head output layer is used to generate the final lesion detection result (including the target bounding box coordinates, confidence score, and class label). The head output layer adopts the structure of a decoupled detection head, and each detection head is divided into two branches, one branch is used to predict the class, and the other branch is used to predict the position information.
[0035] In this embodiment, the feature extraction networks of N levels all include a downsampling convolution to obtain N feature maps of different scales. Only the N - 1 feature maps obtained by the feature extraction networks of levels 2 to N in the backbone network are extracted for subsequent fusion processing by the neck network. Among the N - 1 extracted feature maps, excluding the feature maps of levels 2 and N, any one feature map can be selected as the reference feature map, and the N - 2 feature maps other than the reference feature map among the N - 1 feature maps can be all or part of them as non-reference feature maps. Preferably, the reference feature map is the feature map at the middle level among the N - 1 feature maps, which is beneficial to reducing information loss in feature map alignment.
[0036] In one example, as Figure 2 shown, N is 6, and the feature maps B2, B3, B4, B5, and B6 obtained by the feature extraction networks of levels 2 to 6 in the backbone network are input into the neck network, where the i'-th feature map N represents the batch size; represents the number of channels, the size of the feature map is R Bi = W × H, W represents the width, H represents the height, and i' ∈ [2, 6]. Among them, the feature map B4 of level 4 is selected as the reference feature map, and the feature maps B2, B3, B5, and B6 are all used as non-reference feature maps. Figure 2 In it, assuming the size of the chest X-ray input into the backbone network is 1280×1280×3, the size of the feature map B2 of level 2 is 320×320×128, the size of the feature map B3 of level 3 is 160×160×256, the size of the feature map B4 of level 4 is 80×80×512, the size of the feature map B5 of level 5 is 40×40×768, and the size of the feature map B6 of level 6 is 20×20×1024.
[0037] In this embodiment, the neck network includes two-stage processing.
[0038] In the first-stage processing: with reference to the size of the reference feature map, perform size transformation on the non-reference feature maps so that the sizes of N-1 feature maps are the same and are all the size of the reference feature map; concatenate the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a concatenated feature map; perform fusion processing on the concatenated feature map to obtain a fused feature map; perform distribution processing on the fused feature map to obtain hierarchical distribution feature maps at one or more preset distribution levels; after performing inverse size transformation on the hierarchical distribution feature maps at one or more preset distribution levels, fuse them with the feature maps corresponding to the preset distribution levels respectively to obtain fused features at one or more preset distribution levels, where the hierarchical index of the preset distribution level is greater than 2 and less than N; the size of the hierarchical distribution feature map at the preset distribution level after inverse size transformation is the same as the size of the feature map corresponding to this preset distribution level, which is convenient for fusion. The number of preset distribution levels is greater than or equal to 2 and less than or equal to N-3. Preferably, the preset distribution levels include the level where the reference feature map is located, and at least one level upstream and / or at least one level downstream of the level where the reference feature map is located. In the above example, the predicted distribution levels include level 3, level 4, and level 5. The large-size feature maps are fused through the first-stage processing.
[0039] The second-stage processing: with reference to Figure 2 , for top-down feature aggregation, fuse the fused features at one or more preset distribution levels and the feature map at level N to obtain multiple detection features. Preferably, fuse the fused features at all preset distribution levels and the feature map at level N, and discard the feature map at level 2 for fusion, which can avoid the traditional method directly adding a P2 detection head by using the feature map B2 in order to improve the detection performance of small targets. Because the feature map B2 has a high resolution and requires a large number of pixels to process each target, the traditional method greatly increases the model calculation complexity. Therefore, in this application, only the feature map B2 at level 2 is introduced into the fusion part, and no hierarchical distribution feature map at level 2 is generated. It can obtain high-resolution details, retain more subtle features, which helps to improve the detection accuracy without significantly increasing the inference time. The small-size feature maps are fused through the second-stage processing. In the above example, fuse the fused features at level 3, level 4, and level 5 and the feature map at level 6 to obtain detection features at 4 levels from level 3 to level 6, and the detection features at the 4 levels are respectively input into the detection heads corresponding to the levels in the head output layer.
[0040] In the above example, select the feature maps B2, B3, B4, B5, and B6 for fusion. In order to improve the detection performance of small targets, the higher-resolution feature B2 is specially selected. The high-resolution feature map can retain more image detail information, the positioning information is accurate but the semantic information is not rich. Fusing with the global context information can improve the model's understanding ability, which is crucial for improving the detection performance of small targets.
[0041] In this embodiment, through the first-stage processing and the second-stage processing, the detection ability of the target detection model for targets of different sizes in chest X-rays is improved and the network computing overhead is reduced.
[0042] In a preferred embodiment, to facilitate the extraction of detailed features of chest X-rays. Referring to Figure 2 , the feature extraction network of level 1 includes downsampling convolution; the feature extraction network of level N includes a downsampling convolution, a C2F module, and an SPPF module connected in sequence; the feature extraction network of any level between level 1 and level N includes a downsampling convolution and a C2F module connected in sequence.
[0043] In this embodiment, both the C2F module and the SPPF module are existing modules in the YOLOv8 network. The C2F module (Cross Stage Partial Fusion) can better extract features, has stronger gradient propagation, and higher computational efficiency. The SPPF module (Spatial Pyramid Pooling Fast) can perform multi-scale feature extraction with high efficiency and is suitable for real-time detection tasks.
[0044] In a preferred embodiment, in order to make the target detection model pay more attention to the abnormal areas in the chest X-ray (CXR) and reduce the interference of background features, this application introduces an improved global coordinate attention mechanism module (G-CA module). Specifically, in the backbone network, the feature extraction network of at least one level from level 3 to level N-1 is connected to the feature extraction network of the next level through the G-CA module, and / or the feature extraction network of level N further includes a G-CA module located between the C2F module and the SPPF module. In the above example, referring to Figure 2 , the feature extraction network of level 3 is connected to the feature extraction network of level 4 through the G-CA module, and the feature extraction network of level 4 is connected to the feature extraction network of level 5 through the G-CA module.
[0045] Among them, referring to Figure 3 , the G-CA module includes:
[0046] The first branch is used to obtain the global channel attention weight g of the input feature map of the G-CA module, and the global channel attention weight of the c-th channel can be represented by g c .
[0047] The second branch is used to obtain the height-direction attention weight g h and the width-direction attention weight g w of the input feature map of the G-CA module. The height and width direction attention weights of the c-th channel can be represented by and representation
[0048] The multiplication weighted operation unit is used to multiply the input feature map of the G-CA module element-wise with the global channel attention weight, the height direction attention weight, and the width direction attention weight to obtain the output feature y of the G-CA module. The feature map y of the c-th channel output by the G-CA module c The eigenvalue y at the i-th row and j-th column c The eigenvalue of (i, j) can be represented as:
[0049]
[0050] where x c (i, j) represents the eigenvalue of the feature map point (i, j) of the c-th channel x in the input feature map of the G-CA module c of the feature map, represents the weight value of the i-th row in the height direction attention weight of the c-th channel and represents the weight value of the j-th column in the width direction attention weight of the c-th channel of the c-th channel
[0051] In this embodiment, referring to the attached Figure 3 , preferably, the first branch includes:
[0052] The compression unit compresses the height information and width information of each channel of the input feature map of the G-CA module into channel eigenvalues, and the channel eigenvalues of all channels form a feature vector z.
[0053] In one example, the input feature map of the G-CA module is represented as a tensor The output feature of the G-CA module is a transformed tensor of the same size as X The compression unit compresses the spatial information (i.e., height and width) of each channel into a single value (i.e., channel eigenvalue), thereby obtaining a feature vector z containing the global information of all channels:
[0054]
[0055] The first branch convolution unit performs convolution processing on the feature vector to obtain the first branch feature map F. The convolution kernel size of the first branch convolution unit is 1×1.
[0056] The first branch activation function processing performs activation function processing on the first branch feature map F to obtain the weight of each channel (i.e., predicts the importance of each channel), and the weights of all channels form the global channel attention weight g. The activation function σ is preferably but not limited to the Sigmoid activation function.
[0057] g = σ(F(z)).
[0058] In this embodiment, referring to the attached Figure 3 , preferably, the second branch includes:
[0059] A global average pooling unit that performs global average pooling processing on the input feature map of the G-CA module in the height direction and the width direction to obtain the height-direction pooled feature z h and the width-direction pooled feature z w . Therefore, the output of the c-th channel at height h can be expressed as:
[0060]
[0061] Similarly, the output of the c-th channel at width w can be expressed as:
[0062]
[0063] where W and H respectively represent the width and height of the input feature map of the G-CA module, and x c (h, j) represents the feature value of the c-th channel of the feature map x of the input feature map X of the G-CA module c at the j-th column and the h-th row (height h), and x c (i, w) represents the feature value of the c-th channel of the feature map x of the input feature map X of the G-CA module c at the i-th row and the w-th column (width w).
[0064] A second branch concatenation unit that concatenates the height-direction pooled feature and the width-direction pooled feature to obtain a concatenated pooled feature. The height-direction pooled feature z h and the width-direction pooled feature z w in the width and height directions of the global receptive field are concatenated together to obtain the concatenated pooled feature [z h , z w .
[0065] A second branch convolution unit that performs convolution processing on the concatenated pooled feature to obtain a second branch feature map. The concatenated pooled feature [z h , z w is input into a convolution module with a shared 1×1 convolution kernel to reduce its dimension to C / r of the original, obtaining the second branch feature map.
[0066] A distribution unit that performs batch normalization processing (obtaining the feature map F1) and activation function processing on the second branch feature map in sequence, and then obtains a feature map f in the form of 1×(W + H)×C / r. The activated feature map f is distributed into a height feature sub-map and a width feature sub-map. The activation function is preferably the Sigmoid activation function.
[0067] The height-direction attention weight acquisition unit performs convolution processing (with a convolution kernel of 1×1) and activation processing (such as the Sigmoid activation function) on the height feature sub-map in sequence to obtain the height-direction attention weight g h . After the height feature sub-map is subjected to convolution processing, a feature map f with the same number of channels as the original is obtained h .
[0068] The width-direction attention weight acquisition unit performs convolution processing (with a convolution kernel of 1×1) and activation processing (such as the Sigmoid activation function) on the width feature sub-map in sequence to obtain the width-direction attention weight g w . After the width feature sub-map is subjected to convolution processing, a feature map f with the same number of channels as the original is obtained w .
[0069] The above process can be described as follows:
[0070]
[0071] f = σ(F1([z h ,z w ))
[0072] g h = σ(F h (f h ))
[0073] g w = σ(F w (f w ))
[0074] In this embodiment, by generating global information embeddings, direction-aware and position-sensitive attention maps, the model can more accurately identify and locate the target of interest, improving the model's understanding ability of the spatial structure in CXR anomaly detection. In the G-CA module, on the one hand, two one-dimensional features are used to encode position information along the height and width directions of the image, and on the other hand, a parallel global attention block is used to extract context features, so as to obtain the spatial position information and global context feature information of the image region of interest, thereby accurately locating the distinguishable regions in the image
[0075] In a preferred embodiment, referring to Figure 2 as shown, in the first-stage processing of the neck network, taking the size of the reference feature map as a reference, the size of the non-reference feature map is transformed, including:
[0076] For a non-reference feature map whose size is larger than that of the reference feature map, the size transformation process is as follows: Convolve the non-reference feature map using a first size transformation convolution based on a preset stride. The preset strides of non-reference feature maps at different levels are different, and the preset stride of a relatively lower level is smaller than that of a relatively higher level. The convolution kernel size of the first size transformation convolution is 3×3.
[0077] For a non-reference feature map whose size is smaller than that of the reference feature map, the size transformation process is as follows: Perform bilinear interpolation (Bilinear) and a second size transformation convolution on the non-reference feature map in sequence. The convolution kernel size of the second size transformation convolution is 1×1.
[0078] In the above example, please refer to Figure 2 , the size transformation process for the non-reference feature map is as follows: Use a convolution kernel of size 3×3 and set the strides to 2 and 4 respectively to resize the non-reference feature maps B2 (320×320) and B3 (160×160) to the same size as the reference feature map B4 (80×80). Then, for the non-reference feature maps B5 (40×40) and B6 (20×20), use the bilinear interpolation method combined with a 1×1 convolution operation to also make them reach the size of the reference feature map B4. After these steps, the feature maps of the five levels are unified in width and height. Since the chest X-ray is an overlapping image and there are many small target abnormalities, the above size transformation process can reduce the information loss in the algorithm information transmission process, retain more subtle features, and help improve the accuracy of lesion detection.
[0079] In a preferred embodiment, in the first-stage processing of the neck network, please refer to Figure 2 as shown, perform a size inverse transformation on the hierarchical distribution feature maps of one or more preset distribution levels, including:
[0080] For a preset distribution level whose feature map size is larger than that of the reference feature map, the size inverse transformation process of the hierarchical distribution feature map of the preset distribution level is as follows: Perform bilinear interpolation (Bilinear) and a second size transformation convolution in sequence. The convolution kernel size of the second size transformation convolution is 1×1. In the above example, the preset distribution level whose feature map size is larger than that of the reference feature map is level 3.
[0081] For a preset distribution level where the size of the feature map is smaller than the size of the reference feature map, the inverse size transformation process of the level distribution feature map of the preset distribution level is as follows: Convolve the level distribution feature map of the preset distribution level using a first size transformation convolution based on a preset step size. In the preset distribution levels where the size of the feature map is smaller than the size of the reference feature map, the preset step sizes of different preset distribution levels are different, and the higher the level, the larger the preset step size. No inverse size transformation is performed on the level distribution feature map of the preset distribution level where the reference feature map is located.
[0082] In a preferred embodiment, in the first-stage processing of the neck network, please refer to Figure 2 , and perform fusion processing on the concatenated feature map to obtain a fused feature map, including: sequentially performing convolution processing, batch normalization processing BN, and activation function processing on the concatenated feature map. In the activation function processing, the activation function is preferably but not limited to the ReUL activation function. The convolution kernel for the convolution processing is preferably but not limited to 1×1.
[0083] In this embodiment, although the first size transformation convolution or bilinear interpolation + second size transformation convolution is used to change the size of the feature map in the size transformation, which reduces information, it will increase the number of parameters. Therefore, to avoid introducing too many additional parameters, in the fusion processing of the concatenated feature map, no reparameterization is required, and only a 1×1 convolution is used, combined with batch normalization processing BN and activation function processing, which can also effectively fuse the features.
[0084] In the above example, after performing size transformation on the non-reference feature map, in the channel dimension, the feature maps of the five levels after size transformation are concatenated to complete the alignment of the feature maps and obtain the concatenated feature map F align , perform fusion processing on the concatenated feature map F align to obtain the fused feature map F fuse , perform distribution processing (Split operation) on the fused feature map F fuse to obtain the level distribution feature maps of 3 preset distribution levels from level 3 to level 5, which are respectively denoted as F fuse_P3 、F fuse_P4 and F fuse_P5 . Since the sizes of the level distribution feature maps distributed at this time are all the same as the reference feature map B4 of level 4, it is necessary to perform bilinear interpolation in combination with a 1×1 convolution operation on the level distribution feature map F fuse_P3 of level 3 to enlarge it to 160×160, and use a 3×3 convolution operation with a stride of 2 on the level distribution feature map F fuse_P5 of level 5 to shrink it to 40×40, and finally fuse them with the corresponding level 3, level 4, and level 5 feature maps (B3, B4, and B5) respectively to obtain the level 3, level 4, and level 5 fused features P3, P4, and P5.
[0085] In the above example, the processing process of the first stage of the neck network can be expressed by the following formula:
[0086] F align = Concat([Conv s=4 (B2), Conv s=2 (B3), B4,
[0087] Conv 1×1 (Billnear(B5), Conv 1×1 (Billnear(B6))])
[0088] F fuse = ReUL(BN(Conv 1×1 (F align )))
[0089] F fuse_P3 , F fuse_P4 , F fuse_P5 = Split(F fuse )
[0090]
[0091] In this embodiment, in the second-stage processing of the neck network, a top-down path aggregation module is used to fuse the fusion features of more than one preset distribution level and the feature map of level N to obtain multiple detection features.
[0092] Please refer to Figure 2 , in the above example, the path aggregation module includes N - 2 aggregation sub-networks, that is, 4 aggregation sub-networks, which correspond one-to-one with the N - 2 detection heads in the head output layer and also correspond one-to-one with levels 3 to N in the backbone network.
[0093] The aggregation sub-network corresponding to level 3 includes a cascaded C2F module and an aggregation convolution. Among them, the output feature of the C2F module is the detection feature of level 3 and is correspondingly input into the detection head Detect-P3.
[0094] The aggregation sub-network corresponding to level 4 includes a cascaded first C2F module, a connection unit, a second C2F module, and an aggregation convolution. Among them, the input end of the first C2F module is input with the fusion feature P4 of level 4. The connection unit is used to connect the output feature of the aggregation sub-network corresponding to level 3 and the output feature of the first C2F module to obtain a connection feature, and use the connection feature as the detection feature of level 4, and input the detection feature of level 4 into the detection head Detect-P4.
[0095] The aggregation sub-network corresponding to level 5 includes a cascaded third C2F module, a connection unit, a fourth C2F module, and an aggregation convolution. The input end of the third C2F module is input with the fused feature P5 of level 5. The connection unit is used to connect the output feature of the aggregation sub-network corresponding to level 4 and the output feature of the third C2F module to obtain a connection feature, and input the connection feature into the fourth C2F module. The output feature of the fourth C2F module is used as the detection feature P5 of level 5, and the detection feature P5 of level 5 is input into the detection head Detect-P5.
[0096] The aggregation sub-network corresponding to level 6 includes a cascaded connection unit and a fifth C2F module. The connection unit is used to connect the feature map of level 6 and the feature output by the aggregation sub-network corresponding to level 5. The output feature of the fifth C2F module is used as the detection feature P6 of level 6, and the detection feature P6 of level 6 is input into the detection head Detect-P6.
[0097] In this embodiment, the neck network realizes high-quality fusion of features at each level in the first-stage processing. Using convolution operations to change the size of the feature map increases the computational cost to a certain extent. However, in the second-stage processing, in order to reduce the model computational overhead, a top-down path aggregation module is adopted. Specifically, this module starts from the high-resolution feature map and passes information downward layer by layer through downsampling operations. At each level, the high-resolution feature map is horizontally connected and concatenated with the low-resolution feature map.
[0098] Through the coordinated operation of the two-stage processing of the neck network, not only the loss in the information transmission process is effectively reduced, but also the high-efficiency performance of the YOLO series algorithms is ensured. This design strategy finally screens and retains more critical and useful feature information for the detection branch of the head, which has far-reaching significance for significantly improving the accuracy of chest radiograph abnormality detection and reducing the false positive rate.
[0099] The training process of the object detection model of this application includes:
[0100] Step 1, construct a chest radiograph sample set. Each chest radiograph sample in the chest radiograph sample set includes a chest radiograph and the corresponding annotation data. The annotation data includes the bounding box coordinates, class labels, and confidence scores of one or more objects in the chest radiograph. The chest radiograph sample set is divided into a training set, a test set, and a validation set.
[0101] Step 2, construct the network structure of the object detection model.
[0102] Step 3: Iteratively train the object detection model using the training set. In each training, the object detection model outputs the lesion detection results of the chest X-rays in the training samples. Calculate the loss function based on the lesion detection results and the annotation data of the training samples, and update the network parameters of the object detection model using the gradient descent method according to the loss function until the training stop condition is reached. The training stop condition is that the number of training times reaches the preset maximum number of training times, or the value of the loss function converges. The loss function includes bounding box localization loss, confidence loss, and classification loss, and the classification loss can be the existing cross-entropy loss.
[0103] Step 4: Test and validate the object detection model after reaching the training stop condition using the test set and the validation set. When the test and validation are passed, obtain the final object detection model. If the test and validation are not passed, return to Step 3.
[0104] The present invention also discloses a chest X-ray lesion detection device based on enhanced feature extraction and fusion for implementing the above-mentioned chest X-ray lesion detection method based on enhanced feature extraction and fusion, including:
[0105] A chest X-ray acquisition module for acquiring chest X-rays;
[0106] A detection module for inputting the chest X-ray into the trained object detection model to obtain the lesion detection results; a marking output module for marking the lesion detection results on the chest X-ray and outputting the marked chest X-ray.
[0107] The object detection model includes:
[0108] A backbone network: including feature extraction networks of N levels connected in sequence. The feature extraction networks of N levels respectively extract feature maps of different scales of the chest X-ray; input the N - 1 feature maps obtained by the feature extraction networks from level 2 to level N into the neck network; where the level index of the feature extraction network is from 1 to N, the N - 1 feature maps include 1 reference feature map and N - 2 non-reference feature maps, the level index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3.
[0109] A neck network: taking the size of the reference feature map as a reference, performing size transformation on the non-reference feature maps; splicing the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a spliced feature map; performing fusion processing on the spliced feature map to obtain a fused feature map; performing distribution processing on the fused feature map to obtain hierarchical distribution feature maps of one or more preset distribution levels; performing inverse size transformation on the hierarchical distribution feature maps of one or more preset distribution levels and then fusing them with the feature maps corresponding to the preset distribution levels respectively to obtain fused features of one or more preset distribution levels, where the level index of the preset distribution level is greater than 2 and less than N; fusing the fused features of one or more preset distribution levels and the feature map of the Nth level to obtain multiple detection features.
[0110] Head output layer: Obtain the lesion detection result based on multiple detection features.
[0111] In this embodiment, the chest X-ray acquisition module, the detection module, and the marking output module correspond one-to-one to steps S1, S2, and S3 in the above-mentioned chest X-ray lesion detection method based on enhanced feature extraction and fusion, and the object detection model is the same as the object detection model in the above-mentioned chest X-ray lesion detection method based on enhanced feature extraction and fusion, which will not be elaborated here.
[0112] The present invention also discloses an electronic device. In one embodiment, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein,
[0113] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the chest X-ray lesion detection method provided by the present invention based on enhanced feature extraction and fusion.
[0114] As Figure 4 shown, it is a schematic structural diagram of an electronic device for the chest X-ray lesion detection method based on enhanced feature extraction and fusion provided by an embodiment of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as the chest X-ray lesion detection method program based on enhanced feature extraction and fusion.
[0115] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the chest X-ray lesion detection method based on enhanced feature extraction and fusion, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device and process data.
[0116] The memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In other embodiments, the memory 11 can also be an external storage device of the electronic device, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device. Further, the memory 11 can also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed in the electronic device and various types of data, such as the code of the chest radiograph lesion detection method program based on enhanced feature extraction and fusion, but also to temporarily store data that has been output or will be output.
[0117] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11 and at least one processor 10, etc.
[0118] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the electronic device and to display a visual user interface.
[0119] Figure 4 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 4The structures shown do not constitute a limitation on the electronic device, and it may include fewer or more components than those shown, or combine certain components, or have different component arrangements.
[0120] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0121] It should be understood that the embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.
[0122] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0123] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0124] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A chest radiograph lesion detection method based on enhanced feature extraction and fusion, characterized in that, Including: Obtain a chest X-ray; Input the chest X-ray into a trained object detection model to obtain lesion detection results; Mark the lesion detection results on the chest X-ray and output the marked chest X-ray; The object detection model includes: Backbone network: including N levels of feature extraction networks connected in sequence. The N levels of feature extraction networks extract feature maps of different scales of the chest X-ray respectively. Input the N - 1 feature maps obtained by the feature extraction networks from level 2 to level N into the neck network. Among them, the level index of the feature extraction network is from 1 to N, and the N - 1 feature maps include 1 reference feature map and N - 2 non-reference feature maps. The level index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3; Neck network: With the size of the reference feature map as a reference, perform size transformation on the non-reference feature maps. Concatenate the reference feature map and the size-transformed non-reference feature maps in the channel dimension to obtain a concatenated feature map. Perform fusion processing on the concatenated feature map to obtain a fused feature map. Perform distribution processing on the fused feature map to obtain level distribution feature maps of one or more preset distribution levels. After performing inverse size transformation on the level distribution feature maps of one or more preset distribution levels, fuse them with the feature maps corresponding to the preset distribution levels respectively to obtain fused features of one or more preset distribution levels. The level index of the preset distribution level is greater than 2 and less than N. Fuse the fused features of one or more preset distribution levels and the feature map of level N to obtain multiple detection features; Head output layer: Obtain lesion detection results based on multiple detection features.
2. The method for detecting chest radiograph lesions based on enhanced feature extraction and fusion according to claim 1, wherein, The performing size transformation on the non-reference feature maps with the size of the reference feature map as a reference includes: For non-reference feature maps with a size larger than the size of the reference feature map, the size transformation process is: perform convolution processing on the non-reference feature map using a first size transformation convolution based on a preset stride; For non-reference feature maps with a size smaller than the size of the reference feature map, the size transformation process is: perform bilinear interpolation processing and second size transformation convolution processing on the non-reference feature map in sequence.
3. The method for detecting chest radiograph lesions based on enhanced feature extraction and fusion according to claim 2, characterized in that, The performing inverse size transformation on the level distribution feature maps of one or more preset distribution levels includes: For a preset distribution level with a feature map size larger than the size of the reference feature map, the inverse size transformation process of the level distribution feature map of the preset distribution level is: perform bilinear interpolation processing and second size transformation convolution processing in sequence; For a preset distribution level with a feature map size smaller than the size of the reference feature map, the inverse size transformation process of the level distribution feature map of the preset distribution level is: perform convolution processing on the level distribution feature map of the preset distribution level using a first size transformation convolution based on a preset stride; No inverse size transformation is performed on the level distribution feature map of the preset distribution level where the reference feature map is located.
4. A chest radiograph lesion detection method based on enhanced feature extraction and fusion according to claim 1 or 2 or 3, characterized in that The performing fusion processing on the concatenated feature map to obtain a fused feature map includes: performing convolution processing, batch normalization processing, and activation function processing on the concatenated feature map in sequence.
5. The method for detecting chest radiograph lesions based on enhanced feature extraction and fusion according to claim 4, wherein The feature extraction network of level 1 includes a downsampling convolution; The feature extraction network of level N includes a downsampling convolution, a C2F module, and an SPPF module connected in sequence; The feature extraction network at any level between level 1 and level N includes a downsampling convolution and a C2F module connected in sequence.
6. The method for detecting chest radiograph lesions based on enhanced feature extraction and fusion according to claim 5, wherein, The feature extraction network at at least one level from level 3 to level N-1 is connected to the feature extraction network of the next level through a G-CA module, and / or the feature extraction network of level N further includes a G-CA module located between the C2F module and the SPPF module; Wherein, the G-CA module includes: A first branch for obtaining the global channel attention weight of the input feature map of the G-CA module; A second branch for obtaining the height-direction attention weight and the width-direction attention weight of the input feature map of the G-CA module; A multiplication weighting operation unit for element-wise multiplying the input feature map of the G-CA module with the global channel attention weight, the height-direction attention weight, and the width-direction attention weight.
7. The method for detecting chest radiograph lesions based on enhanced feature extraction and fusion according to claim 6, characterized in that, The first branch includes: A compression unit that compresses the height information and width information of each channel of the input feature map of the G-CA module into channel feature values, and the channel feature values of all channels form a feature vector; A first-branch convolution unit that performs convolution processing on the feature vector to obtain a first-branch feature map; A first-branch activation function processing that performs activation function processing on the first-branch feature map to obtain the weight of each channel, and the weights of all channels form the global channel attention weight.
8. A chest radiograph lesion detection method based on enhanced feature extraction and fusion according to claim 6 or 7, characterized in that, The second branch includes: A global average pooling unit that respectively performs global average pooling processing on the input feature map of the G-CA module in the height direction and the width direction to obtain a height-direction pooled feature and a width-direction pooled feature; A second-branch splicing unit that splices the height-direction pooled feature and the width-direction pooled feature to obtain a spliced pooled feature; A second-branch convolution unit that performs convolution processing on the spliced pooled feature to obtain a second-branch feature map; A distribution unit that sequentially performs batch normalization processing and activation function processing on the second-branch feature map, and then distributes the feature map after the activation function processing into a height feature sub-map and a width feature sub-map; A height-direction attention weight acquisition unit that sequentially performs convolution processing and activation processing on the height feature sub-map to obtain the height-direction attention weight; A width-direction attention weight acquisition unit that sequentially performs convolution processing and activation processing on the width feature sub-map to obtain the width-direction attention weight.
9. An enhanced feature extraction and fusion-based chest radiograph lesion detection device for implementing the enhanced feature extraction and fusion-based chest radiograph lesion detection method according to any one of claims 1-8, characterized in that, Includes: A chest X-ray acquisition module that acquires a chest X-ray; A detection module that inputs the chest X-ray into a trained object detection model to obtain a lesion detection result; A marking output module that marks the lesion detection result on the chest X-ray and outputs the marked chest X-ray; The object detection model includes: A backbone network: including a feature extraction network of N levels connected in sequence, and the feature extraction networks of N levels respectively extract feature maps of different scales of the chest X-ray; input the N-1 feature maps obtained by the feature extraction networks from level 2 to level N into the neck network; wherein, the level index of the feature extraction network is from 1 to N, the N-1 feature maps include 1 reference feature map and N-2 non-reference feature maps, the level index of the reference feature map is greater than 2 and less than N, and N is a positive integer greater than or equal to 3; Neck network: Taking the size of the reference feature map as a reference, perform size transformation on the non-reference feature map; concatenate the reference feature map and the size-transformed non-reference feature map in the channel dimension to obtain a concatenated feature map; perform fusion processing on the concatenated feature map to obtain a fused feature map; perform distribution processing on the fused feature map to obtain hierarchical distribution feature maps of more than one preset distribution level; after performing inverse size transformation on the hierarchical distribution feature maps of more than one preset distribution level, fuse them with the feature maps corresponding to the preset distribution levels respectively to obtain fused features of more than one preset distribution level, where the hierarchical index of the preset distribution level is greater than 2 and less than N; fuse the fused features of more than one preset distribution level and the feature map of level N to obtain multiple detection features. Head output layer: Obtain lesion detection results based on multiple detection features.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a chest radiograph lesion detection method based on enhanced feature extraction and fusion as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image recognition method, electronic equipment and storage medium
CN114821554A
Remote sensing image target detection method and system based on multi-scale semantic features
CN117079139A
Target detection method and system for cervical cell image
CN119559632A
Commutator inner side image defect detection method based on fusible feature pyramid
WO2024208100A1