Fluorescence in-situ hybridization image result analysis method and system based on deep learning and medium
By constructing SEAM-Unet++ and YOLO-SEM models, cell contour segmentation and fluorescent probe detection were performed on fluorescent in situ hybridization images, which solved the problems of low image analysis efficiency and insufficient accuracy in the prior art, and achieved high-precision cell and probe distribution analysis.
Patent Information
- Application Number
- CN202510070853.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to efficiently analyze and process complex fluorescence in-situ hybridization images, resulting in low cell contour segmentation accuracy and long manual detection time, increasing artificial error and eye fatigue.
Using a deep learning-based method, SEAM-Unet++ image segmentation model and YOLO-SEM small object detection algorithm model were constructed. Through these models, cell contour segmentation and fluorescent probe detection were performed on fluorescent in-situ hybridization images, and the relationship diagram of cell and probe distribution was merged into.
It significantly improves the segmentation accuracy of cell contours and the detection accuracy of small targets in FISH images, reduces manual processing time, reduces artificial errors, and improves the efficiency of medical diagnosis.
Smart Images

Figure CN120013883A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image analysis technology, and more specifically, to a method, system and medium for analyzing fluorescence in situ hybridization image results based on deep learning. Background Art
[0002] Fluorescence in situ hybridization can identify and detect precise DNA sequences or RNA molecules in cell or tissue samples through staining. This technology can quantify and locate specific genes, providing key information for disease diagnosis and treatment. However, in FISH images taken under a microscope, the number of cells is large and the nucleic acid sequences are chaotic. For researchers, processing and analyzing images is a time-consuming and laborious task because it can easily cause eye fatigue and lead to misjudgment.
[0003] In recent years, the combination of medical imaging and computer science has become increasingly close, and medical image detection methods based on deep learning have become a hot research topic. In the development history of medical image analysis, the initial methods can be traced back to the 1970s to the 1990s. During that period, medical image analysis mainly applied multiple low-level pixel processing techniques (such as edge and line detection filters, region growing algorithms) and mathematical modeling (such as fitting lines, circles, and ellipses) to build a composite rule-based system to solve specific tasks. There are many similarities with the expert systems with multiple conditional judgment statements that were popular in the field of artificial intelligence at the same time. These methods are called old-fashioned artificial intelligence, and their robustness and fault tolerance are usually low.
[0004] By the late 1990s, supervised learning techniques based on training data became increasingly popular in medical image analysis. Some representative approaches include active shape models, atlas methods, and the use of feature extraction and statistical classifiers. These pattern recognition or machine learning methods are still very popular today and form the basis of many commercial medical image analysis systems.
[0005] With the rise of deep learning technology, especially the successful application of deep learning models such as convolutional neural networks (CNN), medical image analysis has ushered in a new era. Deep learning methods not only perform well in medical image classification and segmentation by learning features in data, but also have made significant progress in tasks such as target detection and lesion recognition. These methods not only improve accuracy, but also have higher robustness and generalization capabilities, making them the cutting-edge technology in the current field of medical image analysis. Summary of the invention
[0006] The purpose of the embodiments of the present application is to provide a method, system and medium for analyzing fluorescence in situ hybridization image results based on deep learning, which can effectively process complex image scenes, support the analysis of various medical images, improve the efficiency of medical diagnosis, reduce human errors, and enhance the cell analysis and pathology detection capabilities in scientific research and clinical applications.
[0007] To achieve the above objectives, this application provides the following technical solutions:
[0008] In a first aspect, the present application provides a method for analyzing fluorescence in situ hybridization image results based on deep learning, comprising the following specific steps:
[0009] S1. Build the SEAM-Unet++ image segmentation model;
[0010] S2. Build the YOLO-SEM small target detection algorithm model;
[0011] S3. The original image of the B channel of the fluorescence in situ hybridization image is used as input and input into the SEAM-Unet++ image segmentation model for segmentation, the outline of each cell is segmented, and they are extracted separately to obtain the accurate cell segmentation map of the B channel;
[0012] S4. The original image of the R channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a red fluorescent probe distribution map;
[0013] S5. The original image of the G channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a green fluorescent probe distribution map;
[0014] S6. The obtained cell precise segmentation map, red fluorescence distribution map and green fluorescence distribution map are combined into one map to obtain a relationship map of cell and probe distribution.
[0015] The SEAM-Unet++ segmentation model includes an encoder module, a decoder module, a dense skip connection module and a SEAM module in the order of image processing;
[0016] The encoder module consists of multiple convolutional layers and downsampling layers. The input image first passes through the convolutional layer to extract image features such as cell shape, texture, and cell outline, and then passes through the downsampling layer to reduce the size of the feature map;
[0017] The decoder module consists of multiple upsampling layers and convolutional layers. The feature map output by the encoder module passes through the convolutional layer to restore the spatial resolution of the image, and then passes through the upsampling layer to restore the size of the feature map;
[0018] The dense skip connection module performs multiple processing and fusion between the features of different resolutions output by all convolutional layers of the decoder module and the encoder module. These skip connections enhance the communication between features of different resolutions and make feature fusion more complete.
[0019] The SEAM module divides the feature map into small blocks and performs embedding processing on each block.
[0020] The YOLO-SEM small target detection algorithm model includes a feature extraction network module, a feature fusion module and a detection head module. The feature extraction network module gradually extracts feature information of different levels through a series of convolutional layers, and extracts multi-scale semantic features from the input image. The feature fusion module integrates features from different levels of the feature extraction network module through a feature pyramid to form a top-down feature fusion method, while enhancing the model's perception ability of multi-scale targets for fusing these features. The detection head module generates a bounding box and category prediction of the target based on the features fused by the feature fusion module.
[0021] The construction of the SEAM-Unet++ image segmentation model is specifically as follows:
[0022] Obtain n high-resolution images to form an original high-resolution image set;
[0023] Divide the high-resolution image collection into training and validation sets;
[0024] Preprocessing each high-resolution image in the training set and each high-resolution image in the validation set to obtain a preprocessed training set and a preprocessed validation set;
[0025] Input the training set into the neural network model, train the model and update the model parameters;
[0026] The validation set is input into the model neural network model after parameter update to verify the model parameters and complete the construction of the SEAM-Unet++ image segmentation model.
[0027] In a second aspect, an embodiment of the present application provides a fluorescence in situ hybridization image result analysis system based on deep learning, the system comprising: a memory and a processor, the memory comprising a program of a fluorescence in situ hybridization image result analysis method based on deep learning, and the program of the fluorescence in situ hybridization image result analysis method based on deep learning is executed by the processor to implement the following steps: construct a SEAM-Unet++ image segmentation model; construct a YOLO-SEM small target detection algorithm model; take the original image of the B channel of the fluorescence in situ hybridization image as input, input it into the SEAM-Unet++ image segmentation model for segmentation, segment the outline of each cell, and extract them separately to obtain an accurate cell segmentation map of the B channel; take the original image of the R channel of the fluorescence in situ hybridization image as input, input it into the YOLO-SEM small target detection algorithm model for detection, and obtain a red fluorescent probe distribution map; take the original image of the G channel of the fluorescence in situ hybridization image as input, input it into the YOLO-SEM small target detection algorithm model for detection, and obtain a green fluorescent probe distribution map; merge the obtained cell accurate segmentation map, red fluorescence distribution map and green fluorescence distribution map into one map to obtain a relationship map of cell and probe distribution.
[0028] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a program code. When the program code is executed by a processor, the steps of the fluorescence in situ hybridization image result analysis method based on deep learning as described above are implemented.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The present invention significantly improves the segmentation accuracy of cell contours in FISH images by introducing the SEAM-UNet++ algorithm with integrated attention mechanism, and can effectively segment cells that adhere to each other. At the same time, the YOLO-SEM algorithm significantly improves the detection accuracy of small targets in low-resolution and noisy images. This method greatly reduces the manual processing time, reduces the eye fatigue caused by long-term manual inspection of images by researchers, and improves the accuracy of detection. In addition, the algorithm of the present invention can effectively process complex image scenes, support the analysis of various medical images, improve the efficiency of medical diagnosis, reduce human errors, and enhance the cell analysis and pathology detection capabilities in scientific research and clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 It is a flow chart of the method for detecting the fusion state of fluorescent probes for fluorescence in situ hybridization and analyzing the results provided by the present invention;
[0033] Figure 2 This is the structure diagram of SEAM-Unet++;
[0034] Figure 3 It is the structure diagram of SEAM module;
[0035] Figure 4 This is a comparison chart of cell segmentation experiments;
[0036] Figure 5 This is the structure diagram of the basic module of YOLO;
[0037] Figure 6 It is the working mechanism of the SPD module;
[0038] Figure 7 It is the structural diagram of the ECA module;
[0039] Figure 8 It is the CE module structure diagram;
[0040] Fig. 9 It is the structural diagram of the YOLO-SEM model;
[0041] Fig.10 This is a comparison chart of the results of the YOLO-SEM model and other models. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0043] The terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0044] The terms "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and should not be understood as indicating or implying relative importance, nor should they be understood as requiring or implying any such actual relationship or order between these entities or operations.
[0045] like Fig. 9 As shown, an embodiment of the present invention provides a method for analyzing fluorescence in situ hybridization image results based on deep learning, comprising the following specific steps:
[0046] S1. Build the SEAM-Unet++ image segmentation model;
[0047] S2. Build the YOLO-SEM small target detection algorithm model;
[0048] S3. The original image of the B channel of the fluorescence in situ hybridization image is used as input and input into the SEAM-Unet++ image segmentation model for segmentation, the outline of each cell is segmented, and they are extracted separately to obtain the accurate cell segmentation map of the B channel;
[0049] S4. The original image of the R channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a red fluorescent probe distribution map;
[0050] S5. The original image of the G channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a green fluorescent probe distribution map;
[0051] S6. The obtained cell precise segmentation map, red fluorescence distribution map and green fluorescence distribution map are combined into one map to obtain a relationship map of cell and probe distribution.
[0052] The construction of the SEAM-Unet++ image segmentation model is specifically as follows:
[0053] S1-1, obtain n high-resolution images to form an original high-resolution image set I, I∈(I1,I2,…,I i ,…,I n ),Ii is the i-th high-resolution image, i∈(1,…,n);
[0054] S1-2, divide the high-resolution image set I into training set I t and validation set I v , For training set I t The i-th image in, i∈(1,…,m), m is the training set I t The number of images in For the validation set I v The i-th image in , i∈(1,…,u), u is the number of images in the validation set;
[0055] S1-3, training set I t Each high-resolution image and validation set I v Each high-resolution image is preprocessed to obtain the preprocessed training set h t And the preprocessed validation set h v , is the preprocessed training set h t The i-th preprocessed image in is the preprocessed validation set h v The i-th preprocessed image in ;
[0056] S1-4, the initial features Figure X 0(W(width)×H(height)×C(number of channels)) is input into the neural network model, W is the width, H is the height, C is the number of channels, and two 3×3 convolution kernels W1 are used to convolution the features. Figure X 0 performs convolution operation to obtain features Figure X ′(W×H×C), and then X′ is subjected to maximum pooling to reduce the spatial resolution to obtain the feature map
[0057] S1-5, the characteristics Figure X 1 Upsampling to restore features Figure X The spatial resolution of 1 is used to obtain the features Figure X 11 (W×H×C), then the feature Figure X 0 and features Figure X 11 Perform skip connection to enhance feature information fusion to obtain features Figure X 12 (W×H×C);
[0058] S1-6, the characteristics Figure X 1 Use two 3×3 convolution kernels W1 to perform feature Figure X 1Perform convolution operation to obtain feature map Then X′ is subjected to maximum pooling to reduce the spatial resolution and obtain the feature map The characteristics Figure X 2 Upsampling to restore features Figure X The feature map is obtained with a spatial resolution of 2 Then the feature Figure X 1 and Features Figure X 21 Perform skip connection to enhance feature information fusion and obtain feature map Then the features Figure X 22 Upsampling to get features Figure X 23 (W×H×C) and X 12 , X0 is skipped to obtain the feature Figure X 24 (W×H×C);
[0059] S1-7, the characteristics Figure X 2 Use two 3×3 convolution kernels W1 to perform feature Figure X 2Perform convolution operation to obtain feature map Then X′ is subjected to maximum pooling to reduce the spatial resolution and obtain the feature map The characteristics Figure X 3 Upsampling to restore features Figure X The spatial resolution of 3 obtains the feature map Then the feature Figure X 2 and Features Figure X 31 Perform skip connection to enhance feature information fusion and obtain feature map Then the features Figure X 32 Upsampling to get feature map After and X 22 , X1 performs skip connection to obtain feature map Then the features Figure X 34 Upsampling to get features Figure X 35 (W×H×C), then the characteristics Figure X 24 With features Figure X 35 , X0 is skipped to obtain the feature Figure X 36 (W×H×C);
[0060] S1-8, the characteristics Figure X 3 Use two 3×3 convolution kernels W1 to perform feature Figure X 2Perform convolution operation to obtain feature map Then X′ is subjected to maximum pooling to reduce the spatial resolution and obtain the feature map The characteristics Figure X 4 Upsampling to restore features Figure X The spatial resolution of 4 is used to obtain the feature map Then the feature Figure X 3 and feature map Perform skip connection to enhance feature information fusion and obtain feature map Then the features Figure X 42 Upsampling to get feature map After and X 32 , X2 is skipped to obtain the feature map Then the features Figure X 44 Upsampling to get feature map Then the features Figure X 34 With features Figure X 45 , X1 performs skip connection to obtain feature map feature Figure X 46 Upsampling to get features Figure X 47 (W×H×C), then the characteristics Figure X 36 With features Figure X 47 , X0 is skipped to obtain the feature Figure X 48 (W×H×C);
[0061] S1-9, the characteristics Figure X 48 That is, X of U-net++ 0,4 After the feature map of the layer is passed to the SEAM module, the Patch Embedding module in each of the three parallel CSMM modules in the SEAM module divides the feature map into small blocks for embedding, introduces nonlinear activation and normalizes it;
[0062] S1-10. After embedding, depth-wise convolution is used in the CSMM module to reduce the amount of multiplication and addition operations, and then the nonlinear expression ability is enhanced through GELU function and normalization. The feature information is fused through point-by-point convolution, and the final normalized output is performed. The outputs of the three CSMM modules are summarized and globally averaged pooled to obtain the feature map y (W×H×C). The specific formula is as follows:
[0063]
[0064] Among them, Con i Indicates the first feature Figure X 48 Through a depth-wise separable convolution, the depth-wise separable convolution is composed of depth-wise convolution and point-wise convolution in sequence, where the depth-wise convolution converts the feature Figure X 48 The C channels in are divided into C groups for convolution, and then the output of the depth-wise convolution is used as the input of the point-wise convolution to obtain the output of the CSMM module. avg_pool means performing an average pooling operation on the outputs of the three CSMMs to obtain a feature map y.
[0065] S1-11. Calculate the attention weight g(x) of the feature map y, which reflects the importance of the feature.
[0066] The formula is as follows
[0067] g(x)=Sigmoid(Linear fc (y))
[0068] Among them, Linear fc Perform a linear transformation on the average feature map y;
[0069] The Sigmoid activation function is mainly used to map the input to the output in the range of (0,1):
[0070]
[0071] Sigmoid is a smooth differentiable function that can effectively calculate the gradient in the back-propagation algorithm to update the model parameters;
[0072] S1-12, weight the obtained attention weight g(x) and y calculated by S1-8 to obtain the weighted feature map y′. The specific formula is as follows:
[0073] y′=g(x)⊙y
[0074] Here, the ⊙ symbol represents element-by-element multiplication. Finally, the natural exponential function is used to process the weighted feature map y′:
[0075] output = x⊙exp(y′)
[0076] Output is the output result of the model, namely the cell segmentation map.
[0077] The specific method of constructing the YOLO-SEM small target detection algorithm model is:
[0078] S2-1. Convolution operation is performed on the feature map T with a size of W×H×C using a 1×1 convolution kernel, changing the size of the feature map channel C to C out , but the size of the feature image is not changed, and the feature map T0 (W×H×C out );
[0079] S2-2, extract cell features from feature map T0 using a convolution kernel size of 3×3 to obtain feature map T1 (W×H×C out );
[0080] S2-3, the feature map T1 is divided into sub-feature sequence blocks F through the SPD module n,m , where n represents the horizontal coordinate and m represents the vertical coordinate.
[0081] F 0,0 =F[0:H:step,0:W:step],F 1,0 =F[1:H:step,0:W:step],...,F step-1,0 =F[step-1:H:step,0:W:step]…
[0082] …
[0083] F 0,step-1 =F[0:H:step,step-1:W:step],F 1,step-1 =F[1:H:step,step-1:W:step],...,F step-1,step-1 =F[step-1:H:step,step-1:W:step]
[0084] The expression represents the extraction of feature blocks from the input feature map through different starting positions and step sizes.
[0085] where F 0,0 It means extracting feature blocks from the coordinate of feature map F(0,0), which is the upper left corner of feature map T1, with step=2 as the step length. H and W input the height and width of the feature map.
[0086] These sub-feature sequences are arranged along channel C out The image size is obtained by splicing in the direction of The feature map T 11 ;
[0087] S2-4, feature map Enter the C2f module and transform the feature map T along the direction of channel C 11 Divided into 2 feature maps feature Figure X0 Without any processing, the feature map Y0 is convolved twice with a convolution kernel of size 3×3 to obtain the feature map Then jump connect with feature map Y0 to get feature map The feature map Y1 is convolved twice with a convolution kernel of size 3×3 to obtain the feature map Then jump connect with feature map Y1 to get feature map The feature map Y2 is convolved twice with a convolution kernel of size 3×3 to obtain the feature map Then jump connect with feature map Y2 to get feature map After n times, we get the feature map Then the features Figure X 0 and Y1 to Y n After splicing along the channel direction, a 1×1 convolution operation is performed to obtain the feature map. The specific formula is as follows:
[0088]
[0089] T1 is divided into two parts, X0 and Y0, conv 3×3 Indicates a convolution operation with a convolution kernel size of 3×3, conv 1×1 Represents a convolution operation with a convolution kernel size of 1×1, and Concat represents the feature map channel C that will be obtained out Direction dimension stitching;
[0090] S2-5, the feature map T 12 Use a convolution kernel size of 3×3 to get the feature map
[0091] The feature map T2 is obtained by S2-3 Then the feature map T 21 After S2-4, the feature map is obtained
[0092] S2-6, feature map T 22 Use a convolution kernel size of 3×3 to get the feature map
[0093] Pass the feature map T3 through S2-3 to obtain the feature map Then the feature map T 31 After S2-4, the feature map is obtained
[0094] S2-7, feature map T 32 Use a convolution kernel size of 3×3 to get the feature map The feature map T4 is obtained by S2-3 Then the feature map T 41 After S2-4, the feature map is obtained
[0095] S2-8, the feature map T 42 The input is sent to the SPFF module and passed through a 3×3 convolution kernel W2 to obtain the feature map S1 for maximum pooling to obtain the feature map S 11 Perform maximum pooling to obtain S 12 After performing a maximum pooling to the feature map After three times of maximum pooling, three feature maps with different resolutions are obtained. These three feature maps are combined with the feature map T 42 Splicing along the channel direction and then undergoing a 1×1 convolution operation to obtain the feature map The specific formula is as follows:
[0096]
[0097] where conv 3×3 Indicates that the feature map T 42 Perform convolution operation, MaxPool2d represents the maximum pooling operation, conv 1×1 represents a convolution operation with a convolution kernel size of 1×1.
[0098] S2-9, the feature map S 14 Upsample to restore the resolution and get the feature map With the feature map T 31 After splicing, a 1×1 convolution operation is performed to obtain the feature map
[0099] T 51 =conv 1×1 (Concat(up(T5),T 31 ))
[0100] Where up represents the upsampling operation, conv 1×1 Represents a convolution operation with a convolution kernel size of 1×1, and Concat represents the feature map channel C that will be obtained out Direction dimension stitching;
[0101] S2-10, the feature map T 51 Input into the C2f module to get the feature map Then the feature map T52 Upsample to feature map After that, the feature map T 21 After splicing along the channel C direction, a 1×1 convolution operation is performed to obtain the feature map
[0102] S2-11, the feature map T 54 Input to the CE module to transform the feature map T 54 Divided into 2 feature maps S0 S0 is not processed in any way, and F0 is input into the ECA module after a convolution operation. The feature map F is obtained by a global average pooling. 01 (1×1×4C out ), and then the feature map F 01 Perform a one-dimensional convolution, the size of the one-dimensional convolution kernel k is adaptively determined according to the number of channels C. The specific steps are as follows:
[0103] The nonlinear function selected is C = φ(k) = 2 γ×K-b This method captures the relationship between features and uses the inverse mapping function to infer the appropriate convolution kernel k size from the number of channels C.
[0104]
[0105] Among them, K represents the size of the one-dimensional convolution kernel, C represents the number of channels, φ is the mapping function, λ is the scaling coefficient in the linear function, b is the offset in the linear function, γ is the scaling coefficient in the nonlinear function, ψ is the inverse mapping function, |t| odd represents the odd number closest to t,
[0106] After obtaining the convolution kernel size, the feature map F 01 Convolution is performed to obtain feature maps Then perform a convolution operation with a convolution kernel size of 3×3 to obtain the feature map Then jump connect with feature map F0 to get feature map F1 is input into the ECA module after another convolution operation, and the feature map is obtained after a global average pooling. Then the feature map F 11 Perform a one-dimensional convolution to get Then a convolution operation with a convolution kernel size of 3×3 is performed to obtain the feature map Then F 13 Jump connection with feature map F1 to obtain feature map F2 is input into the ECA module after another convolution operation, and the feature map is obtained after a global average pooling. Then the feature map F 21 Perform a one-dimensional convolution to get Then a convolution operation with a convolution kernel size of 3×3 is performed to obtain the feature map Then F 23 Perform a jump connection with the feature map F2 to obtain the feature map After n times, we get the feature map Then the feature maps S0 and F1 to F n Along channel C out After the directions are concatenated, a 1×1 convolution operation is performed to the feature map The specific formula is as follows:
[0107]
[0108] T1 is divided into two parts, S0 and F0, conv 3×3 represents a convolution operation with a convolution kernel size of 3×3, conv represents a one-dimensional convolution operation, conv 1×1 Represents a convolution operation with a convolution kernel size of 1×1, and Concat represents the convolution of the obtained feature map along the channel direction C out Perform splicing Perform splicing;
[0109] S2-12, the feature map T 54 Upsample to restore the resolution and get the feature map With the feature map T 11 Along channel C out After the directions are spliced, a 1×1 convolution operation is performed to obtain the feature map The specific formula is as follows:
[0110] T 61 =conv 1×1 (Concat(up(T 54 ),T 11 )
[0111] Where up represents the upsampling operation, conv 1×1 Represents a convolution operation with a convolution kernel size of 1×1, and then the feature map T 71 After steps S2-3, the feature map is obtained
[0112] S2-13, the feature map T 62 A convolution operation with a convolution kernel size of 3×3 is used to obtain the feature map The feature map T7 is obtained by S2-3 Then the feature map T 71 With the feature map T 55 Along the channel direction C out After splicing, a 1×1 convolution operation is performed to obtain the feature map Feature map T 72 S2-4 obtains the feature map
[0113] S2-14, the feature map T 73 A convolution operation with a convolution kernel size of 3×3 is used to obtain the feature map The feature map T8 is obtained by S2-3 Then the feature map T 81 With the feature map T 52 Along channel C out After the directions are spliced, a 1×1 convolution operation is performed to obtain the feature map Feature map T 82 S2-11 obtains the feature map
[0114] S2-15, the feature map T 83 A convolution operation with a convolution kernel size of 3×3 is used to obtain the feature map The feature map T9 is obtained by S2-3 Then the feature map T 91 With the feature map S 14 Along channel C out After the directions are spliced, a 1×1 convolution operation is performed to obtain the feature map Feature map T 92 S2-4 obtains the feature map
[0115] S2-16, the feature map T 62 The input is sent to the Detece layer, which contains two branches, the regression branch and the classification branch. The regression branch outputs the position and size of the bounding box, and the classification branch outputs the category and confidence to determine whether the predicted box is a fluorescent spot or background. The size can be obtained as The bounding box location, size, category and confidence;
[0116] S2-17, the feature map T 73 After step S2-16, the resolution is The bounding box location, size, category and confidence;
[0117] S2-18, the feature map T 83 After step S2-16, the resolution is The bounding box location, size, category and confidence;
[0118] S2-19, the feature map T 93 After step S2-16, the resolution is The bounding box location, size, category and confidence;
[0119] S2-20. After obtaining the categories, confidences of the classification branches of different sizes and the bounding box coordinates of the regression branches, sort them in descending order according to the confidence of each bounding box, select the bounding box with the highest confidence from the sorted list, mark it as selected, and add it to the final detection result list. For each remaining bounding box, calculate the IOU of each bounding box with the selected bounding box.
[0120] The calculation formula for calculating the IOU of each bounding box and the selected bounding box is as follows:
[0121] The coordinates of the upper left corner and lower right corner of the predicted box and the true label box are MPDIoU is a loss function based on IoU, and the calculation formula is as follows;
[0122]
[0123]
[0124] L MPDIoU It is the loss function based on MPDIoU. The specific formula is as follows:
[0125] The coordinates of the upper left and lower right corners are The width and height of the image are H and W,
[0126]
[0127] So MPDIoU is
[0128] L MPDIoU The calculation formula of L MPDIoU =1-MPDIoU
[0129] Among them, H and W represent the height and width of the image respectively.
[0130] After obtaining the RGB three-channel FISH image, split the image into three single-channel images of R, G, and B. FISH image detection can be divided into two steps. The first step is to segment the outline of each cell on the B channel and extract them separately. After obtaining the mask of the selected cell, cover it on the R and G channels respectively, detect the fluorescent spots, and finally merge them into one picture, so as to obtain the chromosome distribution status of this cell. Figure 1 As shown,
[0131] UNet++ is an improvement on the UNet network structure, and its design goal is to improve the performance and accuracy of image segmentation tasks. They have similar structures, but Unet++ introduces several key features to better capture the features and semantic information of the image. UNet++ consists of an encoder part and a decoder part. The encoder part consists of convolutional layers, downsampling layers, and jump connections, which are used to extract image features and gradually reduce the size of feature maps. The goal of this part is to capture features at different levels from low to high. The decoder introduces multi-level feature fusion, that is, feature fusion is performed at different levels of the decoder, which can more comprehensively understand and capture information at different scales. The feature maps at each level are fused through densely connected blocks, allowing the network to more comprehensively utilize multi-level feature information. UNet++ also uses jump connections to connect the feature maps in the encoder with the feature maps in the decoder. This jump connection mechanism helps to combine low-level detail information and high-level semantic information to improve the accuracy of the segmentation results. This feature is particularly useful when processing images with rich details and complex semantic structures, but the effect is still not ideal when processing cells that stick to each other in cell contour images. To solve this problem, we noticed that Unet++'s X 0,4 The layer fuses the features of all previous layers. In order to enhance feature learning and representation, its output is passed to the SEAM module. Considering reasons such as preventing overfitting, we did not add too many attention mechanisms in other positions. SEAM can reduce the negative impact of mutual occlusion between cells on the segmentation results. By enhancing the attention to occluded cells, it helps to accurately segment the occluded cell contours, thereby improving the robustness of the segmentation task. In addition, the cell contour image is blurry and affected by noise. The attention mechanism of SEAM can significantly reduce the interference of background noise on the segmentation results, because it can enhance the detail information in the cell contour image, improve the detail and accuracy of the segmentation, and enable the model to better understand and capture the tiny features in the image. The structure of SEAM-Unet++ proposed in this paper is as follows: Figure 2 .
[0132] The first part of the SEAM module is to obtain multi-scale features using different patches. Patch refers to a small area or sub-image in an image. It contains local information and can be used to detect features, textures, edges, etc. in the image. Then, a depth-separable convolution with residual connection is used, which can reduce the number of parameters and prevent the gradient vanishing problem. The outputs of the convolutional layers of different depths are then combined using point-by-point convolutions. Then, two fully connected layers are used to fuse the information of each channel, which enables the model to learn the relationship between occluded contours and unoccluded contours to compensate for the loss of occluded scenes. Finally, the output of the fully connected layer is processed with an exponential function to expand the range from [0,1] to [1,e], which can make the result more tolerant to position errors.
[0133] Since there is only one category in the cell contour segmentation task, we only used one patch_size. Experiments have shown that when the patch_size is 7, the model's segmentation effect on the cell contour is most significant. To enhance the anti-interference performance, we used three parallel CSMM modules. This strategy helps to reduce certain noise or uncertainties, thereby improving the output stability of the model, while avoiding adding more convolution kernels to prevent the model from being too complex and reducing the risk of overfitting. Then, we use depthwise separable convolution to learn the relationship between spatial dimensions and channels, and finally average their outputs. The above steps help improve the accuracy of cell contour segmentation when cells are occluded by each other. The SEAM structure is as follows: Figure 3 As shown,
[0134] Through comparative experiments, it can be seen that SEAM-Unet++ performs best in terms of accuracy, proving that it successfully captures the area of the target object and accurately distinguishes the target from the background. It should be noted that accurately segmenting the stacked cell contours is a key step in extracting the contours of individual cells, which directly affects the determination of cell properties. Figure 4 As shown in the figure, when segmenting clustered cells, other models all show different degrees of contour adhesion on the mask image, while seam-unet++ has a significant improvement in the ability to completely segment contours.
[0135] The YOLO-SEM model includes several basic modules: C2f (Convolution to Fully Connected), Bottleneck, SPPF (Spatial Pyramid pool-fast), Detect, and Conv. The C2f module uses cross-stage partial feature fusion to integrate low-level and high-level feature maps. This integration significantly improves the detection accuracy and processing speed of the model. The bottleneck architecture reduces the feature map channels and reduces the computational burden. It combines residual connections to alleviate vanishing gradients and includes a 3x3 convolution layer to expand the receptive field. SPPF is a key component with a spatial pyramid pooling function that enables the model to handle various object sizes in a single image by aggregating features from multiple receptive fields. The Detect module processes the output from the Neck module, which integrates the inputs from the C2f, Bottleneck, SPPF, Detect, and Conv modules, such as Figure 5 shown.
[0136] SPD (space-to-depth) is a space-to-depth layer. The SPD module downsamples the feature map but retains all the information in the channel dimension, so there is no information loss.
[0137] Assuming that the size of each feature map F is W×H×C, the feature map is divided into sub-feature sequences F m,n :
[0138]
[0139] Next, these sub-feature sequences are connected along the channel dimension to obtain a sequence of size The feature map F′ is then passed to the next layer. The SPD module is shown as follows Figure 6 .
[0140] The ECA module significantly improves the performance of the model with relatively small additional computational effort. First, the global average pooling (GAP) method is used to obtain the aggregate feature map Box1, which is extracted into a single real value to obtain the feature χ avg ∈R (W×H×C) ,
[0141]
[0142] ECA generates channel weights by performing a fast 1D convolution of size k, where k is adaptively generated through a mapping of the channel dimension C. The coverage of the interaction (i.e., the size k of the 1D convolution kernel) is proportional to the channel dimension, i.e., there is a mapping φ:
[0143] C=φ(k).
[0144] If the mapping is a linear function, that is, φ(k) = λ×kb, its ability to represent feature relationships is too limited, so the linear function is extended to a nonlinear function. As we all know, the channel dimension C (the number of filters) is usually set to 2, so the mapping relationship between the two is as follows:
[0145] C=φ(k)=2 γ×K-b
[0146] The adaptive method of kernel size K can be determined as:
[0147]
[0148] where |t odd Represents the odd number closest to t, and then uses the Sigmoid function to get the activation value of the one-dimensional convolution output. ECA does not reduce the channel dimension, and considers the correspondence between channels, reducing the interference caused by noise and thus improving the noise reduction ability of the model. ECA modules are as follows Figure 7 shown.
[0149] In order to better capture the correlation information between different channels, we introduce the ECA module in C2f. Since the C2f module contains multiple convolutional layers and complex feature transformations, by introducing the ECA module, we can enhance the mutual correlation between different channels, thereby obtaining richer and more discriminative feature representations. This change can enhance feature representation and performance in complex networks, thereby helping to improve accuracy and generalization capabilities.
[0150] Bottleneck is the core component of C2f, which usually operates on local feature maps. We embed ECA in Bottleneck, which means that the network can selectively enhance task-related channels to better capture the key features of the object, and the number of repetitions n of Bottleneck can be flexibly controlled, so that the number of times ECA is executed can be determined according to specific tasks and computing resources. Bottleneck contains two convolution modules. The convolution operation is responsible for extracting and transforming features, while ECA is responsible for weighting the channel attention of these features. Therefore, ECA is placed between the two convolution modules to ensure that it can operate on the output of the convolution layer, and then the weighted features are input to the next convolution layer for further processing, so that the attention mechanism of the ECA module can affect the entire feature map. The CE module structure is as follows: Figure 8 shown.
[0151] YOLO-SEM consists of a backbone, a neck, and a head. The backbone network is used to extract relevant image information for use by subsequent network layers. This component improves efficiency and performance while reducing the computational complexity of feature extraction. The neck is located between the backbone and the head, which optimizes the use of extracted features and facilitates feature fusion. The head uses the features extracted by the neck to identify the target. The YOLO-SEM model structure is as follows Fig. 9 shown.
[0152] Comparative experiments show that the YOLO-SEM model used in this study performs best among all the comparative models in detecting fluorescent spots. Due to the small size of fluorescent spots, low image resolution and the presence of noise, the YOLO model that is not optimized for small target detection shows a significant performance degradation when facing these challenges. Although YOLOv8-P2 introduced a small target detection head and made some improvements in detecting small targets, there are still a large number of missed detections, which cannot meet the needs of high-precision detection. In comparison, the YOLO-SEM model has been specially optimized to more effectively cope with challenges such as small fluorescent spots, low-resolution images and noise interference. It not only excels in detecting tiny fluorescent spots, but also maintains a high accuracy and recall rate under these difficult conditions, demonstrating its adaptability and reliability. Experimental comparisons such as Fig.10 shown.
[0153] The embodiment of the present application provides a fluorescence in situ hybridization image result analysis system based on deep learning, the system comprising: a memory and a processor, the memory comprising a program of a fluorescence in situ hybridization image result analysis method based on deep learning, and the program of the fluorescence in situ hybridization image result analysis method based on deep learning is executed by the processor to implement the following steps: construct a SEAM-Unet++ image segmentation model; construct a YOLO-SEM small target detection algorithm model; take the original image of the B channel of the fluorescence in situ hybridization image as input, input it into the SEAM-Unet++ image segmentation model for segmentation, segment the outline of each cell, and extract them separately to obtain an accurate cell segmentation map of the B channel; take the original image of the R channel of the fluorescence in situ hybridization image as input, input it into the YOLO-SEM small target detection algorithm model for detection, and obtain a red fluorescent probe distribution map; take the original image of the G channel of the fluorescence in situ hybridization image as input, input it into the YOLO-SEM small target detection algorithm model for detection, and obtain a green fluorescent probe distribution map; merge the obtained cell accurate segmentation map, red fluorescence distribution map and green fluorescence distribution map into one map to obtain a relationship map of cell and probe distribution.
[0154] An embodiment of the present application provides a computer-readable storage medium, which stores program code. When the program code is executed by a processor, the steps of the fluorescence in situ hybridization image result analysis method based on deep learning as described above are implemented.
[0155] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0156] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0157] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0159] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0160] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0161] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0162] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for analyzing fluorescence in situ hybridization image results based on deep learning, characterized in that: The specific steps include: S1. Build the SEAM-Unet++ image segmentation model; S2. Build the YOLO-SEM small target detection algorithm model; S3. The original image of the fluorescence in situ hybridization image B channel is used as input and input into the SEAM-Unet++ image segmentation model for segmentation, the outline of each cell is segmented, and they are extracted separately to obtain the accurate cell segmentation map of the B channel; S4. The original image of the R channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a red fluorescent probe distribution map; S5. The original image of the G channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a green fluorescent probe distribution map; S6. The obtained cell precise segmentation map, red fluorescence distribution map and green fluorescence distribution map are combined into one map to obtain a relationship map of cell and probe distribution.
2. The method for analyzing fluorescence in situ hybridization images based on deep learning according to claim 1, characterized in that: The SEAM-Unet++ segmentation model includes an encoder module, a decoder module, a dense skip connection module and a SEAM module in the order of image processing; The encoder module consists of multiple convolutional layers and downsampling layers. The input image first passes through the convolutional layer to extract image features such as cell shape, texture, and cell outline, and then passes through the downsampling layer to reduce the size of the feature map; The decoder module consists of multiple upsampling layers and convolutional layers. The feature map output by the encoder module passes through the convolutional layer to restore the spatial resolution of the image, and then passes through the upsampling layer to restore the size of the feature map; The dense skip connection module performs multiple processing and fusion between the features of different resolutions output by all convolutional layers of the decoder module and the encoder module. These skip connections enhance the communication between features of different resolutions and make feature fusion more complete. The SEAM module divides the feature map into small blocks and performs embedding processing on each block.
3. The method for analyzing fluorescence in situ hybridization images based on deep learning according to claim 1, characterized in that: The YOLO-SEM small target detection algorithm model includes a feature extraction network module, a feature fusion module and a detection head module. The feature extraction network module gradually extracts feature information of different levels through a series of convolutional layers, and extracts multi-scale semantic features from the input image. The feature fusion module integrates features from different levels of the feature extraction network module through a feature pyramid to form a top-down feature fusion method, while enhancing the model's perception ability of multi-scale targets for fusing these features. The detection head module generates a bounding box and category prediction of the target based on the features fused by the feature fusion module.
4. The method for analyzing fluorescence in situ hybridization images based on deep learning according to claim 1, characterized in that: The construction of the SEAM-Unet++ image segmentation model is specifically as follows: Obtain n high-resolution images to form an original high-resolution image set; Divide the high-resolution image collection into training and validation sets; Preprocessing each high-resolution image in the training set and each high-resolution image in the validation set to obtain a preprocessed training set and a preprocessed validation set; Input the training set into the neural network model, train the model and update the model parameters; The validation set is input into the model neural network model after parameter update to verify the model parameters and complete the construction of the SEAM-Unet++ image segmentation model.
5. A fluorescence in situ hybridization image result analysis system based on deep learning, characterized in that: The system includes: a memory and a processor, wherein the memory includes a program of a fluorescence in situ hybridization image result analysis method based on deep learning, and when the program of the fluorescence in situ hybridization image result analysis method based on deep learning is executed by the processor, the following steps are implemented: constructing a SEAM-Unet++ image segmentation model; constructing a YOLO-SEM small target detection algorithm model; taking the original image of the B channel of the fluorescence in situ hybridization image as input, inputting it into the SEAM-Unet++ image segmentation model for segmentation, segmenting the outline of each cell, and extracting them separately, to obtain an accurate cell segmentation map of the B channel; taking the original image of the R channel of the fluorescence in situ hybridization image as input, inputting it into the YOLO-SEM small target detection algorithm model for detection, and obtaining a red fluorescent probe distribution map; The original image of the G channel of the fluorescence in situ hybridization image is used as input and input into the YOLO-SEM small target detection algorithm model for detection to obtain a green fluorescent probe distribution map; the obtained cell precise segmentation map, red fluorescence distribution map and green fluorescence distribution map are merged into one map to obtain a relationship map of cell and probe distribution.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, and when the program code is executed by a processor, the steps of the fluorescence in situ hybridization image result analysis method based on deep learning are implemented as described in any one of claims 1 to 4.
Citation Information
Cited By
Intelligent full-automatic diagnosis and analysis method and device for anterior segment diseases, electronic equipment and storage medium
CN120953259A