An image segmentation method and an image segmentation apparatus
Patent Information
- Application Number
- CN202512050767.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-31
AI Technical Summary
尤其是针对于医疗影像,例如眼底图像,存在病变的目标部位往往只有几百微米的尺寸,在图像中的像素占比小于百分之一甚至千分之一,往往很难被精确识别到
(1)本申请的分割方法能在复杂的目标类别和目标形态情况下,准确地识别出小目标,可以实现对微米级尺寸目标的精准识别;
Smart Images

Figure CN121861282B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image segmentation method and an image segmentation apparatus. Background Technology
[0002] In image segmentation, small targets are a significant challenge due to their extremely small pixel footprint. This is especially true for medical images, such as fundus images, where lesions are often only a few hundred micrometers in size, occupying less than one percent or even one-thousandth of the image's pixels, making them difficult to identify accurately. Furthermore, the mixed and blurred shapes of different types of lesions can result in indistinct target edges, making it difficult for artificial intelligence (AI) models to detect lesions during the learning process. In other words, AI models may be dominated by large, non-lesion-related targets in the image, neglecting smaller lesion targets, leading to a failure to learn effective small target information and resulting in missed or misdiagnosed diseases. Summary of the Invention
[0003] This application provides an image segmentation method and an image segmentation apparatus, which are capable of learning effective small target information in small target images.
[0004] To address the aforementioned technical problems, this application provides the following technical solutions: In a first aspect, embodiments of this application provide an image segmentation method. This method can be executed by an image segmentation device, or by components of the image segmentation device, such as the processor, chip, or chip system of the image segmentation device. It can also be implemented by a logic module or software capable of realizing all or part of the functions of the image segmentation device. The image segmentation method provided in the first aspect includes: Small target regions in the image are cropped and enlarged multiple times in successive steps to obtain a segmented dataset of the enlarged target; the sampling window used for each step of cropping is increased sequentially, and the image is enlarged to the same fixed size after each cropping. The segmented dataset of the magnified target is input into the frozen AI model, and the network layers are gradually unfrozen to perform segmented training using the segmented data. The AI model outputs the image segmentation results corresponding to the small target.
[0005] In one possible implementation of the first aspect, the step of performing multiple progressive cropping and magnification processes on small target regions in the image includes: After cropping the small target area, a first cropped image is obtained. It is then determined whether the first cropped image includes a large target. If a large target is found in the first cropped image, the first cropped image is discarded, and the process continues to obtain the next first cropped image.
[0006] In one possible implementation of the first aspect, the step of performing multiple progressive cropping and magnification processes on small target regions in the image includes: When there are multiple small targets, first identify each small target separately from the image; During cropping, if multiple small target regions belong to the same sampling window and / or the distance between multiple small targets is less than the first threshold, the sampling windows of multiple small target regions are merged to obtain the second cropped image. The second cropped image is subjected to multiple progressive cropping and magnification processes to obtain a segmented dataset of the magnified target.
[0007] In one possible implementation of the first aspect, after obtaining the second cropped image, it is determined whether the second cropped image contains a large target; if the second cropped image contains a large target, the second cropped image is discarded, and the next second cropped image is obtained.
[0008] In one possible implementation of the first aspect, the stepwise unfreezing of network layers and segmented training using segmented data includes: First, the first network layer is unfrozen. The first image obtained after cropping and enlarging the small target region using the first sampling window is used as the first segment data and input into the AI model for training. The second network layer is then unfrozen to obtain a second image after cropping and enlarging the small target region using the second sampling window. This second image is then combined with the first image as the second segmented data and input into the AI model for training. The second sampling window is larger than the first sampling window. Then, the third network layer is unfrozen to obtain the third image after cropping and enlarging the small target region using the third sampling window. This image is then combined with the second image as the third segment data and input into the AI model for training. The third sampling window is larger than the second sampling window. Continue to unfreeze different network layers in stages for training. The sampling window corresponding to the segmented data used in each stage increases in size, and the segmented data includes the current sampling window data and the sampling window data of the previous stage; until all network layers are unfrozen and the segmented data training is completed.
[0009] In one possible implementation of the first aspect, the stepwise unfreezing of network layers and segmented training using segmented data includes: Large target data is added to the segmented data at any stage. The large target data is the data obtained by cropping and enlarging the large target region using the sampling window corresponding to the segmented data used in the current stage. The large target data is then input into the AI model for training. When training different network layers by unfreezing them in stages for small targets, the same training process is performed for large target data; however, in the last unfreezing training stage, the current sampling window data in the segmented data is the smallest cropped image that includes both small and large target regions. The entire original image and the minimally cropped image are input into the fully thawed AI model for training; The entire original image is then input into the fully thawed AI model for training, and the image segmentation results corresponding to small and large targets are output.
[0010] In one possible implementation of the first aspect, the method further includes: Obtain the current training order for training the AI model; Obtain the learning rate and regularization coefficient corresponding to the current training order; wherein, the learning rate decreases exponentially with the number of current training orders, and the regularization coefficient increases exponentially with the number of current training orders; The training level of the AI model is controlled based on the learning rate and regularization coefficient.
[0011] In one possible implementation of the first aspect, the AI model includes two UNET network architectures and two segmentation heads, wherein the first UNET network architecture and the first segmentation head complete the coarse training phase of the segmentation training, and the second UNET network architecture and the second segmentation head complete the fine training phase, which is used to further adjust and optimize the segmentation results of the coarse training.
[0012] In one possible implementation of the first aspect, the method further includes: Based on the segmentation result obtained from the first segmentation head of the current AI model, obtain the coordinates of the first k local maximum values, where k is a positive integer; Based on the coordinates of the k local maxima, the corresponding k regions of interest are cropped from the segmentation result; The k regions of interest are magnified to obtain the magnified k regions of interest; The amplified k regions of interest are input into the second UNET network architecture for fine training. The output of the second segmentation head is sent to the first segmentation head, which outputs the final segmentation result.
[0013] Secondly, embodiments of this application also provide an image segmentation apparatus, the image segmentation apparatus comprising: The magnification module is used to perform multiple progressive cropping and magnification processes on small target regions in an image to obtain a segmented dataset of magnified targets. The sampling window used for progressive cropping increases sequentially, and the image is magnified to the same fixed size after each cropping. The segmented training module is used to perform segmented training using the segmented dataset of the magnified target by gradually unfreezing the network layers. The prediction module is used by the AI model to output image segmentation results corresponding to small targets.
[0014] Compared with the prior art, this application has the following advantages: (1) The segmentation method of this application can accurately identify small targets in complex target categories and target shapes, and can achieve accurate identification of micron-sized targets; (2) By first increasing the pixel ratio of small targets, the AI model can learn the feature information of small targets in the early stage. In the subsequent training process, the information of small targets is effectively maintained and will not be lost due to interference from the information of large targets, thus solving the problem that existing technologies cannot learn the information of small targets.
[0015] (3) By using a phased training and coarse-fine joint training method, the segmentation accuracy of this application can be improved by at least 5% compared with the existing direct training method.
[0016] (4) Combine the thawing strategy for training so that the AI model can iterate stably. Attached Figure Description
[0017] Figure 1 A flowchart illustrating an image segmentation method provided in an embodiment of this application; Figure 2 A schematic diagram of the segmentation training process for cracks provided in an embodiment of this application; Figure 3 A schematic diagram of a raw fundus image to be segmented, provided for an embodiment of this application; Figure 4 The embodiments provided in this application provide for the application of the following: Figure 3 A schematic diagram of the image segmented according to the segmentation labels; Figure 5 This application provides a schematic diagram of a segmented thawing training process. Figure 6 This is a schematic diagram of the structure of an AI model provided in an embodiment of this application; Figure 7 This is a schematic diagram of the composition structure of an image segmentation device provided in an embodiment of this application. Detailed Implementation
[0018] This application provides an image segmentation method and an image segmentation apparatus, which can learn effective information about small targets in an image and accurately segment the small targets.
[0019] The embodiments of this application will now be described with reference to the accompanying drawings.
[0020] In the specification, claims, and accompanying drawings of this application, "at least one" refers to one or more items, and "more than one" refers to two or more items. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0021] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0022] like Figure 1 As shown, this application provides an image segmentation method that can be applied in engineering, medical, and detection fields. The method includes: 101. Perform multiple progressive cropping and magnification processes on small target regions in the image to obtain a segmented dataset of magnified targets; wherein, the sampling window used for progressive cropping increases sequentially, and the image is magnified to the same fixed size after each cropping.
[0023] The image includes at least one type of small target. Small targets typically refer to targets with a small pixel area in the image; in this embodiment, the pixel area of the small target occupies less than 1% of the total pixel area of the image. The type of small target is not limited; for example, a small target could be a small nodule in a lung image, a retinal tear lesion in a fundus image, or other small targets such as microaneurysms. The following example uses a retinal tear as an illustration. Small targets in the image can be obtained based on segmentation labels, such as background, optic disc, and retinal tear. The AI model in this embodiment can be an image segmentation model; the specific artificial intelligence algorithm used in the AI model is not limited here.
[0024] In this embodiment, the small target region in the image is first cropped multiple times in a progressive manner. Each cropping uses a sampling window of a different size, and the window size increases sequentially, without limiting the magnitude of each increase. After each cropping, images with different percentages of small target pixel area are obtained. Then, these images with different percentages of small target pixel area are enlarged until each image is enlarged to the size of the training sample image (e.g., 1024*1024 pixels), thus obtaining images at different cropping sizes, forming a segmented dataset of enlarged targets. It should be noted that as the sampling window size increases progressively, the percentage of pixels of the small target in the entire training image gradually decreases, and thus, its percentage in the entire training sample gradually decreases back to its original proportion. Through multiple cropping and enlargement processes, the AI model can be trained using the enlarged small targets, thereby improving the AI model's ability to segment small targets.
[0025] It should be noted that this application trains the AI model by enlarging small targets, but it cannot learn indefinitely under enlarged conditions. Otherwise, when a small target of the original size is input, the AI model will still not be able to perceive the information of the small target. Therefore, it is necessary to gradually reduce the pixel proportion of the small target in the entire training sample to the original proportion, which is equivalent to guiding the AI model step by step from the enlarged small target to the original small target.
[0026] In some embodiments of this application, a first cropped image is obtained after cropping the small target region. The method provided in the embodiments of this application further includes: Determine whether the first cropped image includes a large target. A large target usually refers to a target with a large pixel area in the image. In this embodiment, a large target refers to a target whose pixel area occupies a larger or much larger proportion of the pixel area of the entire image than a small target. It has more obvious characteristics than a small target. If a large target is found in the first cropped image, the first cropped image is discarded, and the process continues to obtain the next first cropped image.
[0027] In this method, the cropped image of the small target does not include the large target, thus ensuring the independence of small and large target information during the data acquisition stage and preventing the features of the large target from interfering with the AI model's feature learning of the small target. In fundus images, large targets can be structures such as the optic disc.
[0028] In some embodiments of this application, when there are multiple small targets, step 101 specifically includes: A1. Identify multiple small targets from the image; A2. When multiple small target regions are located within the same sampling window, and / or the distance between multiple small targets is less than the first threshold, the sampling windows of the multiple small target regions are merged to obtain the second cropped image; A3. Perform multiple successive cropping and magnification processes on the second cropped image to obtain a segmented dataset of the magnified target.
[0029] This process involves identifying multiple small targets in an image. Taking two small targets as an example, we obtain a first small target and a second small target. Next, we determine whether the two small target regions can be merged and cropped. The merging conditions are that the first and second small target regions belong to the same sampling window, and / or the distance between the first and second small targets is less than a first threshold, meaning the distance between the first and second small targets is very close. In this case, the first and second small target regions can be merged and cropped to obtain a second cropped image. Then, the second cropped image undergoes progressive cropping and magnification processing to obtain a segmented dataset of magnified targets. The specific methods for progressive cropping and magnification are the same as described above. This application's merging and cropping of multiple small targets ensures that the model perceives complete small target features, improving learning effectiveness.
[0030] In some embodiments of this application, further, after obtaining the second cropped image as described above, the method provided in the embodiments of this application also includes: Determine whether the second cropped image contains a large target; if a large target exists in the second cropped image, discard the second cropped image and continue to obtain the next second cropped image.
[0031] 102. Input the segmented dataset of the magnified target into the frozen artificial intelligence AI model, and perform segmented training using the segmented data by gradually unfreezing the network layers.
[0032] In this embodiment, the small target image is cropped and enlarged in step 101, and then the enlarged small target image is input into the AI model, making it easier for the AI model to learn the features of the enlarged target. In this embodiment, all network layers in the AI model are frozen and then trained in stages, i.e., partially unfrozen network layers are trained gradually. The weight parameters of the unfrozen network layers are updated during model training, while the weight parameters of the frozen network layers are not updated. This embodiment does not limit the number of unfrozen network layers, their distribution position in the AI model, or their function.
[0033] For example, in this embodiment of the application, a phased thawing training scheme can be adopted. Taking the hole lesion structure in the fundus image as a small target, in the initial case, the image is first cropped using the smallest sampling window that can completely contain the hole lesion structure. The cropped image is then enlarged to the size of the training sample image (1024*1024 pixels) to obtain the enlarged hole image. Then, the enlarged hole image is input into the AI model so that the AI model can learn the features of the hole from the beginning.
[0034] Because the edges of small targets in images are not obvious, the number of effective image data annotations is very small. If a complete model architecture is used for training from the beginning, it is easy to cause the model to overfit on the training set. In this embodiment, the complexity of the model can be gradually increased in stages, and training can be carried out with segmented datasets. That is, a simple model is used at the beginning to learn relatively simple crack features. As the complexity of crack samples increases, the complexity of the model also increases, thereby reducing the risk of model overfitting.
[0035] For example, suppose an AI model has 12 network layers. Using pre-trained weights trained on the large ImageNet dataset, the training is divided into six stages for unfreezing. The direction of unfreezing is not limited; the lower layers (those with the highest resolution) can be unfrozen first, or the higher layers (those with the lowest resolution) can be unfrozen first. For instance, the top two layers can be unfrozen first, leaving the subsequent 10 layers frozen, and trained using the first segment of data from the segmented dataset. Then, the remaining layers are unfrozen sequentially, and the remaining five stages are trained using segmented data from different stages. This allows the AI model to accurately segment small objects in an image. The weights of the unfrozen layers can be updated, while the weights of the frozen layers are not updated.
[0036] In some embodiments of this application, the specific steps of phased unfreezing of training include: B1. First, unfreeze the first network layer to obtain the first image after cropping and enlarging the small target area using the first sampling window. Use this image as the first segment data and input it into the AI model for training. B2. Unfreeze the second network layer to obtain the second image after cropping and enlarging the small target region using the second sampling window. Combine the second image with the first image as the second segmented data and input it into the AI model for training. The second sampling window is larger than the first sampling window. B3. Then, the third network layer is unfrozen to obtain the third image after cropping and enlarging the small target region using the third sampling window. This image is then combined with the second image as the third segment data and input into the AI model for training. The third sampling window is larger than the second sampling window. At this stage, the data from the first sampling window is discarded. B4. Continue to unfreeze different network layers in stages for training. The sampling window corresponding to the segmented data used in each stage increases in size, and the segmented data includes the current sampling window data and the sampling window data of the previous stage; until all network layers are unfrozen and the segmented data training is completed.
[0037] During the model training process using segmented data, data from other target categories, such as large targets, can be added. For example, a large target could be a spectral disk or other target data in the image; no limitation is made here. To illustrate that this application can achieve accurate segmentation of multi-class targets, an example of a segmented training method is given below when both small and large targets exist in the image.
[0038] B5. Add large target data to the segmented data in any of steps B1-B4. The large target data is the data obtained by cropping and enlarging the large target area using the sampling window corresponding to the segmented data used in the current stage. Input the large target data into the AI model for training. B6. When training different network layers by unfreezing them in stages for small targets, the same training process is performed for large target data; however, in the last unfreezing training stage, the current sampling window data in the segmented data is the smallest cropped image that includes both small and large target regions, that is, this stage takes into account the coexistence of different types of targets. B7. Input the entire original image and the minimally cropped image into the fully thawed AI model for training; then input the entire original image into the fully thawed AI model for training.
[0039] In this embodiment, by inputting both large and small target data into the AI model, the AI model can learn the respective features of small and large targets, avoiding the model being dominated by the larger target image with more obvious features too early. Then, the training data is transitioned to the complex situation where the two types of targets coexist, enabling the AI model to effectively segment small and large targets, thereby achieving accurate segmentation of different types of targets.
[0040] In some embodiments of this application, in addition to the aforementioned method steps, the image segmentation method may also include the following steps: C1. Get the current training order for training the AI model; C2. Obtain the learning rate and regularization coefficient corresponding to the current training order; where the learning rate decreases exponentially with the number of current training orders, and the regularization coefficient increases exponentially with the number of current training orders; C3. Control the training level of the AI model based on the learning rate and regularization coefficient.
[0041] In this embodiment, the AI model can be trained in segments, with each segment representing a training order. The learning rate and regularization coefficient are adjusted based on the current training order. The learning rate decreases exponentially with the number of current training orders, while the regularization coefficient increases exponentially with the number of current training orders. Therefore, the training level of the AI model is controlled by the learning rate and regularization coefficient, thereby improving the training efficiency of the AI model.
[0042] Next, we will explain the learning rate and regularization coefficient of AI models. The learning rate is used to control the convergence speed of parameter updates. The range of the learning rate is not limited and needs to be adjusted according to the specific task. The regularization coefficient is used to control the model training complexity by limiting the size of the model weights through a penalty term to prevent overfitting. The learning rate affects the direction of parameter updates, while the regularization coefficient constrains the parameter space; together, they determine the model's generalization ability.
[0043] 103. The AI model outputs the image segmentation results corresponding to the small target.
[0044] After segmenting the data and training it using an AI model, image segmentation results for small targets can be obtained. When multiple small targets exist, image segmentation results for each small target can be obtained. When a large target exists, image segmentation results for both small and large targets can be obtained simultaneously.
[0045] In some embodiments of this application, the AI model includes multiple network layers. Taking the first network layer (i.e., the part of the network layer that is unfrozen first) as an example, it includes: a residual block, a transform block, and a loss function. The residual block includes at least two convolutional layers, and the residual result is obtained by adding the convolution result output from at least two convolutional layers to the input features. The transform block includes a multi-head attention mechanism and a feedforward network. The multi-head attention mechanism generates global context features based on the normalized residual results, adds the global context features to the residual results to obtain the attention output, and inputs the attention output into the feedforward network. The feedforward network performs nonlinear transformations based on the attention output to obtain the output features. The loss function is used to calculate the loss based on the label and prediction result to obtain the total loss result. The loss function is used to calculate the loss value generated during the model training process. In this embodiment, the total loss result may include: the loss result corresponding to the background, the loss result corresponding to the large target, and the weighted fusion result of multiple loss results corresponding to the small target.
[0046] In the network structure of the aforementioned AI model, the model architecture used in this embodiment is a combination of a Convolutional Neural Network (CNN) and a transform. The aforementioned AI model can be implemented through the combination of CNN and transform. Figure 6 As shown, the AI model in this embodiment consists of two UNET network architectures and two segmentation heads. The first UNET network architecture and segmentation head segHead1 complete the coarse training stage (including the above-mentioned segmented training process), and the second UNET network architecture and segmentation head segHead2 complete the fine training stage. The fine training stage is used to further adjust and optimize the segmentation results of the coarse training.
[0047] In some embodiments of this application, the method provided in this application may further include the following steps: E1. Based on the segmentation result obtained from the segmentation head segHead1 of the current AI model, obtain the coordinates of the first k local maxima, where k is a positive integer; E2. Based on the coordinates of the k local maxima, crop out the corresponding k regions of interest from the segmentation result; E3. Enlarge the k regions of interest to obtain the enlarged k regions of interest; E4. Use the amplified k regions of interest as inputs to the second UNET network architecture for fine training. Send the output of the segmentation head segHead2 to the segmentation head segHead1, and the segmentation head segHead1 outputs the final segmentation result.
[0048] Specifically, in the coarse training stage, the segmentation head seghead1 outputs a channel prediction probability map. This map is set according to the data category input to the model and can be single-channel or multi-channel. In this embodiment, the data categories include holes, visual disks, and background, thus obtaining a three-channel prediction probability map. Then, the coordinates of the top k local maxima for each target category (excluding the background) are extracted from the probability map. k is a positive integer and can be set empirically. Next, k regions of interest corresponding to the target are cropped, enlarged, and then input into the AI model for fine training. In this embodiment, through coarse and fine training stages, the AI model is able to accurately segment the target.
[0049] The following section will illustrate the solution of this application with detailed application scenario examples.
[0050] Please see Figure 2 As shown, taking a small target as a crack and a large target as a visual disk as an example, the overall implementation process of this application embodiment may include: segmented dataset preparation, segmented unfreezing training, and fine training.
[0051] First, we will explain how to prepare the segmented dataset.
[0052] Figure 3 and Figure 4 As shown, Figure 3 This is the original fundus image. Figure 4 Is with Figure 3 The corresponding segmentation labels used are: background, optic disc, and eye hole. Other possible additions include the shallow detachment area (a structure associated with fundus lesions) and the fovea centralis. The background is defined as (0, ...). Figure 4 (White). The display panel 401 is (1, Figure 4 (Lower left part), fovea 404 is located in Figure 4 In the lower left part, crack 402 is (2, Figure 4 The upper right part), that is, the crack is Figure 4 The black part within the shallow detachment zone 403, where shallow detachment zone 403 is (3, Figure 4 (Upper right part). In cases where both a perforation and a shallow detachment area are present, the shallow detachment area and the perforation area are often inseparable. That is, the perforation and the shallow detachment area are related, with the shallow detachment area distributed around the perforation. This makes the edges of the perforation and / or the shallow detachment area very indistinct, making it difficult to accurately separate the perforation and / or the shallow detachment area.
[0053] This application uses a three-category classification of background, viewing disc, and crack as an example for illustration. The processing method for adding a shallow detachment area classification is the same.
[0054] Because the features of cracks are not as obvious as those of the view disk, and in many cases, the area occupied by cracks on the entire image is very small, possibly even less than 0.01% in extreme cases, this poses a significant challenge to accurate crack segmentation. If the AI model is relatively simple, its ability to extract and segment such small targets is very limited; if the model is more complex, it will be limited by the video memory of the graphics processing unit (GPU). Assuming the model uses a 1024*1024 pixel size as input, if the crack accounts for 0.01% and is 10*10 pixels in size, the model will easily ignore these inconspicuous crack features during training, resulting in a failure to learn crack information effectively.
[0055] In this embodiment, sampling windows of different sizes are used to crop and enlarge small target crack regions, resulting in an enlarged crack segment dataset. This transforms small target cracks, which are difficult to learn, into enlarged targets that are easy for the model to learn. For example, in this embodiment, a relatively small sampling cropping size (128*128) can be used initially to crop the complete crack, which is then enlarged to a size of 1024*1024. The overall cropping strategy is as follows: (1) The cutting size will be compared with the complete size of the crack, and the size that can include the complete size of the crack will be used as the cutting size; (2) If two crack areas are too close together, they will be merged into one trimming area; (3) If the clipping area contains a large target such as a viewing plate, then the clipping area is abandoned, thereby ensuring the independence of the viewing plate and the hole.
[0056] It should be noted that if other small target data or large target sampling data such as the optic disc are used in subsequent segmented training, the same cropping and magnification process as the above-mentioned hole should be adopted.
[0057] exist Figure 5 The diagram illustrates the segmented datasets used in different training phases, gradually transitioning from small to large sampling windows. Each sampled data point is enlarged to a fixed size required by the AI model; essentially, this reduces the area proportion of pores in the training data step by step.
[0058] After preparing the segmented dataset, the AI model is used for segmented unfreezing training. There are no restrictions on the direction and number of unfreezing layers in the model's network layers; the decision depends on the specific results. For example, regarding the number of unfreezing layers, if the model includes 4 blocks with a total of 12 convolutional layers, and if it is divided into 4 training stages, then 3 convolutional layers are unfrozen in each stage. Furthermore, the number of layers unfrozen each time can differ between the early, middle, and late stages of training. For instance, more layers might be unfrozen in the early stages, with a total of 4 unfreezing sessions, resulting in 4, 4, 2, and 2 network layers being unfrozen respectively. The number of unfreezing layers depends on the chosen training strategy and is not fixed; this embodiment does not impose any limitations. The following example illustrates the unfreezing training of 12 network layers, with two layers unfrozen sequentially and divided into 6 stages.
[0059] like Figure 5 As shown, the training data for the first thawing stage consists of small target rift images with a crop window size of 128*128. The training data for the second thawing stage uses small target rift images with crop windows of 192*192 and 128*128. From this stage onwards, the training data from the previous stage is mixed in for training to achieve a smooth transition. The training data for the third thawing stage uses small target rift images with crop windows of 256*256 and 192*192. Simultaneously, sampling data from a large target viewport is added at this stage, processed in the same way as for the rift images, i.e., obtaining viewport images with crop windows of 256*256 and 192*192 for training. Of course, viewport data can be added for training in any thawing stage; this embodiment uses the third stage as an example. The training data for the fourth and fifth thawing stages use crop images of 384*384 and 256*256, respectively. The training data includes sampled data of the hole and view disk at sizes of 512*512 and 384*384. Since this embodiment incorporates data on the large target view disk, in the sixth thawing stage, the training data used includes the minimum cropped region (including the hole and view disk) that simultaneously includes both small and large targets, as well as sampled data of the hole and view disk at a crop size of 512*512. After segmented training, the entire original image and the minimum cropped region (including the hole and view disk) are input into the fully thawed model for training. Then, the entire original image is input into the fully thawed model for training again, and the segmentation result is output. In this embodiment, the minimum cropped region refers to the smallest sampling window that can simultaneously and completely contain both the hole and view disk.
[0060] To avoid the visual disc competing with the crack for the AI model's attention from the very beginning, this embodiment adds large target image data for joint training starting in the third unfreezing stage. This prevents the AI model from being dominated by the visual disc, which has more obvious features, in the early stages of operation. After learning the crack features in the first two stages, the AI model has a certain segmentation ability for the crack.
[0061] For example, in this embodiment of the application, the learning rate and regularization coefficient of the AI model are designed according to the following formula: lr = max(1e-4 / pow(1.2,current_stage-2), 5e-6); decay = min(2e-5 * pow(3,current_stage-1),8e-4); Where *lr* is the initial learning rate for each stage, decreasing exponentially with the number of stages; *decay* is the regularization coefficient for each stage, used to prevent overfitting, increasing exponentially with the number of stages; *current_stage* is the current stage number; and *pow* refers to exponentiation.
[0062] The model architecture used in this application embodiment is a combination of Convolutional Neural Network (CNN) and transform.
[0063] Taking the Unet architecture as an example, it consists of residual blocks (resBlock), transform blocks (transformBlock), and a segmentation head (SegHead). The first two residual blocks maintain high-resolution local detail information, have strong local perception, and clear edge response, providing fine segmentation information for small targets and edge regions. The subsequent transformBlock provides global modeling capabilities, which can capture long-distance dependencies through a self-attention mechanism to model the whole image result, and can handle images with complex structures and strong background interference. Then, Unet's skip connections are used to integrate high and low layer semantic information to provide more refined segmentation results. Transform will generate huge GPU memory usage at high resolution, so the first two layers use residual blocks to compensate for the high-resolution information.
[0064] In this embodiment, transformBlock comprises two modules: a multi-head attention mechanism and a feed-forward network (FFN). The multi-head attention mechanism and the feed-forward network are wrapped by two normalization modules (LayerNorm) and feature fusion is performed through residual connections. The specific process is as follows: The input features are first normalized by LayerNorm and then fed into the multi-head attention mechanism to generate global context features; then, the global context features are added to the original input through residual connections. Subsequently, LayerNorm is applied to the result again, and the normalized result is sent to the feed-forward network. This feed-forward network can contain two fully connected layers and a Gaussian Error Linear Unit (GELU) activation. Finally, the result is output again through residual connections.
[0065] In addition, to enhance training stability and prevent overfitting, a random dropout layer can be added to the AI model in this embodiment to randomly discard some paths, thereby improving generalization ability.
[0066] The residual block structure is as follows: the input features first pass through two convolutional layers sequentially, and then the convolution result is added to the original input to form the output. Each of the two convolutional layers is followed by batch normalization (BatchNorm) and rectified linear function (ReLU) activation. If the input and output dimensions are inconsistent, this embodiment can also perform dimensionality matching through a 1×1 convolution.
[0067] The loss function is designed as follows: The segmentation head SegHead1 can output 3*1024*1024 prediction results, representing the background, optic disc, and pore, respectively. A Dice coefficient-based loss function is calculated for each of the three channels. The Dice coefficient is a statistical metric that measures the similarity between two sets or strings, ranging from 0 to 1, and is used in text similarity analysis, image segmentation (especially medical imaging), and other fields.
[0068] For the pore channel, additional Focal Loss and Tversky Loss are calculated. The coefficients of Tversky Loss are set to 0.3 and 0.7, which indicates enhanced segmentation ability for samples that are easily missed. If the coefficients of Tversky Loss are 0.5 and 0.5, then it is the normal dice loss.
[0069] Then, different weighting coefficients are applied to the losses for the three categories: Total loss = a * loss_dice 背景+b*loss_dice 视盘 +c*loss_dice 裂孔 +d*loss_focal 裂孔 +e*loss_tver 裂孔 .
[0070] The coefficient ae automatically adjusts with the number of training epochs. For example, in the first 300 epochs, a=0.3, b=1.0, c=3.0, d=2.0, e=2.0 can be set to strengthen the focus on the gradient of the lesion feature and reduce the interference of the background and optic disc feature gradients. In the middle 300 epochs, the weight coefficients of the lesion can be slightly adjusted to c=2.0, d=1.5, e=1.5, while a=0.3 and b=1 remain unchanged. In the later stage, the weight coefficients of the background and the lesion can be adjusted to a=0.5, d=1.0, e=1.0, while b=1 and c=2.0 remain unchanged. In addition, the Tversky Loss coefficient can be adjusted to 0.4, 0.6, etc. For the Tversky Loss coefficient that needs to be adjusted, the focus on false positives and false negatives can be different at different stages. For example, in the early stage, more attention is paid to false negatives, hoping to cover all lesion segments, and the model tries to "encircle" all possible targets. In later stages, less attention is paid to missed detections, and the focus shifts towards false positives, allowing the model to converge to a more accurate boundary. Specific weight parameters can be determined based on the specific training conditions.
[0071] The coarse-fine two-stage training mode used in this embodiment will be explained next, such as... Figure 6 As shown: In order to more accurately segment the background, holes, and viewing disc, embodiments of this application may employ a coarse-fine two-stage training mode.
[0072] The coarse-stage training includes the phased unfreezing training method described above, guiding the model to learn to segment cracks and achieving a certain degree of accuracy in segmenting cracks, but the segmentation effect for edges or the entire crack is not good enough. Then, a fine-stage training process needs to be added. The specific process is as follows: S01. Based on the segmentation prediction results obtained from the segmentation head segHead1 of the current AI model, obtain the coordinates of the first k local maximum values, where k is a positive integer, for example, k can be set to 10.
[0073] S02. Based on the coordinates of the k local maximum values, crop out the corresponding k regions of interest (ROIs) from the segmentation results. The size of the ROI can be set to 128*128. In this embodiment, the size of the ROI is not limited and can be flexibly set according to the application scenario.
[0074] S04. The ROI region is interpolated and enlarged to 1024*1024, and input into the second UNET network architecture corresponding to the fine-tuning stage for fine-tuning training. The output of segmentation head segHead2 is sent to segmentation head segHead1, which outputs the final segmentation result. SegHead1 outputs the loss value from the coarse-stage segmentation. 粗 The segmentation head segHead2 outputs the loss value for the fine-grained stage. 细 Then, the loss value corresponding to the fine-stage is added to the loss value of the coarse-stage according to a certain ratio. For example, the following calculation method can be used: Total loss = m * loss 粗 +n*loss 细 ; Here, m and n can be set and varied according to the number of iterations.
[0075] In the later inference stage, you can choose the prediction results output by the fine stage, reduce the prediction results to 128*128, and then fill and cover the corresponding positions in the segmentation head according to the local maximum coordinates, or you can directly use the prediction results of the segmentation head.
[0076] fine-stage loss value 细 The configuration can be achieved by using a combination of Dice Loss and Edge Loss, which allows the model to focus more on local edges and overall shape.
[0077] As illustrated by the foregoing embodiments, this application embodiment can employ a phased transition training method from magnified targets to smaller targets, combined with a defrosting strategy and variations in the loss function setting. Furthermore, this application embodiment can also employ a coarse-fine two-stage training method and network architecture to improve the detection performance for small targets.
[0078] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.
[0079] like Figure 7 As shown, an image segmentation apparatus 700 provided in this application embodiment includes: The magnification module 701 is used to perform multiple progressive cropping and magnification processes on small target regions in an image to obtain a segmented dataset of magnified targets; wherein, the sampling window used for progressive cropping increases sequentially, and the image is magnified to the same fixed size after each cropping. Segmented training module 702 is used to perform segmented training using the segmented dataset of the magnified target by gradually unfreezing the network layers. The prediction module 703 is used by the AI model to output the image segmentation results corresponding to small targets.
[0080] The methods for executing the various modules of the image segmentation apparatus can also be found in the foregoing. Figure 1 The image segmentation method shown will not be elaborated here.
[0081] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0082] Those skilled in the art will understand that the flowchart shown is merely an example in which the embodiments of this application can be implemented, and the scope of application of the embodiments of this application is not limited by any aspect of the flowchart.
[0083] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, and devices can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0085] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image segmentation method, characterized in that, include: Small target regions in the image are cropped and enlarged multiple times in successive steps to obtain a segmented dataset of the enlarged target; the sampling window used for each step of cropping is increased sequentially, and the image is enlarged to the same fixed size after each cropping. The segmented dataset of the magnified target is input into the frozen AI model, and the network layers are gradually unfrozen to perform segmented training using the segmented data. The AI model outputs the image segmentation result corresponding to the small target; The step of progressively unfreezing network layers and performing segmented training using segmented data includes: First, the first network layer is unfrozen. The first image obtained after cropping and enlarging the small target region using the first sampling window is used as the first segment data and input into the AI model for training. The second network layer is then unfrozen to obtain a second image after cropping and enlarging the small target region using the second sampling window. This second image is then combined with the first image as the second segmented data and input into the AI model for training. The second sampling window is larger than the first sampling window. Then, the third network layer is unfrozen to obtain the third image after cropping and enlarging the small target region using the third sampling window. This image is then combined with the second image as the third segment data and input into the AI model for training. The third sampling window is larger than the second sampling window. Continue to unfreeze different network layers in stages for training. The sampling window corresponding to the segmented data used in each stage increases in size, and the segmented data includes the current sampling window data and the sampling window data of the previous stage; until all network layers are unfrozen and the segmented data training is completed.
2. The method according to claim 1, characterized in that, The process of performing multiple progressive cropping and magnification operations on small target regions in the image includes: After cropping the small target area, a first cropped image is obtained. It is then determined whether the first cropped image includes a large target. If a large target is found in the first cropped image, the first cropped image is discarded, and the process continues to obtain the next first cropped image.
3. The method according to claim 1, characterized in that, The process of performing multiple progressive cropping and magnification operations on small target regions in the image includes: When there are multiple small targets, first identify each small target separately from the image; During cropping, if multiple small target regions belong to the same sampling window and / or the distance between multiple small targets is less than the first threshold, the sampling windows of multiple small target regions are merged to obtain the second cropped image; The second cropped image is subjected to multiple progressive cropping and magnification processes to obtain a segmented dataset of the magnified target.
4. The method according to claim 3, characterized in that, After obtaining the second cropped image, it is determined whether the second cropped image contains a large target; if the second cropped image contains a large target, the second cropped image is discarded and the next second cropped image is obtained.
5. The method according to claim 1, characterized in that, The step of progressively unfreezing network layers and performing segmented training using segmented data includes: Large target data is added to the segmented data at any stage. The large target data is the data obtained by cropping and enlarging the large target region using the sampling window corresponding to the segmented data used in the current stage. The large target data is then input into the AI model for training. When training different network layers by unfreezing them in stages for small targets, the same training process is performed for large target data; however, in the last unfreezing training stage, the current sampling window data in the segmented data is the smallest cropped image that includes both small and large target regions. The entire original image and the minimally cropped image are input into the fully thawed AI model for training; The entire original image is then input into the fully thawed AI model for training, and the image segmentation results corresponding to small and large targets are output.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the current training order for training the AI model; Obtain the learning rate and regularization coefficient corresponding to the current training order; wherein, the learning rate decreases exponentially with the number of current training orders, and the regularization coefficient increases exponentially with the number of current training orders; The training level of the AI model is controlled based on the learning rate and regularization coefficient.
7. The method according to any one of claims 1 to 6, characterized in that, The AI model includes two UNET network architectures and two segmentation heads. The first UNET network architecture and the first segmentation head complete the coarse training phase of the segmentation training, while the second UNET network architecture and the second segmentation head complete the fine training phase. The fine training phase is used to further adjust and optimize the segmentation results of the coarse training.
8. The method according to claim 7, characterized in that, The method further includes: Based on the segmentation result obtained from the first segmentation head of the current AI model, obtain the coordinates of the first k local maximum values, where k is a positive integer; Based on the coordinates of the k local maxima, the corresponding k regions of interest are cropped from the segmentation result; The k regions of interest are magnified to obtain the magnified k regions of interest; The amplified k regions of interest are input into the second UNET network architecture for fine training. The output of the second segmentation head is sent to the first segmentation head, which outputs the final segmentation result.
9. An image segmentation apparatus, characterized in that, include: The magnification module is used to perform multiple progressive cropping and magnification processes on small target regions in an image to obtain a segmented dataset of magnified targets. The sampling window used for progressive cropping increases sequentially, and the image is magnified to the same fixed size after each cropping. The segmented training module is used to perform segmented training using the segmented dataset of the magnified target by gradually unfreezing the network layers. The prediction module is used by the AI model to output image segmentation results corresponding to small targets; The segmented training module is specifically used for: First, the first network layer is unfrozen. The first image obtained after cropping and enlarging the small target region using the first sampling window is used as the first segment data and input into the AI model for training. The second network layer is then unfrozen to obtain a second image after cropping and enlarging the small target region using the second sampling window. This second image is then combined with the first image as the second segmented data and input into the AI model for training. The second sampling window is larger than the first sampling window. Then, the third network layer is unfrozen to obtain the third image after cropping and enlarging the small target region using the third sampling window. This image is then combined with the second image as the third segment data and input into the AI model for training. The third sampling window is larger than the second sampling window. Continue to unfreeze different network layers in stages for training. The sampling window corresponding to the segmented data used in each stage increases in size, and the segmented data includes the current sampling window data and the sampling window data of the previous stage; until all network layers are unfrozen and the segmented data training is completed.
Citation Information
Patent Citations
Training method of image semantic segmentation model and server
CN110363210A
Medical image segmentation model training method, medical image segmentation method and medical image segmentation device
CN113205528A