Segmentation model training method and apparatus, and computer readable recording medium

By amplifying data and generating mixed samples for medical image segmentation models, the effectiveness reduction caused by insufficient sample number is solved, and efficient training effect is achieved in the case of limited samples.

CN120014254APending Publication Date: 2025-05-16INVENTEC PUDONG TECH CORPOARTION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311532180.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

Smart Images

  • Figure CN120014254A_ABST
    Figure CN120014254A_ABST
Patent Text Reader

Abstract

The invention provides a segmentation model training method. The segmentation model training method comprises the following steps: inputting a plurality of first sample groups in a large sample set into a data amplification model to generate a plurality of amplified sample groups; generating a plurality of mixed sample groups based on a plurality of second sample groups in the small sample set; inputting the plurality of mixed sample sets into a data amplification model to generate a plurality of amplified mixed sample sets; and training a segmentation model according to the plurality of amplified sample groups and the plurality of amplified mixed sample groups, including: pre-training the segmentation model according to the plurality of amplified sample groups; and performing fine tuning training on the segmentation model corresponding to the plurality of amplified mixed sample groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a segmentation model training method, device and non-volatile computer-readable recording medium, and in particular to a segmentation model training device, segmentation model training method and non-volatile computer-readable recording medium for deep neural network (deep learning). Background Art

[0002] In recent years, in the field of machine learning, segmentation models are trained to automatically segment specific areas in images. The training of segmentation models can be applied to different fields. For example, in the field of medicine, accurate image segmentation of special patterns (such as wounds or lesions) on human skin plays a key role in skin-related care and diagnosis in the medical field. The technology of automatic segmentation has made great progress through deep neural networks (DNN). However, since the acquisition and annotation of medical images requires a lot of time, the current challenge in training segmentation models is that the number of samples is insufficient and training can only be performed on limited samples, which results in reduced performance.

[0003] Therefore, how to train the segmentation model when the number of samples is limited is one of the problems to be solved in this field. Summary of the invention

[0004] One aspect of the present application discloses a segmentation model training method, comprising the following steps: inputting multiple first sample groups in a large sample set into a data augmentation model to generate multiple augmented sample groups; generating multiple mixed sample groups based on multiple second sample groups in a small sample set; inputting the multiple mixed sample groups into the data augmentation model to generate multiple augmented mixed sample groups; and training a segmentation model based on the multiple augmented sample groups and the multiple augmented mixed sample groups, comprising: pre-training the segmentation model based on the multiple augmented sample groups; and fine-tuning the segmentation model corresponding to the multiple augmented mixed sample groups.

[0005] In some embodiments, generating a plurality of the mixed sample groups based on the plurality of second sample groups in the small sample set includes: capturing a target image in the first image based on a first image and a first mask of one of the plurality of the second sample groups; acquiring a target area in the second image based on a second image and a second mask of another of the plurality of the second sample groups; and pasting the target image to the target area in the second image to generate a third image of one of the plurality of the mixed sample groups; a ratio value between an area of ​​the target image and an area of ​​the target area is within a ratio value range.

[0006] In some embodiments, it also includes: adjusting at least one of the size and angle of the target image according to an adjustment parameter to generate an adjusted target image, an area of ​​the adjusted target image is the same as an area of ​​the target area; pasting the adjusted target image to the target area in the second image to generate the third image of the plurality of mixed sample groups; and superimposing the first mask with the second mask to generate a third mask of the plurality of mixed sample groups, including: adjusting the first mask according to the adjustment parameter to generate an adjusted mask; and superimposing the adjusted mask with the second mask to generate the third mask of the plurality of mixed sample groups.

[0007] In some embodiments, the method further comprises: adjusting at least one of the sizes, angles, colors and positions of the plurality of first images of the plurality of mixed sample groups according to the data augmentation model to generate a plurality of first augmented images of the plurality of amplified mixed sample groups; adjusting at least one of the sizes, angles, colors and positions of the plurality of second images of the plurality of first sample groups according to the data augmentation model to generate a plurality of second augmented images of the plurality of amplified sample groups; adjusting at least one of the sizes, angles, colors and positions of the plurality of first masks of the plurality of mixed sample groups corresponding to the plurality of first images of the plurality of mixed sample groups to generate a plurality of first augmented masks of the plurality of amplified mixed sample groups; and adjusting at least one of the sizes, angles, colors and positions of the plurality of second masks of the plurality of first sample groups corresponding to the plurality of second images of the plurality of first sample groups to generate a plurality of second augmented masks of the plurality of second amplified sample groups.

[0008] Another aspect of the present application discloses a segmentation model training device. The segmentation model training device includes a memory and a processor. The memory is used to store the segmentation model and the data augmentation model. The processor is coupled to the memory to perform the following steps: inputting multiple first sample groups in a large sample set into the data augmentation model to generate multiple augmented sample groups; generating multiple mixed sample groups based on multiple second sample groups in a small sample set; inputting multiple mixed sample groups into the data augmentation model to generate multiple augmented mixed sample groups; and training the segmentation model based on the multiple augmented sample groups and the multiple augmented mixed sample groups, including: pre-training the segmentation model based on the multiple augmented sample groups; and fine-tuning the segmentation model corresponding to the multiple augmented mixed sample groups.

[0009] In some embodiments, the processor is further used to: capture a target image in the first image based on a first image of one of the plurality of second sample groups and a first mask; obtain a target area in the second image based on a second image of another of the plurality of second sample groups and a second mask; and paste the target image to the target area in the second image to generate a third image of one of the plurality of mixed sample groups; a ratio value between an area of ​​the target image and an area of ​​the target area is within a ratio value range.

[0010] In some embodiments, the processor is further used to: adjust at least one of the size and angle of the target image according to an adjustment parameter to generate an adjusted target image, an area of ​​the adjusted target image is the same as an area of ​​the target area; paste the adjusted target image to the target area in the second image to generate the third image of the plurality of mixed sample groups; and superimpose the first mask with the second mask to generate a third mask of the plurality of mixed sample groups, including: adjusting the first mask according to the adjustment parameter to generate an adjusted mask; and superimposing the adjusted mask with the second mask to generate the third mask of the plurality of mixed sample groups.

[0011] In some embodiments, the processor is further configured to: adjust at least one of the sizes, angles, colors, and positions of the plurality of first images of the plurality of mixed sample groups according to the data augmentation model to generate a plurality of first augmented images of the plurality of amplified mixed sample groups; adjust at least one of the sizes, angles, colors, and positions of the plurality of second images of the plurality of first sample groups according to the data augmentation model to generate a plurality of second augmented images of the plurality of amplified sample groups; adjust at least one of the sizes, angles, colors, and positions of the plurality of first masks of the plurality of mixed sample groups corresponding to the plurality of first images of the plurality of mixed sample groups to generate a plurality of first augmented masks of the plurality of amplified mixed sample groups; and adjust at least one of the sizes, angles, colors, and positions of the plurality of second masks of the plurality of first sample groups corresponding to the plurality of second images of the plurality of first sample groups to generate a plurality of second augmented masks of the plurality of second amplified sample groups.

[0012] Another aspect of the present application discloses a non-volatile computer-readable recording medium for storing a computer program. When the computer program is executed, one or more processing components will be caused to perform multiple operations including: inputting multiple first sample groups in a large sample set into a data augmentation model to generate multiple augmented sample groups; generating multiple mixed sample groups based on multiple second sample groups in a small sample set; inputting multiple mixed sample groups into the data augmentation model to generate multiple augmented mixed sample groups; and training a segmentation model based on the multiple augmented sample groups and the multiple augmented mixed sample groups, including: pre-training the segmentation model based on the multiple augmented sample groups; and fine-tuning the segmentation model corresponding to the multiple augmented mixed sample groups.

[0013] In some embodiments, the operations further include: capturing a target image in the first image based on a first image of one of the second sample groups and a first mask; obtaining a target area in the second image based on a second image of another of the second sample groups and a second mask; and pasting the target image to the target area in the second image to generate a third image of one of the mixed sample groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Shown is a schematic diagram of a segmentation model training device illustrated in some embodiments of the present application.

[0015] Figure 2 Shown is a schematic diagram of a segmentation model training method illustrated in some embodiments of the present application.

[0016] Figure 3 Some embodiments of the present application are shown Figure 2 A diagram of one of the steps in .

[0017] Figure 4 Shown are schematic diagrams of images of a sample set and an amplified image of an amplified sample set according to some embodiments of the present application.

[0019] Figure 5 Some embodiments of the present application are shown Figure 2 A flowchart showing one of the steps in .

[0020] Figure 6 Shown is a schematic diagram of a small sample set illustrated in some embodiments of the present application.

[0021] Figure 7 Shown is a schematic diagram of capturing a target image according to some embodiments of the present application.

[0022] Figure 8Shown is a schematic diagram of a target sample set illustrated in some embodiments of the present application.

[0023] Fig. 9 Shown is a schematic diagram of a method for obtaining a target area according to some embodiments of the present application.

[0024] Fig.10 Shown is a schematic diagram of a mixed sample group illustrated in some embodiments of the present application.

[0025] Fig.11 Some embodiments of the present application are shown Figure 2 Schematic diagram of one of the steps in .

[0026] Fig.12 Some embodiments of the present application are shown Figure 2 Schematic diagram of one of the steps in .

[0027] Component number description

[0028] 100 Segmentation Model Training Device

[0029] 110 Processor

[0030] 120 Memory

[0031] AM Data Augmentation Model

[0032] SM segmentation model

[0033] 200 Segmentation Model Training Method

[0034] S210,S230,S250,S270 Steps

[0035] LI1,LI2,LI3,LI4 Image

[0036] LM1,LM2,LM3,LM4 Shield

[0037] LD1,LD2,LD3,LD4 sample group

[0038] LS Large Sample Set

[0039] ALI1, ALI2, ALI3, ALI4 amplification images

[0040] ALM1,ALM2,ALM3,ALM4 Amplification Shield

[0041] ALD1, ALD2, ALD3, ALD4 amplification sample set

[0042] ALS Amplification Sample Set

[0043] Pa1, pa2, pa3, pa4 positioning points

[0044] Pb1, pb2, pb3, pb4 positioning points

[0045] A0,A1,A2,A11,A12 area

[0046] S232, S234, S236 Steps

[0047] SS small sample set

[0048] SD1,SD2,SD3,SD4 sample group

[0049] SI1,SI2,SI3,SI4 images

[0050] SM1,SM2,SM3,SM4 Shielded

[0051] TP Area

[0052] TS target sample set

[0053] TD1, TD2, TD3, TD4 target sample group

[0054] TI1,TI2,TI3,TI4 target images

[0055] TM1, TM2, TM3, TM4 Shielding

[0056] TA Target Area

[0057] MI images

[0058] MM Shield

[0059] MD mixed sample group

[0060] AMI Amplified Image

[0061] AMM Amplification Shield

[0062] AMD Amplified Pooled Panel

[0063] AMS Amplification Pool

[0064] PW pre-training parameters DETAILED DESCRIPTION

[0065] The following disclosure provides many different embodiments or examples for implementing different features of the present application. The components and configurations in the specific examples are used to simplify the present case in the following discussion. Any examples discussed are only used for illustrative purposes and do not limit the scope and significance of the present application or its examples in any way.

[0066] Please refer to Figure 1. Figure 1The diagram is a schematic diagram of a segmentation model training device 100 according to some embodiments of the present application. In some embodiments, the segmentation model training device 100 includes a processor 110 and a memory 120. In terms of connection, the processor 110 is coupled to the memory 120.

[0067] The segmentation model training device 100 shown in FIG. 1 is for illustration purposes only, and the implementation of the present invention is not limited to FIG. 1. The segmentation model training device 100 may further include other components required for operation and application. For example, the segmentation model training device 100 may further include an output interface (e.g., a display panel for displaying information), an input interface (e.g., a touch panel, a keyboard, a microphone, a scanner or a flash reader), and a communication circuit (e.g., a WiFi communication model, a Bluetooth communication model, a wireless telecommunication network communication model, etc.). In some embodiments, the segmentation model training device 100 may be established by a computer, a server, or a processing center.

[0068] In some embodiments, the memory 120 may be a flash memory, a HDD, an SSD (solid state drive), a DRAM (dynamic random access memory) or an SRAM (static random access memory). In some embodiments, the memory 120 may be a non-volatile computer-readable recording medium storing at least one instruction associated with the segmentation model training method. The processor 110 may access and execute the at least one instruction.

[0069] In some embodiments, the processor 110 may be, but is not limited to, a single processor or a collection of multiple microprocessors, such as a CPU or a GPU. The microprocessor is electrically coupled to the memory 120 to access and execute the segmentation model training method according to at least one instruction. For ease of understanding and description, the details of the segmentation model training method will be described in the following paragraphs.

[0070] In some embodiments, the memory 120 stores a segmentation model SM and a data augmentation model AM. The segmentation model SM and the data augmentation model AM can be read and executed by the processor 110.

[0071] Details of the implementation of the present application are disclosed below with reference to the segmentation model training method 200 in FIG. 2 , which is a flow chart of the segmentation model training method 200 applicable to the segmentation model training device 100 in FIG. 1 . However, the implementation of the present application is not limited thereto.

[0072] Please refer to Figure 2. Figure 2 The diagram is a schematic diagram of a segmentation model training method 200 illustrated in some embodiments of the present application. However, the implementation of the present application is not limited thereto.

[0073] It should be noted that the segmentation model training method can be applied to a system having the same or similar structure as the segmentation model training device 100 in FIG. 1. To simplify the description, the segmentation model training method will be described below using FIG. 1 as an example, but the present application is not limited to the application of FIG. 1.

[0074] It should be noted that in some embodiments, the segmentation model training method can also be implemented as a computer program and stored in a non-transitory computer-readable recording medium, so that a computer, an electronic device, or the processor 110 as shown in FIG. 1 reads the recording medium and executes the operation method. The non-transitory computer-readable recording medium can be a read-only memory, a flash memory, a floppy disk, a hard disk, an optical disk, a flash drive, a magnetic tape, a database accessible by a network, or a non-transitory computer-readable recording medium having the same function that can be easily conceived by a person skilled in the art.

[0075] In addition, it should be understood that the operations of the operation method mentioned in this embodiment, except for those whose order is specifically described, can be adjusted in order according to actual needs, and can even be executed simultaneously or partially simultaneously.

[0076] Furthermore, in different embodiments, these operations may be adaptively increased, replaced, and / or omitted.

[0077] Please refer to FIG. 2. The segmentation model training method 200 includes the following steps. For the sake of convenience and clarity of description, the following refers to FIG. 1 and FIG. 2 at the same time, and the detailed steps of the segmentation model training method 200 shown in FIG. 2 are explained by the operation relationship between the components in the segmentation model training device 100.

[0078] In step S210 , a plurality of sample groups in the large sample set are input into a data augmentation model to generate a plurality of augmented sample groups. In some embodiments, step S210 is executed by the processor 110 as shown in FIG. 1 .

[0079] Please also refer to Figure 3. Figure 3 Shown is a schematic diagram of step S210 illustrated in some embodiments of the present application.

[0080] As shown in FIG. 3 , in step S210 , the large sample set LS includes a plurality of sample groups LD1 to LD4 . The sample groups LD1 to LD4 include images and masks corresponding to the images, respectively. For example, the sample group LD1 includes the image LI1 and the mask LM1 corresponding to the image LI1 , the sample group LD2 includes the image LI2 and the mask LM2 corresponding to the image LI2 , and so on. For the portion of the image LI2 containing the wound, the mask LM2 includes the white portion with the pixel being 1, and for the other portions of the image LI2 (i.e., the portion not containing the wound), the mask LM2 includes the black portion with the pixel being 0.

[0081] In some embodiments, the processor 110 generates an amplified sample set ALS after inputting the large sample set LS into the data augmentation model AM as shown in FIG. 1. The amplified sample set ALS includes a plurality of amplified sample groups ALD1 to ALD4. The amplified sample groups ALD1 to ALD4 include amplified images and amplified masks corresponding to the amplified images, respectively. For example, the amplified sample group ALD1 includes an amplified image ALI1 and an amplified mask ALM1 corresponding to the amplified image ALI1, the amplified sample group ALD2 includes an amplified image ALI2 and an amplified mask ALM2 corresponding to the amplified image ALI2, and the rest are analogous.

[0082] In some embodiments, the data augmentation model AM adjusts at least one of the size, angle (including rotation angle), color, and position of the images LI1 to LI4 in the sample groups LD1 to LD4 to generate the amplified images ALI1 to ALI4 in the amplified sample groups ALD1 to ALD4, and correspondingly adjusts the masks LM1 to LM4 in the sample groups LD1 to LD4 to generate the amplified masks ALM1 to ALM4 in the amplified sample groups ALD1 to ALD4.

[0083] For example, in one embodiment, the processor 110 rotates the image LI1 of the sample group LD1 by 45 degrees clockwise to generate the amplified image ALI1 in the amplified sample group ALD1, and accordingly, the processor 110 correspondingly rotates the mask LM1 of the sample group LD1 by 45 degrees clockwise to generate the amplified mask ALM1 in the amplified sample group ALD1. That is, when the image LI1 of the sample group LD1 is adjusted according to the adjustment parameter P1 (not shown) to generate the amplified image ALI1 of the amplified sample group ALD1, the mask LM1 of the sample group LD1 is adjusted according to the same adjustment parameter P1 to generate the amplified mask ALM1 of the amplified sample group ALD1.

[0084] Please also refer to Figure 4. Figure 4 Schematic diagram showing an image LI2 of a sample set LD2 and an amplified image ALI2 of an amplified sample set ALD2 according to some embodiments of the present application.

[0085] As shown in FIG. 4 , in one embodiment, after the data augmentation model AM rotates and displaces the image LI2 of the sample group LD2, the image LI2 is moved from the original area A0 defined by the positioning points pa1, pa2, pa3, and pa4 to the area A1 defined by the positioning points pb1, pb2, pb3, and pb4. In the implementation of the present case, the processor 110 deletes the portion of the image that exceeds the area A0 (i.e., the portion of the area A11), and only retains the portion of the image that falls in the area A12. In addition, the processor 110 fills the area A2 (i.e., the portion of the area A11 minus the area A12) with black, for example, filling the area A2 with a pixel value of 0. Finally, the processor 110 generates an augmented image ALI2 of the augmented sample group ALD2 based on the areas A12 and A2.

[0086] The generation method of the amplification mask ALM2 of the amplification sample set ALD2 corresponds to the generation method of the amplification image ALI2, which will not be described in detail here.

[0087] It should be noted that the sample groups LD1 to LD4 and the amplified sample groups ALD1 to ALD4 shown in FIG. 3 are only for illustration purposes, and more sample groups and amplified sample groups are within the implementation of the present invention. In addition, one sample group can be adjusted in different ways to generate multiple amplified sample groups.

[0088] In this way, the diversity of the large sample set LS can be increased through the data augmentation model AM.

[0089] Please refer back to FIG. 2. In step S230, multiple mixed sample groups are generated based on multiple sample groups in the small sample set. In some embodiments, step S230 is performed by the processor 110 in FIG. 1. The detailed implementation of step S230 will be described in detail in FIG. 5. Figure 1 And explain.

[0090] Please refer to Figure 5. Figure 5 The flowchart of step S230 in FIG. 2 is shown in some embodiments of the present application. Step S230 includes steps S232 to S236.

[0091] In step S232, a target image in the image is captured based on images of a plurality of sample groups of small samples and masked images.

[0092] Please also refer to Figure 6. Figure 6It is a schematic diagram of a small sample set SS illustrated in some embodiments of the present application. As illustrated in FIG. 6 , the small sample set SS includes a plurality of sample groups SD1 to SD4. The sample groups SD1 to SD4 include images and masks corresponding to the images, respectively. For example, the sample group SD1 includes an image SI1 and a mask SM1 corresponding to the image SI1, the sample group SD2 includes an image SI2 and a mask SM2 corresponding to the image SI2, and the rest are analogous.

[0093] In some embodiments, the processor 110 in FIG. 1 captures a target image in the image based on the images in the sample set and the mask. For example, in one embodiment, the processor 110 captures a target image in the image SI1 based on the image SI1 in the sample set SD1 and the mask SM1 corresponding to the image SI1.

[0094] Please also refer to Figure 7. Figure 7 The diagram shows a schematic diagram of a captured target image according to some embodiments of the present application. In one embodiment, the processor 110 in FIG. 1 overlaps the image SI1 with the mask SM1 to obtain the target image TI1. Specifically, the mask SM1 includes a white portion with a pixel of 1 and a black portion with a pixel of 0. The portion of the image SI1 overlapping the white portion with a pixel of 1 in the mask SM1 is the captured target image TI1. In some embodiments, the processor 110 also performs background removal on the portion of the image SI1 that is not captured to remove the background. That is, the portion of the image SI1 that is not captured is filled with a pixel of 0.

[0095] In one embodiment, the processor 110 extracts a portion of the area TP from the mask SM1 corresponding to the target image TI1 to obtain a mask TM1 corresponding to the target image TI1 , wherein the mask TM1 extracted by the processor 110 includes a white portion of the mask SM1 where the original pixel is 1.

[0096] Please refer to Figure 8. Figure 8 Schematic diagram of a target sample set TS illustrated in some embodiments of the present application. In the target sample set TS of FIG. 8 , a plurality of target sample groups TD1 to TD4 are included. The target sample groups TD1 to TD4 respectively include a target image and a mask corresponding to the target image. For example, the target sample group TD1 includes a target image TI1 and a mask TM1 corresponding to the target image TI1, the target sample group TD2 includes a target image TI2 and a mask TM2 corresponding to the target image TI1, and the rest are analogous.

[0097] Please refer back to FIG. 5. In step S234, a target region in the image is selected based on an image of one of the plurality of sample groups in the small sample set and a mask.

[0098] Please refer to Figure 6 and Figure 8. In one embodiment, the processor 110 in Figure 1 selects one of the plurality of sample groups SD1 to SD4 in Figure 6, and obtains the target area according to the image of the selected one of the plurality of sample groups SD1 to SD4.

[0099] Please also refer to Figure 9. Fig. 9 Schematic diagrams of methods for obtaining a target area according to some embodiments of the present application are shown. As shown in FIG. 9, in one embodiment, the processor 110 in FIG. 1 selects a sample group SD4 from a plurality of sample groups SD1 to SD4, and obtains a target area TA according to an image SI4 in the sample group SD4 and a mask SM4 corresponding to the image SI4. In some embodiments, the target area TA is an area located in the white portion of the mask SM1 where the pixel is 1 when the image SI4 and the mask SM4 are overlapped.

[0100] Please refer back to FIG. 5. In step S236, the target image is pasted to the target area to generate an image in the mixed sample set.

[0101] Please also refer to Figure 10. Fig.10 Shown is a schematic diagram of a mixed sample group illustrated in some embodiments of the present application.

[0102] In one embodiment, the processor 110 in FIG. 1 selects the target sample set TD2 in FIG. 8 in step S232, and selects the sample set SD4 in FIG. 6 in step S234. Then, in step S236, the processor 110 pastes the target image TI2 in the target sample set TD2 in FIG. 8 to the target area (such as the target area TA shown in FIG. 9) in the image SI4 in the sample set SD4 in FIG. 6 to generate the image MI in the mixed sample set MD in FIG. 10.

[0103] It should be noted that the sample group corresponding to the target sample group selected by the processor 110 in step S232 and the sample group selected by the processor 110 in step S234 must be different sample groups. For example, if the processor 110 selects the target sample group TD2 in FIG. 8 in step S232, since the sample group TD2 corresponds to the sample group SD2 in FIG. 6, the processor needs to select a sample group other than the sample group SD1 in step S234.

[0104] In some embodiments, after the processor 110 selects the target sample group TD1 in FIG. 8 and the sample group SD4 in FIG. 6, the processor 110 calculates the area of ​​the target image TI2 in the target sample group TD2 and the area of ​​the image SI4 in the sample group SD4. Then, the processor 110 calculates the ratio between the area of ​​the target image TI1 and the area of ​​the image SI4. If the ratio between the area of ​​the target image TI1 and the area of ​​the image SI4 is outside the ratio range, the processor 110 re-executes step S234 to select another sample group.

[0105] In some embodiments, before the processor 110 pastes the target image TI2 in the target sample group TD2 in FIG. 8 to the target area (such as the target area TA shown in FIG. 9 ) in the image SI4 in the sample group SD4 in FIG. 6 , the processor 110 first adjusts at least one of the size and angle (including the rotation angle) of the target image TI2 according to the adjustment parameter (such as the adjustment parameter P2) so that the area of ​​the adjusted target image (not shown) is the same as the area of ​​the target area TA, or the area of ​​the adjusted target image (not shown) is different from the area of ​​the target area TA within a preset difference range. Then, the processor 110 pastes the adjusted target image to the target area (such as the target area TA shown in FIG. 9 ) in the image SI4 in the sample group SD4 in FIG. 6 .

[0106] In some embodiments, the processor 110 also superimposes the mask TM2 in the target sample set TD2 in FIG. 8 and the mask SM4 in the sample set SD4 in FIG. 6 to generate the mask MM in the mixed sample set MD in FIG. 10 .

[0107] In one embodiment, the processor 110 is further configured to adjust at least one of the size and angle (including the rotation angle) of the mask TM1 according to the adjustment parameter P2 in response to the image MI to generate an adjusted mask (not shown). Then, the processor 110 superimposes the adjusted mask (not shown) with the mask SM4 to generate the mask MM.

[0108] Please refer back to Figure 2. In step S250, a plurality of mixed sample sets are input into the data augmentation model AM to generate a plurality of augmented mixed sample sets.

[0109] Please also refer to Figure 11. Fig.11 The schematic diagram of step S250 in FIG. 2 illustrated in some embodiments of the present application is shown. After the processor 110 inputs the mixed sample set MD into the data augmentation model AM illustrated in FIG. 1 , an augmented mixed sample set AMD is generated. The mixed sample set MD includes an image MI and a mask MM corresponding to the image MI. The augmented mixed sample set AMD includes an augmented image AMI and an augmented mask AMM corresponding to the augmented image AMI.

[0110] In some embodiments, the data augmentation model AM adjusts at least one of the size, angle (including rotation angle), color and position of the image MI in the mixed sample group MD to generate the augmented image AMI in the amplified mixed sample group AMD, and correspondingly adjusts the mask MM in the mixed sample group MD to generate the augmented mask AMM in the mixed sample group MD.

[0111] That is, when the image MI in the mixed sample group MD is adjusted according to the adjustment parameter P3 (not shown) to generate the amplified image AMI in the amplified mixed sample group AMD, the data augmentation model AM adjusts the mask MM in the mixed sample group MD according to the same adjustment parameter P3 to generate the amplified mask AMM in the mixed sample group MD.

[0112] It should be noted that FIG. 10 and FIG. 11 only use the mixed sample group MD and the amplified mixed sample group AMD as examples for description, and more different mixed sample groups and amplified mixed sample groups can be generated, which will not be described in detail here.

[0113] Please refer back to FIG. 2 . In step S270 , the segmentation model is trained based on the amplified sample set and the amplified mixed sample set. In some embodiments, step S270 is executed by the processor 110 in FIG. 1 .

[0114] Please also refer to Figure 12. Fig.12 It is a schematic diagram of step S270 in FIG. 2 illustrated in some embodiments of the present application. As illustrated in FIG. 12, in an embodiment of the present application, a large sample set LS and a small sample set SS are included. After the large sample set LS passes through step S210 in FIG. 2, an augmented sample set ALS is generated. As shown in FIG. 1, the processor 110 pre-trains the segmentation model SM based on the augmented sample set ALS. After the segmentation model SM is pre-trained based on the augmented sample set ALS, the pre-training parameter PW is generated.

[0115] In some embodiments, the large sample set LS is a sample set with a larger number of samples, and the small sample set SS is a sample set with a smaller number of samples. In one embodiment, the large sample set LS is a public sample set with a wider domain, and the small sample set SS is a private sample set with a narrower domain.

[0116] On the other hand, the small sample set SS generates an augmented mixed sample set AMS after the steps S230 and S250 in FIG. 2, wherein the augmented mixed sample set AMS includes a plurality of augmented mixed sample groups (including, for example, the augmented mixed sample group AMD in FIG. 11). As shown in FIG. 1, the processor 110 performs fine-tune training on the segmentation model SM including the pre-training parameters PW based on the augmented mixed sample set AMS.

[0117] In summary, the embodiments of the present application provide a segmentation model training method, a segmentation model training device, and a non-volatile computer-readable recording medium. By clipping, a target image in one of the images in the small sample set is captured and pasted to a target area in another image in the small sample set to generate a mixed sample, which can increase the number of samples with high reliability and increase the diversity of the small sample set when the number of samples is small. In addition, the mixed sample and the large sample set are amplified to further increase the diversity of the samples. Furthermore, when training the segmentation model, the amplified sample group generated according to the large sample set is first pre-trained, and then the mixed sample group generated according to the mixed sample set is fine-tuned. When there is a difference in the domain between the large sample set and the small sample set, the segmentation model can be effectively trained to solve the problem of a small number of samples in the same domain as the small sample set. Accordingly, in the embodiments of the present application, when the number of samples in the small sample set is small, the performance of the segmentation model can be improved by pre-training the large sample set and then fine-tuning the mixed sample generated by the small sample set.

[0118] It should be noted that although the above-mentioned implementation is described by taking skin wound as an example, the implementation of this case is not limited to skin wounds, and various image segmentation training (such as image segmentation of damaged electronic devices, etc.) are all within the implementation of this case.

[0119] The above examples include sequential exemplary steps, but the steps do not have to be performed in the order shown. Performing the steps in different orders is within the scope of the present application. Within the spirit and scope of the embodiments of the present application, the steps may be added, replaced, changed in order, and / or omitted as appropriate.

[0120] Although specific embodiments of the present application have been disclosed in relation to the above embodiments, these embodiments are not intended to limit the present application. Various substitutions and improvements may be performed in the present application by a person of ordinary skill in the relevant art without departing from the principles and spirit of the present application. Therefore, the scope of protection of the present application is determined by the scope of the attached patent application.

Claims

1. A segmentation model training method, characterized in that: Include: Inputting a plurality of first sample groups in a large sample set into a data augmentation model to generate a plurality of augmented sample groups; generating a plurality of mixed sample groups based on a plurality of second sample groups in a small sample set; Inputting a plurality of the mixed sample groups into the data augmentation model to generate a plurality of augmented mixed sample groups; and Training a segmentation model based on the plurality of amplified sample groups and the plurality of amplified mixed sample groups comprises: Pre-training the segmentation model based on the plurality of amplified sample groups; and The segmentation model is fine-tuned and trained corresponding to the plurality of amplified mixed sample groups.

2. The segmentation model training method according to claim 1, characterized in that: Generating a plurality of the mixed sample groups based on a plurality of the second sample groups in the small sample set comprises: capturing a target image in the first image based on a first image of one of the second sample groups and a first mask; Acquire a target area in the second image based on a second image of another one of the plurality of second sample groups and a second mask; as well as Pasting the target image onto the target area in the second image to generate a third image of one of the plurality of mixed sample sets; A ratio value between an area of ​​the target image and an area of ​​the target region is within a ratio value range.

3. The segmentation model training method according to claim 2, characterized in that: Also includes: adjusting at least one of a size and an angle of the target image according to an adjustment parameter to generate an adjusted target image, wherein an area of ​​the adjusted target image is the same as an area of ​​the target region; Pasting the adjusted target image to the target area in the second image to generate the third image of the one of the plurality of mixed sample groups; as well as Superimposing the first mask and the second mask to generate a third mask of the plurality of the mixed sample sets comprises: adjusting the first mask according to the adjustment parameter to generate an adjusted mask; as well as The adjusted mask is superimposed with the second mask to generate the third mask in a plurality of the mixed sample groups.

4. The segmentation model training method according to claim 1, characterized in that: Also includes: Adjusting at least one of the size, angle, color and position of the plurality of first images of the plurality of mixed sample groups according to the data augmentation model to generate a plurality of first augmented images of the plurality of augmented mixed sample groups; Adjusting at least one of the size, angle, color, and position of the plurality of second images of the plurality of first sample groups according to the data augmentation model to generate a plurality of second augmented images of the plurality of augmented sample groups; Adjusting at least one of the size, angle, color and position of a plurality of first masks of a plurality of the mixed sample groups corresponding to a plurality of the first images of a plurality of the mixed sample groups to generate a plurality of first augmented masks of a plurality of the augmented mixed sample groups; as well as At least one of the size, angle, color, and position of a plurality of second masks of a plurality of the first sample groups is adjusted corresponding to a plurality of the second images of a plurality of the first sample groups to generate a plurality of second amplified masks of a plurality of the two amplified sample groups.

5. A segmentation model training device, comprising: a memory for storing a segmentation model and a data augmentation model; and A processor, coupled to the memory, configured to execute: Inputting a plurality of first sample groups in a large sample set into the data augmentation model to generate a plurality of augmented sample groups; generating a plurality of mixed sample groups based on a plurality of second sample groups in a small sample set; Inputting a plurality of the mixed sample groups into the data augmentation model to generate a plurality of augmented mixed sample groups; and Training the segmentation model based on the plurality of amplified sample groups and the plurality of amplified mixed sample groups comprises: Pre-training the segmentation model based on the plurality of amplified sample groups; and The segmentation model is fine-tuned and trained corresponding to the plurality of amplified mixed sample groups.

6. The segmentation model training device according to claim 5, characterized in that: The processor is further configured to: capturing a target image in the first image based on a first image of one of the second sample groups and a first mask; Acquire a target area in the second image based on a second image of another one of the plurality of second sample groups and a second mask; as well as pasting the target image to the target area in the second image to generate a third image of one of the plurality of mixed sample sets; A ratio value between an area of ​​the target image and an area of ​​the target region is within a ratio value range.

7. The segmentation model training device according to claim 6, characterized in that: The processor is further configured to: adjusting at least one of the size and the angle of the target image according to an adjustment parameter to generate an adjusted target image, wherein an area of ​​the adjusted target image is the same as an area of ​​the target region; Pasting the adjusted target image to the target area in the second image to generate the third image of the one of the plurality of mixed sample groups; as well as Superimposing the first mask and the second mask to generate a third mask of the plurality of the mixed sample sets comprises: adjusting the first mask according to the adjustment parameter to generate an adjusted mask; as well as The adjusted mask is superimposed with the second mask to generate the third mask in a plurality of the mixed sample groups.

8. The segmentation model training device according to claim 5, characterized in that: The processor is further configured to: Adjusting at least one of the size, angle, color and position of the plurality of first images of the plurality of mixed sample groups according to the data augmentation model to generate a plurality of first augmented images of the plurality of augmented mixed sample groups; Adjusting at least one of the size, angle, color, and position of the plurality of second images of the plurality of first sample groups according to the data augmentation model to generate a plurality of second augmented images of the plurality of augmented sample groups; Adjusting at least one of the size, angle, color and position of the plurality of first masks of the plurality of mixed sample groups corresponding to the plurality of first images of the plurality of mixed sample groups to generate a plurality of first augmented masks of the plurality of augmented mixed sample groups; as well as At least one of the size, angle, color, and position of the second masks of the first sample groups is adjusted corresponding to the second images of the first sample groups to generate the second amplified masks of the two amplified sample groups.

9. A non-volatile computer-readable recording medium for storing a computer program which, when executed, causes one or more processing components to perform a plurality of operations comprising: Inputting a plurality of first sample groups in a large sample set into a data augmentation model to generate a plurality of augmented sample groups; generating a plurality of mixed sample groups based on a plurality of second sample groups in a small sample set; Inputting the plurality of mixed sample groups into the data augmentation model to generate a plurality of augmented mixed sample groups; as well as Training a segmentation model based on the plurality of amplified sample groups and the plurality of amplified mixed sample groups comprises: Pre-training the segmentation model based on the multiple amplified sample groups; and The segmentation model is fine-tuned and trained corresponding to the multiple amplified mixed sample groups.

10. The non-volatile computer readable recording medium according to claim 9, wherein: The multiple operations also include: capturing a target image in the first image based on a first image of one of the plurality of second sample groups and a first mask; Acquire a target area in a second image of another one of the plurality of second sample groups based on a second image and a second mask; as well as The target image is pasted onto the target area in the second image to generate a third image of one of the plurality of mixed sample groups.