Model training method, image recognition method, terminal device and computer medium

By constructing a segmented pseudo-label and feature similarity loss function based on the image training set, the segmented model is trained, and the model accuracy is solved due to the complexity of the image data, and the model classification accuracy and local information positioning ability are improved.

CN114548213BActive Publication Date: 2025-07-22ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111636815.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-22
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

When applying natural language processing models to industrial practice, the morphological distribution of image data is complex, resulting in blurred edges between target prospects and backgrounds, affecting the accuracy of semantic segmentation of the model.

Method used

By obtaining the segmentation pseudo-label of a single category training image in the image training set, the first loss function is constructed, and the similarity of the pairs of image feature of the same and different categories is extracted, the second loss function is constructed, and the segmentation model is trained to constrain the similarity and dissimilarity of image features.

Benefits of technology

The model's processing ability of edge blur images is improved, the model's classification accuracy and local image information are enhanced, and the model's accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548213B_ABST
    Figure CN114548213B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, an image recognition method, a terminal device, and a computer medium. The method includes: obtaining an image training set including a plurality of first training images of a single category; obtaining first segmentation pseudo-labels of the first training images of each category; inputting the plurality of first training images into a segmentation model to be trained to obtain first prediction labels; constructing a first loss function based on the first prediction labels and the first segmentation pseudo-labels; extracting first image feature pairs of the first training images of the same category and second image feature pairs of the first training images of different categories; obtaining a first similarity of the first image feature pairs and a second similarity of the second image feature pairs, constructing a second loss function, and training the segmentation model using the first loss function and the second loss function. The image recognition method of the present application constrains the dissimilarity of the image feature pairs of different categories and the similarity of the image feature pairs of the same category, thereby improving the model accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and particularly to a model training method, an image recognition method, a terminal device, and a computer medium. Background Art

[0002] In recent years, natural language processing technologies represented by pre-training in the field of artificial intelligence have achieved explosive development, with new technologies and new models emerging in an endless stream. Under the background of the new era, how to efficiently apply diverse and advanced scientific research achievements in the field of natural language processing to industrial practice and solve practical problems is the core issue in the field of natural language processing.

[0003] However, in the process of applying various models to industrial practice, the complex application scenarios make the morphological distribution of image data complex. When processing images, the edges between the target foreground and the image background are blurred, and over-segmentation is likely to occur when performing semantic segmentation on the target object, affecting the accuracy of the model. Summary of the Invention

[0004] This application provides a model training method, an image recognition method, a terminal device, and a computer medium to solve the technical problem of low model accuracy in the prior art.

[0005] To solve the above problems, the first technical solution provided by this application is: to provide a model training method, which includes:

[0006] Obtain an image training set, where the image training set includes a number of first training images of a single category;

[0007] Obtain the first segmentation pseudo-labels of the first training images of each category;

[0008] Input a number of the first training images into the segmentation model to be trained, and obtain the first prediction labels of the first training images;

[0009] Construct a first loss function based on the first prediction labels and the first segmentation pseudo-labels;

[0010] Extract the first image feature pairs of the first training images of the same category, and the second image feature pairs of the first training images of different categories;

[0011] Obtain the first similarity of the first image feature pairs and the second similarity of the second image feature pairs;

[0012] Construct a second loss function based on the first similarity and the second similarity, and use the first loss function and the second loss function to train the segmentation model.

[0013] To solve the above technical problems, the second technical solution provided by this application is: to provide an image recognition method, and the image recognition method includes:

[0014] Input the image to be recognized into the segmentation model to obtain the image recognition category of the image to be recognized, where

[0015] The segmentation model is obtained by using the model training method described above.

[0016] To solve the above technical problems, the third technical solution provided by this application is: to provide a terminal device, and the terminal device includes a processor and a memory connected to the processor, where

[0017] The memory stores program instructions;

[0018] The processor is used to execute the program instructions stored in the memory to implement the model training method described above.

[0019] To solve the above technical problems, the fourth technical solution provided by this application is: to provide a computer-readable storage medium, and the computer-readable storage medium stores program instructions, and when the program instructions are executed, the model training method described above is implemented.

[0020] In the model training method provided by this application, the terminal device obtains an image training set, and the image training set includes a number of first training images of a single category; obtains the first segmentation pseudo-labels of the first training images of each category; inputs a number of first training images into the segmentation model to be trained to obtain the first prediction labels of the first training images; constructs a first loss function based on the first prediction labels and the first segmentation pseudo-labels; extracts the first image feature pairs of the first training images of the same category and the second image feature pairs of the first training images of different categories; obtains the first similarity of the first image feature pairs and the second similarity of the second image feature pairs; constructs a second loss function based on the first similarity and the second similarity, and uses the first loss function and the second loss function to train the segmentation model. In the image recognition method of this application, by using the similarity of the first image feature pairs and the second image feature pairs to construct a second loss function to train the segmentation model, the similarity of the image feature pairs of the same category and the dissimilarity of the image feature pairs of different categories are constrained, and the accuracy of the segmentation model is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0022] Figure 1 It is a schematic flowchart of the first embodiment of the model training method provided by this application;

[0023] Figure 2 It is a schematic structural diagram of an embodiment of the model provided by this application;

[0024] Figure 3 It is a schematic flowchart of the second embodiment of the model training method provided by this application;

[0025] Figure 4 is Figure 3 a schematic flowchart of obtaining the first segmentation pseudo-label in;

[0026] Figure 5 It is a schematic flowchart of the third embodiment of the model training method provided by this application;

[0027] Figure 6 It is a schematic flowchart of the fourth embodiment of the model training method provided by this application;

[0028] Figure 7 It is a schematic flowchart of the fifth embodiment of the model training method provided by this application;

[0029] Figure 8 is Figure 7 a schematic structural diagram of an embodiment of the second training image in;

[0030] Figure 9 It is a schematic flowchart of the sixth embodiment of the model training method provided by this application;

[0031] Figure 10 It is a schematic flowchart of the seventh embodiment of the model training method provided by this application;

[0032] Figure 11 It is a schematic flowchart of the eighth embodiment of the model training method provided by this application;

[0033] Figure 12 It is a schematic flowchart of the ninth embodiment of the model training method provided by this application;

[0034] Figure 13 is Figure 12 a schematic flowchart of obtaining the class response in Figure 1 the embodiment;

[0035] Figure 14 It is a schematic structural diagram of an embodiment of the terminal device provided by this application;

[0036] Figure 15 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. DETAILED DESCRIPTION

[0037] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. It is to be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some structures related to the present application are shown in the accompanying drawings, rather than all structures. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0038] The terms "first", "second", etc. in this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "provided with" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0039] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0040] See also Figure 1-2 , Figure 1 is a flowchart of the first embodiment of the model training method provided by this application, Figure 2 It is a structural schematic diagram of a model 1 embodiment provided in this application.

[0041] like Figure 1 As shown, the specific steps of the model training method are as follows:

[0042] S11: Acquire an image training set, wherein the image training set includes a plurality of first training images of a single category.

[0043] In the embodiment of the present application, the image training set can be obtained by using an image acquisition device to collect data from relevant areas, or it can be obtained from a variety of standard test databases. In a specific implementation, the image training set can be pathological image slices, medical imaging images, etc., or other image data that requires semantic segmentation. Here, the model training method of the present application is described using pathological image slices as a representative.

[0044] As an example, a user can customize pathological images of any part of the body, that is, the user can, but is not limited to, determine a certain organ or body part as the diseased part and collect data of pathological image slices based on the diseased part as an image training set.

[0045] The model used in the embodiments of the present application can be, but is not limited to, a weakly supervised network model. Specifically, as Figure 2 shown, the weakly supervised network model includes n convolutional layers for feature extraction. After the nth convolutional layer outputs a feature map of n channels, the upper branch passes through a Global Average Pooling Layer (GAP) for regularization processing, averages the feature maps of each layer, and uses the average value to represent the parameters of each layer, so that the feature map of n channels becomes n features, reducing the number of parameters and preventing the model from overfitting during training. After the GAP layer outputs features of nx1, classification is performed through a Fully Connected layer (FC). The FC layer maps the features of nx1 to C categories to obtain features of nxC. After the features pass through the fully connected layer, they pass through the class activation map (CAM) classification network of the model to output the region in each target category that is most discriminative for that target category, and the image training set to be segmented is segmented to obtain a number of first training images of a single category.

[0046] When the terminal device trains the segmentation model to be trained, its training process is divided into a pre-training stage and a fine-tuning stage. When entering the pre-training stage, the terminal device obtains an image training set, where the image training set includes a number of first training images of a single category.

[0047] S12: Obtain the first segmentation pseudo-label of the first training image of each category.

[0048] Specifically, the terminal device obtains the first training image of each category and uses a threshold to segment the first training image to obtain the first segmentation pseudo-label of the first training image of each category.

[0049] S13: Input a number of first training images into the segmentation model to be trained and obtain the first prediction label of the first training image.

[0050] After the terminal device obtains the first training image of each category and its first segmentation pseudo-label, it inputs the first training image and the first segmentation pseudo-label into the segmentation model to be trained to obtain the first prediction label of the first training image.

[0051] S14: Construct a first loss function based on the first prediction label and the first segmentation pseudo-label.

[0052] After the terminal device obtains the first prediction label and the first segmentation pseudo-label for each category, it can calculate the first loss function, as shown in the following formula:

[0053]

[0054] Among them, Loss seg is the first loss function; C is the number of categories of the first training images; Z C is the first prediction label; l c is the first segmentation pseudo-label.

[0055] S15: Extract the first image feature pairs of the first training images of the same category, and the second image feature pairs of the first training images of different categories.

[0056] To improve the classification accuracy of the model and make the model applicable to processing images with blurred edges, the terminal device extracts the first image feature pairs of the first training images of the same category, and the second image feature pairs of the first training images of different categories. Among them, the image feature pair is a pair of image features extracted according to the image pixel points predicted as the corresponding category in the first prediction label.

[0057] S16: Obtain the first similarity of the first image feature pairs and the second similarity of the second image feature pairs.

[0058] Since there is similarity between different first training images of the same category, the terminal device obtains the first similarity of the first image feature pairs and the second similarity of the second image feature pairs. Specifically, the method for measuring similarity includes but is not limited to using cosine similarity to measure, and it can also be other similarity measurement methods.

[0059] S17: Construct a second loss function based on the first similarity and the second similarity, and use the first loss function and the second loss function to train the segmentation model.

[0060] After the terminal device obtains the first similarity and the second similarity, it constructs a second loss function based on the first similarity and the second similarity. After the terminal device obtains the first loss function and the second loss function, it trains the segmentation model based on the first loss function and the second loss function, as shown in the following formula:

[0061] Loss total1 = Loss seg + Loss conw ;

[0062] Among them, Loss total1 is the loss function for training the model with the first training images, Loss seg is the first loss function, Loss conwis the second loss function.

[0063] In the embodiment of the present application, the terminal device obtains an image training set, where the image training set includes a number of first training images of a single category; obtains the first segmentation pseudo-labels of the first training images of each category; inputs the number of first training images into a segmentation model to obtain the first prediction labels of the first training images; constructs a first loss function based on the first prediction labels and the first segmentation pseudo-labels; extracts the first image feature pairs of the first training images of the same category and the second image feature pairs of the first training images of different categories; obtains the first similarity of the first image feature pairs and the second similarity of the second image feature pairs; constructs a second loss function based on the first similarity and the second similarity, and uses the first loss function and the second loss function to train the segmentation model. In the model training method of the present application, by using the first segmentation pseudo-labels to train the model, the processing ability of the model for images with blurred edges is improved; the second loss function is constructed using the similarities of the first image feature pairs and the second image feature pairs to constrain the similarity of the image feature pairs of the same category and the dissimilarity of the features between different first training images of different categories, thereby improving the accuracy of the model.

[0064] Please refer to Figure 3-4 , Figure 3 which is a schematic flowchart of the second embodiment of the model training method provided by the present application. Figure 4 is Figure 3 a schematic flowchart of obtaining the first segmentation pseudo-labels in Figure 3 As shown in

[0065] S21: Normalize the first training images.

[0066] After the terminal device obtains a number of first training images of a single category, due to factors such as image acquisition and imaging, the gray-scale information of the same acquisition part in the images will be inconsistent. Therefore, in this embodiment, the first training images are normalized. Optionally, the normalization method can be the maximum-minimum normalization method, or other data normalization methods; in a specific implementation manner, the first training images can also be grayscale processed, for example, using the mean-variance normalization or gray-scale transformation normalization method for grayscale processing. The normalization method is not specifically limited herein.

[0067] S22: Based on a preset segmentation threshold, distinguish the foreground region and the background region in the normalized first training images, so as to obtain the first segmentation pseudo-labels of the first training images.

[0068] As shown in Figure 4As shown in the figure, after the terminal device normalizes the first training image, it inputs the normalized first training image into the segmentation model, and uses a preset threshold to segment the first training image to distinguish the foreground region and the background region. The foreground region of each category is the first segmentation pseudo-label of the first training image. In a specific implementation, the first training image is a sample of three single categories. After the terminal device inputs the three first training images into the segmentation model, the segmentation model segments the images based on the preset threshold to obtain the foreground regions and background regions of the three single categories.

[0069] In the embodiment of the present application, the terminal device normalizes the first training image; based on a preset segmentation threshold, it distinguishes the foreground region and the background region in the normalized first training image, so as to obtain the first segmentation pseudo-label of the first training image. Through the method of this embodiment, the segmentation pseudo-label of the image can be obtained from the training image, and the segmentation model can be trained using the segmentation pseudo-label, improving the efficiency of model training.

[0070] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of the third embodiment of the model training method provided by the present application. As Figure 5 shown, step S16 further includes the following steps:

[0071] S31: respectively obtain the first image feature and the second image feature of two first training images of the target category.

[0072] In this embodiment, since there is similarity between different first training images of the same category, the terminal device respectively obtains the first image feature and the second image feature of two first training images of the target category.

[0073] S32: obtain the first similarity between the first image feature and the second image feature.

[0074] After the terminal device obtains the first image feature and the second image feature of two first training images of the target category, it calculates the first similarity between the first image feature and the second image feature, as shown in the following formula:

[0075]

[0076] where x c , are the first image feature and the second image feature of two first training images of the target category; is the first similarity between the first image feature and the second image feature.

[0077] S33: respectively obtain the third image feature of one first training image of the target category and the fourth image feature of one first training image of other categories.

[0078] Due to the dissimilarity between different first training images of different categories, the terminal device respectively obtains the third image feature of a first training image of the target category and the fourth image feature of a first training image of other categories.

[0079] S34: Obtain the second similarity between the third image feature and the fourth image feature.

[0080] After the terminal device obtains the third image feature and the fourth image feature of different first training images of different categories, it calculates the second similarity between the third image feature and the fourth image feature. The calculation formula is similar to the formula for calculating the first similarity between the first image feature and the second image feature, and will not be elaborated here.

[0081] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of the fourth embodiment of the model training method provided by this application. As Figure 6 shown, step S17 further includes the following steps:

[0082] S41: Calculate the sum of the first similarity and the second similarity.

[0083] After the terminal device obtains the first similarity and the second similarity, it calculates the sum of the first similarity and the second similarity.

[0084] S42: Construct a second loss function based on the ratio of the first similarity and the sum.

[0085] The terminal device calculates the sum of the first similarity and the second similarity, and constructs a second loss function based on the ratio of the first similarity and the sum of the first similarity and the second similarity, as shown in the following formula:

[0086]

[0087] where Loss conw is the second loss function; C is the number of categories of the first training images; x c , are the features of different first training images of the same category, that is, the first image feature pairs; x c , are the features of different first training images of different categories, that is, the second image feature pairs; is the first similarity of the first image feature pair; is the second similarity of the second image feature pair.

[0088] In the embodiment of the present application, the terminal device calculates the sum of the first similarity and the second similarity, and constructs a second loss function based on the ratio of the first similarity and the sum. The model training method of this embodiment constructs a second loss function by obtaining the first image feature pairs of two first training images of the target category and the second image feature pairs of two first training images of different categories, so as to constrain the similarity of the features between different first training images of the same category and the dissimilarity of the features between different first training images of different categories, improve the classification effect of the model, and improve the accuracy of the model.

[0089] Please refer to Figure 7-8 , Figure 7 which is a schematic flowchart of the fifth embodiment of the model training method provided by the present application, Figure 8 is Figure 7 a schematic structural diagram of a first embodiment of the second training image in Figure 7 As shown in

[0090] S51: Mix the first training images of different categories to obtain a second training image and a second segmentation pseudo-label, where the second segmentation pseudo-label is obtained by mixing the first segmentation pseudo-labels of the first training images of different categories.

[0091] In this embodiment, the terminal device mixes the first training images of different categories to enhance the first training images, and then obtains the second training images, as shown in Figure 8 the schematic diagram on the left. The terminal device performs threshold segmentation on the second training images to obtain the second segmentation pseudo-labels, where the second segmentation pseudo-labels are obtained by mixing the first segmentation pseudo-labels of the first training images of different categories, as shown in Figure 8 the schematic diagram on the right.

[0092] Specifically, as shown in Figure 8 the terminal device crops a part of the area of the first training image and randomly fills the area pixel values of the first training images of different categories to obtain a second training image allocated proportionally, where the second training image is composed of areas of the first training images of the same category and different categories allocated proportionally.

[0093] S52: Input the second training image into the segmentation model to obtain the second prediction label of the second training image.

[0094] After the terminal device obtains the second training image and the second segmentation pseudo-label, it inputs the second training image and the second segmentation pseudo-label into the segmentation model for training to obtain the second prediction label of the second training image.

[0095] S53: Construct a third loss function based on the second prediction label and the second segmentation pseudo-label.

[0096] Based on the second prediction label and the second segmentation pseudo-label obtained by the terminal device, the terminal device constructs a third loss function. The construction process of the third loss function is similar to step S14 and will not be elaborated here.

[0097] S54: Extract third image feature pairs of the same class and fourth image feature pairs of different classes in the second training image.

[0098] To improve the model's ability to localize local image information, the terminal device further extracts third image feature pairs of the same class and fourth image feature pairs of different classes in the second training image.

[0099] S55: Obtain the third similarity of the third image feature pair and the fourth similarity of the fourth image feature pair.

[0100] Based on the third image feature pair and the fourth image feature pair extracted by the terminal device, the terminal device obtains the third similarity of the third image feature pair and the fourth similarity of the fourth image feature pair.

[0101] S56: Construct a fourth loss function based on the third similarity and the fourth similarity, and use the third loss function and the fourth loss function to train the segmentation model.

[0102] After the terminal device obtains the third similarity and the fourth similarity, it constructs a fourth loss function based on the third similarity and the fourth similarity, and uses the third loss function and the fourth loss function to train the segmentation model. The training process is similar to step S17 and will not be elaborated here. The process of constructing the fourth function is shown in the following formula:

[0103]

[0104] Where C represents the number of target classes in the second training image; x(u, v) represents the image feature at the (u, v) position extracted by the terminal device; represents the image features at two positions that belong to the same class as x(u, v) in the second training image; represents the image features at two positions that belong to different classes from x(u, v) in the second training image; sim(x(u, v) c , is the third similarity of the third image feature pair; is the fourth similarity of the fourth image feature pair.

[0105] The terminal device constructs a fourth loss function by obtaining the third similarity of the third image feature pair and the fourth similarity of the fourth image feature pair, so as to constrain the similarity of features between the same-class regions of the second training image and the dissimilarity of features between different-class regions of the second training image, thereby improving the classification effect of the model.

[0106] In the embodiment of the present application, the terminal device mixes first training images of different classes to obtain a second training image and a second segmentation pseudo-label, where the second segmentation pseudo-label is obtained by mixing the first segmentation pseudo-labels of the first training images of different classes; inputs the second training image into a segmentation model to obtain a second predicted label of the second training image; constructs a third loss function based on the second predicted label and the second segmentation pseudo-label; extracts third image feature pairs of the same class and fourth image feature pairs of different classes in the second training image; obtains the third similarity of the third image feature pair and the fourth similarity of the fourth image feature pair; constructs a fourth loss function based on the third similarity and the fourth similarity, and trains the segmentation model using the third loss function and the fourth loss function. Through the method of this embodiment, the first training images of different classes are mixed to obtain a second training image, the data of the image training set is enhanced, the training efficiency of the model is improved, and the localization ability of the model for local image information is further enhanced.

[0107] Further, in this model training method, the third image feature pair includes high-probability point features and uncertainty point features, where a high-probability point is an image pixel point whose prediction confidence is higher than a preset confidence threshold, and an uncertainty point is an image pixel point whose prediction confidence is lower than the preset confidence threshold; the fourth image feature pair includes high-probability point features of different-class regions.

[0108] Specifically, the third image feature pair is composed of high-probability point features and uncertainty point features at two positions in the same-class region of the second training image. The high-probability point feature is an image pixel point whose predicted confidence at this position is higher than the preset confidence threshold. Here, the preset confidence threshold can be 0.9 or other suitable preset confidence thresholds; the uncertainty point is an image pixel point whose predicted confidence at this position is lower than the preset confidence threshold.

[0109] Optionally, since the probability that an uncertainty point belongs to a certain class is too low when its prediction confidence is too low, in order to strengthen the constraint on the uncertainty point, the uncertainty point can be set as an image pixel point around the first preset confidence threshold. For example, when the first preset confidence threshold is 0.5, the uncertainty point can be a feature whose prediction confidence is in the range of 0.4 - 0.6. The first preset confidence threshold can be other preset thresholds. Here, the first preset confidence threshold and the range of the uncertainty point are not specifically limited.

[0110] In this embodiment, by obtaining the third similarity between the high-probability point features and the uncertainty point features at two positions in the same-class regions of the second training image, and the fourth similarity between the high-probability point features in the different-class regions of the second training image, the similarity of the same-class regions in the second training image can be effectively restricted, and the dissimilarity of the different-class regions in the second training image can be restricted, further improving the model's ability to distinguish local image information and the accuracy of the model.

[0111] Please refer to Figure 9 , Figure 9 which is a schematic flowchart of the sixth embodiment of the model training method provided by this application. As Figure 9 shown, the model training method further includes the following steps:

[0112] S61: Determine the foreground regions of each category in the second training image based on the second segmentation pseudo-label.

[0113] After the terminal device obtains the second training image, it performs threshold segmentation on the second training image to obtain the second segmentation pseudo-label, and uses the foreground regions in the first segmentation pseudo-label of each category after segmentation mixing as the foreground regions of each category in the second training image.

[0114] S62: Obtain the number of pixels in the foreground regions of each category based on the second prediction label, and the sum of the pixel prediction probabilities of each category in the second training image.

[0115] After the terminal device determines the foreground regions of each category in the second training image, it obtains the number of image pixel points in the foreground regions of each category; and obtains the sum of the confidences that each pixel point in the second training image is predicted as each category, that is, the sum of the pixel prediction probabilities of each category in the second training image, as shown in the following formula:

[0116]

[0117] where S k is the sum of the pixel prediction probabilities, k is the target category in the image training set, represents the confidence that the i-th pixel point in the second training image is predicted as each category.

[0118] S63: Construct a fifth loss function based on the number of pixels and the sum of the pixel prediction probabilities.

[0119] The terminal device obtains the number of pixels and the sum of the pixel prediction probabilities, and constructs a fifth loss function based on the number of pixels and the second pixel prediction probability sum, as shown in the following formula:

[0120]

[0121] Among them, Loss area is the fifth loss function; T represents the second training image; k is the target category in the image training set; A k is the number of pixels in the foreground region of category k, and S k is the sum of pixel prediction probabilities predicted as category k.

[0122] S64: Train the segmentation model using the third loss function, the fourth loss function, and the fifth loss function.

[0123] The terminal device obtains the third loss function, the fourth loss function, and the fifth loss function, and trains the segmentation model using the third loss function, the fourth loss function, and the fifth loss function, as shown in the following formula:

[0124] Loss total2 = Loss seg + Loss conl + Loss area ;

[0125] Among them, Loss total2 is the total loss function of the second training image, Loss seg is the third loss function; Loss conl is the fourth loss function, and Loss area is the fifth loss function.

[0126] Specifically, during the training process, in order to constrain the size of the cutting region in the second training image, the terminal device obtains the number of pixels in the foreground region of each category and the sum of pixel prediction probabilities of each category in the second training image, and constructs the fifth loss function to train the model, effectively improving the accuracy of the first training image segmentation mixture, and further improving the accuracy of model prediction.

[0127] In the embodiment of the present application, the terminal device determines the foreground region of each category in the second training image based on the second segmentation pseudo-label; obtains the number of pixels in the foreground region of each category, and the sum of pixel prediction probabilities of each category in the second training image based on the second prediction label; constructs the fifth loss function based on the number of pixels and the sum of pixel prediction probabilities; trains the segmentation model using the third loss function, the fourth loss function, and the fifth loss function. Through the method of this embodiment, the model can be trained using the second training image composed of the first training images spliced together, and the fifth loss function is introduced to constrain the size of the cutting region in the second training image, effectively improving the model's ability to distinguish local image information and improving the accuracy of the model.

[0128] Please refer to Figure 10 , Figure 10It is a schematic flowchart of the seventh embodiment of the model training method provided by this application. The image training set further includes a third training image and a true image-level label annotated for the third training image, where the true image-level label annotates the categories included in the third training image. As Figure 10 shown, before training the model using the first training image, the model training method further includes:

[0129] S71: Input the third training image into the segmentation model to obtain the predicted image-level label of the third training image.

[0130] Specifically, the terminal device obtains the third training image in the image training set. The third training image includes images of multiple target categories. Among them, the third training image is set with a true image-level label to represent the target categories included in the third training image. The terminal device inputs the third training image into the segmentation model to obtain the predicted image-level label of the third training image.

[0131] S72: Construct a sixth loss function based on the true image-level label and the predicted image-level label, and use the sixth loss function to train the segmentation model.

[0132] Further, after inputting the third training image into the segmentation model and the fully connected layer outputs features of nxC, the terminal device inputs the features into the CAM classification network of the model for training. There is a class response map set on the CAM classification network. By calculating the weighted sum of the average values of each feature map, the class activation map is upsampled to the size of the input image, and the image region most relevant to the preset category is identified. Its loss function is shown as follows:

[0133]

[0134] Among them, Loss cls is the sixth loss function; T represents the target categories existing in the third training image; is the target category that does not exist in the third training image; S k represents the probability score that the target category is k.

[0135] Since the sixth loss function can suppress the categories that do not exist in the image and increase the prediction probability of the categories that exist in the image, the CAM classification network can output the region with the most significant discriminability from the target category after training. After the terminal device obtains the region with the most significant discriminability from each target category, it can segment the image training set to be segmented to obtain several first training images of a single category.

[0136] In an embodiment of the present application, the terminal device inputs a third training image into a segmentation model to obtain a predicted image-level label of the third training image; constructs a sixth loss function based on the true image-level label and the predicted image-level label, and uses the sixth loss function to train the segmentation model. Through the method of this embodiment, the terminal device can construct a sixth loss function based on the true image-level label and the predicted image-level label to train the segmentation model and improve the accuracy of model prediction.

[0137] Please refer to Figure 11 , Figure 11 which is a schematic flowchart of the eighth embodiment of the model training method provided by the present application. As Figure 11 shown, the model training method further includes:

[0138] S81: Input a third training image into the segmentation model to obtain the image region corresponding to each category in the predicted image-level label.

[0139] After the terminal device obtains the third training image, it can input the third training image into the segmentation model to obtain the predicted image-level label of the third training image, and segment the image region corresponding to each category in the predicted image-level label.

[0140] S82: Use the image region corresponding to each category to segment a first training image of a single category.

[0141] After the terminal device obtains the image region corresponding to each category in the predicted image-level label, it can segment a first training image of a single category from the image region corresponding to each category to expand the number of first training images, and then use the segmented first training images to train the model to improve the training effect of the model.

[0142] In an embodiment of the present application, the terminal device inputs a third training image into the segmentation model to obtain the image region corresponding to each category in the predicted image-level label; uses the image region corresponding to each category to segment a first training image of a single category. Through the method of the present application, the terminal device can segment a first training image of a single category from a multi-category training image, improve the sample diversity of the first training images, and improve the training effect of the model.

[0143] Please refer to Figure 12-13 , Figure 12 which is a schematic flowchart of the ninth embodiment of the model training method provided by the present application, Figure 13 is Figure 12 the schematic flowchart of the embodiment for obtaining the category response in Figure 1 As Figure 12 shown, the image training set further includes fourth training images of multiple categories, and the model training method further includes:

[0144] S91: Concatenate the first training images of a single category with the fourth training images of multiple categories to obtain the fifth training image.

[0145] After training the segmentation model using the first loss function and the second loss function, the model can extract similar features for the same semantic regions, and then obtain the same predicted categories, completing the pre-training phase. To further improve the prediction accuracy of the model, the terminal device performs fine-tuning on the model.

[0146] Specifically, as Figure 13 shown, the terminal device obtains the fourth training images, where the fourth training images include training images of multiple categories. The terminal device mixes and concatenates the fourth training images of multiple categories with the first training images of a single category to obtain the fifth training image.

[0147] S92: Input the fifth training image into the segmentation model to obtain the category response map of the fifth training image, where the category response map includes the third predicted label of the fifth training image.

[0148] The terminal device inputs the fifth training image into the segmentation model to obtain the category response map of the fifth training image. The category response map represents the response of the feature map of the fifth training image to a single category. Among them, the category response map includes the third predicted label of the fifth training image.

[0149] S93: Obtain the first category response map corresponding to the first training image and the second category response map corresponding to the fourth training image by concatenating the category response map.

[0150] After the terminal device obtains the category response map of the fifth training image, since the fifth training image is an image formed by concatenating the first training images of a single category and the fourth training images of multiple categories, the terminal device can concatenate the category response map again to restore the category response map into the first category response map corresponding to the first training image and the second category response map corresponding to the fourth training image.

[0151] S94: Construct the seventh loss function based on the segmentation pseudo-label of the first training image and the predicted label of the first category response map.

[0152] After the terminal device obtains the first category response map corresponding to the first training image, it constructs the seventh loss function based on the segmentation pseudo-label of the first training image and the predicted label of the first category response map, as shown in the following formula:

[0153]

[0154] Among them, Loss seg is the seventh loss function; C is the number of target categories in the image training set; Z Cis the predicted label of the first category response map; l c is the segmentation pseudo-label of the first training image.

[0155] S95: Construct the eighth loss function based on the segmentation pseudo-label of the fourth training image and the predicted label of the second category response map.

[0156] The terminal device obtains the second category response map corresponding to the fourth training image, and constructs the eighth loss function based on the segmentation pseudo-label of the fourth training image and the predicted label of the second category response map. The construction process is similar to step S94 and will not be elaborated here.

[0157] S96: Use the seventh loss function and the eighth loss function to train the segmentation model.

[0158] After the terminal device obtains the seventh loss function and the eighth loss function, it can use the seventh loss function and the eighth loss function to train the segmentation model, as shown in the following formula:

[0159] Loss total3 = Loss seg-k + Loss seg-uk ;

[0160] where, Loss total3 is the total loss function of the fifth training image; Loss seg-k is the seventh loss function of the first training image; Loss seg-uk is the eighth loss function of the fourth training image.

[0161] Optionally, as Figure 13 shown, the fifth training image can also be composed of spliced fourth training images of different categories, and the category response maps of the fifth training image are spliced again to obtain the first category response map and the second category response map of different fourth training images. The training process is similar to steps S91 - S96 and will not be elaborated here.

[0162] In an embodiment of the present application, the terminal device splices the first training images of a single category with the fourth training images of multiple categories to obtain the fifth training images; inputs the fifth training images into the segmentation model to obtain the category response maps of the fifth training images, where the category response maps include the third predicted labels of the fifth training images; obtains the first category response map corresponding to the first training images and the second category response map corresponding to the fourth training images by splicing the category response maps; constructs the seventh loss function based on the segmentation pseudo-labels of the first training images and the predicted labels of the first category response maps; and trains the segmentation model using the seventh loss function and the eighth loss function. Through the method of this embodiment, the terminal device can use the fifth training images spliced from the fourth training images of multiple categories and the first training images of a single category to train the model, so as to improve the prediction accuracy of the model for the training images of multiple categories and improve the classification accuracy of the model.

[0163] Optionally, step S96 further includes the following steps: adjusting the weight of the eighth loss function using the adjustment parameter; training the segmentation model using the seventh loss function and the adjusted eighth loss function; where the value of the adjustment parameter is determined by a preset growth function.

[0164] Specifically, since in the early stage of model training, the accuracy of the prediction results output by the model is relatively low, and the accuracy of the segmentation pseudo-labels generated by the model for the input fourth training images of multiple categories is not high either. In order to reduce the adverse effects caused by the low accuracy of the fourth training images of multiple categories, an adjustment parameter is introduced to adjust the weight of the eighth loss function when constructing the segmentation model, and the segmentation model is trained using the seventh loss function and the adjusted eighth loss function, as shown in the following formula:

[0165] Loss total3 =Loss seg-k +w*Loss seg-uk ;

[0166] where, Loss total3 is the total loss function of the fifth training images; Loss seg-k is the seventh loss function of the first training images; Loss seg-uk is the eighth loss function of the fourth training images; w is the adjustment parameter.

[0167] In a specific implementation manner, the value of the adjustment parameter can be determined by a preset growth function. In the initial stage of model training, the adjustment parameter can be set to 0.2 to reduce the adverse effects of the eighth loss function on model training. As the accuracy of the segmentation pseudo-labels generated by the model during the iteration process increases, the adjustment parameter can be gradually increased to 1. The value of the adjustment parameter can also be set to other parameters. Here, the growth process of the adjustment parameter is not specifically limited.

[0168] The present application also provides an image recognition method, and the steps of the image recognition method include: inputting an image to be recognized into a segmentation model to obtain the image recognition category of the image to be recognized, where the segmentation model is obtained by using the model training method described in the above embodiments.

[0169] Specifically, the segmentation model obtained by training with the model training method described in the above embodiments can be used for image recognition. The user inputs the image to be recognized into the segmentation model, and the segmentation model can perform semantic segmentation on the image to be recognized to output the image recognition category of the image to be recognized. The user can perform operations such as disease diagnosis based on the obtained image recognition category.

[0170] Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of an embodiment of a terminal device provided by the present application. The terminal device includes a memory 52 and a processor 51 connected to each other.

[0171] The memory 52 is used to store program instructions for implementing the model training method described in any of the above embodiments.

[0172] The processor 51 is used to execute the program instructions stored in the memory 52.

[0173] Among them, the processor 51 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 51 may be an integrated circuit chip with signaling processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0174] The memory 52 may be a memory stick, a TF card, etc., and can store all information in the terminal device. All input original data, computer programs, intermediate operation results, and final operation results are stored in the memory. It stores and retrieves information according to the positions specified by the controller. With the memory, the string matching prediction device has a memory function and can ensure normal operation. The memory of the string matching prediction device can be classified into a main memory (memory) and an auxiliary memory (external memory) according to its use, or there is also a classification method of external memory and internal memory. The external memory is usually a magnetic medium or an optical disc, etc., which can store information for a long time. The memory refers to the storage component on the motherboard, which is used to store the data and programs being currently executed, but only temporarily stores programs and data. When the power is turned off or cut off, the data will be lost.

[0175] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division manners. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0176] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0178] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a system server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application.

[0179] Please refer to Figure 15 , Figure 15It is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by this application. The computer-readable storage medium of this application stores a program file 61 that can implement all the above model training methods. Among them, the program file 61 can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage device includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0180] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.

Claims

1. A model training method, characterized in that, The model training method includes: Obtain an image training set, where the image training set includes a number of first training images of a single category; Obtain the first segmentation pseudo-labels of the first training images of each category; Input a number of the first training images into the segmentation model to be trained, and obtain the first prediction labels of the first training images; Construct a first loss function based on the first prediction labels and the first segmentation pseudo-labels; Extract the first image feature pairs of the first training images of the same category, and the second image feature pairs of the first training images of different categories; Obtain the first similarity of the first image feature pairs, and the second similarity of the second image feature pairs; Construct a second loss function based on the first similarity and the second similarity, and use the first loss function and the second loss function to train the segmentation model; Among them, obtaining the segmentation pseudo-labels of the first training images of each category includes: Normalize the first training images; Based on a preset segmentation threshold, distinguish the foreground region and the background region in the normalized first training images, so as to obtain the first segmentation pseudo-labels of the first training images; The constructing the second loss function based on the first similarity and the second similarity includes: Calculate the sum of the first similarity and the second similarity; Construct the second loss function based on the ratio of the first similarity and the sum.

2. The model training method according to claim 1, wherein The obtaining the first similarity of the first image feature pairs, and the second similarity of the second image feature pairs includes: Respectively obtain the first image features and the second image features of two first training images of the target category; Obtain the first similarity between the first image feature and the second image feature; Respectively obtain the third image features of one first training image of the target category and the fourth image features of one first training image of other categories; Obtain the second similarity between the third image feature and the fourth image feature.

3. The model training method according to claim 1, wherein After obtaining the segmentation pseudo-labels of the first training images of each category, the model training method further includes: Mix the first training images of different categories to obtain second training images and second segmentation pseudo-labels, where the second segmentation pseudo-labels are obtained by mixing the first segmentation pseudo-labels of the first training images of different categories; Input the second training images into the segmentation model to obtain the second prediction labels of the second training images; Construct a third loss function based on the second prediction labels and the second segmentation pseudo-labels; Extract the third image feature pairs of the same category in the second training images, and the fourth image feature pairs of different categories; Obtain the third similarity of the third image feature pairs, and the fourth similarity of the fourth image feature pairs; Construct a fourth loss function based on the third similarity and the fourth similarity, and use the third loss function and the fourth loss function to train the segmentation model.

4. The model training method according to claim 3, wherein the third image feature pair includes high-probability point features and uncertainty point features. Among them, the high-probability points are image pixel points with a prediction confidence higher than a preset confidence threshold, and the uncertainty points are image pixel points with a prediction confidence lower than the preset confidence threshold; the fourth image feature pair includes high-probability point features of different category regions.

5. The model training method according to claim 3 or 4, wherein the model training method further includes: determining the foreground region of each category in the second training image based on the second segmentation pseudo-label; obtaining the number of pixels in the foreground region of each category based on the second prediction label, and the sum of pixel prediction probabilities of each category in the second training image; constructing a fifth loss function based on the number of pixels and the sum of pixel prediction probabilities; training the segmentation model using the third loss function, the fourth loss function, and the fifth loss function.

6. The model training method according to claim 1, wherein the image training set further includes a third training image and a true image-level label annotated for the third training image, wherein the true image-level label annotates the categories included in the third training image; the model training method further includes: inputting the third training image into the segmentation model to obtain a predicted image-level label of the third training image; constructing a sixth loss function based on the true image-level label and the predicted image-level label, and training the segmentation model using the sixth loss function.

7. The model training method according to claim 6, wherein the model training method further includes: inputting the third training image into the segmentation model to obtain the image region corresponding to each category in the predicted image-level label; using the image region corresponding to each category to segment a single-category first training image.

8. The model training method according to claim 1, wherein the image training set further includes multi-category fourth training images; the model training method further includes: stitching a single-category first training image and multi-category fourth training images to obtain a fifth training image; inputting the fifth training image into the segmentation model to obtain a category response map of the fifth training image, wherein the category response map includes a third prediction label of the fifth training image; obtaining a first category response map corresponding to the first training image and a second category response map corresponding to the fourth training image by stitching the category response map; constructing a seventh loss function based on the segmentation pseudo-label of the first training image and the prediction label of the first category response map; constructing an eighth loss function based on the segmentation pseudo-label of the fourth training image and the prediction label of the second category response map; training the segmentation model using the seventh loss function and the eighth loss function.

9. The model training method according to claim 8, wherein Training the segmentation model by using the seventh loss function and the eighth loss function includes: Adjusting the weight of the eighth loss function by using an adjustment parameter; Training the segmentation model by using the seventh loss function and the adjusted eighth loss function; wherein, the value of the adjustment parameter is determined by a preset growth function.

10. An image recognition method, characterized in that, It includes: Inputting an image to be recognized into the segmentation model to obtain the image recognition category of the image to be recognized, wherein the segmentation model is obtained by using the model training method according to any one of claims 1-9.

11. A terminal device, characterized in that, The terminal device includes a processor and a memory connected to the processor, wherein the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the model training method according to any one of claims 1-9 and / or the image recognition method according to claim 10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and when the program instructions are executed, the model training method according to any one of claims 1-9 and / or the image recognition method according to claim 10 are implemented.

Citation Information

Patent Citations

  • Neural network training method, image processing method and device

    CN112990211A

  • Target recognition model training method and device and electronic equipment

    CN112990432A