Image classification machine learning method and system based on region of interest and smooth annotation
By dividing the image into multiple areas of interest and smoothly annotating, the problem of lack of intermediate state for the annotation of image classification data sets in the prior art is solved, and the classification accuracy of the image classification model is improved.
Patent Information
- Application Number
- CN202510232524.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing image classification data set annotation method lacks intermediate states, resulting in low accuracy of image classification models.
By dividing the image into multiple areas of interest and setting classification labels for each area of interest, a smooth annotation method is used to generate an annotation data set for training the image classification model.
The training effect and classification accuracy of the image classification model are improved, and the situation where some of the images belong to a certain classification can be better handled.
Smart Images

Figure CN119963923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and in particular to an image classification machine learning method and system based on focus area and smooth annotation. Background Art
[0002] The prior art generally includes the following steps: (1) Annotate the image dataset to indicate the category of each image; each image can belong to one or more categories. (2) Use machine learning methods to train an image classification model on the aforementioned dataset. (3) Use the aforementioned model to classify images that are not in the aforementioned dataset.
[0003] There are some shortcomings in the existing technology: In the traditional image classification dataset annotation method, for any image in the dataset, it either belongs to a certain category or not, with no intermediate state. This is often too rigid and cannot well reflect the actual situation, that is, some images only partially belong to a certain category.
[0004] Image classification models trained using rigidly labeled datasets will have limited accuracy. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides an image classification machine learning method and system based on focus region and smooth annotation to solve the technical problem that the existing image classification data set annotations lack intermediate states, resulting in low accuracy of image classification models trained based on the existing image classification data set.
[0006] To achieve the above object, in a first aspect, a machine learning method for image classification based on region of interest and smooth annotation is provided, comprising the following steps:
[0007] S10: Acquire some image data and
[0008] generating an image data set according to the plurality of graphic data;
[0009] S20: Divide each image in the image data set into one or more regions of interest, and set one or more classification labels for each region of interest;
[0010] S30: annotating each image in the image data set according to the focus area and the classification label to obtain an annotated data set, wherein the annotation includes a confidence level and an image classification;
[0011] S40: Using a machine learning algorithm, train an image classification model based on the image dataset and the annotated dataset.
[0012] A further improvement of the present invention is that: the step S10 specifically includes:
[0013] S11: performing image enhancement on the plurality of image data to obtain a plurality of enhanced image data;
[0014] S12: unifying the sizes of the plurality of enhanced image data by using an interpolation method to obtain a plurality of unified image data;
[0015] S13: performing normalization processing on the plurality of unified image data to obtain a plurality of normalized image data;
[0016] S14: Generate an image data set according to the plurality of normalized image data.
[0017] A further improvement of the present invention is that: the step S20 specifically includes:
[0018] S21: extracting features according to the image data set to obtain a number of basic features;
[0019] S22: generating a plurality of classification labels according to the plurality of basic features;
[0020] S23: for dividing each of the images into one or more regions of interest according to the image data set and a preset image division model;
[0021] S24: According to the plurality of classification labels, one or more labels are set for each region of interest.
[0022] A further improvement of the present invention is that: the step S30 specifically includes:
[0023] S31A: According to one or more regions of interest in each of the images, select a number of images without overlap of the regions of interest as a number of first images;
[0024] S32A: Calculate the areas of the plurality of first images and the area of each focus region in each first image;
[0025] S33A: In each of the first images, calculate the relative area of each focus region in the first image according to the area of the first image and the area of each focus region in the first image;
[0026] S34A: In each of the first images, the relative areas of the plurality of focus regions corresponding to the same classification label are added to obtain a first weight of the same classification label, and all classification labels are traversed to obtain a first weight of each classification label in the first image;
[0027] S35A: Calculate a first confidence level and a first image classification according to first weights of the classification labels in the first image;
[0028] S36A: Annotate the first image according to the first confidence level and the first image classification to obtain an annotated data set.
[0029] A further improvement of the present invention is that: the step S30 specifically includes:
[0030] S31B: screening out, according to the plurality of interest regions in each of the images, a plurality of images having overlapping interest regions as a plurality of second images;
[0031] S32B: according to the plurality of regions of interest in each of the second images and one or more classification labels corresponding to each of the regions of interest, using the plurality of classification labels existing in each of the second images as the plurality of classification labels to be processed of each of the second images;
[0032] S33B: In each of the second images, one or more of the to-be-processed classification labels are selected as feature classifications, and the rest are selected as target classifications, the focus area corresponding to the target classification is the target area, and the focus area corresponding to the feature classification is the feature area;
[0033] S34B: In each of the second images, for each target category, the number of overlaps between the target region and the feature region is screened out as a second weight;
[0034] S35B: In each of the second images, calculating a second confidence level of the target classification according to the second weight and a preset correlation coefficient;
[0035] S36B: Annotate the second image according to the target classification and the second confidence level of the target classification to obtain an annotated data set.
[0036] A further improvement of the present invention is that in step S40, the following steps are specifically included:
[0037] S41: Establish an initial image classification model,
[0038] And set the loss function;
[0039] S42: inputting the image data set into an initial image classification model to generate a prediction value;
[0040] S43: Calculate the difference between the labeled data set and the predicted value according to the loss function to obtain a numerical residual;
[0041] S44: Update the initial image classification model according to the numerical residual and the predicted value, repeat steps S42 to S43 until the loss function meets the preset cutoff condition, and output the initial image classification model at this time as the image classification model.
[0042] In a second aspect, the present invention provides an image classification machine learning system based on focus regions and smooth annotations, comprising:
[0043] A data acquisition module, used to acquire a number of image data and generate an image data set according to the number of graphic data;
[0044] A division module, used for dividing each image in the image data set into one or more regions of interest, and setting one or more classification labels for each region of interest;
[0045] A labeling module, used to label each image in the image data set according to the focus area and the classification label to obtain a labeling data set, wherein the labeling includes a confidence level and an image classification;
[0046] The model training module is used to adopt a machine learning algorithm to train an image classification model based on the image data set and the labeled data set.
[0047] In a third aspect, an electronic device is provided, comprising:
[0048] one or more processors;
[0049] a storage device for storing one or more programs,
[0050] When the one or more programs are executed by the one or more processors, the one or more processors implement the image classification machine learning method and system based on focus area and smooth annotation as described in the first aspect.
[0051] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the image classification machine learning method and system based on focus area and smooth annotation as described in the first aspect are implemented.
[0052] The above technical solution has the following beneficial technical effects:
[0053] The present invention divides an image into several regions of interest and annotates the image according to the regions of interest and their classification labels, thereby efficiently producing smooth annotated data, thereby improving the training effect of the image classification model and effectively improving the classification accuracy of the image classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0055] Figure 1 is a flow chart of the image classification machine learning method based on focus area and smooth annotation of the present invention;
[0056] Figure 2 is a flow chart of step S10 in the present invention;
[0057] Figure 3 is a flow chart of step S20 in the present invention;
[0058] Figure 4 is a flow chart of step S30A in the present invention;
[0059] Figure 5 is a flow chart of step S30B in the present invention;
[0060] Figure 6 is a flow chart of step S40 in the present invention;
[0061] Figure 7 is a flow chart of the image classification machine learning system based on focus region and smooth annotation of the present invention;
[0062] Figure 8 It is a schematic diagram of the structure of a computer system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0064] Embodiment 1
[0065] like Figure 1 As shown, the embodiment of the present invention provides an image classification machine learning method based on focus area and smooth annotation, which includes the following steps:
[0066] S10: Acquire a plurality of image data, and generate an image data set according to the plurality of graphic data;
[0067] S20: Divide each image in the image data set into one or more regions of interest, and set one or more classification labels for each region of interest;
[0068] S30: annotating each image in the image data set according to the focus area and the classification label to obtain an annotated data set, wherein the annotation includes a confidence level and an image classification;
[0069] S40: Using a machine learning algorithm, train an image classification model based on the image dataset and the annotated dataset.
[0070] Specifically, Figure 2 As shown, step S10 includes:
[0071] S11: performing image enhancement on the plurality of image data to obtain a plurality of enhanced image data;
[0072] S12: unifying the sizes of the plurality of enhanced image data by using an interpolation method to obtain a plurality of unified image data;
[0073] S13: performing normalization processing on the plurality of unified image data to obtain a plurality of normalized image data;
[0074] S14: Generate an image data set according to the plurality of normalized image data.
[0075] Specifically, in step S11, the image enhancement includes contrast enhancement, sharpening, denoising and color enhancement, etc. The contrast enhancement makes the bright part of the image brighter and the dark part darker by adjusting the histogram of the image. The sharpening makes the image clearer by enhancing the edges and details of the image. The denoising is used to remove noise in the image, such as Gaussian noise or pepper and salt noise. The color enhancement improves the visual effect by adjusting the color balance or saturation of the image. Step S11 improves the visual effect of the image and enhances certain features in the image, making it easier for the machine to recognize.
[0076] Specifically, in step S12, the interpolation method is used to adjust all images to the same resolution so that they can be processed uniformly. The interpolation method includes nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.
[0077] For example, when the nearest neighbor interpolation is used to resize an image, the value of the target pixel is directly taken from the value of the closest source image pixel. This method does not introduce new pixel values, so the calculation speed is fast. The nearest neighbor interpolation adjusts the image to the specified resolution through Python and OpenCV libraries. The bilinear interpolation is suitable for image scaling, with fast calculation speed and good image quality. The bicubic interpolation (cv2.INTER_CUBIC) has higher image quality, but the calculation speed is slow, which is suitable for scenes with high image quality requirements.
[0078] Specifically, in step S13, the normalization is used to adjust the pixel values of the image to a specific range (for example, [0, 1] or [-1, 1]) to eliminate the difference in the range of pixel values between different images. Normalization is used for the input of deep learning models because the model usually has certain requirements on the range of input data. The normalization methods include linear normalization and mean normalization. The current normalization is used to scale the pixel values from the original range to [0, 1]. The mean normalization is used to subtract the mean from the pixel value and divide it by the standard deviation. The linear normalization is suitable for data with strict requirements on the data range and relatively uniform data distribution, but is sensitive to outliers. The mean normalization is suitable for scenarios with high requirements on data centralization and relatively stable data distribution (close to normal distribution), and is more robust to outliers. Make a selection based on the quality of the acquired image data.
[0079] Specifically, Figure 3 As shown, step S20 includes:
[0080] S21: extracting features according to the image data set to obtain a number of basic features;
[0081] S22: generating a plurality of classification labels according to the plurality of basic features;
[0082] S23: for dividing each of the images into one or more regions of interest according to the image data set and a preset image division model;
[0083] S24: According to the plurality of classification labels, one or more labels are set for each region of interest.
[0084] Specifically, in step S21 and step S22, the feature extraction includes color feature extraction, texture feature extraction, shape feature extraction or feature extraction based on transformation. The color feature extraction is used to extract the histogram or color moment of the image. The texture feature extraction is performed by using the Gray-level Co-occurrence Matrix (Gray-level Co-occurrence Matrix,
[0085] The shape feature extraction is used to extract contour or boundary information. The transformation-based feature extraction is performed by using wavelet transform or Fourier transform. In step S21, a deep learning method can also be used to establish a neural network to automatically learn the feature representation of the image, thereby realizing feature extraction.
[0086] Specifically, in steps S23 and S24, the preset image segmentation model is based on Labelme or other segmentation tools, and a classification label is added to each area. The Labelme generates a JSON file, which contains the coordinates and classification labels of each area of interest. For example, it is necessary to obtain an image classification model that can classify "eating" and "reading" scenes. When dividing images, there are 10 people in one of the images, 7 of whom are eating and 3 are reading. Draw a rectangular box in Labelme to surround the 7 people eating (this is a "area of interest"), and set the classification label of this box to "eating". Draw another rectangular box that just surrounds the 3 people reading (this is also a "area of interest"), and set the classification label of this box to "reading". If the two areas of interest overlap, they can be split into more rectangular areas, or polygonal areas of interest can be used to eliminate overlap; or overlap is allowed.
[0087] Specifically, Figure 4 As shown, step S30 includes:
[0088] S31A: According to one or more regions of interest in each of the images, select a number of images without overlap of the regions of interest as a number of first images;
[0089] S32A: Calculate the areas of the plurality of first images and the area of each focus region in each first image;
[0090] S33A: In each of the first images, calculate the relative area of each focus region in the first image according to the area of the first image and the area of each focus region in the first image;
[0091] S34A, in each of the first images, adding the relative areas of the plurality of focus regions corresponding to the same classification label to obtain a first weight of the same classification label, traversing all classification labels to obtain a first weight of each classification label in the first image;
[0092] S35A, calculating a first confidence level and a first image classification according to first weights of the classification labels in the first image;
[0093] S36A: Annotate the first image according to the first confidence level and the first image classification to obtain an annotated data set.
[0094] Specifically, in step S31A, the image with only one area of interest does not need to be screened and can be directly used as the first image to participate in subsequent steps. During screening, whether there is overlap is determined by judging the geometric data (such as coordinates or boundaries, etc.) between the two areas of interest.
[0095] Specifically, in step S33A, the calculation formula of the relative area is as follows:
[0096] Relative area = area of the region of interest / area of the image where the region of interest is located;
[0097] Specifically, in step S35A, the confidence calculation formula is as follows:
[0099] First confidence = first weight α ;
[0100] In the formula, α is the first confidence constant, preferably 0.5, and the first confidence formula is used for the same classification label. Assuming that the first weight of the "eating" area is 0.4, the first confidence that this image belongs to the "eating" category is 0.4 0.5 =0.63; the first weight of the "reading" area is 0.2, so the first confidence that this image belongs to the "reading" category is 0.2 0.5 =0.45. The annotation includes a plurality of first confidences and a plurality of first image classifications, wherein the plurality of image classifications are classification labels corresponding to the plurality of first confidences. For example, in image A, the first confidence of the classification label "reading" is 0.45, and the image label of image A contains "reading".
[0101] Specifically, Figure 5 As shown, step S30 includes:
[0102] S31B: screening out, according to the plurality of interest regions in each of the images, a plurality of images having overlapping interest regions as a plurality of second images;
[0103] S32B: according to the plurality of regions of interest in each of the second images and one or more classification labels corresponding to each of the regions of interest, using the plurality of classification labels existing in each of the second images as the plurality of classification labels to be processed of each of the second images;
[0104] S33B: In each of the second images, one or more of the to-be-processed classification labels are selected as feature classifications, and the rest are selected as target classifications, the focus area corresponding to the target classification is the target area, and the focus area corresponding to the feature classification is the feature area;
[0105] S34B: In each of the second images, for each target category, the number of overlaps between the target region and the feature region is screened out as a second weight;
[0106] S35B: In each of the second images, calculating a second confidence level of the target classification according to the second weight and a preset correlation coefficient;
[0107] S36B: Annotate the second image according to the target classification and the second confidence level of the target classification to obtain an annotated data set.
[0108] Specifically, the calculation formula of the second confidence level is as follows:
[0109]
[0110] Wherein, β is the second confidence constant. Assuming that the feature classification is "person", the target classification is "eating" and "reading", the correlation coefficient between "person" and "eating" is 1, and the correlation coefficient between "person" and "reading" is 1. The number of "person" regions intersecting with the "eating" region is 7, and the number of "person" regions intersecting with the "reading" region is 3. Then the confidence that this image belongs to the "eating" classification is Sigmoid (7*1 / 5) = 0.80, and the confidence that it belongs to the "reading" classification is: Sigmoid (3*1 / 5) = 0.65. The confidence in step S30 includes the first confidence and the second confidence.
[0111] Specifically, Figure 6 As shown, step S40 includes:
[0112] S41: Establish an initial image classification model and set the loss function;
[0113] S42: inputting the image data set into an initial image classification model to generate a prediction value;
[0114] S43: Calculate the difference between the labeled data set and the predicted value according to the loss function to obtain a numerical residual;
[0115] S44: Update the initial image classification model according to the numerical residual and the predicted value, repeat steps S42 to S43 until the loss function meets the preset cutoff condition, and output the initial image classification model at this time as the image classification model.
[0116] Specifically, in step S41, the establishment of the initial image classification model includes the selection of an optimizer, the setting of a learning rate and a learning rate decay strategy, the setting of a batch size, and the setting of the number of iterations, etc. The optimizer includes stochastic gradient descent (SGD), adaptive matrix estimation (Adam), or root mean square propagation (RMSprop). The stochastic gradient descent calculates the gradient by randomly selecting one or more samples (small batches) to update the parameters of the model. The advantage of SGD is that it has high computational efficiency and is suitable for large-scale data sets, but the convergence speed may be slow and it is easy to fall into a local optimal solution. The adaptive moment estimation dynamically adjusts the learning rate of each parameter by calculating the first-order moment (mean) and second-order moment (uncentered variance) of the gradient. The advantage of Adam is that it converges quickly, is suitable for non-convex optimization problems, and is robust to the selection of hyperparameters. The root mean square propagation dynamically changes the learning rate by adjusting the sliding average of the square of the gradient. It is particularly suitable for dealing with problems with large gradient changes and can effectively prevent gradient explosion or gradient disappearance. The core idea of RMSprop is to normalize the gradient by dividing by the mean of the square of the gradient, thereby stabilizing the optimization process. The batch size is used to determine the number of samples for each training, and the number of iterations is used to determine the total number of training. The loss function adopts the cross entropy loss function.
[0117] Specifically, in step S42, the image data set is divided into a training set, a test set and a validation set according to a preset ratio (e.g., 7:1.5:1.5). The image data input into the initial image classification model in step S42 are all training set data, and only the training set data are involved in the subsequent iterative update process. The validation set is used to evaluate the performance of the model during the training process. Through the validation set, the performance of different hyperparameters can be compared to select the optimal model and hyperparameters. The validation set helps detect whether the model is overfitting (i.e., it performs well on the training set but poorly on new data). If the model performs well on the training set but performs poorly on the validation set, it means that the model may be overfitting. The validation set is also used to adjust the hyperparameters of the model (e.g., learning rate, regularization parameter, number of network layers, etc.). By evaluating the performance of different hyperparameters on the validation set, the optimal hyperparameter combination is selected. The test set is used to evaluate the performance of the model on completely unseen data. The test set is used only after the model training and validation are completed, and is used to evaluate the generalization ability of the model. The purpose of the test set is to simulate the performance of the model in actual applications. It provides an independent evaluation environment to ensure the performance of the model on real data.
[0118] Specifically, the numerical residual is the difference between the predicted value and the data in the labeled data set.
[0119] The network weights are updated through back propagation.
[0120] Specifically, in step S44, each time the initial image classification model is updated, data is saved regularly to avoid data loss.
[0121] Embodiment 2
[0122] like Figure 7 As shown, an embodiment of the present invention provides an image classification machine learning system based on focus area and smooth annotation, which includes the following modules:
[0123] A data acquisition module, used to acquire a number of image data and generate an image data set according to the number of graphic data;
[0124] A division module, used for dividing each image in the image data set into one or more regions of interest, and setting one or more classification labels for each region of interest;
[0125] A labeling module, used to label each image in the image data set according to the focus area and the classification label to obtain a labeling data set, wherein the labeling includes a confidence level and an image classification;
[0126] The model training module is used to adopt a machine learning algorithm to train an image classification model based on the image data set and the labeled data set.
[0127] Specifically, the data acquisition module includes:
[0128] An enhancement unit, used for performing image enhancement on the plurality of image data to obtain a plurality of enhanced image data;
[0129] A size unification unit, used to unify the sizes of the plurality of enhanced image data by using an interpolation method to obtain a plurality of unified image data;
[0130] A normalization unit, used for performing normalization processing on the plurality of unified image data to obtain a plurality of normalized image data;
[0131] A generating unit is used to generate an image data set according to the normalized image data.
[0132] Specifically, the division module includes:
[0133] A feature extraction unit, used to extract features according to the image data set to obtain a number of basic features;
[0134] The label generating unit is used to generate a label according to a plurality of
[0135] Generate several classification labels based on the basic features;
[0136] A division unit, used for dividing each of the images into one or more regions of interest according to the image data set and a preset image division model;
[0137] The corresponding unit is used to set one or more labels for each focus area according to the plurality of classification labels.
[0138] Specifically, the annotation module includes:
[0139] A first screening unit, configured to screen out a plurality of images without overlapping the regions of interest according to one or more regions of interest in each of the images, as a plurality of first images;
[0140] an area calculation unit, used for calculating the areas of the plurality of first images and the area of each region of interest in each of the first images;
[0141] a relative area calculation unit, configured to calculate, in each of the first images, a relative area of each focus region in the first image according to an area of the first image and an area of each focus region in the first image;
[0142] A first weight calculation unit is used to add the relative areas of the plurality of focus regions corresponding to the same classification label in each of the first images to obtain a first weight of the same classification label, and traverse all classification labels to obtain a first weight of each classification label in the first image;
[0143] A first confidence calculation unit, used to calculate a first confidence and a first image classification according to a first weight of each classification label in the first image;
[0144] The first labeling unit is used to label the first image according to the first confidence level and the first image classification to obtain a labeling data set.
[0145] Specifically, the marking module also includes:
[0146] A second screening unit is used for screening out a plurality of images having overlapping regions of interest according to the plurality of regions of interest in each of the images as a plurality of second images;
[0147] a label screening unit, configured to use, according to the plurality of focus regions in each of the second images and one or more classification labels corresponding to each of the focus regions, the plurality of classification labels existing in each of the second images as the plurality of classification labels to be processed for each of the second images;
[0148] a label division unit, configured to select, in each of the second images, one or more of the to-be-processed classification labels as feature classifications and the rest as target classifications, wherein the focus area corresponding to the target classification is the target area and the focus area corresponding to the feature classification is the feature area;
[0149] A second weight calculation unit is used to filter out the number of overlaps between the target area and the feature area for each target classification in each of the second images as a second weight;
[0150] a second confidence calculation unit, configured to calculate, in each of the second images, a second confidence of the target classification according to the second weight and a preset correlation coefficient;
[0151] The second labeling unit is used to label the second image according to the target classification and the second confidence of the target classification to obtain a labeling data set.
[0152] Specifically, the training module includes:
[0153] The initial modeling unit is used to build the initial image classification model and set the loss function;
[0154] A prediction unit, used for inputting the image data set into an initial image classification model to generate a prediction value;
[0155] A residual calculation unit, used to calculate the difference between the labeled data set and the predicted value according to the loss function to obtain a numerical residual;
[0156] An updating unit is used to update the initial image classification model according to the numerical residual and the predicted value, repeatedly execute the prediction unit to the residual calculation unit until the loss function meets the preset cutoff condition, and output the initial image classification model at this time as the image classification model.
[0157] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only
[0158] For the convenience of mutual distinction, it is not used to limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the above method embodiment, which will not be repeated here.
[0159] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, any one of the above methods is implemented.
[0160] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. Of course, there are other ways of readable storage media, such as quantum memory, graphene memory, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electrical carrier signals and telecommunication signals.
[0161] The present invention also provides an electronic device. The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the image classification machine learning method based on focus area and smooth annotation provided by the present invention.
[0162] Reference below Figure 8 , which shows a schematic diagram of the structure of a computer system 800 of an electronic device suitable for implementing an embodiment of the present invention. Figure 8 The electronic device shown is only an example and should not be construed as to the functions and scope of use of the embodiments of the present invention.
[0163] Any restrictions.
[0164] like Figure 8 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0165] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom is installed into the storage section 808 as needed.
[0166] In particular, according to the embodiments disclosed in the present invention, the process described in the main step diagram above can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, the above functions defined in the system of the present invention are executed.
[0167] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: having one or more
[0168] The invention relates to an electrical connection of a conductor, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus or a device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus or a device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0169] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0170] The units involved in the embodiments of the present invention may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor including a pre-response
[0171] Unit, receiving unit and requesting unit. The names of these units do not constitute limitations on the units themselves under certain circumstances.
[0172] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A machine learning method for image classification based on region of interest and smooth annotation, characterized in that: The following steps are involved: S10: Acquire a plurality of image data, and generate an image data set according to the plurality of graphic data; S20: Divide each image in the image data set into one or more regions of interest, and set one or more classification labels for each region of interest; S30: annotating each image in the image data set according to the focus area and the classification label to obtain an annotated data set, wherein the annotation includes a confidence level and an image classification; S40: Using a machine learning algorithm, train an image classification model based on the image dataset and the annotated dataset.
2. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: The step S10 specifically includes: S11: performing image enhancement on the plurality of image data to obtain a plurality of enhanced image data; S12: unifying the sizes of the plurality of enhanced image data by using an interpolation method to obtain a plurality of unified image data; S13: performing normalization processing on the plurality of unified image data to obtain a plurality of normalized image data; S14: Generate an image data set according to the plurality of normalized image data.
3. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: The step S20 specifically includes: S21: extracting features according to the image data set to obtain a number of basic features; S22: generating a plurality of classification labels according to the plurality of basic features; S23: for dividing each of the images into one or more regions of interest according to the image data set and a preset image division model; S24: According to the plurality of classification labels, one or more labels are set for each region of interest.
4. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: The step S30 specifically includes: S31A: According to one or more regions of interest in each of the images, select a number of images without overlap of the regions of interest as a number of first images; S32A: Calculate the areas of the plurality of first images and the area of each focus region in each first image; S33A: In each of the first images, calculate the relative area of each focus region in the first image according to the area of the first image and the area of each focus region in the first image; S34A: In each of the first images, the relative areas of the plurality of focus regions corresponding to the same classification label are added to obtain a first weight of the same classification label, and all classification labels are traversed to obtain a first weight of each classification label in the first image; S35A: Calculate a first confidence level and a first image classification according to first weights of the classification labels in the first image; S36A: Annotate the first image according to the first confidence level and the first image classification to obtain an annotated data set.
5. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: The step S30 specifically includes: S31B: screening out, according to the plurality of interest regions in each of the images, a plurality of images having overlapping interest regions as a plurality of second images; S32B: according to the plurality of regions of interest in each of the second images and one or more classification labels corresponding to each of the regions of interest, using the plurality of classification labels existing in each of the second images as the plurality of classification labels to be processed of each of the second images; S33B: In each of the second images, one or more of the to-be-processed classification labels are selected as feature classifications, and the rest are selected as target classifications, the focus area corresponding to the target classification is the target area, and the focus area corresponding to the feature classification is the feature area; S34B: In each of the second images, for each target category, the number of overlaps between the target region and the feature region is screened out as a second weight; S35B: In each of the second images, calculating a second confidence level of the target classification according to the second weight and a preset correlation coefficient; S36B: Annotate the second image according to the target classification and the second confidence level of the target classification to obtain an annotated data set.
6. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: In step S40, it specifically includes: S41: Establish an initial image classification model and set the loss function; S42: inputting the image data set into an initial image classification model to generate a prediction value; S43: Calculate the difference between the labeled data set and the predicted value according to the loss function to obtain a numerical residual; S44: Update the initial image classification model according to the numerical residual and the predicted value, repeat steps S42 to S43 until the loss function meets the preset cutoff condition, and output the initial image classification model at this time as the image classification model.
7. The image classification machine learning method based on region of interest and smooth annotation according to claim 1, characterized in that: The method further includes step S50: acquiring image data to be classified, inputting the image data to be classified into the image classification model, and obtaining an image classification result.
8. Image classification machine learning system based on attention region and smooth annotation, characterized in that: include: A data acquisition module, used to acquire a number of image data and generate an image data set according to the number of graphic data; A division module, used for dividing each image in the image data set into one or more regions of interest, and setting one or more classification labels for each region of interest; A marking module is used to mark the area of interest according to the domain and the classification label, annotating each image in the image dataset to obtain an annotated dataset, wherein the annotation includes a confidence level and an image classification; The model training module is used to adopt a machine learning algorithm to train an image classification model based on the image data set and the labeled data set.
9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image classification machine learning method based on focus area and smooth annotation as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, an image classification machine learning method based on focus area and smooth annotation is implemented as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Object detection model training method and device
CN111709471A
Label smoothing method and device for target detection, equipment and storage medium
CN112418283A
Multi-label classification method and system based on deformable NTS-NET neural network
CN114821155A
Cervical cytology image abnormal region positioning method and device based on fusion attention
CN114897779A
Emotion recognition intelligent contract construction method based on cross fusion and confidence evaluation
CN118799948A