Image recognition model training method and device based on machine learning and medium

By extracting boundary description features through a semantic segmentation model and using a generative adversarial network to generate an extended sample image set, combined with a lightweight classifier to optimize the classification weights, the problems of existing technologies such as dependence on large-scale labeled data and insufficient dynamic adaptability are solved, achieving efficient and accurate image recognition.

CN120599341APending Publication Date: 2025-09-05天元大数据信用管理有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684647.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies rely on large-scale annotated data in image recognition, resulting in high manpower and time costs. In addition, generative adversarial networks lack constraints on the target semantic structure, affecting recognition accuracy. Lightweight classifiers are computationally inefficient on resource-constrained devices and have difficulty dynamically adapting to user-specified content.

Method used

The semantic segmentation model is used to extract the boundary description features of the specified content, and the generative adversarial network is used to generate an extended sample image set that retains the boundary features. The lightweight classifier is combined to transfer the classification weights and optimize the classifier to adapt to the user-specified content.

Benefits of technology

It reduces the dependence on large-scale labeled data, improves the model's dynamic adaptability and recognition accuracy to user-specified content, and enhances the model's generalization ability and classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599341A_ABST
    Figure CN120599341A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition model training method and device based on machine learning and a medium, and relates to the field of image recognition model training. The method comprises the following steps: receiving specified content description information input by a user, and extracting boundary description characteristics of the specified content description information by utilizing a semantic segmentation model; wherein the specified content description information comprises an example image with a specified identification content label and text description for specified identification content in the example image; injecting the boundary description features as constraint conditions into a preset generative adversarial network to generate an extended sample image set with the boundary description features reserved; inputting the example image and the extended sample image set into a preset image classification detection model to extract to-be-processed high-dimensional feature vectors of the extended sample image set and the example image; and according to the to-be-processed high-dimensional feature vector, performing classification weight parameter transmission to a preset lightweight classifier to obtain a to-be-applied classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition model training, and in particular to a method, device, and medium for image recognition model training based on machine learning. Background Art

[0002] In recent years, image recognition technology based on machine learning has been widely used in many fields. Traditional methods usually rely on large-scale annotated datasets to train deep models to achieve the recognition of specific targets. However, such methods have significant limitations: First, labeling massive amounts of data requires a lot of manpower and time costs, especially when users need to temporarily specify new recognition targets, it is difficult to quickly obtain sufficient annotated samples; second, the existing generative adversarial network (GAN) lacks constraints on the semantic structure of the target during sample expansion, resulting in a low degree of match between the boundary features of the generated image and the user-specified content, affecting the subsequent recognition accuracy. In addition, although the pre-trained model can extract common image features, its high-dimensional feature vector often contains redundant information. Direct use in lightweight classifiers can easily lead to overfitting or low computational efficiency, making it difficult to adapt to resource-constrained edge devices.

[0003] Existing research has attempted to train models using a small number of samples, such as those based on transfer learning or data augmentation techniques. However, these methods still have shortcomings when it comes to dynamically adapting to specific content: the diversity of samples generated by data augmentation is limited and cannot guarantee consistency with user-defined structural features; while transfer learning can reuse pre-trained model parameters, it does not focus on optimizing the core features of the specified content, making the model sensitive to subtle differences. Furthermore, classifier weight adjustment often relies on a fixed loss function and lacks an adaptive mechanism based on feature distribution, making it difficult to balance recognition accuracy and generalization ability.

[0004] Therefore, how to reduce dependence on large-scale labeled data and improve dynamic adaptability to user-specified content has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, and medium for training an image recognition model based on machine learning to solve the technical problems of how to reduce dependence on large-scale annotated data and improve dynamic adaptability to user-specified content.

[0006] In the first aspect, an embodiment of the present application provides an image recognition model training method based on machine learning, the method comprising: receiving specified content description information input by a user, and extracting boundary description features of the specified content description information using a semantic segmentation model; wherein the specified content description information includes an example image annotated with specified recognition content and a text description of the specified recognition content in the example image; injecting the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; inputting the example image and the extended sample image set into a preset image classification detection model to extract high-dimensional feature vectors to be processed of the extended sample image set and the example image; and transferring classification weights to a preset lightweight classifier based on the high-dimensional feature vectors to obtain the classifier to be applied.

[0007] In one implementation of the present application, the method also includes: constructing an image classification detection model, specifically including: obtaining sample image data and preprocessing the sample image data to obtain a standard sample image database; wherein the preprocessing includes normalization, resizing, and data enhancement; based on the standard sample image database, training a preset deep convolutional neural network to obtain a pretrained model; deleting the fully connected layer of the last layer of the pretrained model to obtain an image classification detection model.

[0008] In one implementation of the present application, a semantic segmentation model is used to extract boundary description features of specified content description information, specifically including: calling a pre-trained semantic segmentation model to perform pixel-level semantic segmentation processing on the specified content description information, identifying objects or areas corresponding to the annotated content and text description; extracting the boundary feature vector of the object or area, and converting the boundary feature vector into a boundary description feature based on the semantic conversion rules preset in the semantic segmentation model.

[0009] In one implementation of the present application, boundary description features are injected as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features, specifically including: converting the boundary description features into standard description format information that can be recognized by the generative adversarial network; performing feature fusion of the standard description format information with the input layer of the generator of the generative adversarial network to constrain the image features generated by the generator; based on the constraints, generative training is performed through the generative adversarial network, and after the training based on the constraints converges, an extended sample image set is generated based on a preset number of extended samples; wherein, based on the constraints, generative training is performed through the generative adversarial network, specifically including: calculating the difference between the image features generated by the generator and the boundary description features through a loss function, and adjusting the network parameters to minimize the difference until the training based on the constraints converges.

[0010] In one implementation of the present application, after generating an extended sample image set that retains boundary description features, the method further includes: performing image quality assessment on each image in the extended sample image set; wherein the image quality assessment includes image clarity, feature integrity, and semantic consistency; based on the quality assessment results, images that do not meet the quality standards are screened out, and the screened out images that do not meet the quality standards are fed back to the generative adversarial network to perform enhanced training on the generative adversarial network.

[0011] In one implementation of the present application, based on the high-dimensional feature vector to be processed, classification weights are transferred to a preset lightweight classifier to obtain the classifier to be applied, specifically including: dimensionality reduction processing of the high-dimensional feature vector to be processed, and using principal component analysis and / or linear discriminant analysis to map the high-dimensional feature vector to a low-dimensional feature space; inputting the reduced-dimensional high-dimensional feature vector into the lightweight classifier; wherein the reduced-dimensional high-dimensional feature vector is used by the classifier to be applied to determine content recognition requirements; grouping the reduced-dimensional feature vectors through a clustering algorithm, and dynamically adjusting the classification weights according to the grouping situation; wherein the adjustment of the classification weights is based on the cluster center and category distribution probability of the reduced-dimensional feature vectors; iteratively optimizing the parameters of the lightweight classifier to obtain the classifier to be applied.

[0012] In one implementation of the present application, after the classification weights are transferred to a preset lightweight classifier based on the high-dimensional feature vector to be processed to obtain the classifier to be applied, the method further includes: applying the classifier to be applied to the image to be identified to obtain a recognition result; wherein the recognition result includes whether the image to be identified contains specified recognition content and its corresponding confidence score; feeding back the recognition result to the user end, and receiving the user end's annotation feedback on the misidentified sample; and optimizing and updating the classifier to be applied based on the annotation feedback.

[0013] In one implementation of the present application, the classifier to be applied is optimized and updated based on the annotation feedback, specifically including: based on the annotation feedback of the misidentified samples received from the user end, preprocessing and analyzing the annotation feedback to determine the main error types and causes of the classifier to be applied in the recognition process; using the error types and causes, adjusting the model parameters and structure of the classifier, and retraining and / or fine-tuning the classifier to obtain an iteratively upgraded classifier; based on the iteratively upgraded classifier, using a new test set for testing, and feeding back the test results to the user end until the number of misidentified samples is minimized, thereby completing the optimization and update of the classifier to be applied.

[0014] In the second aspect, an embodiment of the present application also provides an image recognition model training device based on machine learning, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to: receive specified content description information input by a user, and extract boundary description features of the specified content description information using a semantic segmentation model; wherein the specified content description information includes an example image annotated with specified recognition content and a text description of the specified recognition content in the example image; the boundary description features are injected as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; the example image and the extended sample image set are input into a preset image classification detection model to extract the high-dimensional feature vectors to be processed of the extended sample image set and the example image; and according to the high-dimensional feature vectors to be processed, the classification weights are transferred to a preset lightweight classifier to obtain the classifier to be applied.

[0015] In a third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for training an image recognition model based on machine learning, storing computer executable instructions, and the computer executable instructions are configured to: receive specified content description information input by a user, and extract boundary description features of the specified content description information using a semantic segmentation model; wherein the specified content description information includes an example image annotated with specified recognition content and a text description of the specified recognition content in the example image; inject the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; input the example image and the extended sample image set into a preset image classification detection model to extract the high-dimensional feature vectors to be processed of the extended sample image set and the example image; and transfer classification weights to a preset lightweight classifier based on the high-dimensional feature vectors to obtain the classifier to be applied.

[0016] The embodiment of the present application provides a method, device and medium for training an image recognition model based on machine learning. By extracting boundary description features through a semantic segmentation model, the method can accurately capture the key features of the specified recognition content in the example image. On this basis, a generative adversarial network is used to generate an extended sample image set that retains boundary features, effectively expanding the diversity of the training data while ensuring that the newly added samples are consistent with the original example images in key features, thereby improving the quality of the training data. By inputting the example image and the extended sample image set into the image classification detection model, the feature vector to be processed is extracted, providing the classifier with more comprehensive and representative feature information. Based on these feature vectors, the classification weights of the lightweight classifier are transferred, so that the classifier to be applied can learn more accurate classification boundaries, thereby more accurately identifying and classifying the image content in practical applications, thereby enhancing the generalization ability and classification accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A flowchart of a method for training an image recognition model based on machine learning provided in an embodiment of the present application;

[0019] Figure 2 A schematic diagram of the internal structure of an image recognition model training device based on machine learning provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0021] The embodiments of the present application provide a method, device, and medium for training an image recognition model based on machine learning to solve the technical problems of how to reduce dependence on large-scale annotated data and improve dynamic adaptability to user-specified content.

[0022] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0023] Figure 1 This is a flow chart of a method for training an image recognition model based on machine learning provided in an embodiment of the present application. Figure 1 As shown, the embodiment of the present application provides an image recognition model training method based on machine learning, which specifically includes the following steps:

[0024] Step 10: Receive the specified content description information input by the user, and use the semantic segmentation model to extract the boundary description features of the specified content description information.

[0025] In this embodiment, the designated content description information includes an example image annotated with designated recognition content and a text description of the designated recognition content in the example image.

[0026] As an optional embodiment, using a semantic segmentation model to extract boundary description features of the specified content description information may specifically include: Step 101: calling a pre-trained semantic segmentation model to perform pixel-level semantic segmentation processing on the specified content description information, and identifying objects or areas corresponding to the annotated content and text description.

[0027] In this step, the semantic segmentation model is a deep learning model that is pre-trained using a large amount of labeled data and can classify each pixel in the input image into a specific semantic category. The example images included in the specified content description information are user-provided images that contain objects or areas that need to be identified. The specified identification content has been annotated on these images. The purpose of the annotation is to clearly indicate which parts of the image are objects or areas that the model needs to focus on and identify. The pre-trained semantic segmentation model assigns a category label to each pixel in the image to accurately delineate the boundaries of objects or areas.

[0028] Step 102: extracting the boundary feature vector of the object or region, and converting the boundary feature vector into a boundary description feature based on the semantic conversion rule preset in the semantic segmentation model.

[0029] In this step, the boundary feature vector is a quantitative representation of the boundary characteristics of an object or region. It is composed of multiple parameters that can reflect the boundary characteristics. These parameters are derived from the geometric properties of the boundary pixels, such as coordinates, direction, curvature, and other information, as well as information such as the relationship between the boundary and surrounding pixels. The boundary description feature is a feature obtained by further processing and conversion based on the boundary feature vector through the preset semantic conversion rules in the semantic segmentation model. It not only contains the basic geometric information of the boundary, but also integrates the semantic understanding of the object or region by the semantic segmentation model. It is a more semantically meaningful boundary representation form that can more intuitively reflect the boundary characteristics of the object or region at the semantic level.

[0030] Step 20: Inject the boundary description features as constraints into the preset generative adversarial network to generate an extended sample image set that retains the boundary description features.

[0031] As an optional embodiment, the boundary description features are injected as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features, specifically including: Step 201: converting the boundary description features into standard description format information that can be recognized by the generative adversarial network.

[0032] In this step, the Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and a discriminator. To enable the GAN to effectively recognize and process boundary descriptors, these features must be converted into a standard descriptive format that the GAN can understand and recognize. This conversion involves unifying the data format, adjusting the dimensions, and performing data normalization or standardization. These conversions enable the GAN to recognize and process boundary descriptors, laying the foundation for the GAN to generate realistic and compliant extended sample image sets.

[0033] Step 202: Feature fusion is performed on the standard description format information and the input layer of the generator of the generative adversarial network to constrain the image features generated by the generator.

[0034] In this step, the generator in the generative adversarial network is responsible for generating new image samples based on the specified content description information of the input; its input layer is the layer that receives the original input data and performs preliminary processing. The quality and feature richness of the input data directly affect the output effect of the generator; the input layer needs to be combined with specific feature information to generate images with specific features.

[0035] Step 203: Based on the constraints, generative training is performed through a generative adversarial network, and after the training based on the constraints converges, an extended sample image set is generated based on a preset number of extended samples.

[0036] As an optional embodiment, generative training is performed through a generative adversarial network based on constraints, specifically including: Step 2031: calculating the difference between the image features generated by the generator and the boundary description features through a loss function, and adjusting the network parameters to minimize the difference until the training based on the constraints converges.

[0037] In this step, the constraints play a key guiding role in the training process of the generative adversarial network (GAN). They provide the generator with a clear generation goal and direction, ensuring that the generated image meets specific requirements and characteristics. The generative adversarial network (GAN) consists of a generator and a discriminator. The two compete with each other and optimize together during the training process. The goal of the generator is to generate realistic images based on the specified content description information of the input, while the discriminator is responsible for distinguishing between the generated image and the real image. The difference between the image features generated by the generator and the boundary description features is calculated through the loss function. According to the value of the loss function, the network parameters (such as the weights and biases in the generator) will be adjusted to reduce the difference. Through continuous iteration and optimization, the generator gradually learns how to generate images that meet the boundary description features.

[0038] As an optional embodiment, after generating an extended sample image set that retains the boundary description features, the method further includes: step 204: performing image quality assessment on each image in the extended sample image set; wherein the image quality assessment includes image clarity, feature integrity, and semantic consistency.

[0039] In this step, image clarity reflects the degree of recognizability of details and textures in the image. When evaluating image clarity, we can observe whether the edges of objects in the image are sharp, whether the details are rich, and whether the texture is clear. In addition, we can quantify the clarity of the image by calculating indicators such as image contrast and noise intensity. Feature integrity refers to whether the key features of objects or regions in the image are fully presented, including appearance features such as the shape, color, and texture of the object, as well as spatial features such as the position and posture of the object in the image. When evaluating feature integrity, we need to check whether all parts of the object in the image are clearly visible, whether there are occlusions, missing or deformations, and whether the features of the object are consistent with the actual situation and whether they can accurately reflect the category and attributes of the object. Semantic consistency refers to the degree of match between image content and expected semantics. When evaluating semantic consistency, we can check whether the category labels of objects in the image are consistent with the actual content, whether the relative positions and relationships between objects are reasonable, and whether the scene expressed by the image is consistent with the expected application scenario. This helps the model learn correct semantic information and avoid misrecognition caused by semantic confusion.

[0040] Step 205: Screen out images that do not meet the quality standards based on the quality assessment results, and feed the screened out images that do not meet the quality standards back to the generative adversarial network to perform enhanced training on the generative adversarial network.

[0041] In this step, after performing image quality assessment on each image in the extended sample image set, the performance results of each image in various evaluation dimensions such as image clarity, feature integrity, and semantic consistency will be obtained. The evaluation results of each image will be judged according to the preset quality assessment standards or thresholds. If the image does not meet the requirements in one or more key quality indicators, it will be screened out and considered to be an image of substandard quality. The substandard image will be fed back to the generative adversarial network to improve and optimize the ability to generate image quality on the original basis.

[0042] Step 30: Input the example image and the extended sample image set into a preset image classification detection model to extract the high-dimensional feature vectors to be processed of the extended sample image set and the example image.

[0043] As an optional embodiment, the method also includes: Step 301: constructing an image classification detection model, specifically including: Step 3011: obtaining sample image data and preprocessing the sample image data to obtain a standard sample image database; wherein the preprocessing includes normalization, size adjustment, and data enhancement.

[0044] In this step, the image classification detection model aims to identify the category and location of objects in the image; during construction, sample image data must be prepared and processed to form a standard sample image database to provide high-quality input for model training; normalization is to map the pixel values ​​of image data to a specific range, unify the scale of image data, avoid excessive eigenvalues ​​affecting model training efficiency and stability, and accelerate training convergence; resizing is to uniformly adjust the image to the input size required by the model to ensure that the image can be correctly input into the model for processing; data enhancement expands the number of training samples and improves the model's generalization ability through operations such as rotation, flipping, translation, scaling, and cropping, so that the model maintains good performance in different scenarios.

[0045] Step 3012: Based on the standard sample image database, the preset deep convolutional neural network is trained to obtain a pre-trained model; Step 3013: The fully connected layer of the last layer of the pre-trained model is deleted to obtain an image classification detection model.

[0046] In this step, the deep convolutional neural network (DCNN) is a deep learning model specially designed for processing grid-structured data. It consists of multiple convolutional layers, pooling layers, activation functions, fully connected layers, etc. The last fully connected layer is generally used to map the high-dimensional features extracted by the previous layers to specific category labels and output category probability distribution; the purpose of deleting the last fully connected layer of the pre-trained model is to free the model from the special image classification task, so that it can be used as a general feature extractor to adapt to a wider range of image classification and detection tasks; at the same time, the pre-trained model has learned rich image feature representations. After deleting the last fully connected layer, only the new classification layer or detection layer needs to be trained in the new task. Compared with training the entire model from scratch, this method only requires training fewer parameters, the training process will converge faster, and it can more efficiently utilize computing resources and time costs, accelerating the model training and optimization process.

[0047] Step 40: Based on the high-dimensional feature vector to be processed, the classification weights are transferred to the preset lightweight classifier to obtain the classifier to be applied.

[0048] As an optional embodiment, based on the high-dimensional feature vector to be processed, the classification weights are transferred to a preset lightweight classifier to obtain the classifier to be applied, specifically including: Step 401: dimensionality reduction processing is performed on the high-dimensional feature vector to be processed, and principal component analysis and / or linear discriminant analysis are used to map the high-dimensional feature vector to a low-dimensional feature space.

[0049] In this step, a high-dimensional feature vector refers to a vector with a large number of dimensions (features), each dimension representing a specific feature in the image data. Although the high-dimensional feature vector contains rich information, it also brings a series of problems. On the one hand, the computational complexity of high-dimensional data is high, and the processing and storage costs are high. On the other hand, data in high-dimensional space is prone to sparsity, which makes model training difficult and prone to overfitting. In addition, there may be redundant information and noise in the high-dimensional feature vector, which will affect the performance and generalization ability of the model. Dimensionality reduction processing aims to map high-dimensional feature vectors to low-dimensional feature space while retaining the key information and feature structure in the original data as much as possible. Through dimensionality reduction, the complexity of the data can be reduced, the computational efficiency can be improved, the storage requirements can be reduced, the overfitting problem can be alleviated, and it is helpful to extract more representative and discriminative features. In image classification and detection tasks, the low-dimensional feature vectors after dimensionality reduction can be more efficiently used for classifier training and reasoning.

[0050] Step 402: Input the high-dimensional feature vector after dimensionality reduction into a lightweight classifier; wherein the high-dimensional feature vector after dimensionality reduction is used by the classifier to be applied to determine content recognition requirements.

[0051] In this step, the lightweight classifier is a classification model with a simple structure, low computational complexity, and a small number of parameters. It can quickly learn and judge based on the input feature vector. The high-dimensional feature vector after dimensionality reduction serves as the input of the lightweight classifier. Its core function is to help the classifier determine the content recognition requirements of the image. The lightweight classifier analyzes and processes the input feature vector, matches and compares the feature pattern in the feature vector with the feature pattern learned by the classifier during training, and finally determines whether the image contains specific content.

[0052] Step 403: Grouping the feature vectors after dimensionality reduction by a clustering algorithm, and dynamically adjusting the classification weights according to the grouping situation; wherein the adjustment of the classification weights is based on the cluster centers and category distribution probabilities of the feature vectors after dimensionality reduction.

[0053] In this step, the clustering algorithm groups similar samples into the same group based on the similarity or distance between data samples. Samples in different groups have lower similarity. By applying the clustering algorithm to group the feature vectors after dimensionality reduction, we can discover potential structures and patterns in the data, grouping data with similar characteristics into one category, and making the data organization more rational. During the classification process, the classification weights are dynamically adjusted based on the importance of different categories. The adjustment of classification weights is based on the cluster center and category distribution probability of the feature vectors after dimensionality reduction. The cluster center is the center point or representative point of the samples in each cluster group, which can reflect the main characteristics and central tendency of the data group. The category distribution probability describes the distribution of the probability of occurrence of each category in the data set. Combining the adjustment of classification weights with these two methods can make the classification model more rational in weight allocation.

[0054] Step 404: iteratively optimize the parameters of the lightweight classifier to obtain a classifier to be applied.

[0055] As an optional embodiment, after the classification weights are transferred to a preset lightweight classifier based on the high-dimensional feature vector to be processed to obtain the classifier to be applied, the method also includes: step 405: applying the classifier to be applied to the image to be identified to obtain a recognition result; wherein the recognition result includes whether the image to be identified contains specified identification content and its corresponding confidence score.

[0056] In this step, the classifier to be applied is obtained by the above method. When the classifier to be applied is applied to the image to be identified, the identification result is obtained. First, the classifier determines whether there is pre-set specific content that needs to be identified in the image. The confidence score is a quantitative evaluation of the classifier's judgment result. A higher confidence score indicates that the classifier is more certain about the judgment that the image contains the specified content, while a lower confidence score indicates that the classifier is less certain about the judgment result and there may be a certain degree of ambiguity or uncertainty.

[0057] Step 406: Feedback the recognition result to the user end, and receive the user end's annotation feedback on the misrecognized sample.

[0058] In this step, during the image recognition process, after completing the analysis and judgment of the image, the corresponding recognition results will be generated and fed back to the user end. After the user discovers the recognition error, these misrecognized image samples are marked, clearly indicating the correct recognition results, and then these marked feedback information is sent back to the image recognition model, which helps to further optimize the model.

[0059] Step 407: Optimize and update the classifier to be applied based on the annotation feedback.

[0060] As an optional embodiment, the classifier to be applied is optimized and updated based on the annotation feedback, specifically including: Step 4071: Based on the annotation feedback received from the user end on the misrecognized sample, the annotation feedback is preprocessed and analyzed to determine the main error types and causes of the classifier to be applied in the recognition process.

[0061] In this step, the features extracted by the classifier are compared with the actual features of the misidentified samples to find out the deviations or inaccuracies in the feature extraction process, and the decision path and basis of the classifier when dealing with misidentified samples are studied in depth to identify logical loopholes or irrationalities in the decision-making process. Furthermore, the training data is checked to see if there are similar situations to the misidentified samples to determine whether the training data is comprehensive and representative. If the training data lacks positive or negative examples similar to the misidentified samples, it may cause the classifier to make errors in actual recognition.

[0062] Step 4072: Using the error type and cause, adjust the model parameters and structure of the classifier, and retrain and / or fine-tune the classifier to obtain an iteratively upgraded classifier.

[0063] In this step, by analyzing the error type and cause, we can determine which parameters need to be adjusted and the direction and magnitude of the adjustment. In some cases, simply adjusting the parameters may not effectively solve the problem of the classifier, and the structure of the model needs to be modified, such as adding or reducing certain layers, nodes or connections in the model to enable it to better capture the features and patterns in the data.

[0064] Step 4073: Based on the iteratively upgraded classifier, a new test set is used for testing, and the test results are fed back to the user end until the number of misidentified samples is minimized, thereby completing the optimization update of the classifier to be applied.

[0065] In this step, after adjusting the model parameters and structure of the classifier and going through an iterative upgrade process of retraining or fine-tuning, the classifier is tested using a new test set to verify the actual performance of the optimized classifier on untrained data, and the test results are fed back to the user end. Through continuous iterative optimization and test feedback, the performance of the classifier is gradually improved, and ultimately the number of misidentified samples is minimized, which means that the recognition accuracy of the classifier is constantly improving.

[0066] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides an image recognition model training device based on machine learning, whose structure is as follows Figure 2 shown.

[0067] Figure 2 This is a schematic diagram of the internal structure of an image recognition model training device based on machine learning provided in an embodiment of the present application. Figure 2 As shown, the equipment includes:

[0068] at least one processor 201;

[0069] and, a memory 202 communicatively coupled to the at least one processor;

[0070] In which, the memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 201 so that the at least one processor 201 can: receive specified content description information input by a user, and use a semantic segmentation model to extract boundary description features of the specified content description information; wherein the specified content description information includes an example image annotated with specified recognition content and a text description of the specified recognition content in the example image; inject the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; input the example image and the extended sample image set into a preset image classification detection model to extract the high-dimensional feature vectors to be processed of the extended sample image set and the example image; and transfer classification weights to a preset lightweight classifier based on the high-dimensional feature vectors to obtain the classifier to be applied.

[0071] Some embodiments of the present application provide corresponding Figure 1A non-volatile computer storage medium for training an image recognition model based on machine learning stores computer-executable instructions, wherein the computer-executable instructions are configured to: receive specified content description information input by a user, and extract boundary description features of the specified content description information using a semantic segmentation model; wherein the specified content description information includes a sample image annotated with specified recognition content and a text description of the specified recognition content in the sample image; inject the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; input the sample image and the extended sample image set into a preset image classification detection model to extract high-dimensional feature vectors to be processed of the extended sample image set and the sample image; and transfer classification weights to a preset lightweight classifier based on the high-dimensional feature vectors to obtain a classifier to be applied.

[0072] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the IoT device and media embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.

[0073] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0074] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0075] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0076] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0078] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0079] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0080] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0081] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0082] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for training an image recognition model based on machine learning, characterized in that: The method comprises: Receive designated content description information input by a user, and extract boundary description features of the designated content description information using a semantic segmentation model; wherein the designated content description information includes an example image annotated with designated identification content and a text description of the designated identification content in the example image; Injecting the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features; Inputting the example image and the extended sample image set into a preset image classification detection model to extract high-dimensional feature vectors to be processed of the extended sample image set and the example image; According to the high-dimensional feature vector to be processed, classification weights are transferred to a preset lightweight classifier to obtain a classifier to be applied.

2. The image recognition model training method based on machine learning according to claim 1, characterized in that: The method further comprises: Build an image classification detection model, including: Acquire sample image data and preprocess the sample image data to obtain a standard sample image database; wherein the preprocessing includes normalization, size adjustment, and data enhancement; Based on the standard sample image database, a preset deep convolutional neural network is trained to obtain a pre-trained model; The fully connected layer of the last layer of the pre-trained model is deleted to obtain an image classification detection model.

3. The image recognition model training method based on machine learning according to claim 1, characterized in that: Extracting boundary description features of the specified content description information using a semantic segmentation model specifically includes: Calling a pre-trained semantic segmentation model to perform pixel-level semantic segmentation processing on the specified content description information to identify objects or regions corresponding to the annotated content and text description; Boundary feature vectors of the object or region are extracted, and based on semantic conversion rules preset in the semantic segmentation model, the boundary feature vectors are converted into boundary description features.

4. The image recognition model training method based on machine learning according to claim 1, characterized in that: Injecting the boundary description features as constraints into a preset generative adversarial network to generate an extended sample image set that retains the boundary description features, specifically including: Converting the boundary description features into standard description format information that can be recognized by a generative adversarial network; Performing feature fusion on the standard description format information and the input layer of the generator of the generative adversarial network to constrain the image features generated by the generator; Based on the constraints, generative training is performed by the generative adversarial network, and after the training based on the constraints converges, the extended sample image set is generated based on a preset number of extended samples; Wherein, based on the constraint conditions, generative training is performed by the generative adversarial network, specifically including: The difference between the image features generated by the generator and the boundary description features is calculated through a loss function, and the network parameters are adjusted to minimize the difference until the training converges based on the constraint conditions.

5. The image recognition model training method based on machine learning according to claim 1, characterized in that: After generating the extended sample image set retaining the boundary description feature, the method further includes: Performing image quality assessment on each image in the extended sample image set; wherein the image quality assessment includes image clarity, feature integrity, and semantic consistency; Based on the quality assessment result, images that do not meet the quality standards are screened out, and the screened out images that do not meet the quality standards are fed back to the generative adversarial network to perform enhanced training on the generative adversarial network.

6. The image recognition model training method based on machine learning according to claim 1, characterized in that: According to the high-dimensional feature vector to be processed, the classification weight is transferred to a preset lightweight classifier to obtain a classifier to be applied, specifically including: Performing dimensionality reduction processing on the high-dimensional feature vector to be processed, and mapping the high-dimensional feature vector to a low-dimensional feature space using principal component analysis and / or linear discriminant analysis; Inputting the high-dimensional feature vector after dimensionality reduction into the lightweight classifier; wherein the high-dimensional feature vector after dimensionality reduction is used by the classifier to be applied to determine content recognition requirements; Grouping the reduced-dimensional feature vectors using a clustering algorithm and dynamically adjusting classification weights based on the grouping; wherein the adjustment of the classification weights is based on the cluster centers and category distribution probabilities of the reduced-dimensional feature vectors; The parameters of the lightweight classifier are iteratively optimized to obtain the classifier to be applied.

7. The image recognition model training method based on machine learning according to claim 1, characterized in that: After transferring classification weights to a preset lightweight classifier based on the high-dimensional feature vector to be processed to obtain a classifier to be applied, the method further includes: Applying the classifier to be applied to the image to be identified to obtain a recognition result; wherein the recognition result includes whether the image to be identified contains the specified identification content and its corresponding confidence score; Feedback the recognition result to the user terminal, and receive annotation feedback from the user terminal on the misrecognized samples; The classifier to be applied is optimized and updated according to the annotation feedback.

8. The image recognition model training method based on machine learning according to claim 7, characterized in that: Optimizing and updating the classifier to be applied according to the annotation feedback specifically includes: Based on the received user terminal's annotation feedback on the misidentified samples, preprocessing and analyzing the annotation feedback to determine the main error types and causes in the recognition process of the classifier to be applied; Using the error type and cause, adjusting the model parameters and structure of the classifier, and retraining and / or fine-tuning the classifier to obtain an iteratively upgraded classifier; Based on the iteratively upgraded classifier, a new test set is used for testing, and the test results are fed back to the user end until the number of misidentified samples is minimized, thereby completing the optimization update of the classifier to be applied.

9. A machine learning-based image recognition model training device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image recognition model training method based on machine learning as described in any one of claims 1-8.

10. A non-volatile computer storage medium for training an image recognition model based on machine learning, storing computer-executable instructions, characterized in that: When the computer-executable instructions are executed, an image recognition model training method based on machine learning as described in any one of claims 1 to 8 is implemented.