Target classification method and device and electronic equipment
By using pre-trained target classification model to extract, screen and combine textile fabric images, the problem of low classification efficiency of textile fabric images in the prior art is solved, and high-precision and high-efficiency classification results are achieved.
Patent Information
- Application Number
- CN202510103023.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing textile fabric image classification method has complex calculation process and high time cost, resulting in low classification efficiency.
The pre-trained target classification model is adopted to extract features, filter, feature combination and recognition of the classified images, and remove noise and redundant information through feature screening and feature combination to improve classification accuracy and efficiency.
High-precision textile fabric image classification is achieved, complex calculation processes are avoided, time costs are reduced, and classification efficiency is improved.
Smart Images

Figure CN120014352A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, for example, to a target classification method, device and electronic device. Background Art
[0002] Textile fabrics are essential basic materials for many fields such as clothing, home decoration, and industrial applications. They are of various types and have ever-changing fabric styles and patterns. In order to efficiently manage these various and large quantities of fabrics, textile fabric factories usually classify them according to the patterns of the fabrics.
[0003] For example, the related art discloses a method for classifying textile fabric images based on enhanced depth features, including: obtaining a textile fabric image; using a two-dimensional discrete Fourier transform formula to convert the textile fabric image from the spatial domain to the frequency domain to obtain a first frequency domain image of the textile fabric; constructing a corresponding band-stop filter based on the textile fabric image, and performing band-stop filtering on the first frequency domain image according to the band-stop filter to obtain a second frequency domain image with the background removed; graying the second frequency domain image after the background is removed by a grayscale conversion algorithm to obtain a grayscale frequency domain image, extracting features from the grayscale frequency domain image to obtain image texture features, and extracting RGB histogram statistical features of the second frequency domain image in the RGB color space; inputting the second frequency domain image after the background is removed into a deep feature extractor of a pre-constructed convolutional neural network, and performing multi-layer encoding on the second frequency domain image after the background is removed using a convolutional neural network with a residual structure to extract image depth features of the second frequency domain image; performing feature fusion on the image texture features, RGB histogram statistical features, and image depth features to obtain enhanced depth features, and inputting the enhanced depth features into an image classification model for image classification to determine the classification category of the textile fabric image.
[0004] Although the relevant technology has improved the accuracy of textile fabric classification, the calculation process is complicated and the time cost is high, resulting in low classification efficiency. Summary of the invention
[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0006] The embodiments of the present disclosure provide a method, device and electronic device for target classification, which avoid complicated calculation processes, reduce time costs and improve classification efficiency.
[0007] In some embodiments, a target classification method is provided, including: acquiring an image to be classified; inputting the image to be classified into a pre-trained target classification model; performing feature extraction on the image to be classified to obtain initial features; performing feature screening on the initial features to obtain target features; performing feature combination on the target features to obtain a target feature group; and identifying the target feature group to obtain a classification result.
[0008] Optionally, the target classification model includes a feature extraction model; performing feature extraction on the image to be classified to obtain initial features includes: inputting the image to be classified into the feature extraction model to perform feature extraction to obtain initial features.
[0009] Optionally, the target classification model includes a feature screening model; performing feature screening on the initial features to obtain the target features includes: inputting the initial features into the feature screening model to perform feature screening to obtain the target features.
[0010] Optionally, the target classification model includes a feature combination model; performing feature combination on the target features to obtain a target feature group includes: inputting the target features into the feature combination model to perform feature combination to obtain the target feature group.
[0011] Optionally, the target classification model includes an image classification model; identifying the target feature group to obtain a classification result includes: inputting the target feature group into the image classification model for identification to obtain a classification result.
[0012] Optionally, the feature extraction model includes a convolutional neural network model; and / or, the feature screening model includes a support vector machine; and / or, the feature combination model includes a random forest; and / or, the image classification model includes a support vector machine or a random forest.
[0013] Optionally, a pre-trained target classification model is obtained in the following manner: obtaining an initial training image set; performing image expansion on the initial training image set to obtain a training image set; and training the target classification model using the training image set to obtain a trained target classification model.
[0014] Optionally, the initial training image set is expanded to obtain a training image set, including: performing image transformation on the training images in the initial training image set to obtain expanded images; the image transformation includes one or more of image rotation, image cropping, dithering, and enhancement filtering; and obtaining the training image set based on the initial training image set and the expanded images.
[0015] Optionally, the initial training image set is expanded to obtain a training image set, including: obtaining training images in the initial training image set and description texts corresponding to the training images; inputting the training images and the description texts corresponding to the training images into a comparative language-image pre-training model to obtain expanded images; and obtaining a training image set based on the initial training image set and the expanded images.
[0016] Optionally, the initial training image set is expanded to obtain a training image set, including: inputting the training images in the initial training image set into an image-to-text model to obtain description texts corresponding to the training images; performing image retrieval based on the description texts to obtain expanded images; and obtaining a training image set based on the initial training image set and the expanded images.
[0017] Optionally, the initial training image set is expanded to obtain a training image set, including: obtaining text labels corresponding to the training images in the initial training image set; inputting the text labels into the trained enhanced stable diffusion model to obtain expanded images; and obtaining the training image set based on the initial training image set and the expanded images.
[0018] Optionally, the image to be classified is input into a pre-trained target classification model to obtain a classification result, including: inputting the image to be classified into the pre-trained target classification model to obtain a preliminary result output by the target classification model; performing uncertainty evaluation on the preliminary result to obtain an uncertainty value; when the uncertainty value is less than or equal to an uncertainty threshold, using the preliminary result as the classification result; when the uncertainty value is greater than the uncertainty threshold, updating the target classification model; and inputting the image to be classified into the updated target classification model to obtain the preliminary result again.
[0019] Optionally, updating the target classification model includes: sending the image to be classified and the classification result to the target terminal; receiving classification information fed back by the target terminal based on the image to be classified and the classification result; the classification information includes whether the classification result is correct or the classification result is wrong; when the classification information indicates that the classification result is correct, training the target classification model according to the image to be classified and the classification result to update the target classification model; when the classification information indicates that the classification result is wrong, obtaining the type label corresponding to the image to be classified; training the target classification model according to the image to be classified and the type label to update the target classification model.
[0020] Optionally, updating the target classification model includes: obtaining a training image set for the target classification model; classifying the training image set according to type labels of the training images in the training image set to obtain multiple sub-training image sets; training multiple target classification models respectively using the sub-training image sets to obtain multiple sub-target classification models; and fusing the multiple sub-target classification models to update the target classification model.
[0021] In some embodiments, a target classification device is provided, including: an acquisition module, configured to obtain an image to be classified; a classification module, configured to input the image to be classified into a pre-trained target classification model: perform feature extraction on the image to be classified to obtain initial features; perform feature screening on the initial features to obtain target features; perform feature combination on the target features to obtain a target feature group; and identify the target feature group to obtain a classification result.
[0022] In some embodiments, a target classification device is provided, including a processor and a memory storing program instructions, wherein the processor is configured to execute the target classification method as described in the above embodiments when running the program instructions.
[0023] In some embodiments, an electronic device is provided, comprising: a device body; and a target classification device as described in the above embodiments, installed on the device body.
[0024] The target classification method, device, and electronic device provided by the embodiments of the present disclosure can achieve the following technical effects:
[0025] In the disclosed embodiment, the target classification model is obtained by training a large amount of labeled image data in advance, and can learn the complex features of target objects of different categories, thereby achieving high-precision classification. By inputting the image to be classified into the pre-trained target classification model for recognition, the classification result can be quickly obtained. Compared with the related art, the complex calculation process is avoided, the time cost is reduced, and the classification efficiency is higher. In addition, in the disclosed embodiment, through feature extraction, feature screening and feature combination, the target classification model can learn more accurate and discriminative feature representations, thereby improving the accuracy of classification. Among them, the feature screening and feature combination steps help to remove noise and redundant information, further improve the classification efficiency, and enhance the robustness, generalization ability and classification efficiency of the target classification model.
[0026] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:
[0028] Figure 1 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0029] Figure 2 is a schematic diagram of a target classification method provided by an embodiment of the present disclosure;
[0030] Figure 3 is a schematic diagram of a target classification model provided by an embodiment of the present disclosure;
[0031] Figure 4 is a schematic diagram of a method for training a target classification model provided by an embodiment of the present disclosure;
[0032] Figure 5 is a schematic diagram of a target classification method provided by another embodiment of the present disclosure;
[0033] Figure 6 is a schematic diagram of a target classification device provided by an embodiment of the present disclosure;
[0034] Figure 7 It is a schematic diagram of a target classification device provided by another embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0036] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.
[0037] Unless otherwise stated, the term "plurality" means two or more.
[0038] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.
[0039] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A, B, A and B.
[0040] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0041] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0042] Combination Figure 1 As shown, the embodiment of the present disclosure provides an electronic device 1, comprising a device body 10 and a target classification device 60 (70). The target classification device 60 (70) is installed on the device body 10.
[0043] In the electronic device 1 provided in the embodiment of the present disclosure, the target classification device 60 (70) is installed in the device body 10. The installation relationship described here is not limited to being placed inside the device body 10, but also includes installation connections with other components of the electronic device 1, including but not limited to physical connections, electrical connections or signal transmission connections, etc. It can be understood by those skilled in the art that the target classification device 60 (70) can be adapted to a feasible device body 10, thereby realizing other feasible embodiments.
[0044] Optionally, the target classification device 70 includes a processor 700. The processor 700 can obtain an image to be classified; can input the image to be classified into a pre-trained target classification model: perform feature extraction on the image to be classified to obtain initial features; perform feature screening on the initial features to obtain target features; perform feature combination on the target features to obtain a target feature group; and identify the target feature group to obtain a classification result.
[0045] Combination Figure 1 and Figure 7 The electronic device 1 shown in the present disclosure provides a target classification method, such as Figure 2 As shown, the target classification methods include:
[0046] S201: The processor obtains an image to be classified.
[0047] The image to be classified refers to the image containing the target object that needs to be classified.
[0048] In the step, an image of a target object to be classified, i.e., an image to be classified, can be obtained through an image acquisition device, such as a camera. In the process of applying the target classification method to the classification of textile fabrics, the target object is the textile fabric to be classified, and the image to be classified is an image containing the textile fabric to be classified.
[0049] S202, the processor inputs the image to be classified into a pre-trained target classification model: performs feature extraction on the image to be classified to obtain initial features; performs feature screening on the initial features to obtain target features; performs feature combination on the target features to obtain a target feature group; and identifies the target feature group to obtain a classification result.
[0050] In the target classification method provided by the embodiment of the present disclosure, the target classification model is obtained by training a large amount of labeled image data in advance, and can learn the complex features of target objects of different categories, thereby achieving high-precision classification. By inputting the image to be classified into the pre-trained target classification model for recognition, the classification result can be quickly obtained. Compared with the related art, the complex calculation process is avoided, the time cost is reduced, and the classification efficiency is higher.
[0051] In the disclosed embodiment, the initial features are obtained by first performing feature extraction on the image to be classified. The obtained initial features are then further subjected to feature screening to obtain the target features. Furthermore, the obtained target features are subjected to feature combination to obtain the target feature group. Finally, the obtained target feature group is identified, and the category to which the target object in the image to be classified belongs is output to obtain the classification result. In this embodiment, through feature extraction, feature screening and feature combination, the target classification model can learn a more accurate and discriminative feature representation, thereby improving the accuracy of classification. Among them, the feature screening and feature combination steps help to remove noise and redundant information, further improve the classification efficiency, and enhance the robustness, generalization ability and classification efficiency of the target classification model.
[0052] In some embodiments, Figure 3 As shown, the target classification model includes a feature extraction model. Extracting features from the image to be classified to obtain initial features includes: inputting the image to be classified into the feature extraction model to extract features to obtain initial features.
[0053] In this embodiment, the feature extraction model is used to extract feature data related to the target object from the input image to be classified. These feature data can be color, shape, texture, etc., so as to help the target classification model better understand and identify the target object in the image to be classified. Through feature extraction, the target classification model can capture the key information in the image to be classified and provide strong support for subsequent classification tasks. In addition, the feature extraction model can also reduce the input data, that is, the dimension of the image to be classified, and reduce the computational complexity.
[0054] Optionally, the feature extraction model includes a convolutional neural network model.
[0055] The convolutional neural network model in the embodiment of the present disclosure is a model built based on the convolutional neural network (CNN) algorithm. CNN can automatically and hierarchically extract features in the image, such as edges, corners, textures, etc. through a series of learning layers. The convolutional neural network model in the embodiment of the present disclosure includes an input layer, a convolution layer, and a pooling layer. Among them, the input layer is used to receive the original image data (in this embodiment, the image to be classified). The convolution layer uses multiple convolution kernels (also called filters) to perform a sliding window operation on the input data, that is, the image to be classified, and each convolution kernel is responsible for extracting a specific feature. After the convolution operation, a nonlinear activation function (such as ReLU) is applied to increase the nonlinear expression ability of the target classification model. The pooling layer downsamples the output of the convolution layer to reduce the dimension and calculation amount of the input data while retaining the most important feature information.
[0056] In this embodiment, a convolutional neural network model is used to extract features of the image to be classified, so as to fully utilize the ability of the convolutional neural network algorithm to automatically learn high-level features and obtain rich feature data.
[0057] In some embodiments, Figure 3 As shown, the target classification model includes a feature screening model. Performing feature screening on the initial features to obtain the target features includes: inputting the initial features into the feature screening model to perform feature screening to obtain the target features.
[0058] In this embodiment, the feature screening model is used to screen the initial features of the input to remove redundant features and features that contribute less to the classification task, and retain the most representative features. Feature screening can reduce the interference of noise and redundant information, and further improve the classification performance of the target classification model. At the same time, the feature set after screening is more concise, which is conducive to the training and reasoning of the target classification model, and further improves the classification efficiency of the target classification model.
[0059] Optionally, the feature screening model includes a support vector machine.
[0060] Support Vector Machine (SVM) is a supervised learning method. The basic idea of SVM is to find an optimal segmentation hyperplane to maximize the interval between data points, so that data points of different categories can be separated to the greatest extent. This optimal segmentation hyperplane is called a decision boundary, which divides the samples in the data set into two categories, one on one side of the decision boundary and the other on the other side.
[0061] In this embodiment, the feature screening model will assign weights to the input feature data in the process of finding the segmentation hyperplane. The weights of the features that contribute less to the determination of the segmentation hyperplane will be relatively small, and the weights of the features that contribute more to the determination of the segmentation hyperplane will be relatively large. In this embodiment, a weight threshold is pre-set, and by comparing the weight corresponding to each feature with the weight threshold, the features with a smaller weight, that is, less than the weight threshold, are removed to retain the most discriminative, that is, representative features, to achieve feature screening.
[0062] In some embodiments, Figure 3 As shown, the target classification model includes a feature combination model. Combining the target features to obtain a target feature group includes: inputting the target features into the feature combination model to perform feature combination to obtain the target feature group.
[0063] In this embodiment, the feature combination model is used to combine the input target features to form a more complex and rich feature representation. According to the specific design of the feature combination model, the combined feature representation can be linear or nonlinear. The combined feature representation is more comprehensive and in-depth. Through feature combination, the target classification model can capture the association and interaction between different features, thereby improving the classification accuracy of the target classification model.
[0064] Optionally, the feature combination model includes a random forest.
[0065] Random Forest is an integrated learning method that improves the accuracy and robustness of the model by constructing multiple decision trees and summarizing their prediction results. In this embodiment, the screened features are used as the original data set of the random forest, and then multiple training subsets are constructed by random sampling with replacement (called bootstrap sampling) from the original data set, and each training subset is used to generate a decision tree. In the process of generating a decision tree, the average value of the reduction in impurity caused by each feature when the node is split in all decision trees is calculated to evaluate the importance of the feature, and the impurity can be expressed by the Gini coefficient or entropy. The higher the importance score of the feature, the greater the contribution of the feature to reducing the impurity of the model, that is, the greater the contribution of the feature to the classification result. Random Forest can select important features for combination according to the importance score of the features in each decision tree, realize feature group output, and improve the accuracy and efficiency of classification while removing unimportant features.
[0066] In some embodiments, Figure 3 As shown, the target classification model includes an image classification model. Identifying the target feature group to obtain a classification result includes: inputting the target feature group into the image classification model for identification to obtain a classification result.
[0067] In this embodiment, the image classification model is used to classify the input image to be classified according to the input target feature group. The image classification model can directly output the classification result and provide intuitive and accurate classification information.
[0068] Optionally, the image classification model includes a support vector machine or a random forest.
[0069] Both support vector machines and random forests can be applied to classification tasks. In this embodiment, an image classification model is constructed based on a support vector machine or a random forest to classify the target feature group as input to achieve classification of the image to be classified.
[0070] In some embodiments, Figure 3 As shown, the target classification model includes a feature extraction model, a feature screening model, a feature combination model and an image classification model which are connected in sequence.
[0071] The target classification model in the embodiment of the present disclosure includes a feature extraction model, a feature screening model, a feature combination model and an image classification model connected in sequence, and is composed of multiple sub-models, each of which performs its corresponding function. In the process of adjusting or optimizing the target classification model, the classification performance of the target classification model can be improved by retraining or fine-tuning the sub-models in the target classification model, so that the target classification model can adapt to new application scenarios. Therefore, the application and deployment efficiency of the target classification model in the embodiment of the present disclosure is higher, which further improves the classification efficiency in practical applications.
[0072] The target classification model provided by the embodiment of the present disclosure integrates the feature extraction model, feature screening model, feature combination model and image classification model in sequence, realizes the combination of steps such as feature extraction, feature screening, feature combination and image classification, and forms an efficient and accurate target classification model, so that the target classification model can not only further reduce the computational complexity and improve the classification efficiency through feature extraction and feature screening, but also improve the classification accuracy and generalization ability of the target classification model through feature combination.
[0073] In this embodiment, the image to be classified is input into a pre-trained target classification model, including: inputting the image to be classified into a feature extraction model for feature extraction to obtain initial features; inputting the initial features into a feature screening model for feature screening to obtain target features; inputting the target features into a feature combination model for feature combination to obtain a target feature group; and inputting the target feature group into an image classification model for recognition to obtain a classification result.
[0074] In this embodiment, the image to be classified is first input into the feature extraction model for feature extraction, and the features output by the feature extraction model are used as initial features to achieve the acquisition of initial features. The obtained initial features are then further input into the feature screening model for feature screening, and the features output by the feature screening model are used as target features to achieve the acquisition of target features. Furthermore, the obtained target features are input into the feature combination model for feature combination, and the feature group output by the feature combination model is used as the target feature group to achieve the acquisition of the target feature group. Finally, the obtained target feature group is input into the image classification model for recognition, thereby outputting the category to which the target object in the image to be classified belongs, and obtaining the classification result.
[0075] Optionally, before inputting the image to be classified into a pre-trained target classification model, the method further includes: preprocessing the image to be classified; and inputting the pre-processed image to be classified into the pre-trained target classification model to obtain a classification result.
[0076] In this embodiment, preprocessing may include but is not limited to operations such as scaling, cropping, and normalization of the image. Preprocessing is performed to ensure that the image input into the target classification model is consistent with the image used in training the target classification model in terms of size, resolution, and pixel value, thereby ensuring that the quality of the image data meets the input requirements of the target classification model and reducing classification errors caused by poor image quality.
[0077] In some embodiments, a pre-trained target classification model is obtained in the following manner. Specifically, Figure 4 As shown, the present disclosure provides a method for training a target classification model, including:
[0078] S401: The processor obtains an initial training image set.
[0079] The initial training image set is a data set containing images of target objects of different types.
[0080] In this step, an initial training image set can be constructed based on commercial images publicly available on the Internet.
[0081] S402: The processor performs image expansion on the initial training image set to obtain a training image set.
[0082] S403: The processor trains the target classification model using the training image set to obtain a trained target classification model.
[0083] The target classification model training method provided by the embodiment of the present disclosure can, before training the target classification model, first image the initial training image set to increase the diversity and scale of the data set, and obtain a richer and more diverse training image set, so as to improve the generalization ability and robustness of the target classification model and improve the classification accuracy of the target classification model.
[0084] Optionally, the initial training image set is expanded to obtain a training image set, including: performing image transformation on the training images in the initial training image set to obtain expanded images; the image transformation includes one or more of image rotation, image cropping, dithering, and enhancement filtering; and obtaining the training image set based on the initial training image set and the expanded images.
[0085] In this embodiment, one or more image transformation operations are performed on each training image in the initial training image set to obtain an expanded image. Specifically, the image transformation operations include but are not limited to one or more of image rotation, image cropping, dithering, and enhancement filtering.
[0086] Among them, image rotation: use rotation algorithms (such as affine transformation, bilinear interpolation, etc.) to rotate the training images. By changing the direction of the training images, different angles of observation are simulated to increase the diversity of the data set. Data enhancement libraries such as OpenCV (Open Source Computer Vision Library) and PIL (Python Imaging Library, also known as Pillow) can be used to easily implement image rotation operations.
[0087] Image cropping: extract a sub-region from the original image (the training image in this embodiment) to generate more samples. Cropping can be random or based on specific strategies, such as center cropping or edge cropping. Image cropping can not only simulate changes in camera viewing angles, but also help the target classification model focus on different image areas and improve its ability to understand local features.
[0088] Jitter: refers to the small random changes in the pixel values of the original image, usually refers to adding noise, such as salt and pepper noise: randomly setting some pixel values to the maximum or minimum value to simulate the point noise in the image, Gaussian noise: adding random noise that follows a Gaussian distribution to the image pixel values. By adding noise to the original image to simulate image noise in the real world, the robustness of the target classification model in a noisy environment can be improved.
[0089] Enhanced filtering: By performing convolution operations on the original image, blurring, sharpening, edge detection and other effects are achieved, and certain features of the original image are enhanced or suppressed. For example, Gaussian blurring and other operations are performed through blur filters to reduce image details and simulate low-quality images; edge information of the original image is processed through sharpening filters to improve the details and sharpness of the original image. Enhanced filtering can help the target classification model better understand the detailed features of the image.
[0090] In this embodiment, new images different from the initial training image set are generated through image transformation, and these new images are called expanded images to increase the diversity and scale of the training data set. The initial training image set and all generated expanded images are merged to form a final training image set for training the target classification model. Image expansion increases the diversity and scale of the initial training images, helps the target classification model learn more generalized feature representations, reduces the risk of overfitting the target classification model to a specific data set, and improves the performance of the target classification model on new data. By training the target classification model with a training image set containing a variety of transformed images, the target classification model can better cope with various image transformations and noise interference, enhance the robustness of the target classification model in practical applications, and improve the accuracy of the target classification model in classification tasks, so that it can more accurately identify target objects under different conditions and of different types.
[0091] Optionally, the initial training image set is expanded to obtain a training image set, including: obtaining training images in the initial training image set and description texts corresponding to the training images; inputting the training images and the description texts corresponding to the training images into a comparative language-image pre-training model to obtain expanded images; and obtaining a training image set based on the initial training image set and the expanded images.
[0092] The Contrastive Language–Image Pre-training (CLIP) model is a multimodal vision and text learning framework that can learn the joint representation of images and text from a large number of image-text pairs. Through contrastive learning, the CLIP model can learn to embed images and related texts into the same high-dimensional space, so that similar image and text vectors are close to each other in this space, and dissimilar ones are far away from each other, thus achieving similar image sample generation.
[0093] In this embodiment, the initial training image set contains multiple types of target object images, and each image corresponds to one or more descriptive text labels or descriptions. These descriptive text labels or descriptions provide semantic information for the image. Each training image in the initial training image set and its corresponding descriptive text are input into the CLIP model. After the CLIP model learns the joint representation between the training image and its corresponding descriptive text, it can generate a text vector similar to the training image, and retrieve the image most similar to the given descriptive text from the text vector space through vector cosine similarity. In this way, image samples similar to the training image are automatically generated, that is, the image is expanded, thereby expanding the initial training image set and improving the diversity and richness of the training data.
[0094] In this embodiment, by introducing description text and CLIP model, the training process is not only concerned with the visual features of the image, but also with the semantic association between the image and the text. This helps the target classification model to learn a more in-depth semantic feature representation, and improves the target classification model's ability to understand the image content. The expanded image is semantically similar to the original image (training image in this embodiment) but visually changed, thereby increasing the diversity of training data, allowing the target classification model to be exposed to more images of different styles and conditions during the training process, thereby improving the generalization ability of the target classification model on new data. A more diverse and rich training image set helps the target classification model to learn a more accurate and discriminative feature representation, improves the accuracy of the target classification model in the classification task, and enables the target classification model to more accurately distinguish different types of target objects.
[0095] Optionally, the training images in the initial training image set and the descriptive text corresponding to the training images are obtained in the following manner: the training images in the initial training image set are retrieved, the training images are input into the image-to-text model, and the descriptive text corresponding to the training images is obtained. The image-to-text model, such as the BLIP (Bidirectional Language-Image Pre-training) model, can generate descriptive text related to the input image. In this embodiment, by inputting the training image into the image-to-text model and obtaining the descriptive text corresponding to the training image, the training image-descriptive text combination can be quickly obtained for image expansion.
[0096] Optionally, the initial training image set is expanded to obtain a training image set, including: inputting the training images in the initial training image set into an image-to-text model to obtain description texts corresponding to the training images; performing image retrieval based on the description texts to obtain expanded images; and obtaining a training image set based on the initial training image set and the expanded images.
[0097] In this embodiment, each training image in the initial training image set is first input into the image-to-text model. The image-to-text model can automatically generate accurate and descriptive text by analyzing the visual features of the image, so as to obtain the description text corresponding to the training image, which closely corresponds to and describes the content in the image. Subsequently, the generated description text is used as a keyword or query condition to search in a large image database to efficiently find images that are highly semantically matched with the description text, so as to obtain the expanded image. These images may be visually different from the initial training images, but they are closely related in content. The retrieved expanded images are merged with the initial training image set to form the final training image set. This training image set not only includes the original image (training image in this embodiment), but also incorporates the expanded images that are closely related to the original image in content but visually changed, thereby greatly enriching the diversity and coverage of the training data.
[0098] In this embodiment, the description text generated by the image-to-text model can closely combine the visual features of the image with the semantic information, so that the target classification model can learn a deeper semantic feature representation during the training process. The expanded image is closely related to the original image in content, but visually different, which increases the diversity of the training data. When the target classification model is exposed to more images of different styles and conditions, it can improve its generalization ability and better adapt to new data. In addition, a more diverse and rich training data set helps the target classification model learn more precise and discriminative feature representations, thereby improving the accuracy of classification tasks.
[0099] Optionally, the initial training image set is expanded to obtain a training image set, including: obtaining text labels corresponding to the training images in the initial training image set; inputting the text labels into the trained enhanced stable diffusion model to obtain expanded images; and obtaining the training image set based on the initial training image set and the expanded images.
[0100] In this embodiment, the initial training image set includes multiple types of target object images, and each image corresponds to one or more descriptive text labels or descriptions. These descriptive text labels or descriptions provide semantic information for the image. The enhanced stable diffusion model includes a Stable Diffusion XL model, referred to as the SDXL model. The SDXL model can generate corresponding images based on a given prompt word (a text label in this embodiment).
[0101] In this embodiment, the text labels of the initial training image set are input into the trained enhanced stable diffusion model. The enhanced stable diffusion model can generate extended images that are highly matched with the label content but visually changed by analyzing the semantic information of the text labels. These images show more details and style changes while maintaining consistency with the original labels. The generated extended images are merged with the initial training image set to form a final training image set. This training image set not only includes the original images (i.e., training images) and their corresponding text labels, but also incorporates extended images that are closely related to the original labels but visually changed, thereby greatly enriching the diversity and coverage of the training data.
[0102] By leveraging text labels and the trained enhanced stable diffusion model, the generated augmented images are semantically consistent with the original images. This helps the target classification model learn more accurate and consistent semantic feature representations during training, thereby improving the accuracy of the classification task. Although the augmented images are visually changed, they remain consistent with the original labels, which increases the diversity of the training data. When the target classification model is exposed to more images of different styles and conditions, it can improve its generalization ability and better adapt to new data.
[0103] Optionally, the enhanced stable diffusion model is trained in the following manner: obtain image prompt annotation files and trigger words that meet the training requirements of the enhanced stable diffusion model; use the LoRA method to inject a trainable low-rank matrix into the fully connected layer of the enhanced stable diffusion model to fine-tune the enhanced stable diffusion model; extract prompt words from the image prompt annotation file according to the trigger words; and input the prompt words into the fine-tuned enhanced stable diffusion model for training.
[0104] The enhanced stable diffusion model can generate corresponding images according to given prompt words. In order to train the enhanced stable diffusion model, it is necessary to prepare an image prompt annotation file that meets the requirements of the enhanced stable diffusion model. The image prompt annotation file contains a series of prompt words or descriptive texts corresponding to the image. The prompt words can accurately reflect the content of the image, such as the specific name, color, shape, etc. of the target object in the image. In the training process of the enhanced stable diffusion model, text drift is a common problem. By setting trigger words, text drift can be reduced. The trigger words are closely related to the content of the image and the training objectives. For example, when the trigger word is color, the prompt word can be red, blue or green, etc.; when the trigger word is shape, the prompt word can be triangle, rectangle or hexagon, etc. In this embodiment, the image prompt annotation file and trigger words that meet the training requirements of the enhanced stable diffusion model can be obtained by technicians, or by inputting the training image into the image-generated text model to obtain the descriptive text and text label corresponding to the training image.
[0105] LoRA (Low-Rank Adaptation) is a parameter efficient fine-tuning method. The core idea of LoRA is to simulate the change of parameters through low-rank decomposition, so as to realize indirect training of large models with extremely small parameters. In this embodiment, the application process of the LoRA method is as follows: the fully connected layer of the enhanced stable diffusion model is selected as the layer that needs to be fine-tuned; for the fully connected layer, the parameter matrix of the fully connected layer is low-rank decomposed by singular value decomposition to obtain two decomposition matrices; wherein the result of multiplying the two decomposition matrices is the same as the original parameter matrix of the fully connected layer; the two decomposition matrices obtained are added to the fully connected layer as additional trainable parameters to obtain the enhanced stable diffusion model after fine-tuning. Compared with traditional fine-tuning methods, the LoRA method has the advantages of fewer parameters, high computational efficiency, and easy deployment, while being able to maintain most of the original capabilities of the model.
[0106] According to the determined trigger words, the prompt words related thereto are extracted from the image prompt word annotation file. These prompt words are used as the input of the enhanced stable diffusion model to guide the enhanced stable diffusion model to generate images that are highly related to the prompt words. The extracted prompt words are input into the enhanced stable diffusion model after fine-tuning, and further training is performed so that the enhanced stable diffusion model can learn how to generate images with specific content and style according to the prompt words. In this embodiment, by introducing the trigger words and the image prompt word annotation file, the enhanced stable diffusion model with customized generation capability can be trained so that images highly related thereto can be generated according to the specified trigger words to meet the personalized image generation requirements.
[0107] Optionally, the target classification model is trained using a training image set to obtain a trained target classification model, including: the target classification model is trained using a reinforcement learning mechanism using the training image set to obtain a trained target classification model.
[0108] Reinforcement Learning (RL) is a machine learning paradigm in which an agent learns how to make decisions to maximize a certain long-term reward by interacting with the environment. In this embodiment, the reinforcement learning mechanism regards the target classification model as an agent, and during the training process, the agent learns how to make correct classification decisions by interacting with the environment (i.e., the training image set). Specifically, the agent receives an input image (training image) and outputs a classification prediction (classification result). Then, based on the accuracy of the prediction, the agent receives a reward signal, which is used to guide the agent to adjust its internal parameters to optimize future predictions. Through continuous trial and error and learning, the agent (i.e., the target classification model) gradually learns how to make accurate classification decisions based on the input image. When the performance of the agent on the training image set reaches a preset threshold or converges, it is considered that the target classification model has been trained and a trained target classification model is obtained.
[0109] Combination Figure 5 As shown, the embodiment of the present disclosure provides another target classification method, including:
[0110] S501: The processor obtains an image to be classified.
[0111] The image to be classified refers to the image containing the target object that needs to be classified.
[0112] S502, the processor inputs the image to be classified into a pre-trained target classification model to obtain a preliminary result output by the target classification model.
[0113] Among them, the target classification model includes a feature extraction model, a feature screening model, a feature combination model and an image classification model which are connected in sequence.
[0114] S503: The processor performs uncertainty evaluation on the preliminary result to obtain an uncertainty value.
[0115] In this step, the purpose of uncertainty assessment is to measure the confidence of the target classification model in the preliminary results, that is, how reliable the target classification model believes its classification results are. Uncertainty values can be calculated by a variety of methods, such as statistics such as entropy and variance based on probability distribution, or confidence intervals based on the output of models (such as Bayesian deep learning models).
[0116] S504: When the uncertainty value is less than or equal to the uncertainty threshold, the processor uses the preliminary result as the classification result.
[0117] S505: When the uncertainty value is greater than the uncertainty threshold, the processor updates the target classification model.
[0118] S506: The processor inputs the image to be classified into the updated target classification model to obtain a preliminary result again.
[0119] The target classification method provided by the embodiment of the present disclosure is pre-set with an uncertainty threshold value, which is used to determine whether the uncertainty of the preliminary result is within an acceptable range. If the uncertainty value is less than or equal to the uncertainty threshold value, the preliminary result is considered to be reliable and is used as the final classification result. If the uncertainty value is greater than the uncertainty threshold value, the reliability of the preliminary result is considered to be low, and the target classification model update mechanism is triggered to update the target classification model to improve the classification accuracy.
[0120] The embodiment of the present disclosure introduces an uncertainty assessment mechanism to quantitatively assess the reliability of preliminary results, thereby avoiding unreliable classification results as final output. In the case of high uncertainty values, the accuracy of the classification results can be further improved by updating the target classification model and reclassifying. The uncertainty assessment mechanism provides a self-checking and self-correction mechanism for the target classification model, so that the target classification model can make more robust decisions when facing uncertain or complex situations, enhances the reliability of the target classification model, and makes the target classification model more stable and credible in practical applications.
[0121] In some embodiments, updating the target classification model includes: sending the image to be classified and the classification result to the target terminal; receiving classification information fed back by the target terminal based on the image to be classified and the classification result; the classification information includes whether the classification result is correct or the classification result is wrong; when the classification information indicates that the classification result is correct, training the target classification model according to the image to be classified and the classification result to update the target classification model; when the classification information indicates that the classification result is wrong, obtaining the type label corresponding to the image to be classified; training the target classification model according to the image to be classified and the type label to update the target classification model.
[0122] In this embodiment, the target terminal can be any device or system that can receive and process data, such as a smart phone, a tablet computer, a server, etc. After the target classification model classifies the image to be classified, the image to be classified and its classification result are sent to the target terminal. After receiving the image to be classified and the classification result, the target terminal can display it to the user, and obtain classification information based on user interaction to determine the correctness of the classification result. Afterwards, the target terminal sends the classification information (including whether the classification result is correct or the classification result is wrong) back to the update system of the target classification model. If the classification information indicates that the classification result is correct, the target classification model is trained according to the image to be classified and the classification result to consolidate the recognition ability of the target classification model for the correct classification, so that the target classification model can classify more accurately when facing similar images in the future. If the classification information indicates that the classification result is wrong, the correct type label corresponding to the image to be classified is obtained (the correct type label corresponding to the image to be classified can be obtained from the classification information obtained by the target terminal and the user interaction), and the target classification model is trained according to the image to be classified and the type label to correct the errors of the target classification model in the classification process, and enable the target classification model to learn the correct classification rules.
[0123] In some embodiments, updating the target classification model includes: obtaining a training image set for the target classification model; classifying the training image set according to type labels of the training images in the training image set to obtain multiple sub-training image sets; training multiple target classification models separately using the sub-training image sets to obtain multiple sub-target classification models; and fusing the multiple sub-target classification models to update the target classification model.
[0124] In this embodiment, the training image set is classified according to the type labels of the training images in the training image set to obtain multiple sub-training image sets. Each sub-training image set contains training images belonging to the same category, so that the target classification model can learn more refined and specific feature representations. A sub-target classification model is trained with each sub-training image set. The sub-target classification models will learn the unique features of their respective categories during the training process, thereby improving the classification accuracy of the sub-target classification model for specific categories. Afterwards, multiple sub-target classification models are fused to obtain an updated target classification model. Among them, the fusion process can be fused using methods such as weighted averaging, voting mechanism or ensemble learning methods in deep learning. By fusing the prediction results of multiple sub-target classification models, the classification performance and generalization ability of the target classification model can be further improved.
[0125] In some embodiments, updating the target classification model includes: sending the image to be classified and the classification result to the target terminal; receiving classification information fed back by the target terminal based on the image to be classified and the classification result; the classification information includes whether the classification result is correct or incorrect; when the classification information indicates that the classification result is correct, training the target classification model according to the image to be classified and the classification result to update the target classification model; when the classification information indicates that the classification result is incorrect, obtaining the type label corresponding to the image to be classified; training the target classification model according to the image to be classified and the type label to update the target classification model and obtain a first classification model; obtaining a training image set for the target classification model; classifying the training image set according to the type labels of the training images in the training image set to obtain multiple sub-training image sets; using the sub-training image sets to train multiple target classification models respectively to obtain multiple sub-target classification models; fusing the multiple sub-target classification models to update the target classification model and obtain a second classification model; fusing the first classification model and the second classification model to obtain an updated target classification model.
[0126] By sending the image to be classified and the classification result to the target terminal and receiving feedback, the target classification model can understand its own classification performance in real time and make rapid adjustments accordingly, which helps the target classification model maintain a high classification accuracy in practical applications while reducing the risk of misclassification. Subdividing the training image set into multiple sub-training image sets and training multiple sub-target classification models separately helps the target classification model learn more refined and specific feature representations, improves the target classification model's classification ability for complex images, and enables the target classification model to better adapt to different classification tasks. In this embodiment, by further fusing the second classification model obtained by fusing multiple sub-target classification models and the preliminarily adjusted first classification model, a more comprehensive and accurate updated target classification model can be obtained to fully utilize the advantages of the first classification model and the second classification model to improve the overall classification performance and generalization ability.
[0127] Optionally, the first classification model and the second classification model are fused to obtain an updated target classification model, including: inputting the training images in the training image set into the first classification model and the second classification model respectively for feature extraction; splicing the initial features extracted by the first classification model and the second classification model to obtain a first intermediate feature; continuing to input the first intermediate feature into the first classification model or the second classification model for recognition, so as to retrain the first classification model or the second classification model; and using the retrained first classification model or the second classification model as the updated target classification model.
[0128] The first classification model and the second classification model both include a feature extraction model, a feature screening model, a feature combination model and an image classification model. In this embodiment, the training images in the training image set are respectively input into the feature extraction model of the first classification model and the feature extraction model of the second classification model for feature extraction, the initial features output by the two feature extraction models are spliced, and then the spliced features are continuously input into the feature screening model, feature combination model and image classification model of the first classification model, or the feature screening model, feature combination model and image classification model of the second classification model for recognition, so as to retrain the first classification model or the second classification model and update the target classification model.
[0129] In this embodiment, the initial features extracted by the first classification model and the second classification model are spliced to form a feature vector containing richer information, so as to integrate the feature extraction capabilities of the two first classification models and the second classification model and obtain a more comprehensive image representation. The spliced initial features, i.e., the first intermediate features, are further input into the first classification model or the second classification model for recognition. During the recognition process, the model will learn and adjust according to the spliced feature vector to optimize its classification performance and realize the updating and fusion of the target classification model.
[0130] Optionally, the first classification model and the second classification model are fused to obtain an updated target classification model, including: inputting the training images in the training image set into the first classification model and the second classification model respectively for feature extraction and feature screening; splicing the screened target features output by the first classification model and the second classification model to obtain a second intermediate feature; continuing to input the second intermediate feature into the first classification model or the second classification model for recognition, so as to retrain the first classification model or the second classification model; and using the retrained first classification model or the second classification model as the updated target classification model.
[0131] In this embodiment, the first classification model and the second classification model both include a feature extraction model, a feature screening model, a feature combination model and an image classification model. In this embodiment, the training images in the training image set are respectively input into the feature extraction model and the feature screening model of the first classification model for feature extraction and feature screening in sequence, and are input into the feature extraction model and the feature screening model of the second classification model for feature extraction and feature screening in sequence, the target features output by the two feature screening models are spliced, and then the spliced features are further input into the feature combination model and the image classification model of the first classification model, or the feature combination model and the image classification model of the second classification model for identification, so as to retrain the first classification model or the second classification model and update the target classification model.
[0132] In this embodiment, the target features extracted by the first classification model and the second classification model are spliced to form a feature vector containing richer information, so as to integrate the feature extraction and feature screening capabilities of the two first classification models and the second classification models and obtain a more comprehensive image representation. The spliced target features, i.e., the second intermediate features, are further input into the first classification model or the second classification model for recognition. During the recognition process, the model will learn and adjust according to the spliced feature vector to optimize its classification performance and realize the updating and fusion of the target classification model.
[0133] Optionally, the first classification model and the second classification model are fused to obtain an updated target classification model, including: inputting the training images in the training image set into the first classification model and the second classification model respectively for feature extraction, feature screening and feature combination; combining the target feature groups output by the first classification model and the second classification model to obtain a collection feature group; continuing to input the collection feature group into the first classification model or the second classification model for recognition, so as to retrain the first classification model or the second classification model; and using the retrained first classification model or the second classification model as the updated target classification model.
[0134] In this embodiment, the first classification model and the second classification model both include a feature extraction model, a feature screening model, a feature combination model and an image classification model. In this embodiment, the training images in the training image set are respectively input into the feature extraction model, the feature screening model and the feature combination model of the first classification model for feature extraction, feature screening and feature combination in sequence, and are input into the feature extraction model, the feature screening model and the feature combination model of the second classification model for feature extraction, feature screening and feature combination in sequence, and the target feature groups output by the two feature combination models are combined, and then the combined feature groups are further input into the image classification model of the first classification model, or the image classification model of the second classification model for identification, so as to retrain the first classification model or the second classification model and update the target classification model.
[0135] In this embodiment, the target features output by the first classification model and the second classification model are combined to form a feature group containing richer information, so as to integrate the feature extraction ability, feature screening ability and feature combination ability of the two first classification models and the second classification model to obtain a more comprehensive image representation. The combined feature group, i.e., the set feature group, is further input into the first classification model or the second classification model for recognition. During the recognition process, the model will learn and adjust according to the combined feature vector to optimize its classification performance and realize the updating and fusion of the target classification model.
[0136] It should be noted that the text vector space, image database, thresholds, etc. mentioned in the embodiments of the present disclosure are pre-set by technical personnel according to specific target objects to be classified and are not specifically limited here.
[0137] Combination Figure 6 As shown, the embodiment of the present disclosure provides a target classification device 60, including an acquisition module 610 and a classification module 620. The acquisition module 610 is configured to acquire an image to be classified; the classification module 620 is configured to input the image to be classified into a pre-trained target classification model: perform feature extraction on the image to be classified to obtain initial features; perform feature screening on the initial features to obtain target features; perform feature combination on the target features to obtain a target feature group; and identify the target feature group to obtain a classification result.
[0138] The target classification device 60 provided in the embodiment of the present disclosure is capable of executing the target classification method described in the above embodiment. Therefore, the technical effects possessed by the target classification method described in the above embodiment are also possessed by the embodiment of the present disclosure and will not be repeated here.
[0139] Optionally, the classification module 620 is further configured to input the image to be classified into a feature extraction model to perform feature extraction to obtain initial features.
[0140] Optionally, the classification module 620 is further configured to input the initial features into a feature screening model for feature screening to obtain target features.
[0141] Optionally, the classification module 620 is further configured to input the target features into a feature combination model to perform feature combination to obtain a target feature group.
[0142] Optionally, the classification module 620 is further configured to input the target feature group into an image classification model for identification to obtain a classification result.
[0143] Optionally, the target classification device 60 also includes a training module 630, which is configured to obtain an initial training image set; perform image expansion on the initial training image set to obtain a training image set; and use the training image set to train the target classification model to obtain a trained target classification model.
[0144] Optionally, the training module 630 is also configured to perform image transformation on the training images in the initial training image set to obtain an expanded image; the image transformation includes one or more of image rotation, image cropping, dithering, noise addition, and enhancement filtering; and a training image set is obtained based on the initial training image set and the expanded image.
[0145] Optionally, the training module 630 is also configured to obtain training images in the initial training image set and descriptive texts corresponding to the training images; input the training images and the descriptive texts corresponding to the training images into a comparative language-image pre-training model to obtain expanded images; and obtain a training image set based on the initial training image set and the expanded images.
[0146] Optionally, the training module 630 is also configured to input the training images in the initial training image set into the image-to-text model to obtain the descriptive text corresponding to the training images; perform image retrieval based on the descriptive text to obtain the expanded images; and obtain the training image set based on the initial training image set and the expanded images.
[0147] Optionally, the training module 630 is further configured to obtain text labels corresponding to training images in the initial training image set; input the text labels into the enhanced stable diffusion model to obtain expanded images; and obtain a training image set based on the initial training image set and the expanded images.
[0148] Optionally, the classification module 620 is also configured to input the image to be classified into a pre-trained target classification model to obtain a preliminary result output by the target classification model; perform uncertainty evaluation on the preliminary result to obtain an uncertainty value; when the uncertainty value is less than or equal to an uncertainty threshold, use the preliminary result as the classification result; when the uncertainty value is greater than the uncertainty threshold, update the target classification model; and input the image to be classified into the updated target classification model to obtain the preliminary result again.
[0149] Optionally, the classification module 620 is also configured to send the image to be classified and the classification result to the target terminal; receive classification information fed back by the target terminal based on the image to be classified and the classification result; the classification information includes whether the classification result is correct or the classification result is wrong; when the classification information indicates that the classification result is correct, the target classification model is trained according to the image to be classified and the classification result to update the target classification model; when the classification information indicates that the classification result is wrong, the type label corresponding to the image to be classified is obtained; the target classification model is trained according to the image to be classified and the type label to update the target classification model.
[0150] Optionally, the classification module 620 is also configured to obtain a training image set for a target classification model; classify the training image set according to the type labels of the training images in the training image set to obtain multiple sub-training image sets; use the sub-training image sets to train multiple target classification models respectively to obtain multiple sub-target classification models; and fuse the multiple sub-target classification models to update the target classification model.
[0151] Combination Figure 7As shown, the embodiment of the present disclosure provides a target classification device 70, including a processor (processor) 700 and a memory (memory) 701. Optionally, the device 70 may also include a communication interface (CommunicationInterface) 702 and a bus 703. Among them, the processor 700, the communication interface 702, and the memory 701 can communicate with each other through the bus 703. The communication interface 702 can be used for information transmission. The processor 700 can call the logic instructions in the memory 701 to execute the target classification method of the above embodiment.
[0152] In addition, the logic instructions in the memory 701 described above may be implemented in the form of software functional units and when sold or used as independent products, may be stored in a computer-readable storage medium.
[0153] The memory 701 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 700 executes the function application and data processing by running the program instructions / modules stored in the memory 701, that is, implementing the target classification method in the above embodiment.
[0154] The memory 701 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 701 may include a high-speed random access memory and may also include a non-volatile memory.
[0155] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned target classification method.
[0156] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes.
[0157] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.
[0158] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0159] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
[0160] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A target classification method, characterized in that: include: Get the image to be classified; Input the image to be classified into the pre-trained target classification model: Perform feature extraction on the image to be classified to obtain initial features; Perform feature screening on the initial features to obtain the target features; Combining target features to obtain a target feature group; Identify the target feature group and obtain the classification result.
2. The method according to claim 1, characterized in that The target classification model includes a feature extraction model; Extracting features of the image to be classified to obtain initial features, including: inputting the image to be classified into a feature extraction model to extract features to obtain initial features; and / or, The target classification model includes a feature screening model; performing feature screening on the initial features to obtain the target features, including: inputting the initial features into the feature screening model to perform feature screening to obtain the target features; and / or, The target classification model includes a feature combination model; combining the target features to obtain a target feature group includes: inputting the target features into the feature combination model to perform feature combination to obtain the target feature group; and / or, The target classification model includes an image classification model; identifying the target feature group to obtain a classification result includes: inputting the target feature group into the image classification model for identification to obtain a classification result.
3. The method according to claim 2, characterized in that The feature extraction model includes a convolutional neural network model; and / or, the feature screening model includes a support vector machine; and / or, the feature combination model includes a random forest; and / or, the image classification model includes a support vector machine or a random forest.
4. The method according to any one of claims 1 to 3, characterized in that: Obtain a pre-trained target classification model as follows: Obtain an initial training image set; Perform image expansion on the initial training image set to obtain a training image set; The target classification model is trained using a training image set to obtain a trained target classification model.
5. The method according to claim 4, characterized in that The initial training image set is expanded to obtain a training image set, including: Performing image transformation on the training images in the initial training image set to obtain an expanded image; the image transformation includes one or more of image rotation, image cropping, dithering, and enhancement filtering; obtaining a training image set based on the initial training image set and the expanded image; and / or, Obtaining training images and description texts corresponding to the training images in an initial training image set; inputting the training images and description texts corresponding to the training images into a contrastive language-image pre-training model to obtain an expanded image; obtaining a training image set based on the initial training image set and the expanded image; and / or, Input the training images in the initial training image set into the image-to-text model to obtain the description text corresponding to the training image; perform image retrieval based on the description text to obtain the expanded image; obtain the training image set based on the initial training image set and the expanded image; and / or, Obtain text labels corresponding to training images in the initial training image set; input the text labels into the trained enhanced stable diffusion model to obtain expanded images; and obtain a training image set based on the initial training image set and the expanded images.
6. The method according to any one of claims 1 to 3, characterized in that: Input the image to be classified into the pre-trained target classification model to obtain the classification results, including: Input the image to be classified into the pre-trained target classification model to obtain the preliminary results output by the target classification model; Conduct uncertainty assessment on preliminary results and obtain uncertainty values; When the uncertainty value is less than or equal to the uncertainty threshold, the preliminary result is taken as the classification result; When the uncertainty value is greater than the uncertainty threshold, the target classification model is updated; and the image to be classified is input into the updated target classification model to obtain the preliminary result again.
7. The method according to claim 6, characterized in that Update the target classification model, including: Sending the image to be classified and the classification result to the target terminal; receiving classification information fed back by the target terminal based on the image to be classified and the classification result; the classification information includes whether the classification result is correct or incorrect; when the classification information indicates that the classification result is correct, training the target classification model according to the image to be classified and the classification result to update the target classification model; when the classification information indicates that the classification result is incorrect, obtaining the type label corresponding to the image to be classified; training the target classification model according to the image to be classified and the type label to update the target classification model; and / or, Obtain a training image set for a target classification model; classify the training image set according to the type labels of the training images in the training image set to obtain multiple sub-training image sets; use the sub-training image sets to train multiple target classification models respectively to obtain multiple sub-target classification models; fuse the multiple sub-target classification models to update the target classification model.
8. A target classification device, characterized in that: include: An acquisition module is configured to acquire an image to be classified; The classification module is configured to input the image to be classified into a pre-trained target classification model: extract features from the image to be classified to obtain initial features; Perform feature screening on the initial features to obtain the target features; Combining target features to obtain a target feature group; Identify the target feature group and obtain the classification result.
9. A target classification device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the target classification method according to any one of claims 1 to 7 when running the program instructions.
10. An electronic device, characterized in that: include: Equipment body; The target classification device as described in claim 8 or 9 is installed on the device body.