Fine-grained image classification method, device and equipment

By combining image segmentation, feature extraction and deep learning classification models, the problems of low efficiency and high error rate of fine-grained image classification in traditional methods are solved, and efficient and accurate fine-grained image classification is achieved, which is suitable for a variety of complex application scenarios.

CN120147704APending Publication Date: 2025-06-13EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510206271.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional fine-grained image classification method relies on manual experience, has low efficiency and high error rate, making it difficult to achieve satisfactory results when there are many categories.

Method used

Through the combination of image segmentation, feature extraction and deep learning classification models, efficient classification of fine-grained targets can be achieved. The specific steps include installing a macro lens on the front end of the camera, using an APP with target pre-positioning function for image acquisition, using the segmentation model for precise segmentation and standardization, and training the image classification model based on the class-center two-dimensional loss function, and finally input the fused feature into the model for inference.

Benefits of technology

It realizes efficient and accurate fine-grained image classification, which can automatically segment, standardize processing and feature extraction on mobile devices, eliminate useless background information, significantly improve the intelligence level of image classification, and adapt to complex application scenarios with a wide variety of targets and small differences in details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147704A_ABST
    Figure CN120147704A_ABST
Patent Text Reader

Abstract

According to the fine-grained image classification method, device and equipment provided by the invention, efficient and accurate target category identification is realized by utilizing a target segmentation algorithm, an image classification algorithm and computing equipment. Through the method, automatic segmentation, standardization processing, feature extraction and useless background information elimination of a target image can be completed on a mobile device, fusion processing of geometric features and image features is combined, fusion features are input into a neural network model, and a target category result with the highest confidence coefficient is accurately output. The method has the advantages of being high in portability, wide in application range, high in recognition efficiency, standardized in processing flow and the like, the intelligent level of image classification can be remarkably improved, the method adapts to complex application scenes with various target types and tiny detail differences, meanwhile, the algorithm can be deployed to mobile terminal equipment, and the high compatibility of the mobile terminal equipment is utilized, so that the recognition efficiency is improved. High-efficiency operation on various devices is realized, and important technical support is provided for intelligentization and popularization of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a fine-grained image classification method, apparatus, and device. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technologies, fine-grained image classification has become an important research direction in the field of image analysis. Fine-grained classification aims to distinguish targets with high similarity but subtle category differences, such as birds, vehicle models, insects, pathological classification in medical images, etc. Due to the tiny differences between categories in such tasks, traditional classification methods are difficult to achieve satisfactory results. Taking lock core recognition as an example, traditional methods mainly rely on manual experience, which is not only inefficient but also has a high error rate. When the number of categories is large, the learning cost and error rate of manual recognition increase exponentially. In addition, in other fine-grained classification application scenarios, such as industrial quality inspection, ecological protection, and medical diagnosis, efficient and accurate classification technologies are also required. To solve the above problems, fine-grained image classification has gradually shifted from solely relying on manual experience to an intelligent method that combines deep learning models with image feature extraction. The model based on edge computing can integrate segmentation, feature extraction, and classification, and can improve the processing efficiency while maintaining high classification accuracy, adapting to various complex scenarios. Summary of the Invention

[0003] The object of the present invention is to provide a fine-grained image classification method, apparatus, and device, which can achieve efficient classification of fine-grained targets through the combination of image segmentation, feature extraction, and deep learning classification models, and the fine-grained image classification method can be deployed to a mobile terminal.

[0004] The specific technical solutions for achieving the object of the present invention are as follows:

[0005] A fine-grained image classification method, the method comprising:

[0006] Step S100: Install a macro lens in front of the camera to ensure that the captured target image can obtain sufficient clear detail information, and the detail information includes the key features of the target object; the key features are the beak, feathers, and wings of a bird; the shape and internal structure of a lock core; the veins and petal textures of a plant;

[0007] Step S200: Use an APP with a target pre-positioning function to collect images, ensure that the target is within a preset area, and save the image with the captured detail information;

[0008] Step S300: Use a segmentation model to accurately segment the target area, remove the irrelevant background, and only retain the effective area of the target;

[0009] Step S400: Scale the segmented target image proportionally while keeping its aspect ratio unchanged, center it and overlay it on a pure white background of the same size, and finally uniformly adjust all images to a fixed size;

[0010] Step S500: Train an image classification model based on the class center two-dimensional loss function, fuse the geometric features and pixel features of the image obtained in Step S400, input the fused features into the image classification model, and finally output the classification result.

[0011] Further, Step S300 specifically includes:

[0012] Step S310: Convert the segmentation model into a format suitable for mobile devices and load the model through OpenCV;

[0013] Step S320: According to the pre-positioning information, select the positions of the 1 / 4, 2 / 4, and 3 / 4 points on the central axis, set these points as mask points, and set the label value of the mask points to 1;

[0014] Step S330: Use the segmentation model to perform segmentation processing on the input image to generate a segmentation image containing only the target area.

[0015] Further, Step S400 specifically includes:

[0016] Step S410: Perform morphological operations on the segmented image, including dilation, erosion, and filtering, to optimize the edge quality and detail performance of the target image;

[0017] Step S420: Scale the image proportionally according to the aspect ratio of the segmented target image;

[0018] Step S430: Create a white background and place the scaled target image in the center of the background to generate a standardized target image.

[0019] Further, Step S500 specifically includes:

[0020] Step S510: Train an image classification model based on the class center two-dimensional loss function;

[0021] Step S520: Use the Canny algorithm to extract the edge information of the target;

[0022] Step S530: Calculate the perimeter of the target based on the target edge;

[0023] Step S540: Use the minimum inscribed circle method to calculate the major axis length of the target;

[0024] Step S550: Calculate the geometric features of the target by the ratio of the target edge perimeter to the major axis.

[0025] Step S560: Load a pre-trained image classification model through OpenCV;

[0026] Step S570: Input the target image into the trained image classification model, extract the pixel features of the layer before the fully connected layer, expand the geometric features into two-dimensional vectors with the same shape, and then add them to the pixel features with weights to obtain the fused features;

[0027] Step S580: Input the fused features into the fully connected layer for inference and output the classification result.

[0028] Furthermore, the step S510 specifically includes:

[0029] Take multiple target images of various categories to construct a training dataset;

[0030] Use a segmentation model to segment the target area, remove background interference, and extract the effective area of the target;

[0031] Perform normalization processing on the segmented target images to provide a consistent data format;

[0032] Extract the geometric features and image features of the target, and fuse the two to generate a feature vector containing multi-dimensional information;

[0033] Input the fused target features into the neural network model to be trained, and iteratively optimize based on the class center two-dimensional loss function to finally obtain the image classification model; where:

[0034] The class center two-dimensional loss function consists of two parts, L feature and L class , where L feature aims to minimize the within-class variance and improve the between-class separability; L feature The calculation method is as follows:

[0035]

[0036] where sim i represents the class index most similar to the target class i, represents the similarity between X i and , Peason represents the absolute value of the Pearson correlation coefficient, X i represents the current class center feature of class i, N represents the number of samples in a batch, and assume the current sample feature vector is f;

[0037] L class uses the JS divergence to measure the similarity between the predicted probability distribution of the samples and the soft label distribution; L classExpressed as:

[0038]

[0039] where M represents the sample feature and the average value of the soft label distribution , and N represents the number of samples in a batch; and represents the probability distribution and M, and the KL divergence between and M, represents the probability distribution and the JS divergence between;

[0040] The final class center two-dimensional loss function is expressed as:

[0041] L cc = a 1 L feature + a 2 L class + a 3 L CE

[0042] where L cc represents the class center two-dimensional loss function, L CE represents the cross-entropy loss function, and a represents the weight value.

[0043] A fine-grained image classification device, the device includes:

[0044] An image acquisition module, configured to acquire a target image to be recognized, ensure that the target is located within a preset area, and provide accurate input data for subsequent processing;

[0045] An image segmentation module, configured to perform segmentation processing on the target image, remove irrelevant background information, and only retain the effective area of the target;

[0046] An image normalization module, configured to perform morphological processing on the target image, remove noise and edge burrs, and adjust the target image to a unified standard size;

[0047] An image classification module, configured to input the target fusion feature into a pre-trained image classification model and output a classification result.

[0048] Further, the image classification module includes:

[0049] A geometric feature extraction unit, configured to extract the edge information of the target, calculate the perimeter and major axis length of the target edge, and generate the geometric feature of the target through the ratio of the two;

[0050] An image feature extraction unit for extracting deep-level image features from a target image and capturing the texture and morphological information of the target;

[0051] A feature fusion unit for fusing the geometric features and image features of the target to generate a feature vector that comprehensively expresses the characteristics of the target;

[0052] A model inference unit for inputting the fused features into a pre-trained recognition model to infer and obtain a classification result.

[0053] An edge computing device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0054] An edge computing device-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0055] The technical effects achieved by the present invention are as follows:

[0056] The present invention provides a fine-grained image classification method, device, and device, which utilize a target segmentation algorithm, an image classification algorithm, and a computing device to achieve efficient and accurate target category recognition. Through this method, the target image can be automatically segmented, standardized, and feature-extracted on a mobile device, eliminating useless background information. By combining the fusion processing of geometric features and image features, the fused features are input into a neural network model to accurately output the target category result. The invention has the advantages of strong portability, wide application range, high recognition efficiency, and standardized processing process, can significantly improve the intelligent level of image classification, adapt to complex application scenarios with a wide variety of target types and small detail differences. At the same time, the fine-grained image classification method can be deployed to mobile devices, and by utilizing the high compatibility of mobile devices, it realizes efficient operation on various devices, providing important technical support for the intelligence and popularization of image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the application scenario of the fine-grained image classification method in an embodiment of the present invention;

[0058] Figure 2 It is a schematic flowchart of the fine-grained image classification method in an embodiment of the present invention;

[0059] Figure 3 It is a structural block diagram of the image classification device in an embodiment of the present invention;

[0060] Figure 4 It is an internal structure diagram of the device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] To help those skilled in the art better understand one or more technical solutions in the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the specific embodiments described here are only used to explain the content of the present invention, rather than limiting its scope.

[0062] An embodiment of the present invention provides a fine-grained image classification method, device, and equipment, which can be applied to a shooting scene as Figure 1 shown. Specifically, the staff can implement an image classification algorithm on a mobile device, and after installing a macro lens, capture a clear target image according to the shooting requirements. Subsequently, the image classification algorithm will automatically process and output the image classification result, simplifying the operation process and improving the recognition efficiency.

[0063] In this embodiment, as Figure 2 shown, a fine-grained image classification method is provided, and the method includes:

[0064] Step S100: Install a macro lens in front of the camera to ensure that the captured target image can capture sufficient clear detail information;

[0065] Specifically, the macro lens can effectively magnify the details of the target image and improve the resolution. Especially when the target size is small and the detail features are complex, the macro lens can significantly enhance the image clarity and ensure the accuracy of subsequent segmentation and feature extraction. In addition, the macro lens can also reduce image distortion and provide higher-quality raw data input for the extraction of geometric and pixel features of the target.

[0066] Step S200: Use an image localization algorithm with a target pre-localization function to collect images, ensure that the target is within a preset area, and save the captured images for subsequent processing;

[0067] Specifically, the image localization algorithm provides an intuitive rectangular frame localization guide, which is convenient for the staff to quickly complete the calibration of the target position. During the shooting process, the built-in camera will automatically adjust the focus to ensure that the target image is clearly visible. The saved images will be stored in a standard format locally or in the cloud, providing reliable input data for subsequent segmentation, feature extraction, and recognition operations.

[0068] Step S300: Use a segmentation model to accurately segment the target area, remove the irrelevant background, and only retain the effective area of the target;

[0069] Specifically, first, the captured image is input into the segmentation model, and global and local features are extracted through the built-in visual architecture of the model. Then, the interference background around the target, such as door frames, light reflections, etc., is removed using the segmentation mask, and only the target main body part is retained. The segmented target image provides high-quality input for subsequent processing and reduces the influence of background interference on the extraction of geometric features and pixel features.

[0070] Step S400: Scale the segmented target image while keeping the aspect ratio unchanged, and uniformly adjust the image to the same size.

[0071] Specifically, first, morphological processing is performed on the segmented image, including dilation and erosion operations, to remove burrs and noise points on the segmentation edges. Then, proportional scaling is performed according to the actual aspect ratio of the image. Finally, the target image is centered on a white background to eliminate the interference of the background color on subsequent feature extraction and ensure the consistency and standardization of image input.

[0072] Step S500: Fuse the geometric features and pixel features of the target, and input the fused features into a pre-trained image classification model, and finally output the classification result.

[0073] Specifically, the target edge is extracted through the Canny algorithm, the perimeter and the major axis length are calculated, and the geometric features of the target are generated according to the ratio of the two. At the same time, the standardized target image is input into the neural network model to extract the high-dimensional feature vector of the image. Then, the geometric features are extended into two-dimensional vectors with the same shape, and then weighted and added to the pixel features of the image to obtain the fused features. Finally, the fused features are input into the pre-trained image classification model for inference analysis, and the target classification result and its corresponding confidence level are output, providing a reliable basis for subsequent decision-making.

[0074] In this embodiment, in step S300: Use the segmentation model to accurately segment the target area, remove the irrelevant background, and only retain the effective area of the target; specifically including:

[0075] Step S310: Convert the segmentation model into a format suitable for mobile devices, and load the model through OpenCV.

[0076] Specifically, use the conversion algorithm to convert the model from the original framework format into a format suitable for mobile devices to achieve cross-platform deployment; load the model efficiently through OpenCV to ensure its running performance on mobile devices.

[0077] Step S320: According to the pre-positioning information, select the positions of the 1 / 4, 2 / 4, and 3 / 4 points on the central axis, set these points as mask points, and set the label value of the mask points to 1.

[0078] Specifically, calculate the coordinates of the center point of the target image, and according to the segmentation requirements, generate several mask points at appropriate positions offset downward by the height of the image to provide initial guiding information for model segmentation and ensure the accuracy of the segmentation results.

[0079] Step S330: Use the segmentation model to perform segmentation processing on the input image to generate a segmentation image that only contains the target area.

[0080] Specifically, input the preprocessed image and the mask points into the model together. The model removes the background area according to the input information and only retains the effective area of the target, and the output result is used as the basic data for subsequent processing.

[0081] In this embodiment, Step S400: Scale the segmented target image while keeping the aspect ratio unchanged, and uniformly adjust the image to the same size; specifically including:

[0082] Step S410: Perform morphological operations on the segmented image, including dilation, erosion, and filtering, to optimize the edge quality and detail performance of the target image;

[0083] Specifically, fill in the small broken parts in the image through dilation operation to ensure the integrity of the target area; remove the noise points at the edges of the image through erosion to further highlight the main features; use Gaussian filtering to smooth the image to reduce the interference of texture details while retaining the key features of the target.

[0084] Step S420: Scale the image proportionally according to the aspect ratio of the segmented target image;

[0085] Specifically, determine the width and height of the target image, and adjust the image size according to the principle that the larger side is equal to the specified number of pixels, ensure that the length and width of the scaled image are both less than or equal to the specified number of pixels, and at the same time keep the image ratio unchanged to avoid deformation.

[0086] Step S430: Create a white background and place the scaled target image in the center of the background to generate a standardized target image.

[0087] Specifically, calculate the center position of the target image relative to the background, fill in the blank area in the background, and make the target image located in the center, so as to generate an image with standardized resolution pixels and provide data in a unified format for subsequent processing.

[0088] In this embodiment, Step S500: Train an image classification model based on the class center two-dimensional loss function, fuse the geometric features and pixel features of the image obtained in Step S400, input the fused features into the image classification model, and finally output the classification result; specifically including:

[0089] Step S510: obtaining an image classification model based on training of a class center two-dimensional loss function;

[0090] Specifically, the image classification model is used to identify the target type in the target image, which can be pre-trained based on the neural network model. The neural network model can be trained by the following method to obtain the image classification model:

[0091] Step 1: Take multiple target images of various categories to build a rich training data set;

[0092] Specifically, to ensure the diversity of the data set, it is necessary to take target images at different angles and under different lighting conditions. The samples are accurately classified and annotated through image annotation tools to generate a standardized data set for training. In addition, to enhance the robustness of the model, data augmentation operations can be performed on the image samples, such as rotation, cropping, mirror flipping, and brightness adjustment, thereby improving the model's adaptability to various complex scenes.

[0093] Step 2: Use the segmentation model to segment the target area, remove background interference, and extract the effective area of ​​the target;

[0094] Specifically, a segmentation model is introduced to accurately segment the target area. First, the captured target image is input into the model. According to the pre-positioning information, the appropriate point on the central axis is selected and set as the mask point to optimize the segmentation effect. The model uses a deep learning algorithm to remove useless background in the image and retain only the effective part of the target. The segmented target area has a clear outline and a clean background, providing high-quality input for subsequent feature extraction.

[0095] Step 3: Standardize the segmented target image to provide a consistent data format for model input;

[0096] Specifically, morphological processing and size normalization are performed on the segmented target image. Morphological processing includes operations such as erosion, dilation, and denoising to further optimize the edge quality of the segmented area. Subsequently, the target image is resized to a fixed size and centered on a white background to ensure size consistency and center alignment of all image input data.

[0097] Step 4: Extract the geometric features and image features of the target, and fuse them to generate a feature vector containing multi-dimensional information;

[0098] Specifically, the geometric feature extraction module extracts the edge of the target through the Canny algorithm, calculates its perimeter and major axis length, and generates a geometric feature vector through ratio operation. The image feature extraction module uses a deep neural network to extract high-level pixel features from the image. Subsequently, the geometric features and image features are fused through the feature fusion module to generate a multi-dimensional feature vector for classification.

[0099] Step Five: Input the fused target features into the neural network model to be trained, and iteratively optimize based on the class-center two-dimensional loss function we proposed, finally obtaining an image classification model with excellent performance;

[0100] Specifically, input the fused features into the designed deep neural network architecture. Through the backpropagation algorithm, use the training dataset to iteratively optimize the model and adjust the network weights to minimize the loss function value. During the training process, use the class-center two-dimensional loss function to evaluate the generalization performance of the model, and adjust the hyperparameters (such as the learning rate and batch size) to obtain the best classification effect. Finally, the fully trained image classification model can quickly and accurately predict the category of the target, and output the classification result and its corresponding confidence distribution, meeting the actual application requirements.

[0101] Among them, the principle of the class-center two-dimensional loss function is as follows:

[0102] The class-center two-dimensional loss function consists of two parts, L feature and L class , where L feature aims to minimize the within-class variance and improve the between-class separability. L feature The calculation method is as follows:

[0103] Let X i represent the current class-center feature of class i, and K i represent the updated weight. Given a new sample feature vector f i , the updated class-center feature X i+1 and the corresponding weight K i+1 are calculated as follows:

[0104] X i+1 K i+1 = X i K i + f i

[0105] K i+1 = K i + 1

[0106] To maximize the between-class separability, first identify the class most similar to the target class i. The similarity between class centers is determined using the Pearson correlation coefficient, and the class most similar to i is calculated as:

[0107] sim i = argmax(s i,j )(i ≠ j)

[0108] s h,w = |Peason(X h , Xw )|

[0109] where sim i represents the class index most similar to the target class i, argmax represents taking the maximum value, and s i,j represents X i and X j the similarity between them, and Peason represents the absolute value of the Pearson correlation coefficient. L feature is finally expressed as:

[0110]

[0111] L class aims to generate soft labels to alleviate overfitting. These soft labels are derived from the class center distribution. Similar to the class center features, during the training phase, by averaging the distributions of all input samples belonging to class y i the class center distribution of the y i -th class is calculated The update process is defined as:

[0112]

[0113] where represents the state value, and the symbol corresponds to the updated state. FC represents the feature vector obtained from the final fully connected layer. The probability distribution of the sample feature is calculated as follows:

[0114]

[0115] The soft label distribution is:

[0116]

[0117] L class uses the JS divergence to measure the similarity between the predicted probability distribution of the sample and the soft label distribution. L class can be expressed as:

[0118]

[0119] where KL represents the KL divergence;

[0120] The final class center two-dimensional loss function can be expressed as:

[0121] L cc = a 1 L feature + a 2 L class + a 3 L CE

[0122] Among them, L cc represents the class - center two - dimensional loss function, and L CE represents the cross - entropy loss function, and a represents the weight value.

[0123] The main purpose of the class - center two - dimensional loss function is to reduce the within - class variance while increasing the between - class separability, thereby promoting the learning of more discriminative features. In addition, by generating soft labels, this loss effectively alleviates overfitting, enabling the model to better generalize different samples. The design of the class - center two - dimensional loss function ensures a balanced optimization process, promoting robust feature representation for the classification task.

[0124] Step S520: Use the Canny algorithm to extract the edge information of the target;

[0125] Specifically, the Canny edge detection algorithm is a commonly used edge detection method that can accurately extract significant edge features in an image. First, convert the target image to a grayscale image and smooth the image through Gaussian filtering to reduce noise interference. Then, use the double - threshold method to calculate the gradient values of the image and retain the local maximum values of the edges through non - maximum suppression, thereby extracting clear edge information of the target. These edge data provide the basis for subsequent geometric feature calculations.

[0126] Step S530: Calculate the perimeter of the target based on the target edge;

[0127] Specifically, using the edge map extracted in step S520, locate the boundary contour of the target through the contour detection algorithm, and then calculate the perimeter of the contour. Through accurate edge data, it can be ensured that the calculated perimeter accurately reflects the actual physical characteristics of the target, providing key parameters for subsequent geometric feature extraction.

[0128] Step S540: Calculate the major axis length of the target using the minimum - inscribed - circle method;

[0129] Specifically, the minimum - inscribed - circle is the smallest circular region centered on the target, which can effectively reflect the geometric symmetry and size characteristics of the target. Use OpenCV to fit the target contour and calculate and output the radius and center coordinates of the minimum - inscribed - circle. The major axis length of the target can be approximated as twice the diameter of the inscribed circle, and this parameter provides an important dimension for the geometric features of the target.

[0130] Step S550: Calculate the geometric feature of the target by the ratio of the perimeter of the target edge to the major axis;

[0131] Specifically, perform a ratio operation on the perimeter calculated in step S530 and the major axis length calculated in step S540. This ratio can reflect the complexity of the target shape. For example, for regular-shaped targets (such as circles or rectangles), the ratio has high consistency; while for targets with more complex shapes, the ratio will fluctuate significantly. This geometric feature can provide key discriminant information in target classification.

[0132] Step S560: Load a pre-trained image classification model through OpenCV;

[0133] Specifically, use the OpenCV library to load a neural network model suitable for the mobile format. The model architecture can be based on ResNet, where ResNet is a commonly used deep learning network architecture with excellent feature extraction capabilities, capable of finding key patterns in the multi-level features of the target image. Through this model, deep learning inference tasks can be efficiently completed on mobile devices.

[0134] Step S570: Input the target image into the trained image classification model, extract the pixel features of the layer before the fully connected layer, and expand the geometric features into two-dimensional vectors of the same shape, and then add them to the pixel features with weights to obtain the fused features;

[0135] Specifically, first, pass the standardized target image through the convolutional layer and pooling layer of the neural network model to extract deep high-dimensional pixel features. Subsequently, use the geometric features through linear transformation or dimensionality increase operations to convert them into feature vectors that match the pixel features. Finally, combine the geometric features and pixel features through a feature fusion module (such as weighted summation or fully connected fusion) to generate a comprehensive feature vector, providing a more comprehensive input for the classification task.

[0136] Step S580: Input the fused features into the fully connected layer for inference and output the classification result.

[0137] Specifically, pass the fused features through the fully connected layer and use the activation function to generate the probability distributions of various categories. The model will calculate the top five target categories with the highest confidence and their corresponding probability values according to the similarity between the fused features and the target categories. The results will be output in a list form for the user to refer to for the final determination and confirmation of the target category.

[0138] In this embodiment, as Figure 3 shown, a fine-grained image classification device based on edge learning, the device includes:

[0139] An image acquisition module, used to obtain the target image to be recognized, ensure that the target is within the preset area, and provide accurate input data for subsequent processing;

[0140] An image segmentation module, configured to perform segmentation processing on a target image, remove irrelevant background information, and only retain the valid region of the target;

[0141] An image normalization module, configured to perform morphological processing on the target image, remove noise and edge burrs, and adjust the target image to a unified standard size;

[0142] An image classification module, configured to input the target fusion features into a pre-trained image classification model and output a classification result.

[0143] In this embodiment, the image classification module includes:

[0144] A geometric feature extraction unit, configured to extract edge information of the target, calculate the perimeter and major axis length of the target edge, and generate geometric features of the target through the ratio of the two;

[0145] An image feature extraction unit, configured to extract deep-level image features from the target image and capture texture and morphological information of the target;

[0146] A feature fusion unit, configured to fuse the geometric features and image features of the target to generate a feature vector comprehensively expressing the characteristics of the target;

[0147] A model inference unit, configured to input the fused features into a pre-trained recognition model to infer and obtain a classification result.

[0148] In this embodiment, as Figure 4 shown, an edge computing device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned fine-grained image classification method are implemented.

[0149] An edge computing device-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned fine-grained image classification method are implemented.

[0150] Those skilled in the art should be aware that all or part of the processes in the above embodiments can be implemented by computer program instructions on hardware devices. These computer programs can be stored in a non-volatile computer-readable storage medium and, when executed, complete the processes described in the above methods. The memories, storage media, databases, or other devices mentioned in the embodiments of the present invention may include non-volatile memory and / or volatile memory. For example, non-volatile memory includes read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory; volatile memory includes random access memory (RAM) or external cache memory, etc. The forms of RAM include, but are not limited to, static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus RAM (Rambus RDRAM), direct memory bus dynamic RAM (DRDRAM), etc.

[0151] The technical features in the embodiments of the present invention can be flexibly combined. For the sake of concise description, not all possible combinations of technical features are elaborated in detail, but as long as these combinations do not have logical conflicts, they fall within the scope of this specification.

[0152] In addition, the above embodiments are only illustrations of several specific implementation manners of the technical solutions of the present invention, intended to provide clear and detailed content descriptions, and do not mean to limit the scope of patent protection. Without departing from the basic concept of the present invention, those skilled in the art can make various changes and optimizations to it, and these changes and optimizations should also be regarded as the content protected by the present invention. Therefore, the final scope of protection shall be subject to the appended claims.

Claims

1. A fine-grained image classification method, characterized in that: The method comprises: Step S100: A macro lens is installed at the front end of the camera to ensure that the captured target image can obtain sufficiently clear detail information, the detail information including key features of the target object; the key features are the beak, feathers and wings of a bird; the shape and internal structure of a lock core; the veins and petal textures of a plant; Step S200: Use an APP with a target pre-positioning function to capture images, ensure that the target is located in a preset area, and save the image with detailed information; Step S300: Use the segmentation model to accurately segment the target area, remove irrelevant background, and only retain the effective area of ​​the target; Step S400: scaling the segmented target image in equal proportions, keeping its aspect ratio unchanged, and superimposing it on a pure white background of the same size in the center, and finally adjusting all images to a fixed size; Step S500: an image classification model is obtained based on training of a class center two-dimensional loss function, the geometric features and pixel features of the image obtained in step S400 are fused, the fused features are input into the image classification model, and finally a classification result is output.

2. The fine-grained image classification method according to claim 1, characterized in that: The step S300 specifically includes: Step S310: converting the segmentation model into a format suitable for the mobile terminal, and loading the model through OpenCV; Step S320: According to the pre-positioning information, select the positions of the 1 / 4, 2 / 4 and 3 / 4 points on the central axis, set these points as mask points, and set the label values ​​of the mask points to 1; Step S330: Use the segmentation model to segment the input image to generate a segmented image containing only the target area.

3. The fine-grained image classification method according to claim 1, characterized in that: The step S400 specifically includes: Step S410: performing morphological operations on the segmented image, including dilation, erosion and filtering, to optimize the edge quality and detail performance of the target image; Step S420: scaling the image in proportion according to the aspect ratio of the segmented target image; Step S430: Create a white background and place the scaled target image at the center of the background to generate a standardized target image.

4. The fine-grained image classification method according to claim 1, characterized in that: The step S500 specifically includes: Step S510: obtaining an image classification model based on training of a class center two-dimensional loss function; Step S520: using the Canny algorithm to extract edge information of the target; Step S530: Calculate the perimeter of the target edge according to the target edge; Step S540: Calculate the major axis length of the target using the minimum inscribed circle method; Step S550: Calculate the geometric features of the target through the ratio of the perimeter of the target edge to the major axis; Step S560: Load the pre-trained image classification model through OpenCV; Step S570: Input the target image into the trained image classification model, extract the pixel features of the previous layer of the fully connected layer, and expand the geometric features into a two-dimensional vector of the same shape, and then perform weighted addition with the pixel features to obtain a fusion feature; Step S580: Input the fused features into the fully connected layer for reasoning and output the classification results.

5. The fine-grained image classification method according to claim 4, characterized in that: The step S510 specifically includes: Take multiple target images of various categories and build a training data set; Use the segmentation model to segment the target area, remove background interference, and extract the effective area of ​​the target; Standardize the segmented target image to provide a consistent data format; Extract the geometric features and image features of the target, and fuse them to generate a feature vector containing multi-dimensional information; The fused target features are input into the neural network model to be trained, and the image classification model is finally obtained based on the iterative optimization of the class center two-dimensional loss function; wherein: The class center two-dimensional loss function consists of two parts, L feature With L class , where L feature Aims to minimize intra-class variance and improve inter-class separability; L feature The calculation is as follows: Among them, sim i represents the class index that is most similar to the target class i, Represents X i and The similarity between them, Peason represents the absolute value of the Pearson correlation coefficient, X i represents the current class center feature of class i, N represents the number of samples in a batch, and assumes that the current sample feature vector is f; L class JS divergence is used to measure the similarity between the predicted probability distribution of the sample and the soft label distribution; L class It is expressed as: Where M represents the sample feature P Xi With soft label distribution Q yi The average value of N represents the number of samples in a batch; KL(P Xi ||M) and KL(Q yi ||M) represents the probability distribution P Xi and M, and Q yi and the KL divergence between M, JS(P Xi ||Q yi ) represents the probability distribution P Xi With Q yi JS divergence between; The final class center two-dimensional loss function is expressed as: <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> cc <h2 style=";text-align:left;direction:ltr"> =a1L<h2 style=";text-align:left;direction:ltr"> feature <h2 style=";text-align:left;direction:ltr"> +a2L<h2 style=";text-align:left;direction:ltr"> class <h2 style=";text-align:left;direction:ltr"> +a3L<h2 style=";text-align:left;direction:ltr"> CE Where L cc represents the class center two-dimensional loss function, L CE represents the cross entropy loss function, and a represents the weight value.

6. A fine-grained image classification device, characterized in that: The device comprises: Image acquisition module, used to obtain the target image to be identified, ensure that the target is located in the preset area, and provide accurate input data for subsequent processing; Image segmentation module, used to segment the target image, remove irrelevant background information, and retain only the effective area of ​​the target; The image standardization module is used to perform morphological processing on the target image, remove noise and edge burrs, and adjust the target image to a uniform standard size; The image classification module is used to input the target fusion features into the pre-trained image classification model and output the classification results.

7. The fine-grained image classification device according to claim 6, characterized in that: The image classification module comprises: A geometric feature extraction unit is used to extract the edge information of the target, calculate the perimeter and the major axis length of the target edge, and generate the geometric features of the target through the ratio of the two; An image feature extraction unit is used to extract deep image features from the target image and capture the texture and morphological information of the target; The feature fusion unit fuses the geometric features of the target with the image features to generate a feature vector that comprehensively expresses the characteristics of the target; The model inference unit inputs the fused features into the pre-trained recognition model and infers the classification results.

8. An edge computing device, comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor implements the steps of the method described in claims 1 to 5 when executing the computer program.

9. An edge computing device readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in claims 1 to 5 are implemented.